Indus Valley Script · structural analysis 27 July 2026
Undeciphered Scripts · c. 2600–1900 BCE

The Indus Script Is Not an Emblem System Sequential dependency in 2,536 inscriptions, measured against matched linguistic and non-linguistic controls

2536 inscriptions11,135 sign tokens 591 signs27 July 2026
An Indus Valley stamp seal in stone beside its modern clay impression. A row of script signs runs along the top edge above a one-horned bull.
A stamp seal and its modern impression. The row of signs along the top is the script; the animal below it is the motif. Our measurements show the two channels carry almost no information about each other. Public domain, released under CC0.

The Indus Valley script has resisted decipherment for a century: the inscriptions are short, no bilingual key exists, and the underlying language is unknown. This analysis does not decipher it. It measures the structure of 1902 real inscriptions against controls whose nature is known in advance: natural language, modelled administrative records, modelled ownership marks, and random signs matched on every property that drives the statistics.

The Indus corpus carries more dependency between successive signs than English prose, while a modelled emblem system carries none at all. Inscriptions are closed units: information does not run across the boundary between two objects. Both results survive deduplication of repeated seal impressions and replicate independently at Mohenjo-daro and Harappa. Directionality and positional fixity are strong but are shown not to discriminate writing from record-keeping, and one earlier finding, induced grammatical categories, is withdrawn here, having failed its own control test.

IThe Corpus

Every figure on this page comes from excavated inscriptions in the published digitisation of the Corpus of Indus Seals and Inscriptions. No reconstructed, simulated or generated text enters the analysis at any point. This is worth stating plainly because the alternative is common in this field and hard to detect from outside.

Inscriptions2536 1902 after collapsing duplicate seal impressions
Sign tokens11,135 mean 4.39 signs per inscription, maximum 17
Distinct signs591 34% occur exactly once
Sites10 Mohenjo-daro, Harappa, Lothal, Dholavira, Kalibangan and others
Table 1. Corpus composition.

IIThe Finding

At this corpus size every measure of structure is inflated. With 591 distinct signs and fewer than nine thousand tokens, the table of sign pairs is about two per cent occupied, and in that regime a corpus shuffled into complete meaninglessness still returns a large positive score. A raw figure therefore means nothing. What follows is in every case the margin over a matched control: the same corpus, its structure destroyed, measured the same way.

Excess sequential dependency
1.033
bits per adjacent sign pair above a corpus of the same signs re-dealt at random. English prose, matched on every relevant property, returns +0.579.
MeasureObservedControl Excessz
Sequential dependency
order carries information
3.1692.1361.03371.5
Order alone
controlling for co-occurrence
3.1692.2690.90064.9
Boundary closure
negative = inscriptions are closed
2.2263.089-0.863-18.9
Directionality
opening vs closing sign sets
0.8310.2060.62561.7
Positional fixity
signs hold their region
0.6600.4780.18235.1
Table 2. Structural measures on the deduplicated corpus, each against its matched control. Boundary closure is negative by construction of the test: sign pairs spanning two different objects share less information than chance.

The second row is the one that rules out a weaker reading. Scrambling the order of signs within an inscription, while keeping exactly which signs co-occur, still destroys 0.900 bits. The information is in the sequence, not merely in the company.

The boundary result is the cleanest in the study. Whatever an inscription says, it finishes saying it: dependency across the join between two artifacts falls below chance. Each object carries one complete statement, and the corpus is not the scattered fragments of a longer running text.

IIIThe Controls

A structural result is uninterpretable until you know what other systems score. The identical analysis was therefore run on four control corpora, each matched to the Indus corpus on text count, length distribution, token count and vocabulary size. Without that matching one measures corpus size rather than writing system, which is the error that makes cross-corpus comparisons in this field unreliable.

CorpusSequentialBoundary PositionalDirectional
Indus script
the real corpus
+1.033-0.863+0.181+0.624
Natural language
real text, matched
+0.579+0.304+0.000-0.008
Administrative records
modelled
+0.537-1.965+0.354+0.758
Ownership marks
modelled
-0.468-1.433+0.221+0.766
Random signs
structure floor
-0.007+0.083-0.007-0.006
Table 3. Matched controls. Read the last row first: randomly ordered signs score zero on everything, which is the check that the method is calibrated and not manufacturing structure.
What separates Indus

Sequential dependency, and only sequential dependency. Indus scores above every control including real natural language, while the ownership-mark model scores negative: emblems carry no information in their ordering. This is a genuine constraint on any account that treats the script as a set of marks.

What does not

Directionality and positional fixity. Both are strong in Indus and stronger still in the administrative and emblem models. They establish that inscriptions are ordered and read in a fixed direction; they cannot tell writing from record-keeping. Any argument resting on them alone is unsafe. That includes ours, until this test was run.

Where that leaves the question

Indus sits cleanly with neither family. It is high on sequential dependency like a language, discrete at its boundaries like an administrative record, and positionally rigid like both. On the profile as a whole it is intermediate, which is independently the same conclusion reached by a separate 2026 analysis of this corpus by different means.

The pure emblem hypothesis is the one that struggles. It predicts no sequential dependency, and the Indus corpus has more of it than English prose.

A limitation we have not solved

The natural-language control is running prose cut into short segments, so it has no genuine text beginnings or endings. That makes it a poor comparator for directionality in particular, and it is why no claim is made that Indus directionality is language-like, only that it is real. A control built from genuinely short complete texts would settle it. We have not built one.

IVWhat Survived

The findings were attacked before they were published. The strongest objection was that duplicate seal impressions, the same seal pressed many times or mass-produced copies of one design, could manufacture all of these results from nothing. Collapsing every repeated inscription removes 634 texts. Each measure was then recomputed on the reduced corpus, and separately at each major site, so that no finding rests on pooling two traditions that might differ.

Excess over controlAll 2,536Dedup.Mohenjo-daroHarappa
Sequential dependency1.3961.0330.9320.746
Order alone1.0730.9000.8370.692
Boundary closure-0.848-0.863-1.210-1.140
Directionality0.6870.6250.6050.554
Positional fixity0.2130.1820.1990.219
Grammatical categories (withdrawn, z)2.611.471.562.35
Table 4. Every headline measure across corpus variants. The struck row is the withdrawn result described below.

The effects persist after deduplication and hold independently at Mohenjo-daro and Harappa, two sites some six hundred kilometres apart, showing the same structure.

How many independent inscriptions are there really?

Collapsing only identical inscriptions is the weak version of this control. Indus seals were also copied, and near-identical texts are not independent observations either. Grouping inscriptions that differ by one or two sign edits reduces 1,902 distinct sequences to 1,141 and then to 558. Intervals computed on the raw count are therefore too narrow, which is a problem for most published statistics on this corpus, and was for ours.

Measure (excess, with z)1,902 exact 1,141 at edit 1558 at edit 2
Sequential dependency+1.033  71.5 +0.875  54.5 +0.665  34.0
Directionality+0.625  61.7 +0.561  45.5 +0.464  26.7
Positional fixity+0.182  35.1 +0.151  24.8 +0.125  18.0
Opening vs closing classes z 6.0z 4.6z 0.5
Table 5. Every headline measure recomputed on progressively stricter definitions of an independent inscription. The struck row is the one that does not survive.

The three main results hold on 558 inscriptions, under a quarter of what we started with. That is the number they should be judged on.

Is the headline number an artifact of the estimator?

This is the sharpest objection available and it has been made before, against earlier entropy work on this corpus. At 11,135 tokens across 591 sign types the table of sign pairs is 0.84 per cent occupied, 2,919 filled cells out of 349,281. In that regime the naive maximum-likelihood estimator overstates dependency badly, and a figure quoted without correction cannot be trusted.

Our figure is already a margin over a permutation control computed with the same estimator, which cancels most of that bias because the control shares the sample size, alphabet and marginal frequencies. It does not cancel it exactly. So the same quantity was recomputed under three published bias-corrected estimators, each applied to the observed data and to the control alike.

EstimatorObservedControl Excess95% interval
Maximum likelihood3.1692.136 +1.033+0.873 to +0.978
Miller-Madow2.9711.778 +1.192+1.014 to +1.124
Chao-Shen3.0011.566 +1.435+1.370 to +1.511
Chao-Wang-Jost2.5380.861 +1.677+1.495 to +1.742
Table 6. The same excess under four estimators. Intervals come from subsampling inscriptions without replacement, which is the correct resampling unit here: drawing them with replacement reintroduces the duplicate structure the analysis exists to control for.

Correction moves the result upward, not downward. The excess rises from 1.033 to 1.677 bits as the correction gets stronger, because the correction reduces the control far more than it reduces the observed data. The shuffled corpus is dominated by sign pairs seen once, which is exactly the situation these estimators exist to fix. The uncorrected figure we publish is therefore the conservative one, and every estimator tested puts the interval clear of zero.

A result we withdrew

An earlier pass found that signs sort into grammatical categories which then line up in order along the inscription. It was the most interesting result of the project and the closest thing to a grammar. It did not survive. Once duplicate impressions were collapsed the effect fell to chance, and it failed to reproduce consistently across sites. The duplicates had been manufacturing it.

It is recorded here because a finding that dies under its own control test is worth more to other researchers than one that was never tested, and because anyone running similar analyses on this corpus should expect the same trap.

What was left when we looked again

Withdrawing a result is not the same as showing there is nothing there. The original test asked one broad question (do a dozen induced categories arrange themselves in order?) and spent all its evidence on it. Three narrower tests were run afterwards on the deduplicated corpus, each able to fail independently.

Second-pass testResultAgainst
Cluster stability under resampling0.2660.049
Category model vs full sign model (bits, held-out)+0.0284t = 1.89
Opening vs closing sign contextsz = 6.02p < 0.001
Table 7. Second-pass tests, deduplicated corpus only.

The categories are more stable under resampling than chance, but not by much. Compressing 591 signs into a dozen categories costs nothing measurable on unseen text. The interval crosses zero, so the honest reading is no detectable loss, not an improvement.

The third test held up at first. The signs that open inscriptions and those that close them differ in the company they keep, and not merely in where they sit: the comparison uses each sign's neighbours with position information removed. Twenty-one openers and forty-eight closers form two distributionally distinct groups at z = 6.02.

It does not survive the harder duplicate control. Collapsing near-identical inscriptions as well as exact ones takes it to z = 4.6 at edit distance one, and to z = 0.5 at edit distance two, which is nothing at all. We are leaving it here as a reported result rather than a claim: it may be real and merely under-powered once the corpus is cut to genuinely independent inscriptions, or it may be another artifact of repeated seals. On this corpus we cannot tell.

VWhat This Settles

Supported
  • Sign order carries information, and more of it than in running prose
  • Inscriptions are complete, self-contained statements
  • The script is directional, with distinct opening and closing positions
  • Signs hold fixed positions rather than floating freely
  • The same system is in use at sites hundreds of kilometres apart
Not established
  • What any inscription says
  • What language, if any, underlies the script
  • The phonetic or semantic value of any sign
  • Whether the script is logographic, syllabic or mixed
  • Any connection to a known language family
  • Whether signs group into a full system of grammatical categories
    though opening and closing signs do form two distinct classes
This is not a decipherment

No one has deciphered the Indus script, and this work does not either. Claims to have read it appear regularly and are not supported by evidence of this kind; neither is anything here. Not one inscription is translated. One sign opens roughly a third of the corpus and we do not know what it means.

The strongest honest statement the data supports is this: the Indus script carries ordered, self-contained, positionally organised structure, and more dependency between successive signs than English prose. That is a real constraint, and the emblem and ownership-mark readings have to answer it. It is not the same as showing the script encodes a language, and we are not showing that.

VIThe Picture and the Object

Every Indus seal carries an image, most often a one-horned bull, sometimes an elephant, a rhinoceros, a tree or a human figure, set above the row of signs. The natural reading is that the signs caption the picture. If so, the two channels must share information, and the image would become a weak bilingual: a known concept standing beside an unknown word.

ChannelObservedControl Excessz
Iconographic motif0.29780.2814+0.01642.0
Object type0.20850.1437+0.064814.0
Material0.18890.1525+0.03647.1
Table 8. Information shared between the sign string and three non-textual channels, deduplicated corpus. Labels were permuted across objects to build each control.

The caption reading is not supported. Before deduplication the motif appears to carry substantial information about the signs; after repeated seal impressions are collapsed, almost all of it disappears. What remains does not clear the bar once the three channels tested here are accounted for. The picture and the text are, as far as this measurement goes, independent channels. Whatever the signs record, it is not a description of the animal above them.

The object itself is a different story. What a piece is, whether seal, tablet, tag or potsherd, genuinely conditions which signs appear on it, and survives deduplication comfortably. Text length tracks it too: seals average 5.06 signs, tablets 3.70, pottery 2.78. The corpus is not one uniform practice, and analyses that pool it are averaging over at least two.

Whether it could be a list of names

A standing objection to the whole enterprise is that these are ownership marks, personal names, lineages and titles, rather than text. That hypothesis makes a sharp prediction. A corpus of one-off identifiers never settles down: every new object introduces new signs, and the vocabulary grows about as fast as the corpus itself. Reused language does the opposite, because a finite lexicon is being drawn on repeatedly.

CorpusVocabulary growthReading
Indus script0.488finite vocabulary, reused
English prose (matched)0.415finite vocabulary, reused
Indus, seals only0.506
Indus, tablets only0.581
Table 9. Heaps' law exponent. Near 0.5, a finite vocabulary is being reused; approaching 1.0, it is not.

Indus sits at 0.488, against 0.415 for matched English prose. Vocabulary growth is slightly faster, consistent with more one-off signs, but nowhere near the behaviour of a list of unique identifiers. The sign inventory is genuinely reused. Whatever the signs are, they are drawn from a working vocabulary rather than minted per object, and the pure name-list reading has to account for that.

VIIHow Close Is This To A Decipherment

Reading an unknown script is not one problem but a sequence of them, and the later ones are far harder than the earlier ones. Setting them out in order makes it possible to say where this work actually sits, and where the field sits, without either being flattered.

Machine-readable corpus2,536 inscriptions in hand. The fuller ICIT corpus is held behind a permission request.
field: done
Corpus is not randomLong established in the literature, and reconfirmed here against shuffled controls.
field: done
Ordered, directional, closed unitsSign order carries information, and inscriptions do not run on across objects.
field: done
Not a pure emblem systemModelled ownership marks score negative on sequential dependency where Indus scores above English prose. The non-linguistic case still has serious advocates.
field: contested
Vocabulary is reused, not minted per objectHeaps exponent 0.488 against 0.415 for matched English. Argues against a pure list of names.
field: open
~
Sign classes and grammarTwo classes, openers against closers, survive at z = 6.02. The twelve-category version was withdrawn. Nobody has demonstrated a grammar.
field: contested
Word or morpheme segmentationNo agreed way to divide Indus strings into units exists.
field: contested
Language family identifiedDravidian, Indo-Aryan and isolate all have advocates. None is established.
field: contested
Phonetic values for signsA full Sanskrit mapping has been claimed and is not independently confirmed. Auditing it is open work.
field: claimed
Read an arbitrary unseen inscriptionThis is the actual bar for decipherment. Nobody is here.
field: not achieved
Survives independent verificationWould require everything above to hold up under outside scrutiny.
field: not achieved

5 of 11 rungs cleared here, one partial, which is roughly 50% of the way up. That figure should be read with care. The rungs are not equal in size, and the remaining ones are much larger than the ones already climbed.

The honest position

Nobody has verifiably cleared rung nine or beyond, ourselves included. Rung nine has been claimed more than once and never independently confirmed. The last three rungs are what the word decipherment actually refers to, and no one is standing on them.

What this work adds is confined to rungs four and five, and to narrowing rung six. That is a genuine contribution to a hundred-year-old problem and it is also a long way from reading the script. Both halves of that sentence matter.

VIIIOn Method

The analysis pipeline is not published. What is published is every result it produced, the control each result was measured against, the significance of each margin, the corpus variants each was recomputed on, and the finding that failed. That is what a reader needs in order to hold these claims to account.

Corpus provenance is public and named above, so the inputs are independently obtainable by anyone wishing to measure them by their own means. That is the test we would rather face: not agreement about procedure, but agreement about numbers arrived at separately. Researchers wanting to compare results are welcome to make contact.