The Indus Script Is Not an Emblem System Sequential dependency in 2,536 inscriptions, measured against matched linguistic and non-linguistic controls
The Indus Valley script has resisted decipherment for a century: the inscriptions are short, no bilingual key exists, and the underlying language is unknown. This analysis does not decipher it. It measures the structure of 1902 real inscriptions against controls whose nature is known in advance: natural language, modelled administrative records, modelled ownership marks, and random signs matched on every property that drives the statistics.
The Indus corpus carries more dependency between successive signs than English prose, while a modelled emblem system carries none at all. Inscriptions are closed units: information does not run across the boundary between two objects. Both results survive deduplication of repeated seal impressions and replicate independently at Mohenjo-daro and Harappa. Directionality and positional fixity are strong but are shown not to discriminate writing from record-keeping, and one earlier finding, induced grammatical categories, is withdrawn here, having failed its own control test.
IThe Corpus
Every figure on this page comes from excavated inscriptions in the published digitisation of the Corpus of Indus Seals and Inscriptions. No reconstructed, simulated or generated text enters the analysis at any point. This is worth stating plainly because the alternative is common in this field and hard to detect from outside.
| Inscriptions | 2536 | 1902 after collapsing duplicate seal impressions |
|---|---|---|
| Sign tokens | 11,135 | mean 4.39 signs per inscription, maximum 17 |
| Distinct signs | 591 | 34% occur exactly once |
| Sites | 10 | Mohenjo-daro, Harappa, Lothal, Dholavira, Kalibangan and others |
IIThe Finding
At this corpus size every measure of structure is inflated. With 591 distinct signs and fewer than nine thousand tokens, the table of sign pairs is about two per cent occupied, and in that regime a corpus shuffled into complete meaninglessness still returns a large positive score. A raw figure therefore means nothing. What follows is in every case the margin over a matched control: the same corpus, its structure destroyed, measured the same way.
| Measure | Observed | Control | Excess | z |
|---|---|---|---|---|
| Sequential dependency order carries information | 3.169 | 2.136 | 1.033 | 71.5 |
| Order alone controlling for co-occurrence | 3.169 | 2.269 | 0.900 | 64.9 |
| Boundary closure negative = inscriptions are closed | 2.226 | 3.089 | -0.863 | -18.9 |
| Directionality opening vs closing sign sets | 0.831 | 0.206 | 0.625 | 61.7 |
| Positional fixity signs hold their region | 0.660 | 0.478 | 0.182 | 35.1 |
The second row is the one that rules out a weaker reading. Scrambling the order of signs within an inscription, while keeping exactly which signs co-occur, still destroys 0.900 bits. The information is in the sequence, not merely in the company.
The boundary result is the cleanest in the study. Whatever an inscription says, it finishes saying it: dependency across the join between two artifacts falls below chance. Each object carries one complete statement, and the corpus is not the scattered fragments of a longer running text.
IIIThe Controls
A structural result is uninterpretable until you know what other systems score. The identical analysis was therefore run on four control corpora, each matched to the Indus corpus on text count, length distribution, token count and vocabulary size. Without that matching one measures corpus size rather than writing system, which is the error that makes cross-corpus comparisons in this field unreliable.
| Corpus | Sequential | Boundary | Positional | Directional |
|---|---|---|---|---|
| Indus script the real corpus | +1.033 | -0.863 | +0.181 | +0.624 |
| Natural language real text, matched | +0.579 | +0.304 | +0.000 | -0.008 |
| Administrative records modelled | +0.537 | -1.965 | +0.354 | +0.758 |
| Ownership marks modelled | -0.468 | -1.433 | +0.221 | +0.766 |
| Random signs structure floor | -0.007 | +0.083 | -0.007 | -0.006 |
Sequential dependency, and only sequential dependency. Indus scores above every control including real natural language, while the ownership-mark model scores negative: emblems carry no information in their ordering. This is a genuine constraint on any account that treats the script as a set of marks.
Directionality and positional fixity. Both are strong in Indus and stronger still in the administrative and emblem models. They establish that inscriptions are ordered and read in a fixed direction; they cannot tell writing from record-keeping. Any argument resting on them alone is unsafe. That includes ours, until this test was run.
Indus sits cleanly with neither family. It is high on sequential dependency like a language, discrete at its boundaries like an administrative record, and positionally rigid like both. On the profile as a whole it is intermediate, which is independently the same conclusion reached by a separate 2026 analysis of this corpus by different means.
The pure emblem hypothesis is the one that struggles. It predicts no sequential dependency, and the Indus corpus has more of it than English prose.
A limitation we have not solved
The natural-language control is running prose cut into short segments, so it has no genuine text beginnings or endings. That makes it a poor comparator for directionality in particular, and it is why no claim is made that Indus directionality is language-like, only that it is real. A control built from genuinely short complete texts would settle it. We have not built one.
IVWhat Survived
The findings were attacked before they were published. The strongest objection was that duplicate seal impressions, the same seal pressed many times or mass-produced copies of one design, could manufacture all of these results from nothing. Collapsing every repeated inscription removes 634 texts. Each measure was then recomputed on the reduced corpus, and separately at each major site, so that no finding rests on pooling two traditions that might differ.
| Excess over control | All 2,536 | Dedup. | Mohenjo-daro | Harappa |
|---|---|---|---|---|
| Sequential dependency | 1.396 | 1.033 | 0.932 | 0.746 |
| Order alone | 1.073 | 0.900 | 0.837 | 0.692 |
| Boundary closure | -0.848 | -0.863 | -1.210 | -1.140 |
| Directionality | 0.687 | 0.625 | 0.605 | 0.554 |
| Positional fixity | 0.213 | 0.182 | 0.199 | 0.219 |
| Grammatical categories (withdrawn, z) | 2.61 | 1.47 | 1.56 | 2.35 |
The effects persist after deduplication and hold independently at Mohenjo-daro and Harappa, two sites some six hundred kilometres apart, showing the same structure.
An earlier pass found that signs sort into grammatical categories which then line up in order along the inscription. It was the most interesting result of the project and the closest thing to a grammar. It did not survive. Once duplicate impressions were collapsed the effect fell to chance, and it failed to reproduce consistently across sites. The duplicates had been manufacturing it.
It is recorded here because a finding that dies under its own control test is worth more to other researchers than one that was never tested, and because anyone running similar analyses on this corpus should expect the same trap.
What was left when we looked again
Withdrawing a result is not the same as showing there is nothing there. The original test asked one broad question (do a dozen induced categories arrange themselves in order?) and spent all its evidence on it. Three narrower tests were run afterwards on the deduplicated corpus, each able to fail independently.
| Second-pass test | Result | Against |
|---|---|---|
| Cluster stability under resampling | 0.266 | 0.049 |
| Category model vs full sign model (bits, held-out) | +0.0284 | t = 1.89 |
| Opening vs closing sign contexts | z = 6.02 | p < 0.001 |
The categories are more stable under resampling than chance, but not by much. Compressing 591 signs into a dozen categories costs nothing measurable on unseen text. The interval crosses zero, so the honest reading is no detectable loss, not an improvement.
The third test is the one that holds. The signs that open inscriptions and those that close them differ in the company they keep, and not merely in where they sit: the comparison uses each sign's neighbours with position information removed. Twenty-one openers and forty-eight closers form two distributionally distinct groups at z = 6.02. That is a real class distinction in the script: two categories, not twelve, and far short of a grammar, but it survives the control that killed the larger claim.
VWhat This Settles
- Sign order carries information, and more of it than in running prose
- Inscriptions are complete, self-contained statements
- The script is directional, with distinct opening and closing positions
- Signs hold fixed positions rather than floating freely
- The same system is in use at sites hundreds of kilometres apart
- What any inscription says
- What language, if any, underlies the script
- The phonetic or semantic value of any sign
- Whether the script is logographic, syllabic or mixed
- Any connection to a known language family
- Whether signs group into a full system of grammatical categories
though opening and closing signs do form two distinct classes
No one has deciphered the Indus script, and this work does not either. Claims to have read it appear regularly and are not supported by evidence of this kind; neither is anything here. Not one inscription is translated. One sign opens roughly a third of the corpus and we do not know what it means.
The strongest honest statement the data supports is this: the Indus script carries ordered, self-contained, positionally organised structure, and more dependency between successive signs than English prose. That is a real constraint, and the emblem and ownership-mark readings have to answer it. It is not the same as showing the script encodes a language, and we are not showing that.
VIThe Picture and the Object
Every Indus seal carries an image, most often a one-horned bull, sometimes an elephant, a rhinoceros, a tree or a human figure, set above the row of signs. The natural reading is that the signs caption the picture. If so, the two channels must share information, and the image would become a weak bilingual: a known concept standing beside an unknown word.
| Channel | Observed | Control | Excess | z |
|---|---|---|---|---|
| Iconographic motif | 0.2978 | 0.2814 | +0.0164 | 2.0 |
| Object type | 0.2085 | 0.1437 | +0.0648 | 14.0 |
| Material | 0.1889 | 0.1525 | +0.0364 | 7.1 |
The caption reading is not supported. Before deduplication the motif appears to carry substantial information about the signs; after repeated seal impressions are collapsed, almost all of it disappears. What remains does not clear the bar once the three channels tested here are accounted for. The picture and the text are, as far as this measurement goes, independent channels. Whatever the signs record, it is not a description of the animal above them.
The object itself is a different story. What a piece is, whether seal, tablet, tag or potsherd, genuinely conditions which signs appear on it, and survives deduplication comfortably. Text length tracks it too: seals average 5.06 signs, tablets 3.70, pottery 2.78. The corpus is not one uniform practice, and analyses that pool it are averaging over at least two.
Whether it could be a list of names
A standing objection to the whole enterprise is that these are ownership marks, personal names, lineages and titles, rather than text. That hypothesis makes a sharp prediction. A corpus of one-off identifiers never settles down: every new object introduces new signs, and the vocabulary grows about as fast as the corpus itself. Reused language does the opposite, because a finite lexicon is being drawn on repeatedly.
| Corpus | Vocabulary growth | Reading |
|---|---|---|
| Indus script | 0.488 | finite vocabulary, reused |
| English prose (matched) | 0.415 | finite vocabulary, reused |
| Indus, seals only | 0.506 | |
| Indus, tablets only | 0.581 |
Indus sits at 0.488, against 0.415 for matched English prose. Vocabulary growth is slightly faster, consistent with more one-off signs, but nowhere near the behaviour of a list of unique identifiers. The sign inventory is genuinely reused. Whatever the signs are, they are drawn from a working vocabulary rather than minted per object, and the pure name-list reading has to account for that.
VIIHow Close Is This To A Decipherment
Reading an unknown script is not one problem but a sequence of them, and the later ones are far harder than the earlier ones. Setting them out in order makes it possible to say where this work actually sits, and where the field sits, without either being flattered.
5 of 11 rungs cleared here, one partial, which is roughly 50% of the way up. That figure should be read with care. The rungs are not equal in size, and the remaining ones are much larger than the ones already climbed.
Nobody has verifiably cleared rung nine or beyond, ourselves included. Rung nine has been claimed more than once and never independently confirmed. The last three rungs are what the word decipherment actually refers to, and no one is standing on them.
What this work adds is confined to rungs four and five, and to narrowing rung six. That is a genuine contribution to a hundred-year-old problem and it is also a long way from reading the script. Both halves of that sentence matter.
VIIIOn Method
The analysis pipeline is not published. What is published is every result it produced, the control each result was measured against, the significance of each margin, the corpus variants each was recomputed on, and the finding that failed. That is what a reader needs in order to hold these claims to account.
Corpus provenance is public and named above, so the inputs are independently obtainable by anyone wishing to measure them by their own means. That is the test we would rather face: not agreement about procedure, but agreement about numbers arrived at separately. Researchers wanting to compare results are welcome to make contact.