Short Texts, Not Absent Structure Linear B, cut to Indus dimensions, reproduces the statistical profile that has been used to argue the Indus script is not writing
The Indus script has resisted decipherment for a century. A recurring argument holds that it never encoded language at all: its texts are too short, its dependencies too local, its structure too thin. That argument has never been tested against the obvious control, which is a script we can already read, cut to the same dimensions.
We did that. Linear B, deciphered in 1952 and known to encode Greek, was subsampled to the Indus token count and length distribution and run through the identical procedure, alongside proto-cuneiform, which is attested notation that was not yet writing. On the information carried purely by sign order, Indus scores 0.901 bits and Linear B 0.913. Proto-cuneiform manages 0.405. Both Indus and Linear B lose almost all dependency beyond adjacent signs, which is the observation most often cited as evidence against Indus being writing.
The absence of long-range structure in the Indus corpus is therefore a consequence of short texts, not a property of the script. This does not show that the Indus script is writing, and nothing here reads a single sign. It shows that the strongest statistical argument that it is not does not survive its own control.
IThe Control That Was Missing
Every published structural analysis of the Indus corpus, including our own until now, has compared it against systems chosen to be non-linguistic: heraldry, ownership marks, tallies, invented baselines. None has compared it against a known writing system measured at the same size. Without that, a result showing Indus to be statistically impoverished says nothing, because nobody had established what real writing looks like when you only ever get four signs at a time.
Three corpora, all attested, all openly published. Every one cut to the Indus token count and the Indus length distribution, sampled without replacement and averaged over repeated draws.
| Measure | Indus | Linear B | Proto-cuneiform |
|---|---|---|---|
| texts / tokens | 1,902 / 8,839 | 1,893 / 8,841 | 1,907 / 8,710 |
| mean length | 4.65 | 4.67 | 4.57 |
| information in sign order alone | 0.901 | 0.913 | 0.405 |
| dependency at distance 1 | +0.901 | +0.912 | +0.404 |
| dependency at distance 2 | +0.077 | +0.095 | −0.118 |
| dependency at distance 3 | −0.106 | −0.054 | −0.421 |
| recurring three-sign blocks | 0.235 | 0.330 | 0.046 |
| vocabulary growth | 0.488 | 0.294 | 0.431 |
| signs occurring once | 0.364 | 0.193 | 0.210 |
The two highlighted rows carry the argument. On information carried by sign order alone, Indus and Linear B are within 1.3 per cent of each other. And both collapse to almost nothing at distance two, which is precisely the observation that has been offered as evidence that Indus cannot be language. A script we have read for seventy years does the same thing at this text length.
Proto-cuneiform, which really is a notation rather than a script, sits at less than half the Indus figure and turns negative at distance two. Whatever the Indus script is, it is not behaving like that.
It does not show that the Indus script encodes language. It shows that one specific argument that it does not, the argument from missing long-range structure, fails its own control. That is a narrower claim and it is the one the data supports.
One limitation matters and is not hidden here. The arms are matched on text count, length and token count, but they cannot be matched on inventory size: Indus has 591 sign types where Linear B offers 211, because Linear B does not contain 591 signs. A larger inventory means sparser data and more estimator bias, which plausibly explains part of the remaining gap in the total figures. It does not explain the ordering result or the decay pattern, which are the two rows the argument rests on.
IIThe Corpus
Every figure on this page comes from excavated inscriptions in the published digitisation of the Corpus of Indus Seals and Inscriptions. No reconstructed, simulated or generated text enters the analysis at any point. This is worth stating plainly because the alternative is common in this field and hard to detect from outside.
| Inscriptions | 2536 | 1902 after collapsing duplicate seal impressions |
|---|---|---|
| Sign tokens | 11,135 | mean 4.39 signs per inscription, maximum 17 |
| Distinct signs | 591 | 34% occur exactly once |
| Sites | 10 | Mohenjo-daro, Harappa, Lothal, Dholavira, Kalibangan and others |
IIIThe Finding
At this corpus size every measure of structure is inflated. With 591 distinct signs and fewer than nine thousand tokens, the table of sign pairs is about two per cent occupied, and in that regime a corpus shuffled into complete meaninglessness still returns a large positive score. A raw figure therefore means nothing. What follows is in every case the margin over a matched control: the same corpus, its structure destroyed, measured the same way.
| Measure | Observed | Control | Excess | z |
|---|---|---|---|---|
| Sequential dependency order carries information | 3.169 | 2.136 | 1.033 | 71.5 |
| Order alone controlling for co-occurrence | 3.169 | 2.269 | 0.900 | 64.9 |
| Boundary closure negative = inscriptions are closed | 2.226 | 3.089 | -0.863 | -18.9 |
| Directionality opening vs closing sign sets | 0.831 | 0.206 | 0.625 | 61.7 |
| Positional fixity signs hold their region | 0.660 | 0.478 | 0.182 | 35.1 |
The second row is the one that rules out a weaker reading. Scrambling the order of signs within an inscription, while keeping exactly which signs co-occur, still destroys 0.900 bits. The information is in the sequence, not merely in the company.
The boundary result is the cleanest in the study. Whatever an inscription says, it finishes saying it: dependency across the join between two artifacts falls below chance. Each object carries one complete statement, and the corpus is not the scattered fragments of a longer running text.
IVInscription Families
A separate question, and one that does not depend on any of the entropy machinery: do inscriptions repeat each other? A workshop, a single scribe, or a fixed formula should leave relatives behind, sharing long runs of signs more often than chance permits.
| Shared run | Pairs observed | Expected by chance | Effect size |
|---|---|---|---|
| three or more signs | 4,064 | 147 | h = 0.08 |
| four or more signs | 467 | 1 | h = 0.03 |
Four hundred and sixty-seven pairs of inscriptions share a run of four or more identical signs, where the control predicts one. That is a real relationship and not a marginal one.
The effect size says the other half of it. Those 467 pairs are three hundredths of one per cent of all possible pairs, so on the rate the effect is negligible. The corpus contains a small number of genuine families, not a corpus-wide system of formulae. Both halves of that sentence come from the same measurement, and reporting only the first would misrepresent it.
The longest shared runs reach seven signs, on pairs of inscriptions that are otherwise unrelated objects. Those are the best candidates in the corpus for two artifacts carrying the same text, and they are a natural target for anyone with access to the originals.
The families are local
If these families are the work of individual hands, they should betray it physically. A single workshop stands in one place, draws on one stock of material and makes one kind of object, so two inscriptions from the same hand ought to agree about where they were found far more often than two inscriptions picked at random. If instead the shared runs are a standard formula circulating across the whole civilisation, they should be scattered.
They are not scattered. Pairs sharing a run of four signs come from the same site 59.1% of the time against 36.5% expected by chance, and that holds after near-identical inscriptions are collapsed at edit distance two. Agreement on object type and material runs the same way but weakens once pairs from different sites are compared directly, which says the object and the material are largely following the geography rather than adding to it. Whatever these repeated sequences are, they are local. That is what a workshop looks like and it is not what a civilisation-wide formula looks like.
The position of the repeated run points the same way. When the same four signs turn up in two inscriptions, they sit at the same place in both 59.3% of the time, against 37.0% expected, a ratio of 1.60. A fixed formula occupies a fixed slot, so that ratio is a measure of how formulaic a script is. The comparison is the informative part: Linear B reaches 2.10 and Ur III Sumerian 3.43 on the same measurement. Indus repeated sequences are the least positionally fixed of the writing systems tested here. Whatever recurs in this corpus recurs more freely than the recurring material in two readable administrative archives, which is the opposite of what the objection that these are rote formulae predicts. The comparison corpora were size-matched but not near-deduplicated, so this is suggestive rather than settled.
Significance and effect size answer different questions, and at this corpus size they disagree constantly. The four-sign result carries a z of 375, which is enormous, next to an effect size that is negligible. Quoting either alone would be misleading. Every headline measure on this page has been rescored the same way, and the two strongest of them turn out to be medium effects rather than large ones.
VThe Controls
A structural result is uninterpretable until you know what other systems score. The identical analysis was therefore run on four control corpora, each matched to the Indus corpus on text count, length distribution, token count and vocabulary size. Without that matching one measures corpus size rather than writing system, which is the error that makes cross-corpus comparisons in this field unreliable.
| Corpus | Sequential | Boundary | Positional | Directional |
|---|---|---|---|---|
| Indus script the real corpus | +1.033 | -0.863 | +0.181 | +0.624 |
| Natural language real text, matched | +0.579 | +0.304 | +0.000 | -0.008 |
| Administrative records modelled | +0.537 | -1.965 | +0.354 | +0.758 |
| Ownership marks modelled | -0.468 | -1.433 | +0.221 | +0.766 |
| Random signs structure floor | -0.007 | +0.083 | -0.007 | -0.006 |
Sequential dependency, and only sequential dependency. Indus scores above every control including real natural language, while the ownership-mark model scores negative: emblems carry no information in their ordering. This is a genuine constraint on any account that treats the script as a set of marks.
Directionality and positional fixity. Both are strong in Indus and stronger still in the administrative and emblem models. They establish that inscriptions are ordered and read in a fixed direction; they cannot tell writing from record-keeping. Any argument resting on them alone is unsafe. That includes ours, until this test was run.
Indus sits cleanly with neither family. It is high on sequential dependency like a language, discrete at its boundaries like an administrative record, and positionally rigid like both. On the profile as a whole it is intermediate, which is independently the same conclusion reached by a separate 2026 analysis of this corpus by different means.
The pure emblem hypothesis is the one that struggles. It predicts no sequential dependency, and the Indus corpus has more of it than English prose.
An attack on the central result, and what happened to it
The sharpest objection this work has faced is that it is one slot deep. The Indus opening is unusually restricted, with five signs covering well over half of all first positions. If that single restricted position is generating the sequential dependency, then the whole resemblance to Greek and Sumerian is an artifact of one slot and the central result should be withdrawn. The objection is testable in the bluntest possible way: delete the opening sign from every inscription in all five corpora and measure again.
| Corpus | Intact | Opening removed | Closing removed | Random interior removed |
|---|---|---|---|---|
| Indus | 0.900 | 0.727 | 0.861 | 0.454 |
| Linear B | 0.922 | 0.809 | 0.734 | 0.461 |
| Ur III Sumerian | 0.908 | 0.741 | 0.974 | 0.435 |
| Proto-Elamite | 0.769 | 0.824 | 0.591 | 0.419 |
| Proto-cuneiform | 0.417 | 0.201 | 0.409 | 0.196 |
The objection does not hold. Removing the opening costs Indus 0.173, while removing an arbitrary interior sign costs 0.446, more than twice as much. The restricted opening is not where the dependency lives. More to the point, the readable scripts behave the same way: Linear B and Ur III Sumerian both lose far more from an interior deletion than from an edge one, in the same proportion. If the Indus figure were an artifact of its opening, Indus would have separated from them under this test. It did not.
One corpus does break the pattern, and it breaks it usefully. Proto-Elamite gains information when its opening is removed, because proto-Elamite is the one system in the panel that locks its final position rather than its first. That is an independent confirmation the measurement is sensitive to where a script's constraints actually sit, rather than returning the same shape regardless.
This test was not our idea. It came out of an adversarial pass in which the machinery was asked, repeatedly and from deliberately unrelated disciplines, for the computation most likely to destroy something we currently believe.
A limitation we have not solved
The natural-language control is running prose cut into short segments, so it has no genuine text beginnings or endings. That makes it a poor comparator for directionality in particular, and it is why no claim is made that Indus directionality is language-like, only that it is real. A control built from genuinely short complete texts would settle it. We have not built one.
VIWhat Survived
The findings were attacked before they were published. The strongest objection was that duplicate seal impressions, the same seal pressed many times or mass-produced copies of one design, could manufacture all of these results from nothing. Collapsing every repeated inscription removes 634 texts. Each measure was then recomputed on the reduced corpus, and separately at each major site, so that no finding rests on pooling two traditions that might differ.
| Excess over control | All 2,536 | Dedup. | Mohenjo-daro | Harappa |
|---|---|---|---|---|
| Sequential dependency | 1.396 | 1.033 | 0.932 | 0.746 |
| Order alone | 1.073 | 0.900 | 0.837 | 0.692 |
| Boundary closure | -0.848 | -0.863 | -1.210 | -1.140 |
| Directionality | 0.687 | 0.625 | 0.605 | 0.554 |
| Positional fixity | 0.213 | 0.182 | 0.199 | 0.219 |
| Grammatical categories (withdrawn, z) | 2.61 | 1.47 | 1.56 | 2.35 |
The effects persist after deduplication and hold independently at Mohenjo-daro and Harappa, two sites some six hundred kilometres apart, showing the same structure.
How many independent inscriptions are there really?
Collapsing only identical inscriptions is the weak version of this control. Indus seals were also copied, and near-identical texts are not independent observations either. Grouping inscriptions that differ by one or two sign edits reduces 1,902 distinct sequences to 1,141 and then to 558. Intervals computed on the raw count are therefore too narrow, which is a problem for most published statistics on this corpus, and was for ours.
| Measure (excess, with z) | 1,902 exact | 1,141 at edit 1 | 558 at edit 2 |
|---|---|---|---|
| Sequential dependency | +1.033 71.5 | +0.875 54.5 | +0.665 34.0 |
| Directionality | +0.625 61.7 | +0.561 45.5 | +0.464 26.7 |
| Positional fixity | +0.182 35.1 | +0.151 24.8 | +0.125 18.0 |
| Opening vs closing classes | z 6.0 | z 4.6 | z 0.5 |
The three main results hold on 558 inscriptions, under a quarter of what we started with. That is the number they should be judged on.
Is the headline number an artifact of the estimator?
This is the sharpest objection available and it has been made before, against earlier entropy work on this corpus. At 11,135 tokens across 591 sign types the table of sign pairs is 0.84 per cent occupied, 2,919 filled cells out of 349,281. In that regime the naive maximum-likelihood estimator overstates dependency badly, and a figure quoted without correction cannot be trusted.
Our figure is already a margin over a permutation control computed with the same estimator, which cancels most of that bias because the control shares the sample size, alphabet and marginal frequencies. It does not cancel it exactly. So the same quantity was recomputed under three published bias-corrected estimators, each applied to the observed data and to the control alike.
| Estimator | Observed | Control | Excess | 95% interval |
|---|---|---|---|---|
| Maximum likelihood | 3.169 | 2.136 | +1.033 | +0.873 to +0.978 |
| Miller-Madow | 2.971 | 1.778 | +1.192 | +1.014 to +1.124 |
| Chao-Shen | 3.001 | 1.566 | +1.435 | +1.370 to +1.511 |
| Chao-Wang-Jost | 2.538 | 0.861 | +1.677 | +1.495 to +1.742 |
Correction moves the result upward, not downward. The excess rises from 1.033 to 1.677 bits as the correction gets stronger, because the correction reduces the control far more than it reduces the observed data. The shuffled corpus is dominated by sign pairs seen once, which is exactly the situation these estimators exist to fix. The uncorrected figure we publish is therefore the conservative one, and every estimator tested puts the interval clear of zero.
An earlier pass found that signs sort into grammatical categories which then line up in order along the inscription. It was the most interesting result of the project and the closest thing to a grammar. It did not survive. Once duplicate impressions were collapsed the effect fell to chance, and it failed to reproduce consistently across sites. The duplicates had been manufacturing it.
It is recorded here because a finding that dies under its own control test is worth more to other researchers than one that was never tested, and because anyone running similar analyses on this corpus should expect the same trap.
What was left when we looked again
Withdrawing a result is not the same as showing there is nothing there. The original test asked one broad question (do a dozen induced categories arrange themselves in order?) and spent all its evidence on it. Three narrower tests were run afterwards on the deduplicated corpus, each able to fail independently.
| Second-pass test | Result | Against |
|---|---|---|
| Cluster stability under resampling | 0.266 | 0.049 |
| Category model vs full sign model (bits, held-out) | +0.0284 | t = 1.89 |
| Opening vs closing sign contexts | z = 6.02 | p < 0.001 |
The categories are more stable under resampling than chance, but not by much. Compressing 591 signs into a dozen categories costs nothing measurable on unseen text. The interval crosses zero, so the honest reading is no detectable loss, not an improvement.
The third test held up at first. The signs that open inscriptions and those that close them differ in the company they keep, and not merely in where they sit: the comparison uses each sign's neighbours with position information removed. Twenty-one openers and forty-eight closers form two distributionally distinct groups at z = 6.02.
It does not survive the harder duplicate control. Collapsing near-identical inscriptions as well as exact ones takes it to z = 4.6 at edit distance one, and to z = 0.5 at edit distance two, which is nothing at all. We are leaving it here as a reported result rather than a claim: it may be real and merely under-powered once the corpus is cut to genuinely independent inscriptions, or it may be another artifact of repeated seals. On this corpus we cannot tell.
VIIWhat This Settles
- Sign order carries information, and more of it than in running prose
- Inscriptions are complete, self-contained statements
- The script is directional, with distinct opening and closing positions
- Signs hold fixed positions rather than floating freely
- The same system is in use at sites hundreds of kilometres apart
- What any inscription says
- What language, if any, underlies the script
- The phonetic or semantic value of any sign
- Whether the script is logographic, syllabic or mixed
- Any connection to a known language family
- Whether signs group into a full system of grammatical categories
though opening and closing signs do form two distinct classes
No one has deciphered the Indus script, and this work does not either. Claims to have read it appear regularly and are not supported by evidence of this kind; neither is anything here. Not one inscription is translated. One sign opens roughly a third of the corpus and we do not know what it means.
The strongest honest statement the data supports is this: the Indus script carries ordered, self-contained, positionally organised structure, and more dependency between successive signs than English prose. That is a real constraint, and the emblem and ownership-mark readings have to answer it. It is not the same as showing the script encodes a language, and we are not showing that.
VIIIThe Picture and the Object
Every Indus seal carries an image, most often a one-horned bull, sometimes an elephant, a rhinoceros, a tree or a human figure, set above the row of signs. The natural reading is that the signs caption the picture. If so, the two channels must share information, and the image would become a weak bilingual: a known concept standing beside an unknown word.
| Channel | Observed | Control | Excess | z |
|---|---|---|---|---|
| Iconographic motif | 0.2978 | 0.2814 | +0.0164 | 2.0 |
| Object type | 0.2085 | 0.1437 | +0.0648 | 14.0 |
| Material | 0.1889 | 0.1525 | +0.0364 | 7.1 |
The caption reading is not supported. Before deduplication the motif appears to carry substantial information about the signs; after repeated seal impressions are collapsed, almost all of it disappears. What remains does not clear the bar once the three channels tested here are accounted for. The picture and the text are, as far as this measurement goes, independent channels. Whatever the signs record, it is not a description of the animal above them.
The object itself is a different story. What a piece is, whether seal, tablet, tag or potsherd, genuinely conditions which signs appear on it, and survives deduplication comfortably. Text length tracks it too: seals average 5.06 signs, tablets 3.70, pottery 2.78. The corpus is not one uniform practice, and analyses that pool it are averaging over at least two.
Whether it could be a list of names
A standing objection to the whole enterprise is that these are ownership marks, personal names, lineages and titles, rather than text. That hypothesis makes a sharp prediction. A corpus of one-off identifiers never settles down: every new object introduces new signs, and the vocabulary grows about as fast as the corpus itself. Reused language does the opposite, because a finite lexicon is being drawn on repeatedly.
| Corpus | Vocabulary growth | Reading |
|---|---|---|
| Indus script | 0.488 | finite vocabulary, reused |
| English prose (matched) | 0.415 | finite vocabulary, reused |
| Indus, seals only | 0.506 | |
| Indus, tablets only | 0.581 |
Indus sits at 0.488, against 0.415 for matched English prose. Vocabulary growth is slightly faster, consistent with more one-off signs, but nowhere near the behaviour of a list of unique identifiers. The sign inventory is genuinely reused. Whatever the signs are, they are drawn from a working vocabulary rather than minted per object, and the pure name-list reading has to account for that.
Whether they are accounts
The third standing reading is that these are administrative records: so many measures of grain, so many head of cattle. That is what the neighbouring systems are. It also makes a prediction that can be tested without reading a single sign, because two of the comparison corpora label their numerals explicitly, and a numeral does not behave like a word.
In proto-cuneiform and proto-Elamite, numerals account for 44% and 42% of all sign tokens respectively. They cluster in one region of the text. Above all they travel in company: where a numeral appears at all, another appears in the same short text 32% of the time, because an account that records one quantity almost always records a second. That combination is a signature, and it can be learned where the answer is known and then hunted where it is not.
A third corpus was added later and agrees. Linear A, undeciphered but certainly writing, keeps its numerals in a separate Unicode block so they identify themselves without any interpretation; its numerals co-occur at 71%, higher than either Mesopotamian system. Three independent corpora, two of them undeciphered, all behave the same way.
Nothing in the Indus corpus matches it. Scoring every sufficiently common sign against the learned profile, the closest candidates reach a co-occurrence rate between 2% and 7%, against roughly 32% for genuine numerals. No sign in this corpus behaves like a number. Either these inscriptions do not record quantities at all, which would sit comfortably with objects used to mark ownership rather than to tally goods, or they record them in a way that shares no statistical property with how every neighbouring bureaucracy did it. The result is negative, and it closes off the one anchor that decipherment of a related script has historically depended on. Proto-Elamite numerals were read long before anything else in proto-Elamite was. That door is not open here.
IXThe Hand and the Stone
Everything so far has treated an inscription as a string of symbols. It is also a physical act: a person cutting into a piece of steatite about the size of a postage stamp, who had to decide, before the first cut, how much would fit. The source database records width, height and thickness for 2,325 seals, and those measurements can be set against the texts they carry.
Longer inscriptions sit on larger objects. Across 1,607 pieces with recorded dimensions, text length and seal width move together at rho 0.445. That number on its own means little, because it could be nothing more than a cataloguing convention: if long texts belong on big copper tablets and short ones on small clay tags, the correlation is a fact about object classes rather than about people.
| Test | n | Correlation | z |
|---|---|---|---|
| Object type and material held fixed lengths permuted within class | 1,607 | 0.445 | 15.4 |
| Face area rather than width the whole writing surface | 1,607 | 0.411 | 13.1 |
| Steatite seals alone one object, one material | 1,030 | 0.406 | 13.0 |
| After collapsing near-duplicates edit distance two | 481 | 0.398 | 8.6 |
It survives all of it. Holding object type and material fixed, the effect is unchanged. Restricted to steatite seals alone, one object in one material, it is unchanged again. It holds separately at Mohenjo-daro and at Harappa, and it is still present after near-identical inscriptions are collapsed at edit distance two, which is the gate that killed five earlier findings in this project. The one place it is absent is clay, where the correlation is 0.038 on 63 pieces.
Whether the text was planned before the first cut
Two quite different craftsmen produce that same correlation. One knows the text before he starts and picks a blank large enough to hold it, so the signs stay the same size and the stone grows with the message. The other starts carving on whatever is to hand and squeezes as he runs out of room, so the stone stays the same and the signs shrink. Only the first is planning, and the two can be separated by asking how much of a longer text is absorbed by a larger surface.
Writing the width of the surface against the number of signs as a power law, the exponent answers the question directly. An exponent of one means every extra sign was paid for with proportionally more stone. An exponent of zero means none of it was, and every extra sign was paid for by cutting smaller.
| Corpus cut | n | Exponent | 95% interval | z |
|---|---|---|---|---|
| Steatite seals | 1,030 | 0.252 | [0.224, 0.282] | 12.4 |
| Tablets | 322 | 0.376 | [0.306, 0.450] | 7.2 |
| All objects | 1,607 | 0.312 | [0.287, 0.341] | 17.4 |
| Near-duplicates collapsed, distance one | 980 | 0.372 | [0.335, 0.409] | 12.9 |
| Near-duplicates collapsed, distance two | 481 | 0.457 | [0.397, 0.524] | 8.8 |
The answer is both, and mostly the second. On steatite seals the exponent is 0.25, rising to 0.46 once near-duplicates are collapsed. Somewhere between a quarter and a half of a longer text was accommodated by choosing a bigger seal; the remainder was accommodated by carving the signs smaller. A carver with ten signs to fit did not simply reach for a stone twice the size. He reached for one slightly larger and cut finer.
That is a modest observation and it is worth being clear about what it does and does not support. It does not show the signs are language. What it does show is that the length of the message was known, at least approximately, before the carving began, because the surface was chosen with reference to it. A sequence produced without foresight, one mark at a time until the maker stopped, has no reason to correlate with the size of the object it is on. This is the first result in this work that comes from the objects rather than from the symbols, and it is independent of every information-theoretic argument above.
The same measurement initially separated the two great cities: Mohenjo-daro at 0.188 and Harappa at 0.476, with non-overlapping intervals, which would have meant the two capitals had measurably different workshop habits. It does not hold. The excavated assemblages differ sharply, 290 tablets and 224 seals at one site against 781 seals and 47 tags at the other, and once object type is held fixed the separation weakens; once the width and length distributions are matched as well, the gap is -0.033 and the significance is gone entirely. What looked like a difference between two cities was a difference between two excavations. It is recorded here rather than deleted, because the reason it failed is more useful than the result would have been.
XThe Frame
This page has said since its first version that the opening slot of an Indus inscription is sharply restricted. That is a statement about one position. The question never asked was whether the other positions are restricted too, and the answer changes what kind of object an inscription is.
The difficulty is that position has to be measured carefully. Inscriptions here average under five signs, so if position is expressed as a fraction of the way through a text, a sign that happens to occur in short inscriptions is handed a middle position by the arithmetic with no scribe involved. The test therefore has to be run inside a single inscription length, where every slot exists and every slot is equally available, against a control that shuffles each inscription's own signs and so holds length, membership and every sign frequency exactly fixed.
| Corpus | 4 signs | 5 signs | 6 signs | 7 signs | Bound to a slot |
|---|---|---|---|---|---|
| Indus | 21/27 | 29/34 | 24/27 | 19/26 | 82% |
| Indus, near-duplicates collapsed | 7/12 | 23/31 | 22/26 | 20/26 | 76% |
| Linear B | 14/34 | 17/39 | 6/34 | 7/25 | 33% |
| Ur III Sumerian | 22/29 | 23/27 | 20/26 | 14/21 | 77% |
| Proto-Elamite | 30/35 | 31/46 | 26/32 | 13/16 | 78% |
| Proto-cuneiform | 24/29 | 29/40 | 23/32 | 15/26 | 72% |
Indus inscriptions are strongly slotted: 82% of testable signs are bound to a position, and 71% survive after near-identical inscriptions are collapsed. This is not the opening restriction in another guise. Delete the first sign of every inscription and sweep again, and the anchors are still there.
The frame has named parts
What makes this a frame rather than a statistic is that individual signs hold specific slots, and hold them across every inscription length. The demonstration is in the two columns below. As an inscription grows from four signs to seven, a sign anchored to the front keeps its position counted from the front while its position counted from the back slides; a sign anchored to the end does the exact reverse. No artifact of length or of measurement produces two groups of signs moving in opposite directions.
| Sign | Slot counted from the front | Slot counted from the back | Anchored to | Share of its uses |
|---|---|---|---|---|
| sign 740 | 1, 1, 1, 1 | 3, 4, 5, 6 | first slot | 68% |
| sign 2 | 3, 4, 5, 6 | 1, 1, 1, 1 | penultimate slot | 57% |
| sign 60 | 3, 4, 5, 6 | 1, 1, 1, 1 | penultimate slot | 83% |
| sign 820 | 4, 5, 6, 7 | 0, 0, 0, 0 | last slot | 77% |
| sign 520 | 1, 1, 1, 1 | 3, 4, 5, 6 | first slot | 82% |
| sign 741 | 3, 4, 5 | 2, 2, 2 | 2 from the end | 59% |
| sign 817 | 4, 5, 6 | 0, 0, 0 | last slot | 90% |
| sign 861 | 5, 6, 7 | 0, 0, 0 | last slot | 70% |
So the corpus has a first-position set, a last-position set and a penultimate-position set, each occupied by particular signs that stay in their part of the inscription whatever its length. An Indus inscription behaves less like a sentence and more like a filled form. That is a structural claim, not a reading, but it is a useful one: it says where to look. If one slot admits only a handful of signs, those signs are a field type, and field types are the first thing anyone reads in a bureaucratic script.
The control this panel was missing
Every comparison so far has set Indus against scripts that can be read, which invites the reply that readability itself is doing the work. There is one script that answers it. Linear A is undeciphered, its language unidentified, and nobody doubts it is writing, because Linear B was adapted from it sign by sign and Linear B is Greek. It is the only available example of an unreadable but unquestionably real writing system, and it belongs beside Indus more than anything else in this panel does.
| Corpus | Status | Signs tested | Bound to a slot |
|---|---|---|---|
| Indus | undeciphered, unknown | 228 | 61% |
| Linear A | undeciphered, certainly writing | 23 | 26% |
| Linear B | deciphered, Greek | 245 | 23% |
| Proto-cuneiform | attested non-language notation | 276 | 41% |
Linear A sits with Linear B, and Indus sits above both. The two Aegean scripts agree with each other, which is what their shared ancestry predicts. The Linear A sample is small, 23 signs testable against 228 for Indus, so this is a difference of degree rather than a settled quantity. It is also, as the next section shows, a comparison between the wrong kinds of object.
The comparison that was on our own disk
Every corpus named so far is an archive of tablets. Indus is a corpus of seals. This work has itself shown that the physical object shapes the text, so comparing seals against tablet archives and concluding something about the script was a mistake, and it was ours.
The correction required no new acquisition. The CDLI catalogue, held here since the beginning of this work, records object type, and 17,024 of its records are seals rather than tablets. Of those, 6,365 are Ur III, which is 2100 to 2000 BC and therefore contemporary with the mature Indus period. They are the same kind of object, from the same centuries, serving the same administrative purpose. And they can be read.
What they say is a formula: personal name, son of personal name, servant of a god or king, often with a title. That is precisely what Indus seals have been supposed to carry for a century, which makes this the one place where the frame machinery can be checked against a known answer. The prediction was written down before the test ran.
| Token | Meaning | Predicted position | Position found | Share of its uses |
|---|---|---|---|---|
| dumu | son of | middle | middle | 57% |
| arad2 | servant of | late | penultimate slot | 75% |
| arad2-zu | your servant | last | last slot | 98% |
| dub-sar | scribe, a title | early | slot 2 | 79% |
| ensi2 | governor, a title | early | slot 2 | 70% |
| lugal | king | late | middle | 45% |
It recovers the formula unprompted. Personal names take the first slot, titles take the second, the word for "son of" sits in the middle, and "your servant" takes the last slot in 98 per cent of its appearances. The machinery that found the Indus frame finds the correct frame in a language it was told nothing about. That is the validation the Indus result needed, and it passes.
One prediction is only half met, and it is left in the table rather than dropped. Lugal, "king", was predicted late and comes out middling. The reason is visible once the legends are read: it appears both as the final royal name and inside compound names earlier in the line, so it genuinely occupies two roles. A test that scored five out of five would be more suspicious than one that scores four.
Then the comparison itself, with both corpora put through the identical near-duplicate collapse, because Ur III seals repeat heavily for the same reason Indus seals do: one official's seal was rolled onto many tablets.
| Corpus | Signs tested | Bound to a slot |
|---|---|---|
| Ur III seal legends, words | 61 | 97% |
| Ur III seal legends, signs | 142 | 75% |
| Indus inscriptions | 95 | 76% |
Read at sign granularity, Indus is slotted to the same degree as contemporary readable seal legends: 76% against 75%. Read at word granularity, the same Ur III corpus scores 97 per cent and Indus sits well below it. Both rows are in the table above and neither can be preferred, because nobody knows whether an Indus sign answers to a Sumerian sign or to a Sumerian word. The word reading is the one whose mean legend length matches an Indus inscription almost exactly, 4.7 tokens against 4.65; the sign reading spreads the same formula over roughly twice as many slots, which is why the figure falls. Neither of those is a reason to choose.
So the result is an interval rather than a point: Indus is slotted at or below the level of readable seal legends, matching them exactly at one end of the interval and falling short at the other. An earlier version of this page reported only the matching end. Three further seal corpora extracted since, Early Old Babylonian, Old Babylonian and Old Assyrian, all score 100 per cent at word granularity, so the upper end of the interval is not a peculiarity of the Ur III chancery but what seal legends do across six centuries. Set beside tablet archives Indus looked like an outlier; set beside the same object from the same centuries it is ordinary or slightly less templated, and the earlier reading on this page has been corrected accordingly.
This does not say Indus seals carry names and titles. It says that if they did, they would look like this, and that nothing in their positional structure argues against it. That is the closest structural analogue this work has found for what an Indus inscription might be, and the hypothesis it favours is the oldest and least exciting one in the field.
The obvious temptation is to read a strongly slotted script as an un-language, and the panel forbids it. Ur III Sumerian is a fully deciphered language and it scores 77% on this measurement, essentially the same as Indus. Linear B is also a fully deciphered language and it scores 33%, less than half as much. Two readable languages sit at opposite ends of the scale.
Slot binding therefore does not separate writing from non-writing at all. It separates formulaic text from free text, and it puts Indus with administrative Sumerian rather than with the more varied Linear B tablets. That is a real finding about what these objects are for. It is not evidence either way about whether the signs encode speech, and anyone quoting the Indus figure without the Ur III figure beside it is misusing it.
XIHow Close Is This To A Decipherment
Reading an unknown script is not one problem but a sequence of them, and the later ones are far harder than the earlier ones. Setting them out in order makes it possible to say where this work actually sits, and where the field sits, without either being flattered.
5 of 11 rungs cleared here, one partial, which is roughly 50% of the way up. That figure should be read with care. The rungs are not equal in size, and the remaining ones are much larger than the ones already climbed.
Nobody has verifiably cleared rung nine or beyond, ourselves included. Rung nine has been claimed more than once and never independently confirmed. The last three rungs are what the word decipherment actually refers to, and no one is standing on them.
What this work adds is confined to rungs four and five, and to narrowing rung six. That is a genuine contribution to a hundred-year-old problem and it is also a long way from reading the script. Both halves of that sentence matter.
XIIWhat Is New Here, And What Is Not
A structural analysis of this corpus published in 2026 by Ashish Nair, How Non-Linguistic Is the Indus Sign System? A Synthetic-Baseline Scorecard (arXiv:2604.17828), works from the same underlying digitisation and computes several of the same quantities. Anything here has to be read against it, so the overlap is set out first rather than left for a reader to find.
Independently replicated
These were computed here before that paper was read, from different sign identifiers and separate code. The agreement is close enough to be worth stating plainly.
| Quantity | Nair 2026 | This work |
|---|---|---|
| Zipf rank-frequency slope | -1.492 | -1.489 |
| Zipf fit R squared | 0.956 | 0.957 |
| Hapax rate | 33.2% | 33.7% |
| Mean inscription length | 4.42 | 4.39 |
| Exact duplicate rate | 24% | 25% |
Extends existing work
The duplicate-structure problem was identified in that paper at the level of exact repeats. Here it is carried further: inscriptions differing by one or two sign edits are also not independent observations, which takes 1,902 distinct sequences down to 1,141 and then 558, and every headline measure is reported at all three levels. Separately, the bias-correction result in section IV appears to run against expectation, since correcting a null-referenced statistic increases the margin rather than shrinking it.
Not previously attempted, as far as we can establish
No dispersion measure appears to have been applied to this corpus. We applied one, asking whether signs divide into a frequent evenly-spread class and a rare bursty class, which is how function words separate from content words in every natural language. They do not. The apparent split is 94 per cent explained by frequency alone, and once each sign is compared against a null matched to its own frequency, not one of 153 signs survives correction. This is a negative result and it is the clearest thing in this work that nobody had checked.
Withdrawn
Five results looked strong and did not survive their own controls: a twelve-category induced grammar, a claim that those categories recovered most of the script's predictability, a correlation between the picture on a seal and the signs above it, the closed-class split above, and a distinction between opening and closing sign classes. Each is described where it arose. None is claimed.
The comparison against English prose in section III uses continuous text cut into short segments. It is matched on text count, length distribution, token count and vocabulary size, but it is not inscriptional writing and it has no genuine text beginnings or endings. Until the same battery is run against a known script of comparable genre and identical sample size, such as Linear B tablet lines, the claim that Indus exceeds real writing on sequential dependency should be read as provisional.
This is the same gap as in the prior work, which uses no linguistic control at all. It is the next thing being built here, and it is the test on which this argument should be judged.
XIIIMethods
The implementation is not published. The protocol is, in enough detail to be attacked.
Corpus
2,536 inscriptions from the published digitisation of the Corpus of Indus Seals and Inscriptions: 11,135 sign tokens, 591 sign types, mean length 4.39, ten sites. Analysis runs on the deduplicated set unless stated. Object type, material and iconographic motif are carried as covariates. No reconstructed or generated text is used anywhere.
Statistics and controls
Every quantity is reported as a margin over a matched null, never as a raw value. Two nulls are used: a global shuffle preserving sign frequencies and length distribution, and a within-inscription shuffle preserving which signs co-occur while destroying their order. The difference between the two isolates ordering from co-occurrence. Nulls run at 1,000 to 5,000 iterations. Significance is a z-score against the null distribution.
Estimator bias
At 0.84 per cent occupancy of the sign-pair table, maximum-likelihood entropy is strongly upward-biased. Referencing every statistic to a null computed with the same estimator removes that bias to first order. Because it does not remove it exactly, the principal result is also reported under Miller-Madow, Chao-Shen and Chao-Wang-Jost corrections.
Independence
Seals were copied, so inscriptions are not independent draws. Results are recomputed after collapsing exact duplicates and then near-duplicates at edit distance one and two, and separately at Mohenjo-daro and Harappa. Intervals come from subsampling inscriptions without replacement; resampling with replacement reintroduces the duplicate structure the analysis exists to control for, and produces intervals that do not bracket their own estimate.
Comparison corpora
Controls are matched to the Indus corpus on text count, length distribution, token count and vocabulary size, then run through the identical procedure. Four arms are used: natural language, a modelled administrative-record system, a modelled ownership-mark system, and randomly ordered signs. The random arm returning approximately zero on every measure is the calibration check.
Discipline
Findings are put to an adversarial review instructed to refute rather than confirm them, and any result that fails deduplication, per-site replication or a matched null is withdrawn and recorded rather than removed. Where a hypothesis is generated and tested automatically, it is registered with its predicted direction before the test runs, and significance is judged against a false-discovery threshold recomputed over every test performed.
Availability
Corpus provenance is public and named above, so the inputs are independently obtainable. Researchers wishing to compare results against their own measurements, or to see the figures behind any table here, are welcome to make contact.