Indus Valley Script · structural analysis 28 July 2026
Undeciphered Scripts · c. 2600–1900 BCE

Short Texts, Not Absent Structure Linear B, cut to Indus dimensions, reproduces the statistical profile that has been used to argue the Indus script is not writing

2536 inscriptions11,135 sign tokens 591 signs28 July 2026
An Indus Valley stamp seal in stone beside its modern clay impression. A row of script signs runs along the top edge above a one-horned bull.
A stamp seal and its modern impression. The row of signs along the top is the script; the animal below it is the motif. Our measurements show the two channels carry almost no information about each other. Public domain, released under CC0.

The Indus script has resisted decipherment for a century. A recurring argument holds that it never encoded language at all: its texts are too short, its dependencies too local, its structure too thin. That argument has never been tested against the obvious control, which is a script we can already read, cut to the same dimensions.

We did that. Linear B, deciphered in 1952 and known to encode Greek, was subsampled to the Indus token count and length distribution and run through the identical procedure, alongside proto-cuneiform, which is attested notation that was not yet writing. On the information carried purely by sign order, Indus scores 0.901 bits and Linear B 0.913. Proto-cuneiform manages 0.405. Both Indus and Linear B lose almost all dependency beyond adjacent signs, which is the observation most often cited as evidence against Indus being writing.

The absence of long-range structure in the Indus corpus is therefore a consequence of short texts, not a property of the script. This does not show that the Indus script is writing, and nothing here reads a single sign. It shows that the strongest statistical argument that it is not does not survive its own control.

IThe Control That Was Missing

Every published structural analysis of the Indus corpus, including our own until now, has compared it against systems chosen to be non-linguistic: heraldry, ownership marks, tallies, invented baselines. None has compared it against a known writing system measured at the same size. Without that, a result showing Indus to be statistically impoverished says nothing, because nobody had established what real writing looks like when you only ever get four signs at a time.

Three corpora, all attested, all openly published. Every one cut to the Indus token count and the Indus length distribution, sampled without replacement and averaged over repeated draws.

MeasureIndusLinear B Proto-cuneiform
texts / tokens1,902 / 8,839 1,893 / 8,8411,907 / 8,710
mean length4.654.67 4.57
information in sign order alone 0.9010.913 0.405
dependency at distance 1+0.901 +0.912+0.404
dependency at distance 2 +0.077+0.095 −0.118
dependency at distance 3−0.106 −0.054−0.421
recurring three-sign blocks0.235 0.3300.046
vocabulary growth0.4880.294 0.431
signs occurring once0.3640.193 0.210
Table 1. Three attested corpora at identical dimensions. Linear B is deciphered administrative writing in a known language. Proto-cuneiform is attested notation that records commodities and quantities without encoding speech.

The two highlighted rows carry the argument. On information carried by sign order alone, Indus and Linear B are within 1.3 per cent of each other. And both collapse to almost nothing at distance two, which is precisely the observation that has been offered as evidence that Indus cannot be language. A script we have read for seventy years does the same thing at this text length.

Proto-cuneiform, which really is a notation rather than a script, sits at less than half the Indus figure and turns negative at distance two. Whatever the Indus script is, it is not behaving like that.

What this does and does not show

It does not show that the Indus script encodes language. It shows that one specific argument that it does not, the argument from missing long-range structure, fails its own control. That is a narrower claim and it is the one the data supports.

One limitation matters and is not hidden here. The arms are matched on text count, length and token count, but they cannot be matched on inventory size: Indus has 591 sign types where Linear B offers 211, because Linear B does not contain 591 signs. A larger inventory means sparser data and more estimator bias, which plausibly explains part of the remaining gap in the total figures. It does not explain the ordering result or the decay pattern, which are the two rows the argument rests on.

IIThe Corpus

Every figure on this page comes from excavated inscriptions in the published digitisation of the Corpus of Indus Seals and Inscriptions. No reconstructed, simulated or generated text enters the analysis at any point. This is worth stating plainly because the alternative is common in this field and hard to detect from outside.

Inscriptions2536 1902 after collapsing duplicate seal impressions
Sign tokens11,135 mean 4.39 signs per inscription, maximum 17
Distinct signs591 34% occur exactly once
Sites10 Mohenjo-daro, Harappa, Lothal, Dholavira, Kalibangan and others
Table 2. Corpus composition.

IIIThe Finding

At this corpus size every measure of structure is inflated. With 591 distinct signs and fewer than nine thousand tokens, the table of sign pairs is about two per cent occupied, and in that regime a corpus shuffled into complete meaninglessness still returns a large positive score. A raw figure therefore means nothing. What follows is in every case the margin over a matched control: the same corpus, its structure destroyed, measured the same way.

Dependency that survives the control
33%
Adjacent signs in this corpus measure 3.169 bits of mutual information. The same signs re-dealt at random measure 2.136. Two thirds of the raw figure is sampling bias; the remaining 1.033 bits is the finding. Matched English prose returns +0.579.
MeasureObservedControl Excessz
Sequential dependency
order carries information
3.1692.1361.03371.5
Order alone
controlling for co-occurrence
3.1692.2690.90064.9
Boundary closure
negative = inscriptions are closed
2.2263.089-0.863-18.9
Directionality
opening vs closing sign sets
0.8310.2060.62561.7
Positional fixity
signs hold their region
0.6600.4780.18235.1
Table 3. Structural measures on the deduplicated corpus, each against its matched control. Boundary closure is negative by construction of the test: sign pairs spanning two different objects share less information than chance.

The second row is the one that rules out a weaker reading. Scrambling the order of signs within an inscription, while keeping exactly which signs co-occur, still destroys 0.900 bits. The information is in the sequence, not merely in the company.

The boundary result is the cleanest in the study. Whatever an inscription says, it finishes saying it: dependency across the join between two artifacts falls below chance. Each object carries one complete statement, and the corpus is not the scattered fragments of a longer running text.

IVInscription Families

A separate question, and one that does not depend on any of the entropy machinery: do inscriptions repeat each other? A workshop, a single scribe, or a fixed formula should leave relatives behind, sharing long runs of signs more often than chance permits.

Shared runPairs observed Expected by chanceEffect size
three or more signs4,064147 h = 0.08
four or more signs4671 h = 0.03
Table 4. Inscription pairs sharing an identical run of signs, against a control preserving sign frequencies and inscription lengths. Effect size is Cohen's h on the rate.

Four hundred and sixty-seven pairs of inscriptions share a run of four or more identical signs, where the control predicts one. That is a real relationship and not a marginal one.

The effect size says the other half of it. Those 467 pairs are three hundredths of one per cent of all possible pairs, so on the rate the effect is negligible. The corpus contains a small number of genuine families, not a corpus-wide system of formulae. Both halves of that sentence come from the same measurement, and reporting only the first would misrepresent it.

The longest shared runs reach seven signs, on pairs of inscriptions that are otherwise unrelated objects. Those are the best candidates in the corpus for two artifacts carrying the same text, and they are a natural target for anyone with access to the originals.

The families are local

If these families are the work of individual hands, they should betray it physically. A single workshop stands in one place, draws on one stock of material and makes one kind of object, so two inscriptions from the same hand ought to agree about where they were found far more often than two inscriptions picked at random. If instead the shared runs are a standard formula circulating across the whole civilisation, they should be scattered.

They are not scattered. Pairs sharing a run of four signs come from the same site 59.1% of the time against 36.5% expected by chance, and that holds after near-identical inscriptions are collapsed at edit distance two. Agreement on object type and material runs the same way but weakens once pairs from different sites are compared directly, which says the object and the material are largely following the geography rather than adding to it. Whatever these repeated sequences are, they are local. That is what a workshop looks like and it is not what a civilisation-wide formula looks like.

The position of the repeated run points the same way. When the same four signs turn up in two inscriptions, they sit at the same place in both 59.3% of the time, against 37.0% expected, a ratio of 1.60. A fixed formula occupies a fixed slot, so that ratio is a measure of how formulaic a script is. The comparison is the informative part: Linear B reaches 2.10 and Ur III Sumerian 3.43 on the same measurement. Indus repeated sequences are the least positionally fixed of the writing systems tested here. Whatever recurs in this corpus recurs more freely than the recurring material in two readable administrative archives, which is the opposite of what the objection that these are rote formulae predicts. The comparison corpora were size-matched but not near-deduplicated, so this is suggestive rather than settled.

Why this is reported with two numbers

Significance and effect size answer different questions, and at this corpus size they disagree constantly. The four-sign result carries a z of 375, which is enormous, next to an effect size that is negligible. Quoting either alone would be misleading. Every headline measure on this page has been rescored the same way, and the two strongest of them turn out to be medium effects rather than large ones.

VThe Controls

A structural result is uninterpretable until you know what other systems score. The identical analysis was therefore run on four control corpora, each matched to the Indus corpus on text count, length distribution, token count and vocabulary size. Without that matching one measures corpus size rather than writing system, which is the error that makes cross-corpus comparisons in this field unreliable.

CorpusSequentialBoundary PositionalDirectional
Indus script
the real corpus
+1.033-0.863+0.181+0.624
Natural language
real text, matched
+0.579+0.304+0.000-0.008
Administrative records
modelled
+0.537-1.965+0.354+0.758
Ownership marks
modelled
-0.468-1.433+0.221+0.766
Random signs
structure floor
-0.007+0.083-0.007-0.006
Table 5. Matched controls. Read the last row first: randomly ordered signs score zero on everything, which is the check that the method is calibrated and not manufacturing structure.
What separates Indus

Sequential dependency, and only sequential dependency. Indus scores above every control including real natural language, while the ownership-mark model scores negative: emblems carry no information in their ordering. This is a genuine constraint on any account that treats the script as a set of marks.

What does not

Directionality and positional fixity. Both are strong in Indus and stronger still in the administrative and emblem models. They establish that inscriptions are ordered and read in a fixed direction; they cannot tell writing from record-keeping. Any argument resting on them alone is unsafe. That includes ours, until this test was run.

Where that leaves the question

Indus sits cleanly with neither family. It is high on sequential dependency like a language, discrete at its boundaries like an administrative record, and positionally rigid like both. On the profile as a whole it is intermediate, which is independently the same conclusion reached by a separate 2026 analysis of this corpus by different means.

The pure emblem hypothesis is the one that struggles. It predicts no sequential dependency, and the Indus corpus has more of it than English prose.

An attack on the central result, and what happened to it

The sharpest objection this work has faced is that it is one slot deep. The Indus opening is unusually restricted, with five signs covering well over half of all first positions. If that single restricted position is generating the sequential dependency, then the whole resemblance to Greek and Sumerian is an artifact of one slot and the central result should be withdrawn. The objection is testable in the bluntest possible way: delete the opening sign from every inscription in all five corpora and measure again.

CorpusIntactOpening removed Closing removedRandom interior removed
Indus0.9000.7270.8610.454
Linear B0.9220.8090.7340.461
Ur III Sumerian0.9080.7410.9740.435
Proto-Elamite0.7690.8240.5910.419
Proto-cuneiform0.4170.2010.4090.196
Table 6. Information carried by sign order after ablating one position. Every corpus is cut to the same token count and length distribution. Removing any sign costs information simply by shortening the text, which is what the last column measures.

The objection does not hold. Removing the opening costs Indus 0.173, while removing an arbitrary interior sign costs 0.446, more than twice as much. The restricted opening is not where the dependency lives. More to the point, the readable scripts behave the same way: Linear B and Ur III Sumerian both lose far more from an interior deletion than from an edge one, in the same proportion. If the Indus figure were an artifact of its opening, Indus would have separated from them under this test. It did not.

One corpus does break the pattern, and it breaks it usefully. Proto-Elamite gains information when its opening is removed, because proto-Elamite is the one system in the panel that locks its final position rather than its first. That is an independent confirmation the measurement is sensitive to where a script's constraints actually sit, rather than returning the same shape regardless.

This test was not our idea. It came out of an adversarial pass in which the machinery was asked, repeatedly and from deliberately unrelated disciplines, for the computation most likely to destroy something we currently believe.

A limitation we have not solved

The natural-language control is running prose cut into short segments, so it has no genuine text beginnings or endings. That makes it a poor comparator for directionality in particular, and it is why no claim is made that Indus directionality is language-like, only that it is real. A control built from genuinely short complete texts would settle it. We have not built one.

VIWhat Survived

The findings were attacked before they were published. The strongest objection was that duplicate seal impressions, the same seal pressed many times or mass-produced copies of one design, could manufacture all of these results from nothing. Collapsing every repeated inscription removes 634 texts. Each measure was then recomputed on the reduced corpus, and separately at each major site, so that no finding rests on pooling two traditions that might differ.

Excess over controlAll 2,536Dedup.Mohenjo-daroHarappa
Sequential dependency1.3961.0330.9320.746
Order alone1.0730.9000.8370.692
Boundary closure-0.848-0.863-1.210-1.140
Directionality0.6870.6250.6050.554
Positional fixity0.2130.1820.1990.219
Grammatical categories (withdrawn, z)2.611.471.562.35
Table 7. Every headline measure across corpus variants. The struck row is the withdrawn result described below.

The effects persist after deduplication and hold independently at Mohenjo-daro and Harappa, two sites some six hundred kilometres apart, showing the same structure.

How many independent inscriptions are there really?

Collapsing only identical inscriptions is the weak version of this control. Indus seals were also copied, and near-identical texts are not independent observations either. Grouping inscriptions that differ by one or two sign edits reduces 1,902 distinct sequences to 1,141 and then to 558. Intervals computed on the raw count are therefore too narrow, which is a problem for most published statistics on this corpus, and was for ours.

Measure (excess, with z)1,902 exact 1,141 at edit 1558 at edit 2
Sequential dependency+1.033  71.5 +0.875  54.5 +0.665  34.0
Directionality+0.625  61.7 +0.561  45.5 +0.464  26.7
Positional fixity+0.182  35.1 +0.151  24.8 +0.125  18.0
Opening vs closing classes z 6.0z 4.6z 0.5
Table 8. Every headline measure recomputed on progressively stricter definitions of an independent inscription. The struck row is the one that does not survive.

The three main results hold on 558 inscriptions, under a quarter of what we started with. That is the number they should be judged on.

Is the headline number an artifact of the estimator?

This is the sharpest objection available and it has been made before, against earlier entropy work on this corpus. At 11,135 tokens across 591 sign types the table of sign pairs is 0.84 per cent occupied, 2,919 filled cells out of 349,281. In that regime the naive maximum-likelihood estimator overstates dependency badly, and a figure quoted without correction cannot be trusted.

Our figure is already a margin over a permutation control computed with the same estimator, which cancels most of that bias because the control shares the sample size, alphabet and marginal frequencies. It does not cancel it exactly. So the same quantity was recomputed under three published bias-corrected estimators, each applied to the observed data and to the control alike.

EstimatorObservedControl Excess95% interval
Maximum likelihood3.1692.136 +1.033+0.873 to +0.978
Miller-Madow2.9711.778 +1.192+1.014 to +1.124
Chao-Shen3.0011.566 +1.435+1.370 to +1.511
Chao-Wang-Jost2.5380.861 +1.677+1.495 to +1.742
Table 9. The same excess under four estimators. Intervals come from subsampling inscriptions without replacement, which is the correct resampling unit here: drawing them with replacement reintroduces the duplicate structure the analysis exists to control for.

Correction moves the result upward, not downward. The excess rises from 1.033 to 1.677 bits as the correction gets stronger, because the correction reduces the control far more than it reduces the observed data. The shuffled corpus is dominated by sign pairs seen once, which is exactly the situation these estimators exist to fix. The uncorrected figure we publish is therefore the conservative one, and every estimator tested puts the interval clear of zero.

A result we withdrew

An earlier pass found that signs sort into grammatical categories which then line up in order along the inscription. It was the most interesting result of the project and the closest thing to a grammar. It did not survive. Once duplicate impressions were collapsed the effect fell to chance, and it failed to reproduce consistently across sites. The duplicates had been manufacturing it.

It is recorded here because a finding that dies under its own control test is worth more to other researchers than one that was never tested, and because anyone running similar analyses on this corpus should expect the same trap.

What was left when we looked again

Withdrawing a result is not the same as showing there is nothing there. The original test asked one broad question (do a dozen induced categories arrange themselves in order?) and spent all its evidence on it. Three narrower tests were run afterwards on the deduplicated corpus, each able to fail independently.

Second-pass testResultAgainst
Cluster stability under resampling0.2660.049
Category model vs full sign model (bits, held-out)+0.0284t = 1.89
Opening vs closing sign contextsz = 6.02p < 0.001
Table 10. Second-pass tests, deduplicated corpus only.

The categories are more stable under resampling than chance, but not by much. Compressing 591 signs into a dozen categories costs nothing measurable on unseen text. The interval crosses zero, so the honest reading is no detectable loss, not an improvement.

The third test held up at first. The signs that open inscriptions and those that close them differ in the company they keep, and not merely in where they sit: the comparison uses each sign's neighbours with position information removed. Twenty-one openers and forty-eight closers form two distributionally distinct groups at z = 6.02.

It does not survive the harder duplicate control. Collapsing near-identical inscriptions as well as exact ones takes it to z = 4.6 at edit distance one, and to z = 0.5 at edit distance two, which is nothing at all. We are leaving it here as a reported result rather than a claim: it may be real and merely under-powered once the corpus is cut to genuinely independent inscriptions, or it may be another artifact of repeated seals. On this corpus we cannot tell.

VIIWhat This Settles

Supported
  • Sign order carries information, and more of it than in running prose
  • Inscriptions are complete, self-contained statements
  • The script is directional, with distinct opening and closing positions
  • Signs hold fixed positions rather than floating freely
  • The same system is in use at sites hundreds of kilometres apart
Not established
  • What any inscription says
  • What language, if any, underlies the script
  • The phonetic or semantic value of any sign
  • Whether the script is logographic, syllabic or mixed
  • Any connection to a known language family
  • Whether signs group into a full system of grammatical categories
    though opening and closing signs do form two distinct classes
This is not a decipherment

No one has deciphered the Indus script, and this work does not either. Claims to have read it appear regularly and are not supported by evidence of this kind; neither is anything here. Not one inscription is translated. One sign opens roughly a third of the corpus and we do not know what it means.

The strongest honest statement the data supports is this: the Indus script carries ordered, self-contained, positionally organised structure, and more dependency between successive signs than English prose. That is a real constraint, and the emblem and ownership-mark readings have to answer it. It is not the same as showing the script encodes a language, and we are not showing that.

VIIIThe Picture and the Object

Every Indus seal carries an image, most often a one-horned bull, sometimes an elephant, a rhinoceros, a tree or a human figure, set above the row of signs. The natural reading is that the signs caption the picture. If so, the two channels must share information, and the image would become a weak bilingual: a known concept standing beside an unknown word.

ChannelObservedControl Excessz
Iconographic motif0.29780.2814+0.01642.0
Object type0.20850.1437+0.064814.0
Material0.18890.1525+0.03647.1
Table 11. Information shared between the sign string and three non-textual channels, deduplicated corpus. Labels were permuted across objects to build each control.

The caption reading is not supported. Before deduplication the motif appears to carry substantial information about the signs; after repeated seal impressions are collapsed, almost all of it disappears. What remains does not clear the bar once the three channels tested here are accounted for. The picture and the text are, as far as this measurement goes, independent channels. Whatever the signs record, it is not a description of the animal above them.

The object itself is a different story. What a piece is, whether seal, tablet, tag or potsherd, genuinely conditions which signs appear on it, and survives deduplication comfortably. Text length tracks it too: seals average 5.06 signs, tablets 3.70, pottery 2.78. The corpus is not one uniform practice, and analyses that pool it are averaging over at least two.

Whether it could be a list of names

A standing objection to the whole enterprise is that these are ownership marks, personal names, lineages and titles, rather than text. That hypothesis makes a sharp prediction. A corpus of one-off identifiers never settles down: every new object introduces new signs, and the vocabulary grows about as fast as the corpus itself. Reused language does the opposite, because a finite lexicon is being drawn on repeatedly.

CorpusVocabulary growthReading
Indus script0.488finite vocabulary, reused
English prose (matched)0.415finite vocabulary, reused
Indus, seals only0.506
Indus, tablets only0.581
Table 12. Heaps' law exponent. Near 0.5, a finite vocabulary is being reused; approaching 1.0, it is not.

Indus sits at 0.488, against 0.415 for matched English prose. Vocabulary growth is slightly faster, consistent with more one-off signs, but nowhere near the behaviour of a list of unique identifiers. The sign inventory is genuinely reused. Whatever the signs are, they are drawn from a working vocabulary rather than minted per object, and the pure name-list reading has to account for that.

Whether they are accounts

The third standing reading is that these are administrative records: so many measures of grain, so many head of cattle. That is what the neighbouring systems are. It also makes a prediction that can be tested without reading a single sign, because two of the comparison corpora label their numerals explicitly, and a numeral does not behave like a word.

In proto-cuneiform and proto-Elamite, numerals account for 44% and 42% of all sign tokens respectively. They cluster in one region of the text. Above all they travel in company: where a numeral appears at all, another appears in the same short text 32% of the time, because an account that records one quantity almost always records a second. That combination is a signature, and it can be learned where the answer is known and then hunted where it is not.

A third corpus was added later and agrees. Linear A, undeciphered but certainly writing, keeps its numerals in a separate Unicode block so they identify themselves without any interpretation; its numerals co-occur at 71%, higher than either Mesopotamian system. Three independent corpora, two of them undeciphered, all behave the same way.

Nothing in the Indus corpus matches it. Scoring every sufficiently common sign against the learned profile, the closest candidates reach a co-occurrence rate between 2% and 7%, against roughly 32% for genuine numerals. No sign in this corpus behaves like a number. Either these inscriptions do not record quantities at all, which would sit comfortably with objects used to mark ownership rather than to tally goods, or they record them in a way that shares no statistical property with how every neighbouring bureaucracy did it. The result is negative, and it closes off the one anchor that decipherment of a related script has historically depended on. Proto-Elamite numerals were read long before anything else in proto-Elamite was. That door is not open here.

IXThe Hand and the Stone

Everything so far has treated an inscription as a string of symbols. It is also a physical act: a person cutting into a piece of steatite about the size of a postage stamp, who had to decide, before the first cut, how much would fit. The source database records width, height and thickness for 2,325 seals, and those measurements can be set against the texts they carry.

Longer inscriptions sit on larger objects. Across 1,607 pieces with recorded dimensions, text length and seal width move together at rho 0.445. That number on its own means little, because it could be nothing more than a cataloguing convention: if long texts belong on big copper tablets and short ones on small clay tags, the correlation is a fact about object classes rather than about people.

TestnCorrelation z
Object type and material held fixed
lengths permuted within class
1,6070.44515.4
Face area rather than width
the whole writing surface
1,6070.41113.1
Steatite seals alone
one object, one material
1,0300.40613.0
After collapsing near-duplicates
edit distance two
4810.3988.6
Table 13. Text length against seal width. The stratified control permutes lengths only within object type and material, so every class convention is preserved in the null and only the pairing within a class is destroyed. Steatite seals alone hold object, material and workshop tradition fixed simultaneously.

It survives all of it. Holding object type and material fixed, the effect is unchanged. Restricted to steatite seals alone, one object in one material, it is unchanged again. It holds separately at Mohenjo-daro and at Harappa, and it is still present after near-identical inscriptions are collapsed at edit distance two, which is the gate that killed five earlier findings in this project. The one place it is absent is clay, where the correlation is 0.038 on 63 pieces.

Whether the text was planned before the first cut

Two quite different craftsmen produce that same correlation. One knows the text before he starts and picks a blank large enough to hold it, so the signs stay the same size and the stone grows with the message. The other starts carving on whatever is to hand and squeezes as he runs out of room, so the stone stays the same and the signs shrink. Only the first is planning, and the two can be separated by asking how much of a longer text is absorbed by a larger surface.

Writing the width of the surface against the number of signs as a power law, the exponent answers the question directly. An exponent of one means every extra sign was paid for with proportionally more stone. An exponent of zero means none of it was, and every extra sign was paid for by cutting smaller.

Corpus cutnExponent 95% intervalz
Steatite seals1,0300.252[0.224, 0.282]12.4
Tablets3220.376[0.306, 0.450]7.2
All objects1,6070.312[0.287, 0.341]17.4
Near-duplicates collapsed, distance one9800.372[0.335, 0.409]12.9
Near-duplicates collapsed, distance two4810.457[0.397, 0.524]8.8
Table 14. Allometric exponent of writing surface against text length. Intervals by subsampling without replacement. An earlier version of this measurement used signs per millimetre against length, which is not a valid statistic here: length appears on both sides of it, so it is positive by construction whatever the carvers did.

The answer is both, and mostly the second. On steatite seals the exponent is 0.25, rising to 0.46 once near-duplicates are collapsed. Somewhere between a quarter and a half of a longer text was accommodated by choosing a bigger seal; the remainder was accommodated by carving the signs smaller. A carver with ten signs to fit did not simply reach for a stone twice the size. He reached for one slightly larger and cut finer.

That is a modest observation and it is worth being clear about what it does and does not support. It does not show the signs are language. What it does show is that the length of the message was known, at least approximately, before the carving began, because the surface was chosen with reference to it. A sequence produced without foresight, one mark at a time until the maker stopped, has no reason to correlate with the size of the object it is on. This is the first result in this work that comes from the objects rather than from the symbols, and it is independent of every information-theoretic argument above.

A result that did not survive

The same measurement initially separated the two great cities: Mohenjo-daro at 0.188 and Harappa at 0.476, with non-overlapping intervals, which would have meant the two capitals had measurably different workshop habits. It does not hold. The excavated assemblages differ sharply, 290 tablets and 224 seals at one site against 781 seals and 47 tags at the other, and once object type is held fixed the separation weakens; once the width and length distributions are matched as well, the gap is -0.033 and the significance is gone entirely. What looked like a difference between two cities was a difference between two excavations. It is recorded here rather than deleted, because the reason it failed is more useful than the result would have been.

XThe Frame

This page has said since its first version that the opening slot of an Indus inscription is sharply restricted. That is a statement about one position. The question never asked was whether the other positions are restricted too, and the answer changes what kind of object an inscription is.

The difficulty is that position has to be measured carefully. Inscriptions here average under five signs, so if position is expressed as a fraction of the way through a text, a sign that happens to occur in short inscriptions is handed a middle position by the arithmetic with no scribe involved. The test therefore has to be run inside a single inscription length, where every slot exists and every slot is equally available, against a control that shuffles each inscription's own signs and so holds length, membership and every sign frequency exactly fixed.

Corpus4 signs5 signs 6 signs7 signs Bound to a slot
Indus21/2729/3424/2719/2682%
Indus, near-duplicates collapsed7/1223/3122/2620/2676%
Linear B14/3417/396/347/2533%
Ur III Sumerian22/2923/2720/2614/2177%
Proto-Elamite30/3531/4626/3213/1678%
Proto-cuneiform24/2929/4023/3215/2672%
Table 15. Signs bound to a particular slot, out of signs common enough to test, within inscriptions of exactly that length. Benjamini-Hochberg correction applied at each length. All corpora cut to the Indus token count and length distribution.

Indus inscriptions are strongly slotted: 82% of testable signs are bound to a position, and 71% survive after near-identical inscriptions are collapsed. This is not the opening restriction in another guise. Delete the first sign of every inscription and sweep again, and the anchors are still there.

The frame has named parts

What makes this a frame rather than a statistic is that individual signs hold specific slots, and hold them across every inscription length. The demonstration is in the two columns below. As an inscription grows from four signs to seven, a sign anchored to the front keeps its position counted from the front while its position counted from the back slides; a sign anchored to the end does the exact reverse. No artifact of length or of measurement produces two groups of signs moving in opposite directions.

SignSlot counted from the front Slot counted from the backAnchored to Share of its uses
sign 7401, 1, 1, 13, 4, 5, 6first slot68%
sign 23, 4, 5, 61, 1, 1, 1penultimate slot57%
sign 603, 4, 5, 61, 1, 1, 1penultimate slot83%
sign 8204, 5, 6, 70, 0, 0, 0last slot77%
sign 5201, 1, 1, 13, 4, 5, 6first slot82%
sign 7413, 4, 52, 2, 22 from the end59%
sign 8174, 5, 60, 0, 0last slot90%
sign 8615, 6, 70, 0, 0last slot70%
Table 16. The strongest slot-bound Indus signs at inscription lengths four through seven, deduplicated corpus. The front-anchored and back-anchored groups move in opposite directions as the inscription lengthens, which is the internal control.

So the corpus has a first-position set, a last-position set and a penultimate-position set, each occupied by particular signs that stay in their part of the inscription whatever its length. An Indus inscription behaves less like a sentence and more like a filled form. That is a structural claim, not a reading, but it is a useful one: it says where to look. If one slot admits only a handful of signs, those signs are a field type, and field types are the first thing anyone reads in a bureaucratic script.

The control this panel was missing

Every comparison so far has set Indus against scripts that can be read, which invites the reply that readability itself is doing the work. There is one script that answers it. Linear A is undeciphered, its language unidentified, and nobody doubts it is writing, because Linear B was adapted from it sign by sign and Linear B is Greek. It is the only available example of an unreadable but unquestionably real writing system, and it belongs beside Indus more than anything else in this panel does.

CorpusStatusSigns tested Bound to a slot
Indusundeciphered, unknown22861%
Linear Aundeciphered, certainly writing2326%
Linear Bdeciphered, Greek24523%
Proto-cuneiformattested non-language notation27641%
Table 17. Slot binding measured at each corpus's own well-populated inscription lengths, since Linear A tablets run about three times longer than Indus inscriptions. Absolute rates depend on how often a sign must occur to be testable; the ORDERING does not, and the ordering is what is being read here.

Linear A sits with Linear B, and Indus sits above both. The two Aegean scripts agree with each other, which is what their shared ancestry predicts. The Linear A sample is small, 23 signs testable against 228 for Indus, so this is a difference of degree rather than a settled quantity. It is also, as the next section shows, a comparison between the wrong kinds of object.

The comparison that was on our own disk

Every corpus named so far is an archive of tablets. Indus is a corpus of seals. This work has itself shown that the physical object shapes the text, so comparing seals against tablet archives and concluding something about the script was a mistake, and it was ours.

The correction required no new acquisition. The CDLI catalogue, held here since the beginning of this work, records object type, and 17,024 of its records are seals rather than tablets. Of those, 6,365 are Ur III, which is 2100 to 2000 BC and therefore contemporary with the mature Indus period. They are the same kind of object, from the same centuries, serving the same administrative purpose. And they can be read.

What they say is a formula: personal name, son of personal name, servant of a god or king, often with a title. That is precisely what Indus seals have been supposed to carry for a century, which makes this the one place where the frame machinery can be checked against a known answer. The prediction was written down before the test ran.

TokenMeaningPredicted position Position foundShare of its uses
dumuson ofmiddlemiddle57%
arad2servant oflatepenultimate slot75%
arad2-zuyour servantlastlast slot98%
dub-sarscribe, a titleearlyslot 279%
ensi2governor, a titleearlyslot 270%
lugalkinglatemiddle45%
Table 18. The slot-binding sweep run on Ur III seal legends, which are readable. It was given no information about Sumerian. Predictions were fixed in advance from the published formula.

It recovers the formula unprompted. Personal names take the first slot, titles take the second, the word for "son of" sits in the middle, and "your servant" takes the last slot in 98 per cent of its appearances. The machinery that found the Indus frame finds the correct frame in a language it was told nothing about. That is the validation the Indus result needed, and it passes.

One prediction is only half met, and it is left in the table rather than dropped. Lugal, "king", was predicted late and comes out middling. The reason is visible once the legends are read: it appears both as the final royal name and inside compound names earlier in the line, so it genuinely occupies two roles. A test that scored five out of five would be more suspicious than one that scores four.

Then the comparison itself, with both corpora put through the identical near-duplicate collapse, because Ur III seals repeat heavily for the same reason Indus seals do: one official's seal was rolled onto many tablets.

CorpusSigns tested Bound to a slot
Ur III seal legends, words6197%
Ur III seal legends, signs14275%
Indus inscriptions9576%
Table 19. Slot binding on seals rather than tablets, every corpus deduplicated identically. Sumerian is written as sign sequences within words and words within lines, and nobody knows which of those an Indus sign corresponds to, so both granularities are given and the answer is bracketed rather than assumed.

Indus inscriptions are slotted to the same degree as contemporary readable seal legends: 76% against 75%. Not more, not less. Set beside tablet archives Indus looked like an outlier; set beside the same object from the same centuries it is entirely ordinary, and the earlier reading on this page has been corrected accordingly.

This does not say Indus seals carry names and titles. It says that if they did, they would look like this, and that nothing in their positional structure argues against it. That is the closest structural analogue this work has found for what an Indus inscription might be, and the hypothesis it favours is the oldest and least exciting one in the field.

What this does not mean, which matters more than what it does

The obvious temptation is to read a strongly slotted script as an un-language, and the panel forbids it. Ur III Sumerian is a fully deciphered language and it scores 77% on this measurement, essentially the same as Indus. Linear B is also a fully deciphered language and it scores 33%, less than half as much. Two readable languages sit at opposite ends of the scale.

Slot binding therefore does not separate writing from non-writing at all. It separates formulaic text from free text, and it puts Indus with administrative Sumerian rather than with the more varied Linear B tablets. That is a real finding about what these objects are for. It is not evidence either way about whether the signs encode speech, and anyone quoting the Indus figure without the Ur III figure beside it is misusing it.

XIHow Close Is This To A Decipherment

Reading an unknown script is not one problem but a sequence of them, and the later ones are far harder than the earlier ones. Setting them out in order makes it possible to say where this work actually sits, and where the field sits, without either being flattered.

Machine-readable corpus2,536 inscriptions in hand. The fuller ICIT corpus is held behind a permission request.
field: done
Corpus is not randomLong established in the literature, and reconfirmed here against shuffled controls.
field: done
Ordered, directional, closed unitsSign order carries information, and inscriptions do not run on across objects.
field: done
Not a pure emblem systemModelled ownership marks score negative on sequential dependency where Indus scores above English prose. The non-linguistic case still has serious advocates.
field: contested
Vocabulary is reused, not minted per objectHeaps exponent 0.488 against 0.415 for matched English. Argues against a pure list of names.
field: open
~
Sign classes and grammarPositional classes are established: particular signs hold the first, penultimate and final slots across every inscription length, and survive deduplication. The twelve-category grammar and the opener-against-closer split were both withdrawn. A positional frame is not a grammar, and nobody has demonstrated a grammar.
field: contested
Word or morpheme segmentationNo agreed way to divide Indus strings into units exists.
field: contested
Language family identifiedDravidian, Indo-Aryan and isolate all have advocates. None is established.
field: contested
Phonetic values for signsA full Sanskrit mapping has been claimed and is not independently confirmed. Auditing it is open work.
field: claimed
Read an arbitrary unseen inscriptionThis is the actual bar for decipherment. Nobody is here.
field: not achieved
Survives independent verificationWould require everything above to hold up under outside scrutiny.
field: not achieved

5 of 11 rungs cleared here, one partial, which is roughly 50% of the way up. That figure should be read with care. The rungs are not equal in size, and the remaining ones are much larger than the ones already climbed.

The honest position

Nobody has verifiably cleared rung nine or beyond, ourselves included. Rung nine has been claimed more than once and never independently confirmed. The last three rungs are what the word decipherment actually refers to, and no one is standing on them.

What this work adds is confined to rungs four and five, and to narrowing rung six. That is a genuine contribution to a hundred-year-old problem and it is also a long way from reading the script. Both halves of that sentence matter.

XIIWhat Is New Here, And What Is Not

A structural analysis of this corpus published in 2026 by Ashish Nair, How Non-Linguistic Is the Indus Sign System? A Synthetic-Baseline Scorecard (arXiv:2604.17828), works from the same underlying digitisation and computes several of the same quantities. Anything here has to be read against it, so the overlap is set out first rather than left for a reader to find.

Independently replicated

These were computed here before that paper was read, from different sign identifiers and separate code. The agreement is close enough to be worth stating plainly.

QuantityNair 2026This work
Zipf rank-frequency slope-1.492-1.489
Zipf fit R squared0.9560.957
Hapax rate33.2%33.7%
Mean inscription length4.424.39
Exact duplicate rate24%25%
Table 20. Independent agreement on shared quantities.

Extends existing work

The duplicate-structure problem was identified in that paper at the level of exact repeats. Here it is carried further: inscriptions differing by one or two sign edits are also not independent observations, which takes 1,902 distinct sequences down to 1,141 and then 558, and every headline measure is reported at all three levels. Separately, the bias-correction result in section IV appears to run against expectation, since correcting a null-referenced statistic increases the margin rather than shrinking it.

Not previously attempted, as far as we can establish

No dispersion measure appears to have been applied to this corpus. We applied one, asking whether signs divide into a frequent evenly-spread class and a rare bursty class, which is how function words separate from content words in every natural language. They do not. The apparent split is 94 per cent explained by frequency alone, and once each sign is compared against a null matched to its own frequency, not one of 153 signs survives correction. This is a negative result and it is the clearest thing in this work that nobody had checked.

Withdrawn

Five results looked strong and did not survive their own controls: a twelve-category induced grammar, a claim that those categories recovered most of the script's predictability, a correlation between the picture on a seal and the signs above it, the closed-class split above, and a distinction between opening and closing sign classes. Each is described where it arose. None is claimed.

The weakness we have not fixed

The comparison against English prose in section III uses continuous text cut into short segments. It is matched on text count, length distribution, token count and vocabulary size, but it is not inscriptional writing and it has no genuine text beginnings or endings. Until the same battery is run against a known script of comparable genre and identical sample size, such as Linear B tablet lines, the claim that Indus exceeds real writing on sequential dependency should be read as provisional.

This is the same gap as in the prior work, which uses no linguistic control at all. It is the next thing being built here, and it is the test on which this argument should be judged.

XIIIMethods

The implementation is not published. The protocol is, in enough detail to be attacked.

Corpus

2,536 inscriptions from the published digitisation of the Corpus of Indus Seals and Inscriptions: 11,135 sign tokens, 591 sign types, mean length 4.39, ten sites. Analysis runs on the deduplicated set unless stated. Object type, material and iconographic motif are carried as covariates. No reconstructed or generated text is used anywhere.

Statistics and controls

Every quantity is reported as a margin over a matched null, never as a raw value. Two nulls are used: a global shuffle preserving sign frequencies and length distribution, and a within-inscription shuffle preserving which signs co-occur while destroying their order. The difference between the two isolates ordering from co-occurrence. Nulls run at 1,000 to 5,000 iterations. Significance is a z-score against the null distribution.

Estimator bias

At 0.84 per cent occupancy of the sign-pair table, maximum-likelihood entropy is strongly upward-biased. Referencing every statistic to a null computed with the same estimator removes that bias to first order. Because it does not remove it exactly, the principal result is also reported under Miller-Madow, Chao-Shen and Chao-Wang-Jost corrections.

Independence

Seals were copied, so inscriptions are not independent draws. Results are recomputed after collapsing exact duplicates and then near-duplicates at edit distance one and two, and separately at Mohenjo-daro and Harappa. Intervals come from subsampling inscriptions without replacement; resampling with replacement reintroduces the duplicate structure the analysis exists to control for, and produces intervals that do not bracket their own estimate.

Comparison corpora

Controls are matched to the Indus corpus on text count, length distribution, token count and vocabulary size, then run through the identical procedure. Four arms are used: natural language, a modelled administrative-record system, a modelled ownership-mark system, and randomly ordered signs. The random arm returning approximately zero on every measure is the calibration check.

Discipline

Findings are put to an adversarial review instructed to refute rather than confirm them, and any result that fails deduplication, per-site replication or a matched null is withdrawn and recorded rather than removed. Where a hypothesis is generated and tested automatically, it is registered with its predicted direction before the test runs, and significance is judged against a false-discovery threshold recomputed over every test performed.

Availability

Corpus provenance is public and named above, so the inputs are independently obtainable. Researchers wishing to compare results against their own measurements, or to see the figures behind any table here, are welcome to make contact.