Indus Valley Script · c. 2600–1900 BCE

The Indus script is not a set of emblems.
What it is remains open.

The Indus Valley script is the last major undeciphered writing system of the ancient world. It has resisted a century of attempts because the inscriptions are short, no bilingual key exists, and the underlying language is unknown. We cannot read it, and neither can anyone else. What we can do is measure its structure against systems whose nature we already know — and it carries more sequential dependency than any of them, including running natural language.

2536 inscriptions · 11,135 sign tokens · 591 distinct signs · published 26 July 2026

01The corpus

Every number on this page comes from real excavated inscriptions — the published digitisation of the Corpus of Indus Seals and Inscriptions. No reconstructed, simulated or generated text is used anywhere in this analysis.

Inscriptions
2536
1902 after collapsing duplicate impressions
Sign tokens
11,135
mean 4.39 signs per inscription
Distinct signs
591
34% appear only once
Sites
10
Mohenjo-daro, Harappa, Lothal, Dholavira, Kalibangan and others

02What we found

Each finding below is stated against a control — the same corpus with its structure removed. That comparison matters more than any raw figure: at this corpus size, measures of structure return large positive values even for text with no structure at all, so only the margin over the control means anything. Every figure quoted is that margin.

1. Sign order carries information

The sequence of signs in an Indus inscription is not arbitrary. Adjacent signs share 1.033 bits more information than the same signs re-dealt at random, and 0.900 bits more than the same inscription with only its order scrambled — so the effect is in the ordering itself, not merely in which signs co-occur.

excess over random control1.033 bits
excess over order-only control0.900 bits
significancez = 71.5, p < 0.001

2. The inscription is a closed unit

Whatever an inscription says, it finishes saying it. Sign pairs that straddle two different objects share less information than chance (0.863 bits below control). Meaning does not run on from one artifact to the next: each object carries one complete, self-contained statement. This also rules out the corpus being fragments of a single longer running text.

deficit vs control-0.863 bits
significancez = -18.9
readinginscriptions are independent

3. Inscriptions are strongly directional

Signs that open an inscription and signs that close it are drawn from almost disjoint sets — a separation of 0.831 where the control gives 0.206. One sign opens roughly a third of all inscriptions; another sits in final position almost every time it appears. This is a real and very strong effect. It is not, on its own, evidence of language: ownership marks and administrative records produce the same asymmetry, and in our tests slightly more of it. It establishes that inscriptions are ordered and read in a fixed direction, nothing further.

observed separation0.831
control0.206
significancez = 61.7, p 0.0033
discriminates language?no — see section 04

4. Sign position is fixed, not free

Individual signs keep to their own regions of an inscription rather than floating freely. Positional fixity runs 0.182 above the control, and this holds at every site tested. As with directionality, this shows the script has assigned positions but does not by itself separate writing from record-keeping.

excess over control0.182
significancez = 35.1, p 0.0033
replicates atMohenjo-daro and Harappa separately
discriminates language?no — see section 04

03What survives scrutiny

The findings were attacked before they were published. The most serious objection was that duplicate seal impressions — the same seal pressed many times — could manufacture every one of these results. Collapsing all duplicate inscriptions removes 634 texts. The findings were then recomputed on the reduced corpus, and separately on each major site on its own, to check that nothing depends on pooling two different traditions.

measure (excess over control)all 2,536deduplicatedMohenjo-daroHarappa
sequential structure (bits)1.3961.0330.9320.746
order alone (bits)1.0730.9000.8370.692
closed-unit boundary (bits)-0.848-0.863-1.210-1.140
directionality0.6870.6250.6050.554
positional fixity0.2130.1820.1990.219

The effects persist after deduplication and hold independently at both Mohenjo-daro and Harappa. Two sites, separated by six hundred kilometres, show the same structure.

04Measured against systems we already understand

A structural finding means nothing until you know what other systems score. We therefore ran the identical analysis on four control corpora, each matched to the Indus corpus on every property that drives these numbers — same number of texts, same length distribution, same token count, same vocabulary size. Without that matching you measure corpus size rather than writing system.

corpussequential
dependency
boundary
closure
positional
fixity
direction­ality
Indus script
the real corpus
+1.033-0.863+0.181+0.624
Natural language
real text, matched
+0.579+0.304+0.000-0.008
Administrative records
modelled
+0.537-1.965+0.354+0.758
Ownership marks
modelled
-0.468-1.433+0.221+0.766
Random signs
structure floor
-0.007+0.083-0.007-0.006

Read the bottom row first. Randomly ordered signs score zero on everything, which is the check that the method is calibrated and not manufacturing structure.

What separates Indus

Sequential dependency. Indus scores higher than every control including natural language, while the emblem model scores negative — ownership marks carry no information in their ordering. Whatever Indus is, the order of its signs matters more than it does in running prose.

What does not separate it

Directionality and positional fixity. Both are strong in Indus — and stronger still in the emblem and administrative models. These measures cannot tell writing from record-keeping, and any argument resting on them alone is unsafe. Including, until this test, ours.

Where that leaves the question

Indus does not sit cleanly with natural language, and it does not sit cleanly with the non-linguistic models either. It is high on sequential dependency like a language, discrete at its boundaries like an administrative record, and positionally rigid like both. On the profile as a whole it is intermediate — which is also, independently, the conclusion reached by a separate 2026 analysis working from the same corpus by different means.

The pure emblem hypothesis is the one that struggles: it predicts no sequential dependency, and Indus has more than English prose.

Known limitation, stated plainly. The natural language control is running prose cut into short segments, so it has no genuine text beginnings or endings. That makes it a poor comparator for directionality specifically, and it is why we do not claim Indus's directionality is language-like — we claim only that it is real. A control built from genuinely short complete texts is the obvious next step and we have not done it yet.

05What this does and does not settle

Supported
  • Sign order is meaningful, not decorative
  • Inscriptions are complete, self-contained statements
  • The script is directional, with distinct opening and closing positions
  • Signs hold fixed positions rather than floating freely
  • The same system is in use across distant sites
Not established
  • What any inscription says
  • What language, if any, underlies the script
  • The phonetic value of any sign
  • Whether the script is logographic, syllabic or mixed
  • Any connection to a known language family
  • Whether signs group into grammatical categories — see below

A result we withdrew

An earlier pass found that signs sort into grammatical categories which line up in order along the inscription — the most exciting result of the project, and the one closest to a grammar. It did not survive. Once duplicate seal impressions were collapsed, the effect fell to chance, and it failed to reproduce consistently across sites.

The duplicates had been manufacturing it. We are reporting it here because a finding that dies under its own control test is worth more to other researchers than a finding that was never tested — and because anyone running similar analyses on this corpus should expect the same trap.

This is not a decipherment

No one has deciphered the Indus script, and this work does not either. Claims to have read it appear regularly; they are not supported by evidence of this kind, and neither is anything here. What is offered is narrower and, we think, more useful: a measurement of the script's structure that any competing account now has to be consistent with.

The strongest honest statement the data supports is this: the Indus script carries ordered, self-contained, positionally organised structure, and it carries more dependency between successive signs than English prose does. That is a real constraint, and the emblem and ownership-mark readings have to answer it. It is not the same as showing the script encodes a language, and we are not showing that.

06Method

The analysis pipeline is not published. What is published is every result it produced, the control each was measured against, the significance of each margin, and the robustness checks each finding survived — which is what any reader needs in order to hold the claims to account. Findings, not machinery.

Corpus provenance is public and standard, so the inputs are independently obtainable. Researchers wishing to compare results against their own measurements are welcome to make contact.