INDUS / ARCHIVE
INDUS / ARCHIVE
03 / THE INDUS ARCHIVE
Imagine finding an administrative form in a language you cannot read. You might recognize a repeated opening, a number field, or a closing mark before you understand a single word.
01 / OBSERVED PATTERN
Some signs prefer the beginning; others favor the end. Position tells us something about use, without telling us pronunciation.
86.72% initial
128 non-solo observations
86.67% terminal
45 non-solo observations
02 / STATISTICALLY REPRODUCED
Numeral-like signs occur before a conservative set of measure-associated signs unusually often. The association is observed; the numerical and metrological interpretations remain external hypotheses.
A separate, train-only bigram model prefers normalized reading order in 79.00% of eligible unseen strings, averaged across folds.
03 / HELD-OUT PREDICTION
Cover one sign in a previously unseen string. Guessing from nearby signs works better than always choosing the most common sign.
| Fold | Eligible tokens | Out of vocabulary |
|---|---|---|
| 1 | 1826 | 40 |
| 2 | 1786 | 61 |
| 3 | 1723 | 58 |
| 4 | 1787 | 46 |
| 5 | 1942 | 42 |
The model recovers constraints in a sequence. It has no verified dictionary, and this task does not test whether it understands the signs.
04 / THE LIMIT OF COVERAGE
Giving every field a tentative label makes a useful index. It does not make the form readable. Rare signs often inherit their annotation from context and the five-role grammar’s assumptions.
Exact duplicate-line share is 35.32%. 284 sequences recur across artifacts; 77 across sites. Repetition is evidence of recurrence, not independent confirmation of meaning.
All displayed metrics are read from checked-in research outputs. Trace the claims and learn how to falsify them ↗