Directly observed
Counts, sign order, repeated strings, recorded metadata.
These are observations in the secondary transcription, subject to source errors.04 / THE INDUS ARCHIVE
INDUS is an evidence-first computational research project. It is not a complete or partial linguistic decipherment.
Six categories of claim, not six steps toward certainty. Each requires a different kind of validation.
Counts, sign order, repeated strings, recorded metadata.
These are observations in the secondary transcription, subject to source errors.Held-out sign prediction, direction preference, permutation controls.
These tests support combinatorial structure. They do not test literal meaning.Published numeral-like and measure-associated sign hypotheses.
Metrological values are proposed calibrations, not established translations.Opening, quantity-like, measure-associated, core, and closing slots.
The grammar includes researcher assumptions and positional priors.An occurrence-level role vector and 90% credible role set for every token.
100% annotated describes software coverage. It does not mean 100% deciphered.Pronunciation, personal names, exact readings, and underlying language.
Unresolved. This project does not establish these claims.Only complete, direction-known horizontal records enter the analysis. IDs 000 and 999 encode damage or gaps and are removed. R/L strings are reversed into reading order; L/R strings are retained. Multi-line records are split.
Exact strings are deduplicated before fitting and held-out testing. Descriptive metadata and contextual coverage still include all analytic lines; the interface labels each denominator.
Five shuffled folds test hidden-sign recovery and direction preference on unseen unique strings. Unseen vocabulary is excluded and counted. Permutation controls preserve appropriate counts while disrupting order.
The exploratory ACGI-5 model uses positional and contextual features, a neighbor graph, and a soft five-slot order. Twenty-four bootstrap refits assess stability. Model mass is uncalibrated; high mass cannot substitute for archaeological evidence.
Full methods ↗ · Allowed and forbidden claims ↗ · Research note ↗
The corpus is a secondary public transcription by Tanishk Tiwari / horus84, pinned to commit ac44f08c835de3b709875441eea9d60d1811e3cf. Its MIT notice is preserved. It is not a substitute for artifact photographs or a primary concordance.
SHA-256: c368f290d8d784093d804243100de27ae6ad09df3ae37cab7432a0205d4ec9ef
The repository contains no authentic sign glyph assets. Every tile is an ID representation. No pseudo-Indus symbols, institutional affiliations, or endorsements have been invented.
References reproduced from the repository’s source bibliography; no new archaeological claims are inferred from these links.
Repeated texts and frequency imbalance inflate apparent coverage. Sign normalization, allograph choices, and damaged records affect the inventory. Missing motif data limits contextual tests. Site and artifact fields are source reports, not independently verified archaeological labels.