INDUSTHE UNREAD ARCHIVERESEARCH EDITION / 01

04 / THE INDUS ARCHIVE

Evidence has levels.
So should our claims.

INDUS is an evidence-first computational research project. It is not a complete or partial linguistic decipherment.

The evidence ladder

Six categories of claim, not six steps toward certainty. Each requires a different kind of validation.

01

Directly observed

Counts, sign order, repeated strings, recorded metadata.

These are observations in the secondary transcription, subject to source errors.
AUDITABLE
02

Statistically reproduced

Held-out sign prediction, direction preference, permutation controls.

These tests support combinatorial structure. They do not test literal meaning.
AUDITABLE
03

Externally anchored

Published numeral-like and measure-associated sign hypotheses.

Metrological values are proposed calibrations, not established translations.
CONDITIONAL
04

Functional hypothesis

Opening, quantity-like, measure-associated, core, and closing slots.

The grammar includes researcher assumptions and positional priors.
CONDITIONAL
05

Contextual inference

An occurrence-level role vector and 90% credible role set for every token.

100% annotated describes software coverage. It does not mean 100% deciphered.
CONDITIONAL
06

Phonetic or language speculation

Pronunciation, personal names, exact readings, and underlying language.

Unresolved. This project does not establish these claims.
UNRESOLVED

From transcription to structure

01 / Normalize and separate

Only complete, direction-known horizontal records enter the analysis. IDs 000 and 999 encode damage or gaps and are removed. R/L strings are reversed into reading order; L/R strings are retained. Multi-line records are split.

Exact strings are deduplicated before fitting and held-out testing. Descriptive metadata and contextual coverage still include all analytic lines; the interface labels each denominator.

02 / Test and annotate

Five shuffled folds test hidden-sign recovery and direction preference on unseen unique strings. Unseen vocabulary is excluded and counted. Permutation controls preserve appropriate counts while disrupting order.

The exploratory ACGI-5 model uses positional and contextual features, a neighbor graph, and a soft five-slot order. Twenty-four bootstrap refits assess stability. Model mass is uncalibrated; high mass cannot substitute for archaeological evidence.

Full methods ↗ · Allowed and forbidden claims ↗ · Research note ↗

The source is part of the evidence.

The corpus is a secondary public transcription by Tanishk Tiwari / horus84, pinned to commit ac44f08c835de3b709875441eea9d60d1811e3cf. Its MIT notice is preserved. It is not a substitute for artifact photographs or a primary concordance.

SHA-256: c368f290d8d784093d804243100de27ae6ad09df3ae37cab7432a0205d4ec9ef

The repository contains no authentic sign glyph assets. Every tile is an ID representation. No pseudo-Indus symbols, institutional affiliations, or endorsements have been invented.

References reproduced from the repository’s source bibliography; no new archaeological claims are inferred from these links.

What could change the conclusion?

Repeated texts and frequency imbalance inflate apparent coverage. Sign normalization, allograph choices, and damaged records affect the inventory. Missing motif data limits contextual tests. Site and artifact fields are source reports, not independently verified archaeological labels.

How to falsify this model

  1. Test prospective, newly transcribed inscriptions: does context still outperform frequency?
  2. Merge or split plausible allographs: do sign roles and ordering results survive?
  3. Remove anchors and positional priors: do the functional role hypotheses remain stable?
  4. Compare predictions against independently labelled object functions and excavation contexts.
  5. Test proposed metrological calibrations against independently measured archaeological objects.
Read the research agenda ↗

A small glossary

Analytic line
A complete, direction-known inscription line after source normalization.
Unique string
One normalized sign sequence, counted once regardless of repetition.
Model mass
A normalized score inside the five-role model; not a calibrated probability of correctness.
Bootstrap stability
How consistently a role assignment recurs across 24 resampled model fits.
90% credible role set
Occurrence-level alternatives accumulating at least 90% model mass. It is conditional on model assumptions.
Gold / Silver
Strict and wider stable-role thresholds, with externally proposed anchors retained. Literal meanings remain unresolved.
Contextual only
A sign outside stable tiers; token roles are inferred from the model and occurrence context.
Unresolved
No secure literal meaning, pronunciation, or underlying language has been recovered.
Open the research workbench ↗