HarappaWorld is the one GEDmatch admixture project with a single calculator, and the focus shows. Zack Ajmal started the Harappa Ancestry Project in 2011 to do for South Asian ancestry what Dodecad was doing for West Eurasia — collect volunteer kits, build reference panels, and fit genomes against them with ADMIXTURE-based tooling. The calculator that resulted, usually just called HarappaWorld, has sixteen components and is still the default recommendation for anyone with South or Central Asian ancestry on GEDmatch.
The family-wide rules apply — an optimiser over fixed profiles, no test and no error bars — so this post is about what the components mean and what has been learned since 2012.
The sixteen components#
S-Indian · Baloch · Caucasian · NE-Euro · SE-Asian · Siberian · NE-Asian · Papuan · American · Beringian · Mediterranean · SW-Asian · San · E-African · Pygmy · W-African.
For South Asian readers, the first two carry most of the story:
S-Indian peaks in southern Indian references and tracks the ancestry the literature now calls AASI-related (Ancient Ancestral South Indian) — the deep indigenous lineage of the subcontinent. No ancient AASI genome had been sequenced in 2012; the component is the calculator's modern-sample shadow of it.
Baloch peaks in the Baloch and Brahui of Balochistan and is nearly the same statistical object as Dodecad's Gedrosia. It tracks Iranian-plateau-related ancestry — the western Eurasian stream that entered South Asia with (and before) food production. The pairing of S-Indian with Baloch is HarappaWorld's rendering of the cline later formalised as ASI–ANI, and the 2019 ancient-DNA work on the Indus periphery gave that cline its historical anchors: Iranian-plateau-related plus AASI ancestry, with steppe-related ancestry layered on later — carried disproportionately on the NE-Euro side.
NE-Euro peaks in northeast European references and, in South Asian profiles, is the visible edge of that steppe-related layer. Caucasian, as in every calculator of this generation, blends what ancient DNA later split into Caucasus and Iranian-related components. The remaining twelve components give the calculator its global reach — East and Southeast Asian, Siberian, African and American profiles that mostly serve as sinks for ancestry the West-Eurasian-plus-South-Asian core cannot absorb.
Reading a HarappaWorld result#
A typical Punjabi profile might read S-Indian ~35, Baloch ~35, Caucasian ~10, NE-Euro ~10, with small change; a typical Tamil profile shifts weight toward S-Indian; a Pashtun profile toward Baloch and Caucasian with more NE-Euro. Three rules keep the reading honest:
- Ratios, not absolutes. The S-Indian : Baloch ratio is the calculator's most informative output, and it is a relative affinity along a real cline — not a census of two ancestral tribes. Every South Asian genome carries both; the mix varies by region, caste history and community.
- Cross-calculator numbers do not transfer. The same kit scores differently in K12b's Gedrosia and HarappaWorld's Baloch, and a British reference runs ~44% North European in Dodecad but ~51% NE-Euro here. Anchoring differs; that is expected, not an error to chase.
- Traces are noise until proven otherwise. Sub-percent Papuan, San or American entries in a South Asian profile are the optimiser distributing noise across sixteen bins. Nothing in the output flags which small numbers are real — no calculator can.
The Oracle utilities — HarappaWorld ships both Oracle and Oracle-4 — turn the profile into ranked nearest-reference lists, which is the most interpretable view of a run, especially for parents-from-different-regions cases where Oracle-4's mixed-fit mode was designed to help.
After HarappaWorld#
The project froze years ago; the ancient-DNA record of South and Central Asia did not. If you want the same exploration against dated ancient panels — Indus-periphery-adjacent sources, steppe pastoralists, era by era — the free Global25 admixture calculators and closest-populations rankings run in the browser, and how to get Global25 coordinates covers the one prerequisite. For a claim that needs defending — whether a steppe-related source is required for your genome, whether a trace component survives testing — the step up is a formal qpAdm model with p-values, standard errors and the ability to reject.
Terms used here are defined in the glossary.
References#
- Narasimhan, V. M. et al. (2019). The formation of human populations in South and Central Asia. Science, 365(6457), eaat7487.
- Shinde, V. et al. (2019). An ancient Harappan genome lacks ancestry from Steppe pastoralists or Iranian farmers. Cell, 179(3), 729–735.
- Alexander, D. H., Novembre, J. & Lange, K. (2009). Fast model-based estimation of ancestry in unrelated individuals. Genome Research, 19(9), 1655–1664.



