Ancestrify
All stories

admixture

By Andi Thomaj
3 min read

HarappaWorld explained: the South Asian admixture calculator and its 16 components

What HarappaWorld's S-Indian, Baloch and NE-Euro components are anchored to, why it remains the reference calculator for South Asian ancestry, and how its 2012 categories map onto what ancient DNA later proved.

admixturegedmatchguidesouth-asia

  1. The sixteen components
  2. Reading a HarappaWorld result
  3. After HarappaWorld
  4. References

HarappaWorld is the one GEDmatch admixture project with a single calculator, and the focus shows. Zack Ajmal started the Harappa Ancestry Project in 2011 to do for South Asian ancestry what Dodecad was doing for West Eurasia — collect volunteer kits, build reference panels, and fit genomes against them with ADMIXTURE-based tooling. The calculator that resulted, usually just called HarappaWorld, has sixteen components and is still the default recommendation for anyone with South or Central Asian ancestry on GEDmatch.

The family-wide rules apply — an optimiser over fixed profiles, no test and no error bars — so this post is about what the components mean and what has been learned since 2012.

The sixteen components#

S-Indian · Baloch · Caucasian · NE-Euro · SE-Asian · Siberian · NE-Asian · Papuan · American · Beringian · Mediterranean · SW-Asian · San · E-African · Pygmy · W-African.

For South Asian readers, the first two carry most of the story:

S-Indian peaks in southern Indian references and tracks the ancestry the literature now calls AASI-related (Ancient Ancestral South Indian) — the deep indigenous lineage of the subcontinent. No ancient AASI genome had been sequenced in 2012; the component is the calculator's modern-sample shadow of it.

Baloch peaks in the Baloch and Brahui of Balochistan and is nearly the same statistical object as Dodecad's Gedrosia. It tracks Iranian-plateau-related ancestry — the western Eurasian stream that entered South Asia with (and before) food production. The pairing of S-Indian with Baloch is HarappaWorld's rendering of the cline later formalised as ASI–ANI, and the 2019 ancient-DNA work on the Indus periphery gave that cline its historical anchors: Iranian-plateau-related plus AASI ancestry, with steppe-related ancestry layered on later — carried disproportionately on the NE-Euro side.

NE-Euro peaks in northeast European references and, in South Asian profiles, is the visible edge of that steppe-related layer. Caucasian, as in every calculator of this generation, blends what ancient DNA later split into Caucasus and Iranian-related components. The remaining twelve components give the calculator its global reach — East and Southeast Asian, Siberian, African and American profiles that mostly serve as sinks for ancestry the West-Eurasian-plus-South-Asian core cannot absorb.

Reading a HarappaWorld result#

A typical Punjabi profile might read S-Indian ~35, Baloch ~35, Caucasian ~10, NE-Euro ~10, with small change; a typical Tamil profile shifts weight toward S-Indian; a Pashtun profile toward Baloch and Caucasian with more NE-Euro. Three rules keep the reading honest:

  • Ratios, not absolutes. The S-Indian : Baloch ratio is the calculator's most informative output, and it is a relative affinity along a real cline — not a census of two ancestral tribes. Every South Asian genome carries both; the mix varies by region, caste history and community.
  • Cross-calculator numbers do not transfer. The same kit scores differently in K12b's Gedrosia and HarappaWorld's Baloch, and a British reference runs ~44% North European in Dodecad but ~51% NE-Euro here. Anchoring differs; that is expected, not an error to chase.
  • Traces are noise until proven otherwise. Sub-percent Papuan, San or American entries in a South Asian profile are the optimiser distributing noise across sixteen bins. Nothing in the output flags which small numbers are real — no calculator can.

The Oracle utilities — HarappaWorld ships both Oracle and Oracle-4 — turn the profile into ranked nearest-reference lists, which is the most interpretable view of a run, especially for parents-from-different-regions cases where Oracle-4's mixed-fit mode was designed to help.

After HarappaWorld#

The project froze years ago; the ancient-DNA record of South and Central Asia did not. If you want the same exploration against dated ancient panels — Indus-periphery-adjacent sources, steppe pastoralists, era by era — the free Global25 admixture calculators and closest-populations rankings run in the browser, and how to get Global25 coordinates covers the one prerequisite. For a claim that needs defending — whether a steppe-related source is required for your genome, whether a trace component survives testing — the step up is a formal qpAdm model with p-values, standard errors and the ability to reject.

€29.99 · one-time
The worked version of this analysis
Distances, admixture models and PCA across six eras against 1,535 curated populations, every source panel published in full, with Notable Matches included free.
See the Global25 analysis

Terms used here are defined in the glossary.

References#

  • Narasimhan, V. M. et al. (2019). The formation of human populations in South and Central Asia. Science, 365(6457), eaat7487.
  • Shinde, V. et al. (2019). An ancient Harappan genome lacks ancestry from Steppe pastoralists or Iranian farmers. Cell, 179(3), 729–735.
  • Alexander, D. H., Novembre, J. & Lange, K. (2009). Fast model-based estimation of ancestry in unrelated individuals. Genome Research, 19(9), 1655–1664.

Related posts

GEDmatch admixture calculators explained: Eurogenes, Dodecad, HarappaWorld, MDLP, puntDNAL
GEDmatch admixture calculators explained: Eurogenes, Dodecad, HarappaWorld, MDLP, puntDNAL

Which GEDmatch admixture project to run for your background, what each calculator's components mean, why the projects disagree with each other, and what has aged since 2012–2016.

4 min read
The best admixture calculator in 2026, by question and by background
The best admixture calculator in 2026, by question and by background

There is no single best admixture calculator — there is a best one per question. An honest decision guide across GEDmatch's classics, Global25 tools and formal methods, from someone who builds one of them.

3 min read
GEDmatch Oracle explained: what the distances mean and how to read Oracle-4
GEDmatch Oracle explained: what the distances mean and how to read Oracle-4

Oracle ranks reference populations by how closely their calculator percentages match yours — not by shared DNA. How the distance is computed, what single and mixed modes tell you, and the misreadings to avoid.

3 min read
Back to all stories
Ancestrify

Combining cutting-edge genomic science with rich historical records to map your ancestry across generations and continents.


© 2026 Ancestrify. All rights reserved. · Ancestrify is a trading name of Andi Thomaj, a sole trader registered in Tiranë, Albania · NUIS M61725001N
Card payments processed by POK Payments (RPay Ltd)VISAMASTERCARD