Ancestrify
All stories

qpadm

By Andi Thomaj
4 min read

qpAdm models for South Asian ancestry: AASI, Indus Periphery and the proxy problem

South Asia is qpAdm's hardest standard fixture: one ancestral stream has no ancient sample at all. The working recipe — Indus Periphery, steppe MLBA, the Onge-as-AASI-proxy problem — with right sets, failure modes and a worked reading.

qpadmguidepopulation-geneticssouth-asia

  1. The streams, and the hole in the record
  2. The standard recipe
  3. The proxy problem, stated honestly
  4. The named failure modes
  5. A worked reading
  6. References

Every region teaches a different qpAdm lesson. Europe teaches the standard recipe; South Asia teaches what to do when a founding ancestry has never been sequenced. The deepest stream in every South Asian genome — the lineage geneticists call AASI — has no ancient sample, and every model of a billion and a half people's ancestry routes through proxies for it. That constraint, honestly handled, is the whole craft here.

The streams, and the hole in the record#

The working framework (Narasimhan and colleagues 2019, still the reference) writes South Asian variation as combinations of:

  • AASI (Ancient Ancestral South Indians) — the subcontinent's deep indigenous hunter-gatherer lineage, distantly related to Andamanese islanders. No ancient genome exists. It is inferred, never sampled.
  • Iranian-plateau-related farmer ancestry — but not proxied by Iran_GanjDareh_N naively: the lineage in South Asia split from Iranian farmers before agriculture reached the plateau, which is why the field's proxy of choice is Indus Periphery (IVC-era individuals from Gonur and Shahr-i-Sokhta genetically continuous with Rakhigarhi), itself already an Iranian-related + AASI mixture.
  • Steppe MLBAKazakhstan_Sintashta_MLBA-related ancestry (not Yamnaya EBA: the steppe stream that reaches South Asia is the later, farmer-admixed one; substituting EBA steppe is the region's classic wrong-era error, temporal logic again).

Modern populations then span the ANI–ASI cline: ASI ≈ AASI + Indus-Periphery-related; ANI ≈ Indus-Periphery-related + steppe MLBA. The 2009 f4-ratio estimate of 39–71% ANI across the cline remains the sanity anchor.

The standard recipe#

Left: target + Indus_Periphery (pooled or West-cluster), Kazakhstan_Sintashta_MLBA (or Central_Steppe_MLBA), and an AASI proxy — in practice Onge.DG, with the caveat that owns the next section.

Right: the O9 spine minus Onge when Onge is a source (never both sides — the cladality rule), plus the contrasts that split Iranian-related from steppe from AASI: Russia_EHG, Georgia_CHG, Anatolia_N, Iran_GanjDareh_N, WSHG/Botai-related for the inner-Asian edge, Mota.DG, Ust_Ishim.DG, MA1. Iran_N moves to the right because Indus Periphery carries the Iranian-related stream on the left — the right-set member that discriminates your sources is worth five generic ones.

Run lowest rank first: many southern targets pass as two-source Indus_Periphery + Onge; northern and upper-caste targets typically require the third steppe stream.

The proxy problem, stated honestly#

Onge are not AASI. They are a sister lineage separated by tens of millennia of island isolation and drift — the least bad sampled relative, not the ancestor. Consequences, in descending order of pain: absolute AASI percentages shift by proxy choice (models swapping Onge for other proxies move weights by real margins, so quote AASI fractions as proxy-conditional); drift accumulated on the Onge branch can push p-values down for reasons that are not model failure; and no right set fully rescues a source that is itself off-topology. The published work handles this with explicit proxy sensitivity checks — running the same model across proxy choices and reporting the spread. Do the same, or read others' absolute percentages with that spread in mind.

The named failure modes#

  • Wrong-era steppe. Yamnaya EBA instead of Sintashta/Central MLBA — passes sometimes, means the wrong thing always.
  • Onge on both sides. Source and O9 outgroup — instant entanglement violation; prune the right set.
  • East Asian edges. Munda-speaking and northeastern targets carry East/Southeast Asian-related ancestry the three-stream model lacks; the rank test will demand a fourth source — give it China_YR_LN- or Austroasiatic-associated proxies rather than torturing the right set.
  • Endogamy noise. Strong founder events in many jati groups inflate drift; SEs widen and p-values roughen even when the model is structurally right. More target individuals help more than more outgroups.

A worked reading#

A northwest-Indian-ancestry target: p = 0.18, Indus_Periphery 0.61 ± 0.030, Steppe_MLBA 0.24 ± 0.019, Onge 0.15 ± 0.026; two-source nested models rejected; ~640k SNPs. Reading: compatible and three-streams-required; the steppe weight is solid ANI-cline-upper territory; and the honest sentence for the report is "15% AASI as proxied by Onge" — with the proxy-swap spread quoted beside it if the number will bear weight. That sentence is the region's entire lesson in miniature: the method is exact about what it tested, and the analyst's job is to keep the words as exact as the arithmetic.

From €29.99 · one-time
The tested version of this question
A qpAdm model composed, run and checked by hand against AADR v66, published with its p-value, every source's standard error and z-score, and the full right set, so the result can be argued with.
See the qpAdm analysis

Terms used here are defined in the glossary.

References#

  • Narasimhan, V. M. et al. (2019). The formation of human populations in South and Central Asia. Science, 365, eaat7487. (The framework, Indus Periphery, steppe MLBA.)
  • Shinde, V. et al. (2019). An ancient Harappan genome lacks ancestry from steppe pastoralists or Iranian farmers. Cell, 179, 729–735. (Rakhigarhi and the pre-agricultural Iranian split.)
  • Reich, D., Thangaraj, K., Patterson, N., Price, A. L. & Singh, L. (2009). Reconstructing Indian population history. Nature, 461, 489–494. (ANI/ASI.)
  • Moorjani, P. et al. (2013). Genetic evidence for recent population mixture in India. American Journal of Human Genetics, 93(3), 422–438. (Dating the cline's formation.)

Related posts

Punjabi DNA: Ancient Origins from the Indus to the Swat Valley
Punjabi DNA: Ancient Origins from the Indus to the Swat Valley

Punjabi ancestry through ancient DNA: the Indus Periphery and Rakhigarhi genomes, AASI, the Steppe MLBA layer, the Swat Iron Age references and qpAdm.

8 min read
Pashtun DNA: Ancient Origins from the Indus to the Swat Valley
Pashtun DNA: Ancient Origins from the Indus to the Swat Valley

What ancient genomes say about Pashtun ancestry: Iranian farmer, AASI and Steppe streams, the Swat valley references, and the legends DNA cannot confirm.

8 min read
Bengali DNA: Ancient Origins of Bangladesh and West Bengal
Bengali DNA: Ancient Origins of Bangladesh and West Bengal

What genomes say about Bengali ancestry: a strong AASI share, a Southeast Asian layer from Austroasiatic and Tibeto-Burman speakers, and the gaps in the data.

8 min read
Back to all stories
Ancestrify

Combining cutting-edge genomic science with rich historical records to map your ancestry across generations and continents.


© 2026 Ancestrify. All rights reserved. · Ancestrify is a trading name of Andi Thomaj, a sole trader registered in Tiranë, Albania · NUIS M61725001N
Card payments processed by POK Payments (RPay Ltd)VISAMASTERCARD