Every region teaches a different qpAdm lesson. Europe teaches the standard recipe; South Asia teaches what to do when a founding ancestry has never been sequenced. The deepest stream in every South Asian genome — the lineage geneticists call AASI — has no ancient sample, and every model of a billion and a half people's ancestry routes through proxies for it. That constraint, honestly handled, is the whole craft here.
The streams, and the hole in the record#
The working framework (Narasimhan and colleagues 2019, still the reference) writes South Asian variation as combinations of:
- AASI (Ancient Ancestral South Indians) — the subcontinent's deep indigenous hunter-gatherer lineage, distantly related to Andamanese islanders. No ancient genome exists. It is inferred, never sampled.
- Iranian-plateau-related farmer ancestry — but not proxied by
Iran_GanjDareh_Nnaively: the lineage in South Asia split from Iranian farmers before agriculture reached the plateau, which is why the field's proxy of choice is Indus Periphery (IVC-era individuals from Gonur and Shahr-i-Sokhta genetically continuous with Rakhigarhi), itself already an Iranian-related + AASI mixture. - Steppe MLBA —
Kazakhstan_Sintashta_MLBA-related ancestry (not Yamnaya EBA: the steppe stream that reaches South Asia is the later, farmer-admixed one; substituting EBA steppe is the region's classic wrong-era error, temporal logic again).
Modern populations then span the ANI–ASI cline: ASI ≈ AASI + Indus-Periphery-related; ANI ≈ Indus-Periphery-related + steppe MLBA. The 2009 f4-ratio estimate of 39–71% ANI across the cline remains the sanity anchor.
The standard recipe#
Left: target + Indus_Periphery (pooled or West-cluster), Kazakhstan_Sintashta_MLBA (or
Central_Steppe_MLBA), and an AASI proxy — in practice Onge.DG, with the caveat that owns the
next section.
Right: the O9 spine minus Onge when Onge is a
source (never both sides — the cladality rule), plus the contrasts that split Iranian-related
from steppe from AASI: Russia_EHG, Georgia_CHG, Anatolia_N, Iran_GanjDareh_N,
WSHG/Botai-related for the inner-Asian edge, Mota.DG, Ust_Ishim.DG, MA1. Iran_N moves
to the right because Indus Periphery carries the Iranian-related stream on the left — the
right-set member that discriminates your sources is
worth five generic ones.
Run lowest rank first: many southern targets pass as two-source Indus_Periphery + Onge; northern and upper-caste targets typically require the third steppe stream.
The proxy problem, stated honestly#
Onge are not AASI. They are a sister lineage separated by tens of millennia of island isolation and drift — the least bad sampled relative, not the ancestor. Consequences, in descending order of pain: absolute AASI percentages shift by proxy choice (models swapping Onge for other proxies move weights by real margins, so quote AASI fractions as proxy-conditional); drift accumulated on the Onge branch can push p-values down for reasons that are not model failure; and no right set fully rescues a source that is itself off-topology. The published work handles this with explicit proxy sensitivity checks — running the same model across proxy choices and reporting the spread. Do the same, or read others' absolute percentages with that spread in mind.
The named failure modes#
- Wrong-era steppe. Yamnaya EBA instead of Sintashta/Central MLBA — passes sometimes, means the wrong thing always.
- Onge on both sides. Source and O9 outgroup — instant entanglement violation; prune the right set.
- East Asian edges. Munda-speaking and northeastern targets carry East/Southeast
Asian-related ancestry the three-stream model lacks; the rank test
will demand a fourth source — give it
China_YR_LN- or Austroasiatic-associated proxies rather than torturing the right set. - Endogamy noise. Strong founder events in many jati groups inflate drift; SEs widen and p-values roughen even when the model is structurally right. More target individuals help more than more outgroups.
A worked reading#
A northwest-Indian-ancestry target: p = 0.18, Indus_Periphery 0.61 ± 0.030, Steppe_MLBA 0.24 ± 0.019, Onge 0.15 ± 0.026; two-source nested models rejected; ~640k SNPs. Reading: compatible and three-streams-required; the steppe weight is solid ANI-cline-upper territory; and the honest sentence for the report is "15% AASI as proxied by Onge" — with the proxy-swap spread quoted beside it if the number will bear weight. That sentence is the region's entire lesson in miniature: the method is exact about what it tested, and the analyst's job is to keep the words as exact as the arithmetic.
Terms used here are defined in the glossary.
References#
- Narasimhan, V. M. et al. (2019). The formation of human populations in South and Central Asia. Science, 365, eaat7487. (The framework, Indus Periphery, steppe MLBA.)
- Shinde, V. et al. (2019). An ancient Harappan genome lacks ancestry from steppe pastoralists or Iranian farmers. Cell, 179, 729–735. (Rakhigarhi and the pre-agricultural Iranian split.)
- Reich, D., Thangaraj, K., Patterson, N., Price, A. L. & Singh, L. (2009). Reconstructing Indian population history. Nature, 461, 489–494. (ANI/ASI.)
- Moorjani, P. et al. (2013). Genetic evidence for recent population mixture in India. American Journal of Human Genetics, 93(3), 422–438. (Dating the cline's formation.)



