Every qpAdm model has two halves, and the right half gets none of the attention. Sources make the story; the right set — the outgroups — makes the story testable, and a wrongly built right set is how bad models pass and good ones fail. The principles are covered in how to choose sources and outgroups; this post is the reference companion: the standard sets the literature actually uses, with the numbers that justify them.
The job, restated in one line#
Right populations supply the contrasts — the f4-statistics of target and sources against them are the equations the model must satisfy. The requirement is differential relatedness: at least some right populations must be closer to some left populations than to others. A right set symmetric to everything on the left yields no usable equations — the auditors' finding is blunt: qpAdm "will not produce meaningful results". And the one prohibition above all: nothing on the right may have received gene flow from the left more recently than the events being modelled, which is why recent, entangled neighbours belong nowhere near the right set.
O9: the spine#
The set the field standardised on comes from Lazaridis et al. 2016, quoted as "the basic set of nine outgroups (o9)" in the Levant Bronze Age literature:
Mbuti, Ust_Ishim, Kostenki14, MA1, Han, Papuan, Onge, Chukchi, Karitiana
The design is a lesson in itself — one deep African anchor (Mbuti); a ~45,000-year-old Eurasian predating the West/East split (Ust_Ishim); an Upper Palaeolithic European (Kostenki14); an Ancient North Eurasian (MA1); East Asian (Han); Oceanian (Papuan); a deeply diverged South Asian lineage (Onge); a Siberian (Chukchi); a Native American (Karitiana). Every major non-African deep branch appears once, so the set breaks symmetry along many independent axes while remaining upstream of the West Eurasian tangles most models live inside.
The extensions: o9a and beyond#
O9 alone often cannot separate West Eurasian sources from each other — its contrasts are too deep. The published fix is era-appropriate additions, and the canonical example carries its own justification. Agranat-Tamir et al. 2020, modelling the Bronze Age Levant, added Neolithic Anatolia to form o9a — and reported that the addition "significantly improved the model by reducing the standard errors of the mixing coefficients by 3-fold (Iran_ChL) and 2.1-fold (Armenia_EBA)". One well-chosen outgroup, threefold sharper weights: that is what differential relatedness buys when it targets exactly the contrast the sources need. Their widest set, o9aamcn, reaches thirteen:
o9 + Anatolia_N, Armenia_MLBA, CHG, Natufian
For European targets the same logic produces the familiar additions — EHG, WHG, CHG, Anatolia_N, Iran_N and a steppe-distinguishing reference — each present to split one specific pair of candidate streams. An addition is also a stress test: put a close relative of a source on the right and the model either sharpens (the Agranat-Tamir case) or gets rejected — which is information, not misfortune, since it means the proxy was standing in for ancestry the new reference resolves.
How many: the floor, the band, the ceiling#
- Arithmetic floor: degrees of freedom are
|right| − |sources|, so a testable model needs at least one more outgroup than sources — and at exactly one degree of freedom the test has almost no power. (The model record prints the dof.) - Working band: the published sets run 9–13; practitioner practice tops out around 13–15. Within the band, composition beats count every time.
- Measured ceiling: Harney et al. 2021 found qpAdm "begins to reject models that would otherwise be deemed plausible when as few as 30 additional populations are added" — with enough outgroups every model fails, true ones included. More right populations is not more rigour; it is a slow-motion rejection of everything.
Rounded out by the standing prohibitions: nothing cladal with a source (it removes the very axis the model needs), nothing descended from the target, no population on both sides, and no careless mixing of ancient and present-day references — differential DNA damage between those classes biases the statistics themselves.
What we run#
Our published models use era-appropriate right sets built on exactly this literature — the deep spine plus contrasts chosen for the specific source pool under test, listed population by population with sample counts in the model record of every report, because a weight without its right set cannot be evaluated by anyone. The worked tutorial shows a full set in action on a real file; the Model Lab lets you rebuild the right set yourself and watch the SEs answer — the Agranat-Tamir experiment, on your own genome. For the right set's place among all the other rules, the best-practices checklist carries the full discipline in one page, and the population-specific recipes show these sets adapted to South Asian and Middle Eastern contrasts.
Terms used here are defined in the glossary.
References#
- Lazaridis, I. et al. (2016). Genomic insights into the origin of farming in the ancient Near East. Nature, 536, 419–424.
- Agranat-Tamir, L. et al. (2020). The genomic history of the Bronze Age Southern Levant. Cell, 181(5), 1146–1157 (supplementary section D).
- Harney, É., Patterson, N., Reich, D. & Wakeley, J. (2021). Assessing the performance of qpAdm. Genetics, 217(4), iyaa045.
- Haak, W. et al. (2015). Massive migration from the steppe was a source for Indo-European languages in Europe. Nature, 522, 207–211.



