Every qpAdm result is a property of three choices: the target, the sources on the left, and the right populations the sources are measured against. Most bad models are bad because of the second and third choices, not because of the software. This guide is about making those choices well. It is the working method behind every qpAdm analysis Ancestrify publishes, and it is also what a customer needs in order to use the Model Lab productively rather than generating rejections at random.
If you have not met the method, Understanding qpAdm covers what it computes; this post assumes you know what an f4-statistic is and want to know how to set one up.
Left and right, in one paragraph#
The left list is the target followed by the candidate sources. The right list is a set of populations that the model uses as reference points: qpAdm computes f4-statistics of the form f4(target, source; right_i, right_j) and asks whether the target's vector of those statistics can be written as a weighted combination of the sources' vectors. The sources do the explaining; the right populations do the measuring. Nothing on the right is ever assigned a weight.
That asymmetry is the whole design. The right populations are there to detect ancestry that the sources cannot account for. If a right population is related to some ancestry the target has and the sources lack, the f4 pattern will be inconsistent and the model will fail, which is exactly what you want. If no right population is related to that missing ancestry, the model will pass anyway, and it will be wrong.
The classic right set#
A right set used across much of the published literature for West Eurasian targets, in AADR labels (its published pedigree, the O9 set and its era extensions, is documented in the standard right sets reference):
| Population | Why it is there |
|---|---|
| Mbuti.DG | Deep African outgroup; anchors the whole set |
| Ust_Ishim.DG | 45,000-year-old Siberian, basal to most non-African lineages |
| Kostenki14 | Early Upper Palaeolithic European |
| MA1 | Mal'ta boy, Ancient North Eurasian |
| Han.DG | East Asian |
| Papuan.DG | Oceanian |
| Onge.DG | Andamanese, a distinct South Asian lineage |
| Karitiana.DG | Indigenous American, carries ANE-related ancestry |
Eight populations, all of them very distant from a West Eurasian target and, crucially, distant in different directions. That diversity is what gives the set power: a source that carries hidden East Asian ancestry will show up against Han.DG; hidden ANE will show up against MA1 and Karitiana.DG.
To this base, add era-appropriate populations that sit closer to the region but are not
plausible sources for the particular model. For a Bronze Age European target modelled from
Anatolia_N, WHG and Yamnaya_Samara, adding Iran_N, Levant_N, Natufian, CHG and EHG on the right
lets the test distinguish, say, a source with Iranian-related ancestry from one without. The
.DG suffix marks high-coverage diploid shotgun genomes in the AADR; the labels themselves are
explained in the AADR explainer.
Rule one: outgroups must not share drift with a source that the target also shares#
This is the rule that is broken most often and hurts the most. The right set has to be "outgroup-like" with respect to the sources and target: any drift a right population shares with one source but not the others will leak into the f4 pattern and either reject a true model or, worse, admit a false one.
The concrete failure: you put Yamnaya_Samara on the left as a source and EHG on the right. EHG is a major component of Yamnaya, so f4(target, Yamnaya; EHG, Mbuti) is large and structured in a way that has nothing to do with whether the target is a mixture. The model can then fit for the wrong reason or fail for the wrong reason. A right population should be related to the sources only through the deep tree, not through recent admixture with one of them.
A quick check before running anything: for each right population, ask "is this a plausible ancestor or close relative of any single source?" If yes, move it or drop it.
Rule two: sources must be distinguishable through the right set#
qpAdm can only assign weights between sources that the right set can tell apart. If two sources look identical to every population on the right, their weights are a single number split arbitrarily, and the standard errors on both will be enormous.
The formal check is the rank test: qpAdm fits the f4 matrix at rank k minus 1 for k sources, and the record reports the chi-square and p-value at every lower rank. If the rank k minus 2 fit also passes, the sources are not all distinguishable and one of them is redundant. The Ancestrify model record prints this table for every published era; it is read in The model record explained.
The practical consequence: if you want to split Anatolia_N from Iberia_N as separate sources, you need something on the right that sees the difference between them. Against the classic eight alone, they are nearly the same population, and the model will tell you so through a rank test that passes one rank too early.
Rule three: temporal logic#
A source should be older than the target, or at least not descended from it. A 2,000-year-old target modelled from a 1,000-year-old source is a chronological impossibility that qpAdm cannot detect, because f-statistics have no clock. The test will happily fit the model if the drift pattern is compatible, which it often is when the "source" is actually a descendant.
For a living person as target, every ancient population is older, so the rule is trivially met on the left. It still bites on the right: a Medieval population placed on the right for a Bronze Age model is not an outgroup, it is a mixture of the very things being modelled. The audit literature has since quantified how much this rule buys: distal versus proximal protocols differ severalfold in measured false-discovery rate.
Proxies, and why the label matters#
Every source is a proxy: a sampled population standing in for an unsampled ancestral one. The AADR gives you many candidate proxies for the same broad ancestry, and they are not interchangeable.
Consider Anatolia_N versus Anatolia_BA. Both are from Anatolia; one is Neolithic farmers around 6500 BC, the other is Bronze Age people three thousand years later who carry additional Iranian-related and Levantine-related ancestry. A model that uses Anatolia_BA as the "farmer" source for a European target is asking the wrong question: it is modelling in a component that European farmers never had, and the CHG or Iran_N weight elsewhere in the model will shrink to compensate. The p-value may still pass. The weights will be wrong.
The same logic separates Iran_N from CHG, Levant_N from Natufian, Steppe_MLBA from Yamnaya_Samara. The last pair is a good example of how much a label carries: Steppe_MLBA populations have a farmer-related layer that Yamnaya_Samara lacks, so swapping one for the other in a European model shifts weight from Anatolia_N to the steppe source by several points.
The rule is to pick the proxy that is closest in time and place to the ancestry you actually mean, and to state the label precisely in the report. Ancestrify publishes the exact AADR panel label beside every source name for this reason.
The rotating strategy#
Harney and colleagues (2021) formalised what practitioners had been doing informally: instead of one fixed right set, take a pool of candidate populations and rotate them, so that each candidate is on the left in some runs and on the right in others. A source that is genuinely needed will be required in every configuration; one that only "works" when a particular competitor is on the right is a proxy artefact.
Rotation is powerful and dangerous in equal measure. Powerful, because it exposes models that pass only by luck of the right set. Dangerous, because if you rotate enough combinations, some will clear p = 0.05 by chance, and picking the best-scoring one is a false-discovery machine. We built an automated rotator, measured its false-discovery rate, and removed it. Rotation is a diagnostic you run on a model you already believe, not a search procedure for finding one. The published false-discovery numbers behind that verdict are in the rotation explainer.
The two-to-four source sweet spot#
One source is a qpWave-style question: is the target a clade with this population? Rarely true for anyone living. Two sources is the most defensible model in the literature, because the rank test has the least room to hide. Three is the workhorse for Holocene West Eurasia (farmer, hunter-gatherer, steppe). Four is possible with a dense file and a strong right set. Beyond four, standard errors balloon, the rank test loses power, and the model starts to describe the reference panel's structure rather than the target.
The publish bar Ancestrify applies at every tier makes this concrete: every source must sit at |Z| > 3 with SE < 0.10, and the model at p > 0.05. Adding a fifth source almost never survives that, and when it does the nested four-source model usually passes too, which means the fifth was not needed.
Common mistakes, collected#
- Putting a source's parent on the right. EHG right, Yamnaya left; Anatolia_N right, Iberia_N left. Rule one.
- Using a descendant as a source. Anatolia_BA to model a Neolithic-era question. Rule three.
- Two proxies for one ancestry on the left. Iran_N and CHG together, against a right set that cannot separate them. Rule two; watch the SEs.
- A thin right set. Six outgroups makes almost everything pass. The p-value is only as strong as the set it was computed against.
- Reading a passing p-value as confirmation. Several contradictory models can pass. A model is "not refuted", never "confirmed".
- Rotating until something passes. See above.
- Ignoring the nested models. If removing a source still passes, publish the simpler model.
A worked illustrative example#
Numbers below are invented to illustrate the reasoning, not a customer's result.
Target: a living person whose file merged to about 210,000 markers with AADR v66.
Right set (11): Mbuti.DG, Ust_Ishim.DG, Kostenki14, MA1, Han.DG, Papuan.DG, Onge.DG, Karitiana.DG, Iran_N, Levant_N, EHG.
Run 1: two sources, Anatolia_N + WHG.
p-value: 0.0007 chi-square: 29.8 dof: 9
Anatolia_N 0.842 SE 0.021 Z 40.1
WHG 0.158 SE 0.021 Z 7.5
Rejected. The right set contains EHG and MA1, both of which carry Ancient North Eurasian ancestry, and the target shares more drift with them than either source can explain. That is the right set doing its job: it has detected something missing.
Run 2: three sources, Anatolia_N + WHG + Yamnaya_Samara.
p-value: 0.31 chi-square: 9.4 dof: 8 f4 rank: 2
Anatolia_N 0.581 SE 0.029 Z 20.0 95% CI 0.524 to 0.638
Yamnaya_Samara 0.271 SE 0.033 Z 8.2 95% CI 0.206 to 0.336
WHG 0.148 SE 0.027 Z 5.5 95% CI 0.095 to 0.201
Passes the bar on every line. Rank test: rank 1 rejected at p = 0.0007 (that was Run 1), so three sources is the minimum. Nested models: removing any one source is rejected below p = 0.001.
Run 3: swap the proxy. Iberia_N + Yamnaya_Samara.
p-value: 0.42 chi-square: 8.1 dof: 9
Iberia_N 0.712 SE 0.030 Z 23.7 95% CI 0.653 to 0.771
Yamnaya_Samara 0.288 SE 0.030 Z 9.6 95% CI 0.229 to 0.347
Also passes, with two sources instead of three. Why? Iberia_N already contains a WHG-related layer that Anatolia_N does not, so the hunter-gatherer share is absorbed into the farmer proxy. Both models are admissible. They say different things: Run 2 says "farmer, hunter-gatherer and steppe, with the hunter-gatherer share estimated separately"; Run 3 says "a western farmer population that had already absorbed hunter-gatherers, plus steppe".
Which to publish depends on the question. If the target's declared background is western European, Run 3 with the geographically closer proxy is the more informative model, and the record should say why. If the point is to estimate the WHG share as its own number, Run 2 is the one. In either case the right set is published in full, the nested table is published, and the analyst's customer explanation says which label was chosen and what it does not mean. That is what the Reading view in every Ancestrify report carries.
Notice what changed between Run 1 and Run 2: nothing about the target, and nothing about the right set. A single source was added and a rejection turned into a pass. Notice what changed between Run 2 and Run 3: one proxy label, and the number of sources needed fell by one. The result is a property of the model, and the model is a set of choices.
Trying your own choices#
After an Ancestrify report is published, the Model Lab unlock (10 EUR, one time) lets you compose your own left and right lists from the same merged panel your report was computed on, with your sample as the target, up to 100 runs per rolling 24 hours, and returns the real statistics. It also includes the EIGENSTRAT bundle download for running ADMIXTOOLS 2 on your own machine. The walkthrough is in Run your own qpAdm models. Before buying anything, the free AdmixTools 2 Lab runs qpAdm over the public panel, which is enough to practise the rules above on real data.
If you would rather have the choices made and defended for you, that is what the qpAdm analysis is: every model composed and checked by a person, exhausting the plausible proxy labels for each source, and published only when it clears p > 0.05, every |Z| > 3 and every SE < 0.10.
References#
- Haak, W. et al. (2015). Massive migration from the steppe was a source for Indo-European languages in Europe. Nature, 522, 207 to 211.
- Harney, É., Patterson, N., Reich, D. & Wakeley, J. (2021). Assessing the performance of qpAdm: a statistical tool for studying population admixture. Genetics, 217(4), iyaa045.
- Lazaridis, I. et al. (2016). Genomic insights into the origin of farming in the ancient Near East. Nature, 536, 419 to 424.
- Maier, R. et al. (2023). On the limits of fitting complex models of population history to f-statistics. eLife, 12, e85492.
- Patterson, N. et al. (2012). Ancient admixture in human history. Genetics, 192(3), 1065 to 1093.



