Type qpAdm into any forum search and a genre appears: the method is broken, the method is fine, the papers using it are all wrong, the critics misunderstand it. Both camps cite the same three audits. This post is the referee's version: what Harney 2021, Williams 2024 and Flegontova 2025 actually established — numbers, not vibes — what survived, and what a careful user does differently now. Spoiler for the impatient: the method survived; several popular workflows did not.
What the audits established#
The machinery is sound. Harney and colleagues' baseline validation still stands un-refuted: on clean simulated data, weights are unbiased (99.3% of estimates within three standard errors of truth), p-values are uniform when the model is true, and the estimates are robust to pseudohaploid data, uniform damage, random missingness and single-individual samples. Nobody in the criticism literature disputes this layer. The identity underneath is arithmetic, and the arithmetic works.
Model selection by p-value does not work. The same audit's most consequential number: among plausible candidates, the true model carries the best p-value in only 48% of cases. Every workflow that sweeps models and keeps the top p — which described a lot of forum practice and some published practice — was measured and found to be roughly a coin flip.
Unstratified screening fails at scale. Flegontova and colleagues put false-discovery rates on whole protocols: 72.5–100% for proximal rotating screens, degrading further as data grows, with 40% of consistently rejected models belonging to targets with zero admixture events — the screens manufacture admixture. Temporally stratified (distal) protocols ran 16.4–31.2%, improvable toward zero with independent corroboration. The full tables are here.
Resolution has a floor. Williams and colleagues measured it: sources separated by FST under roughly 0.002 cannot be distinguished by any right set, and at Iron-Age-scale differentiation the true model is only ~22% of the plausible set. Fine-grained historical claims — this medieval population rather than its neighbour — sit at or below the floor more often than users want.
Single-winner claims overreach. Maier and colleagues found alternative topologies fitting as well as or better than 19 of 22 published admixture graphs — the same lesson at the graph scale: the family of fitting models is the honest result, not one winner.
What that means, and does not mean#
Read carefully, the criticism literature is a manual, not an obituary. None of it found the estimator biased; all of it found usage patterns with measurable error rates. The analogy that fits: nobody concluded thermometers are broken on discovering that waving one around the room does not measure the patient. The audits priced the workflows — and the priced-out ones (p-value ranking, unstratified rotation at scale, fine splits below the FST floor) were exactly the convenient ones.
What genuinely changed in careful practice: composite feasibility criteria replaced p-thresholds alone (84% FDR by itself, severalfold better with weight bounds and nested-model checks); temporal stratification hardened from advice into a rule; right sets are built for the specific contrast rather than by piling on; and corroboration by orthogonal methods — PCA and unsupervised ADMIXTURE — became part of the protocol, because it measurably rescues the distal FDR.
Where that leaves a reader in 2026#
Trust a qpAdm result in proportion to its protocol, and the protocol is checkable from the report: Is it stratified? Is the right set stated in full, with nothing entangled? Were simpler models given first refusal? Is acceptance composite? Are the alternatives that also passed reported? A result carrying all of that — the model record discipline — sits in the best-measured error regime any ancestry method offers. A result carrying none of it inherits the 72–100% band, whatever software logo it wears.
That conditional is also the honest sales pitch for how we run the method: hand-composed, stratified, one bar for every tier, rotation demoted to mapping, the full record published. Not because the critics were wrong — because they were right, and the fixes are known, and applying them is work someone has to actually do.
Terms used here are defined in the glossary.
References#
- Harney, É., Patterson, N., Reich, D. & Wakeley, J. (2021). Assessing the performance of qpAdm. Genetics, 217(4), iyaa045.
- Williams, M. P. et al. (2024). Testing times: disentangling admixture histories in recent and complex demographies using ancient DNA. Genetics, 228(1), iyae110.
- Flegontova, O. et al. (2025). Performance of qpAdm-based screens for genetic admixture on admixture-graph-shaped histories and stepping-stone landscapes. Genetics, 230(1), iyaf047.
- Maier, R. et al. (2023). On the limits of fitting complex models of population history to f-statistics. eLife, 12, e85492.



