Population-genetics papers routinely print both: the colourful stacked bars of ADMIXTURE and tables of qpAdm weights with standard errors. Readers understandably treat them as two brands of the same thing — ancestry percentages — and then cannot make sense of why the numbers differ or why the paper needed both. They are not two brands of one thing. They are answers to different questions, and the difference is the most useful thing an interested reader can internalise about method papers.
What each one estimates#
ADMIXTURE (Alexander, Novembre & Lange 2009; heir to STRUCTURE) is unsupervised description: given all genomes in the dataset and a chosen K, it infers K allele-frequency components and each individual's proportions of them, jointly, by maximum likelihood. The components are properties of the dataset — unlabelled, unanchored to any real population, changing when the sample changes.
qpAdm (Haak et al. 2015) is hypothesis testing: the analyst names an explicit model — this target descends from these named source populations, measured against these named outgroups — and the method solves the f4-statistic system for the proportions, returning a weight, standard error and z-score per source and a p-value for the model as a whole. Nothing is inferred from scratch; a proposed history is measured against the data and can fail.
The one-line version: ADMIXTURE finds structure; qpAdm tests stories.
Why the numbers differ for the same samples#
A "steppe" bar in an ADMIXTURE plot and a Yamnaya-related qpAdm weight are different objects. The bar measures affinity to an inferred component that is modal in steppe samples but built from the whole dataset's variation; the weight measures the fitted contribution of an actual sampled population within an explicit model. The bar moves when you add samples or change K; the weight moves when you change sources or outgroups — and only the second comes with an SE and a test. Neither is a distortion of the other; converting between them is a category error, the same one consumer calculator comparisons run into, since consumer calculators are frozen projections onto ADMIXTURE-style components.
Why papers run both — in a fixed order#
The division of labour is deliberate and directional. ADMIXTURE (with PCA) comes first because it is assumption-light: it shows the clusters, clines and outliers that suggest which models are worth proposing, and flags contaminated or mislabelled samples. qpAdm comes second because it is assumption-heavy and answer-strong: temporally coherent sources, a defensible right set, and in return a testable claim with uncertainties. The 2025 audit literature made the pairing quantitative: corroborating distal qpAdm results with PCA and unsupervised ADMIXTURE drives the false-discovery rate of the screen toward zero — the two methods fail differently, so their agreement is worth more than either alone.
The same order, incidentally, is how a careful hobbyist should work: exploration first, testing second, with the free descriptive tools playing ADMIXTURE's role and a formal model playing qpAdm's.
Which answers your question#
| Your question | The right method |
|---|---|
| "What structure is in this dataset?" | ADMIXTURE / PCA — description |
| "Which populations does my genome resemble?" | Descriptive tools: distances, PCA |
| "Can my genome be explained by sources A + B?" | qpAdm — and it may say no |
| "Is source C required, or is A + B enough?" | qpAdm nested models (the record explains) |
| "What percent am I of component X?" | ADMIXTURE-style tools answer it — read what the number is before quoting it |
| "A claim I would defend in an argument" | qpAdm, every time — that is what the p-value is for |
For your own genome, the practical translation: descriptive percentages are available free and instantly (era-scoped calculators are the modern version); the tested statement — explicit ancient sources, stated outgroups, p-value, per-source SEs and z-scores, nested-model table — is a qpAdm analysis, run against AADR v66 and composed by hand for exactly the reasons the audit papers recommend.
Terms used here are defined in the glossary.
References#
- Alexander, D. H., Novembre, J. & Lange, K. (2009). Fast model-based estimation of ancestry in unrelated individuals. Genome Research, 19(9), 1655–1664.
- Haak, W. et al. (2015). Massive migration from the steppe was a source for Indo-European languages in Europe. Nature, 522, 207–211.
- Lawson, D. J., van Dorp, L. & Falush, D. (2018). A tutorial on how not to over-interpret STRUCTURE and ADMIXTURE bar plots. Nature Communications, 9, 3258.
- Flegontova, O. et al. (2025). Performance of qpAdm-based screens for genetic admixture. Genetics, 230(1), iyaf047.



