Ancestrify
All stories

qpadm

By Andi Thomaj
3 min read

Distal vs proximal qpAdm models: the choice that sets your false-discovery rate

Whether your sources predate your target is not a style preference — measured false-discovery rates run 16–31% for temporally stratified protocols and 72–100% for proximal rotating screens. What each model type is for.

qpadmmethodology

  1. The measured stakes
  2. Why proximity hurts
  3. What each class is for
  4. The consumer footnote worth knowing
  5. References

Two qpAdm models of the same Iron Age genome can both pass every numeric bar and still belong to different risk classes. Model one explains the target from deep-time ancestries — hunter-gatherers, first farmers, steppe pastoralists, all comfortably older than the target. Model two explains it from near-contemporary neighbours — populations a few centuries removed, some possibly younger. The literature calls these distal and proximal models, and the 2025 audit work turned the distinction from taste into measurement: it is the single largest controllable driver of how often qpAdm-based conclusions are wrong.

Definitions first, as Flegontova et al. 2025 fix them: a protocol is distal when "the target postdates or is contemporaneous with all proxy sources" — temporal stratification — and proximal when "the target predates at least one proxy source", or more loosely when no stratification is enforced at all.

The measured stakes#

Flegontova et al. simulated thousands of qpAdm screens over histories where the truth was known, and scored how often accepted models were false:

ProtocolFalse-discovery rate
Proximal, rotating72.5–100%
Proximal, non-rotatingmedians 52–58%
Distal (rotating or not)16.4–31.2%, the lowest of every setup

Two of their findings sharpen the point beyond the headline numbers. More data makes proximal screens worse, not better — FDR grew significantly with data volume for proximal models, while distal FDR was insensitive to it. Tighter standard errors on a mis-specified model reject the true simple story and push the screen up the complexity ladder. And most of the damage is false rejection of simple models: 40% of consistently rejected models had targets with zero admixture events in their history. A proximal screen does not just accept wrong models; it manufactures admixture where none happened.

The rescue also has a number: corroborating distal results with PCA and unsupervised ADMIXTURE drove FDR toward zero in their setups — while for the proximal rotating protocol no such rescue worked.

Why proximity hurts#

Nothing about recency is sinful per se; the mechanism is assumption erosion. qpAdm requires that no right population received gene flow from the left populations after their separation, and that sources stand cleanly for distinct ancestry streams. Near-contemporary populations are connected by exactly the recent, tangled gene flow that violates both — every neighbour has exchanged migrants with every other, sources become near-cladal with each other, and the model's algebra is asked to distinguish streams the history never separated. Distal sources sit upstream of that tangle: fewer shared recent edges with the outgroups, cleaner differential relatedness, an honest rank test. The price is interpretive: "42% Anatolian-farmer-related" is true and unromantic, where a proximal "42% medieval Population X" sounds like history — and is precisely the claim class the FDR table warns about.

What each class is for#

Distal models are the backbone: formation-era proportions, stable across data growth, the right default for any claim that needs defending. Proximal models are hypothesis probes: genuinely valuable when the question itself is recent ("does this early-medieval genome need a Slavic-related stream on top of the local Iron Age base?"), and legitimate when built one at a time with era-appropriate outgroups, temporal ordering intact, and nested simpler models given first refusal — never as an automated screen, which is where the 72–100% band lives.

The consumer footnote worth knowing#

A modern genotype file as target with ancient sources is distal by construction — the sources predate the target unavoidably. Consumer qpAdm done properly therefore starts in the favourable protocol class, which is a quiet structural advantage of the whole product category. Our own publish discipline matches the audit literature's: every published model is temporally stratified, searched lowest rank first, held to one bar (p above 0.05, every source's |Z| above 3, every SE below 0.10) whatever the tier — and composed by hand, because the protocols that fail in the tables above are precisely the automated ones. The worked example shows the sequence on a real file, and the Model Lab lets you probe proximal hypotheses yourself on your own merged dataset, with the distal model as the anchor it should be.

From €29.99 · one-time
The tested version of this question
A qpAdm model composed, run and checked by hand against AADR v66, published with its p-value, every source's standard error and z-score, and the full right set, so the result can be argued with.
See the qpAdm analysis

Terms used here are defined in the glossary.

References#

  • Flegontova, O. et al. (2025). Performance of qpAdm-based screens for genetic admixture on admixture-graph-shaped histories and stepping-stone landscapes. Genetics, 230(1), iyaf047.
  • Harney, É., Patterson, N., Reich, D. & Wakeley, J. (2021). Assessing the performance of qpAdm. Genetics, 217(4), iyaa045.
  • Williams, M. P. et al. (2024). Testing times: disentangling admixture histories in recent and complex demographies using ancient DNA. Genetics, 228(1), iyae110.

Related posts

Learn qpAdm: the complete guide, from zero to defensible models
Learn qpAdm: the complete guide, from zero to defensible models

A structured learning path through everything qpAdm — what to read in what order, from the f4-statistics underneath to running models in R, choosing outgroups, reading results and knowing the method's measured limits.

3 min read
qpAdm best practices: the current checklist, with the numbers behind every rule
qpAdm best practices: the current checklist, with the numbers behind every rule

The definitive working checklist for qpAdm in 2026 — temporal stratification, right-set construction, lowest-rank-first search, composite feasibility and reporting standards — each rule carrying its measured justification from the 2021–2025 audit literature.

4 min read
Is qpAdm still reliable? What the criticism established, honestly assessed
Is qpAdm still reliable? What the criticism established, honestly assessed

qpAdm has been audited harder than any tool in ancient DNA: measured false-discovery rates, resolution floors, protocol failures. What the 2021–2025 criticism literature actually established, what survived it, and how practice changed.

3 min read
Back to all stories
Ancestrify

Combining cutting-edge genomic science with rich historical records to map your ancestry across generations and continents.


© 2026 Ancestrify. All rights reserved. · Ancestrify is a trading name of Andi Thomaj, a sole trader registered in Tiranë, Albania · NUIS M61725001N
Card payments processed by POK Payments (RPay Ltd)VISAMASTERCARD