Choose sources and outgroups by hand and you inherit a suspicion: did the analyst tune the right set until the favourite model passed? The rotating protocol was invented to answer that suspicion with procedure. Take one candidate pool; in each model, some candidates are sources and the rest join the outgroups; test every arrangement. No population is privileged, every model faces comparable opposition, and the search space is explored rather than navigated by preference. Skoglund et al. 2017 introduced the strategy for African population history; Harney et al. 2021 recommended it over the fixed-base alternative, whose models are "not equivalent, and therefore are difficult to compare".
So rotation is the gold standard? It is — for fairness of comparison. For truth of the winners, the 2025 numbers changed the conversation.
What rotation does well#
The fixed-base strategy has a structural flaw the auditors named precisely: any population parked
permanently in the reference set can never be tried as a source, so the search silently excludes
models nobody decided to exclude. Rotation fixes that, and adds a second virtue: moving a
rejected model's sources into the right set of surviving models — model competition, in
Flegontova et al. 2025's terminology — is a genuine stress test, since
a reference differentially related to a source's stream either
sharpens the model or exposes the proxy. In ADMIXTOOLS 2 the
whole machine is one call — qpadm_rotate() — and precomputed f2 statistics make thousands of
models cost seconds.
That cheapness is the trap.
The measured bill#
Flegontova et al. 2025 scored rotating screens against simulated histories where the truth was known. Run without temporal stratification — targets allowed to predate sources, the proximal regime — the rotating protocol's false-discovery rate was 72.5–100%: nearly everything such a screen accepts is wrong, mostly via false rejections of simple true models that push acceptance up the complexity ladder. Temporally stratified rotation landed at 16.4–31.2% — usable, improvable toward zero with independent corroboration (PCA, unsupervised ADMIXTURE). And the failure is not data-starved: proximal screens got worse with more data, as tighter errors rejected true simple models harder.
Two structural cautions complete the bill. A rotating screen ranks its survivors somehow, and the tempting key — the p-value — picks the true model in only 48% of cases (Harney et al. 2021); p ranks nothing. And every screen is a multiple-testing exercise: thousands of models mean dozens of accidental passes at any threshold, which is why the auditors' composite feasibility criteria (weights bounded away from 0 and 1 within error, trailing simpler models rejected) exist at all — Williams et al. 2024 measured p-alone screening at an 84% false-discovery rate.
What rotation is actually for#
Read the numbers as an operating manual rather than a verdict and rotation has a clear, narrow job: map the model space, never crown the winner. Rotate to learn which sources are interchangeable (they swap without moving the fit — a cladality fact worth knowing), which candidate the data genuinely refuses everywhere, whether any temporally legal 2-way family survives at all. Then the analyst work begins: temporal stratification enforced, era-appropriate outgroups chosen for the specific contrast, simple models given first refusal, and the final candidate defended with margins — not with its rank in a screen.
That division of labour is exactly how it works here. Rotation-style enumeration was built into our pipeline early and then deliberately demoted: it proposes, and a person disposes — every published model is composed, run and checked by hand against the one bar (p above 0.05, every |Z| above 3, every SE below 0.10), with roughly four hundred hand-built runs behind a typical order rather than one automated sweep. The Model Lab hands you the same discipline interactively: your merged dataset, any model you can compose, and the numbers to reject most of what you try — which, as the FDR tables show, is the feature.
Terms used here are defined in the glossary.
References#
- Skoglund, P. et al. (2017). Reconstructing prehistoric African population structure. Cell, 171(1), 59–71.
- Harney, É., Patterson, N., Reich, D. & Wakeley, J. (2021). Assessing the performance of qpAdm. Genetics, 217(4), iyaa045.
- Williams, M. P. et al. (2024). Testing times: disentangling admixture histories in recent and complex demographies using ancient DNA. Genetics, 228(1), iyae110.
- Flegontova, O. et al. (2025). Performance of qpAdm-based screens for genetic admixture. Genetics, 230(1), iyaf047.



