Ancestrify
All stories

qpadm

By Andi Thomaj
8 min read

A qpAdm report, read line by line: a worked example

An illustrative qpAdm report for one era, read from the model line through the weights, p-value, right set, nested models and rank test to the explanation.

qpadmguidemethodology

  1. The model line
  2. The weights table
  3. The fit block
  4. The right set
  5. The nested-model table
  6. The rank test
  7. Warnings
  8. Reading your model
  9. The Reading view and the plain-text record
  10. What you can check yourself
  11. Reproducing it
  12. References

Explaining what a qpAdm report contains is one thing; reading an actual one, line by line, is another. This post walks through an illustrative report for one era in the format an Ancestrify qpAdm analysis publishes. Every figure below is invented to be internally consistent and to clear the publish bar; it is not a customer's result. The interactive version, with example data, is on the demo.

The era is Hunter-Gatherer and Neolithic Farmer. The target is a living person whose raw file merged to about 214,000 markers against AADR v66. The report's Reading view is where all of this lives; the Atlas view shows the same model as a map and a story.

The model line#

Era: Hunter-Gatherer and Neolithic Farmer
Model: Anatolia_N + Yamnaya_Samara + WHG
Target: <your sample>    Panel: AADR v66    Run: 7c1e...    Date: 2026-08-30

Three sources. The order is the order the analyst listed them, not a ranking. The panel version is stated because a model is only reproducible against the panel it was run on. The run id is the handle for the exact ADMIXTOOLS 2 run behind the numbers.

What the line does not say: it makes no claim of descent from these three populations. Each is a sampled proxy for an unsampled ancestral population, chosen because it is the closest available stand-in in time and place.

The weights table#

#SourcePanel labelnShareWeightSEZ95% CI
01Anatolian Neolithic FarmerAnatolia_N2448.7%0.4870.02817.40.432 to 0.542
02Western Steppe HerderYamnaya_Samara1033.4%0.3340.03110.80.273 to 0.395
03Western Hunter-GathererWHG817.9%0.1790.0247.460.132 to 0.226

Read the columns right to left, because the rightmost ones decide whether the leftmost mean anything.

95% CI. The weight plus or minus 1.96 standard errors. The steppe share is not "33%"; it is "somewhere between 27% and 40%". None of the three intervals reaches zero, which is what the publish bar is designed to guarantee.

Z. Weight divided by SE. Every source sits well above 3; the smallest, WHG at 7.46, is seven and a half standard errors from nothing. A source at Z = 1.5 would be a source the data cannot certify, whatever its share said.

SE. All three are below 0.10, and in fact below 0.035, which is what a file of this coverage supports. A sparser file would carry SEs of 0.06 to 0.12 on the same model.

Weight and share. The same number twice: raw proportion, and rounded percentage. They sum to 1.000 because qpAdm constrains them to.

n. How many individuals the panel population contains. Eight WHG individuals is a thinner reference than 24 Anatolian farmers, and the record says so.

Panel label. The exact AADR population, beside the catalog name, so the model can be reproduced and so a curated name never hides which samples were used. The labels are explained in the AADR explainer, and each catalog population has a page in the ancestry directory.

The fit block#

p-value: 0.164    chi-square: 11.72    dof: 8    f4 rank: 2
min SNPs per f4: 121,880    merged SNPs: 214,306    jackknife blocks: 711

p = 0.164. Above 0.05, so the model is admissible: the data do not contradict it. Not "true", not "confirmed". A model at p = 0.6 against the same right set would not be stronger evidence; a model at p = 0.6 against a weaker right set would be weaker.

Chi-square 11.72 on 8 degrees of freedom. The p-value is derived from these. Degrees of freedom for k sources fitted at rank r against n right populations are (k minus r) times (n minus 1 minus r): three sources at rank 2 against eleven outgroups gives 1 times 8 = 8. A chi-square of 11.7 on 8 degrees of freedom is close to what chance alone produces, which is what p = 0.164 says in one number.

f4 rank 2. Three sources require rank k minus 1 = 2. The rank test below shows what happened at ranks 1 and 0.

Min SNPs per f4: 121,880. Each f4-statistic is computed on the markers available for that particular quartet, which differ because ancient samples have gaps in different places. The record reports the smallest count, the contrast with the least data behind it.

Merged SNPs: 214,306. The coverage of the file against the panel. This number, not the tier, sets the floor on every SE in the table.

Jackknife blocks: 711. The genome cut into blocks of about 5 centimorgans, the model refitted leaving each out in turn, and the spread of those refits is the SE. Printed so a reader can see the errors were computed the standard way.

The right set#

#Right populationn
1Mbuti.DG4
2Ust_Ishim.DG1
3Kostenki141
4MA11
5Han.DG4
6Papuan.DG14
7Onge.DG2
8Karitiana.DG3
9Iran_N5
10Levant_N12
11EHG3

Eleven populations, listed in run order, each with its sample count. The first eight are the classic distant outgroups; the last three are era-appropriate additions that let the test tell a source with Iranian-related, Levantine-related or Eastern hunter-gatherer ancestry from one without. Mbuti.DG is first because the first right population is the base every f4-statistic is taken against.

This is the section that makes the model checkable. A p-value is only meaningful against the outgroups it was computed with; publish the weights without this list and nobody can evaluate them. The reasoning behind the choice is in How to choose sources and right populations.

The nested-model table#

Sources removedAnatolia_NYamnaya_SamaraWHGp-valueFeasible
none (full model)0.4870.3340.1790.164yes
WHG0.6120.388dropped0.0004yes
Yamnaya_Samara0.803dropped0.1971.1e-19yes
Anatolia_Ndropped0.4170.5836.3e-22yes

Every simpler model is rejected, which is the evidence that each source earns its place. Remove WHG and the model fails at p = 0.0004; remove either of the other two and it fails by an enormous margin. Had any row passed, that row's model is the one that should have been published, and the analyst treats the table exactly that way before proposing a model.

"Feasible" means every refitted weight stayed between zero and one. A "no" in that column is a model that could only fit by pushing a weight past 100%, which is the method's way of saying the removed source was carrying something real.

The rank test#

RankChi-squaredofp-value
211.7280.164
1208.4182.1e-34
02911.0300

The same question asked differently. Rank 2 is the published three-source model. Rank 1 would be any two-source model, and it fails at p = 10 to the minus 34. Rank 0, a single source, is off the scale. Three sources was the minimum, not a choice.

Warnings#

Note: SNP counts vary across f4 contrasts; the reported count is the minimum.

ADMIXTOOLS 2 prints its own diagnostics and the record keeps them verbatim. This one appears on almost every honest run and is the point explained under "min SNPs per f4" above. A warning present on every run is information, not a defect.

Reading your model#

Beside the record sits a written paragraph from the analyst who built the model. In this illustrative report it would read something like:

Your genome in this era resolves into three sources: a Neolithic farmer population sampled in Anatolia around 6500 BC, an early Bronze Age herder population from the Samara steppe, and a Mesolithic hunter-gatherer population of western Europe. Roughly half of the model is the farmer source, a third the steppe source and the remainder the hunter-gatherer source, which is the pattern seen across much of central and western Europe today. The two-source model without the hunter-gatherer source was rejected, so that share is required, not decorative. None of these labels is a place where anyone in your family lived: each is a reference group the model tests against, standing in for an ancestral population that was never sampled directly.

Every explanation Ancestrify publishes carries that last sentence in some form, because a table of weights answers "what" without answering "why", and a label alone invites the wrong reading.

The Reading view and the plain-text record#

Everything above is on screen in the report's Reading view and downloads as a plain-text file, one era or all eras, from the "Keep the record" bar, free at every tier. There is no PDF. The file is the same figures in the same order, with a header (target, era, panel, SNP counts, run id, date), the explanation, the fit block, the weights table with intervals, the right set with counts, the nested table, the rank test and the warnings. It is plain text so that it can be pasted to another analyst or kept beside the EIGENSTRAT bundle. Every field is defined in The model record explained.

What you can check yourself#

Without running anything:

  • Do the weights sum to 1.000? 0.487 + 0.334 + 0.179 = 1.000.
  • Is each Z the weight over its SE? 0.487 / 0.028 = 17.4. 0.334 / 0.031 = 10.8. 0.179 / 0.024 = 7.46.
  • Is each CI the weight plus or minus 1.96 SE? 0.487 minus 1.96 times 0.028 = 0.432.
  • Do the degrees of freedom match the counts? (3 minus 2) times (11 minus 1 minus 2) = 8.
  • Does the rank-2 row of the rank test match the fit block? 11.72, 8, 0.164. It should, because it is the same fit.
  • Is any right population a parent or close relative of a source? EHG is on the right and Yamnaya_Samara on the left. EHG is a component of Yamnaya, which is a real concern in the general case; the analyst's explanation should say why it was kept, typically because the model was also run with EHG removed and the weights did not move. If a report does not address it, ask.
  • Does every source clear the bar? p > 0.05, every |Z| > 3, every SE < 0.10. Yes, yes, yes.

Reproducing it#

Two routes. The Model Lab unlock (10 EUR, one time) lets you rerun this exact model, or any variation of it, on your own merged sample inside the report, up to 100 runs per rolling 24 hours. The same unlock includes the EIGENSTRAT bundle (.geno/.snp/.ind) the report was computed from, for running ADMIXTOOLS 2 on your own machine:

library(admixtools)
f2 <- f2_from_geno("path/to/bundle/prefix")
left  <- c("Target", "Anatolia_N", "Yamnaya_Samara", "WHG")
right <- c("Mbuti.DG", "Ust_Ishim.DG", "Kostenki14", "MA1", "Han.DG",
           "Papuan.DG", "Onge.DG", "Karitiana.DG", "Iran_N", "Levant_N", "EHG")
res <- qpadm(f2, left, right, target = "Target")
res$weights
res$rankdrop
res$popdrop

weights is the table above, rankdrop the rank test, popdrop the nested-model table. The figures should agree with the record to jackknife precision. The walkthrough is in Run your own qpAdm models. Before buying anything, the free AdmixTools 2 Lab runs the same function over the public panel.

A report you can read line by line and reproduce is the product. The qpAdm analysis starts at 29.99 EUR, every model is built and checked by a person, and the bar it is published against is the same at every tier.

From €29.99 · one-time
The tested version of this question
A qpAdm model composed, run and checked by hand against AADR v66, published with its p-value, every source's standard error and z-score, and the full right set, so the result can be argued with.
See the qpAdm analysis

References#

  • Haak, W. et al. (2015). Massive migration from the steppe was a source for Indo-European languages in Europe. Nature, 522, 207 to 211.
  • Harney, É., Patterson, N., Reich, D. & Wakeley, J. (2021). Assessing the performance of qpAdm: a statistical tool for studying population admixture. Genetics, 217(4), iyaa045.
  • Maier, R. et al. (2023). On the limits of fitting complex models of population history to f-statistics. eLife, 12, e85492.
  • Mallick, S. et al. (2024). The Allen Ancient DNA Resource (AADR): a curated compendium of ancient human genomes. Scientific Data, 11, 182.

Related posts

Learn qpAdm: the complete guide, from zero to defensible models
Learn qpAdm: the complete guide, from zero to defensible models

A structured learning path through everything qpAdm — what to read in what order, from the f4-statistics underneath to running models in R, choosing outgroups, reading results and knowing the method's measured limits.

3 min read
qpAdm best practices: the current checklist, with the numbers behind every rule
qpAdm best practices: the current checklist, with the numbers behind every rule

The definitive working checklist for qpAdm in 2026 — temporal stratification, right-set construction, lowest-rank-first search, composite feasibility and reporting standards — each rule carrying its measured justification from the 2021–2025 audit literature.

4 min read
Classic ADMIXTOOLS vs ADMIXTOOLS 2: why your qpAdm numbers differ, and how to match them
Classic ADMIXTOOLS vs ADMIXTOOLS 2: why your qpAdm numbers differ, and how to match them

Same model, different p-value: the real differences between original qpAdm and ADMIXTOOLS 2 — allsnps semantics, fudge_twice, f2 precomputation — and the settings that reproduce classic behaviour when you need to.

3 min read
Back to all stories
Ancestrify

Combining cutting-edge genomic science with rich historical records to map your ancestry across generations and continents.


© 2026 Ancestrify. All rights reserved. · Ancestrify is a trading name of Andi Thomaj, a sole trader registered in Tiranë, Albania · NUIS M61725001N
Card payments processed by POK Payments (RPay Ltd)VISAMASTERCARD