Ancestrify
All stories

qpadm

By Andi Thomaj
11 min read

Every number in your qpAdm report: the model record explained

What the full qpAdm model record in an Ancestrify report means — chi-square, degrees of freedom, f4 rank, SNP counts, jackknife blocks, 95% confidence intervals, the nested-model table and the rank test — and how to read the plain-text download.

qpadmmethodologyadmixturepopulation-geneticsguide

  1. Where the record lives
  2. The analyst note
  3. The fit block
  4. The sources ledger
  5. Right / outgroup populations
  6. Nested simpler models
  7. Rank test
  8. Tool warnings
  9. The download
  10. Why publish all of this
  11. Questions
  12. References

A qpAdm result is usually shown as four numbers: a p-value, and a weight, standard error and z-score per source. Those four decide whether a model is worth reading, and how to read them is its own guide. But the software computes a great deal more than four numbers, and every one of the others is something a second analyst would ask for before accepting the model. Since August 2026 every Ancestrify qpAdm report publishes all of them, on screen and as a plain-text file, at every tier and at no extra cost. This guide reads that record from top to bottom.

Where the record lives#

Open a qpAdm report, and the Ancestry tab has two views: Atlas, the map and the story, and Reading. The Reading view holds three things for the era on screen:

  1. Reading your model — a written explanation, composed for your order by the analyst who built the model and reviewed before publication, of why your genome resolved into these sources in these proportions.
  2. The model record — the complete output of the run the published model came from.
  3. Keep the record — a download of that record as plain text, for the era on screen or for every era at once.

The same view is on the live demo with example data, so you can see the shape of it before buying anything. What follows is the record, section by section, with the figures from a real published two-era model, with every identifying detail removed.

The analyst note#

The note is prose, not statistics, and it exists because a table of weights answers "what" without answering "why". It says what each source population is — the ancient people behind the label — why the proportions look the way they do for this declaration and family history, how the two eras agree with each other, and, most importantly, what a label does not mean.

That last part matters more than it sounds. A model that assigns 18% of someone's ancestry to a source called Arabian Peninsula is not saying an ancestor came from Arabia. It is saying that of the reference populations in the panel, that one explains that part of the genome best. The method cannot place ancestry within a region, and the note says so in plain words, because the label alone invites exactly the wrong reading. A source population is a reference group the model tests against, never a statement about where a particular person lived. Every note we publish carries that sentence in one form or another.

The fit block#

Seven tiles sit at the top of the record. Read them in this order.

p-value. The one number everyone knows. It tests whether the pattern of shared drift in your genome is compatible with the proposed mixture of sources, measured against the outgroups; above 0.05 the model is admissible, below it the model is rejected and the weights should not be read. The example model sits at p = 0.2113. It is not the probability the model is true, and a higher value past the bar is not stronger evidence — see the reading guide for the misreadings.

Chi-square and degrees of freedom. The p-value is derived from these two. qpAdm fits a matrix of f4 statistics between the sources and the outgroups and asks how far the observed matrix sits from the closest matrix of the required rank; that distance is the chi-square, here 9.618. The degrees of freedom count how many independent contrasts the model had to fit, and depend on how many sources and outgroups there are. With the target plus k sources on the left and n populations on the right, qpAdm tests a k × (n − 1) matrix of f4 statistics at rank r, and the degrees of freedom are (kr) × (n − 1 − r). Three sources fitted at rank 2 against ten outgroups gives (3 − 2) × (9 − 2) = 7 degrees of freedom, which is what the example prints; more outgroups mean more degrees of freedom and a harder test. A chi-square of 9.6 on 7 degrees of freedom is close to what chance alone would produce, which is what p = 0.21 says in one number.

f4 rank. For a mixture of k sources, the f4 matrix must have rank k − 1: three sources, rank 2. The record prints the rank the published model was fitted at, and the rank test lower down (see below) shows what happened at every lower rank. If a two-source model had fitted, rank 1 would have passed and the third source would not be there.

Min SNPs per f4. qpAdm computes each f4 statistic on the markers available for that particular quartet of populations, which differ because ancient samples have gaps in different places. The record reports the smallest of those counts, here 96,302, because that is the contrast with the least data behind it and therefore the honest figure. A record that reported the average, or the size of the merged file, would be flattering itself.

Merged SNPs. How many of your markers survived the merge with the ancient panel: 176,468 for this 23andMe kit. This is the coverage that sets the floor on every standard error in the table. It is a property of the file, not of the analyst: a whole-genome kit merges to far more positions and returns visibly tighter intervals, which is the whole argument of uploading a whole-genome VCF.

Jackknife blocks. Standard errors in qpAdm come from a block jackknife: the genome is cut into blocks (by default about 5 centimorgans each), the model is refitted leaving each block out in turn, and the spread of those refits is the error. The record prints how many blocks the genome was cut into, typically about 700. It is there so that a reader can see the errors were computed the standard way, on the standard scale.

The sources ledger#

One row per source, and this is where the four familiar numbers gain some company.

#SourcePanel labelnShareWeightSEZ95% CI
01Gandharan SwatPakistan_Barikot_H.AG444.0%0.43980.06277.020.317 to 0.563
02Arabian PeninsulaBedouinB.DG2118.3%0.18340.03934.670.106 to 0.260
03Peninsular South IndianIrula.SG1037.7%0.37680.036710.260.305 to 0.449

Source and panel label. The catalog name is ours; the panel label is the exact population in the Allen Ancient DNA Resource the model was computed with, suffix and all. The two are printed side by side so the model can be reproduced by anyone with the same panel, and so that a curated name never hides which archaeological context was actually used. Every catalog population has a page in the ancestry directory; the labels are explained in the AADR explainer.

n. How many individuals the source population contains. A source of four samples is a thinner reference than one of twenty-one, and the record says so rather than leaving it to be inferred.

Share and weight. The same number twice: the weight is the raw proportion the model estimated, and the share is that weight as a rounded percentage. They are printed together because a headline percentage should always be traceable to the raw figure it was rounded from.

SE and Z. The standard error of the weight and the weight divided by it. A source at |Z| of 7 sits seven standard errors from zero; the reading guide covers why a source at |Z| of 1 is indistinguishable from nothing whatever its percentage says.

95% CI. New in the published record: the weight plus or minus 1.96 standard errors, the range within which the true proportion would plausibly sit. Read it before the percentage. "44%" invites a precision the data do not have; "between 32% and 56%" is what the model actually established. An interval that reaches down to zero is a source the model cannot certify, and our publish bar (|Z| above 3 for every source) is designed so that no published interval does.

Right / outgroup populations#

The right set is listed in the order the model used it, each with its sample count, because the first right population is the base every f4 statistic is taken against and the count of each one decides how much power it has to reject a wrong source. A right set of well-sampled, genuinely distant populations is what makes a passing model mean something; Understanding qpAdm explains the choice. The example used ten: Mbuti, Han, Karitiana, Tianyuan, Kostenki14, Neolithic Anatolia, Yamnaya, Bronze Age Shahr-i-Sokhta, Bronze Age Gonur and Bronze Age Israel.

Nested simpler models#

This is the table most reports never show and the one a sceptical reader wants first. For every way of removing sources from the published model, the record lists the refitted weights of the sources that remain, the p-value of that simpler model, and whether its weights stayed feasible (inside zero and one).

Sources removedGandharan SwatArabian PeninsulaPeninsular South Indianp-valueFeasible
none (full model)0.4400.1830.3770.2113yes
Peninsular South Indian1.084droppeddropped2.3e-37no
Arabian Peninsula0.687dropped0.3133.2e-13yes
Gandharan Swatdropped0.4160.5843.3e-24yes

Read the p-value column. Every simpler model is rejected by an enormous margin, which is the evidence that each of the three sources is earning its place: remove any one and the data no longer fit. The row that also reads "no" under feasible is a model that could only fit by pushing a weight past 100%, which is the method's way of saying the removed source was carrying something real. When a simpler model passes in this table, the simpler model is the one that should have been published, and our analysts treat the table exactly that way before a model is proposed.

Rank test#

Below it, the rank test shows the same question asked differently: the chi-square, degrees of freedom and p-value of the f4 matrix at every rank from the published one down to zero.

RankChi-squaredofp-value
29.61870.2113
1473.580161.2e-90
04717.319270

Rank 2 is the published three-source model. Rank 1 would be any two-source model, and it fails at p = 10⁻⁹⁰; rank 0, a single source, is off the scale. This is the record's proof that three sources was the minimum, not a choice.

Tool warnings#

ADMIXTOOLS 2 prints its own diagnostics and the record keeps them verbatim. The one you will see on almost every model is that SNP counts vary across f4 contrasts and the reported count is the minimum, which is the point explained under "Min SNPs per f4" above. A warning that is present on every honest run is information, not a defect, and hiding it would be the defect.

The download#

The Keep the record bar at the bottom of the Reading view writes the whole record for the era on screen, or for every era, as a plain-text file named ancient-origins-order-55-v1-classical-antiquity.txt (your order number, the version, the era). It contains the same figures as the screen: the header (target, era, panel, merged and minimum SNPs, run id, run date), the analyst note, the fit block, the sources ledger with confidence intervals, the ordered right set with counts, the nested-model table, the rank test, the warnings, and a short note on how to read the units. It is plain text so that it can be pasted into a message to another analyst, attached to a forum post, or kept beside the EIGENSTRAT bundle that the Model Lab unlock provides for re-running the model yourself.

There is no PDF and never will be; a plain-text ledger of every figure is a more useful thing to keep than a formatted page, and it is what other analysts actually ask for.

Why publish all of this#

Because a number you cannot check is a claim, not a result. The four headline figures let a reader judge a model; the full record lets a reader reproduce it, or argue with it, without asking us for anything. Every published version of a report, including any refined re-analysis, carries its own complete record, so two versions can be compared line by line. The buyer's guide puts this in a checklist; the short version is that a qpAdm product which cannot show you this table has not run the method it is named after.

Questions#

Is the model record an extra charge? No. It is part of every published qpAdm report at every depth tier, and the plain-text download is free.

What is the 95% confidence interval? The weight plus or minus 1.96 standard errors: the range the true proportion would plausibly fall in. Read it before the percentage.

What does "dropped" mean in the nested-model table? That the source in that column was removed for that row's refit; the remaining columns show how the other sources reshuffled, and the p-value shows whether the simpler model survived.

Can I get the record as a file? Yes: the "Download (.txt)" button for the era on screen, or "Download all eras", from the Reading view.

Does the Reading view change my percentages? No. It is the same model, with the numbers behind it and the explanation of them. The Atlas view and the videos are unchanged.

From €29.99 · one-time
The tested version of this question
A qpAdm model composed, run and checked by hand against AADR v66, published with its p-value, every source's standard error and z-score, and the full right set, so the result can be argued with.
See the qpAdm analysis

References#

  • Haak, W. et al. (2015). Massive migration from the steppe was a source for Indo-European languages in Europe. Nature, 522, 207–211. (Supplementary Information 10: the qpAdm method.)
  • Harney, É., Patterson, N., Reich, D. & Wakeley, J. (2021). Assessing the performance of qpAdm: a statistical tool for studying population admixture. Genetics, 217(4), iyaa045.
  • Maier, R., Flegontov, P., Flegontova, O., Işıldak, U., Changmai, P. & Reich, D. (2023). On the limits of fitting complex models of population history to f-statistics. eLife, 12, e85492.
  • Patterson, N. et al. (2012). Ancient admixture in human history. Genetics, 192(3), 1065–1093.

Related posts

How to read qpAdm results: p-value, Z-score and SE
How to read qpAdm results: p-value, Z-score and SE

A plain reading guide to the three numbers in every qpAdm result — what the p-value tests, what a standard error bounds, what a Z-score rules out — with worked examples and the mistakes that make a passing model wrong.

4 min read
What is an admixture calculator? How ancestry percentages are actually computed
What is an admixture calculator? How ancestry percentages are actually computed

Every admixture calculator — GEDmatch's classics, Global25 fits, testing-company estimates — is an optimiser that cannot say no. How the three families work, what the percentages mean, and the questions a calculator can and cannot answer.

5 min read
qpAdm analysis tutorial: from a raw DNA file to a model with a p-value
qpAdm analysis tutorial: from a raw DNA file to a model with a p-value

A step-by-step qpAdm tutorial: check your raw file, merge it into AADR v66, choose sources and outgroups, run it in the browser or in R, and read the result.

14 min read
Back to all stories
Ancestrify

Combining cutting-edge genomic science with rich historical records to map your ancestry across generations and continents.


© 2026 Ancestrify. All rights reserved. · Ancestrify is a trading name of Andi Thomaj, a sole trader registered in Tiranë, Albania · NUIS M61725001N
Card payments processed by POK Payments (RPay Ltd)VISAMASTERCARD