The first time a qpAdm model comes back with p = 0.003, most people assume something went wrong. Nothing did. A rejection is the method working: the mixture you proposed is incompatible with the data, and qpAdm is the only common ancestry tool that can tell you so. The useful question is not "how do I make it pass" but "what is it telling me". This post is the diagnostic list an analyst runs through for every rejection in a qpAdm analysis, and it is the same list a customer should follow in the Model Lab.
What p < 0.05 means, and does not#
The p-value tests one thing: whether the target's f4-statistics against the right populations can be written as a weighted combination of the sources' f4-statistics. Below 0.05, they cannot, at conventional confidence. The model is rejected.
What it does not mean:
- It does not mean the sources are unrelated to the target. A rejected three-source model may be missing a fourth source that accounts for 5% of the genome; the other 95% may be exactly as proposed.
- It does not mean the weights are wrong in any particular direction. They are simply not worth reading, because the model they belong to does not hold.
- It does not mean a higher-coverage file would pass. Coverage mostly sets the standard errors, not the p-value. A well-covered file rejects a wrong model more confidently, not less.
- It is not a probability that the model is false. It is a compatibility statement given this right set. Change the right set and the same left list can pass or fail.
The threshold itself is a convention. A model at p = 0.04 and one at p = 0.06 are almost the same evidence; the bar exists so that a publishing rule can be stated once and applied identically. The reading guide covers the other misreadings.
The five usual causes#
1. A missing source#
The most common cause by far. The right set contains a population that shares drift with the target in a way none of the sources explain, and the f4 pattern is inconsistent. In a European model this is typically the steppe source: Anatolia_N + WHG alone rejects for almost any living European because MA1, EHG and Karitiana.DG on the right detect Ancient North Eurasian ancestry the two sources lack. The rejection is the right set doing exactly what it is for.
Tell: the rejected model has a large chi-square against a modest number of degrees of freedom, and adding one well-chosen source drops it dramatically.
2. The wrong proxy#
The source is the right kind of population but the wrong sample of it. Anatolia_BA standing in for Anatolia_N carries an Iranian-related layer that Neolithic farmers lacked; Steppe_MLBA in place of Yamnaya_Samara carries a farmer-related layer the earlier steppe did not. The model may pass with distorted weights, or reject outright when the right set can see the extra layer.
Tell: the p-value moves a great deal when one label is swapped for a close relative, while the number of sources stays the same. How to choose sources and right populations goes through the label pairs that matter most.
3. A right population too close to a source#
If a right population shares recent drift with one source and not the others (EHG on the right with Yamnaya_Samara on the left; Natufian on the right with Levant_N on the left), the f4-statistics involving that pair are structured by the relationship, not by the target's ancestry, and the model fails for a reason that has nothing to do with the question asked.
Tell: the model passes when that one right population is removed, and the weights barely move. That is a sign the right population was contaminating the test rather than detecting anything.
4. Low coverage, which inflates SE while p can still be fine#
This one is included because people blame it for rejections, and it usually is not the cause. A sparse file produces large standard errors and small Z-scores; it does not, by itself, drive the p-value down. A model on a 90,000-marker file can sit at p = 0.4 with every source at |Z| < 2. That is a different failure: not a rejection, but a model too imprecise to publish. Ancestrify's bar requires every SE < 0.10 and every |Z| > 3 alongside p > 0.05, so a low-coverage file can fail the bar on the error terms while the p-value looks comfortable.
Tell: p above 0.05, standard errors above 0.10, and the free Raw DNA File Check reporting thin overlap with the panel.
5. A target that is itself a mixture of mixtures#
Some genomes are not well described by two to four ancient sources against any right set, because their history involves several admixture events between already-mixed populations. Every proposed model is a little wrong, and a sufficiently strong right set rejects all of them. This is a real finding about the target, not a failure of technique.
Tell: many different plausible models all sit at p between 0.001 and 0.03, none dramatically better than the others, and the four-source models have rank tests that pass one rank too early.
The nested-model table#
Before treating a rejection as informative, check that the passing models around it are honest. For every model, qpAdm can refit each simpler model obtained by removing one or more sources. An illustrative table:
| Sources removed | Anatolia_N | Yamnaya_Samara | WHG | p-value | Feasible |
|---|---|---|---|---|---|
| none (full model) | 0.487 | 0.334 | 0.179 | 0.163 | yes |
| WHG | 0.612 | 0.388 | dropped | 0.0004 | yes |
| Yamnaya_Samara | 0.803 | dropped | 0.197 | 1.1e-19 | yes |
| Anatolia_N | dropped | 0.417 | 0.583 | 6.3e-22 | yes |
Every simpler model rejects, so every source earns its place. When a row in this table passes, the simpler model is the one that should be published, and the extra source was an artefact of the fit. This is the check that stops "add sources until it passes" from working.
Negative weights#
qpAdm does not constrain weights to lie between zero and one. A negative weight means the best linear fit pushes one source below zero and another above one, and it is nearly always a sign that the sources are badly chosen. Illustrative:
p-value: 0.09
Anatolia_N 0.531 SE 0.041 Z 12.9
Yamnaya_Samara 0.372 SE 0.044 Z 8.5
WHG 0.138 SE 0.036 Z 3.8
CHG -0.041 SE 0.038 Z -1.1
The p-value passes, and a careless reader might report "0% CHG". The honest reading is that CHG is not a source here, the model is really a three-source model, and the fourth term is absorbing noise. Run the three-source nested model; if it passes, publish that. A model is marked infeasible in the record when any weight falls outside zero to one, and no Ancestrify model with an infeasible weight is published.
|Z| below 3 with p fine#
The quieter cousin of a rejection. The model passes, but one source's weight is within a few standard errors of zero:
p-value: 0.27
Anatolia_N 0.612 SE 0.033 Z 18.5
Yamnaya_Samara 0.335 SE 0.035 Z 9.6
Levant_N 0.053 SE 0.031 Z 1.7
At |Z| = 1.7, the Levant_N weight is not distinguishable from zero at the confidence the method works at. Ancestrify's bar requires every source above |Z| = 3, which is why this model would not be published as it stands. The correct step is the same as for a negative weight: test the nested two-source model. If it passes, publish it; if it rejects, the Levant_N source is needed but the file cannot pin its weight, and the honest report says so.
What to try next, in order#
- Read the right set first. Is any right population a parent, descendant or close relative of a source? Move it or drop it, rerun, and see whether the weights move. If they do not and the model now passes, the right population was the problem.
- Check chronology. No source should postdate the target's era, and nothing on the right should be a mixture of things on the left.
- Try one more source, chosen by asking which right population the target shares unexplained drift with. If EHG and MA1 are the loud ones, the missing source is steppe-related. If Han.DG is, it is East Asian-related. One source, not three.
- Swap proxies before adding sources. Anatolia_N for Iberia_N, Yamnaya_Samara for Steppe_MLBA, Iran_N for CHG. Rerun with the same right set each time so that p-values are comparable.
- Run the nested models on anything that passes. Publish the simplest model that survives.
- Rotate as a diagnostic, not a search. Move each candidate source to the right in turn and confirm the model still needs it.
- Check the file. If SEs are all above 0.10, no search will fix it; the file sets that floor.
What is not on the list: rerunning combinations until one passes. Run enough models and some will clear any threshold by chance, and the best-scoring model is frequently not the right one. Harney and colleagues (2021) documented this failure mode, and it is why Ancestrify built an automated rotator, measured its false-discovery rate, and removed it.
When to stop#
A rejected model is information. If, after the steps above, every plausible two- to four-source model rejects against a strong right set, the finding is that the target's history is not well-described by the sources available in the panel, and the report should say so rather than publish the least-bad model. Some genomes, at some coverage, with the ancient samples currently excavated, cannot be modelled to the bar. That is an honest result.
The stopping rule for Ancestrify is stated once and applied at every tier: a model is published only when p > 0.05, every source has |Z| > 3 and every SE < 0.10, in every era. A file that cannot meet that is told so before purchase where the file check can see it, and afterwards the report says which era could not be resolved and why.
How Ancestrify handles it#
Every model is composed and checked by a person, not rotated by a script. For each candidate source, the analyst exhausts the plausible AADR proxy labels, runs the nested and rank tests on everything that passes, and publishes only a model that clears the bar with its complete record: weights, SE, Z, 95% CI, the ordered right set with sample counts, the nested-model table, the rank test and the tool's warnings, plus a written explanation of why the model looks the way it does. The record is described in The model record explained, and a sample is on the demo.
After publication, the Model Lab lets you run the diagnostic list above yourself, on your own merged sample, and see the rejections for what they are. If you want the whole process done and defended for you first, the qpAdm analysis starts at 29.99 EUR, and the bar is the same at every tier.
References#
- Harney, É., Patterson, N., Reich, D. & Wakeley, J. (2021). Assessing the performance of qpAdm: a statistical tool for studying population admixture. Genetics, 217(4), iyaa045.
- Maier, R. et al. (2023). On the limits of fitting complex models of population history to f-statistics. eLife, 12, e85492.
- Patterson, N. et al. (2012). Ancient admixture in human history. Genetics, 192(3), 1065 to 1093.



