Run one genome through Eurogenes K13, Dodecad K12b and HarappaWorld and you get three different stories. A British kit runs about 44% North European in K12b and about 51% NE-Euro in HarappaWorld; its Mediterranean-adjacent score drops seven points moving the other way; a testing company's estimate disagrees with all three. The natural conclusion — someone must be wrong — is itself the wrong model. The calculators disagree for reasons that are structural, knowable, and worth understanding, because they are the same reasons no single calculator output should carry an argument on its own.
Reason one: components are anchored, not discovered#
A component is an allele-frequency profile learned by clustering that calculator's reference samples. Change the samples and the profile moves — even when the name stays. K12b's "North European" and HarappaWorld's "NE-Euro" are cousins, not twins: each is modal in northeast Europe, but each absorbs a slightly different mix of the deep ancestries (hunter-gatherer, farmer, steppe) because each was carved from a different panel. The seven-point swing in a British kit is those two anchors, not seven points of ancestry appearing or vanishing.
The practical rule follows directly: numbers never transfer between calculators. Only within-calculator comparisons — you versus a reference, you versus a relative, both run through the same tool — are on one yardstick.
Reason two: K decides how the cake is cut#
The same reference data carved into 9, 13 or 36 slices produces different-looking breakdowns for the same genome, necessarily. With K = 9, most of Europe is one profile; with K = 36, it is a dozen correlated ones, and your single northern ancestry scatters across them. Neither is more true; they are different resolutions of one underlying structure, with different noise floors. A component that exists at one K and not another (K13's East Med, say) cannot have a "true value" across calculators, because it is not the same object anywhere else.
Reason three: the calculator effect#
The bias with a name, coined during the 2012 dispute between the Eurogenes and Dodecad projects: a calculator describes the people whose samples trained it better than everyone else. If your population is well represented in the references, the components will fit you snugly; if it is absent, the optimiser shovels your ancestry into the nearest available profiles — confidently, because confidence is all it can express. The two project authors each claimed their methodology fixed it and the other's retained it; the durable lesson is simpler: every calculator has a home field, and you should know whether you are on it. The GEDmatch-era answer to "which calculator for my background" is really a map of home fields — we keep one here.
Reason four: different questions entirely#
Some disagreement is not about anchoring at all but about what is being estimated. A component calculator fits fixed allele-frequency profiles; a Global25 coordinate fit fits population averages in a PCA space; a testing company assigns chromosome segments against a proprietary panel and adds smoothing and priors of its own. These are three different estimands that all print percentages. Expecting them to agree is expecting the answers to three questions to match because they share a font — the full comparison is in G25 versus a 23andMe-style estimate.
So which number do you believe?#
None of them, in the sense the question intends — and all of them, read correctly:
- Believe the structure that survives every calculator. If farmer-versus-steppe-versus-forager proportions come out broadly similar in K13, K12b and a G25 era fit, that agreement is informative precisely because the tools share so little.
- Distrust everything that appears in only one. A component present in a single calculator, a trace percentage, a fine split between neighbouring profiles — these are the outputs most exposed to anchoring and noise, and no calculator attaches error bars to flag them.
- When it matters, use a method that tests. Disagreement between calculators cannot be settled by a fourth calculator. It can be settled — sometimes — by qpAdm, which proposes an explicit model against dated ancient sources and outgroups and returns a p-value that can reject it. That is a different kind of statement, and it is what the paid analysis exists for.
The disagreement, in other words, is not a scandal. It is the visible signature of what these tools are: optimisers over chosen references. Read three of them side by side with that in mind and they tell you more together than any one of them claims alone.
Terms used here are defined in the glossary.



