Ancestrify
All stories

admixture

By Andi Thomaj
5 min read

What is an admixture calculator? How ancestry percentages are actually computed

Every admixture calculator — GEDmatch's classics, Global25 fits, testing-company estimates — is an optimiser that cannot say no. How the three families work, what the percentages mean, and the questions a calculator can and cannot answer.

admixtureguidemethodologypopulation-genetics

  1. The three families
  2. What every family shares: the optimiser cannot say no
  3. What a percentage actually means
  4. The questions a calculator answers well
  5. The questions it cannot answer
  6. If you want to try one now
  7. References

An admixture calculator takes a genome — or a coordinate standing in for one — and returns a list of percentages: so much of this component, so much of that source, summing to 100%. The percentages look like a measurement. They are actually the solution to an optimisation problem, and almost everything people get wrong about calculators follows from not knowing which problem was solved.

This is the plain version of how the three families of calculator work, what a percentage does and does not mean, and how to decide whether a calculator or a formal method answers your question.

The three families#

Component calculators (ADMIXTURE-style). The classic GEDmatch projects — Eurogenes, Dodecad, HarappaWorld, MDLP, puntDNAL — belong to this family. Each calculator holds K allele-frequency profiles, called components, that were learned once by clustering a reference panel with software in the ADMIXTURE tradition. When you run your kit, the tool finds the mixing proportions of those fixed profiles that best explain your genotypes. K13 means thirteen components; K36 means thirty-six. The components carry geographic names — "North Atlantic", "Gedrosia", "Baloch" — but they are statistical constructs built from whichever samples trained that calculator, not observed ancient populations.

Coordinate calculators (nMonte-style). The Global25 ecosystem works differently. Your genome is first reduced to a 25-number coordinate in a principal-component space, and the calculator then searches for the weighted mixture of reference coordinates whose combined point sits closest to yours. The search is a Monte-Carlo fit in the nMonte tradition, and it reports a fit distance — how far the best mixture still sits from your point. Our own free Global25 admixture calculator is this family: you choose a curated, era-scoped source panel, and the tool fits your row against it in your browser.

Testing-company estimates. 23andMe's Ancestry Composition, AncestryDNA's ethnicity estimate and their peers are proprietary members of the first family with two differences: the reference panels are much larger and non-public, and the assignment is usually made segment by segment along the chromosomes rather than genome-wide. The same interpretive rules apply, which is why an ancient-DNA test is a different product, not a better version of the same one.

What every family shares: the optimiser cannot say no#

Whatever the machinery, the question being answered is the same: given these references, which proportions come closest? Note the two things that are not being asked — whether the references are the right ones, and whether "closest" is close enough to mean anything.

A component calculator with K = 6 will distribute every genome on Earth across its six profiles, because that is the only thing it can do. A coordinate fit will always return a best combination, including from a source panel containing nobody your ancestors ever met. Give a calculator a Yoruba genome and only European references and it returns a confident European breakdown, with no error message, because nothing in the arithmetic knows something is wrong.

This is not a defect; it is what an optimiser is. But it is the single most important fact about every percentage you will ever see from one. A method that cannot fail is a method whose passing means nothing by itself — the surrounding choices, mostly the reference panel, decide whether the number deserves any weight. The contrast is qpAdm, which tests a proposed model against outgroups and rejects it outright when the data are incompatible; that difference is the entire subject of is qpAdm worth it.

What a percentage actually means#

Take a typical line: North_Atlantic 41.2%. The honest reading is: of the K fixed profiles this calculator was trained with, assigning 41.2% of your alleles to the profile someone named "North Atlantic" is part of the combination that best explains your genotypes. Three things follow.

The name is a label, not a place. Components are named by the person who built the calculator, after where the component peaks in the training samples. "Gedrosia" in Dodecad K12b and "Baloch" in HarappaWorld describe nearly the same statistical object under two names — and neither is a population anyone has excavated.

The number moves between calculators. The same British genome scores around 44% North European in Dodecad K12b and around 51% NE-Euro in HarappaWorld, not because either is broken but because similarly named components are anchored differently. Why calculators disagree walks through the mechanics.

Small numbers are usually noise. A 0.9% East Asian sliver in a European profile is far more often the optimiser rounding noise onto the nearest available profile than a great-great-ancestor. No calculator attaches a standard error to a component, so nothing in the output distinguishes the two — that absence is structural, not an oversight you can read around.

The questions a calculator answers well#

Used for what it is, a calculator is a genuinely good instrument:

  • Resemblance. "Which references does my genome lean toward" is exactly the question the optimiser solves. For the sharpest version of it, a plain distance ranking answers without any mixing at all.
  • Comparison across people. Two kits run through the same calculator are measured on the same yardstick; differences between them are informative even when the absolute numbers are not.
  • Deep structure, in outline. Run against ancient references — ancient versus modern panels is the choice that matters — the big strokes (farmer versus hunter-gatherer versus steppe) are real, robust signals that every method recovers.
  • Exploration. Trying five source panels in an evening teaches you more about what shapes the numbers than any single result does. That is what free, instant tools are for.

The questions it cannot answer#

  • "Is this component real?" No error bars, no test. A formal method exists for exactly this; it returns a p-value and a z-score per source and will reject a model that does not hold.
  • "Am I descended from X?" Percentages describe resemblance to references. Descent is a different claim, and no percentage — however large — establishes it.
  • "Which of two close sources is my real ancestor?" Two references that sit close together trade weight almost freely; the split between them is the least stable number on the screen.
  • "What does my 2% mean?" Below a few percent, usually nothing. Treat trace components as unconfirmed until a method with uncertainty attached has looked at them.

If you want to try one now#

Everything needed to explore runs free, in the browser, with no account: the Global25 admixture calculator with curated era panels, the closest-populations distance tool, and the PCA viewer to see the space your coordinate lives in. If you have raw DNA but no coordinates yet, how to get Global25 coordinates covers the official route. And when a percentage starts carrying an argument — a claim you would defend to someone who knows the methods — that is the moment for a formal qpAdm model, because it is the moment you need numbers that can say no.

€29.99 · one-time
The worked version of this analysis
Distances, admixture models and PCA across six eras against 1,535 curated populations, every source panel published in full, with Notable Matches included free.
See the Global25 analysis

Terms used here are defined in the glossary.

References#

  • Alexander, D. H., Novembre, J. & Lange, K. (2009). Fast model-based estimation of ancestry in unrelated individuals. Genome Research, 19(9), 1655–1664.
  • Lawson, D. J., van Dorp, L. & Falush, D. (2018). A tutorial on how not to over-interpret STRUCTURE and ADMIXTURE bar plots. Nature Communications, 9, 3258.
  • Novembre, J. et al. (2008). Genes mirror geography within Europe. Nature, 456, 98–101.

Related posts

Every number in your qpAdm report: the model record explained
Every number in your qpAdm report: the model record explained

What the full qpAdm model record in an Ancestrify report means — chi-square, degrees of freedom, f4 rank, SNP counts, jackknife blocks, 95% confidence intervals, the nested-model table and the rank test — and how to read the plain-text download.

11 min read
How to read qpAdm results: p-value, Z-score and SE
How to read qpAdm results: p-value, Z-score and SE

A plain reading guide to the three numbers in every qpAdm result — what the p-value tests, what a standard error bounds, what a Z-score rules out — with worked examples and the mistakes that make a passing model wrong.

4 min read
How accurate are admixture calculators? An honest audit
How accurate are admixture calculators? An honest audit

Accuracy has three different meanings for an ancestry calculator, and the tools do well on exactly one of them. Where percentages are trustworthy, where they are noise, and how to tell which regime you are in.

4 min read
Back to all stories
Ancestrify

Combining cutting-edge genomic science with rich historical records to map your ancestry across generations and continents.


© 2026 Ancestrify. All rights reserved. · Ancestrify is a trading name of Andi Thomaj, a sole trader registered in Tiranë, Albania · NUIS M61725001N
Card payments processed by POK Payments (RPay Ltd)VISAMASTERCARD