Global25 is the most widely used tool in amateur ancient-ancestry analysis and the most widely over-read. Its outputs look like conclusions — a ranked list of peoples, a tidy percentage breakdown — when they are measurements of position that require interpretation.
This explains what the method computes, what each output supports, and where confident-looking results go wrong.
The 25 numbers#
A Global25 coordinate is a label followed by 25 values. Each value is a position along one axis of genetic variation derived from a large reference set of ancient and modern samples.
Think of it as a location. Your genome sits somewhere in a 25-dimensional space; so does every reference population. Everything Global25 does is measure relationships between positions in that space.
The dimensions are ordered by how much variation they capture: the earliest ones separate the broadest population structure, the later ones progressively finer distinctions. This ordering is the reason the next section matters more than anything else on this page.
Scaled and unscaled — the mistake that silently ruins results#
Coordinates are published in two forms.
Unscaled treats all 25 dimensions equally. Scaled multiplies each dimension by its share of the variance, so the earlier, higher-variance dimensions carry proportionally more weight in a distance calculation.
⚠️ They are not interchangeable, and mixing them produces meaningless numbers that look completely normal. Compare a scaled row against an unscaled panel and you will still get a ranked list, still get plausible-looking distances, and still get an admixture breakdown summing to 100%. Nothing warns you. The arithmetic cannot detect the mismatch.
Before any comparison, confirm both sides are the same form. Scaled is the conventional choice for distance work; the rule that matters is consistency, not which one.
How a closest-population ranking works#
A distance ranking is one operation repeated across a panel: for each reference population, take the difference from your coordinate in each of the 25 dimensions, square each difference, add them up, take the square root. Sort ascending.
That is Euclidean distance, and its plainness is a virtue: it is instant, deterministic, and the same input always produces the same ranking. You can run it free in our G25 distance calculator.
Three things a small distance does not mean:
- It is not descent. Being close to a population is evidence about shared ancestry in general, never a claim that its members were your ancestors.
- It is not a like-for-like comparison. Reference populations are averages of individuals. Individuals scatter widely around their group's centre, so matching an average closely is not the same as matching anyone who actually lived.
- Rank order is not significance. Populations a little further down the list are often statistically indistinguishable from the top one. The gap between ranks 1 and 5 is usually far less meaningful than the ordering suggests.
Era scoping#
A ranking is only as meaningful as the panel it ran against, and panels are scoped by era. Being closest to a modern national average says where you sit among people alive today. Being closest to an Iron Age group says where you sit among the samples excavated from that window.
Comparing across eras is the genuinely interesting exercise — it shows how your position moves relative to different reference sets through time. Treating a single era's ranking as the answer is the common mistake.
How an admixture fit works#
An admixture model asks a different question: what weighted mixture of these source populations lands closest to your coordinate?
The search is a Monte-Carlo fit in the nMonte tradition. It distributes weight across the sources you offered, keeps combinations that reduce the gap between the mixture's combined position and yours, and reports the leftover gap as a fit distance.
Two properties define what this can and cannot support:
It always succeeds. Offer a panel containing nobody your ancestors ever met and it will still distribute 100% of the weight and report a fit distance. There is no p-value, no significance test and no mechanism for saying "these sources are wrong". The fit distance is the only signal that the answer is poor, and a plausible-looking breakdown from a badly chosen panel is the most common way to read a real number wrongly.
The panel does more work than the algorithm. Which sources you offer determines the answer far more than the fitting method does. Two analysts with different panels get different breakdowns from the same coordinate, and both are correct arithmetic.
A third, subtler property: nearby sources trade off almost freely. If two sources sit close together in the reference space, weight can move between them at almost no cost to the fit. Their individual percentages are much less stable than the total they share — which is why a breakdown should usually be read at the level of groupings rather than individual lines.
You can try this free in the admixture calculator, with curated era-scoped panels or your own pasted sources.
PCA#
A PCA scatter flattens many dimensions into a readable plot, positioning samples so the largest axes of variation become the chart's axes.
It is a projection, not a map. The axes are not ancestries, and closeness on a two-dimensional scatter can conceal separation that exists in the dimensions the plot discarded. Read a PCA as orientation — which cluster you fall near — rather than as measurement. You can try this free in the G25 PCA viewer: project your row onto curated era views beside published ancient and modern samples, or fit a custom PCA over rows you paste.
Where confident results go wrong#
An anachronistic panel. Sources drawn from the wrong period will still produce a fit. Check that every source could plausibly have contributed before the target existed.
A proxy standing in for the real source. Coordinate methods happily use a population near the true source without being it. The fit closes; the history is wrong.
A degraded input. A row that has been rounded, hand-edited, or converted from another calculator's output produces clean rankings that describe the conversion rather than your genome. Our authenticity check reads the numeric signature — though it audits the numbers, not the provenance.
Over-reading small differences. Two populations separated by a thousandth in distance are, for practical purposes, tied.
What Global25 is good for#
Given all those caveats, it is worth being clear that this is a genuinely good instrument for what it does: fast, reproducible, and excellent for exploration, ranking, comparison and visualisation. It simply has no mechanism for refusing a hypothesis.
When you need an answer that can come back negative — a test rather than a fit — the method for that is qpAdm, which works from allele-frequency statistics and returns a p-value. The full comparison is in qpAdm vs Global25.
If you do not yet have a coordinate row, start with how to get your Global25 coordinates. For the full report — distances, admixture and PCA across six eras — see our Global25 analysis. Terms are defined in the glossary.



