Ancestrify
All stories

global25

By Andi Thomaj
3 min read

Your closest populations by G25 distance: what the ranking means

What a Global25 distance ranking actually measures, why era choice changes everything, how close is 'close', and the four misreadings that turn a good tool into a wrong conclusion.

global25guidemethodology

  1. What is being computed
  2. Era first, always
  3. How close is "close"
  4. The four misreadings
  5. From ranking to understanding

The first thing everyone does with a fresh Global25 row is rank their closest populations — paste, sort, and stare at the list. It is the right first move: a distance ranking is the most assumption-free calculation in the whole G25 toolkit. It is also where the first misreadings happen, usually within a minute, because the list looks like an answer to "who am I descended from" and is actually an answer to something narrower and better.

What is being computed#

One number per reference population: the Euclidean distance between your row and that population's average across all 25 dimensions — square the 25 differences, sum, take the root. Sort ascending. That is the entire method: no mixture, no model, no fitting, which is exactly why it belongs first. Every result downstream of it (an admixture fit, a PCA position) has more assumptions layered on; the distance list is the raw neighbourhood.

Two technical facts shape everything about how to read it. References are averages — a population entry is the mean row of its member samples, so you are being compared with the centre of a group, not with any individual who lived. And both rows must be scaled — one unscaled row in a scaled comparison produces distances that mean nothing, with no error message; if your nearest modern population sits beyond roughly 0.08, suspect the paste before suspecting your ancestry.

Era first, always#

"Closest populations" is incomplete until you say when. Against the modern era, the ranking says where you sit among people alive today. Against the Iron Age panel, it says whose excavated world your coordinate lands in — populations that may have no living continuation at all. Our free distance tool runs six era panels (Late Bronze Age through Modern, 1,535 curated populations from 30,386 samples) separately on purpose: distances are only comparable within one era's panel, because "closest" is relative to who is in the room.

The cross-era read is the genuinely informative one. A Sicilian row, say, might sit nearest southern Italian averages in the modern era, near Greek colonial and Levantine-adjacent samples in Antiquity, and near Anatolian farmer descendants deeper still — a coherent story told three times at three depths. An incoherent sequence (close to a region in one era, nowhere near anything related in the adjacent ones) is usually the input, not the ancestry.

How close is "close"#

Rules of thumb for scaled distances to modern population averages, from our calibration work: a typical individual sits around 0.011 from their own country's average, and 0.016–0.032 from their own country's other individuals' typical spread. Roughly: under ~0.02 reads as "this average describes me well"; 0.02–0.04 as "same broad region, not this exact group"; beyond ~0.05–0.06 as "not my neighbourhood". Ancient-era panels run systematically wider — every living person is far from every Bronze Age average, because three thousand years of admixture sit in between; comparing your ancient-era distances against each other is meaningful, comparing them against your modern-era numbers is not.

And within any single list: rank order is not significance. Entries separated by a few thousandths are statistically indistinguishable — the top cluster is the finding, the exact winner is not.

The four misreadings#

  • "Closest = descended from." Distance is resemblance. A population can sit near you because you share ancestry — or because it is itself a mixture that happens to average out near your point. The ranking cannot tell those apart; nothing distance-based can.
  • "Closer than my expected population = surprise ancestry." Averages again: an admixed individual routinely lands nearer some third population that sits between their real sources. The PCA view exposes this instantly; the mixture question belongs to an admixture fit, and the tested version of it to qpAdm.
  • "A 0.003 gap between rank 1 and rank 4 means something." It is noise. See above.
  • "My relative's list looks different, so something is wrong." Individual rows scatter around family centres; two siblings' rankings differing in the tail is expected. (Two siblings sitting far apart, by contrast, is a data problem — the authenticity check is for exactly that.)

From ranking to understanding#

The list is triage, and good analysis keeps it in that role: neighbourhood first, then an era-scoped mixture model whose sources should be able to explain the neighbourhood, then — when a claim needs testing — a formal model. That is the exact sequence the paid Global25 analysis walks with a curated panel and a written rationale, and every step short of the last one is free in the browser: distances, admixture, PCA. No coordinates yet? Start here.

€29.99 · one-time
The worked version of this analysis
Distances, admixture models and PCA across six eras against 1,535 curated populations, every source panel published in full, with Notable Matches included free.
See the Global25 analysis

Terms used here are defined in the glossary.


Related posts

What is a good G25 fit distance? The number everyone reads wrong
What is a good G25 fit distance? The number everyone reads wrong

The fit distance measures how far a mixture still sits from your coordinate — not whether the model is right. Calibrated bands for reading one, why lower stops being better, and the checks that actually validate a model.

4 min read
G25 scaled vs unscaled coordinates: which to use, how to tell them apart
G25 scaled vs unscaled coordinates: which to use, how to tell them apart

Every Global25 row exists in two forms, and mixing them is the hobby's most common silent error. What scaling does mathematically, which form each tool expects, and the tell-tale signs of a mixed comparison.

3 min read
How accurate is G25? What the coordinate can and cannot resolve
How accurate is G25? What the coordinate can and cannot resolve

Global25 is a projection, not a test — so 'accuracy' has layers: the coordinate's own fidelity, the references it is compared against, and the models built on top. An honest audit of all three.

3 min read
Back to all stories
Ancestrify

Combining cutting-edge genomic science with rich historical records to map your ancestry across generations and continents.


© 2026 Ancestrify. All rights reserved. · Ancestrify is a trading name of Andi Thomaj, a sole trader registered in Tiranë, Albania · NUIS M61725001N
Card payments processed by POK Payments (RPay Ltd)VISAMASTERCARD