Ancestrify
All stories

global25

By Andi Thomaj
3 min read

How accurate is G25? What the coordinate can and cannot resolve

Global25 is a projection, not a test — so 'accuracy' has layers: the coordinate's own fidelity, the references it is compared against, and the models built on top. An honest audit of all three.

global25methodologyguide

  1. Layer one: the coordinate itself
  2. Layer two: the comparisons
  3. Layer three: the models
  4. Where G25 sits among the instruments

"Is G25 accurate?" gets asked with three different worries behind it: whether the 25 numbers faithfully describe a genome, whether the comparisons built on them are sound, and whether the percentages people derive from them are true. Those are different layers with different answers — and the honest audit is more interesting than either the fan version ("it matches papers!") or the dismissal ("hobbyist toy").

Background if needed: what the coordinates are and where they come from.

Layer one: the coordinate itself#

A Global25 row is a projection of chip-scale genotypes onto 25 fixed PCA axes. Within its design, it is a stable, repeatable measurement: the same genome projects to essentially the same point regardless of which vendor's chip produced the file, siblings land near each other, and your row never changes as references grow. The known sensitivities are inputs, not arithmetic: thin or corrupted files project noisily (pre-flight with a file check), simulated rows are not projections at all, and an unscaled/scaled mix-up invalidates everything downstream while looking like numbers.

The structural limit is compression. Twenty-five dimensions retain the broad and much of the fine structure of West Eurasia — the space was built to — but two populations distinct in allele-frequency statistics can sit close in G25 space. Drifted isolates and thinly sampled regions fold worst. Nothing downstream can recover what the projection folded; that is the ceiling every G25 result lives under, and the reason formal methods work from the genotypes directly.

Layer two: the comparisons#

Distances and PCA positions inherit the coordinate's fidelity plus one more dependency: the reference panel. A closest-populations ranking is exactly as good as who is on the list — a missing population cannot rank, and its absence is invisible. Averages built from two or three individuals wobble (an average of N carries roughly 1/√N of individual noise, which is why our panel curation enforces member floors); era mixing blurs everything. Calibration numbers from our own reference work give the scale of normal: a typical individual sits around 0.011 from their own country's modern average, and single individuals of a country scatter to roughly 0.016–0.032 from a well-built calculator for it. Within those tolerances, distance rankings against curated panels are reliably reproducible and geographically sensible — the genes-mirror-geography result, live in a browser tool.

Layer three: the models#

Here is where "accuracy" usually breaks, and not because the arithmetic fails. nMonte-style fits always return percentages; the fit distance has no significance interpretation; adding sources improves fit mechanically while degrading the story; collinear sources split their signal arbitrarily. A G25 percentage is therefore accurate as a description of a stated fit against a stated panel — and undefined as anything else. The same genome, honestly fitted against two defensible panels, yields two different true descriptions. That is not G25 failing; it is what model-dependence means.

The practical calibration, layer by layer:

ClaimVerdict
"My row places me among these populations"Trustworthy, within panel coverage
"I am closer to X than to Y" (same era, clear gap)Trustworthy; check the gap is not noise-thin
"This fit says 48% source A"A description of one panel, portable nowhere
"My 2% component is real"Unconfirmed — quantisation and noise live at this scale
"This proves descent from X"Never available from coordinates, at any fit

Where G25 sits among the instruments#

Against a testing company's estimate: more transparent and deeper in time, smaller references, no segment view. Against component calculators: current references and open arithmetic versus 2012 constructs. Against qpAdm: G25 is faster, cheaper, always-answering — and unfalsifiable; qpAdm is slower, allele-frequency-based, and able to reject a model, which is why our own G25 conclusions are drawn from G25 evidence alone and claims that must survive argument go to the formal analysis.

Used inside its envelope — curated panels, era discipline, fits read as descriptions — G25 is the best exploration instrument the hobby has, and the free stack (distances, admixture, PCA) plus the worked Global25 report all live inside that envelope deliberately. The inaccuracy people fear mostly enters through the door the tools cannot lock: the reading.

€29.99 · one-time
The worked version of this analysis
Distances, admixture models and PCA across six eras against 1,535 curated populations, every source panel published in full, with Notable Matches included free.
See the Global25 analysis

Terms used here are defined in the glossary.


Related posts

What is a good G25 fit distance? The number everyone reads wrong
What is a good G25 fit distance? The number everyone reads wrong

The fit distance measures how far a mixture still sits from your coordinate — not whether the model is right. Calibrated bands for reading one, why lower stops being better, and the checks that actually validate a model.

4 min read
G25 scaled vs unscaled coordinates: which to use, how to tell them apart
G25 scaled vs unscaled coordinates: which to use, how to tell them apart

Every Global25 row exists in two forms, and mixing them is the hobby's most common silent error. What scaling does mathematically, which form each tool expects, and the tell-tale signs of a mixed comparison.

3 min read
How to read a G25 PCA plot without fooling yourself
How to read a G25 PCA plot without fooling yourself

What a Global25 PCA projection shows, what the axes are and are not, why plot distance lies, the difference between projecting onto a fixed view and fitting your own — and the honest uses of both.

3 min read
Back to all stories
Ancestrify

Combining cutting-edge genomic science with rich historical records to map your ancestry across generations and continents.


© 2026 Ancestrify. All rights reserved. · Ancestrify is a trading name of Andi Thomaj, a sole trader registered in Tiranë, Albania · NUIS M61725001N
Card payments processed by POK Payments (RPay Ltd)VISAMASTERCARD