Ancestrify
All stories

global25

By Andi Thomaj
4 min read

What is a good G25 fit distance? The number everyone reads wrong

The fit distance measures how far a mixture still sits from your coordinate — not whether the model is right. Calibrated bands for reading one, why lower stops being better, and the checks that actually validate a model.

global25guidemethodologyvahaduo

  1. What the number is
  2. Reading the magnitude
  3. Why lower stops being better
  4. What actually validates a model

Run any Global25 admixture fit — Vahaduo, nMonte, our calculators — and above the percentages sits one number: the fit distance. It is the most argued-about figure in the hobby ("distance 0.0192, is that good??") and the most misread, because it looks like a grade for the model and is actually something narrower: the leftover gap. This post is how to read it properly, including the calibrated bands we use when deciding whether a model is publishable at all.

What the number is#

The solver searched for the weighted mixture of your chosen sources whose combined coordinate lands closest to yours. The fit distance is how far that best mixture still sits from your point — a plain Euclidean residual across the 25 dimensions, on scaled coordinates. (Tools differ in display: 0.0192 raw and 1.92 percent-style are the same number.)

So: small fit = the geometry closed well. That is all it certifies. It is not a p-value, carries no standard errors, and cannot tell a historically sound panel from a numerically convenient one — no coordinate method can.

Reading the magnitude#

Bands we calibrated against curated era panels and modern individuals of known origin — for a single modern person fitted against a well-built, era-scoped panel of 3–8 sources, raw scaled distances:

FitReading
≤ 0.020Excellent — the panel describes this coordinate about as well as panels do
0.020–0.035Sound — normal territory for real customers on ethnicity-era panels
0.035–0.045Acceptable for deep-era panels; on a modern-era panel, start asking questions
0.045–0.060Poor — a source axis is probably missing, or the input is atypical (see below)
> 0.060Not a usable model; something is wrong with panel or paste

Population averages fit tighter than individuals (an average has had its personal noise averaged away — our own country calculators fit their country's average around 0.011), and deep-era panels run systematically wider than modern ones because every living person is far from every Bronze Age source. Compare like with like or not at all.

Two disclaimers the bands need. Typicality: some coordinates sit genuinely far from every reference — check your best distance to any modern population in the distance tool; if that baseline is 0.045, no honest panel will fit you to 0.020, and the excess is your coordinate's atypicality (or a kit-quality problem), not the model's failure. And the absurd tell: a nearest-modern-population distance beyond ~0.08 usually means an unscaled or corrupted paste — fix the row before reading anything.

Why lower stops being better#

Here is the property that breaks the "grade" intuition: adding sources can only lower the optimal fit, whether or not they belong in the story. Our calibration demonstration on an Albanian average: a curated 5-source calculator fit at 0.0110 with a clean, readable history; merging five Balkan panels into 12 sources improved the fit to 0.0081 while the lead component's share collapsed from ~56% to ~14%, scattered across near-substitutes; offering all 27 era sources reached 0.0068 as thirteen-component noise. Monotonically better fit, monotonically worse model. The full anatomy of that staircase — and the collinearity trap that makes percentages seed-dependent while the fit barely moves — is in source selection and overfitting.

Consequences worth pinning: fits are only comparable between panels of similar size on the same target; a forum model beating yours by 0.003 with four extra sources has demonstrated nothing; and chasing the last decimal place is how good models are ruined.

What actually validates a model#

Since the fit cannot, validation is structural — the checks are free and any tool supports them:

  • Every source earns its place: real weight (≥ 2%), and dropping it worsens the fit by a stated margin, not by rounding.
  • Stability: re-run unseeded several times; percentages that swing between runs are collinearity talking.
  • The neighbourhood agrees: your closest populations should be explainable by the model that claims to describe you.
  • Era discipline: one dated window per panel, ancient and modern kept apart.
  • And when the claim must survive argument — this source is required, that component is real — the test lives outside coordinate space: qpAdm returns a p-value that can reject the model and a z-score per source, which is the difference in kind.

That is also, plainly, how we work: the paid Global25 analysis publishes its panel, era and fit — bounded from both sides, against a bar where a suspiciously low fit on an oversized panel fails just as surely as a bad one — with a written rationale, because the fit distance was never the argument. It is the gap the argument has to explain.

€29.99 · one-time
The worked version of this analysis
Distances, admixture models and PCA across six eras against 1,535 curated populations, every source panel published in full, with Notable Matches included free.
See the Global25 analysis

Terms used here are defined in the glossary.


Related posts

G25 scaled vs unscaled coordinates: which to use, how to tell them apart
G25 scaled vs unscaled coordinates: which to use, how to tell them apart

Every Global25 row exists in two forms, and mixing them is the hobby's most common silent error. What scaling does mathematically, which form each tool expects, and the tell-tale signs of a mixed comparison.

3 min read
G25 source selection: how good models overfit and honest panels are built
G25 source selection: how good models overfit and honest panels are built

Adding sources always lowers a G25 fit distance and routinely worsens the model — a worked demonstration where the fit improved from 0.0110 to 0.0068 while the story collapsed. The rules that keep a panel honest.

4 min read
How accurate is G25? What the coordinate can and cannot resolve
How accurate is G25? What the coordinate can and cannot resolve

Global25 is a projection, not a test — so 'accuracy' has layers: the coordinate's own fidelity, the references it is compared against, and the models built on top. An honest audit of all three.

3 min read
Back to all stories
Ancestrify

Combining cutting-edge genomic science with rich historical records to map your ancestry across generations and continents.


© 2026 Ancestrify. All rights reserved. · Ancestrify is a trading name of Andi Thomaj, a sole trader registered in Tiranë, Albania · NUIS M61725001N
Card payments processed by POK Payments (RPay Ltd)VISAMASTERCARD