When your Global25 coordinates arrive from the G25 Requests portal, you receive two files: a scaled row and an unscaled row, same genome, same 25 axes, different numbers. Which one you paste matters more than anything else you will do with them — a mixed comparison produces plausible-looking garbage with no error message anywhere. This is the complete version of the scaled/unscaled story: what the transformation is, which form to use where, and how to detect a mix-up after the fact.
What scaling actually does#
The 25 axes of the G25 space are principal components, ranked by how much variation each explains: PC1 carries the most (the deepest continental structure), PC25 the least (fine regional texture). In the unscaled form, each coordinate is the raw projection onto its axis — so PC1 values span a range several times wider than PC20's, and any distance computed across the raw row is dominated by the first few dimensions.
The scaled form multiplies each dimension by a factor tied to that axis's share of variance, rebalancing the row so the later, finer axes contribute meaningfully to distances. In effect: unscaled geometry answers "how far apart are these genomes on the broadest axes"; scaled geometry answers "how far apart are they across the whole structure, fine detail included". For ancestry work within a continent — where all the action is in the fine axes — that is why scaled is the convention almost everywhere: distance rankings, admixture fits, the published reference averages, Vahaduo sheets, and every panel in our tools and the paid analysis.
Unscaled rows are not junk — the raw projection has legitimate uses in some plotting and methodological contexts — but in the consumer toolchain their practical role is: the other file, the one you keep labelled and do not paste.
The one unbreakable rule#
Never let the two forms meet in one calculation. A scaled target against unscaled references (or vice versa) computes fine, sorts fine, and means nothing: the mismatched row's early dimensions are weighted entirely differently from the panel it is being measured against. The arithmetic cannot notice — the numbers are numbers — so the failure is silent, and it poisons everything downstream: distances, fits, PCA positions, averages.
Corollaries worth spelling out: keep both files exactly as delivered, filenames intact; never "convert" one form to the other yourself with a factor found on a forum; and when you build a population average, build it from rows of one form only — an average of mixed rows is mixed garbage with extra steps.
How to tell which row you are holding#
Labels get lost. Three checks, in increasing rigour:
- Eyeball the spread. In an unscaled row, the first few values are conspicuously larger in magnitude than the rest; a scaled row's values are more even across the 25. Suggestive, not proof.
- The absurd-distance tell. Paste the row into a scaled-reference tool like the distance calculator: a genuine scaled row of a modern person lands within ~0.02–0.05 of some modern population. If the nearest population on Earth sits beyond roughly 0.08, you are almost certainly holding the unscaled form (or a corrupted paste) — suspect the row before suspecting your ancestry.
- The fingerprint check. The free G25 authenticity check reads a row's numeric signature and reports what it looks like — scaled, unscaled, rounded, edited or simulated — before you build anything on it.
Answers to the questions people actually ask#
Which form does Ancestrify need? Scaled, everywhere — the free Lab tools and the Global25 analysis alike, matching the convention of every reference panel we publish against. Paste the file labelled scaled and you are done.
Why do my percentages change between forms? Because the geometry changed. An admixture solver fitting unscaled rows is solving a different (broad-axis-dominated) problem; neither result is "the real one" in the abstract, but the scaled fit is the one comparable to every published model and calculator you will ever see.
Both my forms give weird results. Then the row itself is the suspect: thin input file, transcription damage, or a simulated origin. Pre-flight the raw file with the file check and the row with the authenticity check; the full troubleshooting route goes from there.
The two-file design is a small permanent tax the format charges for its portability. Pay it once — label, verify, standardise on scaled — and it never bothers you again.
Terms used here are defined in the glossary.



