Ancestrify
All stories

global25

By Andi Thomaj
3 min read

nMonte explained: the R script behind every G25 admixture fit

Ger Huijbregts's nMonte turned Global25 rows into ancestry percentages and founded a whole tool tradition. How the Monte-Carlo search works, what nMonte3 changed, what the fit distance is, and the discipline the script never enforces.

global25methodologyvahaduo

  1. The problem it solves
  2. nMonte versus nMonte3, and the dials
  3. What the script never enforces
  4. Do you need to run the R script?

Every Global25 admixture percentage you have ever seen — in Vahaduo, in our calculators, in a decade of forum arguments — descends from one R script. Ger Huijbregts wrote nMonte in the mid-2010s, the Eurogenes blog adopted it as the standard companion to its coordinate sheets, and its core idea has been reimplemented so many times that "nMonte-style" is now simply the name of the method. This is what the script actually does, and what it deliberately does not.

The problem it solves#

Given a target row and a set of source rows in the same coordinate space, find non-negative percentages summing to 100% whose weighted average of the sources lands as close to the target as possible. That is a constrained least-squares problem, and nMonte solves it the pragmatic way: Monte-Carlo descent. Start from some allocation of weight across sources; repeatedly propose a small random reallocation (move a slice of weight from one source to another); keep the proposal if the mixture's Euclidean distance to the target shrinks; stop when proposals stop helping. The final allocation is the breakdown, and the residual gap is the fit distance.

Randomised descent has two properties worth knowing. It handles any panel size without matrix algebra, which is why it ports so easily to browsers. And it is a local search: with near-collinear sources, different random runs settle on different splits of the shared signal — same fit, different story. When two sources trade ten points between runs, the instability is information: the panel, not the arithmetic, cannot tell them apart.

nMonte versus nMonte3, and the dials#

The original script fits the target as pasted. nMonte3 added the option everyone now argues about: a penalty term (pen) that trades a slightly worse fit for a sparser, less scattered model, damping the script's tendency to sprinkle 1–2% across many sources. Batch mode, "1 outcome per line" runs and sheet conventions accumulated around it. Every dial is a modelling choice: penalty on and off can move percentages by real amounts with near-identical fits — which is not a bug but the method telling you those models are not distinguishable by distance alone.

Our production engine is a seeded descendant of the same family, with the choices fixed and stated: 500 weight slots (0.2% granularity), five independent restarts, pruning of sub-1.5% components followed by re-solving, a hard cap of eight sources, and a deterministic seed so a published result reproduces byte for byte. None of that changes the mathematics; it changes whether two people running "the same model" get the same answer.

What the script never enforces#

nMonte computes exactly what you asked and nothing about whether the question was sound. The discipline lives outside the script, and forgetting that is the whole failure mode of the genre:

Do you need to run the R script?#

Only if you want the dials or scriptable batch runs. For everything else the browser implementations are the same mathematics with the bookkeeping handled: Vahaduo for bring-your-own-sheet freedom, our free calculators for curated era-scoped panels with the fit stated — and the full toolbox is here. If you do run it: R installed, nMonte3.R plus a data and target file in the working directory, source('nMonte3.R'), and the sheet discipline above observed with the seriousness the script itself will never demand.

A ten-year-old R script with no test, no errors bars and no opinions became the load-bearing tool of an entire hobby — which is exactly why knowing its shape matters. The percentages were never the script's claim. They were always yours.

€29.99 · one-time
The worked version of this analysis
Distances, admixture models and PCA across six eras against 1,535 curated populations, every source panel published in full, with Notable Matches included free.
See the Global25 analysis

Terms used here are defined in the glossary.


Related posts

What is a good G25 fit distance? The number everyone reads wrong
What is a good G25 fit distance? The number everyone reads wrong

The fit distance measures how far a mixture still sits from your coordinate — not whether the model is right. Calibrated bands for reading one, why lower stops being better, and the checks that actually validate a model.

4 min read
G25 source selection: how good models overfit and honest panels are built
G25 source selection: how good models overfit and honest panels are built

Adding sources always lowers a G25 fit distance and routinely worsens the model — a worked demonstration where the fit improved from 0.0110 to 0.0068 while the story collapsed. The rules that keep a panel honest.

4 min read
G25 scaled vs unscaled coordinates: which to use, how to tell them apart
G25 scaled vs unscaled coordinates: which to use, how to tell them apart

Every Global25 row exists in two forms, and mixing them is the hobby's most common silent error. What scaling does mathematically, which form each tool expects, and the tell-tale signs of a mixed comparison.

3 min read
Back to all stories
Ancestrify

Combining cutting-edge genomic science with rich historical records to map your ancestry across generations and continents.


© 2026 Ancestrify. All rights reserved. · Ancestrify is a trading name of Andi Thomaj, a sole trader registered in Tiranë, Albania · NUIS M61725001N
Card payments processed by POK Payments (RPay Ltd)VISAMASTERCARD