Ancestrify
All stories

yamnaya

By Ancestrify
5 min read

Steppe (Yamnaya) ancestry percentage: how to measure

Who the Western Steppe Herders were, how steppe ancestry spread across Europe and Asia after 3000 BC, what a 'steppe percentage' actually measures, and how to estimate yours with qpAdm or Global25 from a raw DNA file.

yamnayaancient-dnasteppebronze-ageqpadmglobal25guide

  1. Who the Western Steppe Herders were
  2. How steppe ancestry spread
  3. What a "steppe percentage" actually measures
  4. Measuring it with qpAdm
  5. Measuring it with Global25
  6. Reading your number honestly
  7. Where to go next
  8. References

"How much steppe ancestry do I have?" is the most-asked quantitative question in consumer ancient DNA, and the one most often answered with a number that means less than it looks. This guide covers who the steppe populations were, what a "steppe percentage" actually measures, and how to estimate yours defensibly from the raw file you already have.

Who the Western Steppe Herders were#

Between roughly 3300 and 2600 BC the Pontic-Caspian steppe — the grassland belt from the Danube delta to the Ural river — was home to the Yamnaya culture: mobile pastoralists who buried their dead under earth mounds (kurgans), used ox-drawn wagons, and herded cattle, sheep and horses across open country. Genetically, Yamnaya people were themselves a mixture: roughly half Eastern Hunter-Gatherer ancestry from the forest-steppe to the north, and half ancestry related to the Caucasus and Iran, with a smaller farmer-related component. That profile is what the term Western Steppe Herder (WSH) names in the reference panels, and its source-population page is at Western Steppe Herders, 5000–2800 BC. The archaeology and genetics are told at length in Yamnaya DNA: where steppe ancestry came from.

How steppe ancestry spread#

Two 2015 papers — Haak and colleagues in Nature, and Allentoft and colleagues in the same journal — showed that after about 3000 BC this steppe profile appears across a huge area where it had been absent: in the Corded Ware cultures of central and northern Europe, then in Bell Beaker groups as far as Britain and Iberia, and eastward into the Afanasievo culture of the Altai. In Britain, Olalde and colleagues (2018) found that the arrival of Beaker-associated people replaced roughly 90% of the earlier gene pool within a few centuries.

The share that arrived varied by region and then changed again with later movements. As a rough picture from the published literature, steppe-related ancestry today is highest in northern and north-eastern Europe, intermediate across central Europe, the British Isles and the Balkans, and lowest in the Mediterranean south — Sardinia in particular retains very little. Southern Asia carries a steppe-related component that arrived by a separate route through Central Asia in the second millennium BC (Narasimhan et al., 2019). None of these are fixed numbers, and the point of running the analysis is to measure your own file rather than to look up a country.

What a "steppe percentage" actually measures#

This is the part to read before the number.

A steppe percentage is the weight assigned to a steppe-related source population in a specific admixture model. Change the model and the number changes — not because your genome changed but because the question did. Four things decide it:

  1. Which population stands in for "steppe". Yamnaya from Samara, Yamnaya from the Caspian shore, Afanasievo and Corded Ware are all steppe-related, but they are not identical, and a model built on one will give a different weight from a model built on another.
  2. Which other sources are offered. Steppe ancestry is measured against the alternatives — typically an Anatolian-farmer source and a hunter-gatherer source in the earliest era. Offer a later, already-mixed population as a source and part of your steppe share is absorbed into it.
  3. Which outgroups the model is tested against (in qpAdm). The right set decides whether the sources can be told apart at all.
  4. How much of your file survives the merge. Coverage sets the standard error; a steppe weight of 0.31 ± 0.03 is a finding, 0.31 ± 0.12 is a range from a fifth to almost a half.

So a defensible steppe figure is always "X% ± Y, with these sources, against these outgroups, in this era" — and two figures from two products with different models are not in disagreement, they are different measurements.

Measuring it with qpAdm#

Formal admixture modelling is the method the papers above used, and it is the one that can tell you whether a three-way farmer–steppe–hunter-gatherer model is even admissible for your genome. In our qpAdm analysis, your raw file is merged with the Allen Ancient DNA Resource v66 and modelled by hand in two eras; the Hunter-Gatherer & Neolithic Farmer era is where the steppe component is resolved most cleanly, because its sources — Western Steppe Herders, Anatolian Neolithic farmers, Western and Eastern Hunter-Gatherers — are genuinely distinct.

The report prints, for every source, the weight, its standard error, its z-score and its 95% confidence interval, plus the model's p-value, the complete right set and the full model record (chi-square, degrees of freedom, the nested-model table) as a plain-text download. Reading those four numbers in the right order is the subject of How to read qpAdm results. The buyer-level overview is What is a qpAdm ancestry test?. After publication, the €10 Model Lab unlock lets you swap the steppe proxy yourself and watch the weight move — the most instructive thing you can do with the number.

Measuring it with Global25#

If you hold Global25 coordinates, the same three-way structure can be fitted as a coordinate mixture. The free admixture calculator does this in your browser against curated per-era source panels; the paid Global25 analysis applies the curated calculators and prints every source the calculator was offered, used and unused. The coordinate method always returns a percentage and has no p-value, so the panel you fit against matters even more — see qpAdm vs Global25 for when each is the right instrument, and Understanding Global25 for reading a fit distance.

Reading your number honestly#

  • Compare within one model, not across products. Your steppe weight is comparable with the same model run on another file, not with a different calculator's output.
  • Do not read a small steppe weight as zero. Check its z-score. A 6% weight with |Z| = 1.1 is indistinguishable from nothing; a 6% weight with |Z| = 4 is real.
  • Do not read a large steppe weight as identity. The Yamnaya were one population in one millennium; a 40% steppe-related weight means your genome is compatible with drawing that share of ancestry from a population like them. It is not a claim that any particular buried steppe individual was related to you, and it says nothing about language or nationality.
  • Expect regional plausibility, not surprise. A result far outside the published range for people of your background is a reason to check the model before believing the number.

Where to go next#

The two components steppe ancestry is measured against have their own guides: Hunter-gatherer ancestry and Neolithic farmer ancestry. The terms are in the glossary. And if you have a raw file and no idea whether it is dense enough to resolve the split, the free Raw DNA File Check will say so before you spend anything.

From €29.99 · one-time
The tested version of this question
A qpAdm model composed, run and checked by hand against AADR v66, published with its p-value, every source's standard error and z-score, and the full right set, so the result can be argued with.
See the qpAdm analysis

References#

  • Haak, W. et al. (2015). Massive migration from the steppe was a source for Indo-European languages in Europe. Nature, 522, 207–211.
  • Allentoft, M. E. et al. (2015). Population genomics of Bronze Age Eurasia. Nature, 522, 167–172.
  • Olalde, I. et al. (2018). The Beaker phenomenon and the genomic transformation of northwest Europe. Nature, 555, 190–196.
  • Narasimhan, V. M. et al. (2019). The formation of human populations in South and Central Asia. Science, 365, eaat7487.
  • Lazaridis, I. et al. (2022). The genetic history of the Southern Arc: a bridge between West Asia and Europe. Science, 377, eabm4247.

Related posts

Ancient DNA test vs 23andMe or AncestryDNA estimates
Ancient DNA test vs 23andMe or AncestryDNA estimates

What a 23andMe or AncestryDNA ethnicity estimate measures, what an ancient-DNA ancestry test measures instead, why the two disagree by design, and how to run the second one on the raw file you already have.

6 min read
Hunter-gatherer ancestry: WHG, EHG and how to test it
Hunter-gatherer ancestry: WHG, EHG and how to test it

Who Europe's Mesolithic hunter-gatherers were — Western, Eastern and Caucasus — how their ancestry survived farming and the steppe migrations, and how to measure your hunter-gatherer share from a raw DNA file.

6 min read
Neolithic farmer ancestry: Anatolian roots, modern share
Neolithic farmer ancestry: Anatolian roots, modern share

Who the Anatolian Neolithic farmers were, how their ancestry spread across Europe from 6500 BC and became the largest component in most southern Europeans, what 'early European farmer' means in a model, and how to measure your share.

5 min read
Back to all stories
Ancestrify

Combining cutting-edge genomic science with rich historical records to map your ancestry across generations and continents.


© 2026 Ancestrify. All rights reserved. · Ancestrify is a trading name of Andi Thomaj, a sole trader registered in Tiranë, Albania · NUIS M61725001N
Card payments processed by POK Payments (RPay Ltd)VISAMASTERCARD