"How much steppe ancestry do I have?" is the most-asked quantitative question in consumer ancient DNA, and the one most often answered with a number that means less than it looks. This guide covers who the steppe populations were, what a "steppe percentage" actually measures, and how to estimate yours defensibly from the raw file you already have.
Who the Western Steppe Herders were#
Between roughly 3300 and 2600 BC the Pontic-Caspian steppe — the grassland belt from the Danube delta to the Ural river — was home to the Yamnaya culture: mobile pastoralists who buried their dead under earth mounds (kurgans), used ox-drawn wagons, and herded cattle, sheep and horses across open country. Genetically, Yamnaya people were themselves a mixture: roughly half Eastern Hunter-Gatherer ancestry from the forest-steppe to the north, and half ancestry related to the Caucasus and Iran, with a smaller farmer-related component. That profile is what the term Western Steppe Herder (WSH) names in the reference panels, and its source-population page is at Western Steppe Herders, 5000–2800 BC. The archaeology and genetics are told at length in Yamnaya DNA: where steppe ancestry came from.
How steppe ancestry spread#
Two 2015 papers — Haak and colleagues in Nature, and Allentoft and colleagues in the same journal — showed that after about 3000 BC this steppe profile appears across a huge area where it had been absent: in the Corded Ware cultures of central and northern Europe, then in Bell Beaker groups as far as Britain and Iberia, and eastward into the Afanasievo culture of the Altai. In Britain, Olalde and colleagues (2018) found that the arrival of Beaker-associated people replaced roughly 90% of the earlier gene pool within a few centuries.
The share that arrived varied by region and then changed again with later movements. As a rough picture from the published literature, steppe-related ancestry today is highest in northern and north-eastern Europe, intermediate across central Europe, the British Isles and the Balkans, and lowest in the Mediterranean south — Sardinia in particular retains very little. Southern Asia carries a steppe-related component that arrived by a separate route through Central Asia in the second millennium BC (Narasimhan et al., 2019). None of these are fixed numbers, and the point of running the analysis is to measure your own file rather than to look up a country.
What a "steppe percentage" actually measures#
This is the part to read before the number.
A steppe percentage is the weight assigned to a steppe-related source population in a specific admixture model. Change the model and the number changes — not because your genome changed but because the question did. Four things decide it:
- Which population stands in for "steppe". Yamnaya from Samara, Yamnaya from the Caspian shore, Afanasievo and Corded Ware are all steppe-related, but they are not identical, and a model built on one will give a different weight from a model built on another.
- Which other sources are offered. Steppe ancestry is measured against the alternatives — typically an Anatolian-farmer source and a hunter-gatherer source in the earliest era. Offer a later, already-mixed population as a source and part of your steppe share is absorbed into it.
- Which outgroups the model is tested against (in qpAdm). The right set decides whether the sources can be told apart at all.
- How much of your file survives the merge. Coverage sets the standard error; a steppe weight of 0.31 ± 0.03 is a finding, 0.31 ± 0.12 is a range from a fifth to almost a half.
So a defensible steppe figure is always "X% ± Y, with these sources, against these outgroups, in this era" — and two figures from two products with different models are not in disagreement, they are different measurements.
Measuring it with qpAdm#
Formal admixture modelling is the method the papers above used, and it is the one that can tell you whether a three-way farmer–steppe–hunter-gatherer model is even admissible for your genome. In our qpAdm analysis, your raw file is merged with the Allen Ancient DNA Resource v66 and modelled by hand in two eras; the Hunter-Gatherer & Neolithic Farmer era is where the steppe component is resolved most cleanly, because its sources — Western Steppe Herders, Anatolian Neolithic farmers, Western and Eastern Hunter-Gatherers — are genuinely distinct.
The report prints, for every source, the weight, its standard error, its z-score and its 95% confidence interval, plus the model's p-value, the complete right set and the full model record (chi-square, degrees of freedom, the nested-model table) as a plain-text download. Reading those four numbers in the right order is the subject of How to read qpAdm results. The buyer-level overview is What is a qpAdm ancestry test?. After publication, the €10 Model Lab unlock lets you swap the steppe proxy yourself and watch the weight move — the most instructive thing you can do with the number.
Measuring it with Global25#
If you hold Global25 coordinates, the same three-way structure can be fitted as a coordinate mixture. The free admixture calculator does this in your browser against curated per-era source panels; the paid Global25 analysis applies the curated calculators and prints every source the calculator was offered, used and unused. The coordinate method always returns a percentage and has no p-value, so the panel you fit against matters even more — see qpAdm vs Global25 for when each is the right instrument, and Understanding Global25 for reading a fit distance.
Reading your number honestly#
- Compare within one model, not across products. Your steppe weight is comparable with the same model run on another file, not with a different calculator's output.
- Do not read a small steppe weight as zero. Check its z-score. A 6% weight with |Z| = 1.1 is indistinguishable from nothing; a 6% weight with |Z| = 4 is real.
- Do not read a large steppe weight as identity. The Yamnaya were one population in one millennium; a 40% steppe-related weight means your genome is compatible with drawing that share of ancestry from a population like them. It is not a claim that any particular buried steppe individual was related to you, and it says nothing about language or nationality.
- Expect regional plausibility, not surprise. A result far outside the published range for people of your background is a reason to check the model before believing the number.
Where to go next#
The two components steppe ancestry is measured against have their own guides: Hunter-gatherer ancestry and Neolithic farmer ancestry. The terms are in the glossary. And if you have a raw file and no idea whether it is dense enough to resolve the split, the free Raw DNA File Check will say so before you spend anything.
References#
- Haak, W. et al. (2015). Massive migration from the steppe was a source for Indo-European languages in Europe. Nature, 522, 207–211.
- Allentoft, M. E. et al. (2015). Population genomics of Bronze Age Eurasia. Nature, 522, 167–172.
- Olalde, I. et al. (2018). The Beaker phenomenon and the genomic transformation of northwest Europe. Nature, 555, 190–196.
- Narasimhan, V. M. et al. (2019). The formation of human populations in South and Central Asia. Science, 365, eaat7487.
- Lazaridis, I. et al. (2022). The genetic history of the Southern Arc: a bridge between West Asia and Europe. Science, 377, eabm4247.



