Where did Indo-European languages come from? A major 2023 study in Science estimates that the family began diverging about 8,120 years before AD 2000, or roughly 6120 BCE. Its authors argue that this chronology fits neither a purely Anatolian farming expansion nor a purely Pontic-Caspian Steppe origin. Instead, they propose a hybrid hypothesis: an initial homeland south of the Caucasus, followed by a movement north onto the steppe, which later became a secondary homeland for some branches spreading into Europe.
The proposal is debated. The study analyzes vocabulary, not newly sequenced DNA; its dates are uncertain, several deep branches are weakly resolved, and later specialists have challenged parts of the analysis.
The short answer: the study supports an early Indo-European divergence and a two-stage south-Caucasus-to-steppe scenario. It does not prove a precise homeland, identify the language of any ancient skeleton, or remove the steppe from Indo-European history.
The debate: Anatolian farmers or steppe pastoralists?#
Languages from Albanian, Greek, English and Spanish to Armenian, Persian and Hindi ultimately descend from a common linguistic ancestor conventionally called Proto-Indo-European. Its location and age remain disputed.
Two broad models dominated the modern discussion:
- The Steppe hypothesis places the homeland on the Pontic-Caspian Steppe, generally no earlier than about 6,500 years ago, with major dispersals associated with mobile pastoralism and later steppe-derived migrations.
- The farming or Anatolian hypothesis connects an earlier language expansion with the spread of agriculture from parts of the Fertile Crescent and Anatolia, beginning roughly 9,000 years ago.
The 2023 paper narrows the linguistic chronology first, then asks which archaeological and demographic scenario fits it. A genetic migration does not automatically reveal the language that moved with it.
What the researchers analyzed#
Paul Heggarty and 32 co-authors created the Indo-European Cognate Relationships database, or IE-CoR. It contains carefully defined core vocabulary from 161 language varieties:
- 109 modern languages
- 52 ancient or historical languages
- 170 core meanings, including basic numbers, body parts, natural features and common actions
- 5,013 cognate sets in the full database
More than 80 specialists contributed to 25,918 lexeme and cognacy determinations. A cognate set groups words inherited from the same ancestral form, even when their modern forms differ.
The non-modern varieties include Hittite, Tocharian, Mycenaean and Attic Greek, Classical Armenian, Latin, Vedic Sanskrit, Avestan, Old English and Old Icelandic. Their dated texts provide historical calibration points.
These are the study's “samples.” It did not sequence skeletons or compare genomes. That makes its approach fundamentally different from qpAdm ancestry modeling, which evaluates genetic mixture models. Neither method can read a person's spoken language directly from DNA.

How a sampled-ancestor language tree works#
Earlier computational studies sometimes forced an ancient written language to be a direct ancestor of a modern group. That can sound intuitive—Latin before Romance, or Old English before modern English—but surviving texts usually preserve one dialect or formal register, not every spoken variety of their period.
The new analysis instead used a sampled-ancestor model. An ancient language could be placed directly on the line leading to later languages, but the model could also treat it as a closely related sister lineage. Written Classical Latin, for instance, can sit beside the spoken form of Latin from which Romance languages developed rather than being assumed to represent that spoken ancestor exactly.
The main analysis encoded cognate presence or absence in a matrix of 161 languages by 4,990 columns. It used BEAST 2.6.5, a relaxed linguistic clock, a birth-death-sampling tree prior and a binary covarion model that allowed change rates to vary across eight groups of meanings. Three independent chains ran for 100 million steps each.
One chronological detail is easy to miss: the paper defines “present” as AD 2000. Its years BP should therefore not be interpreted using the radiocarbon convention of 1950.
The headline result: a root around 8,120 years ago#
The median estimate for the Indo-European root is 8,120 BP, with a 95% credible interval of 6,740–9,610 BP. Converted from the study's AD 2000 baseline, that is approximately:
- Median: 6120 BCE
- 95% interval: 4740–7610 BCE
This is a probability distribution spanning nearly three millennia, not a calendar date for a documented event.
The authors also infer a sequence of early branch separations. Their rounded median estimates include:
| Linguistic split | Estimated date | 95% credible interval |
|---|---|---|
| Indo-European root | 8,120 BP | 6,740–9,610 BP |
| Indo-Iranic separates from the rest | 6,980 BP | 5,650–8,400 BP |
| Balto-Slavic separates from the western European grouping | 6,460 BP | 5,040–7,940 BP |
| Italic separates from Germanic-Celtic | 5,560 BP | 4,230–6,980 BP |
| Indic-Iranic split | 5,520 BP | 4,540–6,800 BP |
| Germanic-Celtic split | 4,890 BP | 3,720–6,190 BP |
The authors conclude that Indo-European had divided into seven major branches by about 6,140 BP, earlier than the large steppe-related genetic expansion into much of Europe.
Few ancient languages were direct ancestors in the model#
Of the 52 non-modern languages, 27 were plausible candidates for direct ancestry. Only four received posterior probabilities above 0.01: Classical Armenian, Mycenaean Greek, Attic Greek and New Testament Greek. Just two exceeded a probability of 50%.
Old English was not inferred as the direct ancestor of modern English because the dataset represents its well-documented West Saxon variety, whereas modern English descends most directly from other dialectal lineages. Written Classical Latin was not placed directly above modern Romance, and Vedic Sanskrit was treated as a sister to the lineages ancestral to later spoken Indic languages.
This does not deny historical continuity. Even one change in the preferred word for one of the 170 meanings creates a split, so a written language can be extremely close to an ancestor without being identical to it. As a check, the model dates the separation of Icelandic and Faroese from mainland Scandinavian lineages to around 830 CE and places initial Romance diversification in the first centuries CE, both compatible with documented history.
What the “hybrid hypothesis” actually proposes#
The model estimates relationships and dates; it does not calculate a homeland from coordinates. The proposal combines its linguistic chronology with published archaeology and ancient-DNA research.
The authors argue for this sequence:
- Proto-Indo-European began diverging south of the Caucasus, in or near the northern Fertile Crescent.
- Early branches separated from about 8,120 BP onward.
- One major lineage moved north through the Caucasus toward the steppe around 7,000–6,500 BP.
- The steppe became a secondary homeland for later Corded Ware-associated expansions into Europe around 5,000–4,500 BP.
In this interpretation, steppe migrations remain central to the spread of several European branches. They are simply too late to explain the first separation of every branch, especially Anatolian and the other early-diverging lineages. Ancient-DNA evidence for the later transformation of southeastern Europe is explored separately in our overview of Roman-era and Slavic-period Balkan ancestry.
Genes and languages can travel together, but they do not have to. Language shift can occur without major genetic replacement, and migrants can adopt a local language. Genetic components such as “steppe ancestry” are not languages in molecular form.
Where Albanian fits#
The study treats Albanian as one of the 12 principal attested Indo-European branches and samples Standard Albanian, Gheg and Arbëresh. Its tree places Albanian, Greek, Armenian and Anatolian deeper than the Germanic-Celtic-Italic grouping, before the main steppe-associated expansion modeled for much of Europe.
The exact position of Albanian among the earliest separations is uncertain, however. In the manuscript's summary table, its estimated split from the rest of Indo-European depends on a node with less than 50% posterior support. A separate estimate for divergence among the sampled Albanian varieties is not a date for the origin of Albanian itself, still less for the origin of Albanian ethnicity.
This linguistic evidence complements but cannot replace the population evidence reviewed in our guide to Albanian ancient DNA. A language tree traces inherited vocabulary; an ancestry study traces biological relationships. Neither alone proves which language an ancient community spoke.
How robust were the dates?#
The researchers tested alternative assumptions about calibrations, loans, missing data, the number of living languages and imposed tree structures. Several changes had limited effects on the root estimate:
- Removing the debated Vedic and Avestan calibrations shifted the median from 8,120 to 8,214 BP.
- Treating parallel loans differently produced 7,934 BP.
- Assuming either 200–400 or 600–800 living Indo-European languages produced 8,064 and 8,177 BP, respectively.
- Removing ten languages with high missing-data rates changed the median by only two years.
- Even forcing all 27 remotely possible ancient ancestors shifted the median to 7,614 BP, still earlier than a strict steppe-only chronology.
The older root is therefore not caused by one calibration or prior choice. The tests do not eliminate uncertainty in the data, deep topology or geographic interpretation.
The most important limitations#
- The root date is model-based and broad. The 8,120 BP median sits inside a 6,740–9,610 BP interval, and alternative assumptions can move both its center and uncertainty.
- Core vocabulary is one part of language history. Traditional classifications also use sound changes and morphology. The authors acknowledge unexpected placements within Nuristani, western Iranic and West Germanic.
- Contact can resemble common ancestry. Recognized loans were marked and tested, but undetected borrowing can still pull neighboring branches together.
- The deepest relationships are weakly resolved. Each of three configurations near the root had less than 26% support. A 2025 peer-reviewed critical reanalysis also challenged several early nodes and identified possible word-selection and loan-coding problems. The hybrid scenario remains debated.
- The written record is uneven. Lost languages leave no wordlists, and too little evidence survives from several Palaeo-Balkan languages for inclusion.
- DNA cannot identify a language. Matching a linguistic date to a genetic movement makes a connection plausible; it cannot demonstrate what every person carrying that ancestry spoke.
Frequently asked questions#
Where did Indo-European languages originate according to this study?#
The authors propose an initial homeland south of the Caucasus, near the northern Fertile Crescent, followed by a movement north onto the steppe. This location is an interpretation combining linguistic chronology with earlier genetic and archaeological evidence, not a geographic result directly calculated by the language model.
How old is Proto-Indo-European?#
The median estimate is approximately 8,120 years before AD 2000, or about 6120 BCE. Its 95% credible interval is approximately 4740–7610 BCE, so the study does not supply an exact birth date for the language.
Does the paper disprove the Steppe hypothesis?#
No. It challenges the steppe as the sole ultimate homeland of every Indo-European branch. The steppe remains a secondary homeland and an important source for later expansions of several European branches.
Is this an ancient-DNA study?#
No. The primary dataset consists of cognate vocabulary from 161 languages. The authors use ancient-DNA findings from other studies to interpret how their linguistic timeline might fit known population movements.
Why are Latin and Sanskrit not direct ancestors in the tree?#
Surviving texts represent particular dialects and written registers. Modern Romance descends from spoken Latin varieties rather than matching Classical literary Latin word for word, while Vedic Sanskrit was a particular early Indic variety. The model permits these samples to be close sister lineages instead of forcing them onto a direct line.
What does the study imply about Albanian?#
It confirms Albanian as an independent major Indo-European branch and places it among branches separating deeper than the main western European group. The exact early branching position has low support and cannot establish a prehistoric ethnic identity, migration route or genetic origin.
A useful result, not the final word#
The study's strongest contribution is a carefully curated dataset and a model that does not assume every famous written language is the direct ancestor of its modern relatives. Its older chronology remains fairly stable across many sensitivity tests. The historical synthesis is more tentative: weak deep branches, contact, incomplete records and the indirect relationship between genes and speech leave room for competing explanations.
The most accurate conclusion is therefore conditional: this language tree supports a hybrid origin scenario, but it does not prove one.
Sources and data#
- Heggarty, P. et al. (2023). Language trees with sampled ancestors support a hybrid model for the origin of Indo-European languages. Science 381(6656), eabg0818. PubMed record.
- Heggarty, P., Anderson, C. and Scarborough, M. IE-CoR: Indo-European Cognate Relationships.
- Heggarty, P. et al. (2023). Supplementary analysis data, result files and reproducibility guide. Zenodo.
- Heggarty, P. et al. Author-accepted manuscript. Universitat Jaume I repository.
- Kassian, A. et al. (2025). Do “language trees with sampled ancestors” really support a “hybrid model” for the origin of Indo-European?. Humanities and Social Sciences Communications.
Editorial note: the hero and section artwork in this article was generated with AI as conceptual illustration. It does not reproduce the study's scientific figures, depict a documented migration route, or reconstruct a specific ancient person, language community or archaeological site.



