Ancestrify
All stories

qpadm

By Andi Thomaj
3 min read

qpAdm models for Middle Eastern ancestry: four streams and the collinearity discipline

The Middle East is where qpAdm's sources crowd closest together: Natufian, Anatolian, Iranian and Caucasus ancestries all interrelated. The working recipe, the right set that splits them, failure modes and a worked reading.

qpadmguidepopulation-geneticsmiddle-east

  1. The four streams
  2. The standard recipe
  3. The named failure modes
  4. A worked reading
  5. References

Europe teaches the standard recipe; South Asia teaches the proxy problem; the Middle East teaches collinearity — what happens when the candidate sources are themselves relatives. The region's four founding streams share deep structure, exchange gene flow from the Neolithic onward, and produce the classic pathology: wild weights, huge errors, models that pass while meaning little. Modelling here is a discipline of making the streams distinguishable before asking about them.

The four streams#

The Epipalaeolithic-to-Neolithic Near East resolves into four deeply divergent but interrelated ancestries:

  • Natufian / Levant_N — the Epipalaeolithic Levant and its early farmers.
  • Anatolia_N (Barcın) — northwest Anatolian farmers, Europe's EEF source.
  • Iran_N (Ganj Dareh) — early Zagros farmers.
  • CHG (Kotias/Satsurblia) — Caucasus hunter-gatherers, Iran_N's closest deep relative.

Iran_N and CHG are the notorious pair — close enough that naive models cannot split them — with Natufian↔Anatolia_N the second-tightest edge. Later prehistory then stirs the pot continuously, which is why this region, more than any other, runs on the distal discipline: all four sources predate the mixing they explain.

The standard recipe#

Left: target + two to four of Levant_N (or Natufian), Anatolia_N, Iran_GanjDareh_N and Georgia_CHG — chosen by geography, and never Iran_N and CHG together in a first model: lowest rank first means starting two-source (Levantine targets: Levant_N + Iran_N; Anatolian/Mesopotamian: Anatolia_N + Iran_N or + CHG) and letting the rank test force additions.

Right: the region's models live or die on the o9a-style extension of the O9 spine: Mbuti.DG, Ust_Ishim.DG, Mota.DG, MA1, Ami.DG, Onge.DG, plus the discriminators — Russia_EHG, WHG, Morocco_Iberomaurusian (the North African edge, essential for Levantine and Egyptian-adjacent targets), and where the source list allows it, whichever of the four streams sits out of the left set. This is the published case where the machinery visibly rewards craft: the Bronze Age Levant work found one added differentially-related outgroup cut standard errors roughly threefold — the exact anti-collinearity lever, applied.

Later-era targets (Bronze Age onward, and all moderns) usually need a steppe-related fourth column (Yamnaya/Sintashta-related, for Anatolian, Iranian-plateau and Levantine-diaspora targets) and, for the Arabian peninsula and Horn-adjacent targets, an African stream (Mota.DG-related or Egypt-adjacent proxies) — each addition paid for by a nested-model rejection, never added on vibes.

The named failure modes#

  • The Iran/CHG seesaw. Both in the left set with a generic right set: weights trade against each other run to run, SEs balloon, sometimes one goes negative. Fix from the right (EHG, steppe-related and Anatolia contrasts split them best), or accept the honest merged reading — "Iranian/Caucasus-related" as one stream — when the data sit below the resolution floor.
  • Missing North Africa. Levantine and coastal targets modelled without an Iberomaurusian-related term fail — or worse, pass with distorted Natufian weights (Natufians themselves relate to North African lineages; the right set must be able to see that edge).
  • Umbrella labels. Iran_N pooling Ganj Dareh with later Chalcolithic plateau samples, or Levant labels pooling PPNB with Bronze Age — check the .anno; era-pure source labels are worth more here than anywhere.
  • Modern reference temptation. Using present-day populations as sources for other moderns — the proximal shortcut — inherits the measured 72–100% screening FDR. The region's dense ancient record makes the distal frame available; use it.

A worked reading#

A Lebanese-ancestry target: p = 0.24, Levant_N 0.58 ± 0.027, Iran_N 0.28 ± 0.031, Anatolia_N 0.09 ± 0.025, Yamnaya-related 0.05 ± 0.014; nested three-source models rejected; ~680k SNPs; Iran_N/CHG swap re-run shifts weights within one SE (stability check passed). Reading: compatible and consonant with the published Bronze-Age-Levant-plus-steppe-trickle picture; the Anatolia term sits barely two SEs from zero, so the honest report flags it as the model's softest claim; and the swap test is the region's signature move — a result that survives proxy rotation is the only kind worth publishing here.

From €29.99 · one-time
The tested version of this question
A qpAdm model composed, run and checked by hand against AADR v66, published with its p-value, every source's standard error and z-score, and the full right set, so the result can be argued with.
See the qpAdm analysis

Terms used here are defined in the glossary.

References#

  • Lazaridis, I. et al. (2016). Genomic insights into the origin of farming in the ancient Near East. Nature, 536, 419–424. (The four-stream framework and O9.)
  • Agranat-Tamir, L. et al. (2020). The genomic history of the Bronze Age Southern Levant. Cell, 181(5), 1146–1157. (o9a and the threefold SE reduction.)
  • Feldman, M. et al. (2019). Late Pleistocene human genome suggests a local origin for the first farmers of central Anatolia. Nature Communications, 10, 1218.
  • Haber, M. et al. (2017). Continuity and admixture in the last five millennia of Levantine history. American Journal of Human Genetics, 101(2), 274–282.

Related posts

qpAdm models for South Asian ancestry: AASI, Indus Periphery and the proxy problem
qpAdm models for South Asian ancestry: AASI, Indus Periphery and the proxy problem

South Asia is qpAdm's hardest standard fixture: one ancestral stream has no ancient sample at all. The working recipe — Indus Periphery, steppe MLBA, the Onge-as-AASI-proxy problem — with right sets, failure modes and a worked reading.

4 min read
qpAdm models for European ancestry: the standard recipe and its regional variations
qpAdm models for European ancestry: the standard recipe and its regional variations

The three-source model that rebuilt European prehistory — WHG, Anatolian farmers, steppe pastoralists — as a working qpAdm recipe: exact source and right-set choices, regional adjustments, failure modes and a worked reading.

3 min read
Jewish ancestry and ancient DNA: what a qpAdm model can tell you
Jewish ancestry and ancient DNA: what a qpAdm model can tell you

What ancient genomes actually say about Jewish ancestry, why a consumer 'Ashkenazi Jewish' percentage answers a different question, and what a formal qpAdm model with a p-value shows for a Jewish genome.

9 min read
Back to all stories
Ancestrify

Combining cutting-edge genomic science with rich historical records to map your ancestry across generations and continents.


© 2026 Ancestrify. All rights reserved. · Ancestrify is a trading name of Andi Thomaj, a sole trader registered in Tiranë, Albania · NUIS M61725001N
Card payments processed by POK Payments (RPay Ltd)VISAMASTERCARD