Skip to content
Ancestrify
All stories

archaeogenetics

By Andi Thomaj
7 min read

What is archaeogenetics? Ancient DNA, the methods, and what it means for your own genome

A plain definition of archaeogenetics, the history from Pääbo's first sequences to the 2010 Neanderthal genome, Haak 2015 and the Allen Ancient DNA Resource, the methods (capture, contamination control, f-statistics, qpAdm, IBD), the big findings by continent, how consumer ancient-DNA services build on the field, and a reading list.

archaeogeneticsancient-dnaguideqpadmaadr

  1. A short history
  2. The methods
  3. The big findings, by continent
  4. How consumer services build on it
  5. Careers and a reading list
  6. Frequently asked questions
  7. What is the difference between archaeogenetics and paleogenetics?
  8. How old is the oldest human DNA sequenced?
  9. Can archaeogenetics tell me who my ancestors were?
  10. What is the AADR?
  11. Sources and further reading

Archaeogenetics is the study of the human past through DNA recovered from archaeological remains. Where archaeology reads pots, bones and buildings, and historical linguistics reads words, archaeogenetics reads the genomes of the people themselves, sampled from teeth and bone that may be forty thousand years old, and asks who they were related to, where their ancestors came from and what happened when populations met. It is a young field; its first whole genomes are barely fifteen years old, and it is the reason every consumer "ancient DNA test" exists at all.

Key facts

Figures as published, one source per row.

Neanderthal ancestry in living peopleThe 2010 draft Neanderthal genome showed that present-day people outside Africa carry about 1 to 4% Neanderthal-derived ancestry, the first direct genomic evidence of interbreeding.Source: A Draft Sequence of the Neandertal Genome (2010), Green et al.
A population known first from DNAA finger bone from Denisova Cave yielded the genome of a previously unknown archaic group, the Denisovans, whose ancestry survives at about 4 to 6% in present-day Melanesians.Source: Genetic history of an archaic hominin group from Denisova Cave in Siberia (2010), Reich et al.
The f-statistics toolkitThe f3, f4 and D statistics and the ADMIXTOOLS software, which formalised how admixture is detected and dated from allele-frequency correlations, were introduced in 2012 and remain the base of qpAdm.Source: Ancient Admixture in Human History (2012), Patterson et al.
The steppe migrationGenome-wide data from 69 ancient Europeans showed that Late Neolithic Corded Ware people derived about three quarters of their ancestry from Yamnaya-related steppe populations; the same paper's supplement introduced qpAdm.Source: Massive migration from the steppe was a source for Indo-European languages in Europe (2015), Haak et al.
One curated reference setThe Allen Ancient DNA Resource collects the published genotypes of more than ten thousand ancient individuals on one shared set of about 1.2 million SNP positions, with dates, locations and quality metadata for each sample.Source: The Allen Ancient DNA Resource (AADR): A curated compendium of ancient human genomes (2024), Mallick et al.
What Ancestrify reportsAn era-by-era qpAdm model of your raw DNA against these ancient sources, published with its p-value, every source's standard error and the full right set, from from €29.99.See the qpAdm analysis

A short history#

The field's first decades were spent on tiny fragments. Svante Pääbo's early work in the 1980s and 1990s showed that DNA could survive in old tissue and, more importantly, showed how easily a modern contaminant could masquerade as an ancient sequence; the discipline's obsession with contamination control dates from those lessons. Mitochondrial sequences came first, because a cell has hundreds of copies of that small genome and only two of each nuclear chromosome.

The turning point was 2010. Green and colleagues published a draft nuclear genome of the Neanderthal from three bones in Vindija Cave, and found that non-African people today carry a small Neanderthal contribution. In the same year Reich and colleagues sequenced a finger bone from Denisova Cave in Siberia and found it belonged to a group nobody had described from fossils, whose genetic trace survives in Melanesians. A human population had been discovered from its DNA before its skeleton.

The next five years scaled the method from single specimens to populations. Haak and colleagues in 2015 assembled genome-wide data from 69 ancient Europeans and showed that a massive movement of people from the Pontic-Caspian steppe, related to the Yamnaya, reshaped the ancestry of Europe after 3000 BCE; the supplement to that paper introduced qpAdm, the model-testing tool this site's paid product is built on. Since then the count of published ancient genomes has climbed past ten thousand and the Allen Ancient DNA Resource, described by Mallick and colleagues in 2024, gathers them on one shared set of positions so that any laboratory, or any consumer service, can compare a new sample with all the old ones.

The methods#

Extraction and capture. Ancient DNA is short, damaged and mostly not human; a bone sample may be 99% microbial. Laboratories extract from the densest bone, often the petrous part of the temporal bone, and enrich the human fraction by capture: baits that fish out a predetermined set of around 1.2 million informative positions. That is why most ancient genomes are reported on the same SNP grid, and why the AADR can stack them.

Authentication and contamination. Real ancient molecules carry characteristic damage at their ends, where cytosine has deaminated over the millennia, and the field uses that damage signature, along with mitochondrial and X-chromosome heterozygosity checks in males, to estimate how much of a sample is contaminant. A sample that fails is dropped, not corrected.

f-statistics. Patterson and colleagues in 2012 formalised a family of statistics that measure shared genetic drift between populations from allele-frequency correlations. f3 detects admixture in a target from two sources, f4 and D test whether four populations form a tree, and the f4-ratio estimates a mixing proportion. They work on low-coverage data, they have standard errors from block jackknife, and they need no model of when things happened.

qpAdm. Built on f4 statistics, qpAdm asks whether a target population can be written as a mixture of a set of sources, using a set of outgroups that separate the sources from one another. It returns proportions with standard errors and z-scores, and a p-value that rejects a model when the target shares drift with the outgroups in a way the sources cannot explain. It is the workhorse of the field's admixture papers, and the tool behind Ancestrify's qpAdm analysis, which publishes those statistics for a customer's own file.

Identity by descent. Where coverage allows, long shared haplotypes between two genomes point to a common ancestor within a few dozen generations. In ancient samples this has found close relatives within cemeteries and, across populations, has dated migrations more finely than allele frequencies can. On array data from consumers the comparable measurement is identity by state, which is what Ancient Matches reports.

The big findings, by continent#

Europe. Three layers: Mesolithic hunter-gatherers, Anatolian-derived Neolithic farmers from around 7000 BCE, and steppe pastoralists after 3000 BCE; nearly every present-day European is a mixture of the three, in proportions that follow geography.

West and Central Asia. Distinct Neolithic populations in Anatolia, the Levant, the Zagros and the Caucasus, later mixing along the trade routes; the Caucasus and Iranian-related components that enter Europe in the Bronze Age.

South Asia. A cline between an Ancestral North Indian pole with Iranian-related and steppe-related ancestry and an Ancestral South Indian pole related to the earliest inhabitants, with the Indus Valley period as the hinge.

East Asia and Siberia. Deep structure between northern and southern populations, the Ancient North Eurasians of the Siberian Upper Palaeolithic who contribute to both Europeans and Native Americans, and the spread of millet and rice farmers.

The Americas. A founding population from Beringia, structured early into northern and southern branches, with later Arctic movements; most published ancient genomes are from the last few thousand years.

Africa. The deepest population structure on Earth, an ancient pastoralist expansion from the northeast, the Bantu-related spread of farmers across the south, and a sampling record still thinner than any other continent's, which limits what any consumer test can say about African ancestry.

Oceania. Denisovan admixture in the ancestors of Papuans and Aboriginal Australians, and the Lapita-associated expansion into remote Oceania a few thousand years ago.

How consumer services build on it#

Everything a consumer ancient-DNA service does is downstream of the published record. The reference panel is the AADR or a derivative of it; the coordinate systems are PCAs of those samples; the methods are f-statistics, qpAdm, distances in PCA space and identity by state. What differs between services is what they publish. Ancestrify runs qpAdm against AADR v66, merges the customer's raw file into that panel, and publishes the p-value, each source's weight, standard error and z-score and the full right set, with a person composing and checking the model; the Global25 analysis runs on the official coordinate row only, which Ancestrify never computes itself; and the free Lab runs the distance, admixture and PCA arithmetic in the browser. The ancient ancestry test post explains what any such result can and cannot claim, and the best ancient DNA test guide compares the services.

Careers and a reading list#

Archaeogenetics is done in a handful of large laboratories (Harvard, Leipzig, Copenhagen, Jena, Vienna, Tartu and others) and a growing number of smaller ones, by people trained in population genetics, bioinformatics, archaeology or molecular biology. A master's in any of those with a strong statistics component is the usual door; the software is open source (ADMIXTOOLS 2, the AADR, the alignment and damage tools), so the analytical half of the field can be learned on a laptop before a laboratory is ever entered. Reading, in order: David Reich's Who We Are and How We Got Here for the arc; the five papers in the facts block for the method; the AADR's own documentation for what a curated panel looks like; and Patterson 2012 again, slowly, for why f4 is the atom everything else is built from.

Frequently asked questions#

What is the difference between archaeogenetics and paleogenetics?#

They overlap almost entirely. Paleogenetics tends to be used for the study of any ancient DNA, including animals, plants and pathogens; archaeogenetics emphasises the human past and its integration with archaeology. In practice the same laboratories do both.

How old is the oldest human DNA sequenced?#

Nuclear DNA has been recovered from hominin remains several hundred thousand years old, and whole-genome data from anatomically modern humans reaches back around forty-five thousand years. Preservation, not age alone, sets the limit: cold, dry and stable is what matters.

Can archaeogenetics tell me who my ancestors were?#

It can tell you which sampled ancient populations your genome resembles, and with qpAdm whether a particular mixture of them is statistically consistent with your file. It cannot name an individual ancestor; thousands of years back, your ancestors are a large share of everyone then living where your lines come from.

What is the AADR?#

The Allen Ancient DNA Resource, a curated compendium of published ancient human genotypes on a shared set of about 1.2 million positions, with metadata for every sample. Ancestrify's qpAdm product runs against version 66 of it plus a documented supplement.

Sources and further reading#

From €29.99 · one-time
The tested version of this question
A qpAdm model composed, run and checked by hand against AADR v66, published with its p-value, every source's standard error and z-score, and the full right set, so the result can be argued with.
See the qpAdm analysis

Terms used here are defined in the glossary.


Related posts

AADR explained: the Allen Ancient DNA Resource
AADR explained: the Allen Ancient DNA Resource

What the Allen Ancient DNA Resource is, who curates it, what a version like v66 contains, the difference between the 1240K and Human Origins panels, and how the dataset becomes the reference behind an ancient-DNA ancestry analysis.

6 min read
Ancient ancestry test: how a DNA test finds the ancient peoples in your genome
Ancient ancestry test: how a DNA test finds the ancient peoples in your genome

What an ancient ancestry test measures, the three methods behind every commercial one (qpAdm modelling, Global25 distances and admixture, shared-segment matching against ancient individuals), the eras a result is split into, how to read one honestly, and which providers do what, checked on their own sites on 2026-09-15.

10 min read
Best ancient DNA test in 2026: every service that models your raw file against ancient genomes, compared
Best ancient DNA test in 2026: every service that models your raw file against ancient genomes, compared

The services that turn a 23andMe, AncestryDNA or whole-genome file into an ancient-ancestry result in 2026, compared on the criteria that decide the answer: a test with a p-value or a ranking, official or simulated Global25 coordinates, published source panels, whole-genome input, and the price model as each site states it.

10 min read
Back to all stories
Ancestrify

Combining cutting-edge genomic science with rich historical records to map your ancestry across generations and continents.


© 2026 Ancestrify. All rights reserved. · Ancestrify is a trading name of Andi Thomaj, a sole trader registered in Tiranë, Albania · NUIS M61725001N
Card payments processed by POK Payments (RPay Ltd)VISAMASTERCARD