---
title: "What is archaeogenetics? Ancient DNA, the methods, and what it means for your own genome"
description: "A plain definition of archaeogenetics, the history from Pääbo's first sequences to the 2010 Neanderthal genome, Haak 2015 and the Allen Ancient DNA Resource, the methods (capture, contamination control, f-statistics, qpAdm, IBD), the big findings by continent, how consumer ancient-DNA services build on the field, and a reading list."
canonical: https://www.ancestrify.io/blog/what-is-archaeogenetics
date: 2026-09-16T06:00:00+00:00
updated: 2026-09-16
author: "Andi Thomaj"
---

# What is archaeogenetics? Ancient DNA, the methods, and what it means for your own genome

Archaeogenetics is the study of the human past through DNA recovered from archaeological
remains. Where archaeology reads pots, bones and buildings, and historical linguistics reads
words, archaeogenetics reads the genomes of the people themselves, sampled from teeth and
bone that may be forty thousand years old, and asks who they were related to, where their
ancestors came from and what happened when populations met. It is a young field; its first
whole genomes are barely fifteen years old, and it is the reason every consumer
"ancient DNA test" exists at all.

## A short history

The field's first decades were spent on tiny fragments. Svante Pääbo's early work in the
1980s and 1990s showed that DNA could survive in old tissue and, more importantly, showed how
easily a modern contaminant could masquerade as an ancient sequence; the discipline's
obsession with contamination control dates from those lessons. Mitochondrial sequences came
first, because a cell has hundreds of copies of that small genome and only two of each
nuclear chromosome.

The turning point was 2010. Green and colleagues published a draft nuclear genome of the
Neanderthal from three bones in Vindija Cave, and found that non-African people today carry
a small Neanderthal contribution. In the same year Reich and colleagues sequenced a finger
bone from Denisova Cave in Siberia and found it belonged to a group nobody had described from
fossils, whose genetic trace survives in Melanesians. A human population had been discovered
from its DNA before its skeleton.

The next five years scaled the method from single specimens to populations. Haak and
colleagues in 2015 assembled genome-wide data from 69 ancient Europeans and showed that a
massive movement of people from the Pontic-Caspian steppe, related to the Yamnaya, reshaped
the ancestry of Europe after 3000 BCE; the supplement to that paper introduced qpAdm, the
model-testing tool this site's paid product is built on. Since then the count of published
ancient genomes has climbed past ten thousand and the Allen Ancient DNA Resource, described
by Mallick and colleagues in 2024, gathers them on one shared set of positions so that any
laboratory, or any consumer service, can compare a new sample with all the old ones.

## The methods

**Extraction and capture.** Ancient DNA is short, damaged and mostly not human; a bone
sample may be 99% microbial. Laboratories extract from the densest bone, often the petrous
part of the temporal bone, and enrich the human fraction by capture: baits that fish out a
predetermined set of around 1.2 million informative positions. That is why most ancient
genomes are reported on the same SNP grid, and why the AADR can stack them.

**Authentication and contamination.** Real ancient molecules carry characteristic damage at
their ends, where cytosine has deaminated over the millennia, and the field uses that damage
signature, along with mitochondrial and X-chromosome heterozygosity checks in males, to
estimate how much of a sample is contaminant. A sample that fails is dropped, not corrected.

**f-statistics.** Patterson and colleagues in 2012 formalised a family of statistics that
measure shared genetic drift between populations from allele-frequency correlations. f3
detects admixture in a target from two sources, f4 and D test whether four populations form a
tree, and the f4-ratio estimates a mixing proportion. They work on low-coverage data, they
have standard errors from block jackknife, and they need no model of when things happened.

**qpAdm.** Built on f4 statistics, qpAdm asks whether a target population can be written as a
mixture of a set of sources, using a set of outgroups that separate the sources from one
another. It returns proportions with standard errors and z-scores, and a p-value that rejects
a model when the target shares drift with the outgroups in a way the sources cannot explain.
It is the workhorse of the field's admixture papers, and the tool behind
[Ancestrify's qpAdm analysis](/qpadm), which publishes those statistics for a customer's own file.

**Identity by descent.** Where coverage allows, long shared haplotypes between two genomes
point to a common ancestor within a few dozen generations. In ancient samples this has found
close relatives within cemeteries and, across populations, has dated migrations more finely
than allele frequencies can. On array data from consumers the comparable measurement is
identity by state, which is what [Ancient Matches](/ancient-matches) reports.

## The big findings, by continent

**Europe.** Three layers: Mesolithic hunter-gatherers, Anatolian-derived Neolithic farmers
from around 7000 BCE, and steppe pastoralists after 3000 BCE; nearly every present-day
European is a mixture of the three, in proportions that follow geography.

**West and Central Asia.** Distinct Neolithic populations in Anatolia, the Levant, the Zagros
and the Caucasus, later mixing along the trade routes; the Caucasus and Iranian-related
components that enter Europe in the Bronze Age.

**South Asia.** A cline between an Ancestral North Indian pole with Iranian-related and
steppe-related ancestry and an Ancestral South Indian pole related to the earliest
inhabitants, with the Indus Valley period as the hinge.

**East Asia and Siberia.** Deep structure between northern and southern populations, the
Ancient North Eurasians of the Siberian Upper Palaeolithic who contribute to both Europeans
and Native Americans, and the spread of millet and rice farmers.

**The Americas.** A founding population from Beringia, structured early into northern and
southern branches, with later Arctic movements; most published ancient genomes are from the
last few thousand years.

**Africa.** The deepest population structure on Earth, an ancient pastoralist expansion
from the northeast, the Bantu-related spread of farmers across the south, and a sampling
record still thinner than any other continent's, which limits what any consumer test can
say about African ancestry.

**Oceania.** Denisovan admixture in the ancestors of Papuans and Aboriginal Australians, and
the Lapita-associated expansion into remote Oceania a few thousand years ago.

## How consumer services build on it

Everything a consumer ancient-DNA service does is downstream of the published record. The
reference panel is the AADR or a derivative of it; the coordinate systems are PCAs of those
samples; the methods are f-statistics, qpAdm, distances in PCA space and identity by state.
What differs between services is what they publish. Ancestrify runs qpAdm against AADR v66,
merges the customer's raw file into that panel, and publishes the p-value, each source's
weight, standard error and z-score and the full right set, with a person composing and
checking the model; the [Global25 analysis](/g25) runs on the official coordinate row only,
which Ancestrify never computes itself; and the free [Lab](/lab) runs the distance,
admixture and PCA arithmetic in the browser. The [ancient ancestry test](/blog/ancient-ancestry-test)
post explains what any such result can and cannot claim, and the
[best ancient DNA test](/blog/best-ancient-dna-test-2026) guide compares the services.

## Careers and a reading list

Archaeogenetics is done in a handful of large laboratories (Harvard, Leipzig, Copenhagen,
Jena, Vienna, Tartu and others) and a growing number of smaller ones, by people trained in
population genetics, bioinformatics, archaeology or molecular biology. A master's in any of
those with a strong statistics component is the usual door; the software is open source
(ADMIXTOOLS 2, the AADR, the alignment and damage tools), so the analytical half of the
field can be learned on a laptop before a laboratory is ever entered. Reading, in order:
David Reich's *Who We Are and How We Got Here* for the arc; the five papers in the facts
block for the method; the AADR's own documentation for what a curated panel looks like; and
Patterson 2012 again, slowly, for why f4 is the atom everything else is built from.

## Frequently asked questions

### What is the difference between archaeogenetics and paleogenetics?

They overlap almost entirely. Paleogenetics tends to be used for the study of any ancient
DNA, including animals, plants and pathogens; archaeogenetics emphasises the human past and
its integration with archaeology. In practice the same laboratories do both.

### How old is the oldest human DNA sequenced?

Nuclear DNA has been recovered from hominin remains several hundred thousand years old, and
whole-genome data from anatomically modern humans reaches back around forty-five thousand
years. Preservation, not age alone, sets the limit: cold, dry and stable is what matters.

### Can archaeogenetics tell me who my ancestors were?

It can tell you which sampled ancient populations your genome resembles, and with qpAdm
whether a particular mixture of them is statistically consistent with your file. It cannot
name an individual ancestor; thousands of years back, your ancestors are a large share of
everyone then living where your lines come from.

### What is the AADR?

The Allen Ancient DNA Resource, a curated compendium of published ancient human genotypes on
a shared set of about 1.2 million positions, with metadata for every sample. Ancestrify's
qpAdm product runs against version 66 of it plus a documented supplement.

## Sources and further reading

- Green et al. 2010, A Draft Sequence of the Neandertal Genome, Science: https://doi.org/10.1126/science.1188021
- Reich et al. 2010, Genetic history of an archaic hominin group from Denisova Cave in Siberia, Nature: https://doi.org/10.1038/nature09710
- Patterson et al. 2012, Ancient Admixture in Human History, Genetics: https://doi.org/10.1534/genetics.112.145037
- Haak et al. 2015, Massive migration from the steppe was a source for Indo-European languages in Europe, Nature: https://doi.org/10.1038/nature14317
- Mallick et al. 2024, The Allen Ancient DNA Resource (AADR): A curated compendium of ancient human genomes, Scientific Data: https://doi.org/10.1038/s41597-024-03031-7

Terms used here are defined in the [glossary](/glossary).
