Open the AADR for the first time and the
population list reads like a code you were supposed to already know: Russia_Sintashta_MLBA,
Iran_GanjDareh_N, Han.DG, Sweden_Motala_HG.SG, French.HO. The words are archaeology; the
suffixes are data provenance — and they matter more than newcomers guess, because qpAdm's
statistics can be biased by exactly the processing differences the suffixes encode. This is the
decoder, plus the selection rules that follow from it.
What the suffixes encode#
- No suffix (most ancients): genotypes from the 1240K capture system — the targeted enrichment assay behind most published ancient genomes — typically pseudohaploid: one allele randomly drawn per site, because low coverage cannot call heterozygotes honestly.
.SG— shotgun sequencing, genotypes called from whole-genome reads rather than capture, usually pseudohaploid as well. Same individual sequenced both ways can appear twice, once plain and once.SG..DG— diploid genotypes: high-coverage genomes (ancient or modern) with real diploid calls. The famous reference individuals live here —Han.DG,Mbuti.DG,Ust_Ishim.DG..HO— genotyped on (or subset to) the Human Origins array, ~600K sites; the standard suffix for the modern reference populations, and the panel to be on when your analysis needs moderns and ancients on a shared, well-behaved site set.- The warning decorations: labels carrying markers like
_contam,_lc(low coverage),_o/outlier tags, or the dataset's explicit ignore/QC flags are the curators telling you a sample failed or strained assessment. The.annometadata file carries the per-sample details — coverage, SNP counts, dating, assessment notes, uniparentals — and reading it before using an unfamiliar population is the habit that separates careful models from lucky ones.
Why it matters: the differential-artefact rule#
Each production route leaves its own tiny systematic fingerprint — ancient-DNA damage profiles, capture bias, diploid-versus-pseudohaploid encoding. The audit literature's finding is the rule to memorise: qpAdm is robust when artefacts are shared (uniform damage barely moves estimates) and biased when they are differential — which is why the caution against co-analysing ancient and present-day data carelessly exists, and why a contrast where one side is capture-pseudohaploid and the other shotgun-diploid can carry a technical signal wearing an ancestry costume.
The practical selection rules that fall out:
- Within a contrast, prefer one data class. Sources being compared against each other — and the right-set members meant to split them — should share production class where the dataset allows it.
- Ancients with ancients, moderns deliberately. When a model genuinely needs modern
references, take them from the
.HOpanel and keep them on the right, not mixed into ancient source roles. - Match your own kit's class. A consumer genotype merged into the panel behaves best compared like-for-like — the merge guide covers the encoding choice.
- Duplicated individuals: pick one. Where a sample appears plain and
.SG, take the version matching your contrast's class — never both, which double-counts one genome. - Respect the flags. Contaminated and ignore-listed samples are excluded from our production panel builds entirely; a hobby model gains nothing by re-including what the curators benched.
Labels are hypotheses too#
One level up from suffixes: the population label itself is a curator's grouping, and
group labels can pool centuries or split arbitrarily.
The .anno file again is the referee — check that the individuals under a label share the
period and place your model assumes, and prefer site-and-period-specific labels over grand
umbrella groupings for source roles. Our own catalog's rule (every source population one real
archaeological context, never an umbrella) exists because the difference is visible in model
stability.
None of this is exotic once seen: the suffixes are just the dataset being honest about how each genome was made, and the rules are one principle worn five ways — never let a processing difference sit where your model expects an ancestry difference.
Terms used here are defined in the glossary.
References#
- Mallick, S. et al. (2024). The Allen Ancient DNA Resource (AADR): a curated compendium of ancient human genomes. Scientific Data, 11, 182.
- Harney, É., Patterson, N., Reich, D. & Wakeley, J. (2021). Assessing the performance of qpAdm. Genetics, 217(4), iyaa045. (Damage robustness and the differential caution.)
- Lazaridis, I. et al. (2014). Ancient human genomes suggest three ancestral populations for present-day Europeans. Nature, 513, 409–413. (The Human Origins array's role.)



