Ancestrify
All stories

raw-dna

By Ancestrify
6 min read

Upload a whole-genome VCF for ancient-DNA ancestry

How to run an ancient-DNA ancestry analysis from a whole-genome sequencing VCF — tellmeGen, Dante Labs, Nebula and similar — what happens to the file at upload, why it is converted to the panel's markers, and what changes versus a chip export.

raw-dnawhole-genomevcfqpadmancient-matchesguide

  1. What a VCF is
  2. How the upload works
  3. What changes versus a chip export
  4. Things that go wrong
  5. Provider notes
  6. Which analyses accept it
  7. What a whole genome does not buy
  8. References

If you had your whole genome sequenced — with tellmeGen, Dante Labs, Nebula Genomics or a similar provider — you were handed a very different file from the one a chip test gives you. It is enormous, it is in a format most consumer ancestry tools reject, and it holds far more information than any of them use. This guide explains how to run an ancient-DNA ancestry analysis on it, what happens to the file when you upload it, and what actually changes compared with a chip export.

What a VCF is#

A VCF — Variant Call Format — lists the positions where your genome differs from a reference genome, with your genotype at each. Whole-genome sequencing produces millions of such positions, usually delivered as a compressed .vcf.gz file of several hundred megabytes; the specification was defined by Danecek and colleagues (2011). A chip export from 23andMe or AncestryDNA, by contrast, is a fixed list of roughly 600,000–700,000 positions the chip was designed to read, whether or not they differ from the reference.

Two consequences matter for ancestry analysis:

  • A VCF names a reference build. Most whole-genome providers deliver against GRCh38; most ancient-DNA reference panels — including the Allen Ancient DNA Resource — are on GRCh37. The same variant has different coordinates on each, so the file must be lifted from one to the other before it can be merged.
  • A VCF only lists differences. Positions where you match the reference may be absent, which means "reference genotype" and "no data" have to be told apart carefully during conversion.

How the upload works#

Whole-genome upload is a €10 add-on on two of our analyses — the qpAdm analysis and Ancient Matches — and accepts .vcf or .vcf.gz up to 1 GB. At upload:

  1. The file is read once and converted to the reference panel's markers — the roughly 1.2 million positions the ancient panel is genotyped at. Everything else in the VCF is discarded; an ancient-DNA analysis cannot use positions the ancient samples were never read at.
  2. Coordinates are lifted from GRCh38 to GRCh37 where the file declares GRCh38.
  3. The converted genotype set is what the analysis runs on. The VCF itself is not stored.

The result is a genotype file in the same shape a chip export produces — which is why everything downstream, from the merge to the report, is identical. The vendor-agnostic detail on what the pipeline detects and refuses is in the free Raw DNA File Check, which is the one place a VCF can be checked without ordering anything.

What changes versus a chip export#

This is the honest part, and it is less dramatic than whole-genome marketing suggests.

Coverage of the panel goes up — moderately. A chip export overlaps the ancient panel at a few hundred thousand positions, depending on the chip. A whole-genome VCF can cover nearly all of the panel's markers, because it read everywhere. More overlapping markers means more f-statistics computed from more data, which means smaller standard errors on the qpAdm weights — and standard error is the number that bounds what a model can resolve. A file that could only support a two-way model at SE 0.09 may support a three-way model at SE 0.05.

Nothing else changes. The reference panel is the same ancient individuals; the model is composed the same way; the p-value tests the same thing. A VCF does not unlock different sources, a different era or a different product. It is a denser measurement of the same genome.

Ancient Matches benefits most. Segment matching against individual ancient genomes is limited by how many informative markers fall inside each shared stretch. A denser file resolves shorter segments and reports marker density on every row — see Ancient DNA matches explained for what those segments mean and do not mean.

Things that go wrong#

  • Low-pass sequencing. Some "whole-genome" products are sequenced at low depth and imputed — the genotypes are statistical guesses at many positions. They still convert, but the standard errors reflect the imputation, not the sequencing.
  • Missing chromosomes. Some providers deliver autosomes only; the Y and mitochondrial chromosomes may be absent, which means the paternal and maternal haplogroup add-ons cannot be run from that file. The File Check reports row counts per chromosome class so you know before you pay.
  • Wrong build declared. A VCF that says GRCh37 but is on GRCh38 will merge at the wrong positions and produce a genome that looks like nobody. The converter reads the header; if the header is wrong, the result is wrong.
  • Files over 1 GB. Some providers ship uncompressed VCFs of several gigabytes. Compress to .vcf.gz first; if it is still over the limit, the file likely contains non-variant records that can be stripped with standard tools.

Provider notes#

The three providers we see most often deliver slightly different things, and the differences matter at upload:

  • Nebula Genomics delivers a whole-genome VCF on GRCh38 as .vcf.gz, typically well under the 1 GB limit; deep-coverage products convert cleanly, low-pass products convert with the imputation caveat above.
  • Dante Labs delivers a large VCF set, sometimes split into SNP and indel files; the SNP file is the one to upload, and it is usually on GRCh38 for recent orders and GRCh37 for older ones — check the header.
  • tellmeGen whole-genome orders deliver a VCF alongside the chip-style raw export; either works, and the VCF is the denser of the two.

In every case the free File Check will state the detected build, the row counts per chromosome class and a verdict per analysis before you commit a euro.

Which analyses accept it#

AnalysisWhole-genome VCFChip export
qpAdmyes, €10 add-onyes
Ancient Matchesyes, €10 add-onyes
Global25no — coordinates come from the Eurogenes serviceno — same
Raw DNA File Checkyes, freeyes, free

Global25 is the odd one out because no one but the independent Eurogenes service produces G25 coordinates, and their input rules are their own — see How to get your Global25 coordinates.

What a whole genome does not buy#

It does not buy certainty about ancestry. Standard errors shrink; the dependence of every weight on which sources and outgroups were chosen does not. A whole-genome qpAdm model still has to clear the same bar as a chip one, still has to be read in the order set out in How to read qpAdm results, and still never names anyone as an ancestor. What it buys is a denser file — which, for a method whose limiting quantity is coverage, is exactly the right thing to buy, and the report's model record shows it directly: the merged SNP count, the minimum SNP count per f4 statistic and the width of every source's 95% confidence interval. The buyer's guide to qpAdm covers the rest, and the terms are in the glossary.

From €29.99 · one-time
The tested version of this question
A qpAdm model composed, run and checked by hand against AADR v66, published with its p-value, every source's standard error and z-score, and the full right set, so the result can be argued with.
See the qpAdm analysis

References#

  • Danecek, P. et al. (2011). The variant call format and VCFtools. Bioinformatics, 27(15), 2156–2158.
  • Mallick, S. et al. (2024). The Allen Ancient DNA Resource (AADR): a curated compendium of ancient human genomes. Scientific Data, 11, 182.
  • Harney, É., Patterson, N., Reich, D. & Wakeley, J. (2021). Assessing the performance of qpAdm: a statistical tool for studying population admixture. Genetics, 217(4), iyaa045.

Related posts

qpAdm data preparation: from a raw DNA file to an AADR merge that works
qpAdm data preparation: from a raw DNA file to an AADR merge that works

The undocumented half of every qpAdm analysis: file formats, genome builds, strand hygiene, convertf and Poseidon, and the merge arithmetic that decides your standard errors before any model runs.

3 min read
G25 coordinates from a whole genome (WGS, VCF, BAM): the routes that work
G25 coordinates from a whole genome (WGS, VCF, BAM): the routes that work

Dante Labs, Nebula, Sequencing.com and other WGS files can become official Global25 coordinates — via the portal's own conversion surcharge or a free WGSExtract conversion. What each route costs and where the traps are.

3 min read
Where to buy a qpAdm analysis in 2026: every service, what it costs, what you get
Where to buy a qpAdm analysis in 2026: every service, what it costs, what you get

Every place that sells or offers qpAdm in 2026: hand-checked analyses, DIY tools and subscriptions, with inputs, reference data and price models side by side.

2 min read
Back to all stories
Ancestrify

Combining cutting-edge genomic science with rich historical records to map your ancestry across generations and continents.


© 2026 Ancestrify. All rights reserved. · Ancestrify is a trading name of Andi Thomaj, a sole trader registered in Tiranë, Albania · NUIS M61725001N
Card payments processed by POK Payments (RPay Ltd)VISAMASTERCARD