A 23andMe raw data file is one of the most common inputs we see for a qpAdm analysis, and it is a good one. It is not a special case: qpAdm does not care which company genotyped you, only which positions your file carries and how many of them survive the merge with the ancient reference panel. This guide covers the 23andMe half of that route, from the download menu to the report, and it tries to be exact about what changes with your chip version and what does not. If you already have the file and want the tiers, they are on the buying page.
Step one: download the raw file#
23andMe does not hand the file over instantly. In your account, open Settings, scroll to
23andMe Data, and choose Download Raw Data. The site asks you to confirm the request and
then sends an email when the file is ready; follow the link in that email to fetch a .zip
holding a single .txt file. Keep the zip as it comes. Our uploader reads .zip, .gz and the
unpacked .txt alike, so there is nothing to unpack, rename or convert.
Two things worth knowing about that file before you upload it anywhere:
- Every line is one position: an rsID, a chromosome, a base-pair position and your two alleles, written in the forward orientation of the GRCh37 build.
- Positions the chip could not call are written as
--. These are not errors; they are simply absent from any analysis, which is why a file's usable marker count is always lower than its line count.
v3, v4, v5: which chip you have and why it matters#
23andMe has shipped several chip generations, and the generation decides how many positions your file shares with the ancient panel. Roughly:
- v3 (2010 to 2013) sat on a large Illumina OmniExpress-derived design with approximately 960k positions and strong overlap with the 1240k capture set that most ancient genomes were sequenced against.
- v4 (2013 to 2017) moved to a custom design of around 600k positions and a different selection philosophy, so a v4 file often overlaps the ancient panel less than a v3 file does despite being newer.
- v5 (2017 onward) is built on Illumina's Global Screening Array, approximately 640k positions. Its overlap with the ancient panel is respectable but not the largest of the three.
These numbers are approximate and the chip designs have been revised within versions, so treat them as orientation rather than a promise. The only number that matters for your model is the one produced after the merge, and that is something you can measure before paying.
What to check first: the free file check#
Upload the zip to the free Raw DNA File Check. It parses the file, detects the format and chip generation, counts usable markers per chromosome, and returns a coverage verdict for qpAdm: how many of your positions intersect the Allen Ancient DNA Resource (AADR) v66 panel, the same panel every paid model is run on. Nothing is ordered and nothing is stored.
That verdict is the honest answer to "will my file work". A file can be perfectly valid and still be a poor qpAdm input if its overlap with the ancient panel is small, and it is better to learn that from a free check than from a report whose standard errors cannot clear the bar.
What a qpAdm order does with the file#
Once you order, your genotypes are converted, filtered of indels and strand-ambiguous positions, and intersected against AADR v66 with Poseidon's trident. The result is a merged sample in which your file and roughly 23,265 ancient and modern reference samples are read at exactly the same positions. From that point the vendor name is gone; your sample is just a target.
An analyst then composes models by hand: a small set of ancient source populations, a set of distant outgroups, and a run of qpAdm that returns a weight, a standard error and a Z-score per source and a single p-value for the whole model. The report publishes all of it for each of two eras, together with the complete outgroup list and the full run record. How those numbers are read is set out in How to read qpAdm results.
What changes with coverage: the standard errors#
The standard error on each weight is bounded by how many positions the model could use. A v3 file with a large overlap will usually produce tighter errors than a v4 file from the same person; a v5 file sits between them. This is not something analyst effort can change. The publish bar is the same for every file and every tier: p > 0.05, and for every source in every era, |Z| > 3 and a standard error below 0.10. When a file's coverage fixes the errors above that line, the model is not published, and the file check is where you find that out before you spend anything.
What does not change: the p-value logic#
Coverage does not change what a p-value means. A passing model is one the data could not refute; a failing one is refuted. That logic is identical whether the target came from a v3 chip, a v5 chip or a whole-genome VCF. Lower coverage makes the test less able to distinguish between close alternatives, which shows up as wider errors, not as a different kind of answer. If you have ever seen a percentage breakdown that never fails, that is a different method entirely; the difference is laid out in qpAdm vs Global25.
Ordering#
The buying page lists the four tiers, from 29.99 EUR. Every tier is held to the same publish bar and returns the same report; what a deeper tier buys is a longer search past the first model that clears the bar, so that more alternatives have been tried and rejected before one is put in front of you. Upload the same zip you ran through the file check, pick the region you want the model scoped to, and the merge starts on our infrastructure. A qpAdm report is reviewed work, so it does not render the moment you pay.
If your 23andMe kit is old enough to be v2 or earlier, or if you sequenced your whole genome
elsewhere, a .vcf or .vcf.gz up to 1 GB is accepted instead through the 10 EUR whole-genome
upload, and a sequencing file usually covers more of the panel than any chip.
Afterwards: the Model Lab#
Once your report is published, a one-time 10 EUR unlock opens the Model Lab, where you run your own qpAdm models on your own merged sample. Your 23andMe genotypes are already sitting in the panel, so you choose sources and outgroups and run up to 100 models per rolling 24 hours. Most of them will fail, which is the method working. You can also download the exact EIGENSTRAT bundle the report was computed from and reproduce everything in ADMIXTOOLS 2 on your own machine. The walkthrough is in Run your own qpAdm models.
And if you want to feel a rejection before ordering anything, the free AdmixTools 2 Lab runs real qpAdm over the public reference panel in your browser. It cannot use your sample, but it teaches what a failing model looks like, which is the most useful thing to know before reading your own.
A short summary#
- Settings, 23andMe Data, Download Raw Data, wait for the email, keep the zip.
- Run the zip through the free file check and read the coverage verdict.
- Order the qpAdm analysis at whichever tier matches how hard you want the search to be.
- Read the report with the p-value, errors and outgroup list in view, not the percentages alone.
- Unlock the Model Lab if you want to run the alternatives yourself.
Terms used here are defined in the glossary.



