The PCA scatter is the most shared artefact in amateur population genetics — your dot among the ancients, screenshot, caption, argument. It is also the most misread, because a PCA plot performs an amputation nobody sees happen: 25 dimensions become 2, and every intuition you form lives in the 2 that remain. Reading one well means knowing, at every moment, what the amputation removed.
What the axes are#
A principal component is the direction along which the plotted samples vary most; PC2 the next, at right angles; and so on. Two properties matter for reading. The axes belong to the samples — they are recomputed facts about a dataset, not fixed features of the world. And they are ranked by variance, not meaning — PC1 of a worldwide panel separates the deepest structure, but nothing guarantees any axis corresponds to a migration, a population, or anything nameable. The axis labels some tools add ("east–west cline") are interpretations, not measurements.
Which is why the same coordinate produces different-looking plots in different views: a West Eurasian view spends its two axes on West Eurasian structure; a worldwide view spends PC1–PC2 on continental splits and crushes Europe into a corner. Neither is wrong. They are different slices of the same 25-dimensional object.
Plot distance lies, in one specific way#
Two points close on the chart are close in the two plotted dimensions. They can be far apart in the other 23 — and routinely are. The classic trap is the projected ancient sample that lands "on top of" a modern population in PC1–PC2 while sitting nowhere near it in full distance. The cure is mechanical: any time plot proximity starts carrying a conclusion, check the actual 25-dimensional distance. A PCA plot is for seeing structure; the distance ranking is for measuring closeness. Confusing the two jobs is the root of most PCA-based wrong conclusions.
The second lie is subtler: a point between two clusters is not necessarily a mixture of them. Intermediate position is consistent with admixture, with belonging to an unsampled third population, or with sitting off-plane in dropped dimensions. Mixture is a model claim — that is what an admixture fit proposes and what qpAdm actually tests.
Projected views versus fitted views#
Our free PCA viewer works in the two modes every serious tool distinguishes:
Era views project. The reference space was computed once over that era's published samples; your pasted row lands in it without moving anything. Positions are comparable across sessions and across people, which makes projected views the right mode for "where do I fall among the ancients" — and the only mode where screenshots from different people belong on one argument.
Custom views fit. Paste your own set of rows and the PCA is recomputed from exactly those rows — add or remove one and every axis can rotate. That is not a defect; it is the right tool for comparing a handful of samples against each other on their own axes of variation. It just means a custom plot is a statement about your input set, never about any fixed reference space, and two custom plots are never comparable.
Swapping the modes silently is the classic tooling mistake — a self-fitted plot read as if it were the standard West Eurasian view supports conclusions it cannot carry.
An honest reading routine#
- Name the view. Which samples, which era, projected or fitted. A screenshot without that caption is unreadable, including your own from last month.
- Read clusters, not points. Population structure is the clouds and clines; a single dot's exact pixel is noise wearing precision.
- Cross-check any conclusion in full distance. One paste into the distance tool settles what the plot only suggests.
- Let mixture claims graduate. Betweenness on a plot → an era-scoped fit with a stated fit distance → and if the claim has to survive argument, a method with a p-value.
Read this way, a G25 PCA is genuinely excellent — the fastest structure-viewer the hobby has, and the paid Global25 analysis leans on exactly these projected era views beside its distances and admixture chapters. The dot is fine. The caption is what separates seeing from fooling yourself.
Terms used here are defined in the glossary.
References#
- Patterson, N., Price, A. L. & Reich, D. (2006). Population structure and eigenanalysis. PLoS Genetics, 2(12), e190.
- Novembre, J. et al. (2008). Genes mirror geography within Europe. Nature, 456, 98–101.
- McVean, G. (2009). A genealogical interpretation of principal components analysis. PLoS Genetics, 5(10), e1000686.



