Why Global25 is one of the strongest ancestry-modelling systems
Global25 was developed by David Wesolowski, better known as Davidski, through the Eurogenes project. He created the original worldwide PCA and the projection system used to turn processed DNA into official Global25 coordinates. The framework is particularly strong for open-ended consumer ancestry modelling because one projected target can be tested against a large and continually expanding collection of modern and ancient references.
Compared with many other third-party DNA-analysis services, AncestralGenome uses Global25 with nMonte, an optimization algorithm that tests weighted combinations of reference populations and searches for the combination with the lowest distance to the customer’s G25 target. The calculation is only historically meaningful when the reference panel contains suitable populations from the customer’s region and period, so AncestralGenome selects its sources using the documented population history of that background rather than applying one generic panel to everyone. The methodology combines official Global25 coordinates, reproducible nMonte modelling, historically selected reference panels and visible genetic distances and model fit. These elements make the result independently checkable, unlike services that do not disclose the method or fit behind their percentages. AncestralGenome also gives customers their own G25 coordinates. They can compare those coordinates independently with published modern reference samples and check whether the populations at the lowest genetic distances agree with the closest populations shown in their report. Our expertise therefore does not depend on asking customers to trust a hidden proprietary system. It lies in building historically appropriate ancestry models: researching the documented population history of each subregion, evaluating the available ancient samples and selecting reference panels that answer the right question for that customer. Some of these panels require months of research and build on modelling practices that have been tested and refined within the Global25 community for years. When Global25 and nMonte are used carefully with appropriate references, their broad ancestry patterns are consistent with results obtained through formal academic methods such as qpAdm. Global25 remains a community-developed framework rather than a peer-reviewed scientific instrument, but reproducible nMonte modelling, historically informed reference panels and broad agreement with formal population-genetic methods make it one of the strongest systems available for third-party ancestry modelling.
This article explains what Global25 is, why it can produce highly accurate ancestry comparisons, where its limitations lie and how AncestralGenome uses official coordinates, curated reference panels and separate historical models to build its ancestry reports.
From hundreds of thousands of markers to 25 coordinates
A consumer DNA file contains results for hundreds of thousands of genetic markers. Comparing those markers one by one is not a useful way to describe the larger population structure carried by the sample. Global25 places the processed DNA on a principal component analysis built from roughly 300,000 markers and records its position along 25 axes of genetic variation.
The result is a row of 25 numbers. Each coordinate describes the sample’s position on one axis; the complete set describes its position inside the Global25 reference space. A two-dimensional PCA image can show only the first two axes. Global25 distance calculations use all 25, retaining information that disappears from a flat PC1-versus-PC2 plot.
Scaled and unscaled Global25
Global25 is distributed in scaled and unscaled form. Scaled coordinates weight the principal components according to the variation represented by each axis. This makes geometric distances more useful for comparing genetic differentiation and generally produces more stable ancestry models.
Scaled targets must be analysed with scaled references. Unscaled targets must be analysed with unscaled references. Mixing the two conventions places the target and sources in different mathematical spaces, so the resulting distances and percentages are invalid. AncestralGenome uses scaled Global25 coordinates and scaled references throughout the ancestry report.
What genetic distance measures
Genetic distance measures how far two Global25 positions lie from one another across all 25 coordinates. Global25 uses Euclidean distance: the difference on every axis is squared, those values are added together and the square root gives the final distance.
A lower value indicates greater overall similarity inside this reference space. It does not represent an ancestry percentage, and the nearest population is not automatically an ancestral source. A person with ancestry from two different regions may sit between them and appear closest to a population occupying a similar intermediate position. The match will often still have a somewhat higher distance than the distances normally seen among people genuinely belonging to that reference population, because the person’s mixture proportions and ancestral sources are unlikely to reproduce its genetic profile exactly.
Closest-population results are therefore best read as a similarity ranking and a plausibility check. They are especially informative when suitable regional references are available, but national averages can hide local variation and some regions remain more densely sampled than others.
Population averages also have limits. A German whose family comes from near the Czech border, for example, may be genetically closer to neighbouring Czech populations than to the national German average. That same person may nevertheless match closely with regional German references from the specific border area their family comes from. AncestralGenome therefore uses regional and subpopulation references where sufficient data are available instead of relying only on country-wide averages. An unexpected national match can reflect genuine local population structure, sparse regional sampling or an average that is too broad; it does not by itself show that the coordinates are wrong.
How Global25 ancestry models work
An ancestry model asks a different question from a distance table. Instead of comparing the target with one population at a time, an optimizer searches for a weighted combination of selected source populations whose combined coordinates reproduce the target as closely as possible. The percentages are the weights assigned to those sources; the remaining distance is the model fit.
A low fit shows that the chosen sources can reproduce the target mathematically. It does not prove that the historical explanation is correct. Closely related or chronologically mismatched sources may overfit the target while producing a mixture that makes little archaeological sense.
Southern Italians provide a clear example. A deep prehistoric panel might contain Anatolian Neolithic farmers, steppe-related ancestry and hunter-gatherer sources. Adding Mycenaeans to that same panel offers the optimizer a later population already formed largely from overlapping older ancestry. The same genetic signal is then presented twice at different historical depths. Reliable modelling keeps sources broadly contemporary, genetically distinguishable and relevant to the region being examined.
AncestralGenome therefore builds separate reference panels for specific ancestry regions using the known population history of each place. A North African model, for example, can include Iberian and Italian sources where historical and genetic evidence supports movement across the western Mediterranean. It would not introduce a Balkan source without comparable evidence of a relevant migration into the population being modelled. This prevents the optimizer from selecting a mathematically convenient source that has no credible place in that population’s history.
The optimizer matters as well. Tools such as nMonte search for the source weights that reduce the distance to the target, but their settings can change the answer. Distance penalties, the number of sources allowed and the decision to remove or retain a weak source can all alter the final percentages. The reference panel and the modelling settings must therefore be evaluated together.
From coordinates to a historical model
Consider a customer with western Sicilian ancestry. The genetic-data file downloaded from the company where the customer tested is converted into 25 scaled Global25 coordinates. AncestralGenome then uses those coordinates as the target for several separate parts of the ancestry report.
The Hunter-Gatherer & Farmer analysis models the customer with a regional source panel selected for the deep ancestry of Sicily. It can show a large Anatolian Neolithic component alongside steppe-related ancestry and smaller Natufian, Iranian Neolithic, Caucasus hunter-gatherer and Western hunter-gatherer components. The percentages and fit belong specifically to this deep model.
Periodical Breakdown uses different source panels for the Bronze Age, Iron Age, Roman Era and Medieval period. A western Sicilian can therefore receive an Iron Age result containing Italic, Phoenician and native Sicilian references, followed by a different result when Roman-period populations are used. Each period answers which populations from that particular era best model the same customer.
Why separate historical periods produce different results
Periodical Breakdown applies separate reference panels to the same target. The Hunter-Gatherer & Farmer model describes deep components. Bronze Age, Iron Age, Roman Era and Medieval models instead use complete populations living during each period. The target does not change; the historical question and available reference populations do.
Migration continually changed the genetic composition of populations between these periods. A Roman population inherited ancestry from Iron Age peoples but could also contain ancestry introduced by later movement across the Roman world. Iron Age populations likewise descended from earlier farmers, hunter-gatherers and steppe groups whose ancestry had spread through previous migrations. Separate panels show which populations from each period best model the customer’s DNA as those population mixtures changed through time.
The period results are therefore separate models of the same measurement, not independent ancestry amounts that should add up across eras. A customer can be modelled one way with Bronze Age references and another way with Roman-era references without the underlying DNA changing.
What determines the quality of a Global25 result?
Ancient genomes vary greatly in coverage. In Davidski’s stated practice, an ancient sample generally needed overlap with roughly 15% of the approximately 300,000 markers used by Global25 before coordinates were published. Low-coverage samples were flagged, unusual outliers were marked with suffixes and samples showing contamination were avoided. These labels matter because a poorly covered or contaminated ancient genome can distort both distances and ancestry models. AncestralGenome considers these quality indicators when selecting the ancient reference samples used in its historical panels and avoids relying on unsuitable samples simply because coordinates are available.
Coordinate provenance matters too. Davidski created Global25 and produces the official coordinates by projecting processed genotype data onto the original Global25 PCA. Other services may simulate or approximate G25-style coordinates, but those estimates are not official Global25 projections and are not compatible with the official datasheets in the same way. AncestralGenome therefore works directly with Davidski for raw-DNA conversion, so the coordinates used in the report come from the original and verifiable Global25 projection rather than a third-party imitation.