Genetic risk models trained with large European datasets can become less accurate for another population once enough local data is available, according to a Google Research study. The finding challenges the assumption that adding a larger outside cohort will always improve prediction.
Researchers examined polygenic risk scores, which combine the effects of many genetic variants to estimate disease-related traits. They compared UK Biobank data with nearly 200,000 participants from Biobank Japan across eight measures, including blood pressure, cholesterol, blood glucose and blood-cell counts.
European data provided a useful baseline when the Japanese training sample was very small. But population-specific models generally pulled ahead at target-cohort sizes of 15,000 or more. For HDL cholesterol, adding more than 5,000 UK samples beyond that crossover reduced performance compared with training only on Japanese data. The same broad pattern appeared across all eight traits.
Differences in genetic architecture, population structure and variant frequency help explain why knowledge does not transfer evenly. The work does not make these scores ready for routine clinical decisions; their adoption remains limited, and the study covers a particular pair of biobanks and selected traits. Its practical lesson is narrower: researchers should measure when outside data stops helping and invest in sufficiently large, representative local cohorts rather than treating European datasets as universally beneficial.