GenPo Science All articles
Clinical Genetics

Scores Without Certainty: The Hidden Limitations of Polygenic Risk Prediction

GenPo Science
Scores Without Certainty: The Hidden Limitations of Polygenic Risk Prediction

When a patient receives a polygenic risk score indicating elevated susceptibility to heart disease, the number feels authoritative—a distillation of thousands of genetic variants into a single, actionable figure. Increasingly, health systems across the United States are incorporating these scores into preventive care workflows, and direct-to-consumer genomics companies are packaging them as windows into future health. The enthusiasm is understandable. The caution, however, is essential.

Polygenic risk scores, commonly abbreviated as PRS, represent one of the most significant methodological advances in complex disease genetics over the past decade. But understanding what these scores can and cannot tell us requires a close look at how they are constructed—and who was in the room when the data were collected.

How a Polygenic Risk Score Is Built

Most common diseases—type 2 diabetes, schizophrenia, breast cancer, atrial fibrillation—are not caused by a single gene variant. They emerge from the cumulative, often interacting effects of hundreds or thousands of small-effect genetic variants scattered across the genome. A polygenic risk score attempts to aggregate this complexity into a single weighted sum.

The construction process begins with a genome-wide association study, or GWAS, in which researchers scan the genomes of large numbers of individuals—cases with the disease and controls without it—searching for single nucleotide polymorphisms, or SNPs, that appear more frequently in one group than the other. Each SNP that clears a statistical threshold is assigned a weight corresponding to the magnitude of its association with the trait. A PRS for a given individual is then calculated by summing those weighted contributions across all included variants.

The resulting score places a person somewhere on a distribution. Those in the uppermost percentiles are considered at elevated risk; those in the lower percentiles are considered relatively protected. In research settings, well-constructed PRS models have demonstrated genuine predictive power for several conditions. For familial hypercholesterolemia and coronary artery disease in particular, high-percentile scores have been shown to confer risk comparable to rare monogenic variants.

The Population Problem

Here is where the architecture of polygenic risk scoring begins to reveal its structural weaknesses. The vast majority of large-scale GWAS datasets—the foundational inputs for PRS construction—have been drawn overwhelmingly from individuals of European ancestry. Estimates consistently suggest that more than 70 percent of GWAS participants have been of European descent, despite that population representing a minority of global genetic diversity.

This imbalance has direct consequences for clinical utility. Allele frequencies differ across ancestral populations, and the linkage disequilibrium patterns that GWAS exploits—the tendency for nearby genetic variants to be inherited together—vary considerably between groups. A score trained on European cohorts may tag a variant that happens to travel alongside a causal variant in that population but not in another. Applied to a patient of West African, South Asian, or Indigenous American ancestry, the score may capture an entirely different set of associations, or none at all.

Studies have confirmed this concern empirically. A PRS for type 2 diabetes derived from European GWAS data performs substantially worse when applied to African American or Hispanic populations. A 2019 analysis published in Nature Genetics quantified this disparity across multiple traits, finding that predictive accuracy degraded systematically as ancestral distance from the training population increased. For clinicians practicing in diverse American cities—where patient populations span dozens of ancestral backgrounds—this is not a peripheral concern. It is a central one.

Correlation Is Not Causation, and Prediction Is Not Certainty

Beyond population stratification, a second class of limitations involves the interpretation of PRS outputs themselves. A score does not identify causative variants. It identifies statistical associations, many of which remain mechanistically unexplained. The variants included in a PRS are often not the functional mutations driving disease biology but rather proxies that happen to be inherited alongside them. This distinction matters when clinicians or patients attempt to draw biological conclusions from score values.

Furthermore, a high PRS does not predict that disease will occur—only that a statistical elevation in probability exists relative to a reference population. The absolute risk increase conferred by even a top-decile score varies enormously depending on baseline population prevalence, age, sex, and environmental context. A score that doubles the relative risk of a condition with a one percent lifetime prevalence still yields a two percent lifetime risk—a figure that may not meaningfully alter clinical management.

Environmental and behavioral factors further complicate the picture. Individuals with high polygenic risk for coronary artery disease who maintain healthy diets, exercise regularly, and avoid smoking may never develop the condition. Conversely, individuals with low scores who accumulate metabolic risk factors over decades may present with myocardial infarction in their fifties. PRS captures one layer of a deeply multifactorial process.

The Path Toward More Equitable Prediction

The scientific community has not been idle in confronting these limitations. Several research initiatives are actively working to diversify GWAS datasets. The All of Us Research Program, funded by the National Institutes of Health and designed to enroll at least one million participants reflecting the full demographic breadth of the United States, represents perhaps the most ambitious domestic effort in this direction. The H3Africa consortium and the Global Biobank Meta-analysis Initiative are expanding representation at the international level.

Methodological innovations are also emerging. Trans-ethnic PRS approaches attempt to leverage shared genetic architecture across populations while accounting for population-specific variation. Bayesian methods and machine learning frameworks are being applied to improve score portability. Some researchers are developing ancestry-specific scores as an interim measure, acknowledging that a single universal model may be insufficient.

Clinical guidelines are beginning to reflect this evolving understanding. The American College of Cardiology has noted that PRS for cardiovascular disease may offer additive value in risk stratification but should not replace established clinical tools—particularly in non-European patients where score performance remains uncertain.

What Clinicians and Researchers Should Take Away

Polygenic risk scores are not pseudoscience. For specific applications, in well-characterized populations, they represent a genuine advance in risk stratification. The challenge lies in the gap between the promise of precision medicine and the current reality of its implementation.

For researchers, the imperative is clear: diversifying training datasets is not merely an equity concern, though it certainly is that. It is a scientific necessity. Scores built on homogeneous data are epistemologically limited regardless of their statistical elegance.

For clinicians, PRS results warrant the same contextual interpretation applied to any probabilistic test. A score is one input among many—not a verdict. Patients receiving high-risk designations deserve counseling that situates the number within the full landscape of their health, their history, and the demonstrated boundaries of the tool being used.

The polygenic era has arrived. Navigating it responsibly requires holding both its genuine potential and its present constraints in view simultaneously.

All Articles

Related Articles

Variants of Uncertain Significance: Navigating the Gray Zone of Genomic Interpretation

Variants of Uncertain Significance: Navigating the Gray Zone of Genomic Interpretation

One Body, Many Genomes: The Science of Somatic Mosaicism and Its Clinical Consequences

One Body, Many Genomes: The Science of Somatic Mosaicism and Its Clinical Consequences

Same Genes, Different Fates: The Science of Why Genetic Risk Is Not Genetic Certainty

Same Genes, Different Fates: The Science of Why Genetic Risk Is Not Genetic Certainty