For generations, the tools of genetic medicine have been calibrated to one kind of human, leaving others to navigate their health futures with instruments not built for them. A team of researchers has now introduced a method — Mondrian Cross-Conformal Prediction — that does not pretend to erase this disparity, but instead names it honestly: flagging when a prediction can be trusted and when it cannot, across White British, South Asian, and African/Caribbean patients with type 2 diabetes. In an era when genomic data remains overwhelmingly European in origin, this act of quantified humility may
New method improves genetic risk prediction for diabetes complications across diverse populations
Explicitly telling clinicians when a prediction is solid and when it isn't
So the core problem here is that genetic risk scores work better for European populations. Why is that?
It's partly about the data used to build the scores. Most genome-wide association studies—the research that identifies which genetic variants matter—have been done in European populations. But there's also real genetic diversity. Different ancestry groups have different allele frequencies, different patterns of genetic linkage, and different environmental exposures. A variant that predicts heart disease in one population might not in another.
But wait—the study shows MCCP performs comparably across populations. So did they solve the problem, or just measure it better?
They didn't solve the underlying genetic differences. What they did was add a layer of honesty. Instead of giving a single risk number that might be wrong, MCCP says: "I'm confident about this prediction" or "I'm uncertain, and here's how uncertain I am." For clinical use, that's actually more useful than a false sense of precision.
The paper mentions that calibration—whether the predicted error rate matches the actual error rate—wasn't uniform. What does that mean practically?
It means that for some outcomes in some populations, the method said "I'm 95 percent confident" but was actually wrong more often than 5 percent of the time. That's a problem if you're relying on that guarantee.
Specifically, they found under-coverage for heart attack in the African/Caribbean population and for kidney function across groups. That's not trivial. If a clinician is told the error rate is 5 percent but it's actually 10 percent, they're making decisions on false confidence.
So the method isn't ready for clinical use yet?
Not without more work. The authors explicitly call for prospective validation—testing it in real clinical settings with real patients—before widespread implementation. They also need to figure out how to handle the "uncertain" predictions. If the model says "I can't confidently predict this patient's risk," what does the doctor actually do?
And there's a sample size issue. The South Asian and African/Caribbean groups were much smaller. That limits how precisely they can estimate performance, especially for rare outcomes. The downsampling analysis showed that differences between populations weren't just about sample size, but it doesn't mean bigger samples wouldn't improve calibration.
The kidney function finding was interesting—they tested both the old formula with a race coefficient and the new race-free formula. Why does that matter?
Because it shows how clinical definitions themselves can embed ancestry assumptions. The old formula assumed kidney function differs by race, which is biologically questionable. When they removed that assumption, 15 patients in the African/Caribbean group got reclassified from normal to abnormal kidney function. The genetic predictions still worked, but the underlying definition of disease changed.
That's a crucial point. You can have a perfectly valid prediction method, but if the thing you're predicting is defined in a way that's biased, you're still perpetuating the bias. The method is only as fair as the outcome definition.
What's the practical takeaway for a clinician reading this?
Don't use this in clinical practice yet. But watch for it. The framework—explicitly quantifying uncertainty and flagging unpredictable cases—is the right direction. It's more honest about what genetic prediction can and can't do.
And be skeptical of any genetic risk score that doesn't tell you how confident it is. If someone gives you a single number without uncertainty bounds, especially in a population different from where the score was developed, you should ask hard questions about what that number actually means.
O Pulso
- Polygenic risk scores — the genetic tools used to forecast heart attack, stroke, and kidney failure — have long performed worse in non-European patients, creating a quiet but consequential inequity in preventive medicine.
- Researchers tested a new framework on nearly 20,000 diabetes patients across three ancestry groups, finding that the method delivered comparable predictive performance where standard approaches have historically faltered.
- The real disruption is not better accuracy but enforced transparency: the model explicitly marks predictions as confident or uncertain, forcing clinicians to confront the limits of what the data can actually say.
- A side finding shook the clinical ground further — removing a race-based coefficient from the standard kidney function formula reclassified 15 additional African/Caribbean patients as having reduced kidney function, exposing how medical definitions themselves carry embedded assumptions.
- The path forward is unfinished: smaller cohorts, unresolved calibration gaps, and the absence of clinical protocols for handling uncertain predictions mean prospective validation is the urgent next frontier.
For generations, the tools of genetic medicine have been calibrated to one kind of human, leaving others to navigate their health futures with instruments not built for them. A team of researchers has now introduced a method — Mondrian Cross-Conformal Prediction — that does not pretend to erase this disparity, but instead names it honestly: flagging when a prediction can be trusted and when it cannot, across White British, South Asian, and African/Caribbean patients with type 2 diabetes. In an era when genomic data remains overwhelmingly European in origin, this act of quantified humility may be as important as any breakthrough in accuracy.
For decades, genetic risk prediction has served people of European descent far better than anyone else. The polygenic risk scores used to anticipate heart disease, stroke, and kidney failure in type 2 diabetes patients were built on European data — and when applied to South Asian, African, or Caribbean patients, they simply perform less well. This is not a minor technical gap. It is a fairness problem with real clinical consequences.
A team of researchers responded not by trying to make one universal model work equally for all — an impossible goal given genuine biological and historical differences — but by building a system that knows what it doesn't know. Their method, Mondrian Cross-Conformal Prediction (MCCP), allows clinicians to set an acceptable error threshold in advance and then receive predictions sorted into two categories: those the model can make confidently, and those it cannot. The study tested this across 19,468 people with type 2 diabetes from the UK Biobank, spanning White British, South Asian, and African/Caribbean populations.
The results held up across all three groups. Predictive performance for stroke, heart attack, and reduced kidney function was comparable whether the model was applied to the largest or smallest cohort — with AUROC scores ranging from 0.69 to 0.87 depending on outcome and ancestry. The method wasn't flawless; calibration varied, and the smaller minority cohorts introduced inherent uncertainty. But the critical innovation was transparency: at a 5 percent error threshold, doctors would know exactly which predictions to act on and which required additional judgment.
A secondary finding added unexpected weight to the study. When researchers replaced the standard kidney function formula's race coefficient with a newer race-free version, the number of African/Caribbean patients classified as having reduced kidney function rose from 47 to 62. MCCP remained valid under both definitions — but the episode illustrated how clinical tools themselves can encode ancestry-based assumptions that deserve scrutiny.
Limitations remain significant. The African/Caribbean cohort in the UK Biobank reflects a specific admixture pattern that does not represent the full breadth of African ancestry globally. Clinical protocols for managing uncertain predictions do not yet exist. And the foundational question — does this approach actually improve patient outcomes in practice — awaits prospective testing. What the study establishes is a principled direction: in genomic medicine, naming uncertainty honestly may be the most equitable thing a model can do.
For decades, genetic risk prediction has worked best for people of European descent. The polygenic risk scores that doctors use to identify who might develop heart disease, stroke, or kidney failure—particularly among those with type 2 diabetes—were built on data from European populations and simply don't perform as well when applied to South Asian, African, or other ancestry groups. This gap isn't a minor technical problem. It's a fairness issue that could leave millions of patients without reliable guidance about their own health risks.
Researchers at multiple institutions set out to fix this. They developed a new approach called Mondrian Cross-Conformal Prediction, or MCCP, that doesn't try to make genetic scores work equally well across all populations—an impossible task given real biological differences. Instead, it does something different: it tells doctors exactly how confident the prediction is, and flags cases where the model simply cannot make a reliable call. The team tested this method on 19,468 people with type 2 diabetes from the UK Biobank: 17,574 who identified as White British, 1,145 South Asian, and 749 African or Caribbean. They were trying to predict three major complications: stroke, heart attack, and reduced kidney function.
The results were striking in their consistency. When the model was trained on European data from the ADVANCE clinical trial and tested on the three UK Biobank populations, MCCP performed comparably to standard prediction methods across all three groups. For stroke in the White British population, it achieved an AUROC of 0.69—a measure of how well it separated high-risk from low-risk individuals. In the South Asian group, that rose to 0.86. In the African/Caribbean group, it reached 0.87. The same pattern held for heart attack and kidney function predictions. The method wasn't perfect—calibration varied by outcome and ancestry—but it was reliable enough to be clinically useful.
What made MCCP different wasn't raw predictive power. It was transparency. The method allowed researchers to set an acceptable error rate in advance—say, 5 percent—and then identify exactly which patients received confident predictions under that threshold and which ones fell into an "uncertain" or "unpredictable" category. In the White British population at a 5 percent error threshold, MCCP issued confident predictions for 10.3 percent of individuals for stroke, 10.4 percent for heart attack, and 20.7 percent for reduced kidney function. Those numbers were higher in the smaller South Asian and African/Caribbean groups, reflecting the inherent uncertainty of working with less data. But the key insight was that doctors would know which predictions to trust and which ones required additional clinical judgment.
The researchers ran multiple tests to understand where differences came from. They trained models directly in the White British population and tested them in South Asian and African/Caribbean groups, removing the confounding effect of different study designs. They downsampled the larger populations to match the smallest group's size, testing whether sample size alone explained the variation. Across all these scenarios, MCCP maintained comparable performance. The differences between populations reflected real genetic and environmental variation, not methodological bias.
One complication emerged around kidney function prediction. The standard medical formula for estimating kidney function, called CKD-EPI, included a race coefficient—a mathematical adjustment based on the assumption that kidney function differs by race. When researchers removed that coefficient and used the newer race-free version, the number of African/Caribbean patients classified as having reduced kidney function increased from 47 to 62. MCCP's predictions remained valid under both definitions, but the finding highlighted how clinical definitions themselves can embed ancestry-related assumptions that need scrutiny.
The limitations were real. MCCP is computationally more demanding than simple logistic regression, though still manageable for research applications. The method produces predictions it cannot confidently make—and those require explicit clinical protocols that don't yet exist. The South Asian and African/Caribbean cohorts, while substantial, remained much smaller than the White British group, limiting the precision of estimates for rare outcomes. Most critically, the African/Caribbean population in the UK Biobank represents a specific subset of African ancestry with particular admixture patterns; the findings don't automatically transfer to continental African populations or other African diaspora communities.
What the study demonstrated was a path forward. Rather than trying to make one-size-fits-all genetic predictions, the field could embrace uncertainty quantification—explicitly telling clinicians when a prediction is solid and when it isn't. This matters most for the populations that have been historically underserved by genomic research. As the researchers noted, continued validation in larger and more diverse cohorts remains essential, particularly as genomic datasets continue to be dominated by people of European ancestry. The next step is prospective clinical testing: Does this approach actually change how doctors make decisions? Does it improve outcomes? Those questions remain open, but the foundation is now in place.
Citações Notáveis
Rather than increasing predictive accuracy, MCCP provides clinicians with explicit information regarding the confidence associated with each prediction, an important property for the future clinical implementation of polygenic risk scores in diverse populations.— Study authors, in discussion of MCCP's primary value
Clinically, under-coverage implies that the observed prediction error may exceed the user-specified error level for these phenotype–population combinations, reducing confidence in the nominal validity guarantee. This observation is particularly important for underrepresented populations, where reliable uncertainty estimation is essential for equitable clinical implementation.— Study authors, on calibration limitations in underrepresented groups