In the pursuit of precision medicine, artificial intelligence has become a powerful instrument for detecting patterns within the vast complexity of human biology — yet a new review reminds us that pattern recognition and clinical wisdom are not the same thing. Researchers publishing in Signal Transduction and Targeted Therapy find that AI-discovered biomarkers frequently falter when moved from the controlled environment of discovery into the varied terrain of real patient care, undone by overfitting, population differences, and the absence of biological understanding. The gap between a promisi
AI-Identified Biomarkers Need More Than Accuracy to Reach Patients
Accurate predictions alone do not establish clinical value.
So the core problem is that AI finds patterns that don't hold up in the real world?
Partly that, yes. But it's more specific. The AI finds a real pattern in the data it was trained on. The problem is that pattern may not work the same way in a different hospital, with different equipment, or in patients who are older or from a different ancestry than the training group.
How often does this actually happen? The review doesn't give a number for how many AI biomarkers fail in validation.
That's a fair point. The review doesn't quantify it. It's more of a conceptual argument about why failure happens and what's needed to prevent it.
What's the difference between a biomarker that's accurate and one that's actually useful in the clinic?
Accuracy means it correctly predicts disease in the dataset you tested it on. Usefulness means it works across different patient groups, you understand why it works, and using it actually changes how doctors treat patients and improves outcomes.
And the review is saying most AI biomarkers are validated only on accuracy, not on those other things?
Exactly. They're validated in discovery cohorts, sometimes in one external cohort, but rarely in prospective studies where you follow patients forward in time and measure whether the biomarker actually improves care.
Why is prospective validation so much harder than retrospective?
Because in retrospective studies you already know the outcome. You can tune your model to fit that outcome. In prospective studies, you make your prediction before you know what happens, so you can't cheat.
The review mentions mechanistic validation—understanding the biology. But it also says mechanistic AI is still mostly correlational. So how much does understanding the mechanism actually help?
It helps you distinguish cause from effect. A correlational biomarker might just be measuring something downstream of the real disease process. If you understand the mechanism, you're more confident the biomarker is measuring something that actually matters.
What's the practical barrier? Why aren't hospitals just doing these prospective studies?
Time, cost, and complexity. Prospective studies take years. They require coordination across multiple sites. And if the biomarker doesn't work, you've invested heavily in something that won't be used.
The review mentions data heterogeneity and workflow integration as barriers. But it doesn't explain what that means in concrete terms.
Data heterogeneity means different hospitals use different equipment, different protocols, different ways of handling samples. A biomarker validated on one scanner may not work on another. Workflow integration means the biomarker has to fit into how doctors actually practice—it can't require equipment or expertise that most hospitals don't have.
El Pulso
- AI models are identifying disease biomarkers with striking accuracy, yet the same markers repeatedly collapse when tested in new patient populations or different clinical settings.
- Overfitting, hidden demographic biases, inconsistent sample handling, and non-standardized measurement methods are quietly corrupting results that look clean on paper.
- The deeper problem is structural: most AI biomarkers reveal correlations without explaining biological cause, leaving clinicians unable to act on them with confidence when treatment decisions are at stake.
- Researchers are calling for a rigorous multi-stage development pathway — from independent external validation to prospective clinical trials — before any AI biomarker guides real patient care.
- The field is converging on a set of urgent reforms: harmonized data standards, diverse study populations, transparent reporting, and interdisciplinary collaboration between AI scientists, biologists, and clinicians.
- Without these safeguards, even the most dazzling predictive performance at the discovery stage offers little protection against failure at the moment that matters most — the clinic.
In the pursuit of precision medicine, artificial intelligence has become a powerful instrument for detecting patterns within the vast complexity of human biology — yet a new review reminds us that pattern recognition and clinical wisdom are not the same thing. Researchers publishing in Signal Transduction and Targeted Therapy find that AI-discovered biomarkers frequently falter when moved from the controlled environment of discovery into the varied terrain of real patient care, undone by overfitting, population differences, and the absence of biological understanding. The gap between a promising signal and a trustworthy clinical tool is wide, and crossing it demands not just accuracy, but reproducibility, mechanistic grounding, and demonstrated benefit to patients. Science, it seems, must still earn the trust of the bedside.
A striking pattern appears in the data: an AI model has found a signal in blood work, imaging, or wearable readings that predicts disease with impressive accuracy. The numbers are compelling. The paper is published. And then, frequently, the biomarker fails — in a different patient population, in a different hospital, or simply because it never changes how any doctor treats any patient. A new review in Signal Transduction and Targeted Therapy examines why this cycle repeats so often, and what must change.
The difficulty is not that AI lacks skill at finding patterns. It is exceptionally good at that. The problem is that finding a pattern and proving it matters in real clinical care are two entirely different achievements. AI can integrate blood samples, imaging, genetic sequences, protein measurements, and wearable data simultaneously, identifying signals no human analyst would catch. But a biomarker that performs well on its training dataset may collapse when the patient population shifts in age, ancestry, or disease stage — or when the equipment changes, or when samples are handled differently between sites. The review's authors argue that predictive accuracy is necessary but not sufficient. A clinically useful biomarker must be reproducible across independent groups, biologically plausible, and demonstrably capable of improving decisions.
The technical obstacles are formidable. Classical machine learning is vulnerable to overfitting — learning noise rather than signal. Deep learning can encode hidden biases related to ancestry or measurement site that masquerade as disease patterns. Multi-modal approaches combining imaging, lab work, and clinical history introduce new error sources: scanner variation, incomplete patient data, and socioeconomic factors shaping who receives which tests. Circulating biomarkers in blood offer a minimally invasive window into disease, but their value depends on sample handling, measurement sensitivity, and timing relative to treatment. Extracellular vesicles carry theoretical promise but suffer from inconsistent isolation methods and absent standardization.
Most AI-derived biomarkers are correlational — they identify associations without explaining mechanisms. For risk stratification this may suffice, but when a biomarker is meant to guide treatment or identify a drug target, correlation alone is insufficient. The review advocates for mechanistic approaches: models that incorporate known biological pathways, distinguish cause from effect, and assess how networks of interacting molecules behave rather than measuring isolated components.
The authors propose a structured development pathway moving from discovery through external validation, diverse population testing, and prospective clinical trials before any deployment. Evaluation must extend beyond standard accuracy metrics to ask whether using the biomarker actually changes clinical decisions and improves outcomes. They call for harmonized data collection, transparent reporting, standardized fairness assessments, regulatory coordination, and genuine interdisciplinary collaboration. The conclusion is clear: AI can expand what biomarker discovery is capable of, but accurate predictions alone do not establish clinical value. Without reproducibility, biological grounding, and demonstrated patient benefit, even the most impressive discovery-phase performance means very little when it finally meets the clinic.
A promising signal emerges from the data. An artificial intelligence model has found a pattern in blood work, imaging scans, or wearable device readings that predicts disease with impressive accuracy. The numbers look good. The researchers publish. And then, often, the biomarker fails to work when tested in a different patient population, or it doesn't actually change how doctors treat anyone, or it simply cannot be reproduced. A new review in Signal Transduction and Targeted Therapy examines why this happens so often—and what needs to change for AI-discovered biomarkers to actually reach patients.
The problem is not that AI is bad at finding patterns. It is excellent at that. The problem is that finding a pattern and proving it matters in real clinical care are two entirely different things. Researchers can combine data from blood samples, imaging studies, genetic sequencing, protein measurements, cellular analysis, and wearable devices. AI can integrate all of it and identify signals that humans would miss. But a biomarker that performs well in the lab—one that correctly predicts disease in the dataset it was trained on—may collapse when applied to a different group of patients, or when the equipment changes, or when the patient population shifts in age, ancestry, or disease stage. The review's authors argue that predictive accuracy alone is not enough. A useful biomarker needs to be reproducible across independent groups, biologically plausible, and proven to actually improve clinical decisions.
The technical obstacles are substantial. Classical machine learning can identify candidate biomarkers from datasets with many variables but few subjects, yet the methods are vulnerable to overfitting—the model learns noise instead of signal. Deep learning works well for images and complex data structures, but it requires large amounts of training data and can encode hidden biases related to ancestry, site of measurement, or batch processing that masquerade as disease signals. Multi-modal approaches that combine imaging, lab work, and clinical history are promising, but they introduce new sources of error: scanner differences, variations in how images are processed, incomplete data from some patients, and socioeconomic factors that affect who has access to certain tests. Circulating biomarkers—cell-free DNA, proteins, metabolites in the blood—offer a minimally invasive window into disease, but their clinical value depends on how the sample is handled before testing, how sensitive the measurement is, how much tumor burden is present, and when the sample is collected relative to treatment. Extracellular vesicles, tiny particles that carry molecular signals between cells, are theoretically valuable but plagued by inconsistent isolation methods and lack of standardized measurement.
Most AI-derived biomarkers are correlational: they identify patterns associated with outcomes without explaining why those patterns occur. That may be sufficient for risk stratification—telling a patient their risk is high or low—but it is insufficient when a biomarker is meant to guide treatment decisions or identify a drug target. The review emphasizes mechanistic approaches: AI models that incorporate known biological pathways, that use data from genetic perturbation experiments to distinguish cause from effect, that assess how networks of interacting molecules behave rather than measuring single components in isolation. Neutrophil extracellular traps, for example, may reflect the interaction between inflammatory and clotting processes, but their value in guiding treatment has not yet been tested. Pathway-state biomarkers that capture functional activity of signaling networks may be more stable and more clinically relevant than measurements of individual molecular abundance, but this remains an area requiring experimental validation.
The gap between discovery and clinical use is wide and well-documented. A biomarker that performs well in the initial study often fails when tested in independent cohorts. Data heterogeneity—differences in how samples are collected, processed, and measured across sites—undermines reproducibility. Limited diversity in the patient populations studied means the biomarker may not work equally well across age groups, ancestries, or disease stages. Incomplete biological annotation means researchers cannot understand what the biomarker actually measures. Treatment-related confounding—the fact that patients receiving different therapies may have different biomarker values for reasons unrelated to disease—is often not controlled for. And even when a biomarker is validated, it may not integrate smoothly into clinical workflows, or it may require equipment or expertise that not all hospitals have.
The authors propose a structured development pathway: discovery, external validation across independent cohorts and subgroups, testing under different conditions and in diverse populations, demonstration of clinical utility through prospective studies, and finally deployment. Evaluation must go beyond standard discrimination measures like the area under the curve to include decision analysis—does using this biomarker actually change what doctors do and improve outcomes?—and workflow integration. Prospective studies are particularly important when biomarkers guide treatment selection or clinical trial enrollment, because retrospective studies can overestimate effectiveness.
The review identifies priorities for moving forward: prospective multicenter validation studies, harmonized data collection across sites, transparent reporting of methods and results, standardized assessment of measurement consistency and fairness, regulatory coordination, and interdisciplinary collaboration between AI researchers, biologists, clinicians, and statisticians. The authors conclude that AI can expand biomarker discovery by integrating diverse data types, but accurate predictions alone do not establish clinical value. A biomarker must be reproducibly measured, independently tested, biologically plausible, and proven to improve care. Without these elements, even the most impressive accuracy in the discovery phase means little when the biomarker reaches the clinic.
Citas Notables
A promising biomarker needs more than strong predictive power to be useful in patient care— Review authors in Signal Transduction and Targeted Therapy
Biological hypotheses should guide multi-modal integration, while experiments must test proposed mechanisms— Review conclusion