From blood samples to scans and wearables, the review follows what happens after an AI model identifies a promising disease signal.

Review: Artificial intelligence in biomarker discovery for diseases: diagnostic and therapeutic prospects. Image Credit: ultramansk / Shutterstock
In a recent selective review published in the journal Signal Transduction and Targeted Therapy, the authors examined how artificial intelligence (AI) can support biomarker discovery, validation, biological interpretation, clinical utility, and therapeutic integration across diseases.
Background
A promising biomarker needs more than strong predictive power to be useful in patient care. Biomarkers can help detect, classify, and guide treatment for health problems, yet many candidates that perform well in initial studies cannot be reproduced in independent cohorts or improve clinical decisions. At the same time, biomarker research relies on combinations of measurements, including genomic, transcriptomic, proteomic, metabolic, imaging, cellular, circulating, and digital data.
AI can integrate these data and identify multiscale patterns, but bias, confounding, data heterogeneity, overfitting, and weak external validation can limit translation. The authors call for biomarkers that are reproducible, biologically relevant, clinically applicable, and validated prospectively.
Biomarker data types
Biomarker discovery encompasses several domains, including molecular, cellular, circulating, imaging, and digital biomarkers. In the molecular domain, researchers use approaches such as genomics, epigenomics, transcriptomics, and proteomics.
Multi-omics strategies integrate these layers, while single-cell and spatial approaches preserve cellular and tissue context. Cellular biomarkers capture immune activation, depletion, lineage shifts, inflammatory programs, and tissue environments. Neutrophil extracellular traps link innate immune activation with vascular injury, thrombosis, and tissue damage.
Circulating biomarkers provide a minimally invasive means of assessing disease biology. Signals such as cell-free deoxyribonucleic acid (cfDNA), which includes circulating tumor DNA, cell-free ribonucleic acid (cfRNA), proteins, and metabolites can support disease detection and monitoring in suitable clinical settings. Their clinical value depends on preanalytical processing, analytical sensitivity, tumor burden, treatment timing, and patient characteristics.
Extracellular vesicles carry molecular signals that reflect their cells of origin and facilitate intercellular communication, but isolation methods, contaminants, quantification standards, and laboratory differences complicate their use as biomarkers.
Imaging biomarkers provide information about tissue structure and function. Radiomics and computational imaging quantify features from computed tomography, magnetic resonance imaging, positron emission tomography, and digital pathology, but scanner differences, acquisition protocols, segmentation, and inter-site harmonization must be controlled.
Digital biomarkers from wearable devices, smartphones, and ambient sensors can capture sleep, mobility, autonomic physiology, and symptoms. Device heterogeneity, incomplete data, behavioral confounding, and socioeconomic bias remain concerns.
AI methods for biomarker discovery
Classical machine learning remains useful when datasets contain many variables but few subjects. Regularized regression, support vector machines, and tree ensembles can identify candidate biomarkers, but leakage, repeated tuning, outcome-based selection, and unstable explanations can produce overly optimistic results.
Deep learning is useful for nonlinear and structured data such as images, arrays, and graphs. Representation learning can generate reusable embeddings, while self-supervised learning can leverage unlabeled data for tasks requiring transfer across settings.
Multi-modal machine learning integrates imaging, laboratory measurements, longitudinal clinical data, and other sources. Transfer learning can use large electronic health record datasets to support smaller paired omics and clinical cohorts.
Graph neural networks can model relationships among genes, proteins, pathways, drugs, and phenotypes and incorporate biological structure into learning.
Foundation models trained on large biological datasets may provide reusable representations for tasks involving omics, drug discovery, and biological analysis. These models may also encode batch, ancestry, or site signals that appear as biomarkers if researchers do not control for them.
Generative and simulation-assisted models can propose candidate signatures, molecular states, perturbation responses, pathway hypotheses, and possible mechanisms. Such outputs should initially be treated as hypotheses rather than evidence.
Taking a candidate further requires internal consistency, external replication, tests of stability across populations and platforms, and clinically meaningful benefit. Explainability methods can support auditing and transparency but cannot substitute for biological validation or experimental testing.
Mechanistic AI and clinical translation
Most AI-derived biomarkers are correlational: they identify patterns associated with outcomes without establishing why those patterns occur. Such markers may still be useful for risk stratification, but mechanistic evidence is important when a marker is used to guide therapy or to identify a treatment target.
Mechanistic AI incorporates biological pathways, networks, perturbation data, or causal structures to connect signatures with disease processes. Pathway-constrained models can encode known signaling relationships, while perturbation datasets from clustered regularly interspaced short palindromic repeats experiments, drug-response assays, and cytokine stimulation can help distinguish causal factors from downstream correlations.
Causal representation learning seeks features that remain stable under intervention, but observational data remain vulnerable to unmeasured confounding and changes in disease over time.
Pathway-state biomarkers may capture functional activity more effectively than isolated molecular abundance measurements. They assess how a network of interacting molecules behaves, rather than measuring a single component.
The article discusses phosphoinositide 3-kinase–protein kinase B–mechanistic target of rapamycin, mitogen-activated protein kinase, nuclear factor kappa B, and Janus kinase–signal transducer and activator of transcription pathways as examples of changing biological systems whose activation cannot always be inferred from individual components.
Neutrophil extracellular traps offer another example: they may reflect the interaction between inflammatory and clotting processes, but their value in guiding treatment still needs to be tested.
Mechanistic approaches can support treatment-response classification, patient stratification, trial enrichment, and therapeutic target prioritization, but plausible targets still require experimental and independent validation.
Validation and implementation
The authors propose a development pathway spanning discovery, external validation, testing under different conditions, clinical utility, and deployment.
External validation across independent cohorts and subgroups is needed because discovery performance can overestimate effectiveness.
Evaluation should extend beyond discrimination measures, such as the area under the curve, to include decision analysis, workflow integration, and patient outcomes. Prospective studies are particularly important when biomarkers guide treatment selection, dosing, or clinical-trial enrichment.
Common barriers include data heterogeneity, limited cohort diversity, incomplete biological annotation, reporting inconsistencies, dataset drift, treatment-related confounding, and inadequate integration into clinical workflows.
The review proposes harmonized data collection, prospective multicenter validation, transparent reporting, and standardized assessment of measurement consistency, calibration, fairness, and clinical utility.
Conclusions
The review concludes that AI can expand biomarker discovery by integrating molecular, cellular, imaging, circulating, and digital data to identify disease signatures. Yet accurate predictions alone do not establish clinical value. A biomarker also needs reproducible measurement, independent testing, and evidence that its use improves care. Biological hypotheses should guide multi-modal integration, while experiments must test proposed mechanisms.
The authors identify prospective studies, harmonized data, transparent reporting, standardized evaluation, regulatory coordination, and interdisciplinary collaboration as priorities for advancing AI-derived biomarkers from discovery settings toward therapeutic development and clinical decision-making.
Journal reference:
- Dinc, R., & Ardic, N. (2026). Artificial intelligence in biomarker discovery for diseases: Diagnostic and therapeutic prospects. Signal Transduction and Targeted Therapy, 11, 407. DOI: 10.1038/s41392-026-02946-4, https://www.nature.com/articles/s41392-026-02946-4