ARTIFICIAL INTELLIGENCE - AI/ML Supporting Biomarker & Target Discovery


Key Points

  • AI/ML provides new capabilities to expedite processes and enable more informed decisions.
  • AI/ML analysis rapidly increases type and amount of biological data employed in identifying biomarker and disease target (BMDT) discovery.
  • Collaboration between domain experts and data scientists is crucial to maintain quality, safety, and reliability of AI systems.

By: William Whitford

INTRODUCTION

There is rapid and successful development of Artificial Intelligence (AI) and Machine Learning (ML) applications in medicine, including in such fields as pharmacology, drug discovery, and bioinformatics.1 In this, AI/ML is revolutionizing biomarker and therapeutic target discovery by providing new capabilities, expediting processes, and enabling more informed decisions.

Due to the successful implementation of the digital transformation in biomarker and disease target (BMDT) discovery, data are now routinely captured and distributed in electronic format. AI supports development of tools that function without explicit (deterministic) programming. These (stoichiometric) tools are increasing used in curating and analyzing data, developing understandings, advising in decisions, and even autonomous agency.

AI/ML provides new capabilities in analysing the rapidly increasing type and amount of biological data employed in identifying potential BMDTs. It can mine research publications and digital libraries to discover and characterize biomolecules or constructs associated with diseases. Its ability to provide deep understanding of relationships is helping in the related disciplines of disease biomarker and therapeutic target investigation and validation.2

Diagnostic, staging, and prognostic biomarkers arise from a broad range of phenomena indicating what is happening in a cell culture or patient. They are objective measurements of cellular, tissue, or organ activity used to measure the presence, state, or progress of disease. Biomarkers improve patient care by enabling earlier diagnosis, more accurate treatment selection, and improved monitoring of disease progression and therapy response.

An example classical biomarker is HbA1c, used to assess long-term glucose control in diabetics. A newer biomarker involves the BRCA1 and BRCA2 gene mutations, which indicate an increased risk of some hereditary cancers.

Biomarker research is constantly evolving, and new biomarkers are being discovered through analysis of bodily tissue, cell, or fluids using such methods as genome-wide association analysis powered by next-generation sequencing and mass spectrometry-based proteomics, as well as by such imaging techniques as MRI and PET.

A therapy target is also a molecule or system in the body directly associated with a particular disease, but one that can be modulated by a therapy. Target identification is increasingly important in therapy discovery following innovations in assay and experimental technologies. But the recent power provided by AI in the newer multiomic studies is promising even more significant gains.

An example of contemporary targets are the mutations in the cystic fibrosis transmembrane conductance regulator (CFTR) gene in patients with cystic fibrosis. The primary gene therapy approach here is to introduce a functional copy of the CFTR gene into patient’s cells.

In personalized medicine, AI can analyze demographic or individual patient biology in the context of published research and diagnostic data to help identify suitable BMDT for specific populations or even patients. This, along with AI-enabled meta-analysis of a candidate therapy’s published primary and off-target activity, can be used to map its overall biological activity, suggesting treatment effects and clinical outcomes.

AI IN BIOMARKER/TARGET DISCOVERY & ANALYSIS

A tremendous amount and type of data is generated in analysing the multiple chemical, physicochemical, and biological systems associated with discovering a candidate biomarker or pursuing its correlation with disease. In target optimization, similar data aids in the identification of both novel therapeutic targets in particular disease, as well as in predicting new targets for existing drugs.

The number of approaches employing such tools as high-throughput screening (HTS) high-content analysis (HCA) is large and growing. While they have already significantly contributed to the field of omics-based biomarker discovery, AI methods are improving their utility and potential.3

AI algorithms help identify candidates due to their ability to handle such large amounts of disparate data, identify patterns, and make classifications or predictions from them.4 These algorithms transform previous biased, slow, computer-assisted but human-based approaches to classifications and predictions into fast and objective systems.

Due to the many factors involved, AI has demonstrated unique value in characterizing a biomolecule’s potential value in reporting on a disease. AI models corelate data from a BMDT candidate’s measured genotype or phenotype properties to the desired ideal properties and predict its value or even suggest alternatives. Dynamic in silico modelling using machine learning and deep learning algorithms are used to model cell and vesicle surface properties, activities, interactions, as well as a drug candidate’s potential disease modulating effects. The power of this related activity is demonstrated by Insilico Medicine in its Phase 2 clinical trial of its AI-discovered drug, INS018055.5

Automated and high-throughput technologies generate measurements for numerous variables across thousands of experimental samples. Because biological systems are inherently interconnected, these data are typically correlated, multi-parametric, and often high-dimensional. Historically, however, such datasets have not always been analyzed in a manner consistent with their complexity. Accurate interpretation increasingly depends upon multivariate analysis (MVA), which examines relationships among multiple variables simultaneously rather than in isolation.

AI-enabled MVA extends these capabilities by identifying nonlinear, higher-order interactions, latent structures, feedback mechanisms, and dependencies that are difficult or impossible to detect using traditional statistical approaches. This is particularly important in complex biological systems that are both dynamic and dynamical, where responses evolve over time and arise from interacting mechanisms that continuously influence one another. Such systems are frequently non-stationary, with relationships among variables changing across developmental stages, physiological states, or environmental conditions. AI methods can model these complex behaviors, integrate heterogeneous data sources, and distinguish meaningful biological signals from noise.

Supported by modern computing and cloud-based infrastructures, AI-powered MVA provides timely predictions, deeper process understanding, and actionable recommendations. By revealing the interconnected factors that govern biological performance, it enables more accurate interpretation and improved decision-making.

Examples of this power in currently available drug discovery tools include an open access machine learning program called ConPlex, developed by researchers at MIT, which predicts drug-target interactions. ConPlex requires only the sequences of the system’s proteins and simple descriptions of the candidate drugs.6 AI-Bind is a model that combines network-based sampling with unsupervised pre-training to improve binding predictions for novel proteins and ligands. It is a deep learning drug/target combination identification algorithm with promise in drug discovery.7

AI is also being exploited in the emerging field of chemoinformatics. Neural networks have been used in in silico screening through chemical modelling for years. However, we are now seeing that synthetic chemistry and chemoinformatics, based upon the application of AI techniques, is promising more powerful analysis, including employing virtual libraries representing a much larger chemical space.8 ML models can now predict the properties of new compounds by employing both measured features of the chemical as well as purely theoretical descriptors derived from its chemical graph (a graph-theoretical representation of a molecule) or 3D structure data.9

EXAMPLE OF AI VALUE IN BIOPHARMACEUTICAL DISCOVERY

We have seen how AI-empowered AlphaFold can predict protein structures from amino-acid sequences. This success has created hope that these neural networks, trained upon well characterized protein sequences and structures, could also help to create novel proteins with understood functionality. The activity of immunogens, receptor traps, and enzymes are often mediated by a small number of functional residues, with their domains properly presented by the overall protein structure.

Open AI’s GPT4 is an AI generative network we are all familiar with, and this type of AI application is now demonstrating the ability to design therapeutically functional proteins from simple molecular specifications. For example, we are seeing the design of novel protein structures containing prespecified functional sites, in an effective orientation, customized for a particular use or disease therapy.10

AI techniques enabled researchers at University of California, San Francisco (UCSF) in collaboration with a team at IBM Research, to expand CAR T technology, and make their design more quantitative and predictive. Using neural networks trained to decode combinatorial CAR signalling motifs they discovered key design rules of the system. AI tools working on libraries supported the engineering of receptors with desired phenotypes.11

ProSurfScan is a popular and commercially available program to model the compatibility and binding mode of a candidate on different regions of a protein surface.12 Here, the protein surface is represented as a graph of nodes defining local supra-molecular interaction (non-ionic binding) features. Convolutional neural networks can then agnostically estimate candidate interactions with distinct regions on the protein surface.

EMERGING THERAPEUTIC MODALITIES & TECHNOLOGIES

We are seeing a remarkable number of emerging developments in the nature of ATMP therapeutic entities, the means to deliver them, and the technologies employed in their production. AI is becoming essential in handling the amount and type of data coming from the many high-throughput and high-density analytics and screenings employed here.13

Deep learning algorithms such as recurrent neural networks (RNNs) are well-suited for analyzing such massive amounts of multivariate time-series data. They assist in tracking of changes over time from assays of cell type-specific responses, including the expression of specific genes from cultures cells and biobank submissions.

AI’s deep learning is supporting a high-level analysis of cell motion, microfluidics, and multiomics. New single-cell analysis techniques allows the study of individual cells’ behavior and omics, leading to a deeper understanding of cellular heterogeneity and disease mechanisms. Here, AI is providing the point-to-point analysis of hundreds of individual cells and yields much deeper knowledge than the results of the averages of millions of cells.

Recently, improvements in single-cell whole-genome sequencing is showing promise in tumors that display a mosaic of genetic variation, or sub-clonal mutations that make the tumor more aggressive.14 One can only imagine the understandings to be revealed as AI classifiers are applied to this data.

There are many other AI-powered applications aiding in therapeutic candidate and BMDT understanding, such as virtual reality (VR) and augmented reality (AR). These are allowing us to visualize and even interact with complex biological structures and datasets. Such capability was reviewed in a Computer Vision and Pattern Recognition Conference which occurred in Vancouver.15

Finally, it is valuable to realize that AI supports not only new, previously impossible, capabilities, but can radically speed-up existing data processing activities and applications. For example, AI-driven natural language processing (NLP) does not really change the nature of literature review, but significantly reduces the time involved in examining and summarizing vast amounts of research and medical literature.

EMPLOYING EVIDENCE-BASED ASSOCIATIONS

By examining the structural, biological, or functional similarity of BMDT in diseases of related ontology, AI-based systems can identify candidates with appropriate genetic, biological pathway, or cellular activity. Existing approved disease modulators can provide data valuable in understanding properties or activity in other disease etologies.

Such use of disease ontology requires an understanding of genetic and clinical data, as well as symptoms through evidence-based associations. A new initiative, The Disease Ontology, promises an open-source tool for the integration of biomedical data associated with human disease, disease features, and mechanisms. It hopes to serve as a reference framework for multiscale biomedical data integration and analysis towards strengthening the disease information ecosystem.16

Existing therapies will have immense amounts of research, clinical, and incremental post-launch data. Such data can reveal much about the nature and associations of a disease with respective phenotype, environment, and genetics. AI can efficiently enable comprehensive multi-parametric collection and analysis of enormous sets of emerging research data with the vast biomedical literature, clinical trial data, electronic health records, and other sources to identify potential BMDT and therapeutic candidates.

The discovery and analysis of various real-world patient data, including published off-label use, is used in the evaluation of the value of proposed BMDT in identified patient populations and even personalized medicine. By searching through historical data pointing out the many primary and off-target clinical effects of a drug, AI can predict candidates in the context of complex endocrine, immune, or metabolic pathways. AI models of the complex interaction networks of existing therapeutic entities with novel candidate BMDT and disease modulators aids in better understanding the relationships between them.

SUPPORT OF TESTING PHASES & CLINICAL TRIALS

Beyond relevant power in data governance and processing, AI’s power in clinical trials per se has been well reported, and examples here include more objective and less biased patient recruitment, trial design optimization, regulatory compliance, and data security. Well-fed AI models help researchers to prioritize and design preliminary studies for these candidates and predict the potential efficacy of biomarker candidates in clinical trials.

Trials can produce diverse data types from researchers, program administrators, and clinicians. Such varied sources provide such data types as labelled and unlabelled, structured and unstructured, numerical and categorical, time series, and text. AI is extremely powerful in aiding in the collection, transmission, and curation of such disparate data types.

CHALLENGES

Some AI algorithms lack of transparency (“black boxes”) makes it all the more difficult to know whether a system is fair, accurate, and complete. An AI system can also be inappropriate based on flawed assumptions about its application context, users, or their current needs and procedures. AI has been reported to contribute to increased cybersecurity risks− on the other hand, it’s been shown to empower a path to enhanced security.

Explainability and transparency in algorithms and their updates is desired to ensure accuracy and compliance in outputs − yet remains difficult to achieve. Collaboration between domain experts, data scientists, and software testers is crucial to establish and maintain the quality, reliability, safety, and alignment of AI systems throughout the entire product lifecycle.17

The dynamic nature of AI models, as well as drift in both the real-world scenarios and user behavior, establish a requirement for continued process monitoring and data or model updating. That this is a developing field can result in ambiguity in user terminology, expectations, and implementations − as well as in model aligned behavior and results.

The growing size and complexity of systems is increasing the difficulty in identifying critical functionalities, defining appropriate testing, and establishing acceptance criteria. SaaS, risk-based validation, CSA, and Pharma 4.0 initiatives determined a significant change in validation requirements and solutions − and now the unique aspects of AI we have mentioned have furthered this need for development in validation approaches.

Deploying AI in compliant environments is a two-edged sword: while it helps ensure such factors as data integrity and increased productivity, questions remain in such topics as how much external information is used in computing results, how much human involvement is necessary for operation, and what the acceptable degrees of algorithm opacity might be. To address this, many authoritative publications are appearing, such as a draft guidance from the FDA18 and a set of principles from the EMA.19

Many have remarked that generating sufficient AI-quality data is the greatest limitation to progress in many applications. And, that we need better orchestration of data generation, capture, storage, and processing. Automated system outcomes can be harmed by data that is flawed, mislabelled, skewed, or unrepresentative – or that incorporates historical bias, or other types of errors.

A PROMISING FUTURE

The combination of modern data science, power computing, system connectivity, and AI is creating a sea-change in our capabilities in therapeutic biomarker and target discovery, development, and testing. Some believe the integration of AI in pharma-related the systems will be more transformative than any other emerging technology.

REFERENCES

  1. Kotaro Miura, et al, Deep learning-based model detects atrial septal defects from electrocardiography: a cross-sectional multicenter hospital-based study, Open Access, 17 Aug, 2023, DOI: https://doi.org/10.1016/j.eclinm.2023.102141.
  2. Hosseini, Maryam, Hammami, Behnaz, Kazemi, Mohammad, Identification of potential diagnostic biomarkers and therapeutic targets for endometriosis based on bioinformatics and machine learning analysis, Journal of Assisted Reproduction and Genetics, https://doi.org/10.1007/s10815-023-02903-y, DO – 10.1007/s10815-023-02903-y.
  3. Unleashing high content screening https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9530837/.
  4. Toni Manzano, William Whitford, Chapter 4 – AI applications for multivariate control in drug manufacturing, Editor(s): Anil Philip, Aliasgar Shahiwala, Mamoon Rashid, Md. Faiyazuddin, A Handbook of Artificial Intelligence in Drug Delivery, Academic Press, 2023, Pages 55-82, ISBN 9780323899253, https://doi.org/10.1016/B978-0-323-89925-3.00023-X.
  5. https://www.clinicaltrialsarena.com/news/insilico-medicine-ins018055-ai/.
  6. https://doi.org/10.1073/pnas.2220778120.
  7. Chatterjee, A., Walters, R., Shafi, Z. et al. Improving the generalizability of protein-ligand binding predictions with AI-Bind. Nat Commun 14, 1989 (2023). https://doi.org/10.1038/s41467-023-37572-z.
  8. Computational approaches streamlining Sadybekov, A.V., Katritch, V. Computational approaches streamlining drug discovery. Nature 616, 673–685 (2023). https://doi.org/10.1038/s41586-023-05905-z.
  9. Artificial Intelligence-Based Drug Design and Discovery DOI: http://dx.doi.org/10.5772/intechopen.89012.
  10. Avery, Chris, John Patterson, Tyler Grear, Theodore Frater, and Donald J. Jacobs. 2022. “Protein Function Analysis through Machine Learning” Biomolecules 12, no. 9: 1246. https://doi.org/10.3390/biom12091246.
  11. https://www.science.org/doi/10.1126/science.abq0225.
  12. https://chemotargets.com/chemotargets-announces-first-ai-designed-drug-for-huntingtons-disease-to-enter-clinical-trials/.
  13. https://www.basetwo.ai/bioprocessing.
  14. Joanna Hård, et al, Long-read whole genome analysis of human single cells, bioRxiv 2021.04.13.439527 doi: https://doi.org/10.1101/2021.04.13.439527.
  15. https://cvpr2023.thecvf.com/.
  16. https://disease-ontology.org/about/.
  17. Manzano T, Fernàndez C, Ruiz T, Richard H. Artificial Intelligence Algorithm Qualification: A Quality by Design Approach to Apply Artificial Intelligence in Pharma. PDA J Pharm Sci Technol. 2021 Jan-Feb;75(1):100-118. doi: 10.5731/pdajpst.2019.011338. Epub 2020 Aug 14. PMID: 32817323.
  18. Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products, January 2025, FDA-2024-D-4689.
  19. Guiding principles of good AI practice in drug development, January 2026, https://www.ema.europa.eu/system/files/documents/other/2026-01-13_ai-guiding-principles_ema-fda_en.pdf.

William Whitford has over 20 years of experience in biotechnology product and process development. He now publishes oral papers, journal articles, and book chapters on such topics as ATMP, AI/ML applications, and environmentally sustainable biomanufacturing practices. He also enjoys serving on a number of biotechnology committees, boards, and panels, such as the MagBIO research consortium Industry Advisory Board.