WHY BIOLOGICAL DATA MATTERS MORE IN AI DRUG DISCOVERY

GSK has expanded its partnership with British biotechnology firm Relation Therapeutics in a research deal worth up to $110 million to use human-cell data and artificial intelligence to discover potential drug targets.

Under the agreement, Relation will generate large-scale datasets that show how human cells respond to genetic changes and drug treatments. The company will use the data to train AI models, including its MORGAN platform, to identify potential targets for new medicines.

The deal brings biological data generation and AI model development into one research programme. Relation combines laboratory experiments with computational analysis to create new information about human cells and disease.

GSK expands work with Relation

Meanwhile, the new agreement builds on earlier partnerships between GSK and Relation that focused on fibrotic diseases and osteoarthritis. Those projects involved observational studies designed to produce two functional disease datasets for analysis through Relation’s Lab-in-the-Loop platform.

The earlier research brought together human genetics, single-cell multi-omics from human tissue, functional assays and machine learning. Researchers used those tools to identify and validate potential disease targets.

Furthermore, the expanded partnership reflects a wider push by pharmaceutical companies to combine human biological data with AI. GSK has increasingly used genetics, functional genomics and machine learning to strengthen its drug discovery pipeline.

How Relation creates biological data

Relation describes its Lab-in-the-Loop system as a combination of laboratory experiments and computational analysis. Its work includes tissue profiling, single-cell and spatial transcriptomics, sequencing and target validation.

The company also uses machine learning to identify, rank and validate potential targets. In addition, it uses the technology to help design new experiments.

In response to the need for more detailed biological information, Relation conducts perturbation experiments that measure how genetic changes affect cellular features linked to disease. Researchers can then compare those results with genetic information and biological data collected from patients.

Public repositories also provide important material for biological foundation models. However, researchers face challenges when they combine data from studies that used different laboratory methods and experimental conditions.

A 2025 review in Experimental & Molecular Medicine highlighted repositories such as CZ CELLxGENE, the Human Cell Atlas and NCBI Gene Expression Omnibus as major sources of single-cell data. The review said CZ CELLxGENE provides access to more than 100 million standardised cells.

More data does not always mean better AI

However, researchers must deal with differences in sampling methods, sequencing protocols, experimental procedures and data-processing systems. Single-cell datasets can also contain technical noise and other artefacts.

Consequently, researchers need careful dataset selection, filtering, composition balancing and quality control when they train foundation models. Poor-quality or inconsistent data can affect how well an AI model represents biological systems.

Dataset duplication creates another problem. Similar cells can appear in several public databases, giving those cells more influence during training than researchers intended.

Furthermore, overlapping training and test datasets can create data-leakage risks. The 2025 review therefore stressed that researchers need high-quality, non-redundant datasets as much as they need strong model architecture.

Research published in Nature Methods in June 2026 also questioned the assumption that bigger datasets will always produce better single-cell AI models.

Researchers used a corpus containing 22.2 million cells to train 400 models and conducted 6,400 experiments. They found that current single-cell foundation models often reach performance plateaus after training on only part of the available data.

Unlike large language models, the researchers found no clear data-scaling pattern in which continually adding more training data consistently improves performance. Instead, they said developers need to balance model capacity, dataset size and computing resources.

Pharma turns to specialised datasets

Meanwhile, another 2025 study in Genome Biology assessed Geneformer and scGPT across several zero-shot tasks. The researchers found that the models did not consistently outperform simpler methods.

The study also identified challenges involving batch effects. It warned researchers against assuming that larger pretrained models automatically create better biological representations.

Relation has already applied its data-generation strategy to Osteomics, which the company describes as a proprietary functional single-cell bone atlas. The project combines patient-derived samples with single-cell and spatial omics, imaging, genomics, proteomics and clinical phenotype data.

The company uses Osteomics to study disease biology, potential drug targets, biomarkers and patient subgroups in osteoporosis. Hospitals and research partners in the UK and Australia participate in the observational study.

Research published in Nature Genetics last month also examined cellular and genetic factors linked to skeletal disease. The study combined single-cell analysis, genetic data and functional validation, with several Relation researchers among its authors.

Big pharma backs AI and biological data

Furthermore, a 2025 analysis in Nature Biotechnology found that specialised dataset providers have become an important part of the AI-focused biopharma market.

The analysis identified several trends in recent partnerships. They included larger upfront payments, new therapeutic approaches and greater involvement from established biotechnology companies.

The report also highlighted the growing importance of disease-specific datasets for causal and generative machine-learning models. It cited GSK’s separate agreement with Ochre Bio, worth $37.5 million for access to human liver single-cell and perfused-organ data.

Another major deal involved AstraZeneca and Pathos AI, which entered a $200 million agreement with Tempus in 2025. Under that arrangement, Pathos planned to develop oncology foundation models using de-identified clinical, genomic and imaging data from more than 150,000 patients.

Access to high-quality biological data remains a major challenge for AI drug discovery. A Nature research highlight on federated learning in pharmaceutical research identified limited access to suitable training data as a major bottleneck.

However, pharmaceutical companies must also deal with restrictions on sharing proprietary information. As a result, partnerships now take different forms, including AI platform access, joint development, data licensing and the creation of new biological datasets.

In the GSK-Relation agreement, the companies will combine data generation with AI model development. Relation will produce human cellular datasets and use them to train AI models that can help identify potential drug targets.

Consequently, the partnership highlights a growing reality in AI drug discovery: better models may depend not only on more computing power, but also on better and more relevant biological data.

GSK, Relation Sign $110m AI Drug Discovery Deal

GSK has expanded its partnership with British biotechnology firm Relation Therapeutics in a research deal worth up to $110 million to use human-cell data and artificial intelligence to discover potential drug targets.

Under the agreement, Relation will generate large-scale datasets that show how human cells respond to genetic changes and drug treatments. The company will use the data to train AI models, including its MORGAN platform, to identify potential targets for new medicines.

The deal brings biological data generation and AI model development into one research programme. Relation combines laboratory experiments with computational analysis to create new information about human cells and disease.

GSK expands work with Relation

Meanwhile, the new agreement builds on earlier partnerships between GSK and Relation that focused on fibrotic diseases and osteoarthritis. Those projects involved observational studies designed to produce two functional disease datasets for analysis through Relation’s Lab-in-the-Loop platform.

The earlier research brought together human genetics, single-cell multi-omics from human tissue, functional assays and machine learning. Researchers used those tools to identify and validate potential disease targets.

Furthermore, the expanded partnership reflects a wider push by pharmaceutical companies to combine human biological data with AI. GSK has increasingly used genetics, functional genomics and machine learning to strengthen its drug discovery pipeline.

How Relation creates biological data

Relation describes its Lab-in-the-Loop system as a combination of laboratory experiments and computational analysis. Its work includes tissue profiling, single-cell and spatial transcriptomics, sequencing and target validation.

The company also uses machine learning to identify, rank and validate potential targets. In addition, it uses the technology to help design new experiments.

In response to the need for more detailed biological information, Relation conducts perturbation experiments that measure how genetic changes affect cellular features linked to disease. Researchers can then compare those results with genetic information and biological data collected from patients.

Public repositories also provide important material for biological foundation models. However, researchers face challenges when they combine data from studies that used different laboratory methods and experimental conditions.

A 2025 review in Experimental & Molecular Medicine highlighted repositories such as CZ CELLxGENE, the Human Cell Atlas and NCBI Gene Expression Omnibus as major sources of single-cell data. The review said CZ CELLxGENE provides access to more than 100 million standardised cells.

More data does not always mean better AI

However, researchers must deal with differences in sampling methods, sequencing protocols, experimental procedures and data-processing systems. Single-cell datasets can also contain technical noise and other artefacts.

Consequently, researchers need careful dataset selection, filtering, composition balancing and quality control when they train foundation models. Poor-quality or inconsistent data can affect how well an AI model represents biological systems.

Dataset duplication creates another problem. Similar cells can appear in several public databases, giving those cells more influence during training than researchers intended.

Furthermore, overlapping training and test datasets can create data-leakage risks. The 2025 review therefore stressed that researchers need high-quality, non-redundant datasets as much as they need strong model architecture.

Research published in Nature Methods in June 2026 also questioned the assumption that bigger datasets will always produce better single-cell AI models.

Researchers used a corpus containing 22.2 million cells to train 400 models and conducted 6,400 experiments. They found that current single-cell foundation models often reach performance plateaus after training on only part of the available data.

Unlike large language models, the researchers found no clear data-scaling pattern in which continually adding more training data consistently improves performance. Instead, they said developers need to balance model capacity, dataset size and computing resources.

Pharma turns to specialised datasets

Meanwhile, another 2025 study in Genome Biology assessed Geneformer and scGPT across several zero-shot tasks. The researchers found that the models did not consistently outperform simpler methods.

The study also identified challenges involving batch effects. It warned researchers against assuming that larger pretrained models automatically create better biological representations.

Relation has already applied its data-generation strategy to Osteomics, which the company describes as a proprietary functional single-cell bone atlas. The project combines patient-derived samples with single-cell and spatial omics, imaging, genomics, proteomics and clinical phenotype data.

The company uses Osteomics to study disease biology, potential drug targets, biomarkers and patient subgroups in osteoporosis. Hospitals and research partners in the UK and Australia participate in the observational study.

Research published in Nature Genetics last month also examined cellular and genetic factors linked to skeletal disease. The study combined single-cell analysis, genetic data and functional validation, with several Relation researchers among its authors.

Big pharma backs AI and biological data

Furthermore, a 2025 analysis in Nature Biotechnology found that specialised dataset providers have become an important part of the AI-focused biopharma market.

The analysis identified several trends in recent partnerships. They included larger upfront payments, new therapeutic approaches and greater involvement from established biotechnology companies.

The report also highlighted the growing importance of disease-specific datasets for causal and generative machine-learning models. It cited GSK’s separate agreement with Ochre Bio, worth $37.5 million for access to human liver single-cell and perfused-organ data.

Another major deal involved AstraZeneca and Pathos AI, which entered a $200 million agreement with Tempus in 2025. Under that arrangement, Pathos planned to develop oncology foundation models using de-identified clinical, genomic and imaging data from more than 150,000 patients.

Access to high-quality biological data remains a major challenge for AI drug discovery. A Nature research highlight on federated learning in pharmaceutical research identified limited access to suitable training data as a major bottleneck.

However, pharmaceutical companies must also deal with restrictions on sharing proprietary information. As a result, partnerships now take different forms, including AI platform access, joint development, data licensing and the creation of new biological datasets.

In the GSK-Relation agreement, the companies will combine data generation with AI model development. Relation will produce human cellular datasets and use them to train AI models that can help identify potential drug targets.

Consequently, the partnership highlights a growing reality in AI drug discovery: better models may depend not only on more computing power, but also on better and more relevant biological data.

External links

Reuters report on the GSK–Relation Therapeutics $110 million deal

GSK partnership information on Relation Therapeutics

Nature Methods study on single-cell foundation-model data scaling

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top