
Internet Comprehensive Service Provider
Using AI to simulate how cells respond to various perturbations is the long-term vision of the AI virtual cell (AIVC).
In reality, however, the two most important types of perturbation data — genetic screening and chemical drug screening — have long been fragmented, making unified modeling difficult.
Facing these industry challenges, Tencent's AI for Life Sciences Lab, in collaboration with the team of Prof. Li Min at Central South University, has built a new AI framework called UniPert-G2CP, and the related paper was recently accepted by "Cell".
The model integrates multimodal representations of molecular perturbation factors (causes) and enables transfer learning from genetic to chemical perturbation phenotypes (outcomes), thereby connecting genetic and chemical screening to predict perturbations and uncover mechanisms of action more accurately, while enabling efficient computer-assisted drug screening.
As the first AI virtual cell research from China published in the main journal "Cell", its groundbreaking contribution lies in finding a new entry point to answer the question: why does the same compound respond so differently across different genetic and cellular backgrounds?
Over the past decade, AI has generated a series of milestone achievements in drug development. AlphaFold solved the problem of protein structure prediction; molecular generation models can create new molecules out of thin air in chemical space; QSAR and ADMET prediction can infer physicochemical properties from molecular structures; and AI models such as DiffDock can predict how compounds bind to given targets.
But all of these achievements share a common blind spot: they model compounds or compound–single-protein-target interactions, with almost no consideration of the downstream effects after a compound enters real cells.
The inability to simulate cellular changes means that excellent laboratory results may show significant efficacy deviations in clinical use,
In the real world, a compound predicted by AI to have "high affinity + good ADMET + excellent drug-likeness" can still be completely ineffective in BT474 breast cancer cells, yet produce significant transcriptomic perturbations in MCF7 cells — because the downstream pathways, gene expression profiles, and mutational backgrounds of these two cell types are completely different.
When computational results cannot be reproduced, pharmaceutical companies often bear huge losses. Multiple studies, including those from Tufts, point out that the average cost of bringing a new drug from project initiation to market is about USD 2.6 billion, with R&D cycles of 10 to 15 years, while the long-term success rate of Phase II clinical trials has remained below 30%. Large numbers of compounds that "look good" at the molecular level fail once they reach cells and patients.
The current cost of failure is concentrated mainly in the cell/animal validation stage between preclinical and clinical development. If computational modeling does not cover cellular changes, AI will struggle to deliver the value expected of it.
UniPert-G2CP, proposed by the joint Tencent–Central South University team, is a solution designed precisely for this dilemma. It is a two-stage deep learning framework that starts from gene-level perturbation information to predict the phenotypic responses of compounds in specific cellular contexts.
UniPert is a multimodal molecular representation model based on contrastive learning, with the goal of pulling functionally similar genes, drugs with the same mechanism, and gene–drug pairs acting on the same pathway into the same semantic space.
For chemical molecules, it uses SMILES strings and featurizes substructures via ECFP (extended connectivity fingerprints); for genetic perturbations, it uses protein sequences as the base input, integrating multiple encoding approaches including protein language models (ESM), multiple sequence alignment (MSA), and graph neural networks (GNN).
On this basis, UniPert uses prior biological network information and contrastive learning to map both into a shared semantic space, bringing functionally similar cross-domain perturbation molecules closer together.
According to the paper's data: on drug mechanism of action (MoA) classification tasks, UniPert's intra-class semantic separability improved from 1.61 to 1.85 compared with traditional ECFP methods, with accuracy up 14.4%; across 35 protein pharmacological classes (PCL), clustering ARI and NMI improved by 76.3% and 47.0%, respectively.
With unified modeling of genetic screening and chemical drugs completed, G2CP (cross-domain phenotype transfer learning) takes over to process the data.
It first pretrains the model on genetic perturbation datasets, enabling the model to learn the basic response patterns of cells under different genetic perturbations; it then fine-tunes on limited chemical perturbation data, transferring genetic knowledge to the drug development domain.
Based on this strategy, G2CP effectively compensates for the scarcity of chemical screening data and significantly improves the model's generalization prediction ability on unseen chemical perturbations.
With the two stages combined, UniPert-G2CP successfully converts the response patterns of genes and drugs into a unified language, then transfers and applies genetic knowledge to the drug development domain.
According to the paper's data: in validation on the LINCS dataset (covering 4,994 genes and 7,860 small molecules across five cancer cell lines), UniPert-G2CP improved prediction accuracy by 375.4% using only 20% of the chemical training data.
In addition, in the ESR1 (estrogen receptor) case, the model successfully reproduced the phenotypic shift caused by breast cancer drug-resistance mutations and identified resistance-related genes and pathways. This means UniPert-G2CP has moved beyond the limitation of traditional models' "prediction" capability, successfully "analyzing" changes and completing a closed loop of biological interpretation.

UniPert-G2CP's breakthrough is no accident. Since 2022, Tencent's Life Sciences Lab has been attempting to describe living systems in a unified language. Over the following years, the lab successively released models including scBERT, scPROTEIN, and scTranslator, with results published in top-tier journals.
This aligns with the direction emphasized by Wu Wenda, president of Tencent Health: under the core philosophy of "Technology for Good", Tencent has continued to invest in six areas — protein structure prediction, virtual screening, molecular design optimization, activity prediction, ADMET druggability prediction, and drug resistance prediction — strongly advancing the development of life sciences globally.
Back to UniPert-G2CP. After validating the results, Tencent's research team plans to make UniPert-G2CP a core foundational capability, integrated into the upcoming AI life sciences Agent platform. This will allow virtual cell capabilities, in the form of Skills, to serve experimental scientists' real research workflows directly.
From the discovery and validation of drug targets, to the evaluation of candidate compounds' efficacy in specific cellular contexts, to the elucidation of drug resistance mechanisms — Tencent's goal is to provide computable, predictable intelligent support for the key stages of drug development, ultimately connecting the entire virtual drug development chain from target discovery to efficacy evaluation.
Because UniPert-G2CP provides the basic computational unit of the virtual cell, its significance is like that of GPUs to deep learning. Once such computational units are scaled, modularized, and agentized, digital twin clinics will no longer be detached from laboratories — they will truly become a reality.
At that point, doctors will be able to build real-time digital twin cell populations of patients in the cloud, run virtual trials of multiple candidate treatment options, and design optimal combination treatment plans for patients based on the feedback.
With greater efficiency and less trial and error, we may see more patients recover at a lower cost, truly realizing the benevolent potential of technology.