
AI Drug Discovery and Development Service Provider
The Trisolarans unfold a proton from its eleven-dimensional higher-dimensional state into a two-dimensional plane. The unfolded membrane spans an area large enough to enclose the entire Trisolaran planet. They etch massive circuits onto this "thicknessless giant film". After refolding, it becomes a "Sophon", a miniature supercomputer endowed with supreme intelligence.
Drawing on this concept, Xiong Zhaoping, CTO of Alphama Biotech, founded Suzhou Proton Unfold Technology Co., Ltd., hereinafter referred to as "Proton Unfold".
He stated: "The logic of scientific research bears striking similarities to the birth of the Sophon. Truly valuable information is often hidden within extremely complex multi-dimensional systems. We aim to 'unfold' this concealed knowledge, enable AI to fully comprehend it, and then refold it into a highly usable system to support researchers in decision-making."
The bottleneck for AI implementation lies not in algorithm accuracy, but in enabling non-technical business personnel without an AI background to intuitively perceive its value, establishing end-to-end workflows with embedded AI, and fostering an AI-First mindset for practical application.
Proton Unfold adheres to the "AI First" philosophy and commits to building a new-generation AI-native computing infrastructure for scientific research. It enables structural analysis, molecular design, virtual screening and R&D decision-making with lower barriers and higher efficiency. Serving as an AI-driven decision hub for research, it undertakes massive simulation, analysis and tool orchestration. Human researchers can therefore focus on high-level scientific creativity, research insights, wet-lab experimental validation and other critical tasks.
SciMiner Agent, its flagship product, was launched in November 2025. Within merely eight months, it surpassed 10,000 registered users, with token consumption witnessing exponential growth for consecutive months.
Amid the AI Agent boom, how will emerging AIDD SaaS architectures evolve? How can they reshape researchers'decision-making workflows? Moreover, how to integrate industrial chains and connect downstream CROs? VCBeat conducts an exclusive interview with Xiong Zhaoping, Co-founder and CEO of Proton Unfold.
1Proprietary High-Real-Time Molecular Knowledge Base: Breaking the R&D "Collision" Bottleneck
VCBeat: Starting from the incubation of Proton Unfold by Alphama Biotech, what new industry-end demands does it aim to meet?
Xiong Zhaoping: Despite years of development in AI-driven drug discovery, a significant pain point remains unresolved: the disconnect between AI and actual R&D workflows. One category of companies emphasizes "automated laboratories," aiming to fully take over wet-lab experiments with robotics, but the high equipment costs are not suitable for individualized or small-team research efforts. Another category focuses on "dry-lab" work, assisting researchers with data analysis, predictions, and experimental design.
The practical pain point addressed by Proton Unfold is that even after completing computations and predictions, scientists still need to switch between various tools, coordinate with CROs, track experimental progress, and manually organize and upload experimental data, thereby fragmenting the entire scientific research workflow into numerous disjointed steps.
These highly complex and coordination-intensive tasks can be fully handled by current AI capabilities, yet within the existing framework, they ultimately still rely on human effort. Therefore, we need a truly full-stack platform that seamlessly integrates the entire end-to-end research workflow.
VCBeat: What are the fundamental differences between SciMiner, the drug discovery decision-making AI agent launched by Proton, and traditional CROs or previous AI tools?
Xiong Zhaoping: The core division of labor and value boundaries among the three are entirely distinct. The core value of traditional CROs lies in their capacity for wet-lab execution, undertaking experimental tasks already defined by clients and delivering experimental data; traditional AI tools serve as point-solution auxiliary software, performing localized computations such as molecule generation and screening, and requiring manual integration of each step by human operators.
SciMiner is neither a tool nor a CRO; it is an intelligent decision-making operating system for drug discovery. Leveraging AI as the entry point for R&D, it autonomously comprehends research objectives, plans research pathways, and orchestrates tools across the entire workflow, fully integrating the iterative DMTA (Design-Make-Test-Analyze) closed loop.
Our model features clear role allocation: AI takes charge of knowledge integration, solution planning and experimental design within the digital realm; scientists retain core scientific creativity and high-level judgment; CROs are responsible for implementing wet-lab experiments at the physical level. Customers hold the initiative in R&D, while SciMiner serves as a process decision-making base that enables efficient connection between digital simulation and physical experiments.
SciMiner is capable of understanding drug discovery questions raised by users, automatically generating research routes, and intelligently invoking modules including information retrieval, target discovery, virtual screening, molecular optimization and synthetic route planning. It seamlessly supports full-modal drug discovery covering macromolecules, polypeptides, nucleic acids and small molecules.
VCBeat: Many AI pharmaceutical companies are developing intelligent agents. How does Proton Unfold build its competitive advantages and establish moats?
Xiong Zhaoping: Many teams across the industry are developing research-oriented intelligent agents, yet technical approaches in this sector are far from converging, with players adopting vastly different entry points. Numerous companies opt for lower-barrier directions at the initial stage, focusing on literature parsing and integrating public data for intelligence gathering. Nevertheless, such capabilities will soon be covered by foundational large model vendors.
As the capabilities of open-source agents continue to improve, many teams have begun leveraging these agents to invoke specialized tools for tasks such as molecular design and virtual screening. I consider these applications to be "low-hanging fruit"; while they can rapidly yield demo-level products, they are highly susceptible to long-term replacement by general-purpose agents like Claude Code or Codex.
A prevalent industry challenge lies in defining technical boundaries. It is difficult to distinguish directions worthy of long-term intensive research, those only achievable in the short term, and those not worth investment at all. Fortunately, the core investment behind Proton Unfold cannot be easily replaced directly by general large models and open-source intelligent agents.
The first barrier is the capability for automated extraction of chemical multimodal data. As early as in the molecular image recognition Kaggle competition hosted by BMS, we secured first place among nearly a thousand participating teams, raising the industry's recognition baseline from 80% to over 95%. Even with hand-drawn sketches, low-resolution scans, and images with complex backgrounds, our system can accurately identify molecular structures.
Even with strong performance metrics, real-world business scenarios are rife with various edge cases, making models prone to errors. Therefore, we have established a comprehensive data extraction and annotation system, PatMap, capable of mining multimodal molecular information — such as activity, targets, synthesis annotations, and ADMET properties — from massive volumes of patents and literature.
The conventional industry approach involves collecting data and training dedicated models for individual biochemical properties in isolation, which often leads to the predicament of having only dozens of usable samples per single task. We adopted an alternative technical pathway: rather than developing fragmented single-task predictors, we directly feed various types of heterogeneous, multimodal information from the literature into a molecular representation model to generate unified, standardized molecular representations. When a user queries a specific molecule, the system performs multidimensional similarity comparisons within the representation space, employing analogical reasoning to infer the potential physicochemical and biological properties of new molecules.
Even the most capable general-purpose large language models currently lack specialized molecular structure recognition capabilities, which is precisely the core gap we aim to fill.
The second barrier stems from scenario-based accumulation and user co-creation. SciMiner has been refined and developed in close collaboration with its users, continuously addressing the genuine tooling needs of frontline researchers and incorporating direct user feedback on a daily basis. Through this process, we have accumulated extensive insights into the practical boundaries of AI-driven drug discovery algorithms, real-world business data, and industry know-how, enabling our agents to execute tool calls and make research decisions with greater precision when tackling tasks.
There are many voices from the outside questioning the value of vertical Agent Harness, but my judgment is that, in professional R&D scenarios, Harness is the only viable solution for the real-world implementation of Agents.
The complexity of actual drug R&D operations increases exponentially with each decision step. Even an AI that never rests cannot exhaust all possibilities. While general-purpose agents can now easily achieve a 70% proficiency level, consistently selecting the highest-probability path at each step typically yields merely mediocre, 70% results.
High-quality decisions that truly determine project outcomes often require the introduction of accurate professional data at critical junctures. The core value of vertical intelligent agents such as SciMiner lies in optimizing decision-making logic at key nodes and deepening insights into professional scenarios.
2Accuracy Competition and Real-Time Practice
VCBeat: Speaking of industry baselines, how does Proton Unfold view the accuracy race within the industry?
Xiong Zhaoping: Accuracy competitions hold positive value by driving the industry to prioritize data accumulation and define scientific problems more clearly. However, it is inherently difficult to establish a fair benchmark in the biopharmaceutical field, and the high accuracy achieved in such competitions is entirely distinct from real-world industrial implementation.
First, the quality of many benchmark datasets is inherently insufficient, and high accuracy rates essentially reflect overfitting to experimental noise. A typical example is the classic virtual screening benchmark DUD-E: its decoy molecules and active compounds differ significantly in basic physicochemical properties such as molecular weight and lipophilicity. Consequently, models can achieve high scores simply by leveraging these basic physical attributes, without needing to learn the true protein–ligand binding mechanisms, rendering them entirely ineffective in real-world screening scenarios.
Second, there is a clear gap between leaderboard accuracy and clinical translation. A 2024 study by Boston Consulting Group revealed that the Phase I success rate for AI-discovered small-molecule drugs stands at 80%–90%, significantly higher than that of traditional R&D. However, upon entering Phase II, the success rate drops to the industry average of 40%. This is because Phase I trials primarily assess safety and pharmacokinetics, areas with abundant data where AI can perform effectively. In contrast, Phase II trials evaluate efficacy, which relies heavily on understanding disease mechanisms; the scarcity of high-quality data in this domain prevents algorithms from bridging the cognitive gap. This implies that, no matter how accurate single-attribute prediction is, it cannot replace the multi-parameter validation required for drug development.
Third, the lack of a universal platform capable of "wholesaling" new drugs means that all technological successes depend on specific systems. For instance, switching targets in ADC development essentially amounts to rebuilding the entire system from scratch; the same applies to PROTACs, molecular glues, and mRNA technologies, none of which are yet capable of delivering a stable pipeline of new drugs.
Currently, protein structure prediction represents the only widely accepted benchmark with clear definitions and robust data quality. CASP conducts double-blind assessments using unpublished experimental structures. This well-vetted standard has earned AlphaFold well-deserved recognition for its breakthroughs. Leveraging high-precision structures, the in-vitro binding accuracy for binder design has improved significantly. Nevertheless, bottlenecks in druggability remain unresolved, as high-quality data are still lacking for critical properties including solubility, metabolic stability and in-vivo toxicity. Many molecules exhibiting excellent in-vitro activity fail to deliver expected outcomes once subjected to animal experiments.
Ultimately, striving for accuracy is fine, but one should not rely solely on benchmark rankings.
VCBeat: Your team has also participated in numerous accuracy competitions. What impact have these competitions had on the deployment of Proton Unfold?
Xiong Zhaoping: Our team has directly benefited from accuracy-focused competitions. During my doctoral studies at the Shanghai Institute of Materia Medica, I secured first place in the Dream Challenge 2018 for multi-target prediction; later, I also took first place among nearly a thousand teams in the Kaggle competition on molecular image recognition hosted by BMS.
The benefits these competitions have provided us are tangible:
First, we developed a robust molecular recognition model, enabling us to enter the industry and gain a thorough understanding of real R&D pain points;
Secondly, through the competition, we have recruited many outstanding technical partners and built a core team;
More importantly, it allowed us to clearly recognize the boundaries of pure algorithmic models early on — even models that performed brilliantly in competitions inevitably encounter various edge cases not covered during training when deployed in real-world business scenarios, leading to unexpected outputs.
VCBeat: Why is "real-time capability" critical? How is the industry's molecular knowledge base with top-tier real-time performance implemented?
Xiong Zhaoping: For drug development, real-time performance directly determines the trial-and-error costs in early-stage R&D. The most typical pain point is falling into the "patent trap": many projects proceed in isolation for over six months, only to discover at the end that the target molecule or core design has already been covered by existing patents, rendering all prior investments futile.
Traditional databases, constrained by the high costs of manual annotation, struggle to incorporate the latest patents and literature in a timely manner. Furthermore, in an effort to standardize formats, they often discard substantial amounts of valuable non-standard information. Such data latency and information gaps directly increase the sunk costs associated with research and development.
To achieve industry-leading real-time performance, the core strategy is to replace "manual curation" with "automated intelligent extraction." Leveraging the aforementioned molecular recognition model, we developed a specialized molecular recognition annotation system, PatMap. This system has evolved into a literature and patent knowledge base built on large language models, enabling rapid and comprehensive extraction of molecular data from the latest publicly available literature and patents. This transforms the traditional "passive query-based" approach into "proactive risk warning," thereby achieving true high real-time performance.
Its core value lies in the continuous and seamless integration of the latest research advancements, experimental data, and patent information into the daily workflows of researchers, thereby directly benefiting frontline scientists. As workflow efficiency improves, user stickiness will naturally increase, creating a natural competitive barrier.
3Leveraging R&D and CRO bilateral resources to address supply-demand mismatches
VCBeat: What kind of "supply-demand mismatch" exists between the research sector and CROs?
Xiong Zhaoping: Many large pharmaceutical companies consider CRO matching a pseudo-demand, as they have established partnerships with top-tier CROs. However, friends from research institutions or startup teams often ask me, "No one seems to take on our small orders involving just dozens of molecules. Which CRO is both specialized and reliable?"
Globally, the combined market share of the top four CROs engaged in early-stage R&D accounts for less than 50%. At the early drug discovery stage, service categories are highly diverse, ranging from protein structure analysis and high-throughput screening to animal experiments. Standardization poses substantial challenges, and low market concentration is a persistent industry feature. A healthy ecosystem ought to accommodate a large number of small and medium-sized CROs that are vertically focused, specialized, and proficient in distinct experimental systems.
VCBeat: The issue of "supply-demand mismatch" is not simple. How does Proton Unfold development build and integrate internal and external workflows?
Xiong Zhaoping: This is precisely the most fundamental distinction between Proton Unfold and traditional internet matchmaking platforms — our entry point is not "transaction matchmaking," but rather establishing a scientific research workflow first, from which a bottom-up network of research services naturally emerges.
Leveraging a comprehensive DMTA cycle system, once the platform accumulates a sufficient number of pharmaceutical companies and researchers, CRO service providers will naturally converge. On the other end, we provide SciMiner as a daily research operating system free of charge to CROs, enabling standardized, visualized management of experimental data and automated progress reporting to clients.
After scientists complete molecular Design via AI on the platform and move on to the Make stage, they can dispatch experimental requirements to CROs integrated on the platform with one click. This enables the workflow to extend naturally along the industrial chain, bridging supply and demand and eliminating disconnects in wet-lab experiments.
VCBeat: How does the platform establish a supply-demand accumulation mechanism and build an ecological closed loop?
Xiong Zhaoping: Even large CROs cannot cover all niche segments, often leaving pharmaceutical companies to bear the costs of trial and error. Once a specific experimental system is proven effective, orders of the same type tend to surge simultaneously — the platform's value lies in intelligently matching these long-tail supply and demand needs.
Under the logic of "building workflows first, then extending service networks," when pharmaceutical companies and CROs collaborate on SciMiner, experimental processes and data delivery are fully digitized and auditable online, eliminating the need for manual monitoring. The system automatically accumulates actual delivery data from each provider, gradually forming capability profiles and credit records for CROs. This ultimately creates a self-driven ecological closed loop: high-quality service providers secure more orders, while pharmaceutical companies minimize their trial-and-error costs.
4Registered Users Surpass 10,000, Making It as Easy to Find a CRO as Ordering Food Delivery
VCBeat: What is the user scale of the SciMiner platform?
Xiong Zhaoping: SciMiner officially launched in November 2025, registered users rapidly surpassed 10,000, achieving in just eight months the user scale that traditional AI drug discovery platforms took six years to accumulate.
VCBeat: Why has such rapid user growth been achieved?
Xiong Zhaoping: Public acceptance of AI has significantly increased. Meanwhile, we have relaxed registration requirements by no longer mandating institutional email addresses, ensuring users can engage with the platform without concerns about privacy intrusion.
VCBeat: What are your current primary customer segments and business model?
Xiong Zhaoping: Many biotech firms cannot afford to maintain in-house computational chemistry teams, yet they urgently need AI to enhance R&D efficiency. Therefore, our current core customers are startup pharmaceutical companies and academic research users, for whom we provide private deployment and system subscription services, along with their partner CROs.
VCBeat: How to accompany scientists and facilitate their transformation into Biotech entrepreneurs?
Xiong Zhaoping: First, foster the workflow habit of "predict first, experiment later". Previously, upon obtaining a target, scientists would instinctively purchase molecules and conduct experiments immediately. Now, SciMiner enables them to recognize that leveraging AI to comb literature, predict properties and design synthetic routes before wet-lab work can avoid numerous detours. Once this workflow habit is established, they will maintain lifelong stickiness to the platform.
VCBeat: Facilitating supply-demand matching across the upstream and downstream sectors, with certain characteristics of third-party intermediary services. How can the profit flywheel be set in motion?
Xiong Zhaoping: We indeed possess the two-sided platform characteristics that connect users with CROs. In the early stage, we generate healthy cash flow through the private deployment and subscription of our AI system. In the mid-term, as scientific research workflows accumulate on the platform, we will activate a larger-scale commercialization flywheel through transaction matching and value-added services within our service network.
VCBeat: The "one-click communication" feature on your website is very internet-native. How do you approach the product logic that combines To-C and To-B elements?
Xiong Zhaoping: In the past, To-B software was widely perceived as difficult to use and offering a poor user experience. The root cause lay not in the capabilities of B-end product managers, but in the prohibitively high costs of development and iteration associated with traditional software. B-end software had to satisfy the complex management and control requirements of leadership while strictly adhering to development budgets, ultimately forcing compromises on the user experience for frontline employees.
However, with the powerful AI-assisted programming capabilities, the cost of software development and iteration has decreased exponentially. This empowers us to achieve one thing: meeting the depth of B-side enterprise-level management while offering a simple and smooth interactive experience akin to To C products.
"One-Click Communication," an extremely intuitive feature, stems from a simple premise: it is designed around the daily usage scenarios of frontline scientists.
Times have changed; it's time to offer B-side users something truly premium.
VCBeat: What plans and progress does Proton Unfold have in terms of international layout and financing?
Xiong Zhaoping: Alphama Biotech, the parent company of Proton Unfold, derives nearly two-thirds of its orders from overseas clients and has established multiple in-depth strategic collaborations with numerous multinational corporations (MNCs).
Therefore, Proton Unfold has possessed an international DNA since its inception. We are highly committed to efficiently exporting and promoting China's premium, cost-effective CRO R&D service resources to global clients through our platform. In the near term, we will also advance a round of financing to operate Proton Unfold as a fully independent entity, thereby accelerating our expansion into overseas markets.
5Value Realization at the PCC Stage: The Expansion Logic of AI Drug Discovery SaaS
VCBeat: How do you view the current entrepreneurial pathways and ecosystem landscape in AI-driven drug discovery?
Xiong Zhaoping: Many AI-driven drug discovery entrepreneurs initially aimed for an asset-light strategy — licensing out their pipelines once they reached the preclinical candidate (PCC) stage. However, a prevailing sentiment later emerged in the market that "PCC-only assets command lower valuations." As a result, many teams gritted their teeth and pushed their candidates into clinical development, only to find that the process became increasingly uncontrollable and capital-intensive as it progressed.
The perception that "assets cannot command high valuations at the PCC stage" is actually a misconception. In the first half of 2026, the total potential transaction value of out-licensing domestic innovative drug pipelines to multinational pharmaceutical companies exceeded 110 billion RMB. Among these licensed pipelines, 75% were at the PCC stage or even earlier stages of research and development. This demonstrates that MNCs are fully willing to invest in high-quality early-stage programs. If a program fails to be out-licensed, the root cause lies in insufficient data quality and project value failing to meet buyers' criteria, rather than a lack of inherent commercial value associated with the PCC stage itself.
In the past, out-licensing deals for AI-designed pipelines were isolated cases, yet such transactions are now becoming increasingly commonplace. The core value of AI-driven drug discovery has undergone industrial "peer review" at the PCC stage — recognition granted by buyers following rigorous, capital-backed due diligence. Once a transaction closes, the value delivered by AI in early-stage molecular discovery and optimization is fully validated. As for the success or failure of pipelines after entering clinical development, this follows inherent objective laws of pharmaceutical R&D. One should not retroactively negate the tangible value created by AI in upstream research based on late-stage outcomes.
Proton Unfold's reconstruction of the scientific research workflow follows the same core logic: empowering pharmaceutical companies to generate more high-quality PCC molecules with higher efficiency and lower costs. These molecules can withstand industrial "peer review" and ultimately realize value through license-out transactions.
VCBeat: What significant advancements will Proton Unfold see over the next three to five years?
Xiong Zhaoping: Within three years, we aim to make SciMiner the go-to research operating system that researchers and medicinal chemistry experts open first thing every workday, making "AI First" their daily habit.
Within five years, we aim to drive a qualitative transformation in the industry's research service network — enabling researchers to search for, access, and track CRO services with the same transparency, traceability, and evaluability as hailing a ride-sharing car on a mobile phone today.
VCBeat: What kind of AI agent story does Proton Unfold aim to tell?
Xiong Zhaoping: The reason we are firmly committed to this path is rooted in a simple judgment: compared to striving for precise perfection within a predetermined narrow niche, aligning with major industry trends and making the right strategic moves in the broader direction delivers far greater value.
Many industry insiders are skeptical about AI-driven drug discovery SaaS, believing that the market ceiling for this sector is low, with Schrödinger representing its upper limit. However, I hold a contrary view: it is not that the potential of SaaS has peaked, but rather that the penetration rate of AI within the pharmaceutical industry chain remains significantly underdeveloped. In the early years, AI-powered drug discovery software was utilized by only about one in a thousand computational chemists and computational biologists. In large pharmaceutical companies with tens of thousands of employees, dedicated computational teams might consist of merely dozens of specialists, while many small and mid-sized pharmaceutical firms did not even have such roles. Given this limited user base, pharmaceutical companies' budgets for AI software naturally remained constrained.
DeepSeek marks a watershed for AI applications. Its chain-of-thought output enables AI to demonstrate human-like logical reasoning and decision-making capabilities, substantially broadening the boundaries of AI adoption. It has begun to be widely integrated into pharmaceutical companies' workflows and evolved into universal infrastructure accessible to all staff. This represents an essential leap at the user scale, rather than a simple quantitative accumulation.
Meanwhile, the transition of AI from a "black-box tool" to an "interpretable expert" is the key to breaking down trust barriers in vertical industries.
The core narrative behind Proton Unfold's vision for AI Agents lies in bringing AI capabilities down to frontline R&D. It aims to equip every bench scientist with a "super digital assistant" proficient in chemistry, biology and end-to-end R&D workflow coordination. It also enables every pharmaceutical enterprise to operate a tireless, continuously iterating R&D decision-making hub.