Company news
Insilico Medicine’s launch of a benchmarking service invites developers to test frontier and foundation models against real drug discovery decisions rather than memorised ‘exam questions’
Artificial intelligence (AI) models have been adopted across drug discovery at pace, yet two questions have become impossible to avoid:
Most of the AI benchmarks currently used to answer that question have been compromised by data contamination, in which a model achieves an artificially inflated score because the answers to its test questions already sit within its training data. Faced with a genuine drug candidate decision, the same model often falters.
To close that gap, Insilico Medicine, a clinical-stage generative AI-driven drug discovery company, has launched the Drug Discovery and Development (DDD) Benchmark as a Service (BaaS). The framework is anchored in decontaminated real-world data and Insilico’s own validated programmes, and it has been built to measure how frontier AI and foundation models, including those built on rival architectures, perform against real tasks across medicinal chemistry, chemical synthesis, disease biology, clinical development and longevity research.
The service is available to any organisation developing frontier AI for drug discovery, or one that uses foundation models in its research, and it provides an independent, real-world measure of a model’s performance.
It is structured around two complementary evaluation suites:
The process to be benchmarked is straightforward. An organisation submits any model served through a standard chat-completions application programming interface to Insilico Medicine, which scores the resulting outputs against expert reference baselines and returns a standardised scorecard comparing performance with leading models.
Results arrive with a verified score report that allows organisations to certify performance for internal teams and for partners. Those seeking public recognition may also publish their results on the DDD Benchmark’s public leaderboard, giving the sector a transparent, like-for-like view of model capability.
AI systems have increasingly begun to act as agents, able to plan experiments, reason over experimental data and call tools compatible with the Model Context Protocol. The DDD Benchmark has been designed to establish whether those capabilities translate into sound, real-world drug discovery decisions, rather than merely strong performance on abstract tasks.
“The rapid progress of AI has made one question more urgent than ever: can these models actually discover drugs? For more than a decade, we have built and validated generative AI across the entire drug discovery value chain, nominating 31 preclinical candidates in six years and advancing an AI-discovered and AI-designed medicine into Phase III.
“The DDD Benchmark converts that real-world experience into a rigorous, standardised yardstick that the whole field can use. As we advance towards pharmaceutical superintelligence, measuring genuine capability, not memorised answers, is how we ensure AI delivers the highest-quality, differentiated medicines and helps extend healthy, productive longevity for people everywhere,” said Alex Zhavoronkov, founder and chief executive of Insilico Medicine.
The DDD Benchmark builds on Insilico’s existing Pharma.AI platform and its MMAI Gym, the company’s post-training environment for scientific AI. Since it was founded, Insilico has nominated 31 preclinical candidates, received more than ten investigational new drug clearances, and compressed the average timeline to PCC nomination to under 18 months, compared with the up to four years required in traditional drug discovery. Its lead programme, Rentosertib (ISM001-055), is a first-in-class, AI-discovered and AI-designed TRAF2- and NCK-interacting kinase inhibitor, now in Phase III development for idiopathic pulmonary fibrosis.
Find out more: dddbench.insilico.com
Lab Asia 33.4 - August 2026