Novel benchmark launched to test if AI models can truly make drug discoveries

Company news

Novel benchmark launched to test if AI models can truly make drug discoveries

14 Aug, 2026


Insilico Medicine’s launch of a benchmarking service invites developers to test frontier and foundation models against real drug discovery decisions rather than memorised ‘exam questions’


Artificial intelligence (AI) models have been adopted across drug discovery at pace, yet two questions have become impossible to avoid: 

    • Is AI capable of finding novel medicines?
    • Has AI simply learned to pass exams? 

Most of the AI benchmarks currently used to answer that question have been compromised by data contamination, in which a model achieves an artificially inflated score because the answers to its test questions already sit within its training data. Faced with a genuine drug candidate decision, the same model often falters.

To close that gap, Insilico Medicine, a clinical-stage generative AI-driven drug discovery company, has launched the Drug Discovery and Development (DDD) Benchmark as a Service (BaaS). The framework is anchored in decontaminated real-world data and Insilico’s own validated programmes, and it has been built to measure how frontier AI and foundation models, including those built on rival architectures, perform against real tasks across medicinal chemistry, chemical synthesis, disease biology, clinical development and longevity research.

The service is available to any organisation developing frontier AI for drug discovery, or one that uses foundation models in its research, and it provides an independent, real-world measure of a model’s performance.

It is structured around two complementary evaluation suites:

    • The first – Drug Discovery Foundations – comprises more than 300 evaluations built from proprietary, out-of-distribution test sets together with rigorously decontaminated public data. It has been designed to measure a model’s core competencies across the discipline, spanning disease biology, molecular property prediction and optimisation, retrosynthesis – working backwards from a target molecule to plan how it might be synthesised – structure-based design and clinical development.
    • The second – Drug Candidate Essentials – assesses a model’s ability to navigate a complete drug discovery programme, from the initial discovery of a promising compound through to preclinical candidate (PCC) nomination. Its reference baselines are anchored in Insilico’s own validated programmes, and it tests whether a model can make the sequential decisions demanded by real drug discovery.

The process to be benchmarked is straightforward. An organisation submits any model served through a standard chat-completions application programming interface to Insilico Medicine, which scores the resulting outputs against expert reference baselines and returns a standardised scorecard comparing performance with leading models.

Results arrive with a verified score report that allows organisations to certify performance for internal teams and for partners. Those seeking public recognition may also publish their results on the DDD Benchmark’s public leaderboard, giving the sector a transparent, like-for-like view of model capability.

AI systems have increasingly begun to act as agents, able to plan experiments, reason over experimental data and call tools compatible with the Model Context Protocol. The DDD Benchmark has been designed to establish whether those capabilities translate into sound, real-world drug discovery decisions, rather than merely strong performance on abstract tasks.

“The rapid progress of AI has made one question more urgent than ever: can these models actually discover drugs? For more than a decade, we have built and validated generative AI across the entire drug discovery value chain, nominating 31 preclinical candidates in six years and advancing an AI-discovered and AI-designed medicine into Phase III.

“The DDD Benchmark converts that real-world experience into a rigorous, standardised yardstick that the whole field can use. As we advance towards pharmaceutical superintelligence, measuring genuine capability, not memorised answers, is how we ensure AI delivers the highest-quality, differentiated medicines and helps extend healthy, productive longevity for people everywhere,” said Alex Zhavoronkov, founder and chief executive of Insilico Medicine.

The DDD Benchmark builds on Insilico’s existing Pharma.AI platform and its MMAI Gym, the company’s post-training environment for scientific AI. Since it was founded, Insilico has nominated 31 preclinical candidates, received more than ten investigational new drug clearances, and compressed the average timeline to PCC nomination to under 18 months, compared with the up to four years required in traditional drug discovery. Its lead programme, Rentosertib (ISM001-055), is a first-in-class, AI-discovered and AI-designed TRAF2- and NCK-interacting kinase inhibitor, now in Phase III development for idiopathic pulmonary fibrosis.


Find out more: dddbench.insilico.com


Latest News

Lab Asia 33.4 - August 2026

Explore our Digital Edition

Discover the latest news and research

Digital edition

Explore Our Other Sites

Envirotech Online
Satellite technology enhances river flow monitors
Explore more Arrow
Pollution Solutions Online
Leading UK biogas operator places first orders for new FlowSep technology
Explore more Arrow
Petro Online
Smart fuel lubricity testers for modern labs
Explore more Arrow
Chromatography Today
Unlock high-resolution analysis of therapeutic oligonucleotides
Explore more Arrow