AI drug discovery still lacks proof of clinical impact, new paper argues
A commentary on Science.org takes up a new paper, linked from Nature.com, that assesses how well AI methods are actually doing in drug discovery. The commentary singles out the paper's authors as people who know the field and are not trying to sell anything, and it endorses their central conclusion: despite a wide variety of AI methods being developed, applied, and benchmarked across the industry, evidence that any of it has translated into clinically relevant impact is, so far, disappointingly limited. The paper describes the field as sitting at a stage of 'absence of evidence' rather than 'evidence of absence,' meaning the effect has not yet been shown, not that it has been shown not to exist.
The piece argues that the decisions which would matter most, choosing disease areas and drug targets, picking lead molecules and clinical candidates, designing trials, are exactly the areas current AI techniques are least equipped to help with. That is not treated as a permanent limit; there is no fundamental reason these tasks cannot improve, but progress is expected to be slower and more expensive than the marketing around AI in drug discovery tends to suggest. Part of the difficulty, the commentary adds, is that measuring real-world effects in drug development has always been hard, AI or not: successful projects tend to get credited to good judgment and hard work, while failed projects, even ones run just as competently, quietly stop getting discussed. That selective storytelling makes it easy for any project with an AI component to claim the credit for success while AI-assisted failures pass unremarked.
The paper singles out Phase II clinical trial success rates as the number that would actually demonstrate whether AI is helping, since early-stage discovery work is a comparatively small share of the time and money spent next to running clinical trials. Non-AI initiatives have tried and failed to move that number before; on the evidence so far, the commentary says the AI era looks like more of the same.
The paper's central prescription is that AI work in drug discovery must move away from modelling whatever data happens to be readily available, since that data is unlikely to move the needle, and toward generating the data actually needed, even where that means committing to substantial new data collection. The commentary calls this a worthy goal that is also expensive and uncertain: choosing to generate new data instead of reusing what is already on hand means accepting murky odds of success alongside a high potential reward, and it means telling investors and the press that results will take longer than they have been led to expect, a message the piece suggests few in the field want to be first to deliver.
The commentary also traces the standard path a drug candidate has to survive: working against an isolated protein is not the same as working in cells, which is not the same as working in rodents, which is not the same as clearing a two-week dog toxicology study, which is still not the same as working in human patients, and AI systems that meaningfully move a candidate along that chain into real clinical success are described as genuinely hard to find. A deeper problem, the paper argues, is the data itself: drug discovery already holds large amounts of compound and assay data, but the number of variables in assay conditions and in biological systems makes it unlikely that machine learning alone can make sense of it, and the field does not know how to clean or categorize that data for AI use, or whether that is even possible. Confounding factors, conditionality, and a lack of transparency in drug-discovery data are not always understood by the people trying to model it, so the limits of labelling and using that labelled data go unrecognized too, and because practitioners tend to trust the labels they are given, that problem carries straight through into the benchmarks used across the field. In areas like image and speech recognition, progress on benchmarks has generally tracked real improvement in actual use; the paper says that correspondence needs to be treated with caution in drug discovery.
The paper closes with recommendations for AI companies and researchers: think about why a given technique or technology is being used rather than adopting it simply because it is newly available, ask whether it can realistically raise the odds of clinical success and how that would be measured, and check whether benchmarks and targets are actually built to reveal progress rather than to manufacture the appearance of it. The commentary argues that taking those recommendations seriously would benefit the field, spare investors some bad experiences, and quiet much of the surrounding hype, while noting that not everyone in the field is necessarily aligned with those goals.
Key facts
- A Science.org commentary highlights a new paper, linked from Nature.com, concluding that AI methods in drug discovery, despite being widely developed and benchmarked, show disappointingly limited evidence of real clinical impact so far.
- The paper frames this as an absence of evidence rather than evidence of absence, and argues AI work must shift from modelling data that is merely available to generating the data actually needed.
- It names Phase II clinical trial success rates, not early discovery metrics, as the number that would prove real impact, since early-stage work is comparatively minor in cost next to running clinical trials.
- The commentary warns that confounding factors and inconsistent labelling in drug-discovery data undermine benchmarks built on it, unlike benchmarks in image or speech recognition, which have generally tracked real-world gains.
- Neither the commentary nor the paper it cites names a specific company, drug, or AI model, and neither gives percentages or success-rate figures, even though Phase II success rates are treated as the metric that matters most.
Why it matters
AI in drug discovery draws a heavy stream of articles, conference talks, and press releases, which makes it hard to tell how much of it is real. This commentary matters because it comes from a source the post itself credits as knowledgeable and not selling anything, and its conclusion cuts against the prevailing pitch: after years of AI methods being developed, applied, and benchmarked, evidence of real clinically relevant impact remains disappointingly limited. Calling the field an 'absence of evidence' rather than an 'evidence of absence' matters too. It resets expectations without closing the door: the effect has not been shown yet, not proven impossible.
Who it affects
AI vendors and drug-discovery startups selling tools into pharma; the pharma companies and researchers deciding whether to trust those tools for target selection, candidate picking, or trial design; and investors funding AI-driven drug pipelines, who the piece suggests should expect longer timelines than they have been told. It also affects the press and publicists behind the wave of promotional coverage the commentary pushes back against.
How to use it
The paper's advice reads like a checklist for evaluating an AI drug-discovery claim: ask why a technique is being used rather than assuming newness is a reason on its own; ask whether it can realistically raise the odds of clinical success and how that would be measured; and check whether a benchmark or target is actually built to reveal progress, rather than to manufacture the look of it. Two anchors help with that check. One is the standard translational chain a candidate has to survive: isolated protein, then cells, then rodents, then a two-week dog toxicology study, then human patients, since a result that only holds at an early step proves little on its own. The other is Phase II clinical trial success rates, which the paper treats as the real scoreboard, ahead of any early-discovery metric a vendor might quote.
How solid is it
The commentary vouches for the paper's authors as domain experts without a product to sell, which is a real claim to credibility even though it cannot be checked independently from the piece alone. What can be checked is that the piece stays argumentative rather than empirical: it names no company, drug, or AI model, and gives no percentages, success rates, or trial counts of its own, even while treating Phase II success rates as the metric that matters most. That makes it a synthesis of an existing evidence gap, not a new dataset closing it.
Risks and caveats
The commentary flags a selection-bias risk in how the field talks about itself: successful projects get credited to skill and leadership, failed ones quietly stop getting discussed, which can make AI's actual contribution hard to isolate either way. The paper's own prescription carries a risk too: generating new data instead of modelling what is already on hand is expensive with uncertain payoff, and it requires telling investors and the press that results will take longer than they have been led to expect, a message the piece suggests few want to deliver first. The underlying data carries a further risk: confounding factors and inconsistent labelling in drug-discovery datasets are not always understood by the people modelling them, which quietly undermines benchmarks built on that data, in a way that image and speech recognition benchmarks, which have generally tracked real-world gains, do not share.
“Although a wide variety of AI methods have been developed, applied and benchmarked, evidence of their clinically relevant impact is, so far, disappointingly limited.”
— the paper's authors, quoted in the Science.org commentary