NomosLogic
Discovery Does Not Fail for Want of Capability
Back to Blog
deeptechdrug discoveryevidence governancedecision-grade evidencepredictive validitytarget validationconstraint inferencegoverned searchkill conditionspre-registrationfalsificationcalibrationabstentionprovenancereproducibilityEroom's lawphase II attritiontranslational medicinepreclinical researchcomputational biologyresearch integritybiopharmaclinical developmentdecision sciencemolecular medicine

Discovery Does Not Fail for Want of Capability

Matt HardySeptember 4, 202610 min read

Discovery Does Not Fail for Want of Capability


Computational Magnitudes of Discovery

Drug discovery has multiplied its computational capability by orders of magnitude and has not improved its success rate. That sentence is not a complaint about the difficulty of biology. It is a measurement, and it carries a conclusion most of the field has not accepted.

Approvals per inflation-adjusted research dollar have halved approximately every nine years since 1950. The remarkable feature of that curve is not its slope. Declining productivity in a maturing industry is unsurprising. The remarkable feature is its indifference to everything that was supposed to bend it. The series passes without inflection through high-throughput screening, combinatorial chemistry, the sequencing of the human genome, systematic RNA interference screening, monoclonal antibody platforms, and the solution of protein structure prediction. Each of those was, at the time, correctly described as a large increase in capability. Several were increases of several orders of magnitude in the throughput of a specific step. None produced a visible discontinuity in approvals per dollar.

If capability increases by orders of magnitude and outcomes do not respond, then capability was not the binding constraint. Something else was, and continued to be, throughout the entire period, unaffected by every one of those interventions.

The Arithmetic That Explains It

There is a formal version of this, and it converts an impression into a number. Scannell and Bosley modelled a discovery cascade as a sequence of noisy filters and asked what happens to yield as a function of two variables: the throughput of the cascade, and the predictive validity of the models used to filter, meaning the correlation between a model's output and the clinical outcome it stands in for.

Yield turns out to be extremely sensitive to predictive validity and comparatively insensitive to throughput. A small decline in the correlation between a preclinical model and the clinical endpoint can offset an increase of orders of magnitude in the number of hypotheses tested. Once stated, the intuition is hard to unsee. If the prior probability that a target hypothesis is correct is low, and the filter applied to it is weakly correlated with truth, then pushing more hypotheses through the filter increases the absolute number of false positives faster than the number of true ones. The false positives are what proceed to consume the expensive part of the pipeline.

Fifty years of investment went to throughput. Predictive validity, which dominates the yield, received far less, partly because it is harder to measure and partly because it does not produce a visible deliverable in a quarter.

Three consequences follow, and the third is the one that makes the problem tractable. Predictive validity is not primarily a property of instruments. It is a property of the relationship between what a model measures and what the clinic requires, and that relationship is established by the design of the evidence, not by the precision of the measurement. A highly precise assay measuring the wrong thing has high reliability and low validity, and the field has systematically confused the two. Which means predictive validity can be raised by governance, with no improvement in capability at all. A requirement that a phenotype be confirmed by an orthogonal manipulation raises the validity of an evidence base without a new instrument. A requirement that a falsification condition be declared before the data are examined raises it further. These are free in capability terms and expensive only in discipline.

Why Capability Downstream Makes Things Worse

Consider where each of those historical gains landed. Target selection precedes screening, precedes structure determination, precedes molecular design, precedes formulation. Every one of the gains above occurred downstream of the decision that determines the outcome. If the target selection was wrong, a downstream capability gain has exactly one effect: it allows the program to reach the point of failure faster, at higher confidence, having spent more.

This is not neutral. It is negative, and the pharmacokinetics case makes the negativity measurable. Pharmacokinetic failure has declined substantially as a cause of attrition, achieved through better absorption and metabolism assays, better allometric scaling, better physiologically based modeling, and property optimization moved earlier. Programs that would once have failed in Phase I for inadequate exposure now largely do not reach Phase I, or reach it with adequate exposure. The consequence is that programs which formerly failed early and cheaply for a pharmacokinetic reason now proceed to fail later and expensively for an efficacy reason. The capability gain converted a cheap failure into an expensive one, while improving the metric it was designed to improve, which is why the effect is invisible to any accounting that tracks pharmacokinetic attrition rather than program economics.

The modal expensive failure in this industry is a compound that was made correctly, behaved as designed in the assays it was designed against, was administered at exposures believed adequate, and did not produce the clinical benefit its mechanism predicted. That is not a failure of chemistry. It is the failure of an inference made years earlier, at a point when the cost of being wrong was still measured in analyst time.

Five Principles

If the binding constraint is evidentiary, the remedy is a standard for what evidence licenses a commitment. Five principles do most of the work, and none of them requires technology that does not already exist.

Decision grade is not publication grade. These are different standards calibrated for different error costs. The scientific literature is designed as a proposal mechanism, and it correctly admits claims that will not survive. A literature claim imported at its published confidence is therefore overconfident by a predictable amount. Worse, a hypothesis resting on three individually plausible findings is much weaker than any of them, and the multiplication is almost never performed. Correlated evidence has correlated failure modes, so three findings resting on one uncertain reagent are one finding. A decision-grade claim requires four properties, as gates rather than trade-offs: provenance that resolves to observations through re-executable transformations, a falsification condition recorded before the evidence was examined, a confidence with a demonstrated relationship to observed outcome frequency, and reproducibility in the narrow sense.

Kill conditions must be typed, owned, and timestamped before the data are seen. A kill condition does not work by making anyone more rigorous. It works by spending the interpreter's discretion over a specific future result in advance. This is why a condition written after the data are seen has no function at all, despite being textually identical to one written before, and why the externally verifiable timestamp is the condition's entire evidentiary content. Six types cover most decisions: threshold, ordering, dose response, replication, orthogonality, and provenance. The provenance condition is the interesting one, because it can fire on a favorable result. Whether it ever has is close to a complete diagnostic of whether a governance regime is real.

Calibrated or abstain. A system required to answer every question will answer the ones it cannot, and its errors will be formally indistinguishable from its valid outputs, because the output format does not encode whether the query fell inside the model's region of support. The classification must therefore admit a third outcome: not decidable at the confidence this decision requires. Abstention is only legitimate when paired with a named, costed, scheduled resolving experiment, which is what separates it from evasion. And most confidences in discovery are verbal registers with no agreed mapping to frequencies, which means they cannot be paired with outcomes, which means an organization cannot compute its own hit rate and therefore cannot distinguish judgment from luck.

Never launder. Laundering is the progressive loss of the qualifications that determine what a claim licenses, through a chain of individually faithful compressions, so that a claim arrives at a decision more confident than any claim beneath it. A word like validated at the top corresponds to nothing at the bottom. It is more damaging than fabrication, because it is performed by everyone as a condition of doing the work, no single step is defective, and the intermediate states that would prove what was lost are destroyed as routine housekeeping. Agreement between models trained on shared data is not confirmation. The remedy is a recomputability rule: the gate accepts only the version that can be recomputed from its inputs.

Inference and search are two acts, not one. Much of the confusion in computational discovery comes from treating them as one activity because one team performs both with one budget on one cluster. Inference determines what is true about a biological system: which components are structurally necessary, which are substitutable, what the system does when a component is removed. Search determines what to build in response. Inference is constrained by data and, given a specified model, produces a determinate answer. Search is constrained by an objective and produces a distribution of candidates from which one must be chosen. A program can characterize a system correctly and fail because the molecule was wrong. A program can generate an excellent molecule and fail because the target was substitutable. From the outside, both look identical: a compound that engaged its target in vitro and did nothing durable in a patient. Governance that cannot distinguish these two cases cannot learn from either.

How to Tell Whether a Regime Is Real

Every one of these practices is cheap to fake. A group required to produce provenance records will produce provenance records, and the records can be complete, well formed, and never consulted. Provenance that is never queried imposes the entire cost and delivers none of the benefit. Kill conditions can be decorative. Conditions can be inflated until a gate becomes a scorecard. Summaries can be marked decision grade until the marking means nothing.

So ignore the documents. Ask three questions. What is the kill rate. Which programs were stopped. Has a provenance condition ever fired on a favorable result. Those three answers describe a regime. The documents describe an intention. An organization that cannot point to a single instance, after two years, in which someone was prevented from advancing a program by reading the record has built a documentation exercise.

The same standard has to apply to one's own work, and applying it is uncomfortable. When I reported the empirical validation of my own inference method under these rules, the result came out weaker than my record of it had settled on. A tie correction to complete separation had to be withdrawn in favor of the raw values. An exclusion of a discordant run turned out to have been triggered by the outcome rather than by a validity screen applied blind, which widens the reported interval considerably. The null construction used could support only the claim that the model responds to interaction structure, not the claim that it identifies the correct one. Those three defects are among the most common in computational biology reporting, and they appeared together, in work arguing for rigor, by someone who knew all three principles. That is the case for structural governance rather than individual care. The principles were not enforced at a gate.

The Adoption Problem

Voluntary adoption of any of this will fail, and it is worth being clear about why, because the reason is not insufficient argument. The costs of this discipline are local, immediate, and visible. The benefits are delayed and accrue partly to other parties. Each practice makes a program's weaknesses visible on the timescale of the person who would have to adopt it, while the benefits arrive on the timescale of a successor. That is a collective action structure, and better arguments do not change collective action structures.

The pressure has to come from the parties that capture the benefit. Funders, who capture the return from not funding irreproducible work and who have exact precedent in clinical trial registration. Regulators, who need only decline to accept a computational claim that arrives without its record. Acquirers, who bear the full cost of claims that fail diligence and who can price the difference without anyone's cooperation. This has happened before, in a neighboring field, within living memory. Clinical research now produces the most trusted evidence in medicine, not because its practitioners are more careful than anyone else, but because its claims are checkable.

Meanwhile, the smallest useful version of all of this fits in three sentences lodged before an analysis is run: what would change my mind, how many analyses of this question I plan to perform, and what I will do if the condition fires. If only one practice can be adopted, state the denominator. Nearly everything else follows from being unable to hide it.

Going Deeper

The full argument, the quantitative treatment of constraint inference and governed search, the validation reported under its own rules with the withdrawn statistic and the disclosed exclusion intact, the governance checklists, the mathematical reference, and the two-page record proposed as a minimum standard for computational discovery claims are all set out in From Deterministic Convergence to Discovery: Designing the Next Generation of Therapeutics. The book names no software and requires none. Every chapter closes with a summary that states its claims compactly, and the twelve summaries read in sequence take about twenty minutes and give the whole argument. It is available here: https://a.co/d/0hBqxmWm.

MH

Matt Hardy

Published on September 4, 2026

Drug discovery has multiplied its computational capability by orders of magnitude and has not improved its success rate. That sentence is not a complaint about the difficulty of biology. It is a measurement, and it carries a conclusion that most of the field has not accepted.