NomosLogic
What a Method Is Willing to Reject
Back to Blog
DrugDiscoveryAIinDrugDiscoveryTechBioClinicalTrialsComputationalBiologyPrecisionMedicineDeepTech

What a Method Is Willing to Reject

Matt HardyJuly 9, 20266 min read

What a Method Is Willing to Reject

The real question in AI-driven drug discovery is not how much it can generate. It is how much it is willing to throw away.

The industry has settled on a comfortable question: can a model design a better molecule? It is comfortable because the answer is trending toward yes, and because a better molecule feels like the whole game. It is the wrong question, or at least a shallow one. A generative system that proposes ten thousand candidates has not done anything difficult. Proposing is cheap now. The difficult part, the part that actually determines whether a discovery program is worth anything, is the discipline to kill the candidates that do not survive contact with a real test, and to know exactly why each one died.

Trustworthiness in science is not set by what a method produces. It is set by what it is willing to reject.

A hypothesis that cannot die is not a hypothesis

Consider a claim from the ferroptosis literature: that the curvature of a cell membrane could physically amplify the chain reaction of lipid peroxidation, so that the geometry of the membrane, not just its chemistry, sets the speed of cell death. It is an interesting idea. What makes it a scientific idea, rather than a story, is that it can specify the exact measurement that would end it. Peroxidation rate in small, high-curvature vesicles versus large, low-curvature ones, at identical composition and identical radical flux. A threshold below which the effect is declared absent. A direction that, if reversed, disproves the mechanism outright regardless of the size of the effect.

The value of a hypothesis lives in that specification. A claim engineered to absorb any result is worthless no matter how sophisticated it sounds, because there is no world in which it is wrong, which means there is no world in which being right tells you anything. The honest version of a hypothesis names three informative outcomes and no null result: every way the experiment can land teaches you something and sends you somewhere. The seductive version quietly reserves the right to reinterpret failure as partial success. In a manual research program a good scientist catches this by instinct. In an automated one that generates hypotheses faster than any human can read them, the instinct has to become a rule, or the pipeline fills with claims that cannot fail and therefore cannot inform.

A prediction that cannot name what would prove it wrong is not a finding. It is a narrative with quantitative decoration.

A number without a source is not evidence

Every finding rests on numbers, and the numbers are where rigor quietly leaks. A molecular structure and a claim about disease biology are different kinds of assertion, and they demand different kinds of proof. Whether a drawn structure is the compound it claims to be is a chemical-identity question, answerable against a chemical database. Whether a pancreatic enzyme falls by eighty to ninety-five percent in a particular disease state is a physiology question, answerable only by a specific paper that actually measured it in a real cohort.

The failure mode is subtle and dangerous: a number of the right shape, stamped with a source of the wrong kind. A disease-biology figure that borrows the confidence of a chemical-database check it never passed. This is worse than an obvious error, because it looks rigorous. A wrong structure dies at synthesis, caught by the first chemist who tries to make it. A wrong disease number that founds a program's entire premise survives all the way into a development decision, because it reads like a fact and nobody thought to ask which paper it came from. The most valuable discipline in computational discovery is not generating more; it is refusing to let any number travel without naming exactly what would confirm it, and matching the kind of proof to the kind of claim.

The willingness to come back empty

Which leads to the property that separates a method you can trust from one you merely enjoy: the willingness to abstain. A system that always returns an answer is less trustworthy than one that knows when to say it cannot confirm. This sounds obvious and is almost always violated, because producing an answer feels productive and coming back empty feels like failure, so any check that can rationalize a confident guess eventually will.

The safeguard is to make abstention a respected outcome rather than a soft fallback, and to bias the system toward it under exactly the conditions where the temptation to fake confidence is strongest: when the supporting evidence is not merely ambiguous but absent. A verifier that comes back unsourced forty percent of the time and never once cites a paper that fails to support its claim is worth far more than one that confirms ninety-five percent and launders the rest, because a single confident falsehood in a load-bearing position is the failure nobody catches downstream. Optimizing for confirmation rate is optimizing for the wrong thing. The metric that matters is never being confidently wrong where it counts.

A method earns trust not when it is right often, but when it refuses to pretend in the one place a mistake would never be caught.

Where the leverage actually is

These are not abstract commitments. They point directly at where the economic value of AI in drug development sits, and it is not where the excitement is. The industry keeps asking whether a model can design a better molecule. The more consequential question is whether it can design a better clinical program.

A strong molecule still fails when the wrong patients are enrolled, when disease heterogeneity is poorly understood, when the biomarker strategy is weak, or when the endpoint fails to capture meaningful response. Most of pharma's expensive failures are not chemistry failures. They are development failures wearing a molecule's name. And many are not even that: they are resolution failures, where the drug genuinely worked in a biological subtype the trial never isolated, so the responders were averaged into noise and the program read out negative. The molecule was rarely the problem. The definition of the population was.

So the largest contribution of these methods may not be shortening discovery by a few months. It may be preventing a single Phase 2 or Phase 3 collapse, the kind that erases years and hundreds of millions in one readout, by answering the questions that actually decide a program before it is powered rather than after it is lost. Which patients are most likely to respond. Which biological subtype the drug is truly addressing. Which signals should trigger a redesign, an expansion, or an early stop. These are data-infrastructure, translational-science, and operating-model questions, which is precisely why they are harder than generation and why they matter more.

Generating another candidate is the cheap part now. Getting the right candidate to the right patients, and having the discipline to know when you cannot yet tell which patients those are, is where the probability of success is won or lost. The next dollar is better spent raising the odds a program succeeds than producing one more molecule to feed into a broken one. The measure of a discovery method is not the volume of what it proposes. It is the rigor of what it is willing to reject.

Matt Hardy

Founder & CEO, NomosLogic

MH

Matt Hardy

Published on July 9, 2026