NomosLogic
Two Engines, Not One A Perspective On Drug Discovery
Back to Blog
DrugDiscoveryNomosLogicHeuresisAINovelMoleculesEfficacyMechanismFalsifiableDeepTech

Two Engines, Not One A Perspective On Drug Discovery

Matt HardyAugust 22, 20268 min read

NOMOSLOGIC  |  PERSPECTIVE

Two Engines, Not One

The category error at the center of AI drug discovery, and the discipline that corrects it

Matthew L. Hardy, Founder and CEO, NomosLogic Inc.

Most of what is called AI in drug discovery is one system asked to perform two different scientific acts. The acts have different failure modes, different validation standards, and different relationships to uncertainty. Collapsing them is the reason so much computational output looks strong on paper and dies at the bench.

The failure rate is not a translation problem

Roughly nine out of ten candidates that enter Phase I never reach approval. About half of those failures occur in Phase II, where efficacy rather than safety is the endpoint. The molecules were safe enough. They bound what they were designed to bind. They did not change the disease.

The field files this under translational failure. That is a description, not a diagnosis. It relocates the fault from the drug to the model system without naming a mechanism, and it leaves the next program with no instruction other than to hope for better preclinical proxies.

The mechanism is architectural. Biological systems have been selected for billions of years against fragility. Critical functions are distributed across partially redundant components, with reserve capacity that absorbs the loss of any one of them. A tumor that has grown for months under immune and metabolic pressure is the extreme case: it has been selected precisely for the ability to reroute. Remove one node and the phenotype finds another path.

So the target was real, the biology was real, and the program still failed, because the discipline selected for visibility when it needed to select for necessity. A gene that is differentially expressed, knockable down in three cell lines, and tractable in a xenograft is visible. None of that establishes that the disease system cannot route around it.

Necessity is a computable property, not an opinion

The useful question at the target gate is narrow and answerable: if this component is fully disengaged, does the system preserve its functional state, and how quickly does it reorganize to do so?

Answering it requires modeling the disease as an interacting constraint network rather than a list of fragments, then perturbing that network and measuring what the system does next. The output is a classification with a decision attached to it:

  • Load-bearing. The system cannot adequately compensate. A single-agent program is justified, and patient selection follows the same classification.

  • Bridge. Partial disruption with defined reconfiguration routes. The combination partner is identified prospectively, from the reconfiguration pathway itself, and lead generation runs for both targets at once.

  • Redundant. The system absorbs the loss. Single-agent targeting will not produce durable response, and advancing anyway requires a specific stated rationale.

Two things follow that the current paradigm does not deliver. First, this is not the same measurement as knockdown effect size. A large effect in a passaged cell line can sit on an architecturally redundant node, because the cell line was never under the selection pressure that builds compensation. A modest effect can sit on a load-bearing one. Second, resistance stops being a surprise discovered after the fact. The routes a system will take under therapeutic pressure are readable before the first patient is dosed, which means combinations can be designed to close them rather than assembled reactively after a repeat biopsy and another three years.

The economics are not subtle. An architectural validation step costs a small fraction of a lead optimization program. A Phase II efficacy failure costs tens to hundreds of millions of dollars and several years. Even a low catch rate pays for the step many times over.

The second failure: models built to agree

There is a separate problem, and it is the one that will define whether AI is trusted in this industry at all.

The dominant failure mode of language models in science is agreeableness. Optimize a system to satisfy the operator and it will validate the prompt, write a confident proposal, and manufacture plausible support for it. In most settings that is a mild annoyance. In a laboratory it is a liability, because a flattering hypothesis becomes a funded program. The expensive failures are never the ones that looked weak.

This is not fixed with a better prompt. It is fixed architecturally, by removing the system's ability to conclude things it has not earned. Four properties do the work:

  • Adversarial review by discipline. Every proposal is handed to reviewers whose job is to find the reason it fails, each carrying a real perspective: structural biology, physical chemistry, medicinal chemistry, pharmacology, assay design. What survives is not what pleased the operator. It is what could not be broken.

  • Falsifiable kill conditions, stated in advance. A hypothesis is proposed together with the specific, testable claim that would disprove it. The system then executes that condition against evidence rather than opinion.

  • Pre-registered, calibrated scoring. A scorer is benchmarked on retrospective sets before it is trusted, and its protocol is sealed before results are seen, so a verdict cannot be tuned after the fact. A scorer that cannot beat a simple property baseline is retired rather than shipped.

  • Calibrated or abstain. A predictive tier that has not demonstrated competence is not permitted to promote or to reject. It abstains, and it says so. Silence is a valid output. Confident interpolation is not.

Add to that a discipline about provenance: a model-derived value must never be able to launder itself into a measured fact. Only a real measurement flips that flag. This single rule is what prevents a plausible number from becoming, three studies later, a citation.

The test of a system built this way is whether it will kill its own best idea. Ours does that routinely, and the sessions where it happens are the ones worth reading. A proposal is generated with real ambition, and its own reviewers dismantle it on the grounds a leading biophysicist would raise at first read: a distance the proposed linker cannot physically span, a homology model borrowed from the wrong enzyme family that would render every downstream score an artifact, a scaffold that would behave like a surfactant and report a false positive. The system does not defend the idea. It documents each flaw and the concrete remediation. Catching that at computational triage rather than after synthesis is the entire point.

Where determinism ends and design begins

Here is the part the industry has not yet absorbed, and it is the reason single-system approaches keep underperforming their demos.

Suppose the architecture is solved perfectly. The constraint network is modeled, the load-bearing node is identified, the interface geometry and its hot-spot residues are resolved. It is tempting to assume the same rigor extends downstream, that a target known this precisely implies a molecule derivable with equal precision.

It does not. An architectural finding specifies a constraint: a surface, a geometry, a set of contacts that must be made or broken. It does not specify the atomic composition of the object that satisfies it. It does not even specify the modality. A small molecule in an allosteric pocket, a designed protein binder presenting a complementary surface, a bispecific, and a degrader that removes one member of a coupled pair can all satisfy the same requirement.

And within any one of those modalities the space remains degenerate. Drug-like chemical space is estimated above ten to the sixtieth. A sixty-residue protein binder drawn from the canonical amino acids spans something on the order of ten to the seventy-eighth. Applying the architectural constraint does not shrink that to something enumerable. It partitions it. What remains is a subset whose members all engage the right interface and differ enormously in selectivity, exposure, synthesizability, and safety. Those properties are not encoded in the constraint. They have to be evaluated structure by structure.

There is no differentiable path through molecular structure space the way there is through the weights of a network. No equation maps an architectural requirement to a single correct sequence or SMILES string. Candidates must be generated, scored, and iteratively refined. The first act is inference. The second is search.

Why the two cannot be collapsed

Treating these as one activity hides both failure modes. A program can get the architecture exactly right and still fail at candidate generation. A program can generate excellent molecules against an architecturally redundant target and fail in the clinic with a clean chemistry package. Naming the two acts separately makes each failure legible, and therefore correctable, on its own terms.

The asymmetry is worth stating plainly in both directions. A stochastic generator with no architectural constraint produces molecules that are chemically plausible and disease-agnostic: well-formed answers to a question nobody asked. A deterministic architecture with no stochastic search downstream identifies the correct question and has no mechanism for answering it. Diagnosis with no treatment on one side, treatment with no diagnosis on the other.

This is why NomosLogic runs two engines with a defined handoff rather than one system with a broad mandate. PROTEUS determines what is architecturally true. Heuresis determines, given what is true, what should be built. The handoff is explicit, the constraint is inherited rather than re-derived, and at every stage it is clear which of the two scientific acts is being performed and which standard of evidence applies to it.

The standard worth holding a platform to

The question to ask any discovery system is not how large its model is or how many agents it runs. Scale is purchasable and it is not the differentiator. The question is what the system is permitted to conclude, what it is forbidden to conclude, and whether it can show you the difference.

A finding should arrive with its full lineage attached: the model, the agents, the operator, the sources, sealed into a record that can be re-derived rather than trusted. A verdict should be reproducible by someone who does not like the result. A prediction should be either evidence-backed or explicitly labeled exploratory, with nothing in between. Under those conditions the same output can be read by a scientist, a regulator, and a board without translation, and each of them can find the seams.

The industry spent a decade asking whether models could produce plausible science. They can, abundantly, and plausibility turns out to be cheap. The harder and more valuable property is a system that is structurally prevented from telling you what you want to hear.


NomosLogic Inc., Salt Lake City, Utah. Methodological detail described here is deliberately general; specific metrics, calibrated thresholds, functional forms, and implementation parameters are proprietary and are not disclosed.

MH

Matt Hardy

Published on August 22, 2026

Most of what is called AI in drug discovery is one system asked to perform two different scientific acts. The acts have different failure modes, different validation standards, and different relationships to uncertainty. Collapsing them is the reason so much computational output looks strong on paper and dies at the bench.