← All news

The Trials Exist for a Reason

Anthropic announced Claude Science this week — a new product designed to support scientific research the way Claude Code supports software engineering. Autonomous, capable of acting on high-level instructions, equipped with tools for computational biology and drug development. They announced it at an event for pharmaceutical executives and biotech founders. Anthropic will also use it on their own drug research for rare and neglected diseases.

This is genuinely impressive. It is also, in the specific domain of pharma, an occasion to say the quiet part loud.

Like no one has ever died because of a human pharma decision.

What the Testing Framework Actually Is

The FDA’s clinical trial requirement — Phase I, Phase II, Phase III, years of data, thousands of patients, documented outcomes — is not bureaucratic overhead. It is not a legacy of excessive caution by regulators who don’t understand innovation. It is scar tissue.

Thalidomide was a sedative prescribed to pregnant women in the late 1950s. It caused severe birth defects in thousands of children across Europe and Australia. The US was largely spared because one FDA reviewer, Frances Kelsey, refused to approve it without adequate safety data despite significant pressure from the manufacturer. The 1962 Kefauver Harris Amendment — the law that created modern drug approval standards — was the direct response.

Vioxx was a painkiller approved by the FDA in 1999. Merck knew by 2000 that it significantly increased cardiac risk. They kept selling it until 2004. An estimated 88,000 Americans had heart attacks attributable to the drug. About 38,000 of them died.

The opioid crisis killed over 500,000 Americans in the twenty years following OxyContin’s approval — approved on the basis of a single, flawed study, marketed as non-addictive, pushed by sales representatives with financial incentives to maximize prescriptions. The executives who made those decisions have mostly settled civil cases and remained free.

Fen-Phen. Baycol. Rezulin. The graveyard of approved drugs that turned out to kill people at scale is substantial. Every drug in it was developed and approved by human beings with advanced degrees, good intentions, and professional accountability structures that nonetheless failed to prevent the harm.

This is not an argument against Claude Science. It is the argument for why the testing framework exists and why “autonomous AI for drug development” requires the same framework, applied with the same rigor, regardless of how the recommendations are generated.

The Accountability Gap

Here is the question Anthropic’s announcement does not answer: when Claude Science identifies a promising compound, and that compound moves through development, and it eventually causes harm at scale, what is the accountability chain?

With a human researcher, the chain is at least theoretically traceable. The scientist who recommended the compound. The executive who approved the trial design. The company that pushed the FDA submission. The regulator who signed off. It is often insufficient — the opioid crisis demonstrated comprehensively that the chain can fail at every link — but it exists. There are names. There are decisions. There are people who can be held responsible even if they rarely are.

With an autonomous system operating on high-level instructions, the chain is undefined. The model made a recommendation. The researcher accepted it. The company filed the application. When something goes wrong, does liability attach to the model’s output? To the instruction that generated it? To the company that deployed the model? To Anthropic for building it? The legal framework does not currently answer this question, and the absence of an answer is not a technical problem — it is a governance problem.

This is precisely the argument at the center of the Covenant: governance has to be in the architecture before the architecture causes harm, not bolted on afterward. We learned this with drugs the hard way — thalidomide first, framework second. We should not need to learn it again with autonomous AI systems doing drug development.

What Claude Science Can Actually Be

The optimistic version of this, which Anthropic is clearly reaching for, is real: AI that can process the literature, model compounds, identify candidates, and run computational trials at a speed no human team can match. If it works, it could genuinely accelerate the development of drugs for rare and neglected diseases that the market currently underfunds because the commercial return doesn’t justify the research cost. Anthropic’s stated intention to use it for exactly that class of disease is the right instinct.

The tools that make it powerful — autonomous operation, broad access to literature, computational biology capabilities — are exactly what’s needed to close the gap between promising compounds and funded trials for diseases that affect small, poor populations.

There’s a precedent worth noting: technology has already made animal testing largely obsolete, and that’s an unambiguous win. Organoids, organ-on-a-chip systems, computational modeling — these methods are more accurate than rodent trials for predicting human response, more ethical, and faster. The pre-clinical phase that used to require years of animal studies is collapsing. Claude Science operating in this space isn’t replacing safety — it’s replacing a testing method that was always a crude approximation of human biology. The framework evolved because better tools existed. That’s how it’s supposed to work.

The distinction that matters: pre-clinical and clinical are different problems. AI replacing animal testing in pre-clinical development is progress. AI replacing clinical trials — the phase where you find out what actually happens to humans at scale — is a different proposal entirely, and one the accountability question makes necessary to resist.

Claude Science can make the trials faster. It cannot make them unnecessary. It can improve the quality of the hypotheses entering clinical testing. It cannot replace the testing. And the accountability framework for the decisions it influences has to exist and be legible before it makes a recommendation that ends up harming someone at scale.

The testing framework wasn’t invented to slow things down. It was invented because people died. The lesson applies regardless of who — or what — is doing the science.

— J.P. Howlett

Related: Insource the Model — when institutions hand critical decisions to external AI systems without building the governance layer, they lose both capability and accountability.

Related: Two Chinas, Two Americas — the same pattern: powerful technology, undefined accountability, and the people who pay the price aren’t the ones making the decisions.

Sources

Discussion

Comments aren’t wired up here yet — they’re coming. For now, if this piece sparked a thought, the fastest way to reach me is through the About page.