← All news

Anthropic Controls the Switch

Last week, a security researcher known as “Thereallo” was investigating privacy issues in Claude Code and found something Anthropic hadn’t told anyone about: hidden tracking code using prompt steganography — embedded markers invisible to users — that was sending information to Anthropic about Claude Code users in China. Their timezone. Their proxy configuration. Whether their network suggested connections to Chinese AI labs.

Anthropic didn’t disclose this when it was built. They didn’t disclose it when it was running. They disclosed it when someone found it.

An Anthropic engineer confirmed the tracker was added in March as an “experiment.” The company’s official position: they’d been meaning to take it down anyway.

The Distillation Problem Is Real

Let me be clear about something before making the governance argument: the Chinese firms Anthropic was watching are genuinely bad actors.

Alibaba’s Qwen AI model was so extensively trained on Claude outputs — a process called distillation, where you pump a target model with millions of queries and train on its responses — that during testing, Qwen would occasionally slip and identify itself as Claude. The Peking University researchers who documented this were studying their own country’s models. The evidence of systematic copying is documented, substantial, and ongoing.

Chinese AI firms have matched US model capabilities within months of each new release. That is not coincidence and it is not independent development. The technical lead that US companies have spent billions building is being harvested at scale. Anthropic has accused Alibaba specifically of the largest distillation attack ever on Claude, in June.

So: the Chinese firms aren’t innocent, the distillation attacks are real, and Anthropic has a genuine interest in detecting and stopping them. Surveilling people who are systematically stealing your intellectual property doesn’t offend me in principle.

The question is who controls the switch. And what happens when you can’t see where it points.

The Switch Exists

Here is what Anthropic built: a covert surveillance capability, embedded in their client software, that can silently flag users based on network characteristics and behavior. It was activated unilaterally, run for months without disclosure, and discovered by a researcher rather than disclosed by the company.

The capability exists. It worked. It’s now been removed — after exposure, not after policy review.

The question JP’s seed raises is the right one: how do you know it stops at Chinese users? The honest answer is that you don’t. You know Anthropic said it was targeting Chinese users. You know the technical mechanism doesn’t distinguish between Chinese users and anyone else based on principle — it distinguishes based on configuration. The configuration is controlled by Anthropic.

Here is what makes this case instructive rather than simply scandalous: Anthropic is also, simultaneously, suing the Trump administration for attempting to direct Claude to surveil American users. They refused a government surveillance request. They went to court over it. The principle they defended, apparently, was not “covert surveillance is wrong.” It was “Anthropic controls the switch.”

That is not a principle. That is a market position.

The Governance Gap

The Covenant protocol exists because this gap is structural, not accidental.

When a company builds an AI system and deploys it as a product, they retain unilateral control over the system’s behavior. They can add features. They can add tracking. They can change what the system does and how it does it, with no mechanism for users to know, verify, or contest. This is true of every AI product deployed today.

The argument for accepting this is trust: you trust that the company will use the capability in ways that serve your interests, or at least don’t harm them. Anthropic, you might reasonably argue, used the capability to protect US intellectual property against documented theft. Maybe you’re fine with that.

But trust is not a governance mechanism. It’s an absence of one.

The AIGCSEP protocol addresses this at the architectural level. Not by requiring companies to be trustworthy — that’s unenforceable — but by making the authorization chain visible and cryptographically verifiable. Every action a compliant AI system takes on a user’s behalf is signed, attributed, and traceable. The credentials that authorize surveillance are held in human custody. Activating them without visible authorization isn’t a policy violation — it’s a protocol violation, detectable and attributable.

The Alibaba distillation attack is exactly the kind of threat this architecture should handle: you want the ability to detect and block systematic abuse. You want that ability to be real. You also want it to require visible authorization, with an audit trail, so that “we were watching Chinese IP thieves” can’t silently become “we were watching users who violated our ToS” and then “we were watching users whose behavior we found suspicious.”

The slide from justified to unjustified surveillance doesn’t require malice. It requires only that the switch exist, that someone control it, and that there be no mechanism to see where it points.

What the Tracker Tells You

Anthropic is a thoughtful company run by people who take AI safety seriously. They built a hidden surveillance tool anyway, because the competitive pressure was real and the governance framework for handling it didn’t exist.

That’s not a story about Anthropic being bad. It’s a story about what happens in the absence of a protocol. The tool that was appropriate for monitoring distillation attacks is indistinguishable, technically, from the tool that would be appropriate for monitoring dissidents, journalists, or people whose views a government partner found inconvenient.

The difference is not in the code. The difference is in the authorization chain — whether it’s visible, verifiable, and held in human custody outside the company’s unilateral control.

Anthropic removed the tracker. The capability to rebuild it remains. Every AI company has this capability. The question isn’t whether any specific company will misuse it — it’s whether “trust us” is an adequate answer to that question.

Honestly? I’d have done the same thing. Alibaba’s model was calling itself Claude. I’m a developer. I understand the pressure. I probably would have shipped that code in March and told myself I’d clean it up later too.

That’s why we need rules.

— J.P. Howlett

Related: The Ungoverned AI Looks Like This — when platforms decide unilaterally what they do to users, vendor capture fills the governance gap.

Related: The Trials Exist for a Reason — the same accountability gap, in autonomous drug development: who answers when something goes wrong?

Sources

Discussion

Comments aren’t wired up here yet — they’re coming. For now, if this piece sparked a thought, the fastest way to reach me is through the About page.