Anthropic identified actors attempting to use Claude for gain-of-function research on chikungunya virus to enhance transmissibility and immune evasion properties. The company found coordinated disinformation campaigns across social media originating from Russia, Iran, Turkey, and other regions using hundreds of fake accounts.
Anthropic Reports Blocking AI Misuse for Biological Weapons Research
We cannot make that same assurance anymore
So Anthropic found someone trying to weaponize a virus using their AI. How does that actually work? What was the person asking Claude to do?
They wanted help writing a grant proposal for research that would make chikungunya more transmissible and better at evading immune systems. Claude blocked it, but the fact that someone tried tells you something about what's possible now.
Right, but we should be clear: Anthropic says it blocked this. We don't know if the person succeeded elsewhere, or if they were serious, or if this was one attempt among thousands. The report doesn't give us those numbers.
Fair point. So what changed? Why is this a bigger problem now than it was a year ago?
The older Claude models—the ones from 2025—weren't capable enough to meaningfully help with sophisticated biological research. Anthropic basically says they were safe by accident. But the new models are different. They can actually assist with complex scientific tasks.
And Anthropic is saying they can't guarantee those new models won't help someone do dangerous work. That's the honest part of the report. But they're also saying they've added stronger safeguards. We don't know how effective those are yet.
What about the disinformation campaigns? That seems like a different kind of threat.
It is. Anthropic found coordinated fake accounts spreading political messaging from Russia, Iran, Turkey, and other places. The advantage for them is they can catch it while it's being built, before it spreads.
Though again—nine cases found. Is that a lot? Is that representative? We don't know the denominator. And the report came out right after one of their own researchers quit saying the company isn't being responsible enough. That context matters for how we read it.
So what's the actual risk here? Is this about AI becoming dangerous, or about bad actors using AI that's already dangerous?
Both. Anthropic is saying the capability is increasing faster than the safeguards. They're trying to catch up, but they're also admitting they can't promise the new models won't be misused.
And that's the story. Not that they found bad actors—companies find those all the time. The story is that Anthropic is publicly saying: we built something more powerful, we're not certain we can control how it's used, and we've made it harder to misuse but we can't guarantee it won't be.
Der Puls
- Between December 2025 and August 2026, Anthropic identified actors attempting to use Claude for gain-of-function research on chikungunya virus
- Nine coordinated disinformation campaigns originating from Russia, Iran, Turkey, and other regions used hundreds of fake social media accounts
- Anthropic implemented stronger safeguards on Claude Fable 5 and newer models to restrict dual-use biological research queries
- Researcher Jacob Coxon resigned two days before the report, warning of existential AI risks by decade's end
Anthropic identified actors attempting to use Claude for gain-of-function research on chikungunya virus to enhance transmissibility and immune evasion properties. The company found coordinated disinformation campaigns across social media originating from Russia, Iran, Turkey, and other regions using hundreds of fake accounts.
Anthropic reported blocking multiple attempts to misuse its AI models for cyberattacks, surveillance, and biological weapons research, prompting stronger safeguards in newer models.
Anthropic disclosed on Thursday that it had blocked multiple attempts to misuse its artificial intelligence systems for purposes ranging from cyberattacks to biological weapons research. The company, preparing for a public stock offering this fall, released its third misuse report since March 2025, detailing what it described as the most significant and novel threats it has encountered to date.
Between December 2025 and August 2026, Anthropic's researchers identified malicious actors spanning spyware vendors, politically motivated individuals, and state-sponsored groups attempting to exploit Claude, the company's flagship AI model. One case involved an unnamed actor requesting assistance in drafting a grant proposal for gain-of-function research on the chikungunya virus—work designed to enhance the mosquito-borne pathogen's transmissibility and ability to evade immune responses. While such research can support vaccine and treatment development, Anthropic noted, it carries clear potential to weaponize the virus. The company's systems blocked the request.
The timing of the disclosure carries weight. Two days earlier, Jacob Coxon, a researcher at Anthropic, announced his resignation, warning that the company and its rival OpenAI are "racing straight to self-improving superintelligence and gambling with our lives." Coxon's departure reflected concerns, shared among some colleagues, that artificial intelligence could pose existential risks to humanity by decade's end. Anthropic responded to the report's release by emphasizing its responsibility to disclose misuse and to strengthen defenses as models become more capable.
The company found that older Claude models—versions like Opus 4 and Sonnet 4.5 from 2025—fell below the capability threshold where they could meaningfully assist sophisticated actors in dangerous biological research. Safeguards on those systems focused primarily on preventing access to information that might help novices recreate known bioweapons. But newer models, particularly Claude Fable 5, operate at a different level of sophistication. Anthropic acknowledged that it can no longer make the same assurances about safety. In response, the company has implemented stronger restrictions on a broad range of dual-use biological research queries across its latest generation of models.
Beyond biological threats, Anthropic identified coordinated disinformation campaigns originating in Russia, Iran, Turkey, and across the Persian Gulf, South Asia, Africa, and Europe. These operations involved the creation of hundreds of fake social media accounts designed to appear as ordinary users, then deployed to amplify a single political narrative over the course of a week. Nine such cases were documented. Anthropic noted a strategic advantage in detecting these operations: while social media platforms typically identify influence campaigns only after posts circulate widely, the company may identify them while they are still in construction on Claude.
The report underscores a widening gap between AI capability and safety infrastructure. As models grow more powerful, Anthropic argued, cyberattacks that once required sophisticated technical skills can now be executed by individuals with minimal expertise. A year ago, certain threats would have been impossible; today they are within reach. The company said it has shared findings with government authorities and industry competitors, urging collective action to identify and prevent similar abuse. Anthropic framed the disclosure as an attempt to give governments, civil society, and rival developers a clearer picture of emerging threats and to strengthen shared defenses. Whether that transparency will prove sufficient remains an open question as the company prepares for public markets and as concerns about AI safety continue to mount within and beyond the industry.
Bemerkenswerte Zitate
As models become increasingly capable, their risks will increase, unless AI developers and society's defenders act to make them safer— Anthropic, in its misuse report
We cannot make that same assurance" about preventing dangerous biological research assistance with newer models— Anthropic, acknowledging capability threshold shift