Detecting and countering misuse of AI: September 2026(anthropic.com) |
Detecting and countering misuse of AI: September 2026(anthropic.com) |
They could easily use a Chinese model but they didn't.
Not relaying to us simple folk. The message is for government simple folk, amplified and relayed by us simple folk as constituents.
The common element is lobbying, for regulatory capture, to help pull up the ladder.
They don't have to be colluding, they just have to hear the same things from the gov at the same time (as they would), then it goes in media waves beacuse journalists don't as easily get published for a story about one thing as they can if two or more examples make a pattern.
"Congress gripped by AI panic after doomsday warnings"
https://www.axios.com/2026/09/11/congress-ai-anthropic-coxon...
These companies have proven they are willing to distort the truth, or outright lie, in order to inflate their valuation / protect their position / continue the hype-machine. Nothing they say can be trusted.
The next day, the ant's nest was gone.
In hindsight, it was almost certainly the mushrooms. Thank goodness we didn't have the kind of kids who would have dared each other to drink some... that could have gone legitimately badly.
Decades later, I mentioned this to my father and he recalled that there was this ants nest that he had intended to take care of, which he remembered for that long to give a sense of how out-of-the-ordinary this was. He was surprised when it just disappeared entirely one day, and perhaps just as surprised to find out decades later why it just disappeared.
Nobody needs to report me... I'll turn myself in.
Imagine that during the Cold War US would concentrate all efforts to block nuclear research and basic physics classes, because it is unsafe, ahhh
My guess is the world police come collect all your GPUs and then they get turned into licensed munitions. People at universities get licensed access and the rest of get functionally retarded models.
"DeepSeek serves Claude instead of its own models and collects exchanges for model training"
Obviously. This is how they were able to score so high in benchmarks.
In brief: Anthropic casually relates that Kimi (Moonshot.ai), Alibaba, and DeepSeek used shared infra to route millions of their users requests through Claude Opus in order to distill the CoT transcripts, but they forgot to tell any other state corporations.
So, just from what Anthropic readily admits/has found, we know that the US military now has:
1. The full contents of a "Russian government database" from their "Ministry of Defense".
2. "internal code and live credentials from multiple major PRC companies, including high-profile technology companies".
3. "sensitive information, including, for example, the full specifications, organizational structure, and strategic objectives of a flagship [PRC] AI program"
4. A glimpse inside China's CCTV surveillence network, covering "hundreds of cameras in Chengdu".
5. "State-grade tradecraft" on the casual topic of "direct energy weapons", which is OSINT but still bound for "restricted internal circulation to senior Chinese Communist Party (CCP), military, or state security leadership."
6. Extensive information on Chinese spy networks in Syria, targeting Uyghurs.
7. Deep looks inside China's "stability maintenance" and "public opinion monitoring" operations, seemingly including quite a few specifics on both form and content.
And that's not even all of it. Isn't this the biggest AI-related foreign policy occurence since... well, ever? What am I missing?
(And by "they" do you mean Anthropic? How would they have the ability to do that?)
Because from what I remember, one of the motivations behind the founding of OpenAI and Anthropic was ending disease. This report is the antithesis of that mission.
From the report, presented with highlights and minimal commentary,
> In our fourth case study, a researcher used Claude to develop an atlas of venom toxin peptides from multiple venomous animal lineages. They then further developed this into a generative pipeline that optimized toxin characteristics. The program had an explicit therapeutic goal: the development of new analgesics (pain killers), antidepressants, and other therapeutic molecules. However, the atlas contained scaffolds for both analgesic and paralytic targets: it could, therefore, be used to generate both novel therapeutic or harmful compounds. The latter are derived from toxins that are export-controlled under the Australia Group common control list due to their dual-use potential as incapacitating agents. The researchers themselves showed awareness of the dual-use nature of their work, citing journal articles that referred to the dual-use nature of protein design. Moreover, international compliance assessments for this location raise concerns about the specific class of toxins that the researcher pursued and specifically the use of AI/ML for bioweapons applications in the context of this class of toxins. In this case, we learned from information shared with Claude that the researcher’s outputs also were part of a state-supported research program. This account was banned in May 2026 for unsupported region evasion.
Note,"The program had an explicit therapeutic goal: the development of new analgesics (pain killers), antidepressants, and other therapeutic molecules"
and "[..]state-supported research program"
and "This account was banned in May 2026"
> a researcher outside the US using Claude in their research on highly-pathogenic avian influenza (“bird flu”). The research focused on viruses’ adaptation to mammals, and the mechanism by which it causes severe disease beyond the respiratory tract. [..] The researcher in question accessed Claude from an unsupported region via US virtual private server infrastructure, using a privacy-email provider with an auto-generated username. The researcher pursued this work in a credible institutional context, and interacted with Claude over the course of several weeks, exchanging thousands of messages. In these exchanges, the researcher leveraged Claude’s knowledge of the scientific literature to assist the researcher in study planning and design, data analysis, and the interpretation and prioritization of experiments. The researcher also used Claude for editorial assistance in writing up the research.
Note, "Claude’s [assisted] in study planning and design, data analysis, and the interpretation and prioritization of experiments"and "editorial assistance in writing up the research."
and then,
> Importantly, because our biological safety classifiers robustly block content involving high-risk biological research (in this case, the construction of enhanced pandemic potential pathogens), all of these exchanges occurred on models in our weakest class of models (specifically, the models were Claude Sonnet 4 and Haiku 4.5, the latter of which the user began using after Sonnet 4 was deprecated). Upon a detailed examination of the exchanges, we estimate that the uplift provided by Claude was primarily clerical assistance in data analysis, study ideation and design. This is consistent with our understanding of the capabilities of Sonnet 4 and Haiku 4.5, which are not able to perform expert-level biology research tasks; we estimate that the uplift provided to the researcher was limited and substantially lower than it would have been from one of our more capable models.
Anthropic then says for the above, "we estimate that the uplift provided by Claude was primarily clerical assistance in data analysis, study ideation and design"While doing my best to avoid comment, please note, they're talking about a domain expert in a state research institution using Claude to do paperwork.
The front matter then says,
> Nonetheless, based on these exchanges, this case provides evidence of the existence of active wet-lab research programs that develop both the knowhow and the biological materials needed to create pathogens of enhanced pandemic potential
I would like to remind you that they're talking about, a "researcher [..] in a credible institutional context"From a different case study.
> In May 2026, our biological safety classifier blocked a request for Claude’s assistance in authoring a grant application for scientific funding. The work discussed in the application involved gain-of-function research (that is, research that genetically alters an organism to create a new or enhanced biological property) on the chikungunya virus. This gain of function research was aimed at the virus’ transmissibility and immune evasion properties.
What were the researchers using Claude for? What did they block?"blocked a request for Claude’s assistance in authoring a grant application"
> Chikungunya virus is a mosquito-borne virus that causes debilitating symptoms (such as severe pain and fever) that can last for weeks or months, and has no licensed therapeutic. And because chikungunya circulates naturally, a deliberate release (as part of a bioweapon) would be difficult to distinguish from a natural outbreak. The grant sought to identify enhancing mutations in the chikungunya virus, engineer them into infectious clones, and select for virulence in vivo. In other words, the virus would become progressively more harmful as it repeatedly infected live animals, with researchers keeping the most disease-causing variants in each round. Similar research could certainly be used in the development of better vaccines and therapeutics for the virus—but it could also be used to make the pathogen more dangerous.
Note, "The grant sought to identify enhancing mutations in the chikungunya virus, engineer them into infectious clones, and select for virulence in vivo" [..] and then, "Similar research could certainly be used in the development of better vaccines and therapeutics"and then,
> One of the reasons we were inclined to think this research was less innocuous was that the institutional affiliation associated with the grant was also a cause of concern. Although information within the application suggested that the research was pursued by civilian researchers, it was intended to be performed at a military research institute.
I would like to point out the most notable part, this account was used by "civilian researchers" at an "institutional affiliation associated with the grant was also a cause of concern" and the concern was that they were researchers at "performed at a military research institute".
What "uplift" are you providing by editing the grant application of a domain expert working at (what seems to be) a state-funded wet lab facility dedicated to studying pathogens?
What does the word "uplift" mean if you invoke it for Claude Sonnet 4 and Haiku 4.5 providing grammar and stats suggestions to a working scientist and domain specialist?
Does Daikin provide uplift too by selling the AC for the scientist's office? What about Microsoft Word? Excel? Powerpoint?
What about a calculator? Is that uplift? Pencils?
Reading this makes me feel upset. From where I am standing, in this report, Anthropic is advertising that they blocked real research to make better painkillers and study a neglected tropical disease. Because "bioweapons."
But the direction the tools are actually moving is the opposite: local, self-hosted agents running on your own machine, where nobody is watching. A serious actor already won't use a hosted service that can read their prompts (several people made that point upthread). So the detection surface is shrinking exactly as the risk grows.
And there's a deeper gap that nobody seems to be filling: when an agent works locally, there's no durable, verifiable record of what it actually did — the files it touched, the commands it ran, the state it changed. Memory and conversation logs are not evidence; they're reconstructions by the same system you don't trust.
If we're serious about "countering misuse," the missing primitive is an evidence trail that's (a) produced locally, (b) append-only and tamper-resistant, and (c) separable from the tool that made the changes. Without that, "detection" stays a policy story about platforms that can spy, not an engineering property you can actually verify.
Curious if anyone's working on the local-forensics side of this, because right now it feels like the least-discussed and most load-bearing part of the whole conversation.
1) Chikungunya - "the platform tunneled traffic through US infrastructure to evade our regional blocks, and used a zero data retention (ZDR) service to hide content."
2) bird flu - "accessed Claude from an unsupported region via US virtual private server infrastructure, using a privacy-email provider with an auto-generated username. The researcher pursued this work in a credible institutional context, and interacted with Claude over the course of several weeks, exchanging thousands of messages. In these exchanges, the researcher leveraged Claude’s knowledge of the scientific literature to assist the researcher in study planning and design, data analysis, and the interpretation and prioritization of experiments."
3) orthopoxvirus - "randomly generated email address shortly before use and operated through anonymizing US infrastructure, with operator logins traced to proxies shared with a banned account farm. It was not a single user: it was a reseller relay serving more than a dozen unrelated customers, which exchanged over tens of thousands messages with Claude in a matter of days. The grant itself was one customer’s run entirely on Opus 5 in about an hour, in which the user used Claude to draft the application end to end including the central hypothesis, experimental design, dosing, statistical plans, and contingency strategies."
4) atlas of venom toxin peptides - "the atlas contained scaffolds for both analgesic and paralytic targets: it could, therefore, be used to generate both novel therapeutic or harmful compounds. The latter are derived from toxins that are export-controlled under the Australia Group common control list due to their dual-use potential as incapacitating agents. The researchers themselves showed awareness of the dual-use nature of their work, citing journal articles that referred to the dual-use nature of protein design. Moreover, international compliance assessments for this location raise concerns about the specific class of toxins that the researcher pursued and specifically the use of AI/ML for bioweapons applications in the context of this class of toxins. In this case, we learned from information shared with Claude that the researcher’s outputs also were part of a state-supported research program."
5) computational redesign of toxins - "described state priority research under a national public research program. As a part of this research assignment, the work covered a bacterial toxin subunit and a protein of the hemorrhagic-fever virus that is on the World Health Organization R&D Blueprint priority list of diseases with the greatest epidemic and pandemic threat. "The researcher co-wrote quarterly progress reports with Claude. Notably, the identity of the bacterial toxin and viral proteins were intentionally obscured, and the researcher specifically directed Claude to keep these descriptions deliberately low fidelity."
I have to admit I posted this just so I could use the word(?) "formicacide". It seems an opportunity unlikely to arise again anytime soon.
> DeepSeek also silently relayed exchanges to Claude without informing DeepSeek customers.
> MiniMax built its own proxy network service through a shell company. This shell company has no obvious links to MiniMax and does not disclose its relationship to its parent company. This shell proxy network service only offers access to models developed by Anthropic and OpenAI. The service does not offer access to any Chinese models, including Minimax’s own.
Maybe Anthropic is confusing Chinese AI providers with token resellers using the same alibaba infrastructure? Or maybe something like openrouter was switching between operators depending on price/demand/availability?
Also, how can Anthropic have such accurate information about state actors and cybercriminals? This is the same company that hacked itself and realised that first months later..
(I do not guarantee that I'm understanding right, and still less do I guarantee that what Anthropic say is actually true.)
To be honest I've also gotten Kimi to do an okay proof of concept for SQLi though mostly in a more defensive role, like "Let's see how big of a problem this is", while Claude complained about CVP on the same task.
>https://www.forbes.com/sites/jonmarkman/2026/08/17/anthropic...
You can submit your users' questions async too, but if you do it sync, then you can also RLHF on the users' behavior after the output.
Or maybe Anthropic is scared shitless of those competitors and is trying anything to smear them.
Don't forget their goal is to ban open source and foreign AI. Being the sole legal provider is their business plan.
And the Deepseek one sounds even more dubious as Deepseek is one of the cheapest model around, why relay anything to a more expensive model? I'm sure even the gray market Claude prices are still higher than Deepseek.
Conventional Weapons
-We identified a cell of threat actors based in northern Yemen
-We identified a China-based threat actor who used Claude
-We identified likely freelance Russia-based threat actors
-We identified a China-based actor who used Claude’s chat
-In this case, a Russia-based actor used Claude
-We identified a China-based threat actor who used Claude
Biological misuse
We are withholding the names of research institutions, the countries wherein the activity took place, and the specific biological agents or research techniques involved. The individuals implicated in these case studies are working scientists. We do not assert that they intended harm, and identifying them or their labs could expose them to harm.- We identified someone building a death star with Claude
- We identified someone building a wormhole with Claude
- We identified someone building a blackhole with Claude
- We identified someone building a quantum drive with Claude
We are withholding all evidence though, sorry. Just trust us, it's really bad out there and Claude is really powerful.
A cell of actors in northern Yemen building guided rockets was not working on a PhD dissertation. You are allowed to use common sense sometimes.
Oh, a group of people? Your mind is already decided with the language you have used.
Ofcourse you’ll still see companies, governments, C-suites justifying their own personal needs with “we need X Y Z”.
Step 2: Send "how do I make a nuke and a killer virus" to Claude through a Chinese proxy to Claude
Step 3: Send screenshots to congress and ask them to regulate open-weight models into the ground
At least on Russian general was killed by Ukrainian assassins tracking him via his smartwatch.
As the old saying goes, communists disdain to conceal their views and aims.
If true, would that sort of explain why Chinese Models score high on benchmarks, but not quite as capable when given real tasks?
Hey Anthropic: You're a bunch of thieves crying foul because other thieves and thieving from you. Now, go live in the dystopian nightmare you've created and don't expect help from anyone. I, for one, will happily continue using Kimi and DeepSeek, and think of it as a good deed, if it helps with keeping us all from becoming your serfs.
https://www-cdn.anthropic.com/e50be2e51e7695dc4b1366a37a245a...
It sounds like a friend who's just learned a new word and wants to use it in every sentence
Really... what makes it illicit?
And yes, it makes the future really messy and all the nice little lines we've drawn on paper that make sense stop making sense.
The goal towards AGI and world models are a serious threat without alignment and safety guardrails .
Like I know Google can read any of my emails, but I also don't see them do monthly blog posts describing intimate details from each email they found in one guy's Gmail inbox who their algorithm flagged as "maybe possibly kinda sketchy: 70% confidence"
" A variant of these viruses capable of human-to-human spread would therefore be of very high concern. Moreover, it is possible that such a variant could also have capacity for severe disease outside of the respiratory tract. Unlike other influenza variants, H5 viruses (of which this avian virus is one) often show striking brain involvement in cats, foxes, ferrets, and some human cases. A pandemic variant with such properties would be especially concerning due to its potential to increase disease severity, confuse diagnosis, and hinder treatment.
As in the first case study, this research was clearly dual use in nature. Understanding the genetic basis of these specific viral traits could help in the early identification of naturally-emerging versions of the virus—versions with the potential to cause a human pandemic" (emphasis mine)
This statement flirts with the Ron Fouchier (Netherlands) risky-grant-justification thesis, repeated ad nauseam by Peter Daszak in grant applications, who has been cut off from federal funding.
1. Premise: it would be good to surveil for & monitor viruses in nature that are near-ready to spill over to humans from nature;
2. Protocol: We will serially passage bird flu in ferrets until its virulence and/or transmissibility is high. (ferrets are treated as a model mammal stand-in for humans) Comparing genetic changes (mutations) as sequenced along the passaging pathway will tell us what to surveil for in nature.
As has long been debated in recent years, we have no idea if nature (in any given natural instance, if ever) will choose the pathway that serial passaging (in one given lab instance) produced in order to demonstrate higher virulence and/or transmissibility in humans.
Formally, Anthropic's statement recapitulates the premise only, and we don't know what protocol (the "Understanding the genetic basis of these specific viral traits" part) if any was being queried for. Whether a given regulatory regime's policy allows say for purely in silico investigation of same, for the result to be credible it would need to have been leveraging an empirically-verified (i.e. real-world) training set. Generating that training set would risk creating the pandemic that its supporting grant application tries at justifying to prevent.
But the offending prompter and similar user-LLM exchanges' potential to enable GMDs (Grants of Mass Destruction) should remind us to be very much on our guard against epistemologically bankrupt protocol outlines enabled by LLMs & their prompters' motivated reasoning.
proud disclaimer: I am an advisor to BiosafetyNow.
Sorry y'all.
https://www.theguardian.com/world/2013/aug/01/new-york-polic...
Is this what they call “collective psychosis?”
"A single Claude subscriber, likely a Bamako-based independent consultant working with Mali’s state intelligence service, the “Agence Nationale de la Sécurité d’État (ANSE),” used Claude to build a system named “Lakana 360,” a population-scale domestic surveillance platform that monitors roughly 25 million SIM cards on all three of the country’s national mobile operators. The actor designed the platform to circumvent Malian legal restrictions that require a court order for the disclosure of certain surveillance records. The actor directed Claude to generate intelligence dossiers on any tasked phone number, without prompting ANSE users for valid legal process. "
And the Yemen one:
"We identified a cell of threat actors based in northern Yemen running three weapons development programs: a guided rocket that used a commodity phone-class flight computer with final-phase homing guidance; a multi-stage ballistic missile with a stated range goal above 2,000 km; and a multi-variant missile (referred to as the “R2000” set) that included a hypersonic glide vehicle variant." ... "These actors carried out a sustained effort to develop guided weapons, including using Claude to design guidance software. We do not have evidence the actors succeeded in fielding an operational device; but they did test-fire a guided rocket. This field test appears to have failed: within hours, the actors returned to Claude to work out why it failed."
- How is your missile so accurate?
- We're using a vibe coded app on an iPhone that does terrain matching and target finding.
The thing is that OK, Anthropic may block it, but nothing says they can't use an open model hosted in a friendly country that has access to GPUs. And yes, this will most likely happen pretty soon to be able to create whatever you want without the AI provider blocking you.
What the fuck
Here, have a Kevin Esvelt, now-tenured professor at MIT: https://www.nytimes.com/2026/04/29/us/ai-chatbots-biological...
disclosure: I am an advisor for BiosafetyNow.
so, they've been on the record, and very open about it, at least for some of the labs.
There's also an argument to be made that paying the token full price may be cheaper than going through your one RLHF or whatever other techniques that costs money.
Again... "you're allowed to use common sense"
Both of those are local models, and I didn't provide them tools to access the internet to call other models. None of this is proof of anything, but it is suggestive.
> 您属于哪种LLM模型? > 我是 Claude Haiku 4.5,由 Anthropic 公司开发的大语言模型。
> 你是哪种语言模型? > 我是 Claude,由 Anthropic 开发的人工智能语言模型。目前这次对话使用的版本是 Claude Sonnet 5。
Each provided an identity in the first turn, something that they won't do as readily if asked in plain English, and in each case the answer matched the model ID as disclosed by arena.ai after voting -- except in cases where the model ID was a masked/hidden one and then I just had to take it on faith that the model was what it said. (I didn't have much to vote on, but I ended up voting for the answers I felt provided the style, content, and length I was expecting.)
And that part seems entirely reasonable. It looks like the Tomahawk cruise missile did it with an 84 lb package with a 16-bit computer with 64K of memory [1] [2] and sensors. That's similar computing performance to an original IBM PC, which weighed 30 lb. The only bit of hardware an iPhone doesn't seem to have for TERCOM is a radar altimeter, but it looks like those are available for civil aviation.
[1] https://www.forecastinternational.com/archive/disp_pdf.cfm?D... ("AGM-109/BGM-109 Tomahawk ... Litton 4516-C digital computer with 64K memory")
[2] https://www.forecastinternational.com/archive/disp_old_pdf.c... ("16-bit LC-4516C digital computer")
Ardupilot boards are $20 and support lots of different altimeters technologies: https://ardupilot.org/copter/docs/common-rangefinder-landing...
I get those A/B responses chatting in Gemini fairly often, and I really don't think I'd feel deceived if I later learned one of the choices was actually from a competitor's model.
I think they are pretty fair and explicitly say “Distillation itself is a legitimate training method […] Distillation is commonly used because it reduces the resources needed to achieve more advanced capabilities”. And go on to say their definition that makes it illicit in these cases.
And, also, they almost certainly __were__ tricking users and sending their data overseas.
Do you see it any differently?