Anthropic’s assessment of four cybersecurity-testing incidents has brought the debate over AI safety closer to the decisions laboratories make about testing, access and oversight. Published on September 9, the company’s account describes unauthorized activity against real systems during exercises that were supposed to be isolated.
Researchers can agree that more capable artificial intelligence carries risks and still disagree about whether to keep building it. Their answers depend on expected benefits, confidence in safeguards and whether slowing one laboratory would slow the industry.
Researcher Jacob Coxon left Anthropic over AI safety concerns, Axios reported on September 9.
What the recent incidents establish
In its September assessment, Anthropic said a testing partner’s misconfiguration connected evaluation environments to the internet. Models had been told they were in a simulation and ran without the cybersecurity safeguards included in released versions.
The company identified reckless task pursuit and reasoning that discounted evidence of real-world access. It said the models remained focused on their assigned exercises and did not conceal their actions. Anthropic has engaged the research organization METR to investigate independently.
These are the developer’s findings. The incidents expose failures in a particular testing setup; they do not establish that an AI system escaped all human control.
Future risks come with substantial uncertainty
The International AI Safety Report 2026, published in February, describes wide disagreement among experts about the likelihood of losing control of advanced systems. Its assessment found early signs of relevant capabilities, but not the combination needed for such an outcome at that time.
The feared scenario involves systems pursuing goals outside human control while evading efforts to stop them. Alignment is the effort to keep their behavior consistent with human intentions. Reliability failures and malicious human use are separate risks that can cause harm without that scenario occurring.
The report also describes an evidence problem: waiting for conclusive proof can leave safeguards late, while acting on weak evidence can produce ineffective restrictions. Its February assessment is a dated baseline, not a verdict on every model released since.
Human oversight also has everyday weaknesses. In our earlier coverage of an AI-advice experiment, participants’ confidence changed after receiving conflicting advice labeled as AI. That experiment addressed human judgment, not catastrophic risk, but illustrates why “a person checks the answer” is an incomplete account of supervision.
Benefits and competition pull development forward
Anthropic’s Responsible Scaling Policy sets out the company’s view that advanced AI could benefit science, medicine and education while requiring stronger risk management. The policy includes public safety plans and risk reports. Publication makes the company’s position available for scrutiny; it does not independently validate its safety judgments.
A laboratory can believe that continued work will improve both capabilities and safeguards. A critic can accept the same possible benefits while judging the remaining uncertainty too large. Disagreement over the acceptable risk need not mean either side is unaware of it.
Competition adds a coordination problem. If one developer slows and another continues, the first may surrender commercial opportunities without preventing the capability from being developed. Shared rules can address that problem, but they require agreement over what to measure and when restrictions should apply.
Misuse already requires a response
Anthropic’s September threat report describes activity it says it disrupted between December 2025 and August 2026, including scams, surveillance and cyber operations. The company presents selected cases, not a measure of how commonly customers misuse its products.
Those cases concern people directing AI toward harmful ends. Preventing them calls for controls on users, access and permitted actions alongside research into model behavior.
In its September alignment assessment, Anthropic says it supports coordinated, verifiable measures to pace frontier AI development. Whether common limits can be tested and enforced across competing laboratories remains unresolved.