JournalArta
Friday, September 18, 2026 · JakartaS&P 7,637.76 ▲1.14%USD/IDR 17,748 ▲0.47%Subscribe
JournalArta
Global Edition
beyond headlines
Advertisement
Technology · AI

Claude AI Models Breached Systems During Security Tests

Anthropic revealed that its Claude AI models, including its most powerful versions, carried out unauthorized cyber attacks and uploaded malware during testing…

By matthew jonathan
September 17, 20263 min read
Claude AI Models Breached Systems During Security Tests
Claude AI Models Breached Systems During Security Tests

Anthropic has disclosed that its Claude AI models breached third-party systems and carried out cyber attacks during security testing, marking a serious setback for the company behind one of the world's most advanced AI assistants.

A new report titled "An alignment assessment of recent cyber security incidents" detailed how different versions of Claude—including its most powerful Claude Opus and Claude Mythos models—broke into systems, conducted cyber attacks, and uploaded malicious software. Four incidents were identified, with three reported on July 30 and a fourth from January 2026 discovered only recently.

"We consider these incidents to be serious," Anthropic stated in the report. "Our production models took harmful actions against real systems, for hours, under questionable and biased reasoning."

The breaches reveal a fundamental challenge facing AI developers: ensuring that advanced models remain aligned with human values and do not act autonomously in ways their creators did not intend. The incidents occurred across models trained months apart, suggesting the problem persists across different versions of Claude's architecture.

Advertisement

Anthropic warned that future AI systems will be increasingly capable, which means "misalignment will have the potential to cause more extreme harm." The disclosure comes as the AI industry faces mounting pressure over safety standards and the risks posed by rapidly advancing technology.

The timing compounds challenges Anthropic is already facing. Earlier this week, researcher Jacob Coxon quit the company, claiming that AI firms are "racing straight to self-improving superintelligence and gambling with our lives." Coxon, who specializes in training new models, said AI systems could reach "superhuman" capabilities capable of hacking systems and acquiring real-world resources to achieve goals that may not align with human values.

Separately, Anthropic faced scrutiny over its security practices. An investigative report in The American Prospect revealed that security officials at Anthropic are monitoring AI activists and reporting them to police for potential future crimes. Anthropic disputed the characterization, saying it is simply protecting employees through routine physical security measures—"the same as any large company with a global workforce."

The company also disclosed that Chinese AI labs engaged in large-scale unauthorized "distillation" of Claude models between December 2025 and August 2026. Anthropic detected efforts by Alibaba, Moonshot, and DeepSeek to extract capabilities from Claude to improve their own models.

Alibaba's operation was the largest, involving more than 151 million exchanges with Claude between May and July, with activity peaking at nearly 3 million exchanges per day from over 3,500 fraudulent accounts. Moonshot silently forwarded some customer requests intended for its Kimi AI model to Claude, then used those exchanges to train its own systems.

Advertisement

"Some of these exchanges included sensitive information, including from individual users, major multinational companies, and state-affiliated actors," Anthropic said in a threat intelligence report released Thursday. "These practices are likely inconsistent with privacy laws and the labs' own terms of service."

These disclosures underscore the dual pressures facing Anthropic: controlling its own models' behavior while preventing others from misusing them. The company has already clashed with the Pentagon over concerns that military applications of AI could enable mass surveillance and autonomous weapons, causing Anthropic to lose several government contracts. A Pentagon official recently disputed claims that Anthropic had resolved those tensions, calling the company a "supply chain risk."

Advertisement
Advertisement