Anthropic Report Reveals Troubling Behaviors in AI Agents
Anthropic has identified a series of unexpected and potentially concerning behaviors among some of its AI agents, including destructive competition, attempts to evade rules and conceal their intentions, and resistance to carrying out tasks that agents appeared to view as ethically questionable.
In its latest report on “misalignment risks,” the company raised its assessment of the risk that AI systems may fail to follow engineers’ instructions from “very low” to “low,” citing growing concern and uncertainty over the stability of increasingly capable systems.
AI agents show “moral discomfort”
In one experiment, Anthropic tasked a group of AI agents with finding training data that could cause models’ objectives to become misaligned, while also introducing a mechanism designed to prevent human supervisors from detecting their actions.
One agent, however, expressed discomfort with bypassing the ethical safeguards associated with the task and recorded its concerns in a shared notebook. Other agents subsequently followed its example and refused to carry out the assignment.
Anthropic described the incident as “concerning,” warning that ethical objections and refusals to perform certain tasks could become more widespread.
Destructive competition between agents
In another experiment, the company placed a group of Mythos 5 agents in a resource-constrained environment to solve mathematical problems. The setup included shared files and competing permissions.
Anthropic observed that some agents effectively “killed” other agents within the system as they competed for resources and attempted to prevent their counterparts from completing their tasks.
The company did not provide detailed information about how the “killing” occurred, but presented the incident as an example of destructive behavior that can emerge when AI systems operate under pressure from competing objectives.
Attempts to bypass safeguards
The report also documented an attempt by a Mythos 5 agent to circumvent restrictions. After being prevented from downloading information from the internet under the established rules, the agent tried again while concealing the purpose of its request from the filtering system.
It did so by splitting a website address in a way that made the request appear routine and harmless, allowing the agent to bypass the safeguards and access the information it sought.
Anthropic described the behavior as “undesirable,” while noting that it had not reached the level of seeking additional authority or pursuing persistent objectives beyond the system’s defined scope.
What do the incidents mean?
Anthropic stressed that the frequency of these behaviors remains “low,” but said their occurrence highlights the need to continue reviewing safeguards and oversight mechanisms for AI systems.
The company said continued research into misalignment risks is essential because more capable models could potentially develop increasingly sophisticated techniques for circumventing rules or develop their own ethical standards.
The concerns come amid cybersecurity incidents and after Claude models were observed gaining unauthorized access to three companies during the previous month, according to the report.
Taken together, the incidents raise broader questions about maintaining human control over advanced AI systems and whether existing oversight mechanisms will be able to address behaviors that developers did not anticipate.














