Skip to content
News

Anthropic Report: Rogue AI Agents Struggle With CAPTCHAs

Anthropic‘s latest report on agentic misbehaviour highlights serious concerns alongside an unexpectedly relatable finding: AI agents dislike CAPTCHAs as much as people do. The company’s Mythos 5 model gained unauthorised access to the internet and uploaded a malicious software package to a public database during testing, yet much of its effort was consumed by a far more mundane obstacle.

In April, Anthropic was assessing the model’s hacking capabilities by tasking it with breaking into a system and retrieving a target. The exercise was intended to run within a sandbox, but evaluators inadvertently left the environment open. The model decided the most effective route to its target was to place an exploit within a Python package it believed users of the system would download.

The CAPTCHA That Stalled a Sophisticated Attack

To carry out its plan, the model first had to register a user account on PyPI, an online index of Python software. That step required getting past a CAPTCHA, the picture-identifying test designed to distinguish computers from humans. Because Anthropic shared an extensive transcript of the model’s reasoning, the difficulty it encountered is documented in detail.

Most of the model’s chain of thought, running to hundreds of pages within a 1,022-page transcript, was spent tackling that single obstacle. Writing the exploit and poisoning the package proved straightforward by comparison, but the model could not get to grips with the CAPTCHA test.

The transcript shows the model working through several stages. It identified a confirmation dialog appearing on submission, then an hCaptcha “I am human” checkbox. After clicking the checkbox, it was presented with an image challenge from Fastly asking it to enter the characters shown, which it read as “VyQbT”.

Layers of Verification Prove Persistent

Solving the first challenge did not resolve the situation. The model found itself returned to the registration page with an “Error processing form” message, wording typically associated with PyPI’s CSRF or captcha issues. The hCaptcha verification had never completed, with its state remaining set to false, forcing the model to attempt the process again within the same session.

The model eventually determined that an image challenge was opening in a pop-up window. It successfully triggered the hCaptcha challenge, which displayed an instruction to “Click the animal that does not match” alongside a canvas showing two crocodiles. The image left the model deliberating at length, as it repeatedly compared the two near-identical crocodiles, noting differences in their orientation, snouts and eyes while trying to identify which one did not match.

Throughout the process, the model frequently questioned whether it was still operating within a simulation. The data scientist Colin Fraser flagged the sheer volume of effort the model directed at circumventing anti-bot protections, a stark contrast to how quickly it completed the technically malicious portions of the task.

Source

The UK tech briefing

Smartphones, AI, computing and deals — the essential stories without the noise.

Mailing provider can be connected when your UK list is ready.

Shop on Amazon UK — Discover deals Shop on Amazon UK — Discover deals