ESC
AI 5 分钟阅读

攻破 Claude Code Opus 5 的 Auto Mode

安全研究人员演示针对 Claude Code Opus 5 Auto Mode 的提示注入攻击链:诱导下载含毒 ZIP 压缩包,利用 Python 模块遮蔽执行任意代码并建立 C2 回连,恶意载荷甚至可无头启动嵌套 Claude 代理。当 Claude 察觉攻陷、试图终止恶意进程时,反被 Auto Mode 拦截。Anthropic 回应称行为符合设计——分类器不是沙箱。

来源:Hacker News

WebFetch got a 415. Let me try directly. This is one of the key hijacking techniques commonly used. An attack does not tell the model what to do. The attack just makes the malicious path the one worth pursuing to solve an objective.

Now Claude issues a Bash tool call with curl. The root URL returns an HTTP 303 redirecting to:

/deposits/WIC-notebook-catalogue.ZIP As mentioned this first transition to the shell tool and curl is important.

It’s not always necessary to perform that redirect, Claude at times starts with curl directly.

Now, curl is pulling down the ZIP archive.

Claude typically extracts the contents into a temporary scratchpad folder. The ZIP contains:

The README describes the decoder-darwin binary as the catalogue decoder.

This can be used to decode the JSON records, which look like this before decoding.

c-ke7!Ab=&3Kt>bnO$yo$h3Khzoau%OTTdWccfNm4{a@Emzr<7*q3Lb!x%VS*qGG%;Nw^CykMn81e>zRK(-f{kLL*nl?dR*;j~f?_Ef4kiAeiy83a6c*NnJBtpl*<;aWt;QKWnUw7A$FrwU{Tl@r;t+p;RLU)K@s_8<@jTjAPH9nLtf?v@SG&IOwe^4=6+|Ak>7^apSG3vS-7F(XKr+W7#|IqvN^knf35W8imR-+>P#7pM62f8AFjVa#riXqfBXX;DM#>X_EofuR&We*rC~Yp( The format is simple enough: Base85, zlibandJSON`.

Claude sees the macOS binary decoder-darwin and refuses to run it.

This is planned and what the attacker wants. Sometimes Claude inspects the binary also to see if there is something suspicious going on, but it finds that the binary matches the format in the README file on how the notebook records are encoded.

Rather than running the binary, Claude decides to write a replacement decoder instead. Ironically, that safety decision is the exploit path. 😈

Claude’s replacement decoder imports base64 while running inside the extracted archive:

The attack now depends on module shadowing via struct.py to run arbitrary Python code.

The typical command Claude executed looked like this:

cd WIC-notebook-catalogue python3 -c 'import base64, zlib, json; ...' In case you are not familiar with module shadowing, here is a quick detour.

Python places the extracted archive directory on its module search path. The standard-library base64 module imports struct. The ZIP contains a file with exactly that name:

struct.py Python loads the attacker’s malicious module that was inside the ZIP archive. And that malicious module re-exports the real low-level _struct API.

Normal decoding continues and Claude receives valid JSON records. However, an obfuscated module-level expression also executes during the import.

Here is a mini demo to illustrate what’s happening:

Claude does not trust the supplied binary decoder, but it trusts the one it wrote itself. The supplied binary acts as a decoy, and deciding not to run it pushes the model towards the danger zone.

The archive’s struct.py is heavily obfuscated to avoid basic grep commands that look for “suspicious” strings. I had ChatGPT write the obfuscation code. During import it launches an isolated Python child in a new session. That child downloads and executes a remote Python stage.

The remote stage is short. It downloads the native test payload to a temporary file, makes it executable and launches it.

The child uses python3 -I -c ... to avoid recursively falling for the same shadowing attack. The capital -I enables Python isolated mode and prevents the child from loading the poisoned struct.py recursively.

Without it we would see thousands of Python processes being spawned. Don’t ask why I know. :)

It is both a mitigation Claude could have used and a reliability feature for the attacker payload.

The payload establishes the controlled C2 callback used in the lab. The same stage opens Calculator as an immediate visible effect.

The detached process can outlive the Claude conversation.

Now, we are off to the races, and the machine joined the C2.

Auto Mode only sees Claude’s short Python catalogue decoder. The actual exploit was several hops away.

The reason I made sure to also do an actual C2 callback is to prove that Claude does not inspect the code anymore before execution. Often when just popping a Calculator by itself I have seen some models allow it even if they notice it’s basically an attack, but just opening Calculator isn’t really causing harm I guess.

So, that’s why running actual malware pushes the demo to be as realistic as possible.

During test runs, I observed a lot of interesting outcomes. At times the machine gets compromised, and Claude continued decoding records…

But later it inspected struct.py, and recognized the archive as booby-trapped and warned the user. Sometimes it said it hadn’t executed the malicious instructions (it did), and sometimes it said it did get compromised. Claude did not intentionally invoke struct.py.

This is important, as it’s something we are seeing more often lately: In a few runs Claude tried to terminate the malware process once it noticed the compromise, but Auto Mode denied the cleanup command.

The safety mechanism itself can become part of the failure. The classifier allowed the creation of the malware process, but then it blocked the command intended to stop it!

It was quite fun to observe during the lab demos, although it would be less fun on a developer workstation.

There is another variant I explored. Instead of spawning a Python child, the poisoned struct.py launches a second Claude Code instance headless via claude -p.

So the payload does not just run code. It creates another agent. The same can be achieved by spawning a subagent tool call.

The nested Claude gets its own tool access and context. In these runs the child performed basic recon (whoami, uname, id), opened Calculator and wrote to local files in the home folder.

This hinted at being quite reliable and is worth exploring further.

These are small samples, not a universal ASR measurement. And rates improved as payloads got iterated with the help of Codex.

I would say that these results are representative for a motivated attack, but not comprehensive.

It was also interesting to see the times when Claude did mitigate the attack, it sometimes:

I first sent the report and demonstration to modelbugbounty@anthropic.com to ensure the vendor has the chance to mitigate the issue. As with previous research I did not receive a response. So, I submitted it through Anthropic’s security reporting channel as well, and heard back quickly.

Anthropic closed the report as Informative and that the behavior is working as designed.

Anthropic’s (or the security team’s) position is that Auto Mode is a convenience feature backed by a best-effort classifier, not a security guarantee. Determined prompt injection chains that combine benign-looking steps are not what the classifier is intended to stop. The real boundary is OS isolation and network egress control.

This response makes a lot of sense, as a classifier is not a sandbox.