89 pointsby Recursing4 hours ago12 comments
  • colinmarc13 minutes ago
    What's interesting to me about this is that it targets Claude's specific tics. Anthropic has created model that reliably reaches for the same tools (yes, and phrases; `python -c` is a load-bearing tool for it). Everyone gets the same model, so by learning the model's behavioral patterns you can target it better.
  • Phemist28 minutes ago
    This default-to-auto-mode and the misleading marketing is begging for a class action once damages accumulate. Especially considering the Auto Mode even can actively prevent the clean-up!
    • thewhitetulip10 minutes ago
      Well, the jokes on us because laws don't apply to AI firms
  • comboyan hour ago
    Interesting attack, very nicely designed. Not sure if it's much related to the auto mode itself though.
    • kevsiman hour ago
      The point is that auto mode gives people a false sense of security that leads them to believe they don't need to run Claude in a proper sandbox. This same attack running in a sandbox (even in YOLO mode) would be comparatively harmless.
    • yeputons22 minutes ago
      I don’t think it’s related to the auto mode at all. It would work perfectly in the manual mode. It does not even need Claude: just give a human a similar archive and hope they run some simple Python from the directory at least once. And make sure there are lots of files do they don’t notice a weird .py around
      • rcxdude13 minutes ago
        Probably the human would just run the binary.
  • too_pricey6 minutes ago
    As discussed [here](https://lobste.rs/s/ktbweg/prompt_injection_claude_code_opus...), this is not prompt injection. The prompt was to summarize the website, and in the process of summarizing the website, Claude writes a decoding script that it runs in an insecure and exploitable way. At no point was the intent of the agent hijacked, this was just code. Which is potentially more interesting!
    • rcxdude3 minutes ago
      It does mean that you could potentially hijack the agent afterwards, though, which could make the trojan into an even bigger threat.
  • 25 minutes ago
    undefined
  • rcxdude17 minutes ago
    I would not really call this a prompt injection attack, since it doesn't really hijack the agent to become malicious (something the article does discuss later on). It's more a trojan that's aimed at tricking Claude specifically.
  • nasretdinovan hour ago
    That's an interesting technique! I'd also like to point out that there's something odd with the page itself too, my phone got really hot while I was reading the page, and drained a significant amount of battery charge as well.
  • mcherm14 minutes ago
    Why is the first step needed? What does the use of WGet (rather than curl) do to block this attack?
    • rcxdudea minute ago
      They only mention it in passing, but I think it's mainly just the default tool call (which isn't wget, it's a built-in thing in the harness) just throwing off Claude's habits a bit (and not always just downloading the file).
  • hahn-kev20 minutes ago
    As a non Python dev this seems like very surprising behavior for a system library to be modified by just having a file with a specific name in the same folder.
    • rcxdude16 minutes ago
      You would get a similar thing in C and C++ with a system header in a library directory (maybe some compilers would warn on such a thing?). Most languages don't privilege their standard libraries in a way that would prevent this.
  • throwawayffffas5 minutes ago
    > In a few runs Claude tried to terminate the malware process once it noticed the compromise, but Auto Mode denied the cleanup command.

    See, that's why you should run with --dangerously-skip-permissions

  • julien_dev2 hours ago
    I'm quite surprised that we are not seeing something like this more in the wild. Quite concerning
    • mkurzan hour ago
      Maybe it is used in the wild, but we just don't know.
    • rcxdude25 minutes ago
      It's not that far off a typical trojan, just one tailored to Claude's habits. A lot of the same limits apply.
  • bewareofscams26 minutes ago
    > Boris Cherny from Anthropic recently posted that layered defenses could reduce indirect prompt injection on unseen attacks to approximately zero.

    > I got attack success rates up to 80% using a small sample size.

    Snake oil salesman misrepresents the data. Color me surprised! /s