4 pointsby Bender2 hours ago1 comment
  • bigyabaian hour ago
    GLM 5.3. Anthropic even made the mistake of comparing it to Mythos (lol): https://www.anthropic.com/research/glm-5-3-and-the-spread-of...

    > Like Claude Mythos Preview, GLM-5.3 has strong capabilities for autonomously building end-to-end cyber exploits. But GLM-5.3 is unlike other frontier models in that it has been released without meaningful safeguards to limit misuse. We find that attackers can bypass GLM-5.3’s safeguards between 64% and 100% of the time with simple techniques in our simulated tests.

    > We find that GLM-5.3 develops end-to-end exploits in 50 of 410 attempts. Claude Mythos Preview did so at a similar rate—in 56 of 410 attempts.

    • Benderan hour ago
      I may end up going that direction. I would ideally like to find something that is purpose built to do code pen-testing so I do not have to bypass anything. There are forks of other models built for this, maybe there is a fork of GLM too.

      Grok is telling me it can do code security reviews but I am not sure I believe it.