52 pointsby behnamoh8 hours ago11 comments
  • cjk37 minutes ago
    Anecdotally, I have also felt many of these same complaints.

    In my case, I have a semi-autonomous loop where Claude writes some code, and uses `codex exec` to do an adversarial review. I had what should have been a trivial feature go for 13 rounds of review/fix before I stopped it, where each round was just flip-flopping the same logic back and forth to try to make the tests pass. Codex kept (correctly) re-identifying the issues Claude was flip-flopping on. Codex even suggested fixes that would have worked; Claude ignored them repeatedly. I never saw anything remotely this bad on Opus 4.8.

    Additionally, I have a CLAUDE.md instruction to not silently defer anything, and ask me anytime it wants to do so. Opus 4.8 paid attention to this rule the overwhelming majority of the time. Opus 5 seemingly cannot be bothered.

    I have tried updating my CLAUDE.md according to Anthropic’s recommendations for Opus 5, but it doesn’t seem to have made any difference.

  • dahdum5 hours ago
    I’ve been using Opus 4.8 heavily with Claude Code heavily and I haven’t noticed any problems at all with 5. It seems a bit more proactive (not as much pre-ban Fable though). My claude.md and agent prompts are both 〜20kb, and I use memories/rules/skills as well.

    Agent to agent communication does appear to have improved considerably. I run the loops to the full context window and with proper working docs I don’t even notice a difference after compaction.

  • K0balt6 hours ago
    My experience with opus 5 has surfaced many of the complaints listed it the tweet. It seems to be slightly brighter than 4.8, but the laziness more than makes up for it.
  • Grimblewald3 hours ago
    mirrors my own findings as well. It royally bungles projects in a way that make git a bigger godsend than it should be (evem though it is).

    I really miss the 4.5 era, that was a magic time.

  • Frannkyan hour ago
    I think they're doing something wrong. I have access to better models, but I just use 4.6 and 4.8. Fable and Opus 5 burn tokens too fast for me. I also switched to a cheaper plan to work less and to experiment with OpenCode and open models when I hit the 5h limits. I'm doing this partly because I feel they've already started the enshittification of the old models, and I want to be able to just stop paying for the subscription. I only use Claude via subscription plan—I've never needed to pay the insane API pricing, and I'm able to do everything with OpenRouter and cheap Chinese models (DeepSeek Pro, Flash, MiMo 2.5 Pro for reasoning, and MiniMax M3 for agents). I'm just still using them for coding, but I'm not sure for how long—I haven't nailed down an open harness with the new models that's cheap, good, and fast. Open to suggestions and ideas.
  • xyzsparetimexyz3 hours ago
    Is it a really bad model in the same way that idk, Gemma 1.1 is a bad model? Stupid hyperbole
  • Bawoosette7 hours ago
    Most of this seems to be related to harness issues, specifically prompts provided by Claude Code. The upside of this is that it is likely cheap and easy for Anthropic to fix the issues by revising their approach to prompting their new models.
  • Cider99863 hours ago
    Was stunningly smart for the one question I asked it and it was on arena.ai so I didn't know which it was until voting for it.
  • dgellow5 hours ago
    Could we replace the link with https://xcancel.com/i/article/2081697911847481502 ? Faster to load, doesn’t nag you to sign up or install an app
    • orphea4 hours ago
      Usually the link remains original but someone helpful posts a comment with a link to xcancel or archive.org.
    • eddyg4 hours ago
      The guidelines⁽¹⁾ make it clear the OP did the correct thing: ”Please submit the original source.”

      ⁽¹⁾ https://news.ycombinator.com/newsguidelines.html

    • Cider99863 hours ago
      Recently the xcancel captcha changed and now I prefer nitter.net. seems the same but no captcha.
    • fragmede5 hours ago
      email hn@ycombinator.com
  • dude2507114 hours ago
    There had been a big pressure to increment the number.

    It's not like users can prove anything about a remote black box anyway.

  • 5 hours ago
    undefined