1 pointby onnies7 hours ago1 comment
  • cloudie786 hours ago
    I might be dumb, can someone explain to me how these are supposed to be useful? What if I want to tackle one of these myself, manually?

    I expand the sections and it’s all full of prompt slop

    • onnies5 hours ago
      This benchmark measures an agents ability to solve SRE/On-call tasks!

      Which expanded sections are slop? Happy to point you in the right direction - we also have a github repo: https://github.com/abundant-ai/incident-arena

      which has the tasks in harbor format (instruction.md files etc)