3 pointsby IreneAI6 hours ago2 comments
  • claudiusa9 minutes ago
    Are the published numbers single-run or averaged, and which model does the judging? With LLM-as-judge scoring I would expect a couple of points of run-to-run noise, which does not matter for the top spot but matters a lot for the middle of the table.
  • MemoryEnthusiasan hour ago
    [dead]