27 pointsby rigelbm7 hours ago15 comments
  • alinebindel6 hours ago
    thorough work, good stuff.. it even runs a selection-bias analysis against their own benchmark and reports that some tasks that were disproportionately hard for a model. Rare to see a benchmark paper attack itself like that.
  • kmiens7 hours ago
    The comparison between harnesses is very nice. Interesting to see that using a different harness can bump the performance of the model as much as a new version (e.g., GPT 5.5+Codex ~= GPT 5.6+Terminus, at lower cost)
    • dvaplima4 hours ago
      That’s a nice discussion. Some people say that with current model capabilities, the real differentiator is the harness. What are the best harnesses you guys are using?
  • MatheusFelipe4 hours ago
    Interesting the idea of treating the benchmark as an evolving system rather than a static dataset.
  • Betaantunes117 hours ago
    The methodology was the most interesting part for me. The paper spends as much time explaining how the benchmark was built as the benchmark itself.
  • aamdias5 hours ago
    Using semantic perturbation to test whether difficulty survives rewording is really smart. Great work!
  • mrhectograma7 hours ago
    Refreshing to see something practical instead of another leaderboard battle. Also, props to the team for being so meticulous.
  • pedroaugusto-me6 hours ago
    This approach of not only producing the benchmark tasks, but also focusing on creating a data engine that will improve over time and produce up-to-date tasks that challenge the cutting-edge models is very interesting and valuable.
  • Luvison_Rafael7 hours ago
    nice!
  • VicElko4 hours ago
    good work!
  • ltononro6 hours ago
    [flagged]
  • caue_gimenez4 hours ago
    [dead]
  • dvaplima6 hours ago
    [flagged]
  • MstTK5 hours ago
    [flagged]
  • gramulho5 hours ago
    [flagged]
  • abernat4 hours ago
    [dead]