1 pointby henryrobbins00an hour ago1 comment
  • henryrobbins00an hour ago
    Some additional details about these results that didn't make it into the main post:

    - The OpenATP "standard provers" were used; see docs [5] for model / harness configuration details

    - Time and cost are function of effort level, which may lead to unfair comparison across provers

    - FATE-X excludes task 10 since claude and grok hit session limits

    - FATE-X excludes leanstral and aristotle due to temporary endpoint failures

    - Deepseek's FATE-X accuracy is corrected from 2 to 3 due to verifier bug (now fixed)

    - 2 FATE-X deepseek misses are sorry-free, but rely on native_decide

    - Claude's FATE-X miss is due to a failed delegation to a background subagent

    - All costs come from underlying CLI, except codex which uses pricing table