3 pointsby sameerhimati7 hours ago1 comment
  • sameerhimati7 hours ago
    I was picking a search API for a research agent and every comparison I could find was run by one of the vendors so I ran my own.

    Most search benchmarks end up grading the answer the model produces, not the actual pages returned.

    Open to feedback!