Hacker News
new
top
best
ask
show
job
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks 2025
(
arxiv.org
)
2 points
by
janandonly
4 hours ago
1 comment
dima853
3 hours ago
How do you handle partial success in the benchmark? If an agent picks the wrong tool mid-chain but fixes it with backtracking, does it still count as a full success?