The benchmark runs Kimi K3 through leading coding harnesses across a range of software tasks to measure how much the harness itself affects performance.
My main takeaway from our results is that verification loops and a few deterministic rules can make a coding harness perform significantly better, even with the exact same model.
More in the blog.