Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers(senior-swe-bench.snorkel.ai)

42 pointsby matt_d2 hours ago11 comments

_345a few seconds ago
This makes so much sense as to why I've always felt that Opus 4.8 was leagues ahead of GPT 5.5. It's so good at taking underspecified requirements and filling in the gaps with sensible approaches for your project
0xbadcafebee11 minutes ago
The "tasteful solves" is codified cargo culting. The software industry has a tendency to anthropomorphize software while playing to the ego of the programmer. The programmer imagines they are creating a "beautiful" subjective artistic expression, rather than a clock. Nobody's paying you to make paintings. They're paying you to build machines. Once you make the machine work, then you can go about gilding the lily.
The big mistake is conflating "making working software" with the "taste" part. These should never be considered the same thing. It devolves into bikeshedding and subjective opinionism, and detracts from the real purpose of the thing. Did you solve the user's problem? If not, shut up and make it work. If you did, then move on.
facorreia4 minutes ago
It's nice to see a new public benchmark from Snorkel. They're doing some pretty sophisticated stuff over there.
magnio11 minutes ago
I saw on Twitter that in an ML course at Tsinghua University, one of the tests asks students to write quizzes that fail the most LLM models as possible.
What if we create a benchmark that works like this and assigns ELO scores? Models fight head-to-head by writing a question, a bug, or an incomplete implementation, which the opponent has to answer, fix, or finish.
jonathanleane2 hours ago
Top solve rate is currently 24% with Opus 4.8... What's a competent human supposed to score?
- lacunary2 hours ago
  presumably whatever the top model uses and then some, since the human can use the model.
  I wonder if a model could score higher if it had a human at its disposal?
  - pishpash6 minutes ago
    Maybe models should ask for human-in-the-loop input, as a matter of convention.
guilhermecgs30 minutes ago
fable 5?
- guessmyname13 minutes ago
  The people who created the benchmark(s) don’t have access to Fable 5.
LiamPowellan hour ago
> You are a senior SWE-Bench reviewer, make no mistakes.
I don't know what a better approach would look like while still remaining feasible, however this approach of telling a LLM to make a subjective judgement seems fundamentally flawed.
Madmallardan hour ago
next round of trust me bro benchmarks
- dozerly21 minutes ago
  Just wait for the next 100 rounds. People love seeing the 65% -> 85% seemingly over and over again for every new model.
danpalmer2 hours ago
Why didn't they just make it "Staff SWE-Bench", would be much better smh. /s
But seriously, as an industry we're terrible at assessing engineering levels, I've worked with "senior engineers" who can't code and I've worked with "junior engineers" who could run rings around them.
Benchmarks like this should be much more precise about what they're actually testing, and what axes they're hard on. We also need to rise above prompts like "you are a senior engineer", it's woo, and it's far better to ask for precise outcomes.
- glaslong33 minutes ago
  Principal-SWE-Bench will take some time to run, because the LLM needs to wait for a crisis to present its solution, having correctly identified that the same solution would have been organizationally impossible to propose until that moment.
- amrrs2 hours ago
  As someone who's trying to get better assessments, I'm struggling to come up with objective coding tasks that evaluates all aspects of real life like planning, design choices, problem solving and context usage. From your experience with humans, Do you have any recommendations on what could be effective in measuring it?
  - allan_san hour ago
    I think the source of your issue is in your statement itself, why do you want a task that evaluate things as broad to be only a coding task ? Shouldn't it be a planning task, documentation task, knowledge retrieval task etc. And very certainly not with just an initial prompt but an existing codebase + existing doc + tickets ?
jocelyneran hour ago
[flagged]
purple-leafy2 hours ago
Benchmarks are great, but I feel like there’s a better way this seems quite subjective.
What you really need is an objective benchmark
- elian hour ago
  I actually really like subjective benchmarks, so long as it's a human (ideally me) grading the results. LLM as judge never made much sense.
  - charcircuit25 minutes ago
    The issue is that you can't do unsupervised learning if you require humans.
- echelonan hour ago
  > What you really need is an objective benchmark
  "When are all the software engineers unemployed?"
  - purple-leafyan hour ago
    Not sure I follow haha