8 pointsby brene7 hours ago3 comments
  • brene7 hours ago
    Eleanor from my team and I collab'ed on this benchmark. A few things worth stating up front because they shape how to read this. We constantly see LLMs saying "THIS IS A CRITICAL SECURITY FINDING" and we wanted to put it to the test. Most LLMs naturally gravitate towards using CVSS for severity scoring.

    We wanted a task where model judgment could be checked against verified results from security engineers (ground truth) and root cause "why" security findings are always inflated. TL;DR LLMs still make too many severity judgements without the proper context, so naturally they bias towards the "worst case" scenario. Happy to answer any questions on this topic

  • fortitudedev5 hours ago
    [dead]
  • 5 hours ago
    undefined