We wanted a task where model judgment could be checked against verified results from security engineers (ground truth) and root cause "why" security findings are always inflated. TL;DR LLMs still make too many severity judgements without the proper context, so naturally they bias towards the "worst case" scenario. Happy to answer any questions on this topic