We scored 2,449 security findings with eight models. Every model inflated severity on average. Explore the results, context failures, and what unnecessary investigations could cost your team.

8 points•brene•1 day ago•1 comment•

1 comment

brene1 day ago
Eleanor from my team and I collab'ed on this benchmark. A few things worth stating up front because they shape how to read this. We constantly see LLMs saying "THIS IS A CRITICAL SECURITY FINDING" and we wanted to put it to the test. Most LLMs naturally gravitate towards using CVSS for severity scoring.

We wanted a task where model judgment could be checked against verified results from security engineers (ground truth) and root cause "why" security findings are always inflated. TL;DR LLMs still make too many severity judgements without the proper context, so naturally they bias towards the "worst case" scenario. Happy to answer any questions on this topic

Read the full thread on Hacker News →

Related stories