CVE-Bench: testing LLM agents on real-world vulnerability patches
↗Five frontier LLMs were tested on 20 real CVEs across three prompt types; no model reliably fixes vulnerabilities, with a best 50% solve rate and significant cross-family differences; token cost varies up to ~4x by model, and locate prompts are the hardest test of genuine security reasoning.
May 29, 20261%