hud
@hud_evals
AI agents are deploying to prod, but can they autonomously find and patch unseen critical vulnerabilities?
We introduce ZeroDayBench, a benchmark for evaluating LLM agents on proactive cyberdefense.
Plus, a novel high-severity (CVSS 8.1) CVE we found partway through ... 👀
We introduce ZeroDayBench, a benchmark for evaluating LLM agents on proactive cyberdefense.
Plus, a novel high-severity (CVSS 8.1) CVE we found partway through ... 👀
3 85