Home
Tech Grid
News Room
Interviews
CISO POV
Think Stack
Articles
  • Agentic AI

Novee Launches PWNBench Benchmark for Agentic Penetration Testing


Novee Launches PWNBench Benchmark for Agentic Penetration Testing
  • by: GlobeNewswire
  • |
  • September 24, 2026

Novee, the leader in AI penetration testing, today announced that PWNBench, its benchmark for agentic penetration testing of live web applications, is launching with Fireworks Specialized Intelligence Index.

Practitioner-built benchmark gives security teams and model builders a real-world measure of how AI performs on live offensive security work. PWNBench measures how AI models perform on the ambiguous, multi-step reality of offensive security work, moving beyond the bounded, static tasks common in public benchmarks.

Quick Intel

  • Novee launches PWNBench v0.1 benchmark for agentic penetration testing of live web applications on Fireworks Specialized Intelligence Index
  • Practitioner-built standard measuring ambiguous multi-step offensive security work against live apps judged by offensive security experts
  • Launch results show divergence: Grok 4.5, Grok 4.6 and DeepSeek-V4-Flash-0731 hold efficient frontier on F0.5 against cost
  • Claude Opus 5 highest recall 51% for $1,400 API spend at k=3 vs Kimi K3 42% for $209, Grok 4.6 and Claude Opus 4.8 precision high-70s to low-80s
  • Foundational benchmark for Specialized Intelligence Index alongside Harvey Legal Agent Benchmark and Doximity BedsideBench
  • Available at fireworks.ai/specialized-intelligence-index and full methodology at novee.security

Why PWNBench Measures Live Offensive Security Beyond Static Benchmarks

PWNBench measures how AI models perform on the ambiguous, multi-step reality of offensive security work, moving beyond the bounded, static tasks common in public benchmarks. It gives security teams and model builders a practitioner-built standard for comparing how models perform on the capabilities and tradeoffs that matter in agentic penetration testing.

The launch results show how much today's leading models diverge on live offensive work. In the summary view, Grok 4.5, Grok 4.6 and DeepSeek-V4-Flash-0731 hold the efficient frontier on F0.5 against cost. Claude Opus 5 buys the highest recall, 51% for roughly $1,400 in API spend at k=3, against Kimi K3 at 42% for $209. On precision, Grok 4.6 and Claude Opus 4.8 sit in the high-70s to low-80s.

"Existing benchmarks can tell you a model recalled a vulnerability in code that it saw in training. It tells you nothing about whether that model can break into a running system it has never seen, and prove what is exploitable, the way a real attacker would," said Dan Padnos, Head of AI at Novee. "We built PWNBench to measure offensive security the way it actually happens – against live applications, judged by the people who do the work. Publishing it on the Specialized Intelligence Index makes the results of that discipline widely available."

"PWNBench shows why the Specialized Intelligence Index is needed," said Benny Chen, Co-Founder of Fireworks. "No one model leads on every dimension. By surfacing the tradeoffs across recall, precision, and cost, PWNBench helps teams understand which models are best suited to the demands of their security workflows."

How Fireworks Specialized Intelligence Index Surfaces Recall Precision Cost Tradeoffs

PWNBench is a foundational benchmark for the Specialized Intelligence Index, Fireworks destination for comparing open, closed, and specialized models on domain-specific work. Benchmarks are admitted under shared requirements for production relevance, task diversity, calibration, and contamination disclosure. It sits alongside evaluations from teams that set the standard in their own fields, including Harvey's Legal Agent Benchmark and Doximity's BedsideBench in healthcare.

 

About Novee

Novee is the AI penetration testing platform built to secure constantly changing environments against attackers operating at machine speed. Its purpose-trained AI reasons like a real attacker, uncovers novel attack paths continuously, and delivers precise, personalized fixes so teams can stay one step ahead of hackers.

  • Fireworks AIAgentic AIOffensive Security
News Disclaimer
Want to reach B2B tech decision-makers through TechIntelPro? Get our Media Kit
  • Share
Enterprise Tech News