Home
Tech Grid
News Room
Interviews
Think Stack
Articles
  • Threat Intelligence

Ridge Security Publishes First-of-Its-Kind AI Red Teaming Benchmark


Ridge Security Publishes First-of-Its-Kind AI Red Teaming Benchmark
  • by: Business Wire
  • |
  • September 4, 2026

Ridge Security, a leader in AI-powered offensive security and continuous security validation, today published the industry's first public benchmark evaluating multiple leading large language models in autonomous penetration-testing workflows. The results show the model is only half the equation. How an AI harness reasons, executes, adapts, and validates its actions matters just as much, especially for enterprise-grade requirements. Using its RidgeGen offensive-security harness, Ridge Security evaluated eight AI models across 96 model-target test runs against intentionally vulnerable environments.

Quick Intel

  • Ridge Security publishes first public benchmark comparing AI models for autonomous red teaming

  • Grok 4.5 led on cumulative coverage at 77%, Claude Opus 4.6 at 63%, Gemini 3 Flash at 52%

  • GPT-OSS-120B delivered peak efficiency at 16.9 findings per million tokens

  • Agent architecture and harness matters as much as model selection for security testing

  • Model alignment can interrupt authorized security testing mid-workflow

  • RidgeGen separates reasoning from execution with deterministic controls and verification

Model Performance Findings

The results showed significant differences in security coverage, cost and efficiency:

  • Grok 4.5 led on cumulative coverage at 77%

  • Claude Opus 4.6 reached 63% coverage, at an estimated $217 per run

  • Gemini 3 Flash achieved 52% coverage at approximately $5.42 per run

  • GPT-OSS-120B delivered the benchmark's highest peak efficiency at 16.9 findings per million tokens, at approximately $2.32 per run

The findings challenge the assumption that the highest-performing general-purpose LLM will automatically deliver the strongest autonomous security agent.

"Security teams are beginning to ask which LLM is smartest, as though that answer determines which AI system will be the best penetration tester," said Lydia Zhang, President of Ridge Security. "Our research shows that autonomous offensive security is a systems problem. The model needs an architecture around it that can manage execution, adapt to what it discovers and verify that a finding is real."

The Harness Matters

Autonomous penetration testing requires more than cybersecurity knowledge or the ability to generate exploit code. An agent must maintain state across an attack sequence, adapt when a path fails, execute tools, chain weaknesses together and ultimately prove that a vulnerability is exploitable. Ridge Security's benchmark shows the agent harness is a must-have layer in AI-powered offensive security. RidgeGen is designed to separate reasoning from execution and verification: the model reasons, the harness controls execution, and independent validation confirms whether a finding is real.

The model-agnostic architecture allows organizations to use different AI models without rebuilding their security platform as the AI landscape changes. RidgeGen combines multi-step attack reasoning with deterministic controls, so findings are backed by reproducible proof rather than an AI model's assertion.

Model Selection and Tradeoffs

Higher-performing frontier models can deliver stronger coverage, but at substantially higher run costs. Smaller and open-source models may offer a different balance of coverage, efficiency and deployment flexibility. And for many organizations, that tradeoff matters more than raw coverage alone.

"The highest-scoring model may not be the right model for every security task," said Nick Mo, CEO at Ridge Security. "A model-agnostic architecture gives security teams the flexibility to make those tradeoffs without locking their offensive security strategy to a single provider."

The benchmark also surfaced a separate issue: model alignment can interrupt authorized security testing. Frontier models may refuse certain actions mid-workflow, including payload generation or other exploitation steps, even when testing is conducted within defined boundaries. Ridge Security addresses this challenge by enforcing authorization and safety controls at the architecture and tool layers rather than relying solely on the model. Target boundaries, sandboxing, deterministic controls and auditability help keep autonomous testing bounded while allowing the AI to reason through complex attack paths.

From Model Selection to Harness Engineering

Ridge Security expects the center of gravity in AI-powered security to shift from picking the most capable model to building the systems that make any model more reliable, repeatable and verifiable.

"A powerful model is only one part of an autonomous security system," said Zhang. "You need the architecture around it to control execution, maintain context and prove the result. AI that finds a vulnerability is table stakes. What matters is whether that finding holds up under verification."

The benchmark gives security teams a framework for evaluating AI models based on how they perform in actual offensive security workflows, not just how they score on general-purpose AI benchmarks.

Read the full benchmark: The Harness Advantage in Autonomous Red Teaming: Why Frontier LLMs Alone Fail Offensive Security and How RidgeGen Solves the Alignment Dilemma, by Ken Huang, CISSP

About Ridge Security

Ridge Security delivers autonomous cybersecurity validation solutions that help organizations proactively manage risk and improve resilience. It develops agentic AI-based adversarial risk platforms that support continuous threat exposure management programs. Ridge has earned industry honors including the 2026 Frost & Sullivan Global New Product Innovation and Top Emerging Cyber Security Company. The company serves customers worldwide across finance, government, telecom, and enterprise sectors.

  • Offensive SecurityLLMCyber Security
News Disclaimer
Want to reach B2B tech decision-makers through TechIntelPro? Get our Media Kit
  • Share
Enterprise Tech News