OpenAI’s GPT-5.6 Sol Scores 73.5% on ExploitBench, Setting New Cybersecurity Benchmark

OpenAI's GPT-5.6 Sol Scores 73.5% on ExploitBench, Setting New Cybersecurity Benchmark

OpenAI’s flagship GPT-5.6 Sol model has posted a 73.5% score on ExploitBench, up from GPT-5.5’s 47.9%, as the company claims state-of-the-art cybersecurity performance.

OpenAI has introduced GPT-5.6 Sol, the flagship model in its new GPT-5.6 family, and the benchmark numbers it has published are striking. On ExploitBench — a standard evaluation for end-to-end exploit progression — Sol scores 73.5% against GPT-5.5’s 47.9% at a comparable output-token budget. That’s a jump of more than 25 percentage points from one model generation to the next.

The GPT-5.6 family comprises three variants: Sol, the most capable; Terra, a lower-cost option; and Luna, the fastest. OpenAI is positioning Sol specifically as its strongest model for cybersecurity work, integrating it with a product called Codex Security, which is designed to apply Sol’s capabilities in practical defensive workflows — think automated secure code review, vulnerability validation, and patch generation in production-grade codebases.

What the Benchmarks Actually Show

The ExploitBench figure isn’t the only data point OpenAI has published. On ExploitGym, which tests a model’s ability to convert real-world vulnerabilities into working exploits within a time limit, GPT-5.6 Sol reaches a 24.9% pass rate within a two-hour cap, rising to 33.7% within six hours. GPT-5.5’s peak pass rate under comparable conditions was 15.1%. On SEC-Bench Pro, which measures proof-of-concept exploit generation, Sol scores 71.2% against GPT-5.5’s 45.8%.

External technical commentary, citing OpenAI’s GPT-5.6 Preview System Card, also reports that Sol achieves 96.7% on OpenAI’s internal cyber challenge set — above the company’s own threshold for a “High” cybersecurity capability rating. That figure should be treated with some caution: the underlying test methodology is not fully public, so it cannot be independently verified at this stage.

OpenAI has also announced that Sol achieves state-of-the-art results on something called “The Last Ones” cyber range, described as an advanced evaluation environment for high-end cybersecurity AI. Public technical documentation on this range remains limited, and specific scoring conditions have not been independently confirmed, so the “state-of-the-art” claim on that particular benchmark rests primarily on OpenAI’s own reporting.

Defensive Tool, Not Autonomous Attacker

It’s worth being precise about what these scores mean in practice. OpenAI’s own documentation is clear that GPT-5.6 Sol is better at finding and fixing vulnerabilities than it is at conducting fully autonomous, end-to-end attacks on hardened software. Testing did not show the model reliably compromising systems such as Chromium or Firefox without human intervention. The emphasis throughout OpenAI’s messaging is on defensive applications: vulnerability research, secure code review, threat modelling, and blue-team operations.

That framing matters. The model family has been rated High for cybersecurity capability under OpenAI’s Preparedness Framework — the same framework that also rates GPT-5.6 models as High in biological and chemical domains. But it has not reached the “Critical” threshold, which would trigger even stricter controls.

Sam Altman, OpenAI’s chief executive, has not yet issued a separate public statement specific to GPT-5.6 Sol’s cybersecurity capabilities at the time of writing. OpenAI’s published materials stress that the model was developed with close coordination with the United States government, and that multi-layered safety controls are in place: model-level training to refuse malicious requests, real-time content classifiers, and account-level behaviour monitoring.

Restricted Access and Safety Controls

GPT-5.6 Sol is not yet widely available. Access is currently restricted to a preview programme covering government-approved organisations and trusted partners, following coordination with the US government. Broader release of Sol, Terra, and Luna is expected in stages, with OpenAI working on standardised release procedures before opening access further.

Some analysts are already raising questions about this approach. Critics point out that relying on proprietary benchmarks — including the 96.7% internal challenge set figure — without full public transparency makes it difficult to independently verify the “state-of-the-art” claims. There are also concerns that even well-controlled releases of high-capability cyber AI could gradually lower the barrier to entry for sophisticated operations if safety measures are bypassed or the model architecture is replicated by others.

Security professionals, on the other hand, may see Codex Security and GPT-5.6 Sol as tools that could meaningfully reduce false positives, accelerate vulnerability discovery, and ease workload pressure in security operations centres. But software developers and DevSecOps teams will need to think carefully about oversight and responsibility when an AI system is proposing changes to production code.

What This Means for Kent Residents

GPT-5.6 Sol and Codex Security are currently in restricted preview, so direct access for Kent-based organisations — whether local councils, NHS Kent and Medway, or the county’s tech firms in Maidstone, Canterbury, and Ashford — is not yet available and will depend on vendor offerings or a future wider release. If security vendors and managed service providers do eventually adopt Sol-based tooling, local businesses and public-sector bodies could benefit indirectly through faster vulnerability detection in the software and cloud services they rely on. Kent organisations should monitor guidance from the National Cyber Security Centre, which has not yet issued specific public advice on GPT-5.6 but remains the primary UK authority on responsible AI use in cybersecurity. Any adoption involving sensitive local authority or health data would also need to comply with UK GDPR and relevant NCSC standards such as Cyber Essentials.

Source: @OpenAI

OpenAI's GPT-5.6 Sol Scores 73.5% on ExploitBench, Setting New Cybersecurity Benchmark Quiz

5 questions