OpenAI has launched a three-model GPT-5.6 family targeting health intelligence, with its cheapest option reportedly beating its predecessor’s best performance for a fraction of the price.
OpenAI has released the GPT-5.6 model family, positioning it as a major advance in health-related artificial intelligence — covering everything from patient-facing symptom guidance to clinician support tools. The headline claim, posted by an OpenAI health lead and shared widely across the company’s communications channels, is striking: GPT-5.6 Luna, running at its *lowest* reasoning effort setting, outperforms GPT-5.5 running at its *highest* reasoning setting on health tasks, while costing around 25 times less per use.
That cost ratio is based on OpenAI’s own internal economic comparisons and has not been independently audited or validated by external bodies. But if the figures hold up to scrutiny, the pricing shift alone could reshape how health-technology companies and institutions approach AI deployment.
Three Models, One Health Focus
The GPT-5.6 family comprises three distinct models, each aimed at a different point on the capability-cost curve. GPT-5.6 Sol sits at the top — the flagship, highest-capability option, and the most expensive. GPT-5.6 Terra occupies the middle ground, offering capable performance at roughly half the cost of GPT-5.5 for comparable tasks, according to OpenAI’s product materials. GPT-5.6 Luna is the fastest and cheapest of the three, yet OpenAI’s internal evaluations place it above GPT-5.5 on health reasoning tasks.
All three models are classified by OpenAI’s deployment safety documentation as *high capability* in biological and chemical domains — but below the company’s highest defined risk threshold. Safeguards are in place, according to those materials, to restrict guidance on engineering dangerous biological threats or conducting fully autonomous cyberattacks, while still permitting legitimate life sciences research.
What the Evaluations Show
OpenAI’s health evaluation work drew on physician-led assessments across around 20,000 individual axis ratings, covering accuracy, communication, completeness, instruction following, and helpfulness for health decisions. Across those axes, GPT-5.6 models reportedly achieved a higher proportion of “perfect” ratings than physician-written responses. The methodology has been described by OpenAI, but the numerical results have not been independently verified by UK regulators or academic bodies.
GPT-5.5 models were already handling over 230 million health-related queries weekly globally before GPT-5.6’s release, according to OpenAI — a figure that is self-reported and not confirmed by external sources. GPT-5.6 is designed to absorb and extend that workload at lower cost.
The improvements go beyond raw accuracy. OpenAI’s evaluation reports state that GPT-5.6 models show better recognition of when urgent care may be needed, clearer handling of uncertainty, and more accessible communication for lay users — areas where earlier models, including GPT-5.5 Instant, had already seen targeted upgrades.
Safety Boundaries and Unanswered Questions
OpenAI has been careful to frame GPT-5.6 as an assistive tool rather than a clinical replacement. The company’s own materials state that the models are not a substitute for professional medical advice or clinical judgement, and that deployment in health workflows should come with appropriate oversight.
That framing is likely to matter a great deal in regulated markets. Some clinicians and digital rights researchers have questioned the weight placed on OpenAI’s own evaluation claims, and the “25x cheaper” cost ratio in particular has yet to be validated by independent economic analysis. The claim that GPT-5.6 outperforms physician-written responses across key quality axes is also likely to draw scrutiny — not because it is implausible, but because the conditions of those evaluations, and how they translate to real-world clinical settings, remain to be tested outside OpenAI’s own research environment.
On cybersecurity, OpenAI’s benchmarks show GPT-5.6 achieving 73.5% on ExploitBench, compared with GPT-5.5’s 47.9% at a comparable output-token budget. These figures are published by OpenAI and have not been independently audited.
Pricing Strategy and Market Implications
The three-tier structure — Sol, Terra, Luna — reflects a deliberate attempt to widen access without flattening the product range. Luna’s low cost per million tokens is positioned to bring advanced health intelligence within reach of smaller organisations, app developers, and institutions that couldn’t justify the price of earlier flagship models. Terra sits in the middle, offering a credible performance-to-cost ratio for mid-scale deployments. Sol targets the most demanding use cases: medical literature review, protocol design, and complex clinical informatics work where GPT-5.6 sets what OpenAI describes as state-of-the-art results on benchmarks including complex browsing, tool use, and computer use.
Meanwhile, the aggressive pricing on Luna in particular is likely to intensify competition among AI providers operating in the health-tech space, pushing rivals to lower costs or differentiate on other grounds — such as on-premise deployment options, open-source transparency, or bespoke clinical fine-tuning.
What This Means for Kent Residents
Kent residents who use ChatGPT or other OpenAI-powered apps for health information may find the experience more accurate and easier to understand as GPT-5.6 rolls out, though these tools remain non-clinical and are not a substitute for NHS care. Any formal use of GPT-5.6 within NHS Kent and Medway services — whether for triage support, administrative drafting, or clinical decision support — would need to pass through NHS England’s clinical safety and AI procurement frameworks, comply with UK GDPR, and meet Care Quality Commission expectations; there are currently no official figures from NHS England or the Kent and Medway Integrated Care Board on GPT-5.6 adoption in local health settings.
Source: @OpenAI
OpenAI's GPT-5.6 Luna Outperforms GPT-5.5 on Health Tasks at 25 Times Lower Cost Quiz
5 questions