Tomorrow, the most powerful flagship GPT-5.6 Sol will debut! The brand-new Ultra mode ushers in the era of multi-agents.
```
OpenAI is about to launch its latest flagship model, marking a new phase in AI capabilities and commercial deployment.
OpenAI CEO Sam Altman announced on social media Tuesday that GPT-5.6 Sol will officially launch this Thursday. This flagship model of the GPT-5.6 series features a brand-new ultra multi-agent mode and max reasoning intensity, setting new records in core benchmarks such as coding, biology, and cybersecurity.

In terms of pricing, Sol is priced at $5 per million input tokens and $30 per million output tokens, making it the most premium of the three products in the GPT-5.6 series. OpenAI also announced that Sol will launch on Cerebras hardware in July, with inference speeds up to 750 tokens per second.
This release adopts a phased strategy, initially opening API and Codex access only to select trusted partners. OpenAI plans to roll out the entire GPT-5.6 series to a broader user base in the coming weeks.
Flagship performance: Sol Ultra leads the industry at 91.9%
On the coding benchmark Terminal-Bench 2.1, GPT-5.6 Sol Ultra ranked first with a score of 91.9%, followed closely by GPT-5.6 Sol at 88.8%. Competitor Claude Mythos 5 was third with 88.0%, a difference of about 0.8 percentage points. Gemini 3.1 Pro Preview came last with 70.7%, showing a significant gap from the first tier.

In terms of tier distribution, GPT-5.6 Sol Ultra and Sol form the first tier; Claude Mythos 5, GPT-5.6 Terra, and Claude Fable 5 are all above 84%, forming the second tier; GPT-5.5 and GPT-5.6 Luna belong to the third tier.
In cost-efficiency, GPT-5.6 Sol consistently scores highest at the same API cost, making it the top value in the series. GPT-5.5 and GPT-5.6 Luna show a clear "cost bottleneck"—extra investment yields only limited performance gains.

For reasoning ability, as the number of output tokens increases, GPT-5.6 Sol's score improves most steeply, indicating it can most effectively leverage complex reasoning to enhance output quality; Luna's curve is relatively flat, and even with more output, quality improvement is very limited.

New mode: Ultra multi-agent architecture surpasses single model limits
GPT-5.6 introduces two key technical upgrades. First is max reasoning intensity, giving Sol ample time for deep reasoning. Second is ultra mode, which calls upon sub-agents to collaborate, breaking through the limits of a single agent's capabilities and designed to accelerate complex tasks.
In biology, Sol achieved better results than GPT-5.5 on GeneBench v1, which evaluates long-term genomics and quantitative biological analysis, using fewer tokens and demonstrating higher computational efficiency.
In cybersecurity, Sol competes with Claude Mythos Preview on ExploitBench using only about one third of the output tokens. On the ExploitGym benchmark created by UC Berkeley researchers in partnership with OpenAI and other frontier labs, all three GPT-5.6 models—Sol, Terra, and Luna—show marked improvement in cybersecurity capabilities as reasoning intensity increases.
OpenAI states that Sol is better at helping users discover and fix vulnerabilities than reliably executing end-to-end attacks, emphasizing that its primary goal is to benefit defenders.
Three-level product matrix: Pricing meets different demand levels
The GPT-5.6 series adopts a new naming system: numerals indicate model generations, while Sol, Terra, and Luna represent three independently evolving capability levels, aiming to provide clearer choices of intelligence, speed, and cost for users and developers.
For pricing, Sol is $5 per million input tokens and $30 per million output; Terra is $2.50 input and $15 output, about half of Sol; Luna is $1 input and $6 output, the lowest-cost option in the series. OpenAI notes that Terra's performance rivals GPT-5.5 at half its cost, while Luna provides basic capability at the lowest price.
Regarding caching, GPT-5.6 introduces more predictable prompt caching, supporting explicit cache breakpoints and a minimum 30-minute cache lifetime. Cache writes are billed at 1.25 times the uncached input rate, while cache reads enjoy a 90% discount.
Additionally, OpenAI plans to launch GPT-5.6 Sol on Cerebras in July, with inference speeds up to 750 tokens per second. Initial access will be limited to some customers and gradually expanded as capacity increases.
Safety mechanisms: Tiered defensive stack and seven million GPU hour red team tests
In light of Sol's strong cybersecurity capabilities, OpenAI has equipped the GPT-5.6 series with its most robust safety system to date and employs a phased release strategy.
The defensive framework includes multiple layers: model-level rejection mechanisms during training, real-time cybersecurity and biology abuse classifiers in the generation process, account-level behavior review, and differentiated access control. For high-risk requests, the system can pause output during generation, a larger reasoning model reviews the context, and can intercept content before it reaches the user.
For stress testing, OpenAI committed over 700,000 A100-equivalent GPU hours for automated red teaming, focusing on discovering universal jailbreak attacks effective across multiple prompts and contexts, supplemented by third-party human expert red teams.
According to OpenAI's readiness framework, GPT-5.6 Sol has not crossed the "critical" limit in cybersecurity. In evaluations involving Chromium and Firefox, Sol identified vulnerabilities and exploit primitives but did not autonomously generate usable complete attack chain exploit in test conditions.
OpenAI states that granting limited preview access to trusted partners is part of ongoing communication with the US government, but clarifies that "we do not believe such government access procedures should become standard practice." OpenAI will work with government to develop repeatable processes for future model releases.
Risk Warning and DisclaimerThe market is risky, investment needs caution. This article does not constitute personal investment advice nor does it consider the individual investment goals, financial status, or needs of any user. Users should consider whether any opinions, views, or conclusions in this article fit their particular situation. Decisions made based on this article are your own responsibility. ```