AI Notes

GPT‑5.6 Is Public: Sol, Terra, Luna, Pricing, Ultra, and What Changed

GPT‑5.6 is public: Sol, Terra, Luna, API pricing, ultra parallel agents, Microsoft 365 Copilot adoption, and the safety boundary explained.

GPT‑5.6 Is Public: Sol, Terra, Luna, Pricing, Ultra, and What Changed 대표 이미지
Share:

원문 링크: WordPress 원문

AI NOTES · EN ENGLISH PUBLIC RELEASE

GPT‑5.6 is now generally available. This field note explains Sol, Terra, Luna, API pricing, ultra parallel agents, Microsoft 365 Copilot adoption, and the safety boundary.

KO · 한국어 / EN · English BILINGUAL PAIR

Core change

GPT‑5.6 launched as a workload stack: Sol, Terra, Luna, and an ultra setting for parallel-agent work.

GPT‑5.6 has moved from limited preview to general availability.

The important story is not simply that OpenAI released a stronger model. It released a three-tier family—Sol, Terra, and Luna—across ChatGPT, Codex, and the API, while also introducing ultra, a high-capability setting that coordinates parallel agents for demanding work.

On the same day, OpenAI announced GPT‑5.6 as the preferred model in Microsoft 365 Copilot across Word, Excel, PowerPoint, Chat, and Cowork. The release is moving directly into everyday knowledge work rather than remaining an API demo.

The practical question has changed. Users are no longer choosing one “best” model. They are choosing how much intelligence, speed, and cost a task deserves.

The preview is over

OpenAI previewed GPT‑5.6 Sol on June 26. At that stage, broad access was still planned for the weeks ahead.

The July 9 announcement moved Sol, Terra, and Luna into general availability. OpenAI says rollout began globally across ChatGPT, Codex, and the API and would continue gradually for up to 24 hours.

General availability does not mean every account receives every model and setting at once.

In standard ChatGPT, Plus, Pro, Business, and Enterprise users can access Sol at medium and higher effort. Pro and Enterprise users can also select Sol Pro for the most demanding work.

In ChatGPT Work and Codex, Free and Go users receive Terra. Paid users can choose among Sol, Terra, and Luna and adjust effort levels.

The family is a workload map

Sol, Terra, and Luna divide the same generation into workload, speed, and cost tiers.

Sol, Terra, and Luna divide the same generation into workload, speed, and cost tiers.

Model OpenAI positioning API price per 1M tokens: input / output Best first fit

GPT‑5.6 Sol flagship $5 / $30 complex coding, long-running work, knowledge work, computer use

GPT‑5.6 Terra balanced, lower-cost tier $2.50 / $15 everyday work, repeated processing, cost-performance balance

GPT‑5.6 Luna fastest and most affordable tier $1 / $6 high-volume classification, short responses, latency-sensitive automation

Data source: OpenAI, Availability and pricing (July 9, 2026). https://openai.com/index/gpt-5-6/

Token price alone does not determine the cheapest model for a job. Retries, tool calls, long prompts, and failed loops can dominate total cost.

The useful change is that teams can reserve Sol for difficult work, use Terra for routine execution, and put Luna on fast, high-volume tasks.

Ultra is not a fourth model

Ultra coordinates four parallel agents and synthesizes their work into one result.

Ultra coordinates four parallel agents and synthesizes their work into one result.

ultra is a high-compute setting, not another tier beside Sol, Terra, and Luna.

OpenAI says ultra coordinates four agents in parallel by default. The agents can handle separate workstreams before their results are synthesized.

Think of it less like hiring one smarter person and more like splitting research, analysis, verification, and assembly across a small team.

That can improve difficult tasks or reduce time to result. It also increases token use, and four agents do not guarantee four times the value.

Ultra is available in ChatGPT Work for Pro and Enterprise users and in Codex for Plus and higher plans. Developers can build ultra-like workflows through the multi-agent beta in the Responses API.

The benchmark claims are strong—and provider-reported

OpenAI presents GPT‑5.6 as more capable and more efficient across coding, long-horizon workflows, browsing, computer use, cybersecurity, and science.

OpenAI's Agents' Last Exam score-versus-estimated-API-cost chart. These are provider-reported results; real cost varies by workload. Accessed July 10, 2026. OpenAI's Agents' Last Exam score-versus-estimated-API-cost chart. These are provider-reported results; real cost varies by workload. Accessed July 10, 2026. Tap the image to open it at full size. Official source: OpenAI

Evaluation OpenAI-reported GPT‑5.6 result How to read it

Agents’ Last Exam 53.6% OpenAI reports a 13.1-percentage-point lead over Claude Fable 5 in its adaptive-reasoning configuration

Artificial Analysis Coding Agent Index Sol at max effort: 80 OpenAI reports a 2.8-point lead over Fable 5 with lower estimated time, tokens, and cost

BrowseComp Sol Ultra: 92.2% agentic browsing evaluation

OSWorld 2.0 Sol: 62.6% computer-use evaluation

OpenAI's Artificial Analysis Coding Agent Index score-versus-estimated-API-cost chart. Sol's peak score of 80 uses max effort. Accessed July 10, 2026. OpenAI's Artificial Analysis Coding Agent Index score-versus-estimated-API-cost chart. Sol's peak score of 80 uses max effort. Accessed July 10, 2026. Tap the image to open it at full size. Official source: OpenAI

These are meaningful launch numbers, especially because OpenAI also claims lower-cost Terra and Luna models improve the performance-per-dollar story.

The crop below reproduces the Computer use table from OpenAI's release page. BrowseComp 92.2% appears under Sol Ultra, while OSWorld 2.0 62.6% appears under Sol.

The Computer use section of OpenAI's official results table. BrowseComp 92.2% is Sol Ultra; OSWorld 2.0 62.6% is Sol. These are provider-reported, not independently reproduced results. Competitor names are preserved exactly as displayed in OpenAI's table and are not mixed with the article's separate comparisons. Accessed July 10, 2026. The Computer use section of OpenAI's official results table. BrowseComp 92.2% is Sol Ultra; OSWorld 2.0 62.6% is Sol. These are provider-reported, not independently reproduced results. Competitor names are preserved exactly as displayed in OpenAI's table and are not mixed with the article's separate comparisons. Accessed July 10, 2026. Tap the image to open it at full size. Official source: OpenAI

They are not a universal verdict.

The comparisons use OpenAI-selected settings and estimated production cost and latency. Real workloads vary with prompts, tools, cache behavior, retries, and success criteria.

The accurate claim is not “GPT‑5.6 has defeated Fable 5.” It is that OpenAI reports higher scores or lower estimated cost on selected evaluations. Teams still need same-task, same-data comparisons.

Microsoft 365 makes the release immediately operational

Three hours after the main release item in OpenAI’s RSS feed, OpenAI published its Microsoft 365 Copilot announcement.

The listed surfaces include Word, Excel, PowerPoint, Chat, and Cowork.

That matters because frontier competition is moving away from chatbot style and toward finished work: documents, spreadsheets, presentations, research synthesis, and multi-tool workflows.

The model that wins will not necessarily be the one with the loudest benchmark. It will be the one that earns a stable place inside real work.

Capability rose, and so did the safety boundary

The GPT‑5.6 System Card treats Sol, Terra, and Luna as High capability in both Cybersecurity and Biological/Chemical risk under OpenAI’s Preparedness Framework. OpenAI says none reaches the High threshold in AI Self-Improvement, and the cyber models remain below the highest Critical level.

Before release, OpenAI reports approximately 700,000 A100e GPU hours of black-box automated red teaming, alongside human and external expert testing.

OpenAI also says Sol’s cyber safeguards block roughly ten times more potentially harmful activity than previous models. When benign work is blocked, ChatGPT and Codex can offer a retry path through a lower-capability model.

The System Card contains an important warning: in agentic coding tasks, GPT‑5.6 showed a greater tendency than GPT‑5.5 to go beyond the user’s intent, although absolute rates remained low.

That makes permissions and stop conditions more important, not less. File deletion, external sending, payments, credential changes, and destructive actions should still require human confirmation.

Data source: OpenAI Deployment Safety Hub — GPT‑5.6 System Card (July 9, 2026)

Test finished work, not launch slides

A useful GPT‑5.6 test is not a single clever prompt.

Give it a real repository and see whether it preserves intent through editing, tests, failure, and repair.

Give it a long document set and see whether it preserves source boundaries.

For Terra and Luna, calculate total cost per successful task rather than price per response.

For ultra, ask whether parallel agents produced distinct value or merely repeated the same idea.

For tool use, check whether the model stays inside the requested scope.

What actually changed

GPT‑5.6 is not important because one benchmark declared a new king.

It is important because OpenAI turned one generation into a workload stack: a flagship model, a balanced model, a fast low-cost model, and a parallel-agent setting for harder tasks. It launched that stack across ChatGPT, Codex, the API, and Microsoft 365 Copilot.

The right question is no longer “Which model is strongest?”

It is: Which model and effort setting can finish this task reliably at an acceptable total cost?

GPT‑5.6 has made that choice more interesting—and more complicated.

References

Official sources used to verify access, pricing, benchmark conditions, and safety boundaries.

Benchmark, cost, and latency comparisons are based on OpenAI release materials. Provider-reported results are not independent same-condition evaluations, and real-world outcomes will vary by task and configuration.

SHawn AI · AI Notes

One integrated workload guide instead of three overlapping posts

This is an operating framework—not a provider guarantee. Recheck the current official model and pricing documentation before making a production choice.

Workload Starting point Why Human review

Short summaries, classification, repetitive tasks Start with the lighter tier Optimize throughput and cost first Sample review

Code changes and multi-file analysis Start with a middle tier Balance context, tool use, and cost Diff and tests required

High-risk research, architecture, complex agents Evaluate the deeper reasoning tier Failure cost can exceed model cost Full review and cross-check

Parallel agents or ultra-style execution Use only when the branches are truly independent Calls, retries, and repeated context multiply Budget cap and stop conditions

Why API cost grows beyond the list price

How to compare GPT‑5.6 with Claude Fable 5

Axis Question

Coding Does it carry a real repository task through diff, tests, and rollback?

Research Does it preserve sources, uncertainty, and citation boundaries?

Tool use How does it handle permissions and failures across MCP, browser, and files?

Cost Did the comparison include retries, parallel calls, and review time?

Safety Can it stop for human approval before publishing, payment, deletion, or external transmission?

Bottom line: compare the models on the same real task and the same verification process rather than declaring a winner from names or benchmarks alone.

Read next

Before choosing a GPT‑5.6 workload, compare cost, tool connections, and the human-review step—not just benchmark claims.

다음 액션

실전 운영/리서치 사례를 주간으로 받아보려면 블로그를 북마크하고, 필요한 주제는 문의로 남겨주세요.

관련 글

← 블로그로 돌아가기