GPT-5.6 Test: We Rate Sol, Terra and Luna

GPT-5.6 Test scene with Sol Terra and Luna running pro workloads on split screen in a lab

OpenAI launched GPT-5.6 on July 8, with a global rollout starting the following day, and for the first time since GPT-5.5 the family splits into three distinct models: Sol, Terra and Luna. Our verdict after framing each model against its real-world use case: Sol targets long, verticalized reasoning, Terra replaces GPT-5.5 on daily pro workloads, and Luna plays the low-cost card on volume.

Key Takeaways

  • Sol is the flagship, tuned for biology, chemistry and cybersecurity, priced at $5 per million input tokens and $30 per million output tokens.
  • Terra replaces GPT-5.5 on everyday marketing and product workflows, at an intermediate $2.50 input and $15 output per million tokens, with no dramatic performance jump.
  • Luna drops to $1 input and $6 output per million tokens to absorb volume workloads, right into the zone where teams rationalize their bill after xAI released Grok 4.5 the day before.

Three models, three workloads: what the GPT-5.6 family says about the OpenAI repositioning

The first thing that strikes us about GPT-5.6 is not a single model but a family segmented by usage intent. Since GPT-5.5, OpenAI had let frontier use cases (long, verticalized reasoning) and volume use cases (agents and high-throughput pipelines) sit inside the same tier. The split between Sol, Terra and Luna gives each workload its own price entry point.

Sol is the flagship. It is the only model in the family with claimed tuning across biology, chemistry and cybersecurity, which signals a clear intent toward verticalized deployments and multi-step tasks that require long reasoning. Pricing at $5 per million input tokens and $30 per million output tokens places Sol in OpenAI’s historical high tier, without the off-catalog premium some competitors set for their frontier tier.

Terra takes the seat GPT-5.5 held until yesterday. It is the daily pro model, with performance comparable to GPT-5.5 on standard tasks and pricing set at $2.50 per million input tokens and $15 per million output tokens. The product note does not promise a spectacular leap on this segment, which is consistent: Terra is not there to impress, it is there to hold production load without forcing anyone to rewrite existing prompts.

Luna is the low-end entry point at $1 per million input tokens and $6 per million output tokens. Its target is explicit: volume. Teams processing thousands of requests per day (agents, classification, summarization, data enrichment) get a model calibrated for the invoice, not the demo. Our earlier verdict on the Claude Sonnet 5 arrival on the Anthropic side already showed that this affordable-volume tier is becoming the real battleground between labs.

The release context matters too. According to the announcement note, the Trump administration cleared the rollout after additional testing beyond the late-June preview. Concretely, that means the calendar was not forced to match a competitor, but the final green light landed on a window where xAI had just shipped Grok 4.5 the day before. The overlap is too clean to ignore on the positioning side.

Our reading: OpenAI is no longer trying to hold “one model that does everything”. The GPT-5.6 family acknowledges that pro workloads are three distinct segments, with three cost curves and three latency profiles, and that it is better to frame each with the right model than to pay Sol prices for a job Luna handles just as well.


GPT-5.6 Test

Hands-on verdict: Sol vs Terra vs Luna on real pro workloads

A methodology note first. Preview access to GPT-5.6 was limited in late June to a narrow partner circle, mainly the Codex partner track and a few external testers. Our verdict therefore relies on public and verifiable elements (official pricing, product positioning, release note) and rational calibration against the previous generation. We are not claiming a heavy test across thousands of requests.

On daily marketing tasks (email sequence drafting, campaign brief, insight extraction from a report), Terra is the default choice. Migrating from GPT-5.5 does not require prompt rewrites, and the $2.50 per million input tokens tag stays within the budget of most content teams. Nothing justifies flipping to Sol for these workloads, unless the reasoning demanded goes beyond the classic brief (multi-step competitive analysis, full CRM synthesis, deep sector-level intelligence).

On dev and product workloads, the split shifts. For a CTO running a backend team, Terra covers 80 % of daily needs (refactoring, standard code review, unit test generation). Sol becomes relevant as soon as reasoning stretches beyond a single step: security audit of a full auth flow, architecture migration, an agent chain coordinating multiple tools. Our Leanstral 1.5 test on the Mistral side already surfaced this gap between “model that answers” and “model that reasons long”.

Luna plays in a different league. Its purpose is clear: absorb volume with minimal unit cost. For a CEO scaling an AI support team, a founder automating lead enrichment, or a data analyst who has to summarize hundreds of tickets a day, Luna at $1 per million input tokens rewrites the feasibility math. It is not a reasoning model, it is an execution model at scale.

For a freelancer, the logic changes again. Sol stays out of budget for most independents ($30 per million output tokens adds up fast on a long writing day), and Luna rarely suffices when the deliverable is a client-facing report. Terra is the default sweet spot for freelancers delivering value-added work without volume to process.

Sol’s specific tuning on biology, chemistry and cybersecurity sits outside the “generalist pro” scope and targets R&D teams, labs and CISOs. For these profiles, Sol becomes a production tool, not a gadget. For everyone else, the tuning brings nothing tangible and does not justify the premium.

One angle worth noting for pro users: the family split also changes how a team should design its stack. Instead of routing every request to the same model, the pragmatic setup is now to route by intent at the pipeline level. Classification and enrichment go to Luna, drafting and standard reasoning go to Terra, and only the tasks that truly require multi-step reasoning go to Sol. This routing logic used to be a nice-to-have. With three price points that vary by a factor of five between input tokens, it becomes a budget lever that a CFO can actually read on the invoice.


Also on Horizon:


Pricing and competition: where GPT-5.6 stands against Grok 4.5 and Claude Sonnet 5

Pricing is the real signal of this release. Sol at $5 input and $30 output per million tokens confirms that OpenAI is not trying to break the high-end market, but to hold the stable flagship seat. Terra at $2.50 input and $15 output settles into an intermediate zone that stays above several generalist competitors. Luna at $1 input and $6 output joins the low end of the market.

Grok 4.5, shipped the day before by xAI at $2 input and $6 output per million tokens, ends up head to head with Terra on daily workloads and head to head with Luna on the low-cost tier. The battle plays out over a few cents per million tokens, which looks trivial from a distance but adds up to thousands of dollars a month for a production agent pipeline. The choice can no longer be a feeling. We ran that rival through the bench too, our Grok 4.5 coding test.

Against Claude Sonnet 5 on the Anthropic side, the reading changes. Sol and Sonnet 5 target similar workloads (long reasoning, verticalization, multi-step tasks), but the product strategies diverge. OpenAI bets on an explicit bio/chem/cyber tuning, Anthropic historically bets on safety and long reasoning. Our Claude vs ChatGPT vs Gemini comparison for pros already covers that ground, and Sol reinforces the OpenAI flagship stance without playing the disruptor.

Our global verdict: GPT-5.6 is a repositioning drop, not a rupture drop. The family clarifies the OpenAI offer and finally provides three coherent entry points. Terra replaces GPT-5.5 without surprise, Luna opens a credible volume door, Sol installs a flagship that speaks explicitly to sensitive verticals. For a pro managing an AI budget, the choice now happens workload by workload, no longer by default on “the latest released model”.

Our short-term advice: if you already run GPT-5.5 with calibrated prompts, switch to Terra without rewriting. If you operate a high-throughput volume pipeline, run Luna in parallel with your Grok 4.5 runs to compare quality per dollar spent. If your workloads touch biology, chemistry or cybersecurity, Sol deserves a serious pilot. The rest of the time, keep Sol on the shelf and pay Terra or Luna at the fair value of the actual need.

Follow the story on Horizon.

Comments

No comments yet. Why don’t you start the discussion?

    Leave a Reply

    Your email address will not be published. Required fields are marked *