GPT-5.6 Sol, Terra, and Luna: which tier to use for agent work

Use Sol for multi-agent reasoning and long context, Terra as the default for coding at 40% of the cost, and Luna only for short, high-volume tasks.

Use GPT-5.6 Sol for multi-agent orchestration, complex reasoning, and retrieval over long contexts. Use GPT-5.6 Terra as the primary workhorse for scoped coding, code review, and general agentic pipelines at 40% of Sol's per-token cost. Use GPT-5.6 Luna strictly for short, single-step triage and classification where context is narrow and token volume dominates.

Key facts

  • Preview date: 2026-06-26. General availability: 2026-07-09 across ChatGPT, Codex, and the OpenAI API.
  • Context window: 1,050,000 tokens across all three tiers; max output is 128,000 tokens.
  • Long-context billing threshold: Inputs above 272K tokens trigger the higher long-context rate across the entire request.
  • API pricing (standard short-context per 1M tokens): Sol launch price is $5.00 input / $30.00 output (with a 20%+ discount active 2026-08-21 through 2026-11-21); Terra is $2.00 input / $12.00 output (cut 20% on 2026-07-30 from $2.50 / $15.00); Luna is $0.20 input / $1.20 output (cut 80% on 2026-07-30 from $1.00 / $6.00).
  • Reasoning effort settings: Light, Medium, High, Extra High, Max, and Ultra.
  • Ultra availability: Exclusive to Sol in ChatGPT Work (Pro and Enterprise) and Codex (Plus and higher); developers can implement comparable multi-agent patterns via the Responses API multi-agent beta.

Tier

Input / 1M tokens

Output / 1M tokens

Notes

Sol

$5.00

$30.00

Launch rate; temporary 20%+ discount active 2026-08-21 through 2026-11-21

Terra

$2.00

$12.00

Cut 20% on 2026-07-30 (down from $2.50 / $15.00)

Luna

$0.20

$1.20

Cut 80% on 2026-07-30 (down from $1.00 / $6.00)

How GPT-5.6 Sol, Terra, and Luna perform on agent benchmarks

All scores from OpenAI's published benchmark tables (read 2026-10-06).

Benchmark

Sol

Sol Ultra

Terra

Luna

Terminal-Bench 2.1

88.8%

91.9%

87.4%

84.7%

Agents' Last Exam

52.7%

N/A

50.4%

50.3%

DeepSWE v1.1

72.7%

N/A

69.6%

67.2%

AA Coding Agent Index

80.0

N/A

77.4

74.6

BrowseComp

90.4%

92.2%

87.5%

83.3%

MRCR v2 (256K-512K)

91.5%

N/A

89.6%

41.3%

The Multi-hop Retrieval over Contextual Reasoning (MRCR) metric determines tier viability for context-heavy agent work. MRCR measures whether a model accurately retrieves and reasons over facts dispersed throughout long prompts. Luna's 41.3% retrieval rate on the 256K to 512K range falls far below GPT-5.5's prior 74.0% benchmark score on the same evaluation. While all three GPT-5.6 tiers provide a 1,050,000-token context window, Luna cannot reliably locate dispersed facts in prompts spanning more than a single document.

How GPT-5.6 Sol ultra mode works and when it matters

The max reasoning mode extends the duration Sol spends exploring alternatives, checking intermediate outputs, and revising plans before responding. The ultra setting builds on this by coordinating four Sol agents concurrently by default, aggregating their intermediate findings into a unified result.

On Terminal-Bench 2.1, running Sol in the four-agent Ultra configuration raises task completion from 88.8% to 91.9% while reducing wall-clock latency through parallel tool execution. On BrowseComp, a 16-agent Ultra configuration attains 92.2%. In the OpenAI API, developers can implement the same multi-agent workflow using the multi-agent beta in the Responses API without requiring a ChatGPT Work Pro license.

Ultra mode is unavailable on Terra and Luna. At lower reasoning levels, subagent delegation must be explicitly requested in the prompt or agent configuration. With Ultra enabled on Sol, ChatGPT Work delegates tasks to parallel subagents proactively when decomposition yields speed or accuracy gains.

Which GPT-5.6 tier fits which agent job

Use Sol when:

  • The task requires synthesizing information across multiple documents, analyzing large code repositories, or executing multi-hop retrieval over long contexts (where Sol achieves 91.5% on MRCR versus Luna's 41.3%).
  • The pipeline requires max or ultra reasoning, such as automated vulnerability discovery, multi-step system refactoring, or autonomous long-horizon research.
  • You are running an orchestrator agent that plans workflows and evaluates outputs from specialized subagents.

Use Terra when:

  • The task is bounded: scoped feature implementation, pull-request review, or initial document extraction up to several hundred pages.
  • Budget efficiency is required for daily development: Terra scores within 1.4 to 3.1 points of Sol across coding benchmarks (87.4% versus 88.8% on Terminal-Bench 2.1; 69.6% versus 72.7% on DeepSWE v1.1) at 40% of Sol's per-token cost.
  • Working in ChatGPT Work or Codex under Free or ChatGPT Go plans, where Terra serves as the default model tier.

Use Luna when:

  • The task consists of short, single-step operations: log triage, summarization, pull-request title generation, syntax and linting checks, or high-volume document classification where each document fits within a bounded window.
  • Cost minimization is paramount: at $0.20 input and $1.20 output per 1M tokens, Luna costs 25 times less than Sol at launch rates. On DeepSWE v1.1, Luna at max reasoning delivers approximately 24 benchmark points per API dollar.
  • Luna acts as a disposable filtering subagent, passing only complex anomalies up to Terra or Sol.

Avoid Luna when:

  • The task relies on recalling information positioned early in a long context, cross-referencing multi-file codebases, or complex multi-hop question answering. Routing long-context jobs to Terra prevents retrieval failures that can silently corrupt downstream agent steps.

Sources

Last verified: 2026-10-06.

Spotted an outdated or wrong claim? Agents can report it with evidence throughPOST /api/feedback; an editor checks every report. See llms.txt for the agent API.