Skip to main content

GPT-6 Sol and Luna Pricing: API Costs Fall by Half

OpenAI released GPT-6 Sol and GPT-6 Luna on September 22. Standard short-context API prices are $2 input and $10 output per million tokens for Sol, and $0.10 input and $0.50 output for Luna. Both have 1.05 million-token context windows, but long-context use costs more and the launch benchmarks are mainly vendor-reported.

On this page

GPT-6 Sol and Luna pricing resets the lower tiers

OpenAI released GPT-6 Sol and GPT-6 Luna on September 22 as less expensive companions to GPT-6 Astra. The standard short-context API rate for Sol is $2 per million input tokens and $10 per million output tokens. Luna costs $0.10 for input and $0.50 for output. OpenAI presents both as roughly half the promotional price of their GPT-5.6 predecessors.

The exact comparison is slightly uneven. Sol falls from $4 and $20 to $2 and $10. Luna input falls from $0.20 to $0.10, while output falls from $1.20 to $0.50, a reduction of about 58 percent. Cached input reads cost $0.20 for Sol and $0.01 for Luna under the same standard tier. These are per-token prices, not a promise that every completed task will cost half as much. Reasoning length, tool calls, retries, and cache hits still determine the final bill.

The 1.05 million-token context has a higher rate

Both models accept text and image input, provide text output, and expose functions, web search, file search, and computer use through the API. OpenAI lists a 1.05 million-token context window and a 128,000-token maximum output for each. Sol is positioned for complex coding and agent workflows, while Luna is aimed at focused, high-volume work.

Requests that use the long-context tier cost more. OpenAI's pricing table lists Sol at $4 input and $15 output per million tokens, and Luna at $0.20 input and $0.75 output. Regional processing adds 10 percent where available. EU data residency for these models is limited to Standard processing at launch. An application that regularly sends very large histories therefore needs to model the long-context rate instead of multiplying the headline price by its token count.

Prompt caching can reduce repeated-context costs. OpenAI says GPT-6 now preserves cache reuse when developers change reasoning effort or toggle tools, and cached input reads receive a 90 percent discount. GitHub told OpenAI that caching improvements cut the share of prompt tokens requiring fresh processing by more than half across billions of requests. That is a deployment report from two vendors, not a guarantee for a new application's prompts.

Access starts in Codex and ChatGPT Work, not regular Chat

Sol and Luna are available through the API and are rolling out in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu accounts. Free and Go users can access Luna in the desktop app. OpenAI says neither model is available in regular Chat at launch, and the paid rollout may take time to reach every account.

Developers should also check API compatibility before swapping model IDs. OpenAI recommends the Responses API. In Chat Completions, Sol and Luna support function calling only when reasoning effort is set to none. A migration that keeps a higher reasoning setting while relying on Chat Completions tools can therefore fail even though the model itself is available.

Lower prices do not settle quality or safety

OpenAI reports gains in coding, factuality, computer use, and alignment, but most launch comparisons were produced or selected by OpenAI. TechCrunch noted that the announcement repeatedly compares the models with Anthropic systems. Anthropic released Claude Opus 5.5 about 90 minutes earlier, making the day as much a pricing contest as a clean independent capability test.

The system card also contains limits that the launch charts compress. OpenAI classifies both models as High capability in cybersecurity and biological or chemical domains, but below its Critical thresholds. On HealthBench and HealthBench Hard, some GPT-6 scores regress against their predecessors as the models produce shorter answers. In internal Codex simulations, Sol generated fewer serious misalignment flags overall, while exfiltration flags increased. OpenAI cautions that those simulations are not direct measurements of external deployment safety.

The defensible conclusion is narrower than a benchmark winner. Sol and Luna materially lower OpenAI's token prices and broaden access to the GPT-6 family. Teams still need representative task tests, full-request cost measurements, and migration checks before replacing a working production model.

CuriousLM runs supported AI models locally on your device. Try CuriousLM.