OpenAI Cuts GPT-5.6 Sol Pricing, Brings Terra and Luna to AWS GovCloud

LLMsDeveloper Tools
Illustration generated by AI: Editorial image for OpenAI Cuts GPT-5.6 Sol Pricing, Brings Terra and Luna to AWS GovCloud

The Core · TL;DR

  • GPT-5.6 Sol's Bedrock pricing drops to $4/M input and $20/M output tokens, cuts of 20% and 33.3%.
  • Promotional pricing holds at least through November 21, 2026.
  • GPT-5.6 Terra and Luna are now generally available on Amazon Bedrock in AWS GovCloud (US-West and US-East).
  • Both support 1M-token context windows and prompt caching with a 90% discount on repeated context.

OpenAI's GPT-5.6 Sol just got substantially cheaper on Amazon Bedrock. Input tokens now run $4 per million, output tokens $20 per million, a drop of 20% and 33.3% respectively from prior rates, according to AWS's own announcement.

The promotional pricing is confirmed to run at least through November 21, 2026, though AWS has not stated whether it reverts afterward or becomes permanent. For teams running high-volume agentic workloads, that gap matters: Sol is positioned as OpenAI's top performer on agentic coding benchmarks, and cheaper output tokens directly lower the cost of long tool-calling chains.

Alongside the pricing news, two other models in the GPT-5.6 family, Terra and Luna, are now generally available on Amazon Bedrock within AWS GovCloud (US-West) and AWS GovCloud (US-East). That's a notable expansion into regulated and government-adjacent environments, where compliance requirements often keep newer frontier models off the table for months after commercial release.

What Terra and Luna Bring

Terra is pitched as a value play: AWS describes it as delivering performance on par with GPT-5.5 at half the cost, effectively pushing the older model's capability down a price tier. Luna sits below it, aimed at high-throughput, latency-sensitive use cases where AWS says it offers the lowest price point in the lineup.

Both models support 1 million token context windows on Bedrock, matching the long-context ceiling increasingly expected from frontier-tier releases. That's paired with prompt caching using explicit cache breakpoints, letting repeated context in a conversation or agent loop get billed at a 90% discount rather than full price on every call.

Taken together, the caching discount and the GovCloud rollout point toward the same audience: engineering teams running sustained, high-volume agentic or RAG pipelines who need predictable costs and, in the GovCloud case, a compliance-cleared deployment path. None of the three models change OpenAI's underlying capabilities so much as make existing capability cheaper to run at scale, and cheaper to deploy in environments that previously lacked access altogether.

WK

WAKIB Editorial Team

This review was prepared and summarized by the WAKIB AI intelligence engine and vetted by our editorial board for accuracy and reliability.

Subscribe to Newsletter

Get a weekly summary of the most promising AI research and tools delivered to your inbox.

Telegram Channel

Join our active community on Telegram for real-time tracking of AI models and trends.

Join us on Telegram