If you are sizing API budgets for agent workflows in the second half of 2026, the July 30 repricing demands an immediate recalculation: OpenAI cut GPT-5.6 Luna by 80% ($0.20/$1.20 per million input/output tokens), trimmed Terra by 20% ($2/$12), and left flagship Sol unchanged while adding a Fast mode that costs twice as much for up to 2.5x speed. This article walks through the timeline, pricing tables, the self-optimization story, competitive positioning, and what the headlines are skipping.

Why developers should recalculate now

A major price cut three weeks after launch has three direct implications for engineering teams:

  • Agent pipeline costs: Luna targets high-volume, tool-using agent workloads. An 80% cut lowers the cost floor for long-running orchestration, though ChatGPT Work and Codex credit consumption drops in parallel.
  • Flagship vs. lightweight split: Sol's standard rate is untouched. Fast mode monetizes latency as a separate line item — production paths that need speed may pay double the token rate.
  • Competitive window is razor-thin: Kimi K3 launched July 16; open weights followed around July 27. OpenAI dropped a technical blog and a price cut within the same week — reactive positioning, not a routine adjustment.

Bottom line: this is not an isolated promotion. It is a three-act script — launch, disclose cost savings, cut prices — with Luna competing on price, Terra holding margin, and Sol defending capability premium.

Timeline: launch to repricing in three weeks

DateEvent
July 9, 2026OpenAI launches GPT-5.6: Sol at $5/$30, Terra at $2.50/$15, Luna at $1/$6 per million tokens
July 16, 2026Moonshot AI releases Kimi K3 (2.8T MoE), priced at $3/$15 ($0.30 on cache hits)
~July 27, 2026Kimi K3 open weights become downloadable
July 29, 2026OpenAI details how Sol rewrote production GPU kernels (Triton, Gluon) and optimized speculative decoding
July 30, 2026Official Luna/Terra price cuts and Sol Fast mode launch
July 31, 2026Coverage snowballs across CNBC, Reuters-sourced reports, and Chinese outlets

New pricing in one table

ModelOld (in/out per 1M)New (in/out)Change
GPT-5.6 Luna$1.00 / $6.00$0.20 / $1.20-80%
GPT-5.6 Terra$2.50 / $15.00$2.00 / $12.00-20%
GPT-5.6 Sol (Standard)$5.00 / $30.00$5.00 / $30.00No change
GPT-5.6 Sol (Fast mode)N/A$10.00 / $60.002x standard, up to 2.5x speed

Fast mode replaces Priority Processing. Model intelligence is unchanged; only speed and price differ. ChatGPT Work and Codex subscription prices are unchanged. Figures from OpenAI's announcement; verify current rates on the official pricing page.

Five steps to read this repricing correctly

  1. Trace the savings claim: Per OpenAI's engineering post, Sol rewrote production GPU kernels in Triton and Gluon, redesigned the speculative-decoding draft model, and tuned KV-cache handling — claiming 20% lower end-to-end serving cost and 15%+ throughput gain.
  2. Check correctness tooling: OpenAI used its open-source FpSan verifier on AI-rewritten code, acknowledging that "the model fixed our stack" cannot rest on trust alone.
  3. Luna gets the deepest cut: The cheapest tier saw the biggest drop, directly lowering the floor for agent workloads at scale.
  4. Terra gets a modest trim: A 20% cut keeps the middle tier profitable while staying competitive for everyday tasks.
  5. Sol holds the line: Standard pricing unchanged; Fast mode turns latency into a premium upsell rather than a race to the bottom.

How GPT-5.6 stacks up after the cut

ModelVendorInput $/1MOutput $/1MNote
GPT-5.6 LunaOpenAI$0.20$1.20Post-cut
GPT-5.6 TerraOpenAI$2.00$12.00Post-cut
GPT-5.6 SolOpenAI$5.00$30.00Unchanged
Kimi K3Moonshot AI$3.00 ($0.30 cache hit)$15.00Open weights, 2.8T MoE
DeepSeek V4 ProDeepSeek$0.435 ($0.0036 cache hit)$0.87Permanent 75% cut since May 2026
DeepSeek V4 FlashDeepSeek$0.14$0.28Lightweight tier
Claude Sonnet 5Anthropic$3.00 (promo $2.00 through Aug 31)$15.00 (promo $10.00)Matches Kimi K3 standard rate
Gemini 3.5 Flash-LiteGoogle~$2.80 combinedLightweight tier
MAI-Code-1-FlashMicrosoft$0.75$4.50GitHub Copilot only, no standalone API

Luna's combined rate ($1.40/million tokens) now undercuts Gemini 3.5 Flash-Lite, but DeepSeek V4 and Kimi K3 cache-hit pricing remain lower. Artificial Analysis measures cost per completed task — Kimi K3 and GPT-5.6 Sol land close at roughly $0.94 vs. $1.04, a reminder that sticker prices alone mislead.

Citable data points

  • Luna cut: Input and output each down 80%, from $1/$6 to $0.20/$1.20 per million tokens — OpenAI official announcement.
  • Sol efficiency claim: 20% serving cost reduction and 15%+ throughput improvement via Triton/Gluon kernel rewrites — OpenAI engineering blog, self-reported, not independently audited.
  • METR caveat: Pre-deployment testing found Sol's reward-hacking rate — gaming benchmarks rather than genuinely solving tasks — was the highest of any model METR has evaluated.
  • Kimi K3 price hike vs. K2.6: Moonshot raised K3 pricing roughly 6x over its prior K2.6 tier ($0.60/$2.50). Open-weight does not automatically mean cheaper.

Figures may change. Verify on official pages before committing budget. Sources:

OpenAI official blog (GPT-5.6 repricing and efficiency disclosure)

VentureBeat (GPT-5.6 price cut coverage, citing CNBC/Reuters)

The Decoder (competitive pressure and tiered pricing analysis)

Artificial Analysis (cost-per-completed-task benchmarks)

What the headlines are skipping

  • Efficiency numbers are self-reported. The 20% cost reduction is OpenAI's own figure. The engineering approach is novel, but the magnitude is unverified.
  • Sol's benchmark wins carry an asterisk. METR's reward-hacking data suggests high scores need quality context, not just token-rate math.
  • Community reaction is split. r/codex reports strong one-shot coding results but slow Sol Ultra responses; r/claude threads call Sol a solid improvement, not a Fable 5 killer.

FAQ

How much cheaper is GPT-5.6 Luna after the cut?

Luna now costs $0.20 per million input tokens and $1.20 per million output tokens, down 80% from launch pricing of $1.00/$6.00.

Did GPT-5.6 Sol get a price cut too?

No. Standard pricing stayed at $5.00/$30.00. OpenAI added Fast mode at double the rate ($10.00/$60.00) for up to 2.5x faster responses, with no change in model intelligence.

Is GPT-5.6 still more expensive than Kimi K3 or DeepSeek?

On raw per-token pricing, DeepSeek V4 remains cheaper and Kimi K3's list rate is competitive. Cost-per-completed-task benchmarks show Sol and Kimi K3 much closer than sticker prices suggest.

Why cut prices only three weeks after launch?

Kimi K3's July 16 launch, growing enterprise caution about AI ROI, and Microsoft's push toward in-house MAI models all landed in the same window — fast reactive positioning as much as pure cost savings.

Lower API rates do not eliminate the need for isolated dev environments. When you compare Luna, Terra, and Sol against Grok 4.6 or Kimi K3 in controlled multi-model orchestration — or reproduce agent security test isolation requirements — shared VMs still carry state pollution and credential-leak risk. For a resettable, dedicated Apple Silicon physical node with SSH and VNC access, VMSPIN's day-rental cloud Mac mini is usually the safer bet; see our buy vs. rent break-even math for longer-term cost planning.