Skip to main content

The AI price war went nuclear this week: DeepSeek V4 Flash at $0.14/M tokens and GPT-5.6 Luna down 80% — in the same 24 hours

Disclosure: this article links to some providers through our referral links. If you sign up through one, we may earn credit or a commission. It never changes what we recommend — the data says what it says.

If you blinked this week you missed the cheapest frontier-class AI has ever been. On July 30, OpenAI cut GPT-5.6 Luna prices by 80%. Less than a day later, DeepSeek's V4 Flash graduated from preview to an official public beta — the '0731' build — at prices starting from $0.14 per million input tokens. We maintain a structured catalog of 42 AI providers with real prices, free tiers, and rate limits, and weeks like this are exactly why. Here's what changed, in numbers, and what it means if you spend money on tokens.

What OpenAI did on July 30

GPT-5.6 Luna dropped from $1.00 to $0.20 per million input tokens and from $6.00 to $1.20 per million output tokens — an 80% cut on both sides, effective immediately. GPT-5.6 Terra got a smaller 20% cut, and the flagship Sol kept its pricing but gained a new Fast mode. The backstory is unusually good: per reporting around the announcement, Sol autonomously rewrote parts of its own production inference stack (GPU kernels and speculative-decoding paths), and the efficiency gains funded the price drop. Watch the fine print if you push long contexts — prompts over ~272K input tokens bill at 2x input and 1.5x output, and cache writes run 1.25x the uncached rate.

What DeepSeek did on July 31

DeepSeek V4 Flash is now an official public-beta release. The architecture is unchanged from the preview — this is DeepSeek's efficiency model, not a new flagship — but the pricing is the point: third-party trackers list it from $0.14 input / $0.28 output per million tokens, which undercuts even the newly-cut Luna on the input side. One thing to watch: DeepSeek has announced a peak-hours pricing policy that is not yet active. If you build on V4 Flash, build with that in mind. We've added V4 Flash to our DeepSeek provider entry today.

DateProviderMoveNew price
Jul 9–10MetaMuse Spark 1.1 — Meta's first paid model API, then the cheapest paid frontier$1.25 / $4.25
Jul 12OpenAIGPT-5.6 Sol usage limits temporarily relaxed
Jul 30OpenAIGPT-5.6 Luna cut 80%; Terra cut 20%; Sol gains Fast mode$0.20 / $1.20 (Luna)
Jul 31DeepSeekV4 Flash '0731' official public betafrom $0.14 / $0.28
The last three weeks of the price war, as tracked in our catalog and public announcements. Prices are per million tokens (input/output).

What this means if you buy tokens

Three practical consequences. First: if you pinned your costs to Luna at $1/$6 in a budget spreadsheet, your projections are now wrong by 5x in your favor — re-run them. Second: the 'which provider is cheapest' question no longer has a stable answer measured in months; it now moves in weeks, which is why we keep the provider directory as data rather than as a blog post someone wrote once. Third: the floor is now low enough that the free tiers matter differently — a $0.14/M model is close enough to free that rate limits, not price, become the constraint.

If you'd rather not re-platform every time the market sneezes, this is the exact case for bring-your-own-key tooling: keep your accounts at multiple providers, and switch models per task. Our own /ai-hub/chat runs entirely in your browser with your keys — you can put DeepSeek in one column and anything else in the other and compare them on your actual workload today, not on someone's benchmark. The free-tier roundup covers what you can use without a card at all.