On August 10, OpenAI made GPT-5.6 Luna — a 1-million-token-context model — the default for every ChatGPT user, free tier included, with no cap on text conversations. On August 16, starting at 16:00 UTC, DeepSeek raises the output price on its V4-Flash model from $0.28 to $1.32 per million tokens at peak hours — a 4.7x increase — and scales V4-Pro output from $0.87 to $3.96 at peak. These two moves, announced six days apart, are being reported as independent news cycles. They are not. Both companies are making rational business decisions that expose the same thing the AI industry rarely says directly: the proposition that AI inference is becoming structurally free is a market phenomenon, not a physics phenomenon. The subsidies behind it are sorting themselves out in real time.
Defending 51 Percent
ChatGPT held roughly 79 percent of the AI chatbot market in mid-2025. By July 2026, FirstPageSage puts it at 51.3 percent, with Google Gemini at 27.7 and Claude at 10.3. The 2026 monthly data shows the erosion was not gradual: January 67.1, February 65.7, March 62.6, April 61.8, May 59.1, June 58.6, July 51.3. Something accelerated in the second half of the year. The broader market offered no relief: July 2026 was the first month in which the overall AI chatbot market contracted in total users — every major platform combined — for the first time since ChatGPT launched. OpenAI is defending a shrinking lead in a market that just went negative.
The Luna move is a direct answer to that position. On July 30, three weeks after Luna launched, OpenAI cut Luna’s API price 80%, taking it to $0.20 per million input tokens and $1.20 per million output — roughly what a developer at mid-tier scale can build against without watching the meter. Then, in stages through the week of August 6 to 10, Luna became the default model for every free and Go-tier ChatGPT user, with unlimited text conversations and a Think button for extended reasoning. The reasoning toggle for paid subscribers got a depth slider. Free users got a button. But both got a model released this July rather than the older GPT-5.5 Instant that had been their ceiling.
The cost of that gift is real. ChatGPT’s reported weekly active user count sits at roughly one billion; giving a billion people access to current-generation inference is not costless. What makes it rational is that the alternative — watching 28 more percentage points of market share drift toward Gemini and Claude — is more expensive still. My read is that this is not a consumer-access initiative. It is a customer acquisition campaign funded by enterprise revenue, with a per-user marginal cost low enough that OpenAI can run it indefinitely as long as enterprise subscriptions hold. The free tier is where OpenAI fights for habit formation. The enterprise contract is where it captures value. The distinction matters because framing this as OpenAI making AI free implies the economics changed. They did not. The subsidy just got more visible.

The Subsidy Withdrawal
When DeepSeek-R1 shipped in January 2025, the prevailing reaction was that it proved frontier-quality AI could be served near-costlessly. That interpretation was always wrong, but it was encouraged by DeepSeek’s pricing: V4-Flash at $0.28 per million output tokens, V4-Pro at $0.87 — both undercutting every major Western competitor by a wide margin, and the gap was large enough that the working assumption among developers who had switched was that DeepSeek had discovered something structural about inference efficiency.
Starting at 16:00 UTC tomorrow, that assumption gets a data point it cannot survive. V4-Flash output rises to $1.32 per million tokens during peak hours — defined as 01:00 to 04:00 UTC and 06:00 to 10:00 UTC — and $0.66 during off-peak. V4-Pro goes to $3.96 at peak and $1.98 off-peak. Cache-hit tokens, the cheapest tier, face the steepest relative increase: V4-Pro cache-hit input at peak goes from $0.003625 per million to $0.044 — 12x higher, which is where the coverage figure of “up to 1,100%” comes from. That specific number applies to the narrowest input category at the most congested hours. The broader picture for a typical developer is a 4-to-5x increase in output token cost, with DeepSeek introducing peak and off-peak tiers simultaneously.
DeepSeek’s stated reason: “to allocate resources more reasonably.” That phrase explains nothing about the direction. The evidence does. Bloomberg reported in July that DeepSeek has begun IPO preparations targeting a valuation of roughly $71 billion, with a potential filing as early as this year and a target private raise of at least 10 billion yuan ahead of it. Founder Liang Wenfeng holds approximately 78 percent of the company. Before a listing at that valuation, the unit economics of serving API traffic need to look sustainable to investors, not subsidized. DeepSeek has not said the price hike is IPO-driven. The timing says it anyway. The evidence points to a connection, not a coincidence.
What the move reveals is this: the prices DeepSeek was charging before were not the product of a revolutionary cost structure. They were below what the infrastructure actually costs to run — either because DeepSeek chose to absorb the difference to drive developer adoption, or because the scale at which they were operating in the early months let them do so with the resources of a well-funded lab. Neither condition holds indefinitely. The adoption window that justified the subsidy has passed. The developer ecosystem is built. The IPO is on the horizon. This is what the exit from a subsidy looks like.
What the Counter-Argument Gets Right
The obvious pushback to this framing is that DeepSeek after the hike remains substantially cheaper than its Western competitors. That pushback is correct and worth stating plainly.

V4-Flash at $1.32 per million output tokens at peak is still 3.8x cheaper than Claude Haiku 4.5 at $5 per million output, which is itself Anthropic’s fastest and most affordable current model. V4-Pro at $3.96 peak output is 2.5x cheaper than Claude Sonnet 5 at $10 per million and 12.6x cheaper than Claude Fable 5 at $50 per million. The off-peak pricing makes the advantage more pronounced: a developer routing non-latency-sensitive work to DeepSeek off-peak pays $0.66 per million output on V4-Flash — 2.4x higher than the old price, but still cheaper than every Anthropic model currently available. The difference between pre-hike and post-hike DeepSeek is not the difference between cheap and expensive. It is the difference between two price points that both happen to dramatically undercut the Anthropic lineup.
What changed is not DeepSeek’s competitive position against Anthropic. What changed is the story about why DeepSeek was cheap. The story was that DeepSeek found a more efficient path. The better story, now visible, is that DeepSeek subsidized access, the subsidized period is ending, and the underlying cost floor — which applies to every lab running this class of infrastructure — has not moved. The market is not going to zero on inference. The market is pricing who absorbs the cost and for how long.
The Floor That Has Not Moved
AI inference costs have fallen dramatically: for comparable capability tiers, the per-token cost fell roughly 10x between early 2025 and mid-2026. That decline is real and will continue, though at a slower pace as software optimization opportunities become scarcer and further reductions require hardware advances on longer timelines. The physical constraints underneath inference — GPU memory bandwidth, power, cooling, networking — do not compress on the same schedule as software margins.
What this week makes visible is a distinction the market has been blurring: the decline in inference cost and the below-cost pricing of inference are not the same thing. OpenAI subsidizes the marginal user of Luna via enterprise revenue. DeepSeek subsidized API access via investor capital, then ended the subsidy as IPO math came into focus. Anthropic prices Fable 5 at $50 per million output tokens. The range between Luna’s free tier and Fable 5’s $50 per million doesn’t represent the range of inference costs. It represents the range of subsidy strategies across labs at different competitive positions, with different pressure to charge what the infrastructure actually requires.
For developers, the practical implication is clear: any architecture built on a single provider’s subsidized pricing is exposed when the subsidy window closes. DeepSeek warned developers on August 6 that changes were coming, with no specific amounts or effective date disclosed. Eight days later, the specifics arrived. A product built entirely around V4-Flash at $0.28 per million output tokens is now a product facing $1.32 at peak, with a week’s warning. Neither OpenAI nor DeepSeek is the villain in this story — both are making rational decisions with their capital and their strategic calendars. But the lesson is not one either company will say out loud: the current pricing landscape is a temporary arrangement of subsidies, not a settled equilibrium, and building on any single provider’s current price as if it reflects the long-run cost is a planning error. Both moves this week make the same point. The floor of AI inference has not moved. The conversation about who is standing on it has.

AI-generated editorial illustration · TemperatureZero · August 15, 2026
Keep reading the signal
Get the Daily Signal — a concise briefing on what actually matters in AI and the systems around it.
Subscribe FreeContinue the archive