The AI Subscription Model Is Splitting in Two
Basic chat is now cheap enough to give away, but autonomous agents burn tokens so fast that flat-rate pricing has become a guaranteed loss — and every major company is converging on the same answer.
In August 2026, OpenAI made two pricing decisions that point in opposite directions. It removed text chat rate limits for free users — basic conversation, effectively unlimited, at zero cost. At the same time, it kept strict caps on Codex, its agentic coding tool, for paying subscribers — caps it had introduced in April with a new $100-per-month tier aimed at heavy users. [1][2] One company, one month, two contradictory answers to the same question: what should AI cost? The contradiction is not confusion. It is a fault line that runs through the entire industry, and it explains why the AI subscription model is not dying — it is splitting. Basic chat has become cheap enough to give away. Consumer spending on ChatGPT subscriptions reached $3 billion by the end of 2025, a 408 percent increase over the previous year, driven almost entirely by flat-rate plans. [3] For the kind of usage those subscribers generate — questions, drafts, summaries — the subscription model still works. Autonomous agents are different. They do not answer a question and stop. They plan, execute, iterate, and call other tools, burning tokens with every step. Nvidia Vice President Bryan Catanzaro put the problem plainly.
For my team, the cost of compute is far beyond the costs of the employees. — Bryan Catanzaro
The numbers bear him out. Uber exhausted its entire 2026 AI coding-tool budget within four months. Microsoft cancelled thousands of internal Claude Code licenses in its Experiences and Devices group and redirected engineers to GitHub Copilot CLI to contain the damage. [4] An agent does not cost marginally more than a chatbot. It costs orders of magnitude more — and a flat monthly fee cannot absorb the difference. Every company that has crossed from assistant to agent is now converging on the same answer, and each arrived there with its own verb. GitHub moved first. In June 2026, it switched Copilot from a flat $29-per-month subscription to token-based AI Credits. Some power users saw their projected costs jump to between $750 and $3,000 a month. GitHub's chief product officer, Mario Rodriguez, was direct about why.
Today, a quick chat question and a multi-hour autonomous coding session can cost the user the same amount. — Mario Alberto Rodríguez
The flat-rate model, he said, was no longer sustainable because Copilot had evolved from an in-editor assistant into an agentic platform. GitHub had been absorbing the escalating inference costs behind that usage. It stopped. [5] Microsoft followed weeks later. In June 2026, it launched Copilot Cowork with pay-as-you-go billing at one cent per credit — the company's first departure from fixed subscription fees for office software in two decades. Charles Lamanna, the corporate vice president who oversees the product, called it a big evolution for us.
This is a big evolution for us ... which has been a user subscription-based business for so long, for really like two decades — Charles Lamanna
Over half the Fortune 500 adopted it in preview. The base subscription remained, but the agentic service was consumption-priced. [6] Anthropic took the hardest line. In March 2026, it throttled Claude subscription session limits during peak hours, with roughly 7 percent of Pro users hitting caps they had not encountered before. [7] In April, it blocked third-party tools like OpenClaw from accessing Claude Pro and Max subscriptions entirely. The company's statement was unambiguous.
We've been working hard to meet the increase in demand for Claude, and our subscriptions weren’t built for the usage patterns of these third-party tools. — Boris Chernyshov
Users who wanted to keep using agentic tools had to switch to pay-as-you-go API keys. [8] The mechanism driving all three decisions is what Wayne Liu, the chief growth officer of Perfect Corp America, calls the token paradox. Per-token inference costs keep falling — DeepSeek cut prices 75 percent in May, OpenAI and Anthropic slashed theirs in August to counter Chinese competitors [9][10] — but total consumption is rising far faster than unit costs are declining. Goldman Sachs forecasts monthly token consumption will increase 24-fold to 120 quadrillion tokens by 2030. [11] Efficiency gains, in this arithmetic, lead to higher total bills.
Cheap intelligence creates value only when it improves judgment. — Liu Ruilin
The financial consequences of this paradox are already visible in the spread between the companies that build AI models and the companies that supply the infrastructure to run them. Apollo Chief Economist Torsten Slok found that silicon and equipment companies maintain a 41 percent profit margin, while model developers operate at a negative 59 percent margin.
The bottom line is that the most profitable part of the AI value chain depends on the least profitable part continuing to grow revenue or raise capital. — Torsten Slok
The infrastructure layer, meanwhile, is building as if the subscription model were still intact. The combined backlogs of the top four hyperscalers have reached $2.3 trillion. Amazon alone raised its 2026 AI capital expenditure forecast to $220 billion, with CEO Andy Jassy saying demand still significantly outstrips supply and that capacity will remain constrained through 2027. [12] Brookfield Asset Management raised a record $77 billion in the second quarter of 2026, deploying aggressively into AI energy and digital infrastructure. [13] The credit market has begun to notice the gap between infrastructure spending and the revenue models meant to pay for it. Oracle's stock has fallen 59 percent from its September peak. The company carries $167 billion in debt and negative free cash flow. Bank of America estimates that OpenAI accounts for more than half of Oracle's AI backlog — meaning the infrastructure provider's solvency depends on a customer that loses money on every query. S&P downgraded Oracle to BBB- in August. [14] Consumption-based pricing solves a real problem for providers. It stops them from subsidizing the heaviest users — the ones whose agentic workloads were, in GitHub's phrase, being absorbed at a loss. But it does not answer the question that matters for the entire buildout: whether the spending is worth it for the customer. Greg Rowell, the CEO of Revenium, which launched a tool in March 2026 to help enterprises measure the return on their AI agent spending, framed the risk directly.
Revenium AI Outcomes changes that. — Greg Rowell
The infrastructure was financed on the assumption that subscription revenue would eventually cover the cost of compute. The new pricing model — metered, per-token, pay-as-you-go — lets providers stop the bleeding. It does not establish whether the spending generates returns. And on current trajectories, the revenue from how AI products are priced cannot cover the compute costs the infrastructure buildout was premised on. [15]
- 1. OpenAI Inc. Removes ChatGPT Text Limits and Updates Models
- 2. OpenAI Launches $100 ChatGPT Pro Tier for Developers
- 3. ChatGPT App Reaches $3 Billion in Consumer Spending
- 4. Microsoft and Uber Cut AI Tool Use Amid Rising Compute Costs
- 5. GitHub Copilot Switches to Token-Based AI Credit Billing
- 6. Microsoft Launches Copilot Cowork With Pay-As-You-Go Pricing
- 7. Anthropic Reduces Claude Session Limits During Peak Hours
- 8. Anthropic Blocks Claude Subscription Access for OpenClaw and Third-Party Tools
- 9. DeepSeek Permanently Cuts V4-Pro AI Model Prices by 75%
- 10. OpenAI and Anthropic Slash Prices to Counter Chinese AI
- 11. AI Firms Struggle With Unpredictable LLM Token Costs
- 12. Amazon Raises 2026 AI Spending Forecast to $220 Billion
- 13. Brookfield Asset Management Raises Record $77 Billion in Q2
- 14. Oracle Stock Plummets 59% Amid AI Debt Concerns
- 15. Revenium Launches AI Outcomes to Track Agent ROI