ThinkPatternGet the app
Perspective
TECHNOLOGY · JUL 21, 2026

The iPhone Apple Is Building to Shrink the Cloud

Apple is re-engineering the iPhone's power layer for on-device AI — bigger batteries, a dual-cell design, model compression — just as model architectures and cloud economics converge toward local inference from opposite directions.

The iPhone has used a single battery for seventeen years. That streak ends with the iPhone Ultra. Code in the iOS 27 beta reveals a dual-cell architecture — a 1,921mAh cell paired with a 2,962mAh cell, totaling 4,883mAh — the first multi-battery design in the device's history [1]. The iPhone 18 Pro, arriving the same cycle, will carry the largest batteries Apple has ever put in a phone, up to 5,567mAh, and the company's own language is unusually direct about why: they are "designed to meet the power demands of on-device artificial intelligence in iOS 27" [2]. The silicon tells the same story. A 2nm A20 Pro chip. Twelve gigabytes of RAM. And PrismML, a compression technology Apple is evaluating that shrinks 27-billion-parameter models from 54 gigabytes to under 4 — small enough to run on an iPhone without a cloud call [3]. PrismML's CEO has been blunt about what matters: the intelligence has to be local and it has to be fast [3]. Craig Federighi, Apple's software chief, has drawn the line even more sharply.

Many of the existing chatbots, they're really focused on engagement to a large degree and sycophancy, right? — Craig Federighi

Apple is spending silicon, battery chemistry, and compression research to make on-device inference a physical reality — and two independent forces, neither of which Apple controls, are now moving in the same direction. The first is happening inside the models themselves. Google DeepMind's DiffusionGemma, released in June, is a 26-billion-parameter model that activates only 3.8 billion parameters per inference step — a design that generates text four times faster and was built with a specific use case in mind [4].

The model iteratively refines its own output, allowing it to evaluate the entire text block at once to fix mistakes in real-time. — Brendan O'donoghue

That 3.8-billion active parameter count closely matches the ~3-billion-parameter on-device Foundation Model Apple already runs on iPhones with a Neural Engine [5]. The model layer is not waiting for Apple's hardware to catch up. It is independently producing architectures — sparse activation, diffusion-based generation, smaller active footprints — that make local inference faster and cheaper by design. This is not a trend Apple created. It is a trend Apple's hardware investment is positioned to capture. The second force is happening on the other side of the equation, in the data centers that cloud AI depends on, and it is less benign. Anthropic reduced Claude session limits during peak hours in March because its compute capacity could not meet demand, throttling roughly 7% of Pro users [6]. OpenAI's adjusted gross margin fell from 40% in 2024 to 33% in 2025 as inference expenses quadrupled, and the company cut its infrastructure spending target from $1.4 trillion to $600 billion [7][8]. Power grid bottlenecks are slowing data center expansion across the US [9]. Google is facing local protests in Alabama, Virginia, and Iowa over water consumption at its data centers [10]. The physical layer of cloud AI — electricity, water, grid capacity, cooling — is becoming a constraint at the same moment Apple is solving the physical layer of on-device AI with bigger batteries and more efficient silicon. These two tracks are not causally linked. DiffusionGemma's architecture was not designed because Anthropic is throttling Claude. But they are converging. Model design is making local inference more capable. Cloud infrastructure is making remote inference more expensive and less reliable. The result is that on-device inference looks less like a choice and more like where the math and the economics are both pointing. What does this convergence mean for Apple's position? It shrinks the cloud surface area the company needs to a residual — work that genuinely cannot run on a phone — and makes that residual a commodity any provider can fill. Apple has already demonstrated this. The redesigned Siri chatbot in iOS 27 runs on a custom Google Gemini model, not OpenAI's [11]. Apple asked Google to install dedicated servers inside Apple's own data centers to support it [12]. And the Extensions framework lets users route queries to Claude or ChatGPT as interchangeable options, with Apple taking a commission on the subscriptions [13]. OpenAI was the launch partner for Apple Intelligence. It is now one interchangeable provider among several, and not even the primary one. The counter-evidence is real and worth stating plainly. Apple still needs the cloud. Its own server capacity is, by one account, underpowered for an AI-driven Siri, with only 10% of Private Cloud Compute capacity in use [12]. The most advanced Siri features — on-screen awareness, real-time camera integration, multi-step reasoning — require hardware newer than the iPhone 15 and still lean on cloud infrastructure for complex conversational tasks [14]. And OpenAI is not standing still: it launched GPT-5.4 Mini and Nano models explicitly targeting edge-based systems and the workloads where latency matters most [15].

These models are built for the kinds of workloads where latency directly shapes the product experience: coding assistants that need to feel responsive, subagents that quickly complete supporting tasks, computer-using systems that capture and interpret screenshots, and multimodal applications that can reason over images in real-time. — OpenAI

But the direction of travel is what matters. Every hardware generation expands the on-device partition. The iPhone 18 Pro's 5,567mAh battery and 2nm chip will run more AI locally than the iPhone 17 could. The iPhone Ultra's dual-cell architecture will run more still. The 2027 smart glasses — display-free, designed for voice- and vision-assisted interfaces — will by definition require local inference, because a device with no screen cannot be a cloud-chat terminal [16]. Each new device Apple ships contracts the cloud partition a little further. OpenAI is betting in the opposite direction. It raised $122 billion at an $852 billion valuation, with Amazon's $50 billion partially contingent on achieving AGI or going public, and it is explicitly pivoting from a research lab toward an infrastructure provider [8]. It unveiled Jalapeño, a custom inference chip developed with Broadcom, as the first phase of a roadmap toward gigawatt-scale data centers by 2029 [17]. Greg Brockman put the ambition plainly.

By designing more of the stack ourselves, we can serve more intelligence with greater efficiency and keep pushing advanced AI toward broader access. — Greg Brockman

That is a $600 billion bet on cloud scale — on the proposition that the future of AI runs in data centers someone else pays to build and cool and power. Meanwhile, inference costs are quadrupling, margins are falling, and DeepSeek is cutting model prices by 75%, making its V4-Pro 12 to 19 times cheaper than GPT-5.5 for equivalent tasks [18]. The model layer is being commoditized from below at the same moment the infrastructure layer is getting squeezed from above. Apple did not plan this convergence. It could not have known, when it began designing the A20 Pro or evaluating PrismML, that DiffusionGemma would ship with a 3.8-billion-parameter active footprint or that Anthropic would be throttling users during peak hours. But the company is now the beneficiary of a coincidence neither side orchestrated: model architectures and cloud economics are both tilting toward the device in your pocket, and Apple is the only company investing billions to make that device ready for the load. OpenAI is investing billions to own the side of the equation that is contracting. Apple is investing billions to own the side that is expanding. The convergence is not an inevitability — but it is not an accident, either.


Sources
  1. 1. iOS 27 Beta Code Reveals Apple iPhone Ultra Multi-Battery System
  2. 2. Apple Leaks Reveal Massive Battery Boost for iPhone 18 Pro
  3. 3. Apple Evaluates PrismML Tech for On-Device iPhone AI
  4. 4. Google DeepMind Releases DiffusionGemma for High-Speed Text Generation
  5. 5. Apple Unveils iOS 27 and Beta Hints at Foldable Devices
  6. 6. Anthropic Reduces Claude Session Limits During Peak Hours
  7. 7. OpenAI Cuts Infrastructure Spend to $600 Billion Ahead of IPO
  8. 8. OpenAI Raises $122 Billion at $852 Billion Valuation
  9. 9. AI Data Center Expansion Hits Power Grid Bottlenecks
  10. 10. Google Faces Local Backlash Over Water and Power Use
  11. 11. Apple to Launch Redesigned Siri Chatbot for iOS 27
  12. 12. Apple Asks Google to Install Servers for Next-Gen Siri
  13. 13. Apple Opens Siri to Third-Party AI in iOS 27
  14. 14. Apple Introduces iOS 27 With Advanced Siri AI
  15. 15. OpenAI Launches GPT-5.4 Mini and Nano AI Models
  16. 16. Apple Develops AI Smart Glasses for 2027 Release
  17. 17. OpenAI and Broadcom Unveil Jalapeño Custom AI Inference Chip
  18. 18. DeepSeek Permanently Cuts V4-Pro AI Model Prices by 75%

Keep reading in the app

The full perspective, free in the app.

Download on the App StoreComing soonGoogle Play