Microsoft Launches Maia 200 AI Chip to Cut Nvidia Reliance
Microsoft deployed the Maia 200 AI inference chip to lower operational costs and increase performance for Azure cloud services and OpenAI models.
On January 26, 2026, Microsoft unveiled the Maia 200, its second-generation in-house AI inference accelerator. Fabricated by TSMC using a 3-nanometer process, the chip features over 140 billion transistors and integrated SRAM to improve throughput and memory efficiency for large-scale reasoning models. The hardware is designed to reduce the company's reliance on Nvidia and lower the cost of AI token generation, offering a 30% improvement in performance per dollar compared to Microsoft's existing fleet.
Deployment began at a data center in Des Moines, Iowa, with subsequent expansion planned for the US West 3 region near Phoenix, Arizona. The Maia 200 will power Microsoft 365 Copilot, Microsoft Foundry, and OpenAI's GPT-5.2 models. To support the hardware, Microsoft introduced a new software development kit and programming tools, including the open-source Triton tool developed with contributions from OpenAI.
Microsoft executives claim the chip outperforms competitors, delivering three times the FP4 performance of Amazon's third-generation Trainium and exceeding the FP8 performance of Google's seventh-generation TPU. While the company has already begun designing a Maia 300 successor, CEO Satya Nadella stated that Microsoft will maintain a dual-track strategy, continuing to purchase hardware from Nvidia and AMD to meet growing computational demands. The launch coincides with a 40% year-on-year increase in Azure and other cloud services revenue for the first quarter of fiscal year 2026.