Microsoft launches Maia 200 AI chip to reduce reliance on Nvidia

The most expensive part of the AI boom isn’t training the model. It’s the millionth time someone asks the AI a question, and the system has to answer instantly, again and again, without the cost exploding.
On January 26, Microsoft announced the second generation of its in-house AI chip, Maia 200. The hardware launch is paired with a software push that aims to take on one of Nvidia’s biggest strengths with developers.
Why Microsoft built Maia 200 for inference first
Maia 200 is an inference accelerator designed to make the economics of token generation more efficient, the step where models produce outputs in real time. It comes online this week at a data center in Iowa, with Arizona slated next.

Microsoft is tying the chip directly to its most prominent products, including Microsoft Foundry and Microsoft 365 Copilot. The clearest validation is the model lineup: the company says Maia 200 will host OpenAI’s latest GPT-5.2 models.
The specs Microsoft keeps repeating
Microsoft has been unusually open about the specs of the chip and how it compares to the previous generation of chips. Here are the most important ones:
Built on TSMC’s 3nm process with over 140 billion transistors per chip
216GB of HBM3e memory at 7 TB/s, plus 272MB of on-chip SRAM
Over 10 petaFLOPS of FP4 and over 5 petaFLOPS of FP8 performance
750W SoC TDP envelope
Claimed 30% better performance per dollar than the latest hardware in Microsoft’s fleet
Microsoft also claims Maia 200 delivers three times the FP4 performance of AWS’s Trainium (Gen 3) and FP8 performance that tops Google’s TPU v7.

The real fight is software, not silicon
The chip announcement comes with a software play aimed at Nvidia’s CUDA advantage. Maia 200 ships with a package that includes Triton, an open-source tool with major contributions from OpenAI that handles the same kind of low-level optimization Nvidia’s CUDA is known for.
The strategy is simple. CUDA locks a lot of performance work to Nvidia hardware, while Triton is meant to make optimized code more portable across different accelerators, reducing the risk of betting on a single chip provider forever.
SRAM and Ethernet are the hidden design tells
The memory and networking choices indicate the workload Microsoft is targeting first: large-scale, high-concurrency inference.
Architecturally, Maia 200 pairs high-bandwidth memory with a large on-chip SRAM pool, even as it uses an older, slower HBM generation than Nvidia’s forthcoming Vera Rubin chips.
The “so what” is latency under load. SRAM is significantly faster than HBM, so keeping frequently reused data closer to the compute can reduce stalls. This helps chat systems respond consistently when a large number of users hit the model at once.
How Maia 200 fits into the broader chip race
This launch lands in the middle of a bigger shift among hyperscalers. Microsoft, Google, and AWS are still major Nvidia customers, but each is trying to control AI costs and supply by designing more of the stack in-house.
This pits Microsoft even more directly against Nvidia’s full stack, especially as it works to bridge the software gap and roll its own silicon into production data centers.
The next question is who gets Maia 200 beyond Microsoft’s own marquee products. Bloomberg reported the chip announcement as part of Microsoft’s effort to reduce reliance on Nvidia, but broader Azure availability is still a moving target.
For now, that uncertainty is part of the point. Maia 200 looks like a margin weapon inside Microsoft’s own stack first, where every Copilot response served on first-party silicon is one less inference bill tied to someone else’s ecosystem.
Y. Anush Reddy is a contributor to this blog.



