Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteAt Microsoft Ignite on November 15, 2023, Microsoft announced two internally designed Azure processors: Maia 100, an accelerator for large-scale AI training and inference, and Cobalt 100, a 64-bit Arm processor for general-purpose cloud computing. They were not retail chips or products that customers could install in their own servers. Microsoft designed them for its Azure datacenters and planned to expose them through cloud services and virtual machines.
The distinction matters: Maia is the AI-focused device, while Cobalt is a conventional server CPU. Together they represent Microsoft’s strategy of building a heterogeneous Azure fleet alongside Nvidia, AMD and other suppliers—not replacing every third-party processor.
The two-chip announcement at a glance
| Processor | Primary role | Typical workloads | How customers access it |
|---|---|---|---|
| Azure Maia 100 | AI accelerator | Large-model training and inference, including workloads behind Azure OpenAI, Bing, GitHub Copilot and ChatGPT | Mainly through Microsoft-managed Azure AI infrastructure; not a generally selectable retail accelerator SKU in the sources reviewed |
| Azure Cobalt 100 | 64-bit Arm CPU | Web and application servers, databases, analytics, caches, microservices and other scale-out workloads | Azure virtual-machine families such as Dpsv6, Dplsv6, Dpdsv6, Dpldsv6, Epsv6 and Epdsv6 |
Microsoft’s original announcement is documented in its Ignite 2023 Book of News. The announcement described deployment in Microsoft datacenters beginning in 2024, not sales of standalone processors.
Why Microsoft designed its own cloud silicon
AI demand increased the need for accelerators while supply, power and datacenter capacity became strategic constraints. Designing the processor, server board, rack, networking, cooling and software together gives a hyperscaler more control over those constraints.
#1 Best Overall
- NVIDIA Volta GV100 Architecture — 4,608 CUDA Cores, 640 1st-Gen Tensor Cores delivering 14 TFLOPS FP32 and 112 TFLOPS deep learning performance for AI training, inference, HPC, and scientific computing workloads
- 32GB HBM2 ECC Memory — 900 GB/s Bandwidth — High-bandwidth memory on a 4096-bit bus with ECC error correction provides the memory capacity and throughput required for the largest AI models, simulations, and datasets
- PCIe 3.0 x16 Interface — 250W TDP — Standard PCIe Gen3 connectivity with passive cooling designed for enterprise rack server deployment in HPE ProLiant, Dell PowerEdge, and Supermicro platforms with adequate chassis airflow
- NVLink — Scale to 96GB Unified Memory — Connect two V100 GPUs via NVLink at 300 GB/s bi-directional bandwidth to scale GPU memory from 32GB to 96GB for larger AI training and HPC workloads
- Multi-Precision Computing — Supports FP64 (7 TFLOPS), FP32 (14 TFLOPS), FP16 (112 TFLOPS) and INT8 precision modes for flexible deployment across training, inference, and scientific simulation workloads
- Workload specialization: Microsoft can optimize for services it operates at enormous scale, including Azure OpenAI and Copilot.
- Power and cooling: Rack-level design can improve energy use and thermal density rather than treating the chip as an isolated component.
- Supply planning: Internal silicon provides another source of capacity alongside external processors.
- Economics: A workload-specific design may improve performance per dollar, although Microsoft’s goal is not proof that it beats every Nvidia or AMD product.
- System control: Microsoft can coordinate firmware, compilers, networking and operations across the entire Azure fleet.
Microsoft has continued to describe custom silicon as complementary to industry partnerships. The strategy is diversification and vertical optimization, not an immediate abandonment of Nvidia or AMD.
Maia 100: Microsoft’s AI accelerator
What Maia was built to do
Maia 100 was Microsoft’s first in-house AI accelerator, designed for cloud-based training and inference of large models. Its intended users were primarily Microsoft’s own services and Azure-managed offerings rather than customers buying a card and managing drivers themselves.
Microsoft later said Maia 100 was operating in the US East Azure region for Azure OpenAI workloads. That is evidence of production use in a particular region, not proof of worldwide availability or a customer-selectable Maia VM.
Rank #2
- High-Performance AI Processing: The MX3 is designed to handle the most demanding AI computer vision workloads, delivering exceptional performance and efficiency.
- Flexible Integration: The MX3 can be easily integrated into your existing systems via its M.2 M-key form factor and support for Linux operating systems.
- Energy Efficient: The MX3 is designed to provide high performance while minimizing power consumption.
- Comprehensive Software Development Kit (SDK): The MX3 is supported by a comprehensive SDK that simplifies development and deployment.
- Hardware compatability: The MX3 is compatible with the PCI-SIG M.2 M-key 2280 Specification. It can be used with the Raspberry Pi 5 with a M-key 2280 HAT.
Published hardware details
Microsoft’s later Hot Chips 2024 disclosure reported that Maia 100 used TSMC’s 5nm process, an approximately 820 mm² die, TSMC CoWoS-S packaging, four HBM2E stacks, 64 GB of HBM and approximately 1.8 TB/s of HBM bandwidth. These details were published after the Ignite announcement in Microsoft’s Inside Maia 100 technical article.
The system was larger than the chip
Microsoft’s Maia design included custom server boards, rack-level power management, closed-loop liquid cooling and a thermal “sidekick” for the accelerator and host CPUs. Microsoft also reported a custom Ethernet-based networking protocol with aggregate bandwidth of 4.8 Tb/s per accelerator. Those are Microsoft-reported design characteristics, not independent application benchmarks.
Software integration
To make a nonstandard accelerator useful, Microsoft worked on support for PyTorch, ONNX Runtime, Triton, libraries, compilers and developer tools. Real-world performance depends on those software layers, model architecture, precision, batch size, memory movement and networking—not just the silicon specification.
Rank #3
- ✅Powered by 26 Tera-Operations Per Second (TOPS) Hailo-8 AI Processor. 2.5W typical power consumption
- ✅Scalable, enabling simultaneous processing of multi-streams & multi-models
- ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
- ✅Supports TensorFlow, TensorFlow Lite, ONNX, Keras, Pytorch frameworks
- ✅Supports Linux and Windows. Supports the temperature range of -40°C to 85°C
Cobalt 100: an Arm CPU for Azure
Architecture and Microsoft’s performance claim
Cobalt 100 is a fully custom 64-bit Arm processor for the Microsoft Cloud, based on Arm’s Neoverse N2 design. Microsoft described it as a 128-core processor and claimed up to 40% better performance than previous generations of Azure Arm processors. “Up to” is a selected-workload claim, not a universal advantage over every x86 server.
What Azure documentation says today
Current Microsoft documentation lists Cobalt 100 at 3.4 GHz, with one physical core corresponding to each vCPU. Documented Cobalt families reach 96 vCPUs per VM, with memory ranging from 2 GiB to 8 GiB per vCPU depending on the family. The current family list and supported images are maintained in Microsoft’s Cobalt processor-based VM documentation.
| Workload characteristic | Cobalt fit | What to verify |
|---|---|---|
| Linux, horizontally scalable services | Often plausible | Arm64 builds for the application and dependencies |
| Web servers, microservices, caches and open-source databases | Common target workloads | Extension, monitoring and security-agent support |
| x86-only commercial software | Potentially poor fit | Vendor certification and licensing terms |
| Local temporary storage | Family-dependent | Check the selected series; some Cobalt families include local NVMe and others do not |
Microsoft lists support for images including Ubuntu 20.04 and later, Debian 11 and later, RHEL 8.6 and later, SLES 15 SP4 and later, AlmaLinux 8 and later, and Azure Linux 3, subject to the current image documentation.
Rank #4
- 48GB AI graphics accelerator
What customers actually receive
Cobalt: an Azure VM, not a processor shipment
Customers obtain Cobalt indirectly by selecting a supported Azure VM family. Azure billing depends on VM size, region, operating system, storage, networking, billing agreement and usage. There is no single global “Cobalt price”; use the Azure Pricing Calculator for the target configuration and date.
Before moving a workload, test the complete software supply chain:
- Build the application and all native dependencies for Arm64.
- Check container base images, sidecars, agents, database extensions and security tools.
- Confirm vendor certification and architecture-specific licensing.
- Benchmark representative traffic and data, not just a synthetic CPU test.
- Compare the result with an equivalent x86 Azure VM in the same region and billing model.
Maia: mostly a managed-service path
Maia 100 was presented primarily as infrastructure behind Microsoft services. A customer using Azure OpenAI or another managed Azure AI service generally consumes an API or managed capacity rather than selecting Maia hardware, installing a driver or controlling its cooling and firmware. Azure AI Foundry is Microsoft’s managed entry point for many such services: Azure AI Foundry.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Best Value
- Professional AI & Creator Workstation: AMD Radeon AI PRO R9700 GPU with 32GB GDDR6 is engineered for AI development, professional content creation, and compute-intensive workloads.
- Massive 32GB Memory Capacity: 32GB of GDDR6 memory on a 256-bit bus provides ample bandwidth for large AI models, 8K video editing, and complex 3D rendering.
- Advanced RDNA 4 with AI Accelerators: 64 Compute Units with 3rd Gen Ray Tracing and dedicated 2nd Gen AI Accelerators for groundbreaking AI performance and visual computing.
- Professional Blower Cooling: Efficient single blower design exhausts heat directly out of the chassis, ideal for multi-GPU workstation and server configurations.
- Enterprise-Grade Thermal Solution: Vapor chamber heatsink with industrial Honeywell PTM7950 thermal interface material ensures reliable cooling under sustained professional loads.
That model can simplify deployment, but it does not provide direct hardware access, CUDA portability or guaranteed Maia capacity in every region.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Maia versus Nvidia and AMD accelerators
| Criterion | Maia 100 | Nvidia accelerators | AMD accelerators |
|---|---|---|---|
| Primary advantage | Co-designed for Microsoft’s Azure systems and selected services | Broad software ecosystem and extensive cloud and enterprise deployment | Alternative accelerator architecture and supplier |
| Portability | Most naturally aligned with Microsoft’s stack | Often strong where CUDA-based software is already established | Depends on framework, compiler and kernel support |
| Public apples-to-apples evidence | Limited in the supplied sources | Varies by GPU, model and platform | Varies by accelerator, model and platform |
| Customer hardware purchase | Not offered as a standalone chip in the reviewed sources | Typically consumed through cloud instances or purchased systems | Typically consumed through cloud instances or purchased systems |
There is no defensible universal winner. Model size, precision, batch size, memory capacity, compiler maturity, region, utilization and contractual pricing determine the result. Microsoft can use Maia for workloads it has optimized while continuing to offer Nvidia and AMD capacity for other models and software stacks.
Arm migration checklist for Cobalt
- Inventory x86-only binaries, plugins and native libraries.
- Confirm Arm64 versions of container images and build tools.
- Check database drivers, language runtimes and cryptographic libraries.
- Verify observability, backup, endpoint-security and vulnerability-scanning agents.
- Review software licenses that price by architecture, socket or core.
- Test performance, startup behavior and autoscaling under production-like load.
- Check the selected VM family’s temporary-disk and networking characteristics.
Cobalt is most compelling when an application is Linux-first, horizontally scalable and already maintained for Arm64. An x86 dependency that cannot be ported or emulated reliably can outweigh any processor-level efficiency benefit.
From Ignite 2023 to Microsoft’s later roadmap
| Date | Development |
|---|---|
| November 15, 2023 | Microsoft announces Maia 100 and Cobalt 100 at Ignite. |
| April 3, 2024 | Microsoft publishes deeper Maia systems, cooling, networking and software details in Azure Maia for the era of AI. |
| 2024 | Microsoft discloses additional Maia 100 process, packaging and HBM specifications. |
| Late 2024 | Microsoft reports Maia 100 live in US East for Azure OpenAI workloads in its Ignite 2024 keynote transcript. |
| 2025 onward | Cobalt 100 appears in customer-facing Azure VM families. |
| January 26, 2026 | Microsoft announces Maia 200, an inference-focused successor with a 3nm process, 216 GB of HBM3e, 7 TB/s memory bandwidth and native FP8/FP4 tensor support: Microsoft’s Maia 200 announcement. |
What the announcement means for cloud buyers and investors
The important change is not that Microsoft created two chips that everyone can buy. It is that Azure can combine internally designed processors with third-party hardware and expose the result as cloud capacity. For buyers, the practical questions are compatibility, regional availability, service-level guarantees and total cost of ownership. For observers of Microsoft’s infrastructure economics, the chips show how hyperscalers are trying to control power, supply and workload-specific efficiency as AI demand expands.
Microsoft’s performance and cost statements should be read as vendor claims. A chip’s result in production depends on the whole system and software stack, and availability in one region does not establish global availability.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




