
NVIDIA vs AMD vs Intel vs Cerebras AI Accelerator Comparison
## Executive Summary
The AI accelerator market in 2026 remains dominated by **NVIDIA**, which reported **\$62.3 billion** in quarterly Data Center revenue for its fiscal fourth quarter, up 75% year over year, and **\$193.7 billion** for the full fiscal year, up 68% <a href="http://nvidianews.nvidia.com/news/nvidia-announces-financial-results-for-fourth-quarter-and-fiscal-2026#:~:text=Record%20quarterly%20Data%20Center%20revenue%20of%20%2462.3%20billion%2C%20up%2022%25%20from%20Q3%20and%20up%2075%25%20from%20a%20year%20ago" title="Highlights: Record quarterly Data Center revenue of $62.3 billion, up 22% from Q3 and up 75% from a year ago" class="citation-link"><sup>[1]</sup></a> <a href="http://nvidianews.nvidia.com/news/nvidia-announces-financial-results-for-fourth-quarter-and-fiscal-2026#:~:text=Full-year%20revenue%20rose%2068%25%20to%20a%20record%20%24193.7%20billion" title="Highlights: Full-year revenue rose 68% to a record $193.7 billion" class="citation-link"><sup>[2]</sup></a>. Total fiscal 2026 revenue reached **\$215.9 billion**, up 65% year over year <a href="http://nvidianews.nvidia.com/news/nvidia-announces-financial-results-for-fourth-quarter-and-fiscal-2026#:~:text=For%20fiscal%202026%2C%20revenue%20was%20%24215.9%20billion%2C%20up%2065%25%20from%20a%20year%20ago" title="Highlights: For fiscal 2026, revenue was $215.9 billion, up 65% from a year ago" class="citation-link"><sup>[3]</sup></a>. Independent market research from Mordor Intelligence estimates NVIDIA retained roughly **80%** of global AI training revenue in 2024, with the broader AI accelerator market worth **\$174.69 billion** in 2026 and projected to reach **\$518.12 billion** by 2031, a **24.30%** compound annual growth rate (CAGR) <a href="https://www.mordorintelligence.com/industry-reports/ai-accelerators-market#:~:text=NVIDIA%20retained%20about%2080%25%20of%20global%20training%20revenue%20in%202024" title="Highlights: NVIDIA retained about 80% of global training revenue in 2024" class="citation-link"><sup>[4]</sup></a> <a href="https://www.mordorintelligence.com/industry-reports/ai-accelerators-market#:~:text=forecast%20to%20reach%20USD%20518.12%20billion%20by%202031%20at%2024.30%25%20CAGR" title="Highlights: forecast to reach USD 518.12 billion by 2031 at 24.30% CAGR" class="citation-link"><sup>[5]</sup></a>. NVIDIA's flagship data center platform, the **GB200 NVL72**, links 72 Blackwell Tensor Core GPUs and 36 Grace CPUs into a single rack delivering **30x** faster real-time trillion-parameter large language model (LLM) inference and 25 times more performance per watt than air-cooled Hopper H100 infrastructure <a href="https://www.nvidia.com/en-us/data-center/gb200-nvl72/#:~:text=delivers%2030x%20faster%20real-time%20trillion-parameter%20large%20language%20model" title="Highlights: delivers 30x faster real-time trillion-parameter large language model" class="citation-link"><sup>[6]</sup></a> <a href="https://www.nvidia.com/en-us/data-center/gb200-nvl72/#:~:text=GB200%20delivers%2025x%20more%20performance%20at%20the%20same%20power" title="Highlights: GB200 delivers 25x more performance at the same power" class="citation-link"><sup>[7]</sup></a>.
**AMD** is the most direct GPU-based challenger, reporting Q1 2026 Data Center segment revenue of **\$5.8 billion**, up 57% year over year, driven by EPYC central processing units (CPUs) and Instinct graphics processing units (GPUs) <a href="https://www.datacenterdynamics.com/en/news/amd-posts-q1-2026-data-center-revenue-of-58bn-forecasts-120bn-server-cpu-income-by-2030/#:~:text=AMD%20saw%20a%2057%20percent%20year-on-year%20%28YoY%29%20revenue%20increase%20for%20its%20data%20center%20segment" title="Highlights: AMD saw a 57 percent year-on-year (YoY) revenue increase for its data center segment" class="citation-link"><sup>[8]</sup></a>. Its **Instinct MI355X** accelerator ships with 288 gigabytes (GB) of HBM3E memory and 8 terabytes per second (TB/s) of bandwidth, exceeding the memory capacity of NVIDIA's B200 <a href="https://www.amd.com/en/products/accelerators/instinct/mi350.html#:~:text=288GB%20of%20HBM3E%20memory%20and%208TB%2Fs%20bandwidth" title="Highlights: 288GB of HBM3E memory and 8TB/s bandwidth" class="citation-link"><sup>[9]</sup></a>, and in February 2026 AMD signed a five-year, up to 6-gigawatt (GW) supply partnership with Meta built around custom MI450-based silicon <a href="https://ir.amd.com/news-events/press-releases/detail/1279/amd-and-meta-announce-expanded-strategic-partnership-to-deploy-6-gigawatts-of-amd-gpus#:~:text=agree%20to%20a%20definitive%20multi-year%2C%20multi-generation%20partnership%20to%20deploy%20up%20to%206%20gigawatts" title="Highlights: agree to a definitive multi-year, multi-generation partnership to deploy up to 6 gigawatts" class="citation-link"><sup>[10]</sup></a>. Independent benchmarking by Artificial Analysis found the MI300X achieved **25 to 35%** higher peak system throughput than comparable [NVIDIA H100 and H200 systems](https://gpusmith.com/articles/gpu-cloud-rental-prices-h100-h200) under high-concurrency inference loads, though NVIDIA retained a latency edge at low concurrency <a href="https://artificialanalysis.ai/articles/independent-analysis-of-leading-gpus-amd-nvidia#:~:text=AMD%20MI300X%20system%20achieved%2025-35%25%20higher%20peak%20system%20output%20throughput" title="Highlights: AMD MI300X system achieved 25-35% higher peak system output throughput" class="citation-link"><sup>[11]</sup></a>.
**Intel's Gaudi 3** accelerator, priced by industry analysts at roughly **\$16,000** per unit in volume against [\$30,000-plus for an NVIDIA H100](https://gpusmith.com/articles/own-vs-rent-gpus-tco-comparison), targets price-conscious inference buyers with 128GB of HBM2e memory and 33% more input/output (I/O) connectivity per accelerator than the H100 <a href="https://www.tweaktown.com/news/98977/intel-discounting-new-gaudi-3-ai-accelerator-16k-against-nvidia-h100-gpu-for-30k/index.html#:~:text=costs%20just%20%2416%2C000%20per%20AI%20accelerator%20with%20128GB%20of%20HBM2e%20memory" title="Highlights: costs just $16,000 per AI accelerator with 128GB of HBM2e memory" class="citation-link"><sup>[12]</sup></a> <a href="https://www.intel.com/content/www/us/en/products/details/processors/ai-accelerators/gaudi.html#:~:text=33%20percent%20more%20I%2FO%20connectivity%20per%20accelerator%20compared%20to%20H100" title="Highlights: 33 percent more I/O connectivity per accelerator compared to H100" class="citation-link"><sup>[13]</sup></a>. Intel's Data Center and AI segment grew to **\$5.1 billion** in Q1 2026 as part of a broader beat that CEO Lip-Bu Tan described as a "sixth consecutive quarter of revenue above our expectations" <a href="https://www.intc.com/news-events/press-releases/detail/1767/intel-reports-first-quarter-2026-financial-results#:~:text=sixth%20consecutive%20quarter%20of%20revenue%20above%20our%20expectations" title="Highlights: sixth consecutive quarter of revenue above our expectations" class="citation-link"><sup>[14]</sup></a>, though official MLPerf Inference v6.0 submissions in April 2026 did not include Gaudi 3 results (Source: [spheron.network](https://www.spheron.network/blog/mlperf-inference-v6-benchmark-results-2026/#:~:text=Intel%27s%20official%20v6.0%20submissions%20covered%20Xeon%206%20CPU%20and%20Arc%20Pro%20GPU%20workloads)).
**Cerebras Systems** takes the most architecturally distinct approach: its third-generation **Wafer-Scale Engine (WSE-3)** is a single 46,225 square millimeter chip carrying 4 trillion transistors and 900,000 AI-optimized cores, delivering 125 petaflops of compute and claimed to offer 19 times more transistors and 28 times more compute than [a single NVIDIA B200](https://gpusmith.com/articles/own-vs-rent-gpus-tco-comparison) <a href="https://www.cerebras.ai/chip#:~:text=The%20WSE-3%20is%20the%20largest%20AI%20chip%20ever%20built" title="Highlights: The WSE-3 is the largest AI chip ever built" class="citation-link"><sup>[15]</sup></a> <a href="https://www.cerebras.ai/chip#:~:text=more%20compute%20than%20the%20NVIDIA%20B200" title="Highlights: more compute than the NVIDIA B200" class="citation-link"><sup>[16]</sup></a>. Cerebras completed the largest initial public offering (IPO) of 2026 in May, pricing at a valuation of up to **\$48.8 billion**, up from a **\$23 billion** private round three months earlier <a href="https://www.cnbc.com/2026/05/11/cerebras-raises-ipo-range.html#:~:text=worth%20up%20to%20%2448.8%20billion%20based%20on%20the%20new%20price%20range" title="Highlights: worth up to $48.8 billion based on the new price range" class="citation-link"><sup>[17]</sup></a> <a href="https://www.cnbc.com/2026/05/11/cerebras-raises-ipo-range.html#:~:text=up%20from%20the%20%2423%20billion%20valuation%20Cerebras%20announced%20in%20February" title="Highlights: up from the $23 billion valuation Cerebras announced in February" class="citation-link"><sup>[18]</sup></a>, buoyed by an OpenAI inference deal worth more than **\$10 billion** covering 750 megawatts of deployed capacity <a href="https://www.cnbc.com/2026/01/16/openai-chip-deal-with-cerebras-adds-to-roster-of-nvidia-amd-broadcom.html#:~:text=OpenAI%27s%20deal%20with%20Cerebras%20is%20worth%20more%20than%20%2410%20billion" title="Highlights: OpenAI's deal with Cerebras is worth more than $10 billion" class="citation-link"><sup>[19]</sup></a> <a href="https://www.cnbc.com/2026/01/16/openai-chip-deal-with-cerebras-adds-to-roster-of-nvidia-amd-broadcom.html#:~:text=an%20agreement%20to%20deploy%20750%20megawatts%20of%20Cerebras%27%20AI%20chips" title="Highlights: an agreement to deploy 750 megawatts of Cerebras' AI chips" class="citation-link"><sup>[20]</sup></a>.
On raw power efficiency, Cerebras's own analysis claims a 2.2 times improvement in performance per watt over a comparable NVIDIA DGX B200 system <a href="https://www.cerebras.ai/blog/cerebras-cs-3-vs-nvidia-b200-2024-ai-accelerators-compared#:~:text=a%202.2x%20improvement%20in%20performance%20per%20watt" title="Highlights: a 2.2x improvement in performance per watt" class="citation-link"><sup>[21]</sup></a>, while an independent, peer-reviewed arXiv comparison finds wafer-scale integration offers genuine advantages in performance per watt and memory scalability but concedes cost-effectiveness at commercial scale remains unproven <a href="https://arxiv.org/html/2503.11698v1#:~:text=advantages%20of%20WSE-3%20in%20performance%20per%20watt%20and%20memory%20scalability" title="Highlights: advantages of WSE-3 in performance per watt and memory scalability" class="citation-link"><sup>[22]</sup></a>. There is no single "best" accelerator for every workload as of July 2026. NVIDIA leads on software maturity, ecosystem breadth, and raw training scale, evidenced by MLPerf Training v6.0 results in which it was "the only platform to submit on every test" <a href="https://developer.nvidia.com/blog/nvidia-blackwell-tops-mlperf-training-6-0-with-industry-leading-scale-and-performance/#:~:text=was%20the%20only%20platform%20to%20submit%20on%20every%20test" title="Highlights: was the only platform to submit on every test" class="citation-link"><sup>[23]</sup></a>; AMD offers the strongest open-source, memory-per-dollar alternative for inference; Intel competes on acquisition price and standard Ethernet networking; and Cerebras wins outright on single-request inference latency and token throughput for models that fit its architecture. This report examines each vendor's capabilities, adoption, and limitations in detail, presents third-party benchmark and financial evidence, and profiles four real-world deployments to help technical buyers match accelerator choice to workload.
## Introduction and Background
Artificial intelligence (AI) accelerators, purpose-built processors designed to speed up the matrix multiplication and tensor operations underlying neural network training and inference, have become the single largest driver of data center capital spending as of mid-2026. What began as a market almost entirely defined by general-purpose graphics processing units (GPUs) has splintered into competing architectural philosophies: NVIDIA's CUDA-centric GPU platform, AMD's open ROCm software stack running on Instinct GPUs, Intel's Ethernet-native Gaudi accelerators, and Cerebras's radically different wafer-scale approach that replaces racks of discrete chips with a single dinner-plate-sized piece of silicon.
The stakes are enormous.
```NVIDIA alone generated **\$215.9 billion** in total revenue for fiscal 2026, up 65% from the prior year, with Data Center products accounting for the overwhelming majority of that growth <a href="http://nvidianews.nvidia.com/news/nvidia-announces-financial-results-for-fourth-quarter-and-fiscal-2026#:~:text=For%20fiscal%202026%2C%20revenue%20was%20%24215.9%20billion%2C%20up%2065%25%20from%20a%20year%20ago" title="Highlights: For fiscal 2026, revenue was $215.9 billion, up 65% from a year ago" class="citation-link"><sup>[3]</sup></a>. Mordor Intelligence's industry analysis puts the entire AI accelerator market at **\$174.69 billion** in 2026, up from **\$140.55 billion** in 2025, expanding at a **24.30%** CAGR through 2031 as hyperscale cloud providers, sovereign AI initiatives, and enterprise buyers race to secure compute capacity <a href="https://www.mordorintelligence.com/industry-reports/ai-accelerators-market#:~:text=grow%20from%20USD%20140.55%20billion%20in%202025%20to%20USD%20174.69%20billion%20in%202026" title="Highlights: grow from USD 140.55 billion in 2025 to USD 174.69 billion in 2026" class="citation-link"><sup>[24]</sup></a>. Within that spending, GPUs still account for the majority of revenue, an estimated 59.20% in 2025, while ASICs are catching up at a projected 27.15% CAGR through 2030 <a href="https://www.mordorintelligence.com/industry-reports/ai-accelerators-market#:~:text=GPUs%20held%2059.20%25%20revenue%20share%20of%20the%20AI%20accelerators%20market%20in%202025" title="Highlights: GPUs held 59.20% revenue share of the AI accelerators market in 2025" class="citation-link"><sup>[25]</sup></a>. Cloud and colocation data centers absorb roughly three-quarters of total deployment (74.30% of 2025 spending), while hyperscale cloud service providers alone controlled 52.40% of overall AI accelerator spending in the same year <a href="https://www.mordorintelligence.com/industry-reports/ai-accelerators-market#:~:text=Cloud%20and%20colocation%20facilities%20accounted%20for%2074.30%25%20of%202025%20spending" title="Highlights: Cloud and colocation facilities accounted for 74.30% of 2025 spending" class="citation-link"><sup>[26]</sup></a>.
This comparison focuses on four contenders that between them account for the overwhelming majority of merchant (non-captive) AI accelerator silicon shipped for training and inference workloads: **NVIDIA** (Hopper H100/H200 and Blackwell B200/GB200), **AMD** (Instinct MI300X, MI325X, and MI350 series), **Intel** (Gaudi 2 and Gaudi 3), and **Cerebras** (the CS-3 system built around the WSE-3 wafer-scale chip). It deliberately excludes hyperscaler-only, non-merchant silicon such as Google's Tensor Processing Units (TPUs), Amazon's Trainium and Inferentia chips, and Microsoft's Maia, which are not generally available for third-party purchase or lease in the same way. It also notes, where relevant, custom application-specific integrated circuit (ASIC) competition from Broadcom, whose custom AI chip pipeline has driven speculation of a **\$60 billion to \$90 billion** opportunity by 2027 according to Mordor Intelligence's industry tracking, and whose market capitalization CNBC reported at over **\$1.6 trillion** as investor enthusiasm for its custom XPU pipeline grew <a href="https://www.mordorintelligence.com/industry-reports/ai-accelerators-market#:~:text=Broadcom%20anticipates%20a%20USD%2060%E2%80%9390%20billion%20ASIC%20opportunity%20by%202027" title="Highlights: Broadcom anticipates a USD 60–90 billion ASIC opportunity by 2027" class="citation-link"><sup>[27]</sup></a> <a href="https://www.cnbc.com/2026/01/16/openai-chip-deal-with-cerebras-adds-to-roster-of-nvidia-amd-broadcom.html#:~:text=Broadcom%20is%20now%20valued%20at%20over%20%241.6%20trillion" title="Highlights: Broadcom is now valued at over $1.6 trillion" class="citation-link"><sup>[28]</sup></a>.
Each vendor is assessed on architecture and capabilities, real-world adoption evidenced by named customer deployments, and documented strengths and limitations, before the report turns to a head-to-head feature matrix, third-party benchmark data (including MLPerf Training v6.0 and Inference v6.0 results published by MLCommons in 2026), financial and market-share evidence, and four detailed case studies. All figures are anchored "as of" their publication or filing date and are drawn from vendor technical documentation, U.S. Securities and Exchange Commission (SEC) filings, peer-reviewed and preprint research, and independent benchmarking organizations, with discrepancies between vendor-reported and independently measured figures flagged explicitly throughout.
## NVIDIA: The Data Center Standard
## Capabilities
NVIDIA's current data center lineup spans the Hopper-generation **H100** and **H200** GPUs and the newer Blackwell-generation **B200** and rack-scale **GB200 NVL72**. The H100 delivers up to 3,958 teraFLOPS (TFLOPS) of FP8 Tensor Core performance with sparsity, 80GB of HBM3 memory, and 3.35TB/s of memory bandwidth, with a configurable thermal design power (TDP) of up to 700 watts (W) <a href="https://www.nvidia.com/en-us/data-center/h100/#:~:text=Up%20to%20700W%20%28configurable%29" title="Highlights: Up to 700W (configurable)" class="citation-link"><sup>[29]</sup></a>. NVIDIA states the H100's fourth-generation Tensor Cores and Transformer Engine deliver up to 4 times faster training than the prior generation for GPT-3 175B parameter models, alongside up to 30 times higher inference performance on the largest models compared with the previous Ampere generation. The GB200 NVL72 connects 72 Blackwell GPUs and 36 Grace CPUs in a single liquid-cooled rack using fifth-generation NVLink, providing 130TB/s of low-latency GPU-to-GPU communication across the rack, 13.4TB of pooled HBM3E memory, and 576TB/s of aggregate memory bandwidth <a href="https://www.nvidia.com/en-us/data-center/gb200-nvl72/#:~:text=130%20terabytes%20per%20second%20%28TB%2Fs%29%20of%20low-latency%20GPU%20communications" title="Highlights: 130 terabytes per second (TB/s) of low-latency GPU communications" class="citation-link"><sup>[30]</sup></a> <a href="https://www.nvidia.com/en-us/data-center/gb200-nvl72/#:~:text=13.4%20TB%20HBM3E" title="Highlights: 13.4 TB HBM3E" class="citation-link"><sup>[31]</sup></a>. NVIDIA states GB200 NVL72 delivers 4 times faster training and 30 times faster real-time trillion-parameter LLM inference than an equivalent H100 cluster, alongside a claimed 10 times uplift for mixture-of-experts (MoE) architectures. The successor **GB300 NVL72** (Blackwell Ultra) integrates 72 Blackwell Ultra GPUs and 36 Arm-based Grace CPUs and targets up to a 50 times increase in AI factory output versus Hopper-based platforms <a href="https://www.nvidia.com/en-us/data-center/gb200-nvl72/#:~:text=up%20to%20a%2050x%20overall%20increase%20in%20AI%20factory%20output%20performance" title="Highlights: up to a 50x overall increase in AI factory output performance" class="citation-link"><sup>[32]</sup></a>.
## Adoption
NVIDIA's ecosystem advantage rests on the maturity of its CUDA software stack, which remains the default target for most machine learning research and production frameworks. On the infrastructure side, the clearest evidence of scale is xAI's **Colossus** supercomputer, built with Supermicro and connecting 100,000 NVIDIA Hopper Tensor Core GPUs over NVIDIA Spectrum-X Ethernet networking, described by Supermicro as one of the largest liquid-cooled AI clusters in the world <a href="https://www.supermicro.com/en/featured/xai-colossus#:~:text=connect%20100%2C000%20NVIDIA%20Hopper%20Tensor%20Core%20GPUs" title="Highlights: connect 100,000 NVIDIA Hopper Tensor Core GPUs" class="citation-link"><sup>[33]</sup></a>. In September 2025, NVIDIA said it would commit \$100 billion to support OpenAI's deployment of at least 10 gigawatts of NVIDIA systems, a project CEO Jensen Huang called "a giant project" that will equate to between 4 million and 5 million GPUs <a href="https://www.cnbc.com/2026/01/16/openai-chip-deal-with-cerebras-adds-to-roster-of-nvidia-amd-broadcom.html#:~:text=commit%20%24100%20billion%20to%20support%20OpenAI%20as%20it%20builds%20and%20deploys%20at%20least%2010%20gigawatts" title="Highlights: commit $100 billion to support OpenAI as it builds and deploys at least 10 gigawatts" class="citation-link"><sup>[34]</sup></a>. NVIDIA's own MLPerf Training v6.0 results, published in June 2026, report the platform was "the only platform to submit on every test," including new DeepSeek-V3 671B mixture-of-experts and GPT-OSS-20B benchmarks <a href="https://developer.nvidia.com/blog/nvidia-blackwell-tops-mlperf-training-6-0-with-industry-leading-scale-and-performance/#:~:text=was%20the%20only%20platform%20to%20submit%20on%20every%20test" title="Highlights: was the only platform to submit on every test" class="citation-link"><sup>[23]</sup></a>, with cloud partners scaling training runs up to 8,192 Blackwell GPUs working in unison across production hyperscale fleets <a href="https://developer.nvidia.com/blog/nvidia-blackwell-tops-mlperf-training-6-0-with-industry-leading-scale-and-performance/#:~:text=scaled%20up%20to%208%2C192%20Blackwell%20GPUs%20working%20in%20unison" title="Highlights: scaled up to 8,192 Blackwell GPUs working in unison" class="citation-link"><sup>[35]</sup></a>. NVIDIA also disclosed a new Rubin platform in its fiscal 2026 results, designed to deliver up to a 10 times reduction in inference token cost compared with Blackwell <a href="http://nvidianews.nvidia.com/news/nvidia-announces-financial-results-for-fourth-quarter-and-fiscal-2026#:~:text=up%20to%20a%2010x%20reduction%20in%20inference%20token%20cost%2C%20compared%20with%20the%20NVIDIA%20Blackwell%20platform" title="Highlights: up to a 10x reduction in inference token cost, compared with the NVIDIA Blackwell platform" class="citation-link"><sup>[36]</sup></a>.
## Strengths and Limitations
NVIDIA's principal strength is ecosystem depth: CUDA, cuDNN, TensorRT-LLM, and the broader NVIDIA AI Enterprise software suite give developers a well-trodden path from research to production that competitors still struggle to match. NVIDIA's own data shows the cost-per-token economics improving generationally: the company states H100 delivers inference at approximately \$0.09 per million tokens for a GPT-OSS-120B workload under vLLM, while B200 cuts that to roughly \$0.02 per million tokens under TensorRT-LLM, a 4.5 times improvement, according to SemiAnalysis InferenceX benchmarks NVIDIA cites as of April 2026 <a href="https://www.nvidia.com/en-us/data-center/h100/#:~:text=approximately%20%240.09%20per%20million%20tokens%20at%2066%20TPS%2Fuser%20for%20GPT-OSS-120B" title="Highlights: approximately $0.09 per million tokens at 66 TPS/user for GPT-OSS-120B" class="citation-link"><sup>[37]</sup></a> <a href="https://www.nvidia.com/en-us/data-center/h100/#:~:text=roughly%204.5x%20cheaper%20than%20H100%20at%20%240.09%20per%20million%20tokens" title="Highlights: roughly 4.5x cheaper than H100 at $0.09 per million tokens" class="citation-link"><sup>[38]</sup></a>. The chief limitations are cost, power draw, and supply constraints tied to geopolitics: the April 2025 U.S. export licensing requirement on H20 products forced NVIDIA to take a **\$4.5 billion** charge in a single quarter for excess China-bound inventory, after H20 sales of \$4.6 billion prior to the new requirement and an inability to ship a further \$2.5 billion of H20 revenue in the same quarter <a href="https://www.sec.gov/Archives/edgar/data/1045810/000104581025000115/q1fy26pr.htm#:~:text=%244.5%20billion%20charge%20in%20the%20first%20quarter%20of%20fiscal%202026%20associated%20with%20H20%20excess%20inventory" title="Highlights: $4.5 billion charge in the first quarter of fiscal 2026 associated with H20 excess inventory" class="citation-link"><sup>[39]</sup></a>. That exposure has only grown: TrendForce reporting in July 2026 cites Bernstein Research projecting NVIDIA's share of the Chinese AI semiconductor market will fall to around 8% in 2026, down from roughly 40% a year earlier, as Chinese buyers shift toward domestic suppliers such as Huawei <a href="https://www.trendforce.com/news/2026/07/07/news-chinese-firms-reportedly-raise-domestic-ai-chip-budget-share-from-30-to-46-amid-shift-from-nvidia/#:~:text=investment%20bank%20Bernstein%20projected%20that%20NVIDIA%E2%80%99s%20share%20of%20the%20Chinese%20AI%20semiconductor%20market%20will%20fall%20to%20around%208%25%20in%202026" title="Highlights: investment bank Bernstein projected that NVIDIA’s share of the Chinese AI semiconductor market will fall to around 8% in 2026" class="citation-link"><sup>[40]</sup></a>.
## AMD: Instinct and the Open Ecosystem Challenge
## Capabilities
AMD's Instinct line is built on the CDNA architecture, with the **MI300X** (launched December 2023) offering 192GB of HBM3 memory, 5.3TB/s of peak memory bandwidth, 750W typical board power, 153 billion transistors, and 19,456 stream processors across 304 compute units <a href="https://www.amd.com/en/products/accelerators/instinct/mi300/mi300x.html#:~:text=750W%20Peak" title="Highlights: 750W Peak" class="citation-link"><sup>[41]</sup></a> <a href="https://www.amd.com/en/products/accelerators/instinct/mi300/mi300x.html#:~:text=153%20Billion" title="Highlights: 153 Billion" class="citation-link"><sup>[42]</sup></a>. The newer **MI350 series**, built on 4th Gen CDNA architecture, pushes memory to 288GB of HBM3E with 8TB/s of bandwidth on the flagship MI355X, and an 8-GPU MI350 platform aggregates 2.3TB of total HBM3E memory with 64TB/s of peak theoretical bandwidth <a href="https://www.amd.com/en/products/accelerators/instinct/mi350.html#:~:text=288GB%20of%20HBM3E%20memory%20and%208TB%2Fs%20bandwidth" title="Highlights: 288GB of HBM3E memory and 8TB/s bandwidth" class="citation-link"><sup>[9]</sup></a> <a href="https://www.amd.com/en/products/accelerators/instinct/mi350.html#:~:text=2.3%20TB%20Total%20HBM3E%20Memory" title="Highlights: 2.3 TB Total HBM3E Memory" class="citation-link"><sup>[43]</sup></a>. AMD's own comparisons claim the MI355X delivers up to 2.2 times the AI performance of a comparable B200 SXM5 on select precision formats, and 1.6 times the memory capacity <a href="https://www.amd.com/en/products/accelerators/instinct/mi350.html#:~:text=Up%20to%202.2X%20the%20AI%20performance%20vs.%20competitive%20accelerators" title="Highlights: Up to 2.2X the AI performance vs. competitive accelerators" class="citation-link"><sup>[44]</sup></a>. For buyers who do not need flagship specifications, AMD also offers the PCIe-based MI350P, positioned around lower cost, with 128 GPU compute units, 144GB of HBM3E memory, and up to 4TB/s of peak theoretical memory bandwidth. Software runs on the open-source **ROCm** stack, which AMD positions as a no-licensing-fee alternative to CUDA.
## Adoption
AMD's Instinct fleet has moved from niche deployments to hyperscaler-scale commitments in the past two years. At its 2024 Advancing AI event, AMD disclosed that MI300X was "serving all live traffic on Llama 405B" for Meta and had already been deployed by cloud, OEM, and ODM partners serving OpenAI's ChatGPT and Hugging Face workloads <a href="https://ir.amd.com/news-events/press-releases/detail/1218/amd-unveils-leadership-ai-solutions-at-advancing-ai-2024#:~:text=MI300X%20serving%20all%20live%20traffic%20on%20Llama%20405B" title="Highlights: MI300X serving all live traffic on Llama 405B" class="citation-link"><sup>[45]</sup></a>, while Databricks reported over 50% performance gains on Llama and its own proprietary models using MI300X's larger memory <a href="https://ir.amd.com/news-events/press-releases/detail/1218/amd-unveils-leadership-ai-solutions-at-advancing-ai-2024#:~:text=over%2050%25%20increase%20in%20performance%20on%20Llama%20and%20Databricks%20proprietary%20models" title="Highlights: over 50% increase in performance on Llama and Databricks proprietary models" class="citation-link"><sup>[46]</sup></a>. AMD CEO Lisa Su used the same event to project "the data center AI accelerator market growing to \$500 billion by 2028" <a href="https://ir.amd.com/news-events/press-releases/detail/1218/amd-unveils-leadership-ai-solutions-at-advancing-ai-2024#:~:text=the%20data%20center%20AI%20accelerator%20market%20growing%20to%20%24500%20billion%20by%202028" title="Highlights: the data center AI accelerator market growing to $500 billion by 2028" class="citation-link"><sup>[47]</sup></a>. That relationship deepened materially in February 2026, when AMD and Meta announced a five-year, multi-generation partnership to deploy up to 6 gigawatts of AMD Instinct GPUs, with the first tranche built on a custom MI450-based GPU and AMD's sixth-generation "Venice" EPYC CPUs running on the Helios rack-scale architecture <a href="https://ir.amd.com/news-events/press-releases/detail/1279/amd-and-meta-announce-expanded-strategic-partnership-to-deploy-6-gigawatts-of-amd-gpus#:~:text=agree%20to%20a%20definitive%20multi-year%2C%20multi-generation%20partnership%20to%20deploy%20up%20to%206%20gigawatts" title="Highlights: agree to a definitive multi-year, multi-generation partnership to deploy up to 6 gigawatts" class="citation-link"><sup>[10]</sup></a> <a href="https://ir.amd.com/news-events/press-releases/detail/1279/amd-and-meta-announce-expanded-strategic-partnership-to-deploy-6-gigawatts-of-amd-gpus#:~:text=custom%20AMD%20Instinct%20GPU%20based%20on%20the%20MI450%20architecture" title="Highlights: custom AMD Instinct GPU based on the MI450 architecture" class="citation-link"><sup>[48]</sup></a>. OpenAI separately committed to deploying six gigawatts of AMD GPUs across multiple hardware generations starting in the second half of 2026, a deal that came with a warrant for up to 160 million shares of AMD stock <a href="https://www.cnbc.com/2026/01/16/openai-chip-deal-with-cerebras-adds-to-roster-of-nvidia-amd-broadcom.html#:~:text=deploy%20six%20gigawatts%20of%20AMD%27s%20GPUs%20across%20multiple%20years" title="Highlights: deploy six gigawatts of AMD's GPUs across multiple years" class="citation-link"><sup>[49]</sup></a>. AMD's own case studies also point to cloud infrastructure specialist TensorWave, which the company says built an AMD Instinct GPU cloud "delivering reliable, resilient AI infrastructure" with meaningful performance and cost advantages versus alternatives <a href="https://www.amd.com/en/products/accelerators/instinct/mi350.html#:~:text=delivering%20reliable%2C%20resilient%20AI%20infrastructure" title="Highlights: delivering reliable, resilient AI infrastructure" class="citation-link"><sup>[50]</sup></a>.
## Strengths and Limitations
AMD's core advantage is memory capacity per dollar: the MI300X's 192GB and the MI355X's 288GB of HBM comfortably exceed the H100's 80GB and even the H200's 141GB, letting a single accelerator host larger models with less multi-GPU partitioning overhead. Independent benchmarking from Artificial Analysis, conducted with AMD's cooperation but using the firm's own published methodology, found the MI300X achieved 25 to 35% higher peak system output throughput than comparable NVIDIA H100 and H200 systems at high concurrency on DeepSeek R1 and Llama 4 Maverick workloads, while NVIDIA retained an edge in per-query latency at low concurrency <a href="https://artificialanalysis.ai/articles/independent-analysis-of-leading-gpus-amd-nvidia#:~:text=AMD%20MI300X%20system%20achieved%2025-35%25%20higher%20peak%20system%20output%20throughput" title="Highlights: AMD MI300X system achieved 25-35% higher peak system output throughput" class="citation-link"><sup>[11]</sup></a> <a href="https://artificialanalysis.ai/articles/independent-analysis-of-leading-gpus-amd-nvidia#:~:text=NVIDIA%20H100%20%26%20H200%20systems%20demonstrate%20marginally%20faster%20average%20per-query%20output%20speeds" title="Highlights: NVIDIA H100 & H200 systems demonstrate marginally faster average per-query output speeds" class="citation-link"><sup>[51]</sup></a>. The persistent limitation remains software maturity. Developer sentiment on Reddit's r/MachineLearning community, while not a benchmark, is illustrative of the gap AMD is closing: one widely upvoted comment noted that "properly written and optimized ROCm code is just as fast as Cuda, right below whatever the maximum number of tflops of your GPU is," but cautioned that most published kernels, including flash-attention implementations, are written and tuned for CUDA first <a href="https://www.reddit.com/r/MachineLearning/comments/1fa8vq5/d_why_is_cuda_so_much_faster_than_rocm/#:~:text=Properly%20written%20and%20optimized%20ROCm%20code%20is%20just%20as%20fast%20as%20Cuda" title="Highlights: Properly written and optimized ROCm code is just as fast as Cuda" class="citation-link"><sup>[52]</sup></a>. Another commenter in the same thread argued the software gap reflects hiring and investment, not architecture, writing that AMD "is not serious about hiring talent" relative to NVIDIA, OpenAI, and Meta <a href="https://www.reddit.com/r/MachineLearning/comments/1fa8vq5/d_why_is_cuda_so_much_faster_than_rocm/#:~:text=Speaking%20from%20my%20own%20experience%2C%20AMD%20is%20not%20serious%20about%20hiring%20talent" title="Highlights: Speaking from my own experience, AMD is not serious about hiring talent" class="citation-link"><sup>[53]</sup></a>. AMD has moved to close this gap through acquisitions, including the **\$4.9 billion** purchase of ZT Systems in August 2024 to build systems-integration expertise and the **\$665 million** acquisition of Silo AI the same month for multilingual model-development capabilities, per Mordor Intelligence's tracking of both deals <a href="https://www.mordorintelligence.com/industry-reports/ai-accelerators-market#:~:text=AMD%20acquired%20ZT%20Systems%20for%20USD%204.9%20billion" title="Highlights: AMD acquired ZT Systems for USD 4.9 billion" class="citation-link"><sup>[54]</sup></a>.
## Intel: Gaudi and the Price-Performance Contender
## Capabilities
Intel's **Gaudi 3** accelerator, generally available since mid-2024, is built for open, Ethernet-native scaling rather than NVIDIA's proprietary NVLink/NVSwitch interconnects. Intel states Gaudi 3 offers 33% more I/O connectivity per accelerator than an H100 <a href="https://www.intel.com/content/www/us/en/products/details/processors/ai-accelerators/gaudi.html#:~:text=33%20percent%20more%20I%2FO%20connectivity%20per%20accelerator%20compared%20to%20H100" title="Highlights: 33 percent more I/O connectivity per accelerator compared to H100" class="citation-link"><sup>[13]</sup></a> and delivers 2 times the FP8 AI compute, 4 times the BF16 (16-bit floating point) compute, and 2 times the network bandwidth of the prior-generation Gaudi 2 accelerator <a href="https://www.intel.com/content/www/us/en/products/details/processors/ai-accelerators/gaudi.html#:~:text=AI%20compute%20%28BF16%29%20vs.%20Intel%20Gaudi%202%20AI%20Accelerators" title="Highlights: AI compute (BF16) vs. Intel Gaudi 2 AI Accelerators" class="citation-link"><sup>[55]</sup></a>. The accelerator ships in mezzanine, PCIe, and OAM (open accelerator module) form factors with a TDP that scales from 450W up to 900W for water-cooled OAM configurations <a href="https://www.tweaktown.com/news/98977/intel-discounting-new-gaudi-3-ai-accelerator-16k-against-nvidia-h100-gpu-for-30k/index.html#:~:text=water-cooled%20Gaudi%203%20in%20OAM%20form%20features%20up%20to%20900W%20TDP" title="Highlights: water-cooled Gaudi 3 in OAM form features up to 900W TDP" class="citation-link"><sup>[56]</sup></a>. Intel markets Gaudi's avoidance of proprietary interconnect lock-in as a core selling point for buyers who want to reuse existing Ethernet infrastructure.
## Adoption
Intel's most concrete enterprise deployment is with **IBM Cloud**, the first cloud service provider to offer Gaudi 3 accelerators as a service to enterprise customers, with support for IBM's watsonx AI platform following in Q2 2025 <a href="https://www.intel.com/content/www/us/en/customer-spotlight/stories/ibm-gaudi-3-customer-story.html#:~:text=over%205%2C000%20tokens%20per%20second%20for%20IBM%27s%20granite-8b%20model" title="Highlights: over 5,000 tokens per second for IBM's granite-8b model" class="citation-link"><sup>[57]</sup></a>. IBM testing showed a single Gaudi 3 card generating over 5,000 tokens per second on IBM's granite-8b model while supporting more than 100 concurrent users with sub-20-millisecond inter-token latency, and the platform is designed to scale from a single 8-accelerator node at 9.6TB/s of throughput up to a 1,024-node cluster of 8,192 accelerators delivering 9.830 petabytes per second (PB/s) of aggregate throughput <a href="https://www.intel.com/content/www/us/en/customer-spotlight/stories/ibm-gaudi-3-customer-story.html#:~:text=1%2C024-node%20cluster%20%288%2C192%20accelerators%29%20with%20a%20throughput%20of%209.830%20PB%2Fs" title="Highlights: 1,024-node cluster (8,192 accelerators) with a throughput of 9.830 PB/s" class="citation-link"><sup>[58]</sup></a>. To accelerate that adoption, Intel and IBM have run promotional pricing offering customers a discount on Gaudi 3 accelerators on IBM Cloud so they can "run twice as long for the same price" <a href="https://www.intel.com/content/www/us/en/products/details/processors/ai-accelerators/gaudi.html#:~:text=for%206%20months.%20Run%20twice%20as%20long%20for%20the%20same%20price." title="Highlights: for 6 months. Run twice as long for the same price." class="citation-link"><sup>[59]</sup></a>. Intel's Data Center and AI segment revenue reached **\$5.1 billion** in Q1 2026, part of an overall quarterly beat that CEO Lip-Bu Tan attributed to a "sixth consecutive quarter of revenue above our expectations" <a href="https://www.intc.com/news-events/press-releases/detail/1767/intel-reports-first-quarter-2026-financial-results#:~:text=sixth%20consecutive%20quarter%20of%20revenue%20above%20our%20expectations" title="Highlights: sixth consecutive quarter of revenue above our expectations" class="citation-link"><sup>[14]</sup></a>.
## Strengths and Limitations
Intel's clearest differentiator is acquisition price: industry analyst estimates cited by trade press put bulk per-unit Gaudi 3 pricing at approximately **\$16,000**, versus **\$30,000-plus** for an NVIDIA H100, a roughly 2 times discount for comparable memory-bound inference workloads <a href="https://www.tweaktown.com/news/98977/intel-discounting-new-gaudi-3-ai-accelerator-16k-against-nvidia-h100-gpu-for-30k/index.html#:~:text=costs%20just%20%2416%2C000%20per%20AI%20accelerator%20with%20128GB%20of%20HBM2e%20memory" title="Highlights: costs just $16,000 per AI accelerator with 128GB of HBM2e memory" class="citation-link"><sup>[12]</sup></a> <a href="https://www.tweaktown.com/news/98977/intel-discounting-new-gaudi-3-ai-accelerator-16k-against-nvidia-h100-gpu-for-30k/index.html#:~:text=compared%20to%20%2430%2C000%2B%20for%20NVIDIA%27s%20Hopper%20H100%20AI%20GPU" title="Highlights: compared to $30,000+ for NVIDIA's Hopper H100 AI GPU" class="citation-link"><sup>[60]</sup></a>. Gaudi's reliance on standard Ethernet rather than proprietary interconnects also lowers switching and networking costs for buyers who do not want NVLink lock-in, per Intel's own positioning. The central limitation is competitive visibility: Intel's official MLPerf Inference v6.0 submissions in April 2026 covered Xeon 6 CPUs and Arc Pro GPUs rather than Gaudi 3, meaning no directly comparable, independently audited per-accelerator throughput figures for Gaudi 3 exist in the latest MLCommons dataset (Source: [spheron.network](https://www.spheron.network/blog/mlperf-inference-v6-benchmark-results-2026/#:~:text=Intel%27s%20official%20v6.0%20submissions%20covered%20Xeon%206%20CPU%20and%20Arc%20Pro%20GPU%20workloads)). This absence of recent third-party benchmark data, combined with revenue that remains a fraction of NVIDIA's, means buyers evaluating Gaudi 3 must weigh cost savings against a comparatively thin body of recent, standardized performance evidence.
## Cerebras: Wafer-Scale Computing
## Capabilities
Cerebras takes a fundamentally different design philosophy. Rather than networking many discrete chips, the **Wafer-Scale Engine 3 (WSE-3)** is fabricated as a single piece of silicon measuring 46,225 square millimeters, the largest AI chip ever built, containing 4 trillion transistors and 900,000 AI-optimized cores <a href="https://www.cerebras.ai/chip#:~:text=The%20WSE-3%20is%20the%20largest%20AI%20chip%20ever%20built" title="Highlights: The WSE-3 is the largest AI chip ever built" class="citation-link"><sup>[15]</sup></a>. An independent, peer-reviewed comparison published on arXiv in March 2025 confirms the WSE-3 integrates "900,000 AI-optimized cores, and 44 GB of on-chip SRAM," achieving "a peak computing performance of 125 petaflops" and a memory bandwidth of 21 petabytes per second <a href="https://arxiv.org/html/2503.11698v1#:~:text=900%2C000%20AI-optimized%20cores%2C%20and%2044%20GB%20of%20on-chip%20SRAM" title="Highlights: 900,000 AI-optimized cores, and 44 GB of on-chip SRAM" class="citation-link"><sup>[61]</sup></a> <a href="https://arxiv.org/html/2503.11698v1#:~:text=a%20peak%20computing%20performance%20of%20125%20petaflops" title="Highlights: a peak computing performance of 125 petaflops" class="citation-link"><sup>[62]</sup></a>. This is the third generation of the architecture: the original 2019 WSE-1 held over 1.2 trillion transistors, 400,000 AI-optimized cores, and 18GB of high-speed SRAM with 9 petabytes per second of memory bandwidth, while the WSE-2 grew to 2.6 trillion transistors, 850,000 cores, and 40GB of on-chip SRAM with 20 petabytes per second of memory bandwidth and 220 petabits per second of fabric bandwidth. The wafer is packaged into the **CS-3** system, which decouples compute from memory using an external "MemoryX" appliance that scales from 1.2 terabytes up to 1.2 petabytes, allowing a single CS-3 to hold models with up to 24 trillion parameters without the multi-GPU partitioning schemes GPU clusters require <a href="https://arxiv.org/html/2503.11698v1#:~:text=advantages%20of%20WSE-3%20in%20performance%20per%20watt%20and%20memory%20scalability" title="Highlights: advantages of WSE-3 in performance per watt and memory scalability" class="citation-link"><sup>[22]</sup></a>.
## Adoption
Cerebras's commercial momentum accelerated sharply in the first half of 2026. In January, OpenAI signed an agreement to deploy 750 megawatts of Cerebras AI chips in tranches through 2028, a deal worth more than **\$10 billion**, with Cerebras CEO Andrew Feldman calling it a way of "bringing the world's leading AI models to the world's fastest AI processor" <a href="https://www.cnbc.com/2026/01/16/openai-chip-deal-with-cerebras-adds-to-roster-of-nvidia-amd-broadcom.html#:~:text=an%20agreement%20to%20deploy%20750%20megawatts%20of%20Cerebras%27%20AI%20chips" title="Highlights: an agreement to deploy 750 megawatts of Cerebras' AI chips" class="citation-link"><sup>[20]</sup></a> <a href="https://www.cnbc.com/2026/01/16/openai-chip-deal-with-cerebras-adds-to-roster-of-nvidia-amd-broadcom.html#:~:text=OpenAI%27s%20deal%20with%20Cerebras%20is%20worth%20more%20than%20%2410%20billion" title="Highlights: OpenAI's deal with Cerebras is worth more than $10 billion" class="citation-link"><sup>[19]</sup></a>. CNBC reported Cerebras's wafer-scale chips can "deliver responses up to 15 times faster than GPU-based systems" according to the company's own release accompanying the OpenAI deal <a href="https://www.cnbc.com/2026/01/16/openai-chip-deal-with-cerebras-adds-to-roster-of-nvidia-amd-broadcom.html#:~:text=deliver%20responses%20up%20to%2015%20times%20faster%20than%20GPU-based%20systems" title="Highlights: deliver responses up to 15 times faster than GPU-based systems" class="citation-link"><sup>[63]</sup></a>. Separately, Amazon Web Services announced a deal to bring Cerebras chips into AWS data centers alongside Amazon's own Trainium3 silicon <a href="https://www.cnbc.com/2026/05/11/cerebras-raises-ipo-range.html#:~:text=Amazon%20Web%20Services%20announced%20a%20deal%20to%20bring%20Cerebras%20chips%20into%20its%20data%20centers" title="Highlights: Amazon Web Services announced a deal to bring Cerebras chips into its data centers" class="citation-link"><sup>[64]</sup></a>. These commitments underpinned Cerebras's IPO, which Nasdaq scheduled for May 14, 2026, pricing at a valuation of up to **\$48.8 billion**, more than double the **\$23 billion** private valuation the company had announced only three months earlier <a href="https://www.cnbc.com/2026/05/11/cerebras-raises-ipo-range.html#:~:text=Nasdaq%20expects%20the%20Cerebras%20IPO%20to%20take%20place%20on%20May%2014" title="Highlights: Nasdaq expects the Cerebras IPO to take place on May 14" class="citation-link"><sup>[65]</sup></a> <a href="https://www.cnbc.com/2026/05/11/cerebras-raises-ipo-range.html#:~:text=worth%20up%20to%20%2448.8%20billion%20based%20on%20the%20new%20price%20range" title="Highlights: worth up to $48.8 billion based on the new price range" class="citation-link"><sup>[17]</sup></a>.
## Strengths and Limitations
Cerebras's headline strength is raw single-model inference speed. In its own benchmarking against a 405-billion-parameter Llama 3.1 model, Cerebras reported output of 969 tokens per second, describing the result as "12x faster than best GPU result," at a published price of \$6 per million input tokens and \$12 per million output tokens <a href="https://www.cerebras.ai/blog/llama-405b-inference#:~:text=12x%20faster%20than%20best%20GPU%20result" title="Highlights: 12x faster than best GPU result" class="citation-link"><sup>[66]</sup></a> <a href="https://www.cerebras.ai/blog/llama-405b-inference#:~:text=%246%20per%20million%20input%20tokens%20and%20%2412%20per%20million%20output%20tokens" title="Highlights: $6 per million input tokens and $12 per million output tokens" class="citation-link"><sup>[67]</sup></a>. On power efficiency, Cerebras's own head-to-head comparison against an NVIDIA DGX B200 found "the CS-3 consumes 23kW peak while the DGX B200 consumes 14.3 kW," but delivers 125 petaflops versus 36 petaflops for the DGX B200, which Cerebras calculates as "a 2.2x improvement in performance per watt" <a href="https://www.cerebras.ai/blog/cerebras-cs-3-vs-nvidia-b200-2024-ai-accelerators-compared#:~:text=The%20CS-3%20consumes%2023kW%20peak%20while%20the%20DGX%20B200%20consumes%2014.3%20kW" title="Highlights: The CS-3 consumes 23kW peak while the DGX B200 consumes 14.3 kW" class="citation-link"><sup>[68]</sup></a> <a href="https://www.cerebras.ai/blog/cerebras-cs-3-vs-nvidia-b200-2024-ai-accelerators-compared#:~:text=providing%20125%20petaflops%20vs.%2036%20petaflops%20of%20the%20DGX%20B200" title="Highlights: providing 125 petaflops vs. 36 petaflops of the DGX B200" class="citation-link"><sup>[69]</sup></a> <a href="https://www.cerebras.ai/blog/cerebras-cs-3-vs-nvidia-b200-2024-ai-accelerators-compared#:~:text=a%202.2x%20improvement%20in%20performance%20per%20watt" title="Highlights: a 2.2x improvement in performance per watt" class="citation-link"><sup>[21]</sup></a>. The chief limitation is cost and manufacturing complexity: the independent arXiv analysis estimates a single CS-3 system costs approximately \$2 million to \$3 million, versus roughly \$350,000 for an NVIDIA DGX H100 and \$500,000 for a DGX B200, a 5 to 7 times price premium per system that only pencils out for buyers whose workloads specifically benefit from wafer-scale memory bandwidth and single-device simplicity <a href="https://arxiv.org/html/2503.11698v1#:~:text=The%20price%20for%20each%20CS-3%20is%20estimated%20to%20be%20%242M%20to%20%243M" title="Highlights: The price for each CS-3 is estimated to be $2M to $3M" class="citation-link"><sup>[70]</sup></a>. The same paper notes each CS-3 server "is said to consume 23KW" of power, a figure that complicates deployment in facilities not already provisioned for high-density liquid cooling <a href="https://arxiv.org/html/2503.11698v1#:~:text=Each%20CS-3%20server%20is%20said%20to%20consume%2023KW" title="Highlights: Each CS-3 server is said to consume 23KW" class="citation-link"><sup>[71]</sup></a>. Cerebras's prior fundraising history also illustrates customer-concentration risk: CNBC reported the company withdrew a 2025 IPO attempt after its prospectus revealed "heavy reliance" on a single customer, Microsoft-backed G42, which was simultaneously a major buyer and a Cerebras investor, before the OpenAI and AWS deals diversified that base ahead of the successful 2026 listing <a href="https://www.cnbc.com/2026/01/16/openai-chip-deal-with-cerebras-adds-to-roster-of-nvidia-amd-broadcom.html#:~:text=heavy%20reliance" title="Highlights: heavy reliance" class="citation-link"><sup>[72]</sup></a>.
## Feature Comparison
_Table 1_ below summarizes the core specifications of each vendor's current flagship data center accelerator, drawing on official vendor datasheets, product pages, and the independent arXiv architectural comparison.
| **Specification** | **NVIDIA GB200 NVL72 (per GPU: B200)** | **AMD Instinct MI355X** | **Intel Gaudi 3** | **Cerebras WSE-3 (CS-3 system)** |
| --- | --- | --- | --- | --- |
| **Architecture** | Blackwell, TSMC 4nm process | CDNA 4, TSMC process | Gaudi 3, matrix and tensor cores | Wafer-scale, TSMC 5nm process |
| **Peak AI compute** | Rack: 1,440 PFLOPS NVFP4 (sparse) | Up to 10.1 PFLOPS FP8 (sparsity) per GPU | Up to 1.8 PFLOPS FP8/BF16 | 125 petaflops per WSE-3 chip <a href="https://arxiv.org/html/2503.11698v1#:~:text=a%20peak%20computing%20performance%20of%20125%20petaflops" title="Highlights: a peak computing performance of 125 petaflops" class="citation-link"><sup>[62]</sup></a> |
| **Memory capacity** | 13.4TB HBM3E across rack | 288GB HBM3E per GPU | 128GB HBM2e per accelerator | 44GB on-chip SRAM plus 1.2TB to 1.2PB external MemoryX |
| **Memory bandwidth** | 576TB/s across rack | 8TB/s per GPU | 3.7TB/s per accelerator | 21 petabytes/s on-wafer |
| **Interconnect** | 130TB/s NVLink Switch across 72 GPUs | Infinity Fabric, 128GB/s per link | Standard Ethernet (RoCE) | On-wafer 2D mesh (no external interconnect needed) |
| **Approximate system power** | 14.3kW (DGX B200, 8 GPUs) <a href="https://arxiv.org/html/2503.11698v1#:~:text=The%20CS-3%20consumes%2023kW%20peak%20while%20the%20DGX%20B200%20consumes%2014.3%20kW" title="Highlights: The CS-3 consumes 23kW peak while the DGX B200 consumes 14.3 kW" class="citation-link"><sup>[73]</sup></a> | 750W per MI300X GPU <a href="https://www.amd.com/en/products/accelerators/instinct/mi300/mi300x.html#:~:text=750W%20Peak" title="Highlights: 750W Peak" class="citation-link"><sup>[41]</sup></a> | Up to 900W (water-cooled OAM) | 23kW per CS-3 system <a href="https://arxiv.org/html/2503.11698v1#:~:text=Each%20CS-3%20server%20is%20said%20to%20consume%2023KW" title="Highlights: Each CS-3 server is said to consume 23KW" class="citation-link"><sup>[71]</sup></a> |
| **Approximate list price** | \~\$500,000 (DGX B200, 8 GPUs, estimate) | Not publicly listed (OEM/hyperscale contracts) | \~\$16,000 per accelerator in volume <a href="https://www.tweaktown.com/news/98977/intel-discounting-new-gaudi-3-ai-accelerator-16k-against-nvidia-h100-gpu-for-30k/index.html#:~:text=costs%20just%20%2416%2C000%20per%20AI%20accelerator%20with%20128GB%20of%20HBM2e%20memory" title="Highlights: costs just $16,000 per AI accelerator with 128GB of HBM2e memory" class="citation-link"><sup>[12]</sup></a> | \~\$2M to \$3M per CS-3 system (estimate) <a href="https://arxiv.org/html/2503.11698v1#:~:text=The%20price%20for%20each%20CS-3%20is%20estimated%20to%20be%20%242M%20to%20%243M" title="Highlights: The price for each CS-3 is estimated to be $2M to $3M" class="citation-link"><sup>[70]</sup></a> |
| **Software stack** | CUDA, TensorRT-LLM, NVIDIA AI Enterprise | ROCm (open source) | SynapseAI, PyTorch integration | Cerebras SDK, layer-by-layer execution model |
_Table 1_ makes the divergence in design philosophy explicit; the underlying specification figures are drawn from the NVIDIA GB200 NVL72 datasheet, AMD's Instinct MI350 series and MI300X product pages, Intel's Gaudi product page, and the independent arXiv wafer-scale comparison cited throughout the Capabilities sections above. NVIDIA and AMD compete primarily on a per-GPU, rack-aggregated basis, where discrete accelerators are networked together and the interconnect fabric becomes a first-order performance variable, whether that is NVIDIA's proprietary 130TB/s NVLink Switch or AMD's Infinity Fabric. Intel's Gaudi 3 undercuts both on acquisition price while trading away proprietary interconnect performance for standard Ethernet compatibility. Cerebras inverts the entire model: instead of scaling out across many chips, it scales a single chip to the size of an entire silicon wafer, eliminating inter-chip communication for workloads that fit within one WSE-3, at the cost of a substantially higher unit price and power draw per system. Budget-tier options exist on every vendor's roadmap, from AMD's MI350P PCIe card (144GB HBM3E, 4TB/s bandwidth) to Intel's lower-wattage 450W Gaudi 3 configurations, giving buyers a way to trade peak throughput for lower acquisition and operating cost within a single vendor family.
## Performance and Benchmarks
MLCommons, the industry consortium behind the MLPerf benchmark suite, published **MLPerf Training v6.0** results in June 2026 and **MLPerf Inference v6.0** results in April 2026, providing the most current independently audited data on training and inference throughput. NVIDIA submitted results across every benchmark and reported the fastest time-to-train at scale on every category it entered, including a DeepSeek-V3 671B mixture-of-experts model trained in 2.02 minutes on a GB300 NVL72 cluster of 8,192 GPUs <a href="https://developer.nvidia.com/blog/nvidia-blackwell-tops-mlperf-training-6-0-with-industry-leading-scale-and-performance/#:~:text=2.02%20mins" title="Highlights: 2.02 mins" class="citation-link"><sup>[74]</sup></a>. NVIDIA also reported that software optimizations alone lifted GB300 training throughput on DeepSeek-V3 from 1,298 to 1,648 TFLOPS per GPU in three months without any hardware changes, describing "DeepSeek-V3 training throughput increasing by 1.3x in just three months" through kernel fusion, CUDA graph, and communication-overlap improvements <a href="https://developer.nvidia.com/blog/nvidia-blackwell-tops-mlperf-training-6-0-with-industry-leading-scale-and-performance/#:~:text=DeepSeek-V3%20training%20throughput%20increasing%20by%201.3x%20in%20just%20three%20months" title="Highlights: DeepSeek-V3 training throughput increasing by 1.3x in just three months" class="citation-link"><sup>[75]</sup></a>.
On the inference side, an April 2026 analysis of MLPerf Inference v6.0 results derived per-GPU throughput by dividing system-level scores by accelerator count. _Table 2_ below summarizes the closed-division, per-GPU results for the GPT-OSS 120B and LLaMA 2 70B language model benchmarks, plus the Stable Diffusion XL (SDXL) image-generation and YOLOv11 object-detection tasks.
| **Accelerator** | **GPT-OSS 120B (tok/s per GPU)** | **LLaMA 2 70B (tok/s per GPU)** | **SDXL (samples/s per GPU)** | **YOLOv11 (QPS per GPU)** |
| --- | --- | --- | --- | --- |
| NVIDIA B200 (GB200 NVL72) | \~7,800 | \~17,500 | \~14.2 | \~28,400 |
| NVIDIA H200 SXM | \~3,100 | \~7,800 | \~6.3 | \~15,200 |
| AMD MI355X (closed division) | \~2,600 | \~6,200 | \~4.8 (open division) | \~11,800 (open division) |
_Source: MLCommons Inference v6.0, closed division unless noted, as analyzed by Spheron Network in April 2026_ (Source: [spheron.network](https://www.spheron.network/blog/mlperf-inference-v6-benchmark-results-2026/#:~:text=roughly%202.5x%20the%20per-GPU%20throughput%20of%20the%20H200)).
On the GPT-OSS-120B benchmark, NVIDIA's B200 delivered roughly 2.5 times the per-GPU throughput of the H200, a gap the analysis attributes primarily to the combination of FP4 precision support and 8.0TB/s of HBM3e bandwidth versus the H200's 4.8TB/s (Source: [spheron.network](https://www.spheron.network/blog/mlperf-inference-v6-benchmark-results-2026/#:~:text=roughly%202.5x%20the%20per-GPU%20throughput%20of%20the%20H200)). AMD's closed-division submission, however, put the MI355X within 20% of the H200 on both LLaMA 2 70B and GPT-OSS-120B, figures the analysis notes are directly comparable because both vendors ran under identical model and precision constraints, with AMD's larger 288GB memory capacity allowing bigger batch sizes on models that fit within it (Source: [spheron.network](https://www.spheron.network/blog/mlperf-inference-v6-benchmark-results-2026/#:~:text=AMD%27s%20closed%20division%20submission%20puts%20MI355X%20within%2020%25%20of%20the%20H200)). Notably, Intel's official v6.0 submissions covered Xeon 6 CPUs and Arc Pro GPU workloads rather than Gaudi 3, leaving a gap in the most current independently audited inference dataset for Intel's flagship AI accelerator (Source: [spheron.network](https://www.spheron.network/blog/mlperf-inference-v6-benchmark-results-2026/#:~:text=Intel%27s%20official%20v6.0%20submissions%20covered%20Xeon%206%20CPU%20and%20Arc%20Pro%20GPU%20workloads)).
Independent, vendor-adjacent benchmarking from Artificial Analysis offers a complementary view focused on realistic production loads rather than the standardized MLPerf harness. Using its own System Load Test methodology on DeepSeek R1 and Llama 4 Maverick across matched 8-GPU NVIDIA H100, H200, and AMD MI300X systems, Artificial Analysis found NVIDIA held an advantage in per-query output speed and latency at low concurrency, while AMD's MI300X sustained near-perfect response rates all the way to a higher peak system throughput, translating into a lower implied cost per token at scale <a href="https://artificialanalysis.ai/articles/independent-analysis-of-leading-gpus-amd-nvidia#:~:text=NVIDIA%20H100%20%26%20H200%20systems%20demonstrate%20marginally%20faster%20average%20per-query%20output%20speeds" title="Highlights: NVIDIA H100 & H200 systems demonstrate marginally faster average per-query output speeds" class="citation-link"><sup>[51]</sup></a>. For accuracy, the same study found no statistically significant difference between AMD and NVIDIA systems on GPQA, MMLU-Pro, and MATH-500 evaluations at a 95% confidence interval, indicating that hardware choice among GPU vendors does not, on its own, materially affect model output quality.
Cerebras does not participate in the standard MLPerf Inference harness, instead publishing benchmarks independently verified by Artificial Analysis. On Llama 3.1 405B, Cerebras reported 969 output tokens per second and a 240-millisecond time to first token, roughly 12 times faster than the best comparable GPU-based result at the time of publication, and 8 times faster than SambaNova and 75 times faster than AWS on the same evaluation <a href="https://www.cerebras.ai/blog/llama-405b-inference#:~:text=12x%20faster%20than%20best%20GPU%20result" title="Highlights: 12x faster than best GPU result" class="citation-link"><sup>[66]</sup></a>. Because Cerebras, NVIDIA, and AMD use different benchmarking harnesses, Cerebras favoring single-request latency and token throughput and MLCommons favoring standardized offline and server scenarios, buyers evaluating "best ai accelerator for training and inference" claims should treat vendor-reported multiples cautiously and, where possible, request workload-specific proof-of-concept testing before committing to a platform.
## Data Analysis and Evidence
The quantitative case for NVIDIA's continued market leadership rests on both revenue and independent market-share estimation. NVIDIA's fiscal 2026 Data Center segment generated **\$193.7 billion**, up 68% year over year, with the fourth quarter alone contributing a record **\$62.3 billion**, up 75% from the prior year <a href="http://nvidianews.nvidia.com/news/nvidia-announces-financial-results-for-fourth-quarter-and-fiscal-2026#:~:text=Full-year%20revenue%20rose%2068%25%20to%20a%20record%20%24193.7%20billion" title="Highlights: Full-year revenue rose 68% to a record $193.7 billion" class="citation-link"><sup>[2]</sup></a> <a href="http://nvidianews.nvidia.com/news/nvidia-announces-financial-results-for-fourth-quarter-and-fiscal-2026#:~:text=Record%20quarterly%20Data%20Center%20revenue%20of%20%2462.3%20billion%2C%20up%2022%25%20from%20Q3%20and%20up%2075%25%20from%20a%20year%20ago" title="Highlights: Record quarterly Data Center revenue of $62.3 billion, up 22% from Q3 and up 75% from a year ago" class="citation-link"><sup>[1]</sup></a>. Mordor Intelligence's independent industry analysis separately estimates NVIDIA held approximately 80% of global AI training revenue in 2024, a concentration the research firm attributes to "CUDA lock-in, integrated software libraries, and a mature partner ecosystem" <a href="https://www.mordorintelligence.com/industry-reports/ai-accelerators-market#:~:text=NVIDIA%20retained%20about%2080%25%20of%20global%20training%20revenue%20in%202024" title="Highlights: NVIDIA retained about 80% of global training revenue in 2024" class="citation-link"><sup>[4]</sup></a>. Even accounting for accelerating diversification toward AMD, Intel, and custom silicon, GPUs as a processor category retained an estimated 59.20% revenue share of the overall AI accelerator market in 2025, with the balance split between application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), and CPU and neural processing unit (NPU) hybrids <a href="https://www.mordorintelligence.com/industry-reports/ai-accelerators-market#:~:text=GPUs%20held%2059.20%25%20revenue%20share%20of%20the%20AI%20accelerators%20market%20in%202025" title="Highlights: GPUs held 59.20% revenue share of the AI accelerators market in 2025" class="citation-link"><sup>[25]</sup></a>.
_Table 3_ below assembles the most recent quarterly financial data for the three publicly traded vendors, alongside Cerebras's newly disclosed public-market valuation, to illustrate the scale gap that still separates NVIDIA from its challengers even as those challengers post strong percentage growth.
| **Company** | **Most Recent Quarterly Data Center / AI Revenue** | **Year-over-Year Growth** | **Source Period** |
| --- | --- | --- | --- |
| NVIDIA | \$62.3 billion (Data Center, Q4 FY2026) <a href="http://nvidianews.nvidia.com/news/nvidia-announces-financial-results-for-fourth-quarter-and-fiscal-2026#:~:text=Record%20quarterly%20Data%20Center%20revenue%20of%20%2462.3%20billion%2C%20up%2022%25%20from%20Q3%20and%20up%2075%25%20from%20a%20year%20ago" title="Highlights: Record quarterly Data Center revenue of $62.3 billion, up 22% from Q3 and up 75% from a year ago" class="citation-link"><sup>[1]</sup></a> | +75% | Quarter ended January 25, 2026 |
| AMD | \$5.8 billion (Data Center segment, Q1 2026), \$10.25 billion total company revenue <a href="https://www.datacenterdynamics.com/en/news/amd-posts-q1-2026-data-center-revenue-of-58bn-forecasts-120bn-server-cpu-income-by-2030/#:~:text=AMD%20saw%20a%2057%20percent%20year-on-year%20%28YoY%29%20revenue%20increase%20for%20its%20data%20center%20segment" title="Highlights: AMD saw a 57 percent year-on-year (YoY) revenue increase for its data center segment" class="citation-link"><sup>[8]</sup></a> | +57% (Data Center); +38% (total) | Quarter ended late March 2026 |
| Intel | \$5.1 billion (Data Center and AI segment, Q1 2026); \$13.6 billion total revenue <a href="https://www.intc.com/news-events/press-releases/detail/1767/intel-reports-first-quarter-2026-financial-results#:~:text=sixth%20consecutive%20quarter%20of%20revenue%20above%20our%20expectations" title="Highlights: sixth consecutive quarter of revenue above our expectations" class="citation-link"><sup>[14]</sup></a> | +22% (DCAI); +7% (total) | Quarter ended March 28, 2026 |
| Cerebras | Not yet reported as a public company; \$48.8 billion IPO valuation <a href="https://www.cnbc.com/2026/05/11/cerebras-raises-ipo-range.html#:~:text=worth%20up%20to%20%2448.8%20billion%20based%20on%20the%20new%20price%20range" title="Highlights: worth up to $48.8 billion based on the new price range" class="citation-link"><sup>[17]</sup></a> | Valuation up from \$23 billion three months prior | IPO priced May 14, 2026 |
_Table 3_ shows NVIDIA's single quarter of Data Center revenue exceeds the combined trailing quarterly Data Center or AI revenue of AMD and Intel by more than a factor of five, underscoring why analysts continue to describe the market as NVIDIA-led rather than genuinely multipolar, even as AMD's 57% and Intel's 22% year-over-year segment growth rates indicate real share gains are occurring at the margin, and AMD's own guidance for the following quarter called for total revenue of \$11.2 billion <a href="https://www.datacenterdynamics.com/en/news/amd-posts-q1-2026-data-center-revenue-of-58bn-forecasts-120bn-server-cpu-income-by-2030/#:~:text=AMD%20saw%20a%2057%20percent%20year-on-year%20%28YoY%29%20revenue%20increase%20for%20its%20data%20center%20segment" title="Highlights: AMD saw a 57 percent year-on-year (YoY) revenue increase for its data center segment" class="citation-link"><sup>[8]</sup></a>. On power efficiency specifically, the most rigorous side-by-side data point comes from the independent arXiv comparison, which normalizes NVIDIA and Cerebras hardware to equivalent rack space and finds a CS-3-based rack (2 systems, 46kW total) delivers substantially more raw FLOPS per rack than a comparably sized H100 rack (32 GPUs, 41.6kW) or B200 rack (24 GPUs, 43.9kW), even though the CS-3 rack draws marginally more total power <a href="https://arxiv.org/html/2503.11698v1#:~:text=Each%20CS-3%20server%20is%20said%20to%20consume%2023KW" title="Highlights: Each CS-3 server is said to consume 23KW" class="citation-link"><sup>[71]</sup></a>. The paper's abstract concludes the results highlight the advantages of wafer-scale integration in performance per watt and memory scalability, while cautioning that "work is required to address cost-effectiveness and long-term viability" of the approach at commercial scale <a href="https://arxiv.org/html/2503.11698v1#:~:text=advantages%20of%20WSE-3%20in%20performance%20per%20watt%20and%20memory%20scalability" title="Highlights: advantages of WSE-3 in performance per watt and memory scalability" class="citation-link"><sup>[22]</sup></a>.
On the training-versus-inference split that determines which architecture matters most for a given buyer, Mordor Intelligence estimates training consumed 57.30% of 2025 AI accelerator revenue, while inference spending is forecast to grow at a faster 26.10% CAGR through 2031 as more models move from research into always-on production serving <a href="https://www.mordorintelligence.com/industry-reports/ai-accelerators-market#:~:text=Training%20consumed%2057.30%25%20of%202025%20revenue" title="Highlights: Training consumed 57.30% of 2025 revenue" class="citation-link"><sup>[76]</sup></a>. That shift favors architectures optimized for latency and cost-per-token, a category where Cerebras's inference-focused wafer-scale design and AMD's memory-dense MI300X and MI350 series are increasingly positioned to compete, even where they trail NVIDIA on aggregate training throughput. Beyond the four vendors profiled here, broader industry data underlines both the scale and the geographic concentration of the AI accelerator buildout: North America commanded a 43.50% share of the global AI accelerator market in 2025, while Asia-Pacific posted the fastest growth at a 27.00% CAGR, driven partly by Chinese firms raising their domestic AI chip budget share from 30% to 46% amid the shift away from NVIDIA described above <a href="https://www.trendforce.com/news/2026/07/07/news-chinese-firms-reportedly-raise-domestic-ai-chip-budget-share-from-30-to-46-amid-shift-from-nvidia/#:~:text=executives%20surveyed%20expect%20domestic%20products%20to%20account%20for%2046%25%20of%20their%20AI%20accelerator%20budgets%20over%20the%20next%2012%20months%2C%20up%20from%2030%25%20currently" title="Highlights: executives surveyed expect domestic products to account for 46% of their AI accelerator budgets over the next 12 months, up from 30% currently" class="citation-link"><sup>[77]</sup></a>.
## Case Studies and Real-World Examples
## xAI's Colossus and the NVIDIA Hopper-to-Blackwell Buildout
xAI's Colossus supercomputer, built in partnership with Supermicro, connects 100,000 NVIDIA Hopper Tensor Core GPUs using NVIDIA's Spectrum-X Ethernet networking platform, and Supermicro describes it as designed "to take xAI's Grok AI to another era" through liquid-cooled, direct-to-chip cooling at unprecedented density <a href="https://www.supermicro.com/en/featured/xai-colossus#:~:text=connect%20100%2C000%20NVIDIA%20Hopper%20Tensor%20Core%20GPUs" title="Highlights: connect 100,000 NVIDIA Hopper Tensor Core GPUs" class="citation-link"><sup>[33]</sup></a>. The cluster illustrates both the scale NVIDIA's ecosystem can support and the infrastructure demands that scale imposes: 100,000 H100-class GPUs at up to 700W TDP each implies tens of megawatts of continuous power draw before accounting for cooling and networking overhead, a demand profile that has become a defining constraint on where hyperscale AI training clusters can physically be sited. The Colossus build also demonstrates the practical value of NVIDIA's Spectrum-X Ethernet and NVLink ecosystem at scale, since coordinating 100,000 GPUs across a single logical training job depends on the kind of mature networking and orchestration software that remains NVIDIA's most defensible advantage against GPU-based challengers.
## Meta and AMD's 6-Gigawatt Instinct Partnership
In February 2026, AMD and Meta announced what AMD called a "definitive multi-year, multi-generation partnership to deploy up to 6 gigawatts of AMD Instinct GPUs," with the first gigawatt of shipments, based on a custom MI450-architecture GPU paired with AMD's sixth-generation "Venice" EPYC CPUs, scheduled to begin in the second half of 2026 <a href="https://ir.amd.com/news-events/press-releases/detail/1279/amd-and-meta-announce-expanded-strategic-partnership-to-deploy-6-gigawatts-of-amd-gpus#:~:text=agree%20to%20a%20definitive%20multi-year%2C%20multi-generation%20partnership%20to%20deploy%20up%20to%206%20gigawatts" title="Highlights: agree to a definitive multi-year, multi-generation partnership to deploy up to 6 gigawatts" class="citation-link"><sup>[10]</sup></a>. Meta CEO Mark Zuckerberg framed the agreement around compute diversification, stating the company was "excited to form a long-term partnership with AMD to deploy efficient inference compute" as part of a broader strategy to avoid single-vendor dependency. AMD structured the deal with a performance-based warrant for up to 160 million shares of AMD common stock, vesting as Meta's Instinct GPU purchases scale from the initial 1 gigawatt toward the full 6-gigawatt commitment <a href="https://ir.amd.com/news-events/press-releases/detail/1279/amd-and-meta-announce-expanded-strategic-partnership-to-deploy-6-gigawatts-of-amd-gpus#:~:text=warrant%20for%20up%20to%20160%20million%20shares%20of%20AMD%20common%20stock" title="Highlights: warrant for up to 160 million shares of AMD common stock" class="citation-link"><sup>[78]</sup></a>. This deal built on Meta's existing production use of MI300X, which by late 2024 was already "serving all live traffic on Llama 405B" according to AMD's own disclosure <a href="https://ir.amd.com/news-events/press-releases/detail/1218/amd-unveils-leadership-ai-solutions-at-advancing-ai-2024#:~:text=MI300X%20serving%20all%20live%20traffic%20on%20Llama%20405B" title="Highlights: MI300X serving all live traffic on Llama 405B" class="citation-link"><sup>[45]</sup></a>, making Meta one of the clearest examples of a hyperscaler moving an AMD accelerator from pilot to primary production inference infrastructure at gigawatt scale.
## IBM Cloud and Intel Gaudi 3 for Regulated-Industry Inference
IBM Cloud became the first cloud service provider to offer Intel Gaudi 3 accelerators as a service, targeting enterprise customers in highly regulated industries such as financial services, government, healthcare, and telecommunications who require hybrid cloud deployment options and strict AI governance <a href="https://www.intel.com/content/www/us/en/customer-spotlight/stories/ibm-gaudi-3-customer-story.html#:~:text=over%205%2C000%20tokens%20per%20second%20for%20IBM%27s%20granite-8b%20model" title="Highlights: over 5,000 tokens per second for IBM's granite-8b model" class="citation-link"><sup>[57]</sup></a>. IBM's own internal testing found a single Gaudi 3 card could generate over 5,000 tokens per second on IBM's granite-8b model while supporting over 100 concurrent users at under 20 milliseconds of inter-token latency, and the architecture is designed to scale linearly from an 8-accelerator single node up to a 1,024-node, 8,192-accelerator cluster capable of 9.830 petabytes per second of aggregate throughput <a href="https://www.intel.com/content/www/us/en/customer-spotlight/stories/ibm-gaudi-3-customer-story.html#:~:text=1%2C024-node%20cluster%20%288%2C192%20accelerators%29%20with%20a%20throughput%20of%209.830%20PB%2Fs" title="Highlights: 1,024-node cluster (8,192 accelerators) with a throughput of 9.830 PB/s" class="citation-link"><sup>[58]</sup></a>. This case demonstrates the specific niche Gaudi 3 is chasing: cost-conscious enterprise inference deployment inside a hybrid cloud, rather than frontier-scale training, where IBM's messaging emphasizes total cost of ownership and an open ecosystem over raw peak performance.
## OpenAI, Cerebras, and the Wafer-Scale Inference Bet
In January 2026, OpenAI signed an agreement to deploy 750 megawatts of Cerebras wafer-scale chips in tranches through 2028, a deal CNBC valued at more than \$10 billion <a href="https://www.cnbc.com/2026/01/16/openai-chip-deal-with-cerebras-adds-to-roster-of-nvidia-amd-broadcom.html#:~:text=an%20agreement%20to%20deploy%20750%20megawatts%20of%20Cerebras%27%20AI%20chips" title="Highlights: an agreement to deploy 750 megawatts of Cerebras' AI chips" class="citation-link"><sup>[20]</sup></a>. Cerebras's release describing the partnership claimed its chips can "deliver responses up to 15 times faster than GPU-based systems," positioning the deployment specifically around a model that "writes code," a latency-sensitive workload where Cerebras's single-request throughput advantage is most pronounced <a href="https://www.cnbc.com/2026/01/16/openai-chip-deal-with-cerebras-adds-to-roster-of-nvidia-amd-broadcom.html#:~:text=deliver%20responses%20up%20to%2015%20times%20faster%20than%20GPU-based%20systems" title="Highlights: deliver responses up to 15 times faster than GPU-based systems" class="citation-link"><sup>[63]</sup></a>. This case is notable for its context: Cerebras had previously withdrawn a planned 2025 IPO amid investor concern over "heavy reliance" on a single customer, Microsoft-backed G42, which was both a major buyer and a Cerebras investor <a href="https://www.cnbc.com/2026/01/16/openai-chip-deal-with-cerebras-adds-to-roster-of-nvidia-amd-broadcom.html#:~:text=heavy%20reliance" title="Highlights: heavy reliance" class="citation-link"><sup>[72]</sup></a>. The OpenAI deal, layered on top of the subsequent AWS partnership, materially diversified Cerebras's customer concentration in the months immediately before its successful May 2026 IPO, illustrating how a single marquee inference deployment can reshape a chip vendor's fundraising trajectory as much as any benchmark result.
## Implications and Future Directions
The near-term trajectory of the AI accelerator market points toward continued architectural specialization rather than convergence on a single winning design. NVIDIA's own roadmap disclosures suggest this: the company's Rubin platform, unveiled alongside fiscal 2026 results, is designed to deliver "up to a 10x reduction in inference token cost, compared with the NVIDIA Blackwell platform," an explicit acknowledgment that inference economics, not just raw training throughput, now drive purchasing decisions at the margin <a href="http://nvidianews.nvidia.com/news/nvidia-announces-financial-results-for-fourth-quarter-and-fiscal-2026#:~:text=up%20to%20a%2010x%20reduction%20in%20inference%20token%20cost%2C%20compared%20with%20the%20NVIDIA%20Blackwell%20platform" title="Highlights: up to a 10x reduction in inference token cost, compared with the NVIDIA Blackwell platform" class="citation-link"><sup>[36]</sup></a>. Mordor Intelligence's forecast that inference will grow at a 26.10% CAGR through 2031, faster than training's growth rate, reinforces the expectation that future accelerator generations across all four vendors will increasingly optimize for cost-per-token and latency rather than peak FLOPS alone.
Power availability is emerging as an equally consequential constraint as raw silicon supply. Both the AMD-Meta and OpenAI-Cerebras deals are structured and reported in gigawatts rather than unit counts, a framing shift that reflects the reality that data center electricity capacity, not chip fabrication, is now frequently the binding constraint on deployment timelines. NVIDIA's own \$100 billion OpenAI commitment is likewise denominated in gigawatts, with the company estimating 10 gigawatts equates to 4 million to 5 million GPUs <a href="https://www.cnbc.com/2026/01/16/openai-chip-deal-with-cerebras-adds-to-roster-of-nvidia-amd-broadcom.html#:~:text=commit%20%24100%20billion%20to%20support%20OpenAI%20as%20it%20builds%20and%20deploys%20at%20least%2010%20gigawatts" title="Highlights: commit $100 billion to support OpenAI as it builds and deploys at least 10 gigawatts" class="citation-link"><sup>[34]</sup></a>. This dynamic advantages accelerators with superior performance-per-watt, a metric where Cerebras's arXiv-verified claims and NVIDIA's own claimed per-watt improvement for GB200 over air-cooled H100 infrastructure, cited earlier in this report, will likely become as commercially important as headline FLOPS figures.
Geopolitical and supply-chain risk will also continue to shape vendor strategy. NVIDIA's \$4.5 billion H20 export-control charge in a single quarter of fiscal 2026 demonstrates how quickly regulatory changes can disrupt even the market leader's revenue recognition <a href="https://www.sec.gov/Archives/edgar/data/1045810/000104581025000115/q1fy26pr.htm#:~:text=%244.5%20billion%20charge%20in%20the%20first%20quarter%20of%20fiscal%202026%20associated%20with%20H20%20excess%20inventory" title="Highlights: $4.5 billion charge in the first quarter of fiscal 2026 associated with H20 excess inventory" class="citation-link"><sup>[39]</sup></a>, and TrendForce's July 2026 reporting that Bernstein projects NVIDIA's Chinese AI semiconductor market share falling to roughly 8% in 2026, with domestic champion Huawei's share predicted to surpass 50%, creates an opening for both Chinese domestic accelerators and, indirectly, for AMD, Intel, and Cerebras to capture non-China demand that might otherwise have concentrated further with NVIDIA <a href="https://www.trendforce.com/news/2026/07/07/news-chinese-firms-reportedly-raise-domestic-ai-chip-budget-share-from-30-to-46-amid-shift-from-nvidia/#:~:text=investment%20bank%20Bernstein%20projected%20that%20NVIDIA%E2%80%99s%20share%20of%20the%20Chinese%20AI%20semiconductor%20market%20will%20fall%20to%20around%208%25%20in%202026" title="Highlights: investment bank Bernstein projected that NVIDIA’s share of the Chinese AI semiconductor market will fall to around 8% in 2026" class="citation-link"><sup>[40]</sup></a>. Finally, custom silicon from hyperscalers and Broadcom's ASIC design partnerships, forecast by Mordor Intelligence to represent a \$60 billion to \$90 billion opportunity by 2027 for a company CNBC now values at over \$1.6 trillion, will continue to erode merchant GPU share at the margin, particularly for steady-state inference workloads where a purpose-built ASIC's efficiency advantage outweighs the flexibility of a general-purpose GPU <a href="https://www.mordorintelligence.com/industry-reports/ai-accelerators-market#:~:text=Broadcom%20anticipates%20a%20USD%2060%E2%80%9390%20billion%20ASIC%20opportunity%20by%202027" title="Highlights: Broadcom anticipates a USD 60–90 billion ASIC opportunity by 2027" class="citation-link"><sup>[27]</sup></a> <a href="https://www.cnbc.com/2026/01/16/openai-chip-deal-with-cerebras-adds-to-roster-of-nvidia-amd-broadcom.html#:~:text=Broadcom%20is%20now%20valued%20at%20over%20%241.6%20trillion" title="Highlights: Broadcom is now valued at over $1.6 trillion" class="citation-link"><sup>[28]</sup></a>.
## Frequently Asked Questions (FAQs)
**Which AI accelerator has the best power efficiency in 2026?** On a raw performance-per-watt basis for training, Cerebras's own analysis claims the CS-3 delivers roughly 2.2 times the performance per watt of an NVIDIA DGX B200 <a href="https://www.cerebras.ai/blog/cerebras-cs-3-vs-nvidia-b200-2024-ai-accelerators-compared#:~:text=a%202.2x%20improvement%20in%20performance%20per%20watt" title="Highlights: a 2.2x improvement in performance per watt" class="citation-link"><sup>[21]</sup></a>, while NVIDIA claims its GB200 NVL72 delivers 25 times more performance at the same power as air-cooled H100 infrastructure specifically. Because these are vendor-reported, non-identical comparisons (one against a same-generation GPU competitor, the other against the vendor's own prior generation), buyers should treat both as directionally informative rather than directly comparable, and should request workload-specific power measurements before finalizing procurement.
**How does NVIDIA H100 compare to AMD MI300X and Cerebras WSE-3 in practice?** The H100 offers 80GB of memory and up to 3,958 TFLOPS of FP8 compute with sparsity at up to 700W <a href="https://www.nvidia.com/en-us/data-center/h100/#:~:text=Up%20to%20700W%20%28configurable%29" title="Highlights: Up to 700W (configurable)" class="citation-link"><sup>[29]</sup></a>, while the MI300X offers 192GB of memory at 750W <a href="https://www.amd.com/en/products/accelerators/instinct/mi300/mi300x.html#:~:text=750W%20Peak" title="Highlights: 750W Peak" class="citation-link"><sup>[41]</sup></a> and independent testing shows it wins on high-concurrency throughput while the H100 wins on low-concurrency latency <a href="https://artificialanalysis.ai/articles/independent-analysis-of-leading-gpus-amd-nvidia#:~:text=AMD%20MI300X%20system%20achieved%2025-35%25%20higher%20peak%20system%20output%20throughput" title="Highlights: AMD MI300X system achieved 25-35% higher peak system output throughput" class="citation-link"><sup>[11]</sup></a>. Cerebras's WSE-3 is not a like-for-like comparison at the chip level since a single WSE-3 replaces an entire rack of GPUs; the arXiv comparison paper found a CS-3-equipped rack outperforms both H100 and B200 racks on raw compute density while costing several times more per system <a href="https://arxiv.org/html/2503.11698v1#:~:text=The%20price%20for%20each%20CS-3%20is%20estimated%20to%20be%20%242M%20to%20%243M" title="Highlights: The price for each CS-3 is estimated to be $2M to $3M" class="citation-link"><sup>[70]</sup></a>.
**Is Intel Gaudi competitive with NVIDIA and AMD for AI training and inference?** Gaudi 3 is positioned primarily for inference and price-sensitive deployments rather than frontier-scale training, with an approximately 2 times lower acquisition cost than an H100 <a href="https://www.tweaktown.com/news/98977/intel-discounting-new-gaudi-3-ai-accelerator-16k-against-nvidia-h100-gpu-for-30k/index.html#:~:text=compared%20to%20%2430%2C000%2B%20for%20NVIDIA%27s%20Hopper%20H100%20AI%20GPU" title="Highlights: compared to $30,000+ for NVIDIA's Hopper H100 AI GPU" class="citation-link"><sup>[60]</sup></a>, but the absence of Gaudi 3 results in the most recent MLPerf Inference v6.0 dataset means buyers cannot currently benchmark it against NVIDIA and AMD using the same independently audited methodology (Source: [spheron.network](https://www.spheron.network/blog/mlperf-inference-v6-benchmark-results-2026/#:~:text=Intel%27s%20official%20v6.0%20submissions%20covered%20Xeon%206%20CPU%20and%20Arc%20Pro%20GPU%20workloads)).
**What is the Cerebras Wafer-Scale Engine and how does it differ from a GPU?** The WSE-3 is a single 46,225 square millimeter chip fabricated from an entire silicon wafer, carrying 900,000 cores and 44GB of on-chip memory, versus a GPU's die size of roughly 800 to 1,600 square millimeters <a href="https://www.cerebras.ai/chip#:~:text=The%20WSE-3%20is%20the%20largest%20AI%20chip%20ever%20built" title="Highlights: The WSE-3 is the largest AI chip ever built" class="citation-link"><sup>[15]</sup></a> <a href="https://arxiv.org/html/2503.11698v1#:~:text=900%2C000%20AI-optimized%20cores%2C%20and%2044%20GB%20of%20on-chip%20SRAM" title="Highlights: 900,000 AI-optimized cores, and 44 GB of on-chip SRAM" class="citation-link"><sup>[61]</sup></a>. Because all cores communicate over an on-wafer mesh rather than external interconnects, the WSE-3 eliminates a major source of latency and power overhead present in multi-GPU clusters, at the cost of a single point of failure per wafer and a substantially higher unit price.
**How big is the overall AI accelerator market, and who is growing fastest?** Mordor Intelligence estimates the market at \$174.69 billion in 2026, growing to \$518.12 billion by 2031 at a 24.30% CAGR <a href="https://www.mordorintelligence.com/industry-reports/ai-accelerators-market#:~:text=forecast%20to%20reach%20USD%20518.12%20billion%20by%202031%20at%2024.30%25%20CAGR" title="Highlights: forecast to reach USD 518.12 billion by 2031 at 24.30% CAGR" class="citation-link"><sup>[5]</sup></a>. By percentage growth, AMD's 57% year-over-year Data Center revenue increase and Intel's 22% Data Center and AI increase in Q1 2026 both outpaced NVIDIA's underlying quarterly growth rate on a relative basis, even though NVIDIA still adds more absolute revenue in a single quarter than either competitor generates in several <a href="https://www.datacenterdynamics.com/en/news/amd-posts-q1-2026-data-center-revenue-of-58bn-forecasts-120bn-server-cpu-income-by-2030/#:~:text=AMD%20saw%20a%2057%20percent%20year-on-year%20%28YoY%29%20revenue%20increase%20for%20its%20data%20center%20segment" title="Highlights: AMD saw a 57 percent year-on-year (YoY) revenue increase for its data center segment" class="citation-link"><sup>[8]</sup></a>.
**Which accelerator is best for training versus inference?** Training workloads at frontier scale remain best served by NVIDIA's GB200/GB300 NVL72 platforms given their independently verified MLPerf Training v6.0 leadership and mature distributed-training software stack <a href="https://developer.nvidia.com/blog/nvidia-blackwell-tops-mlperf-training-6-0-with-industry-leading-scale-and-performance/#:~:text=was%20the%20only%20platform%20to%20submit%20on%20every%20test" title="Highlights: was the only platform to submit on every test" class="citation-link"><sup>[23]</sup></a>. For latency-sensitive, single-model inference serving, Cerebras's CS-3 and AMD's memory-dense MI300X/MI350 series offer compelling cost and speed advantages, per the vendor and independent benchmarks cited throughout this report, while Intel's Gaudi 3 remains the lowest-cost entry point for enterprises prioritizing budget over peak throughput.
## Conclusion
As of July 2026, the "nvidia vs amd vs intel vs cerebras" comparison resolves less into a single winner than into four distinct value propositions matched to different buyer priorities. NVIDIA remains the default choice for organizations that need the broadest software ecosystem, the largest verified training scale, and access to the newest architectures first, backed by fiscal 2026 Data Center revenue of \$193.7 billion and an estimated 80% share of global AI training revenue. AMD has emerged as the most credible GPU-based alternative, winning gigawatt-scale commitments from Meta and OpenAI on the strength of memory-dense Instinct accelerators and an increasingly capable open-source ROCm stack, even as CUDA's software head start persists as a real, if narrowing, gap. Intel's Gaudi 3 occupies a genuine niche among cost-conscious enterprise buyers, particularly in regulated industries served through partners like IBM Cloud, though its absence from the latest independently audited MLPerf Inference results leaves a meaningful evidence gap for buyers who prioritize standardized benchmarking. Cerebras has translated its architecturally unique wafer-scale approach into rapid public-market validation, closing a more than \$10 billion OpenAI deal and completing 2026's largest IPO at a valuation of up to \$48.8 billion, built on demonstrated single-model inference speed rather than training-cluster scale.
For technical buyers, the practical takeaway is that accelerator selection should be workload-driven rather than brand-driven. Frontier pretraining runs at trillion-parameter scale still favor NVIDIA's NVLink-connected GPU racks; memory-bound inference at high concurrency increasingly favors AMD's larger HBM capacity; budget-constrained enterprise inference deployments inside existing Ethernet infrastructure favor Intel's Gaudi 3; and latency-critical, single-model serving at extreme token throughput favors Cerebras's wafer-scale systems. Given how quickly benchmark leadership has shifted across MLPerf Training v6.0 and Inference v6.0 rounds within a single year, buyers should treat any static comparison, including this one, as a snapshot that warrants periodic re-verification against the latest MLCommons results and vendor financial disclosures before committing to multi-year procurement contracts measured in gigawatts rather than individual chips.
External Sources
About GPUSmith
GPU Smith is an independent engineering firm that specifies, procures, integrates and validates private AI compute infrastructure on Nvidia reference architectures, from a single inference node to multi-megawatt compute halls. Every engagement is delivered against written acceptance criteria and an as-built documentation set, with procurement at a disclosed margin and no reseller quota or cloud of its own. Six disciplines: hardware integration and commissioning; cluster architecture and sizing; inference build-out; serving optimization; datacenter operations; and sovereign/air-gapped systems. Core thesis: at sustained load, the amortized cost of owned hardware falls below per-token cloud and API pricing, and GPU Smith locates that crossover for a defined workload and states build/no-build in writing. Sectors served: government and regulated enterprise (bounded inference), scaling AI teams past the ownership crossover, and investors/operators needing technical due diligence.
DISCLAIMER
This document is provided for informational purposes only. No representations or warranties are made regarding the accuracy, completeness, or reliability of its contents. Any use of this information is at your own risk. GPUSmith shall not be liable for any damages arising from the use of this document. This content may include material generated with assistance from artificial intelligence tools, which may contain errors or inaccuracies. Readers should verify critical information independently. All product names, trademarks, and registered trademarks mentioned are property of their respective owners and are used for identification purposes only. Use of these names does not imply endorsement. This document does not constitute professional or legal advice. For specific guidance related to your needs, please consult qualified professionals.