Skip to main content

AI Compute Card Full Comparison Table (100+ Models)

Last updated: 2026-07-22 · Data continuously updated. Spot an error? Submit an Issue.

Quick Filter

ScenarioRecommended Models
Trillion-parameter training (GPT-4 class)Rubin R200, B300 Ultra, MI400, TPU Ironwood
10B–100B parameter trainingH100, H200, B200, MI300X, MI325X
China market (domestic alternatives)Ascend 950DT, Ascend 910C, Ascend 920, MLU690
High-throughput inferenceL40S, L4, H200 (inference mode), Crescent Island
Edge AIJetson Orin, Edge TPU, Hailo-8L

Datacenter Training GPU

Compute unified with precision and sparsity/density labels. ✅ Available · 🔄 Upcoming · 🔮 Forward-looking

ModelFP8/FP4 ComputeFP16 ComputeMemoryMemory BandwidthTDPReleaseStatus
NVIDIA Rubin R20035 PFLOPS (FP4 training)~25 PFLOPS288GB HBM422 TB/s1,800-2,300W2026 H2🔮
NVIDIA B300 Ultra14 PFLOPS (sparse)~7 PFLOPS288GB HBM3e8 TB/s1,400W2026 Q1
NVIDIA GB300288GB HBM3e8 TB/s1,600W2025 H2
NVIDIA B2009 PFLOPS (sparse)4.5 PFLOPS (sparse)192GB HBM3e8 TB/s1,000W2025 Q2
NVIDIA GB200192GB HBM3e8 TB/s1,000W2024 Q4
NVIDIA B1007 PFLOPS (sparse)3.5 PFLOPS (sparse)192GB HBM3e8 TB/s700W2024 Q4
NVIDIA H2003,958 TFLOPS (sparse)1,979 TFLOPS141GB HBM3e4.8 TB/s700W2024 Q2
NVIDIA H100 SXM3,958 TFLOPS (sparse)1,979 TFLOPS80GB HBM33.35 TB/s700W2022 Q3
NVIDIA H20296 TFLOPS148 TFLOPS96GB HBM34.0 TB/s400W2024 Q1
NVIDIA A100312 TFLOPS80GB HBM2e2.0 TB/s300-400W2020 Q3
AMD MI400 Series40 PFLOPS (FP4)~10 PFLOPS432GB HBM419.6 TB/s1,200-1,500W2026 H2🔮
AMD MI455X20 PFLOPS~10 PFLOPS432GB HBM419.6 TB/s1,200-1,500W2026 H2🔮
AMD MI350X10.1 PFLOPS (MXFP6)~5 PFLOPS288GB HBM3e8 TB/s1,400W2025 H2🔄
AMD MI325X2,614 TFLOPS1,307 TFLOPS256GB HBM3e6.48 TB/s750W2024 Q4
AMD MI300X2,614 TFLOPS1,307 TFLOPS192GB HBM35.3 TB/s750W2023 Q4
Huawei Ascend 950PR1 PFLOPS (FP8)~780 TFLOPS128GB HiBL~3 TB/s600W2026 H1🔄
Huawei Ascend 950DT1 PFLOPS (FP8)~500 TFLOPS144GB HiZQ4 TB/s500W2026 H1🔄
Huawei Ascend 9201,800 TFLOPS~96GB HBM3~4 TB/s400W2025 H2
Huawei Ascend 910C780 TFLOPS (BF16)~390 TFLOPS128GB HBM2e (dual-die)~1.2 TB/s310W2025 H1
Huawei Ascend 910B256 TFLOPS64GB HBM2e1.2 TB/s310W2023
Cambricon MLU690700+ TFLOPS196GB HBM33.35 TB/s~500W2025
Moore Threads MTT S5000~63 TFLOPS (FP32)80GB GDDR6X1.6 TB/s300W2025 Q1
MetaX C6001,000 TFLOPS3.6 TB/s400W2025 Q4🔄
Kunlun P800345 TFLOPS96GB HBM3400W2024 Q1
Iluvatar BI-V150~192 TFLOPS64GB HBM2e350W2023

Domestic chip note: Huawei Ascend, Cambricon MLU, Moore Threads MTT, MetaX, Kunlun, and Iluvatar are Chinese domestic AI chip representatives, primarily targeting the China market due to US export controls. MTT S5000 is priced in CNY (¥55,000).

Datacenter Inference GPU

ModelFP8 ComputeINT8 ComputeMemoryTDPUse CaseStatus
NVIDIA L40S733 TFLOPS (sparse)1,466 TOPS48GB GDDR6350WDatacenter inference
NVIDIA RTX 6000 Ada1,458 TFLOPS (sparse)2,905 TOPS48GB GDDR6300WWorkstation inference
NVIDIA RTX Pro 6000 Blackwell96GB GDDR7 ECC600WWorkstation inference
NVIDIA L4485 TFLOPS970 TOPS24GB GDDR672WEdge inference
NVIDIA L296 TFLOPS (sparse)193 TOPS16GB GDDR650WLow-power inference
NVIDIA T465 TFLOPS130 TOPS16GB GDDR670WEntry-level inference
Intel Arc Pro B6024GB GDDR6200WMid-range inference
Intel Arc Pro B5016GB GDDR670WEntry-level inference
Qualcomm AI 200800 TFLOPS280WDatacenter inference🔄

AI Training ASIC (TPU / Gaudi / Trainium)

ModelVendorCompute (BF16/FP8)MemoryInterconnectReleaseStatus
Google TPU Ironwood (v7)Google~2,000 TFLOPS192GB HBM~5 Tb/s2026 H1🔄
Google TPU v6pGoogle96GB HBM22024 Q4
Google TPU v6e (Trillium)Google918 TFLOPS32GB HBM1.6 Tb/s2024 Q4
Google TPU v5pGoogle2023 Q4
Google TPU v5eGoogle16GB HBM22023 Q3
Google TPU v4Google32GB HBM22020 Q3
Google TPU 8t (Training)Google2026 Q2🔮
Google TPU 8i (Inference)Google~1,500 TOPS2026 Q2🔮
Intel Gaudi 3Intel1,600 TFLOPS128GB SRAM2.4 Tb/s2024 Q2
Intel Gaudi 2Intel865 TFLOPS (FP8)96GB HBM2e2.4 Tb/s2022 Q2
Intel Gaudi 4Intel192GB HBM3e2026 Q2🔮
Intel Crescent IslandIntelTBD480GB LPDDR5xTBD2026 H2🔄
AWS Trainium 3AWS~5.7 PFLOPS~144GB~4.5 Tb/s2025 Q4🔄
AWS Trainium 2AWS1,299 TFLOPS (dense)64GB~1.6 Tb/s2024 Q4
AWS Trainium 1AWS191 TFLOPS (FP8)32GB HBM2020 Q4
AWS Inferentia 2AWS190 TFLOPS (FP16)32GB HBM2e2022 Q4
AWS Inferentia 1AWS2019 Q4
Microsoft Maia 200Microsoft5+ PFLOPS2026 Q1🔄
Meta MTIA v3Meta2026 Q3🔮

Wafer-Scale Training

ModelVendorTransistorsOn-Chip MemoryFP8 ComputeReleaseStatus
Cerebras WSE-4Cerebras~5-6 trillion44GB SRAM~400 PFLOPS2027🔮
Cerebras WSE-3Cerebras4 trillion40GB SRAM125 PFLOPS2024 Q1
Cerebras WSE-2Cerebras2.6 trillion40GB SRAM85 PFLOPS2021 Q3

Edge AI & On-Device NPU

ModelVendorCompute (TOPS)PowerUse CaseStatus
NVIDIA Jetson ThorNVIDIA2,070 TOPS130WRobotics / autonomous driving
NVIDIA Jetson Orin AGXNVIDIA275 TOPS60WEdge inference
Qualcomm AI 100Qualcomm70 TOPS15WDatacenter edge inference
Huawei Ascend 310Huawei22 TOPS8WOn-device inference
Hailo-8LHailo13 TOPS1.5WOn-device vision AI
Google Edge TPUGoogle4 TOPS2WIoT on-device inference

Innovative Architectures

ModelArchitecture TypeKey FeatureVendorStatus
Groq LPU v2LPU (Language Processing Unit)Ultra-low latency inference (~500 tok/s)Groq
Graphcore IPU (Bow)IPU (Intelligence Processing Unit)Native graph computing, 1,400 IPU coresGraphcore
Tesla Dojo (D1)Distributed training waferIntegrated auto-labeling + model trainingTesla
Apple M5 UltraSoC + NPUOn-device 50 TOPS, unified memoryApple🔮
BrainChip Akida 2Spiking Neural Network (SNN)Ultra-low-power neuromorphicBrainChip

Pricing Reference

Prices fluctuate with market supply and demand. Purchase prices are affected by export controls. CNY denotes Chinese domestic chip pricing in RMB (reference rate: 1 USD ≈ 7.2 CNY). Data for reference only.

NVIDIA

ModelMSRP (USD)Market Price (USD)Notes
Rubin R200$85,000Estimated
GB300$75,000$72,000
GB200$65,000$62,000
B300 Ultra$55,000$52,000
B200$45,000$42,000
B100$38,000$35,000
H200$40,000$38,000
H100$30,000$28,000
H100 NVL$40,000$38,000
H20$14,000$13,000China-specific
A100$15,000$28,000Secondary market
L40S$7,000$6,500
RTX Pro 6000 Blackwell$6,800$6,500
RTX 6000 Ada$4,500$4,200
RTX 5090$2,000$1,900
L4 / L2$2,500$2,300
T4$2,500$1,800Secondary market
Jetson Thor$800Module
Jetson Orin$400$380Module
H800¥280,000¥250,000CNY, China market

AMD

ModelMSRP (USD)Market Price (USD)Notes
MI400$55,000Estimated
MI350X$40,000$37,000
MI355X$22,000$20,500
MI325X$18,000$16,500
MI300X$15,000$13,500
MI250$12,000$10,000Secondary market
MI210$9,000$8,500

Intel

ModelMSRP (USD)Market Price (USD)Notes
Gaudi 4$25,000Estimated
Gaudi 3$18,000$16,500
Gaudi 2$12,000$11,000
Gaudi 1$8,000$7,000Discontinued
Max Series$4,000$3,700
Flex Series$1,000$900
Arc Pro B60$500$480
Arc Pro B50$350$330

Huawei Ascend

ModelMSRP (USD)Market Price (USD)Notes
Ascend 950DT$22,000Estimated
Ascend 950PR$18,000Estimated
Ascend 910D$18,000$16,000
Ascend 920$25,000$23,000
Ascend 910C$16,000$14,500
Ascend 910B$12,000$10,500

Google TPU / AWS / Cloud

ModelMSRP (USD)Market Price (USD)Notes
TPU Ironwood$45,000Estimated
TPU v6p$40,000
TPU v5p$35,000
TPU v6e$22,000
TPU v5e$18,000
TPU v4$25,000
Trainium 3$30,000
Trainium 2$22,000
Inferentia 2$12,000
Trainium 1$15,000
Inferentia 1$8,000

Chinese Domestic Chips (CNY Pricing)

ModelMSRP (CNY)Market Price (CNY)USD Equivalent
Cambricon MLU690¥150,000¥140,000~$19,444
Moore Threads MTT S5000¥55,000¥50,000~$6,944
Iluvatar TG150¥85,000¥78,000~$10,833
Enflame T21¥70,000¥65,000~$9,028
MetaX C500¥55,000¥52,000~$7,222
Moore Threads S4000¥60,000¥55,000~$7,639

Other Vendors

ModelMSRP (USD)Market Price (USD)Notes
Cerebras WSE-4$8,000,000Rack system
Cerebras WSE-3$5,000,000Rack system
Cerebras WSE-2$3,000,000Rack system
SambaNova SN40L$200,000System
Groq LPU v2$35,000
Rubin Ultra$150,000Estimated
Graphcore IPU$15,000
Tenstorrent Blackhole$20,000
Qualcomm AI 100$8,000$7,000
Hailo-8L$300$280Module

Purchasing Guide

By Model Scale

By Region

  • North America / Europe: NVIDIA + AMD, freely available
  • China: Ascend 950 / 910C / 920 / Cambricon MLU690 (domestic alternatives)
  • Cloud (no hardware preference): Any vendor, choose by price

← Back to Home | Roadmap → | TCO Calculator → | Industry News →