AI Compute Card Full Comparison Table (100+ Models)
Last updated: 2026-07-22 · Data continuously updated. Spot an error? Submit an Issue.
Quick Filter
| Scenario | Recommended Models |
|---|---|
| Trillion-parameter training (GPT-4 class) | Rubin R200, B300 Ultra, MI400, TPU Ironwood |
| 10B–100B parameter training | H100, H200, B200, MI300X, MI325X |
| China market (domestic alternatives) | Ascend 950DT, Ascend 910C, Ascend 920, MLU690 |
| High-throughput inference | L40S, L4, H200 (inference mode), Crescent Island |
| Edge AI | Jetson Orin, Edge TPU, Hailo-8L |
Datacenter Training GPU
Compute unified with precision and sparsity/density labels. ✅ Available · 🔄 Upcoming · 🔮 Forward-looking
| Model | FP8/FP4 Compute | FP16 Compute | Memory | Memory Bandwidth | TDP | Release | Status |
|---|---|---|---|---|---|---|---|
| NVIDIA Rubin R200 | 35 PFLOPS (FP4 training) | ~25 PFLOPS | 288GB HBM4 | 22 TB/s | 1,800-2,300W | 2026 H2 | 🔮 |
| NVIDIA B300 Ultra | 14 PFLOPS (sparse) | ~7 PFLOPS | 288GB HBM3e | 8 TB/s | 1,400W | 2026 Q1 | ✅ |
| NVIDIA GB300 | — | — | 288GB HBM3e | 8 TB/s | 1,600W | 2025 H2 | ✅ |
| NVIDIA B200 | 9 PFLOPS (sparse) | 4.5 PFLOPS (sparse) | 192GB HBM3e | 8 TB/s | 1,000W | 2025 Q2 | ✅ |
| NVIDIA GB200 | — | — | 192GB HBM3e | 8 TB/s | 1,000W | 2024 Q4 | ✅ |
| NVIDIA B100 | 7 PFLOPS (sparse) | 3.5 PFLOPS (sparse) | 192GB HBM3e | 8 TB/s | 700W | 2024 Q4 | ✅ |
| NVIDIA H200 | 3,958 TFLOPS (sparse) | 1,979 TFLOPS | 141GB HBM3e | 4.8 TB/s | 700W | 2024 Q2 | ✅ |
| NVIDIA H100 SXM | 3,958 TFLOPS (sparse) | 1,979 TFLOPS | 80GB HBM3 | 3.35 TB/s | 700W | 2022 Q3 | ✅ |
| NVIDIA H20 | 296 TFLOPS | 148 TFLOPS | 96GB HBM3 | 4.0 TB/s | 400W | 2024 Q1 | ✅ |
| NVIDIA A100 | — | 312 TFLOPS | 80GB HBM2e | 2.0 TB/s | 300-400W | 2020 Q3 | ✅ |
| AMD MI400 Series | 40 PFLOPS (FP4) | ~10 PFLOPS | 432GB HBM4 | 19.6 TB/s | 1,200-1,500W | 2026 H2 | 🔮 |
| AMD MI455X | 20 PFLOPS | ~10 PFLOPS | 432GB HBM4 | 19.6 TB/s | 1,200-1,500W | 2026 H2 | 🔮 |
| AMD MI350X | 10.1 PFLOPS (MXFP6) | ~5 PFLOPS | 288GB HBM3e | 8 TB/s | 1,400W | 2025 H2 | 🔄 |
| AMD MI325X | 2,614 TFLOPS | 1,307 TFLOPS | 256GB HBM3e | 6.48 TB/s | 750W | 2024 Q4 | ✅ |
| AMD MI300X | 2,614 TFLOPS | 1,307 TFLOPS | 192GB HBM3 | 5.3 TB/s | 750W | 2023 Q4 | ✅ |
| Huawei Ascend 950PR | 1 PFLOPS (FP8) | ~780 TFLOPS | 128GB HiBL | ~3 TB/s | 600W | 2026 H1 | 🔄 |
| Huawei Ascend 950DT | 1 PFLOPS (FP8) | ~500 TFLOPS | 144GB HiZQ | 4 TB/s | 500W | 2026 H1 | 🔄 |
| Huawei Ascend 920 | — | 1,800 TFLOPS | ~96GB HBM3 | ~4 TB/s | 400W | 2025 H2 | ✅ |
| Huawei Ascend 910C | 780 TFLOPS (BF16) | ~390 TFLOPS | 128GB HBM2e (dual-die) | ~1.2 TB/s | 310W | 2025 H1 | ✅ |
| Huawei Ascend 910B | — | 256 TFLOPS | 64GB HBM2e | 1.2 TB/s | 310W | 2023 | ✅ |
| Cambricon MLU690 | — | 700+ TFLOPS | 196GB HBM3 | 3.35 TB/s | ~500W | 2025 | ✅ |
| Moore Threads MTT S5000 | — | ~63 TFLOPS (FP32) | 80GB GDDR6X | 1.6 TB/s | 300W | 2025 Q1 | ✅ |
| MetaX C600 | 1,000 TFLOPS | — | — | 3.6 TB/s | 400W | 2025 Q4 | 🔄 |
| Kunlun P800 | — | 345 TFLOPS | 96GB HBM3 | — | 400W | 2024 Q1 | ✅ |
| Iluvatar BI-V150 | — | ~192 TFLOPS | 64GB HBM2e | — | 350W | 2023 | ✅ |
Domestic chip note: Huawei Ascend, Cambricon MLU, Moore Threads MTT, MetaX, Kunlun, and Iluvatar are Chinese domestic AI chip representatives, primarily targeting the China market due to US export controls. MTT S5000 is priced in CNY (¥55,000).
Datacenter Inference GPU
| Model | FP8 Compute | INT8 Compute | Memory | TDP | Use Case | Status |
|---|---|---|---|---|---|---|
| NVIDIA L40S | 733 TFLOPS (sparse) | 1,466 TOPS | 48GB GDDR6 | 350W | Datacenter inference | ✅ |
| NVIDIA RTX 6000 Ada | 1,458 TFLOPS (sparse) | 2,905 TOPS | 48GB GDDR6 | 300W | Workstation inference | ✅ |
| NVIDIA RTX Pro 6000 Blackwell | — | — | 96GB GDDR7 ECC | 600W | Workstation inference | ✅ |
| NVIDIA L4 | 485 TFLOPS | 970 TOPS | 24GB GDDR6 | 72W | Edge inference | ✅ |
| NVIDIA L2 | 96 TFLOPS (sparse) | 193 TOPS | 16GB GDDR6 | 50W | Low-power inference | ✅ |
| NVIDIA T4 | 65 TFLOPS | 130 TOPS | 16GB GDDR6 | 70W | Entry-level inference | ✅ |
| Intel Arc Pro B60 | — | — | 24GB GDDR6 | 200W | Mid-range inference | ✅ |
| Intel Arc Pro B50 | — | — | 16GB GDDR6 | 70W | Entry-level inference | ✅ |
| Qualcomm AI 200 | 800 TFLOPS | — | — | 280W | Datacenter inference | 🔄 |
AI Training ASIC (TPU / Gaudi / Trainium)
| Model | Vendor | Compute (BF16/FP8) | Memory | Interconnect | Release | Status |
|---|---|---|---|---|---|---|
| Google TPU Ironwood (v7) | ~2,000 TFLOPS | 192GB HBM | ~5 Tb/s | 2026 H1 | 🔄 | |
| Google TPU v6p | — | 96GB HBM2 | — | 2024 Q4 | ✅ | |
| Google TPU v6e (Trillium) | 918 TFLOPS | 32GB HBM | 1.6 Tb/s | 2024 Q4 | ✅ | |
| Google TPU v5p | — | — | — | 2023 Q4 | ✅ | |
| Google TPU v5e | — | 16GB HBM2 | — | 2023 Q3 | ✅ | |
| Google TPU v4 | — | 32GB HBM2 | — | 2020 Q3 | ✅ | |
| Google TPU 8t (Training) | — | — | — | 2026 Q2 | 🔮 | |
| Google TPU 8i (Inference) | ~1,500 TOPS | — | — | 2026 Q2 | 🔮 | |
| Intel Gaudi 3 | Intel | 1,600 TFLOPS | 128GB SRAM | 2.4 Tb/s | 2024 Q2 | ✅ |
| Intel Gaudi 2 | Intel | 865 TFLOPS (FP8) | 96GB HBM2e | 2.4 Tb/s | 2022 Q2 | ✅ |
| Intel Gaudi 4 | Intel | — | 192GB HBM3e | — | 2026 Q2 | 🔮 |
| Intel Crescent Island | Intel | TBD | 480GB LPDDR5x | TBD | 2026 H2 | 🔄 |
| AWS Trainium 3 | AWS | ~5.7 PFLOPS | ~144GB | ~4.5 Tb/s | 2025 Q4 | 🔄 |
| AWS Trainium 2 | AWS | 1,299 TFLOPS (dense) | 64GB | ~1.6 Tb/s | 2024 Q4 | ✅ |
| AWS Trainium 1 | AWS | 191 TFLOPS (FP8) | 32GB HBM | — | 2020 Q4 | ✅ |
| AWS Inferentia 2 | AWS | 190 TFLOPS (FP16) | 32GB HBM2e | — | 2022 Q4 | ✅ |
| AWS Inferentia 1 | AWS | — | — | — | 2019 Q4 | ✅ |
| Microsoft Maia 200 | Microsoft | 5+ PFLOPS | — | — | 2026 Q1 | 🔄 |
| Meta MTIA v3 | Meta | — | — | — | 2026 Q3 | 🔮 |
Wafer-Scale Training
| Model | Vendor | Transistors | On-Chip Memory | FP8 Compute | Release | Status |
|---|---|---|---|---|---|---|
| Cerebras WSE-4 | Cerebras | ~5-6 trillion | 44GB SRAM | ~400 PFLOPS | 2027 | 🔮 |
| Cerebras WSE-3 | Cerebras | 4 trillion | 40GB SRAM | 125 PFLOPS | 2024 Q1 | ✅ |
| Cerebras WSE-2 | Cerebras | 2.6 trillion | 40GB SRAM | 85 PFLOPS | 2021 Q3 | ✅ |
Edge AI & On-Device NPU
| Model | Vendor | Compute (TOPS) | Power | Use Case | Status |
|---|---|---|---|---|---|
| NVIDIA Jetson Thor | NVIDIA | 2,070 TOPS | 130W | Robotics / autonomous driving | ✅ |
| NVIDIA Jetson Orin AGX | NVIDIA | 275 TOPS | 60W | Edge inference | ✅ |
| Qualcomm AI 100 | Qualcomm | 70 TOPS | 15W | Datacenter edge inference | ✅ |
| Huawei Ascend 310 | Huawei | 22 TOPS | 8W | On-device inference | ✅ |
| Hailo-8L | Hailo | 13 TOPS | 1.5W | On-device vision AI | ✅ |
| Google Edge TPU | 4 TOPS | 2W | IoT on-device inference | ✅ |
Innovative Architectures
| Model | Architecture Type | Key Feature | Vendor | Status |
|---|---|---|---|---|
| Groq LPU v2 | LPU (Language Processing Unit) | Ultra-low latency inference (~500 tok/s) | Groq | ✅ |
| Graphcore IPU (Bow) | IPU (Intelligence Processing Unit) | Native graph computing, 1,400 IPU cores | Graphcore | ✅ |
| Tesla Dojo (D1) | Distributed training wafer | Integrated auto-labeling + model training | Tesla | ✅ |
| Apple M5 Ultra | SoC + NPU | On-device 50 TOPS, unified memory | Apple | 🔮 |
| BrainChip Akida 2 | Spiking Neural Network (SNN) | Ultra-low-power neuromorphic | BrainChip | ✅ |
Pricing Reference
Prices fluctuate with market supply and demand. Purchase prices are affected by export controls. CNY denotes Chinese domestic chip pricing in RMB (reference rate: 1 USD ≈ 7.2 CNY). Data for reference only.
NVIDIA
| Model | MSRP (USD) | Market Price (USD) | Notes |
|---|---|---|---|
| Rubin R200 | $85,000 | — | Estimated |
| GB300 | $75,000 | $72,000 | |
| GB200 | $65,000 | $62,000 | |
| B300 Ultra | $55,000 | $52,000 | |
| B200 | $45,000 | $42,000 | |
| B100 | $38,000 | $35,000 | |
| H200 | $40,000 | $38,000 | |
| H100 | $30,000 | $28,000 | |
| H100 NVL | $40,000 | $38,000 | |
| H20 | $14,000 | $13,000 | China-specific |
| A100 | $15,000 | $28,000 | Secondary market |
| L40S | $7,000 | $6,500 | |
| RTX Pro 6000 Blackwell | $6,800 | $6,500 | |
| RTX 6000 Ada | $4,500 | $4,200 | |
| RTX 5090 | $2,000 | $1,900 | |
| L4 / L2 | $2,500 | $2,300 | |
| T4 | $2,500 | $1,800 | Secondary market |
| Jetson Thor | $800 | — | Module |
| Jetson Orin | $400 | $380 | Module |
| H800 | ¥280,000 | ¥250,000 | CNY, China market |
AMD
| Model | MSRP (USD) | Market Price (USD) | Notes |
|---|---|---|---|
| MI400 | $55,000 | — | Estimated |
| MI350X | $40,000 | $37,000 | |
| MI355X | $22,000 | $20,500 | |
| MI325X | $18,000 | $16,500 | |
| MI300X | $15,000 | $13,500 | |
| MI250 | $12,000 | $10,000 | Secondary market |
| MI210 | $9,000 | $8,500 |
Intel
| Model | MSRP (USD) | Market Price (USD) | Notes |
|---|---|---|---|
| Gaudi 4 | $25,000 | — | Estimated |
| Gaudi 3 | $18,000 | $16,500 | |
| Gaudi 2 | $12,000 | $11,000 | |
| Gaudi 1 | $8,000 | $7,000 | Discontinued |
| Max Series | $4,000 | $3,700 | |
| Flex Series | $1,000 | $900 | |
| Arc Pro B60 | $500 | $480 | |
| Arc Pro B50 | $350 | $330 |
Huawei Ascend
| Model | MSRP (USD) | Market Price (USD) | Notes |
|---|---|---|---|
| Ascend 950DT | $22,000 | — | Estimated |
| Ascend 950PR | $18,000 | — | Estimated |
| Ascend 910D | $18,000 | $16,000 | |
| Ascend 920 | $25,000 | $23,000 | |
| Ascend 910C | $16,000 | $14,500 | |
| Ascend 910B | $12,000 | $10,500 |
Google TPU / AWS / Cloud
| Model | MSRP (USD) | Market Price (USD) | Notes |
|---|---|---|---|
| TPU Ironwood | $45,000 | — | Estimated |
| TPU v6p | $40,000 | — | |
| TPU v5p | $35,000 | — | |
| TPU v6e | $22,000 | — | |
| TPU v5e | $18,000 | — | |
| TPU v4 | $25,000 | — | |
| Trainium 3 | $30,000 | — | |
| Trainium 2 | $22,000 | — | |
| Inferentia 2 | $12,000 | — | |
| Trainium 1 | $15,000 | — | |
| Inferentia 1 | $8,000 | — |
Chinese Domestic Chips (CNY Pricing)
| Model | MSRP (CNY) | Market Price (CNY) | USD Equivalent |
|---|---|---|---|
| Cambricon MLU690 | ¥150,000 | ¥140,000 | ~$19,444 |
| Moore Threads MTT S5000 | ¥55,000 | ¥50,000 | ~$6,944 |
| Iluvatar TG150 | ¥85,000 | ¥78,000 | ~$10,833 |
| Enflame T21 | ¥70,000 | ¥65,000 | ~$9,028 |
| MetaX C500 | ¥55,000 | ¥52,000 | ~$7,222 |
| Moore Threads S4000 | ¥60,000 | ¥55,000 | ~$7,639 |
Other Vendors
| Model | MSRP (USD) | Market Price (USD) | Notes |
|---|---|---|---|
| Cerebras WSE-4 | $8,000,000 | — | Rack system |
| Cerebras WSE-3 | $5,000,000 | — | Rack system |
| Cerebras WSE-2 | $3,000,000 | — | Rack system |
| SambaNova SN40L | $200,000 | — | System |
| Groq LPU v2 | $35,000 | — | |
| Rubin Ultra | $150,000 | — | Estimated |
| Graphcore IPU | $15,000 | — | |
| Tenstorrent Blackhole | $20,000 | — | |
| Qualcomm AI 100 | $8,000 | $7,000 | |
| Hailo-8L | $300 | $280 | Module |
Purchasing Guide
By Model Scale
- Trillion-parameter (GPT-4 class): B300 Ultra / Rubin R200, AMD MI400 (2026 H2)
- 10B–100B parameters (Llama 70B, Qwen 72B): H100 / H200, AMD MI300X / MI325X
- 1B–10B parameters (Llama 7B–13B): H100, A100 80GB
- Small models / inference: L40S, L4, T4
By Region
- North America / Europe: NVIDIA + AMD, freely available
- China: Ascend 950 / 910C / 920 / Cambricon MLU690 (domestic alternatives)
- Cloud (no hardware preference): Any vendor, choose by price
← Back to Home | Roadmap → | TCO Calculator → | Industry News →