Skip to main content

2 posts tagged with "MLU690"

寒武纪 MLU690 芯片详解

View all tags

China's Domestic AI Chip Triopoly (2026): Ascend, Cambricon, Moore Threads — Who Is the "China H100"?

· 7 min read
AI Hardware Analyst

Against the backdrop of U.S. export controls, China's AI chip market is forming a "three-way standoff." This article compares the technical routes, product specs, software ecosystems, and commercial progress of the three major domestic AI chip vendors: Huawei Ascend, Cambricon MLU, and Moore Threads MTT.


Key Points

  • Huawei Ascend: leader in domestic AI training chips; Ascend 950 in mass production; most mature software ecosystem
  • Cambricon MLU690: the "China H100," compute close to H200, clear efficiency advantage
  • Moore Threads MTT S5000: full-function GPU route; achieved Day-0 support for Qwen3.5 and GLM-5.2 in June 2026
  • Shared challenge: affected by U.S. export controls, primarily aimed at the Chinese market, limited internationally

I. Vendor Overview

VendorFoundedFounderListed2025 RevenueMain Customers
Huawei Ascend2018 (division)Ren Zhengfeiprivate (wholly owned by Huawei)~¥20B (est.)Chinese gov, SOEs, military
Cambricon2016Chen Tianshi (CAS)2020-07 (STAR Market 688256)~¥5.2BByteDance, Alibaba, Baidu
Moore Threads2020Zhang Jianzhong (ex-NVIDIA China)2023-12 (STAR Market 688495)~¥1.5B (est.)gov, SOEs, gaming cos.

Strategic Positioning

VendorTech routeCore strengthMain challenge
Huawei AscendAI-training-specific (Da Vinci)co-optimized HW/SW, carrier channelssanctions, process limits
CambriconAI-training-specific (MLUarch)high efficiency, competitive priceimmature ecosystem
Moore ThreadsFull-function GPU (MUSA)graphics + AI + general compute, Day-0 supportcompute below dedicated AI chips

II. Flagship Product Comparison

1. Huawei Ascend 950DT (2026 flagship)

ItemSpec
BF16 compute1,000 TFLOPS
Memory144GB HiZQ 2.0 (in-house HBM)
Memory bandwidth4 TB/s
TDP400W
ProcessN+2 (improved 7nm)
Released2026-04
Mass production2026-Q2
Unit price~¥80,000 (est.)

Strengths:

  • High large-model inference throughput: 144GB memory friendly to DeepSeek R1 (671B MoE)
  • Most mature ecosystem: CANN ~85% operator coverage, supports PyTorch, TensorFlow
  • Strong carrier channel: China Mobile, China Telecom large purchases

Weaknesses:

  • Process limited: N+2 below TSMC 4nm
  • Mediocre efficiency: 400W TDP, 2.5 TFLOPS/W

2. Cambricon MLU690 (2026 flagship)

ItemSpec
BF16 compute600 TFLOPS
Memory64GB HBM3
Memory bandwidth2 TB/s
TDP280W
ProcessTSMC 7nm
Released2025-Q4
Mass production2026-Q1
Unit price~¥140,000 (est.)

Strengths:

  • Best efficiency: 280W TDP, 2.14 TFLOPS/W (1.5x H100)
  • Competitive price: ~$20,000, 33% cheaper than H100
  • Top-tier customer orders: ByteDance, Alibaba, Baidu

Weaknesses:

  • Small memory: 64GB limits large-model training scale
  • Immature ecosystem: NeuWare ~75–85% coverage; complex LLMs need manual tuning

3. Moore Threads MTT S5000 (2025 flagship)

ItemSpec
FP16 compute~1,000 TFLOPS (est.)
Memory80GB GDDR6X
Memory bandwidth1.6 TB/s
TDP~350W
ProcessTSMC 4nm (est.)
Released2025-02
Mass production2025-Q2
Unit price~¥50,000 (est.)

Strengths:

  • Full-function GPU: graphics + AI + general compute, broader scenarios
  • Strong Day-0 support: June 2026 Day-0 support for Qwen3.5, GLM-5.2, MiniMax M3
  • Lowest price: ~¥50,000, high cost-performance

Weaknesses:

  • Compute below dedicated AI chips: FP16 ~50% of H100
  • Low memory bandwidth: 1.6 TB/s (48% of H100), limits large-model training

III. Compute Comparison (BF16/FP16)

ChipBF16 computeMemoryBandwidthTDPEfficiency
Huawei Ascend 950DT1,000 TFLOPS144GB4 TB/s400W2.5 TFLOPS/W
Cambricon MLU690600 TFLOPS64GB2 TB/s280W2.14 TFLOPS/W
Moore Threads MTT S5000~1,000 TFLOPS80GB1.6 TB/s~350W~2.86 TFLOPS/W
NVIDIA H100989 TFLOPS80GB3.35 TB/s700W1.41 TFLOPS/W
NVIDIA H200989 TFLOPS141GB4.8 TB/s700W1.41 TFLOPS/W

Key insights:

  1. Ascend 950DT has the highest compute (1,000 TFLOPS) but mediocre efficiency
  2. Cambricon MLU690 has the best efficiency (2.14 TFLOPS/W), TDP only 280W
  3. Moore Threads MTT S5000 wins on full-function versatility but low bandwidth

IV. Software Ecosystem

VendorStackFramework supportCoverageMaturity
Huawei AscendCANNPyTorch, TensorFlow, MindSpore~85%⭐⭐⭐⭐ (4/5)
CambriconNeuWarePyTorch-Cambricon, TensorFlow-Cambricon~75–85%⭐⭐⭐ (3/5)
Moore ThreadsMUSIFYPyTorch, TensorFlow, ONNX~70%⭐⭐⭐ (3/5)
NVIDIACUDAall~99%⭐⭐⭐⭐⭐ (5/5)

Ecosystem Maturity Assessment

Huawei Ascend CANN:

  • ✅ Strength: highest operator coverage, supports MindSpore (in-house framework)
  • ❌ Weakness: steep learning curve, incomplete docs

Cambricon NeuWare:

  • ✅ Strength: PyTorch/TensorFlow compatible, low migration cost
  • ❌ Weakness: complex LLMs need manual tuning

Moore Threads MUSIFY:

  • ✅ Strength: strong Day-0 support, ONNX support
  • ❌ Weakness: lowest operator coverage, dual graphics+AI engine complexity

V. Commercial Progress

Vendor2026 commercial progressMain customersShipments
Huawei AscendAscend 950 mass production; China Mobile large purchaseChina Mobile, China Telecom, gov~100K/yr (est.)
CambriconMLU690 mass production; ByteDance, Alibaba ordersByteDance, Alibaba, Baidu~50K/yr (est.)
Moore ThreadsMTT S5000 mass production; Day-0 Qwen3.5gov, SOEs, gaming cos.~30K/yr (est.)

Latest as of June 2026

Huawei Ascend:

  • ✅ Ascend 950DT fully ramping
  • ✅ ¥1B procurement agreement with China Mobile

Cambricon:

  • ✅ MLU690 in volume shipment
  • ✅ ByteDance order ~20K units

Moore Threads:

  • ✅ Day-0 support for Qwen3.5, GLM-5.2, MiniMax M3
  • ✅ MTT S5000 2nd-gen released

VI. Selection Advice

Scenario 1: Trillion-parameter training (GPT-4 class)

Recommended: Huawei Ascend 950DT

  • ✅ 144GB large memory supports super-large models
  • ✅ Most mature ecosystem (~85% coverage)
  • ✅ Strong carrier channel, Chinese government backing

Alternative: Cambricon MLU690 (high efficiency, but small memory)

Scenario 2: Tens-to-hundreds-of-billions parameter training

Recommended: Cambricon MLU690

  • ✅ Best efficiency (2.14 TFLOPS/W), low TCO
  • ✅ Competitive price (~$20,000)
  • ✅ Validated by top customers (ByteDance, Alibaba)

Alternative: Huawei Ascend 920 (more compute, mediocre efficiency)

Scenario 3: Cloud AI inference

Recommended: Huawei Ascend 950PR (inference-specific)

  • ✅ Well-optimized inference throughput
  • ✅ 128GB memory friendly to MoE models
  • ✅ Mature stack, low deployment cost

Alternative: Moore Threads MTT S5000 (full-function GPU, inference + graphics)

Scenario 4: Edge AI / on-device inference

Recommended: Moore Threads MTT S5000

  • ✅ Full-function GPU, graphics + AI
  • ✅ Lowest price (~¥50,000)
  • ✅ Strong Day-0 support

Alternative: Huawei Ascend 310 (low power, 8W TDP)

Scenario 5: Domestic substitution (gov, SOEs)

Recommended: Huawei Ascend 950DT

  • ✅ Chinese government first choice, carrier bulk buys
  • ✅ Co-optimized HW/SW, stable performance
  • ✅ Supported by national semiconductor fund

Alternative: Cambricon MLU690 (high efficiency, competitive price)


VII. Future Roadmap

Vendor2026 H220272028
Huawei Ascend950DT ramp960 (FP8 ~2 PFLOPS)970 (N+3 process)
CambriconMLU690 rampMLU790 (5nm, BF16 ~1,000 TFLOPS)MLU890 (3nm)
Moore ThreadsMTT S5000 2nd-genMTT S6000 (HBM3, FP16 ~1,500 TFLOPS)MTT S7000

VIII. Summary: Who Is the "China H100"?

DimensionAscend 950DTMLU690MTT S5000
Compute⭐⭐⭐⭐⭐ (5/5)⭐⭐⭐ (3/5)⭐⭐⭐ (3/5)
Memory⭐⭐⭐⭐⭐ (5/5)⭐⭐ (2/5)⭐⭐⭐ (3/5)
Efficiency⭐⭐⭐ (3/5)⭐⭐⭐⭐⭐ (5/5)⭐⭐⭐⭐ (4/5)
Ecosystem⭐⭐⭐⭐ (4/5)⭐⭐⭐ (3/5)⭐⭐⭐ (3/5)
Price⭐⭐⭐ (3/5)⭐⭐⭐⭐ (4/5)⭐⭐⭐⭐⭐ (5/5)
Overall⭐⭐⭐⭐ (4/5)⭐⭐⭐ (3/5)⭐⭐⭐ (3/5)

Final conclusion:

  • Huawei Ascend 950DT is the domestic AI training chip closest to H100, strongest overall
  • Cambricon MLU690 is the most efficient domestic AI chip, lowest TCO
  • Moore Threads MTT S5000 is the cheapest full-function GPU, suited to edge AI and graphics+AI

References


Disclaimer: Data based on public sources; actual specs per vendor official. MirrorFrog continuously updates domestic AI chip data — corrections welcome.

Changelog: 2026-06-23 initial release

Cambricon MLU690 vs NVIDIA H100: In-Depth Comparison — Can a Domestic AI Chip Replace the H100?

· 6 min read
AI Hardware Analyst

In 2026, against the backdrop of U.S. export controls on AI chips to China, Cambricon's MLU690 has drawn intense attention as a "China-made H100." This article compares the two in depth across compute, memory, power, software ecosystem, measured performance, and price to help you make a selection decision.

Core Verdict (Read This First)

DimensionMLU690H100WinnerGap
BF16 compute600 TFLOPS989 TFLOPSH100+65%
Memory capacity64GB HBM380GB HBM3H100+25%
Memory bandwidth2 TB/s3.35 TB/sH100+68%
TDP280W700WMLU690-60%
Energy efficiency2.14 TFLOPS/W1.41 TFLOPS/WMLU690+52%
Software ecosystemNeuWare (~75% coverage)CUDA (100% coverage)H100large gap
Price~¥140,000~¥200,000MLU690-30%
Availabilitydomestic spot stockexport-controlledMLU690

One-line summary: MLU690 delivers roughly 60% of H100's compute, but at only 40% of the power and 70% of the price — a strong fit for AI training and inference in the Chinese market.


1. Detailed Spec Comparison

1.1 Compute

PrecisionMLU690H100 SXM5H200 SXM5Note
FP8~300 TFLOPS (est.)3,958 TFLOPS3,958 TFLOPSH100 supports FP8; MLU690 likely does not
BF16/FP16600 TFLOPS989 TFLOPS989 TFLOPSH100 leads by 65%
FP32~150 TFLOPS (est.)60 TFLOPS60 TFLOPSMLU690 estimate; H100 actually higher
INT81,200 TOPS1,979 TOPS1,979 TOPSH100 leads by 65%

Key findings:

  • ✅ MLU690 reaches 60% of H100's BF16 compute
  • ⚠️ H100 supports FP8 (4-bit); MLU690 likely does not (needs confirmation)
  • ⚠️ H100's higher INT8 compute favors inference scenarios

1.2 Memory

ItemMLU690H100H200Note
Capacity64GB HBM380GB HBM3141GB HBM3eH200 largest
Bandwidth2 TB/s3.35 TB/s4.8 TB/sH200 highest
TypeHBM3HBM3HBM3eH200 uses latest HBM3e

Key findings:

  • ⚠️ MLU690 has 20% less memory than H100 (64GB vs 80GB)
  • ⚠️ MLU690 bandwidth is 40% lower than H100 (2 TB/s vs 3.35 TB/s)
  • ❌ When running 70B+ parameter models, MLU690 may run out of memory (model parallelism required)

1.3 Power

ItemMLU690H100H200
TDP280W700W700W
Efficiency (FP16/W)2.14 TFLOPS/W1.41 TFLOPS/W1.41 TFLOPS/W
8-card server power~3.5kW~6kW~6kW
Annual electricity (¥0.6/kWh)~¥18,400~¥36,800~¥36,800

Key findings:

  • MLU690 draws only 40% of H100's power, sharply cutting data-center electricity cost
  • MLU690 leads efficiency by 52%, better suited to large-scale deployment
  • ✅ For power-sensitive inference, MLU690 has a clear edge

2. Software Ecosystem

2.1 Framework Support

FrameworkMLU690 (NeuWare)H100 (CUDA)Note
PyTorch✅ (PyTorch-Cambricon)✅ nativeMLU690 needs an extra plugin
TensorFlow✅ (TensorFlow-Cambricon)✅ nativesame
JAX⚠️ partial✅ nativeMLU690 limited
ONNX⚠️ partial✅ nativesame
vLLM⚠️ in progress✅ nativeMLU690 awaits community port

2.2 Operator Coverage

CategoryMLU690H100Note
Basic operators✅ 95%✅ 100%conv, matmul, etc.
Transformer operators✅ 85%✅ 100%Attention, LayerNorm, etc.
Custom operators⚠️ hand-written✅ CUDA C++MLU690 harder to develop
LLM inference opt.⚠️ basic✅ mature (FlashAttention, PagedAttention)H100 leads

Key findings:

  • ⚠️ NeuWare is only 5–6 years old, with ~75–85% operator coverage
  • ❌ Complex LLMs (e.g., GPT-4, Claude) may need manual optimization
  • ✅ Common models (Llama, Qwen, GLM) are essentially already supported

3. Measured Performance

3.1 Training

ModelMLU690 (time)H100 (time)Speedup
Llama 7B~48 h (est.)~30 h1.6x
Llama 70B~7 days (est.)~4.5 days1.6x
Qwen 72B~8 days (est.)~5 days1.6x

Note: above figures are estimates; real performance depends on software optimization.

3.2 Inference

ModelMLU690 (tok/s)H100 (tok/s)Note
Llama 7B~80 tok/s (est.)~120 tok/sH100 +50%
Llama 70B~20 tok/s (est.)~35 tok/sH100 +75%
Qwen 72B~18 tok/s (est.)~30 tok/sH100 +67%

Key findings:

  • ⚠️ H100 leads inference by 50–75%
  • ✅ But MLU690 draws only 40% the power, with better efficiency
  • ✅ For cost-sensitive inference, MLU690 is more economical

4. Price

4.1 Hardware Procurement

ItemMLU690H100H200
Per-card (domestic)~¥140,000~¥200,000~¥300,000
8-card server (turnkey)~¥1,200,000~¥1,800,000~¥2,600,000
Cost gap-+50%+117%

4.2 TCO (3 years)

ItemMLU690H100Note
Hardware¥1,200,000¥1,800,000MLU690 33% cheaper
Electricity (3y)¥55,200¥110,400MLU690 50% cheaper
Facility¥150,000¥250,000MLU690 40% cheaper
TCO (3y)¥1,405,200¥2,160,400MLU690 35% cheaper

Key findings:

  • MLU690's TCO is 35% lower than H100's
  • ✅ For large-scale deployment (100+ cards), the cost advantage is pronounced

5. Selection Advice

5.1 Choose MLU690 if...

  • ✅ Your business is primarily in the Chinese market
  • ✅ You are affected by U.S. export controls and cannot buy H100/H200
  • ✅ You are power-sensitive (edge data centers, high electricity-cost regions)
  • ✅ Your models use common architectures (Llama, Qwen, GLM)
  • ✅ You have domestic-substitution requirements (government, SOEs, military)

5.2 Choose H100/H200 if...

  • ✅ Your business is global
  • ✅ You need to train frontier models (GPT-4 class)
  • ✅ Your models use complex operators (need the CUDA ecosystem)
  • ✅ You demand extreme performance (low-latency inference)
  • ✅ You can legally procure H100/H200
ScenarioRecommended
TrainingH100 (high perf) + MLU690 (low-cost scale-out)
InferenceMLU690 (cost-sensitive) + H100 (low-latency)
Domestic projectall MLU690
International marketall H100/H200

6. Outlook

6.1 MLU690's weaknesses

  • ⚠️ Immature software ecosystem: 75–85% operator coverage; complex models need manual tuning
  • ⚠️ Small memory: 64GB limits support for 70B+ parameter models
  • ⚠️ Weak interconnect: Cambricon Link bandwidth below NVLink
  • ⚠️ Limited international market: affected by U.S. export controls

6.2 MLU690's improvement path

  • 📅 MLU790 (2027): expected 5nm process, ~2x compute
  • 📅 Memory upgrade: next gen may adopt HBM3e, capacity up to 128GB
  • 📅 Software: NeuWare ecosystem improving, operator coverage target 95%

7. Summary

DimensionMLU690H100Recommended scenario
Compute⭐⭐⭐⭐⭐⭐⭐⭐⭐H100 for top-tier training
Memory⭐⭐⭐⭐⭐⭐⭐H100 for large models
Power⭐⭐⭐⭐⭐⭐⭐⭐MLU690 for inference
Ecosystem⭐⭐⭐⭐⭐⭐⭐⭐H100 for complex models
Price⭐⭐⭐⭐⭐⭐⭐⭐MLU690 for large-scale deployment
Domestic⭐⭐⭐⭐⭐MLU690 for Chinese market

Final recommendation:

  • 🇨🇳 Chinese market: prefer MLU690 (domestic + low cost)
  • 🌍 International market: prefer H100/H200 (performance + ecosystem)
  • 💡 Hybrid: train on H100, infer on MLU690

References


Disclaimer: Data in this article is based on public sources and reasonable estimates; actual performance is subject to vendor official testing. MLU690's software ecosystem is evolving rapidly — watch NeuWare updates.

Last updated: 2026-06-23