Skip to main content

Intel Cancels Falcon Shores, Pivots to Jaguar Shores: From Single-Chip Competition to Rack-Scale Systems

· 5 min read
Industry Research Team

May 14, 2026, Intel disclosed in its Q1 earnings report that it has formally cancelled the Falcon Shores single-chip GPU project and confirmed a new rack-scale AI system project named Jaguar Shores to launch in 2027-2028. This is a major strategic adjustment in Intel's AI roadmap. This article provides an in-depth analysis of the reasons and future implications.

Inference Optimization Technology Evolution: PagedAttention / FlashAttention / Speculative Decoding Deep Dive

· 8 min read
Industry Research Team

LLM inference performance = Algorithm + Software + Hardware. Hardware (H100, B300, Rubin) only determines the theoretical ceiling. Actual inference performance can be improved 5-30× through algorithmic optimization. This article provides a deep analysis of the three major inference optimization technologies: PagedAttention, FlashAttention, and Speculative Decoding.

Apple Silicon Comeback: M3 Ultra 192GB UMA Local LLM Revolution

· 8 min read
Industry Research Team

Apple Silicon is staging a comeback in the AI era. The M3 Ultra in a single Mac Studio packs 192GB unified memory (UMA) and an 80-core GPU, capable of running 70B-200B parameter LLMs locally without quantization. This is a revolution in consumer/workstation-class AI inference. This article provides an in-depth analysis of Apple Silicon's AI advantages, current ecosystem, and future.

NVIDIA Acquires Groq for $20 Billion: LPU Officially Enters the NVIDIA Ecosystem

· 4 min read
Industry Research Team

In Q1 2026, one of the biggest pieces of news in the AI chip industry: NVIDIA acquired Groq for approximately $20 billion in a full acquisition. This means Groq's LPU architecture officially becomes part of NVIDIA's compute landscape, complementing GPUs. This article analyzes the strategic significance of this acquisition in detail.

AWS Trainium 3 GA: 3nm Process + 4.4× Compute + 4× Efficiency + 144-Chip UltraServer

· 4 min read
Industry Research Team

On December 2, 2025, at the re:Invent 2025 conference, AWS formally GA'd its third-generation custom AI training chip Trainium 3. This is a critical upgrade to the AWS compute landscape: 3nm process, 4.4× compute improvement, 4× efficiency improvement, Trn3 UltraServer with 144 chips. This article provides a detailed analysis.

Huawei Ascend 920: China's Highest Bandwidth at 4 Tbps + 3× H20 Compute for Domestic Substitution

· 5 min read
Industry Research Team

Huawei Ascend 920 (昇腾 920) entered large-scale mass production in 2025 H2, representing a major breakthrough for Chinese domestic AI chips. This article analyzes its specifications, comparison with NVIDIA H20, the CloudMatrix 384 Ultra system, and its significance for China's AI industry.