// TOM'S HARDWARE US — HARDWARE & GADGET
Hot Chips 2026: Intel dives deep on Crescent Island AI accelerator — larger caches and deeper XMX engines target maximum AI FLOPS per watt
Xe3P focuses on the most compute-intensive phases of AI inference.
When you purchase through links on our site, we may earn an affiliate commission. Here’s how it works.
Intel shared more details of its Crescent Island AI accelerator, powered by the Xe3P architecture, at the Hot Chips symposium this week. Unlike Nvidia's Rubin and AMD's MI455X GPUs, which are high-power, exclusively liquid-cooled chips with massive pools of HBM4 memory that provide maximum performance across both AI training and inference workloads, Crescent Island is designed to fit into a lower-power, inference-first niche in the AI accelerator market.
As a refresher, Crescent Island is a 350W air-cooled PCIe card that uses up to 480 GB of LPDDR5X memory, meaning it can be deployed in traditional servers without exotic power and cooling requirements. We've already learned about some of Crescent Island's DNA from past disclosures, but Intel went deeper into the chip's architectural details at Hot Chips.
Crescent Island is built up from four Xe3P slices, each containing eight Xe Cores, for a total of 32. Each Xe Core has eight Xe Vector Engines and eight XMX matrix accelerators, for a total of 256 of each resource.
The Xe3 graphics architecture, as seen on Intel's Panther Lake processors, already modified the capacity and flexibility of the GPU cache hierarchy to improve utilization and decrease performance-sapping register spills, and Xe3P further refines that hierarchy.
The Xe3P Xe Core has twice the amount of general register file space for working data versus Battlemage. Each Xe Core now has 1MB of general-purpose register file space, up from 512KB on Battlemage and Xe2. In addition, Xe3P offers 512KB of L1 cache or shared local memory per Xe Core, a structure that started out at 256KB on Battlemage and grew by approximately 1.33x on Panther Lake's Xe3 GPU. The chip also has 32MB of shared L2 cache. These expanded caches are meant to serve the chip's larger matrix accelerators on its AI compute-focused mission.
Xe3P boasts a larger systolic depth in its XMX engines than past Xe GPU designs. Xe3P's XMX systolic engines are a 16-deep design, meaning they can process matrices in much larger chunks than the four-deep systolic design of Xe2 and Xe3. Nvidia doesn't discuss the architecture of its Tensor Cores in anywhere near this level of detail, but as an AI inference-focused part, the fact that Xe3P can theoretically work on more elements at once during general matrix-multiply operations is an important capability boost for Crescent Island's inference ambitions.
Intel is also prioritizing a broad range of data types with this chip, from FP4 formats with microscaling support (aka MXFP4) all the way to what it describes as full-rate double-precision (via 64 FP64 FMA units per Xe Core). FP64 isn't widely used in AI workloads, but Intel says that the inclusion of full-rate processing for that data type makes Crescent Island useful as a converged high-performance computing and AI chip.
Each Xe Core also supports sigmoid and tanh transcendental functions, which are important to a variety of operations during AI inference, especially the softmax function. AMD and Nvidia have prioritized the performance of these functions in their recent architectures as well, so the fact that Xe3P offers support for them is key for its AI-first initiatives.