NVIDIA has expanded NVLink Fusion with NVIDIA NVHBM, a custom high-bandwidth memory approach intended to increase memory bandwidth and reduce power for XPUs used in semi-custom AI infrastructure. Amazon’s Annapurna Labs is set to be the first collaborator working on NVHBM as part of its broader work with NVIDIA around NVLink Fusion.
NVHBM shifts the memory controller out of the XPU die and into the HBM base die. NVIDIA describes standard HBM architectures as placing the memory controller on the XPU die, consuming silicon area that could otherwise be used for compute. With NVHBM, NVIDIA’s custom memory controller is integrated into the 3D HBM stack instead of the XPU.
NVIDIA lists three quantified impacts versus standard HBM4E: up to 30% greater memory bandwidth, 15% lower HBM power consumption, and up to 25% more area freed on the XPU compute die. For data center architects pushing large-model inference and training, those numbers matter because memory bandwidth and the thermal and power costs of feeding accelerators can become limiting factors at rack scale.
NVIDIA said it is establishing a standard NVHBM implementation that will be validated and offered by multiple memory partners. The stated goal is to reduce engineering work to integrate and qualify memory across multiple suppliers, giving NVLink Fusion customers a faster path to bringing custom AI chips to market.
On the hyperscaler side, Annapurna Labs will work with NVIDIA on NVHBM technology and the NVLink scale-up architecture. The companies also tied the effort to Trainium: Annapurna Labs will support NVLink Fusion with its next-generation Trainium chips starting with Trainium4, which is intended to allow Amazon chips and NVIDIA GPUs to operate together within a common rack-scale architecture.
“NVHBM represents a new architectural approach to advancing high-bandwidth memory performance and efficiency,” said Nafea Bshara, vice president of Annapurna Labs at Amazon. “We look forward to this technology collaboration to benefit future AWS infrastructure designs.”
NVLink Fusion is described as a way for partners to connect custom XPUs and CPUs into NVIDIA’s rack-scale platform, with access to NVLink chiplets, NVLink-C2C, NVLink Switches, and NVIDIA MGX systems and racks.
Source: NVIDIA


















