Page 1 of 1

SE Hybrid HBM-HBF Architecture in LLM Inference (University of Oxford)

Posted: Sun Aug 30, 2026 7:06 am
by admin
Researchers at the University of Oxford published a technical paper titled “Hardware-Managed Heterogeneous High-Bandwidth Memory and Flash in LLM Inference Systems.” Abstract Excerpt: “ High-Bandwidth Flash (HBF) offers a denser alternative, providing 16x more capacity per stack at comparable bandwidth. In this work, we show that while replacing HBM with HBF can address the capacity problem, doing so naively severely impacts performance due to HBF’s long tail memory latency starving GPU schedulers. To address this, we propose a Heterogeneous Memory Architecture (HMA) that combines HBM and HBF through a prediction-based migration policy to keep high latency HBF off the GPU’s critical path. “ Find the technical paper here. August 2026. Atassi, Hakam, Noa Zilberman, and Amro Awad. “Hardware-Managed Heterogeneous High-Bandwidth Memory and Flash in LLM Inference Systems.” IEEE Computer Architecture Letters (August 2026). https://doi.org/10.1109/LCA.2026.3723326     The post Hybrid HBM-HBF Architecture in LLM Inference (University of Oxford) appeared first on Semiconductor Engineering.

Source: https://semiengineering.com/hybrid-hbm- ... of-oxford/