We highlight four promising research opportunities to improve large language model inference for datacenter AI: high bandwidth flash for 10X memory capacity with HBM-like bandwidth, processing-near-memory and 3D memory-logic stacking for high memory bandwidth, and low-latency interconnect to speedup communication. We also review their applicability for mobile devices.
Paper
References (14)
09SanDisk's new High Bandwidth Flash memory , 2025, Tom's Hardware
102504 Memory-ExpansionStructera™ A
11Microsoft Is Losing a Staggering Amount of Money onAI
12“Structera™ A 2504 memory-expan-sion controller,”Marvell, Santa Clara, CA, USA
Scroll for more · 2 remaining