CHIME: Chiplet-based Heterogeneous Near-Memory Acceleration for Edge Multimodal LLM Inference
The proliferation of large language models (LLMs) is accelerating the integration of multimodal assistants into edge devices, where inference is executed under stringent latency and energy constraints, often exacerbated by intermittent connectivity. These challenges become particularly acute in the context of multimodal LLMs (MLLMs), as high-dimensional visual inputs are transformed into extensive token sequences, thereby inflating the key-value (KV) cache and imposing substantial data movement overheads to the LLM backbone. We present CHIME, a chiplet-based heterogeneous near-memory accelerator for edge MLLM inference. CHIME pairs monolithic-3D (M3D) DRAM for low-latency, bandwidth-hungry attention with M3D RRAM for dense, non-volatile weight storage, and uses a co-designed mapping framework that executes fused kernels near data to minimize cross-chiplet traffic and maximize effective bandwidth. On FastVLM (0.6B/1.7B) and MobileVLM (1.7B/3B), CHIME achieves up to 54× speedup and 246× energy efficiency per inference over NVIDIA Jetson Orin NX, sustaining 116.5–266.5 token/J vs. 0.7–1.1 token/J. It delivers up to 69.2× higher throughput than FACIL. Compared to an M3D DRAM-only design, heterogeneous memory improves energy efficiency by 7% and performance by 2.4×.
Paper
References (35)
Scroll for more · 23 remaining