First Experimental Demonstration of Disturb-Free 3D Vertical 1T-nC-1T Ferroelectric-based KV Cache with Co-Optimization of Hybrid Analog-Digital CIM and Token-Wise Dynamic Pruning for Efficient Long-Context LLM Inference

For the first time, a ferroelectric (FE)-based key-value (KV) cache for large language models (LLMs) is proposed and experimentally demonstrated. Through device-architecture-algorithm co-optimization, a novel 3D vertical 1T-nC-1T FeRAM structure, featuring orthogonally aligned word-lines/bit-lines and shared FE capacitor (FeCap) strings, as well as a hybrid analog-digital compute-in-memory (CIM) architecture with token-wise dynamic pruning algorithm, are presented. The designed 1T-nC-based analog CIM with small-signal non-destructive read can evaluate similarity in O(1) time for efficient token selection, and the nC-1T-based digital CIM with robust destructive read supports accurate attention computation with a subset of selected KV cache, enabling efficient and accurate processing of dynamically sparse KV cache workloads. Besides, by optimizing data mapping scheme, parallel computation across nC strings is enabled, leading to enhanced throughput and avoided write disturbance. Based on the above design, a 3D 3×32×32 FeCap array is experimentally fabricated with ~10ns switching, 10-year retention, 1016 endurance and good consistency, and the KV cache-based attention is demonstrated with 6×/315× improved performance and energy efficiency over the state-of-the-art designs, along with high accuracy comparable with full attention, showing its great potential for efficient long-context LLM inference.

Paper

Full text

PDF

First Experimental Demonstration of Disturb-Free 3D Vertical 1T-nC-1T Ferroelectric-based KV Cache with Co-Optimization of Hybrid Analog-Digital CIM and Token-Wise Dynamic Pruning for Efficient Long-Context LLM Inference

Semantic Scholar · 2025

Abstract

For the first time, a ferroelectric (FE)-based key-value (KV) cache for large language models (LLMs) is proposed and experimentally demonstrated. Through device-architecture-algorithm co-optimization, a novel 3D vertical 1T-nC-1T FeRAM structure, featuring orthogonally aligned word-lines/bit-lines and shared FE capacitor (FeCap) strings, as well as a hybrid analog-digital compute-in-memory (CIM) architecture with token-wise dynamic pruning algorithm, are presented. The designed 1T-nC-based analog CIM with small-signal non-destructive read can evaluate similarity in O(1) time for efficient token selection, and the nC-1T-based digital CIM with robust destructive read supports accurate attention computation with a subset of selected KV cache, enabling efficient and accurate processing of dynamically sparse KV cache workloads. Besides, by optimizing data mapping scheme, parallel computation across nC strings is enabled, leading to enhanced throughput and avoided write disturbance. Based on the above design, a 3D 3×32×32 FeCap array is experimentally fabricated with ~10ns switching, 10-year retention, 1016 endurance and good consistency, and the KV cache-based attention is demonstrated with 6×/315× improved performance and energy efficiency over the state-of-the-art designs, along with high accuracy comparable with full attention, showing its great potential for efficient long-context LLM inference.

Similar papers

© 2026 NYSGPT2525 LLC