VLSI Design of a Fast and Reprogrammable Decision Tree Inference Accelerator

Tree-based models are highly efficient in terms of prediction accuracy and energy consumption, making them very suitable for embedded devices. However, current ASIC implementations are either not reprogrammable or are not accompanied by a model-to-memory translation method for fast deployment. This work proposes a reprogrammable hardware accelerator for Decision Tree (DT) inference, combined with its memory-mapping procedure, targeting easy deployment for energy-efficient embedded applications. To allow an efficient design space exploration, the architecture was designed to support different levels of parallelism and input bitwidth. Validation was conducted using testbenches and a UVM framework to ensure functional correctness. Experimental results using the SkyWater 130nm PDK show that the accelerator consumes 0.99mW for a 3-bit input, up to 790.73mW for 16-bit inputs and deeper trees. Compared to a CPU implementation, the designed solutions provides speed-ups between $8 \times$ and $25 \times$, operating at 100MHz. This scalable solution demonstrates the viability of fast and energy-efficient decision tree inference for embedded AI systems.

Paper

Full text

PDF

VLSI Design of a Fast and Reprogrammable Decision Tree Inference Accelerator

OpenAlex · Advanced Neural Network Applications · 2025

Abstract

Tree-based models are highly efficient in terms of prediction accuracy and energy consumption, making them very suitable for embedded devices. However, current ASIC implementations are either not reprogrammable or are not accompanied by a model-to-memory translation method for fast deployment. This work proposes a reprogrammable hardware accelerator for Decision Tree (DT) inference, combined with its memory-mapping procedure, targeting easy deployment for energy-efficient embedded applications. To allow an efficient design space exploration, the architecture was designed to support different levels of parallelism and input bitwidth. Validation was conducted using testbenches and a UVM framework to ensure functional correctness. Experimental results using the SkyWater 130nm PDK show that the accelerator consumes 0.99mW for a 3-bit input, up to 790.73mW for 16-bit inputs and deeper trees. Compared to a CPU implementation, the designed solutions provides speed-ups between $8 \times$ and $25 \times$, operating at 100MHz. This scalable solution demonstrates the viability of fast and energy-efficient decision tree inference for embedded AI systems.

Similar papers

© 2026 NYSGPT2525 LLC