Memory-Latency-Accuracy Trade-offs for Continual Learning on a RISC-V\n Extreme-Edge Node

AI-powered edge devices currently lack the ability to adapt their embedded\ninference models to the ever-changing environment. To tackle this issue,\nContinual Learning (CL) strategies aim at incrementally improving the decision\ncapabilities based on newly acquired data. In this work, after quantifying\nmemory and computational requirements of CL algorithms, we define a novel HW/SW\nextreme-edge platform featuring a low power RISC-V octa-core cluster tailored\nfor on-demand incremental learning over locally sensed data. The presented\nmulti-core HW/SW architecture achieves a peak performance of 2.21 and 1.70\nMAC/cycle, respectively, when running forward and backward steps of the\ngradient descent. We report the trade-off between memory footprint, latency,\nand accuracy for learning a new class with Latent Replay CL when targeting an\nimage classification task on the CORe50 dataset. For a CL setting that retrains\nall the layers, taking 5h to learn a new class and achieving up to 77.3% of\nprecision, a more efficient solution retrains only part of the network,\nreaching an accuracy of 72.5% with a memory requirement of 300 MB and a\ncomputation latency of 1.5 hours. On the other side, retraining only the last\nlayer results in the fastest (867 ms) and less memory hungry (20 MB) solution\nbut scoring 58% on the CORe50 dataset. Thanks to the parallelism of the\nlow-power cluster engine, our HW/SW platform results 25x faster than typical\nMCU device, on which CL is still impractical, and demonstrates an 11x gain in\nterms of energy consumption with respect to mobile-class solutions.\n

Paper

References (30)

Scroll for more · 18 remaining

Similar papers

© 2026 NYSGPT2525 LLC