How to Reach Real-Time AI on Consumer Devices? Solutions for Programmable and Custom Architectures
The unprecedented performance of deep neural networks (DNNs) has led to large\nstrides in various Artificial Intelligence (AI) inference tasks, such as object\nand speech recognition. Nevertheless, deploying such AI models across commodity\ndevices faces significant challenges: large computational cost, multiple\nperformance objectives, hardware heterogeneity and a common need for high\naccuracy, together pose critical problems to the deployment of DNNs across the\nvarious embedded and mobile devices in the wild. As such, we have yet to\nwitness the mainstream usage of state-of-the-art deep learning algorithms\nacross consumer devices. In this paper, we provide preliminary answers to this\npotentially game-changing question by presenting an array of design techniques\nfor efficient AI systems. We start by examining the major roadblocks when\ntargeting both programmable processors and custom accelerators. Then, we\npresent diverse methods for achieving real-time performance following a\ncross-stack approach. These span model-, system- and hardware-level techniques,\nand their combination. Our findings provide illustrative examples of AI systems\nthat do not overburden mobile hardware, while also indicating how they can\nimprove inference accuracy. Moreover, we showcase how custom ASIC- and\nFPGA-based accelerators can be an enabling factor for next-generation AI\napplications, such as multi-DNN systems. Collectively, these results highlight\nthe critical need for further exploration as to how the various cross-stack\nsolutions can be best combined in order to bring the latest advances in deep\nlearning close to users, in a robust and efficient manner.\n