AI Compilation and Hardware-Software Co-Design for Resource-Constrained Edge Devices: A Brief Review

Edge artificial intelligence (Edge AI) has emerged as a promising solution to provide real-time and privacy-aware intelligent services at proximity to data sources. However, it is difficult to implement deep learning on edge devices with limited resources, such as microcontrollers, embedded devices, and IoT terminals with low power consumption, because they have limited computing power, memory capacity, and energy supply. This paper introduces a concise overview on AI compilation and hardware co-design methodologies for efficient inference execution on resource-constrained edge devices. This paper provides an overview of the key challenges in deploying deep learning models on resource-constrained devices and presents illustrative solutions in terms of compression techniques and hardware implementation. The discussion includes graph optimization, operator fusion, quantization, memory-aware inference, hardware-specific code generation, accelerator-based inference, latency prediction, and runtime optimization. The reviewed work suggests that high-performance efficient edge AI requires cross-layer rather than isolated model-level compression. AI compilation enables more efficient execution by converting neural networks to optimized hardware-aware code, and hardware- co-design further improves latency, memory, and energy efficiency. These results imply that next-generation edge AI systems should focus on portable compiler toolchains, memory-aware optimization, energy-aware design, and standardized benchmarking methodologies.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC