From Visual Perception to Multimodal Fusion: Application Analysis of Artificial Intelligence in Physical Retail Environment
Under the continuous impact of e-commerce, traditional physical retail needs to complete the digital transformation and upgrading to the "retail 4.0" mode. This paper systematically reviews the application of computer vision and multimodal fusion technology in the real retail scene. Firstly, the research sorted out the five stages of visual perception evolution from single mode to multimodal, and compared the core characteristics and development context of target detection and behavior recognition technology; Then, the mathematical model and collaboration mechanism of the fusion of vision and RFID, audio, radar and other modes are analyzed in depth, and an intelligent retail ecological framework covering bottom perception, middle fusion, and top-level applications is built, covering key application directions such as unmanned retail, digital twins, and supply chain optimization. The research confirms that single vision technology has the limitation of scene adaptation, and multimodal fusion is the necessary path to achieve high-precision customer insight and independent retail. Finally, this paper sorts out the problems of large-scale applications such as privacy protection, model robustness, edge computing resource constraints, and puts forward future research directions, such as lightweight end-to-side model development, interactive system construction of fusion generative large model, and exploration of human-computer deep collaboration mechanism.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex