Multimodal interactive narration with virtual reality and exhibition and transformer-based audience behavior mining and plot adaptive algorithm
The progress of augmented reality (AR) and virtual reality (VR) technology has promoted the development of the virtual-real integration. Because the traditional fixed narrative mode is difficult to adapt to the individual differences of the audience, it leads to insufficient immersion and personalized experience. The existing methods still have limitations in multimodal data fusion, audience intention analysis and dynamic plot generation. Therefore, this study proposes a multimodal fusion interactive narrative model based on Transformer model, and constructs a three-tier architecture of multimodal perception fusion, audience behavior mining and adaptive plot generation. A cross-modal Transformer encoder is designed to deeply integrate visual, auditory, position trajectory and interactive log data to generate high-quality unified representation. The converter decoder is used to model the behavior sequence, perform macro behavior classification, interest regression and understanding estimation in parallel. A comprehensive behavior intention vector is constructed. The audience state is comprehensively captured. Combining the framework of Gated Recurrent Unit and Proximal Policy Optimization, a coherent, personalized and low-repetition plot is dynamically generated through a composite reward function. The final experimental results show that the model is significantly better than the mainstream baseline in indicators, and the ablation experiment verifies the effectiveness of each core module. The overall research can provide a complete technical solution for realizing intelligent and personalized exhibition experience, and has a good application prospect.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex