The proliferation of streaming platforms and smart television content services has led to a large amount of multimedia content availability; thus, it becomes harder for viewers to choose suitable options due to severe content overload. In most cases, traditional methods based on remote control navigation prove to be not only inconvenient but also ineffective, especially during a search across many apps, genres, and customized menus. It becomes clear that there is an urgent need for modern personalized solutions to facilitate intelligent voice interactions that recognize natural speech commands and user contexts. In this paper, we propose a Transformer-Based Voice Understanding System for Smart Television Content Personalization, which provides more efficient content navigation by using state-of-the-art spoken language understanding and fusion of recommendations. Specifically, we apply a transformer encoder to learn semantic intents from user queries and viewing history to fuse the obtained information with recommendations, which can incorporate household preferences, temporal patterns, and content similarities in the process. The effectiveness of the developed model is estimated via benchmarks of smart television interaction data along with simulations of the household viewing preferences. The performance indicators such as accuracy of intent extraction, precision and recall of recommendations, and overall user satisfaction are calculated. The experimental results show that the proposed approach outperforms common methods based on CNN, LSTM, and attention mechanisms.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex