Using Image Recognition and Local Large Language Model-Driven Dialogue for Confirming Dynamic Intentions

Understanding human intention is critical in healthcare, where direct communication might be limited. We developed a dynamic intention recognition system for elderly care, integrating visual data and conversational interactions. Building upon in-home fixed-point cameras and lightweight image recognition systems, facial, skeletal, and hand features were detected to monitor behavior changes. Image-based analysis alone is often insufficient to fully grasp the intentions behind observed actions. Therefore, we developed a multimodal system that combines (1) image-derived contextual cues and (2) dialogue with an interactive agent, analyzed through locally executable large language models (LLMs). Specifically, we utilized locally deployable models such as Meta-Llama and Llama 3.1 within a web-based interface built on AnythingLLM, enabling real-time conversational intention estimation without compromising data privacy. The developed system performs image-based context extraction, conversational prompting with embedded intentions, and time-windowed log analysis using local LLMs for chronological intention estimation and confirmation. Five participants were recruited to compare their self-reported intentions with those inferred by the system using document similarity metrics. The results demonstrated the feasibility of the system and its real-time, privacy-conscious performance in elderly home care. The system contributes to the advancement of human-agent collaboration and lays a foundation for adaptive, intention-aware support systems.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC