Explainable Multimodal Event Knowledge Graph (E²-KG): Modeling Cognitive–Motor Coupling through Manifold Learning and Graph-Based Reasoning

Recent advancements in Artificial Intelligence (AI) and Machine Learning (ML) have achieved remarkable success in natural language processing, computer vision, and reinforcement learning. However, when applied to human behavior understanding and motion prediction, most deep learning architectures remain black-box models—lacking interpretability in how perceptual cues and motor actions interact to form human decision-making patterns. To address this limitation, this study proposes an Explainable Multimodal Event Knowledge Graph (E²-KG) framework that integrates eye-tracking data and OpenPose-based joint angles to construct a semantic representation of cognitive–motor coupling in dynamic tasks. Instead of treating perception and motion as separate streams, E²-KG embeds them within a unified perception–action manifold, where local energy gradients automatically detect event transitions (fixation, preparation, release, follow-through). These events are organized into a knowledge graph structure, enabling interpretable reasoning about causal links between visual attention, body movement, and task outcomes. A perception–action manifold model capturing the joint evolution of gaze and motion through local energy variations. A multimodal event extraction pipeline for synchronized eye–pose data. An event knowledge graph that supports cause–effect reasoning. Validation through a throwing-distance prediction task, demonstrating higher interpretability and transparency than data-driven baselines. E²-KG thus bridges manifold learning and knowledge-graph reasoning, offering a transparent, cognitively grounded approach for explainable human performance modeling—applicable to domains such as human factors, adaptive training, and human–robot collaboration.

Paper

Full text

PDF

Explainable Multimodal Event Knowledge Graph (E²-KG): Modeling Cognitive–Motor Coupling through Manifold Learning and Graph-Based Reasoning

Semantic Scholar · Computer Science · 2025

Abstract

Recent advancements in Artificial Intelligence (AI) and Machine Learning (ML) have achieved remarkable success in natural language processing, computer vision, and reinforcement learning. However, when applied to human behavior understanding and motion prediction, most deep learning architectures remain black-box models—lacking interpretability in how perceptual cues and motor actions interact to form human decision-making patterns. To address this limitation, this study proposes an Explainable Multimodal Event Knowledge Graph (E²-KG) framework that integrates eye-tracking data and OpenPose-based joint angles to construct a semantic representation of cognitive–motor coupling in dynamic tasks. Instead of treating perception and motion as separate streams, E²-KG embeds them within a unified perception–action manifold, where local energy gradients automatically detect event transitions (fixation, preparation, release, follow-through). These events are organized into a knowledge graph structure, enabling interpretable reasoning about causal links between visual attention, body movement, and task outcomes. A perception–action manifold model capturing the joint evolution of gaze and motion through local energy variations. A multimodal event extraction pipeline for synchronized eye–pose data. An event knowledge graph that supports cause–effect reasoning. Validation through a throwing-distance prediction task, demonstrating higher interpretability and transparency than data-driven baselines. E²-KG thus bridges manifold learning and knowledge-graph reasoning, offering a transparent, cognitively grounded approach for explainable human performance modeling—applicable to domains such as human factors, adaptive training, and human–robot collaboration.

References (23)

Scroll for more · 11 remaining

Similar papers

© 2026 NYSGPT2525 LLC