WirelessSenseLLM: Zero-Shot Human Activity Understanding by Bridging Wireless Signals and Human Language
There is a growing interest in enabling wireless sensing systems to interpret human motion from unsegmented wireless signals; however, existing CSI-based applications rely heavily on accurate signal segmentation and predefined action labels, which limit their applicability in zero-shot scenarios. We present WirelessSenseLLM, a language-driven framework that leverages large language models (LLMs)to enable zero-shot human motion understanding from unsegmented Wi-Fi Channel State Information (CSI). To bridge the modality gap between time-series CSI and discrete language representations, we introduce a CSI-to-Language Adapter and a cross-modal projection mechanism that allows the CSI feature to be mapped into a language-aligned semantic space. This design enables the generation of fine-grained natural language descriptions of sequential and overlapping human motions, supporting down-stream reasoning without segmented training data. We address two core technical challenges: modality mismatch between CSI features and language embeddings, and overlapping actions in unsegmented CSI streams. Extensive experiments demonstrate strong performance in zero-shot action understanding ($92 \%$ accuracy and $91 \% \mathrm{~F} 1$-score), language-based reasoning quality ($30 \%$ factual and $15 \%$ reasoning improvements), and multi-person motion explanation with an average $12.33 \%$ improvement over prior methods. These results highlight WirelessSenseLLM’s effectiveness for robust, interpretable human motion understanding from CSI signals.
Paper
References (34)
Scroll for more · 22 remaining