Can Large Language Models Be Used to Conduct Attacks on Smart Home Devices?

Smart Home Systems (SHS) are increasingly vulnerable to adversarial manipulation due to their reliance on networked sensors and automated decision making. This paper explores the misuse potential of Large Language Models (LLMs), specifically GPT-4, to generate context-aware cyber-physical attacks targeting SHS fire alarm subsystems. We design five adversarial scenarios: dynamic sensor tampering, gradual calibration drift, environmental Trojan triggers, missing value injection, and device overload using prompt-based LLM code generation to simulate realistic data manipulation. To evaluate system resilience, we benchmark seven classical machine learning (ML) classifiers on a real-world smoke detection IoT dataset under these adversarial conditions. Our results show that while ensemble-based models such as Random Forest and Gradient Boosting maintain high accuracy in detecting tampered inputs, simpler models like Logistic Regression and Naive Bayes (which are more likely to be employed in SHS applications) degrade significantly. The study reveals that LLMs can automatically synthesize highly effective attack logic with no access to internal model parameters, posing a novel threat vector in IoT cybersecurity. These findings highlight the need to protect ML-based SHS systems against generative AI-based attacks through the incorporation of robust, adaptive, and light-weight defense mechanisms.

Paper

Full text

PDF

Can Large Language Models Be Used to Conduct Attacks on Smart Home Devices?

Semantic Scholar · 2025

Abstract

Smart Home Systems (SHS) are increasingly vulnerable to adversarial manipulation due to their reliance on networked sensors and automated decision making. This paper explores the misuse potential of Large Language Models (LLMs), specifically GPT-4, to generate context-aware cyber-physical attacks targeting SHS fire alarm subsystems. We design five adversarial scenarios: dynamic sensor tampering, gradual calibration drift, environmental Trojan triggers, missing value injection, and device overload using prompt-based LLM code generation to simulate realistic data manipulation. To evaluate system resilience, we benchmark seven classical machine learning (ML) classifiers on a real-world smoke detection IoT dataset under these adversarial conditions. Our results show that while ensemble-based models such as Random Forest and Gradient Boosting maintain high accuracy in detecting tampered inputs, simpler models like Logistic Regression and Naive Bayes (which are more likely to be employed in SHS applications) degrade significantly. The study reveals that LLMs can automatically synthesize highly effective attack logic with no access to internal model parameters, posing a novel threat vector in IoT cybersecurity. These findings highlight the need to protect ML-based SHS systems against generative AI-based attacks through the incorporation of robust, adaptive, and light-weight defense mechanisms.

Similar papers

© 2026 NYSGPT2525 LLC