Privacy Leak Detection in LLM Interactions with a User-Centric Approach

In recent years, services based on Large Language Models (LLMs) have garnered increasing attention leading to more frequent interactions between users and LLMs. However, owing to the inherent characteristics of LLMs, user inputs are at risk of privacy leaks. While previous research has proposed methods for protecting user input privacy, many of these approaches face limitations, particularly in adapting to the dynamic and diverse nature of user interactions with LLMs. To address this challenge, our study approaches privacy protection from a user-centric perspective, employing detection methods to safeguard user inputs. Specifically, we compiled a comprehensive privacy list based on General Data Protection Regulation (GDPR) requirements, defined the detection scope, and developed the PA-BERT model for automatically detecting privacy leaks in LLMs. By utilizing the PA-BERT-BiLSTM-CRF architecture, our method effectively monitors private information in user inputs. In our experiments, we not only utilized real-world datasets but also constructed a user input privacy dataset containing 15,036 privacy entities. Experimental results on the constructed dataset show that our method significantly outperforms commonly used general detection models in privacy detection, with marked improvements in precision, recall, and F1 score, making it more effective for identifying privacy in user inputs.

Paper

Full text

PDF

Privacy Leak Detection in LLM Interactions with a User-Centric Approach

OpenAlex · Privacy-Preserving Technologies in Data · 2024

Abstract

In recent years, services based on Large Language Models (LLMs) have garnered increasing attention leading to more frequent interactions between users and LLMs. However, owing to the inherent characteristics of LLMs, user inputs are at risk of privacy leaks. While previous research has proposed methods for protecting user input privacy, many of these approaches face limitations, particularly in adapting to the dynamic and diverse nature of user interactions with LLMs. To address this challenge, our study approaches privacy protection from a user-centric perspective, employing detection methods to safeguard user inputs. Specifically, we compiled a comprehensive privacy list based on General Data Protection Regulation (GDPR) requirements, defined the detection scope, and developed the PA-BERT model for automatically detecting privacy leaks in LLMs. By utilizing the PA-BERT-BiLSTM-CRF architecture, our method effectively monitors private information in user inputs. In our experiments, we not only utilized real-world datasets but also constructed a user input privacy dataset containing 15,036 privacy entities. Experimental results on the constructed dataset show that our method significantly outperforms commonly used general detection models in privacy detection, with marked improvements in precision, recall, and F1 score, making it more effective for identifying privacy in user inputs.

Similar papers

© 2026 NYSGPT2525 LLC