Learned, Lagged, LLM-splained: LLM Responses to End User Security Questions

Answering end user security questions is challenging. While large language models (LLMs) like GPT, Llama, and Gemini are far from error-free, they have shown promise in answering a variety of questions outside of security. We qualitatively evaluated responses from three popular LLMs to 900 systematically collected security questions in the first such evaluation in the area of end user security. While LLMs demonstrate broad generalist “knowledge” of end user security information, there are patterns of errors and limitations across LLMs-including stale, inaccurate, and incomplete answers, as well as indirect or unresponsive communication styles-which negatively impact the user experience. Based on these patterns, we suggest directions for improved model development and recommend user strategies for interacting with LLMs when seeking assistance with security.

Paper

Similar papers

© 2026 NYSGPT2525 LLC