MINDSTATE: Systematic Analysis of Belief Manipulation Vulnerabilities in Multi-Agent LLM Networks

As large language models (LLMs) are increasingly deployed in multi-agent systems, from serverless architectures to collaborative AI workflows, understanding their social and cognitive vulnerabilities becomes critical. Despite rapid growth, little attention has been paid to how belief manipulation may propagate across interacting LLM agents. To address this gap, this paper presents MINDSTATE, a Python simulation platform for studying belief manipulation in LLM networks through controlled false memory injection experiments. The core methodology includes a memory injection module, multi-agent belief tracking, baseline model comparisons and cross-model validations. False memories are injected into the agent context to assess manipulation resistance, and belief evolution is monitored using embedding analysis. The baseline model comparison tests are conducted against established consensus models, and crossmodel validation experiments are performed across GPT-3.5 and GPT-4. The core discovery from this project reveals AI safety implications: LLM agents exhibit over-agreeability, reaching consensus 2-5x faster than established social dynamics models. They converge in 9-17 iterations compared to 20-50 iterations for baseline models like DeGroot consensus. This excessive convergence reveals systematic vulnerabilities in deployed AI systems. Memory manipulation attacks can compromise AI networks within 3-6 iterations, with false beliefs persisting indefinitely. These results reveal fundamental differences from human social dynamics and raise concerns for deployed multiagent systems due to the implications for AI safety.

Paper

Full text

PDF

MINDSTATE: Systematic Analysis of Belief Manipulation Vulnerabilities in Multi-Agent LLM Networks

Semantic Scholar · 2025

Abstract

As large language models (LLMs) are increasingly deployed in multi-agent systems, from serverless architectures to collaborative AI workflows, understanding their social and cognitive vulnerabilities becomes critical. Despite rapid growth, little attention has been paid to how belief manipulation may propagate across interacting LLM agents. To address this gap, this paper presents MINDSTATE, a Python simulation platform for studying belief manipulation in LLM networks through controlled false memory injection experiments. The core methodology includes a memory injection module, multi-agent belief tracking, baseline model comparisons and cross-model validations. False memories are injected into the agent context to assess manipulation resistance, and belief evolution is monitored using embedding analysis. The baseline model comparison tests are conducted against established consensus models, and crossmodel validation experiments are performed across GPT-3.5 and GPT-4. The core discovery from this project reveals AI safety implications: LLM agents exhibit over-agreeability, reaching consensus 2-5x faster than established social dynamics models. They converge in 9-17 iterations compared to 20-50 iterations for baseline models like DeGroot consensus. This excessive convergence reveals systematic vulnerabilities in deployed AI systems. Memory manipulation attacks can compromise AI networks within 3-6 iterations, with false beliefs persisting indefinitely. These results reveal fundamental differences from human social dynamics and raise concerns for deployed multiagent systems due to the implications for AI safety.

Similar papers

© 2026 NYSGPT2525 LLC