Modern AI systems deployed in large-scale production environments are increasingly prone to performance degradation due to data drift, infrastructure failures, concept shift, and unexpected system interactions. Existing approaches primarily focus on failure detection and alerting, relying heavily on manual intervention for recovery and remediation. This paper proposes an autonomous self-healing AI framework that leverages multiagent reinforcement learning to continuously monitor, diagnose, and remediate system-level failures without human intervention. The proposed system models monitoring, diagnosis, and recovery as cooperative learning agents that operate over a shared system-state graph. Each agent learns adaptive policies to detect anomalies, identify root causes, and execute corrective actions such as model reconfiguration, retraining, or resource reallocation. We introduce a hierarchical coordination mechanism that enables agents to balance short-term recovery actions with longterm system stability. Experimental evaluation across simulated and real-world-inspired failure scenarios demonstrates that the proposed framework reduces recovery time by up to 47%, improves system availability by 32%, and significantly outperforms traditional rule-based and single-agent baselines. These results highlight the feasibility of fully autonomous, self-healing AI systems for mission-critical deployments.
Paper
Full text
Self-Healing AI Systems Using Multi-Agent Learning
Semantic Scholar · 2026
Abstract
Modern AI systems deployed in large-scale production environments are increasingly prone to performance degradation due to data drift, infrastructure failures, concept shift, and unexpected system interactions. Existing approaches primarily focus on failure detection and alerting, relying heavily on manual intervention for recovery and remediation. This paper proposes an autonomous self-healing AI framework that leverages multiagent reinforcement learning to continuously monitor, diagnose, and remediate system-level failures without human intervention. The proposed system models monitoring, diagnosis, and recovery as cooperative learning agents that operate over a shared system-state graph. Each agent learns adaptive policies to detect anomalies, identify root causes, and execute corrective actions such as model reconfiguration, retraining, or resource reallocation. We introduce a hierarchical coordination mechanism that enables agents to balance short-term recovery actions with longterm system stability. Experimental evaluation across simulated and real-world-inspired failure scenarios demonstrates that the proposed framework reduces recovery time by up to 47%, improves system availability by 32%, and significantly outperforms traditional rule-based and single-agent baselines. These results highlight the feasibility of fully autonomous, self-healing AI systems for mission-critical deployments.