AE-RAGX: Combining Autoencoders with Retrieval-Augmented Generation for Explainable Anomaly Detection using LLMs
Detecting anomalies in high-dimensional or unstructured data remains a challenge, as many traditional approaches such as Autoencoders (AEs), Variational Autoencoders (VAEs), and One-Class SVMs rely heavily on reconstruction errors or distance measures. These techniques often overlook contextual information, struggle with subtle deviations, and provide limited interpretability. To address these gaps, we propose AE-RAGX, a hybrid architecture that combines Autoencoders with Retrieval-Augmented Generation (RAG) using Large Language Models (LLMs). By transforming structured data into descriptive text and incorporating rule-based knowledge, the model not only improves anomaly detection but also produces transparent, human-readable explanations. Evaluation on a credit card fraud dataset shows that AE-RAGX enhances recall and provides interpretable outputs, making it well suited for high-stakes applications such as fraud detection.
Paper
Full text
AE-RAGX: Combining Autoencoders with Retrieval-Augmented Generation for Explainable Anomaly Detection using LLMs
Semantic Scholar · Computer Science · 2025
Abstract
Detecting anomalies in high-dimensional or unstructured data remains a challenge, as many traditional approaches such as Autoencoders (AEs), Variational Autoencoders (VAEs), and One-Class SVMs rely heavily on reconstruction errors or distance measures. These techniques often overlook contextual information, struggle with subtle deviations, and provide limited interpretability. To address these gaps, we propose AE-RAGX, a hybrid architecture that combines Autoencoders with Retrieval-Augmented Generation (RAG) using Large Language Models (LLMs). By transforming structured data into descriptive text and incorporating rule-based knowledge, the model not only improves anomaly detection but also produces transparent, human-readable explanations. Evaluation on a credit card fraud dataset shows that AE-RAGX enhances recall and provides interpretable outputs, making it well suited for high-stakes applications such as fraud detection.