A Machine Learning Framework for Detecting and Mitigating Adversarial Prompt Attacks in Large Language Models

Large Language Models LLMs have changed numerous industries including education cybersecurity healthcare robotics and Internet of Things IoT by empowering them to converse and comprehend the context and generate intelligent arguments to power intelligent systems. Since these models are deployed as part of software platforms robotics and linked infrastructures their reliability and trustworthiness becomes a key issue. Nonetheless, the bulk of the available literature is concerned with the optimization of model performances and its usage and security risks and adversarial threats are not given much consideration. New attacks like jailbreaking can bypass safety nets and cause great harm to users and systems based on LLMs like multi turn manipulation and injection by emergency attacks. This paper fills this gap by building and analyzing a varied dataset of prompt and safe and adversarial prompts generated by human professionals and advanced LLMs including GPT4o and Claude 3.5 Sonnet. The controversial intent is detected by use of TF IDF feature extraction and machine learning classifiers such as K Nearest Neighbors Decision Tree Logistic Regression random forest Naive Bayes Support Vector machine and ensemble model. The SVM model is the most successful model with 98.4 percent accuracy and 98.7 percent F1 score as well as 99.7 percent recall.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC