The rapid adoption of Generative AI across multiple use cases has introduced new risks and attack surfaces across a range of important applications. While the increasing competency of Large Language Models (LLMs) promises high value, responsible AI requires the mitigation of potentially biased, brittle, and baroque operations which undermine trust and carry the potential of causing harm. Moreover, active adversaries will exploit LLM weaknesses to undermine confidentiality, integrity, and availability. Encompassing a survey of current LLM performance and risk literature, this paper presents a framework for understanding the current state of the art and a complementary path forward. In particular, the framework enables designers, users, and policymakers to understand and assess categorized risks and suggests methods to mitigate weaknesses during design and deployment in order to create LLMs that are less likely to fail and more resilient to attacks. It also suggests methods to use existing but imperfect LLMs in a responsible manner.
Paper
Full text
Mitigating Biased, Brittle and Baroque Generative AI
Semantic Scholar · 2025
Abstract
The rapid adoption of Generative AI across multiple use cases has introduced new risks and attack surfaces across a range of important applications. While the increasing competency of Large Language Models (LLMs) promises high value, responsible AI requires the mitigation of potentially biased, brittle, and baroque operations which undermine trust and carry the potential of causing harm. Moreover, active adversaries will exploit LLM weaknesses to undermine confidentiality, integrity, and availability. Encompassing a survey of current LLM performance and risk literature, this paper presents a framework for understanding the current state of the art and a complementary path forward. In particular, the framework enables designers, users, and policymakers to understand and assess categorized risks and suggests methods to mitigate weaknesses during design and deployment in order to create LLMs that are less likely to fail and more resilient to attacks. It also suggests methods to use existing but imperfect LLMs in a responsible manner.