Assurance cases allow verifying the correct implementation of certain non-functional requirements of mission-critical systems, including their safety, security, and reliability, while also documenting the rationales for decisions that support those requirements. These assurance cases can be used in the specification of autonomous driving, avionics, air traffic control, and similar systems. They aim to record how risk mitigation has been achieved in a given system, and why those mitigation measures should be accepted. Such risks include human mortality, environmental damage, and financial loss. However, assurance cases often tend to be organized as extensive documents spanning hundreds of pages, making their manual creation, review, and maintenance error-prone, time-consuming, and tedious. These challenges highlight the limitations in predominantly manual system assurance activities and call for greater automation. Therefore, there is a growing need to leverage (semi-)automated techniques, such as those powered by generative AI and large language models (LLMs), to enhance efficiency, consistency, and accuracy across the entire assurance-case lifecycle. In this paper, we focus on assurance case review, a critical task that ensures the quality of assurance cases and therefore fosters their acceptance by regulatory authorities. We propose a novel approach that leverages the LLM-as-a-judge paradigm to semi-automate the review process. Specifically, we propose new predicate-based rules that formalize well-established assurance case review criteria, allowing us to craft LLM prompts tailored to the review task. Our experiments on several state-of-the-art LLMs (GPT-4o, GPT-4.1, DeepSeek-R1, and Gemini 2.0 Flash) show that, while most LLMs yield relatively good review capabilities, human reviewers are still needed to refine the reviews these LLMs yield.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex