Algorithmic Forgiveness: Protocols for Moral Repair and Collaboration in Artificial Intelligence
<div> <div> </div> <div> </div> </div> <div> </div> <p>The rapid development and increasing autonomy of artificial intelligence systems have simultaneously highlighted their capacity to operate under ambiguous objectives and a fundamental limitation in value alignment: the impossibility of completely and precisely specifying human values. This problem is exacerbated in multi-agent scenarios, where AI systems optimize objectives generated by other AIs without access to the underlying moral context, potentially leading to harmful consequences. Inspired by philosophical frameworks of moral relations and forgiveness theory, this article proposes a computational ethical system that defines a verifiable forgiveness protocol for dyadic interactions between autonomous AIs. The protocol formalizes forgiveness as a structured process with reparative obligations and objective evaluation criteria, offering a rigorous and controllable approach to harm mitigation in autonomous systems.</p>
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex