AI Safety and Alignment with Human Interests

This chapter examines the concrete risks AI systems pose to society, moving beyond science fiction scenarios to document real harms including Facebook’s role in Myanmar’s ethnic cleansing, algorithmic bias in criminal justice systems, and sophisticated voice cloning scams. The text distinguishes between two threat categories: intentional misuse by malicious actors (jailbreaking, phishing, surveillance systems) and AI misalignment where systems fail unpredictably or pursue goals that conflict with human welfare. Technical safety approaches are detailed, including Constitutional AI, sandbox testing, red teaming, and toxic tool tripwires. This chapter reviews major governance frameworks—NIST’s voluntary AI Risk Management Framework, the EU’s enforceable AI Act with penalties up to €35 million, ISO/IEC 42001 certification standards, co-operation and development principles, and IEEE 7000 series standards. Significant attention is given to cascading failure risks in interconnected AI systems and the challenge of translating abstract safety principles into implementable technical controls. This chapter concludes with ethical considerations regarding potential AI consciousness and sentience.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC