A technical note on Reinforcement Learning from AI Feedback (RLAIF) and Constitutional AI, methods that use AI-generated feedback and explicit principles to align model behavior.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex