Null-Space Gradient Projection: A Geometric Framework for Provable AI Alignment

The rapid advancement of large language models toward superintelligence exposes fundamental vulnerabilities in prevailing AI alignment techniques. Reinforcement Learning from Human Feedback (RLHF), Direct Preference Optimization (DPO), and Constitutional AI rely on probabilistic suppression of harmful outputs, yet they leave intact the parameter-space directions encoding unsafe capabilities. Consequently, even modest adversarial fine-tuning can reliably restore undesired behaviors, as optimization trajectories remain free to traverse the very directions that alignment failed to block. This work introduces Null-Space Gradient Projection (NSGP), a novel geometric framework for training-time safety. NSGP learns a harm direction and orthogonally projects every gradient update onto its null space, rendering the unsafe subspace mathematically inaccessible to learning dynamics. Training is formalized as gradient descent on a degenerate Riemannian manifold We prove five rigorous guarantees: 1. Absolute Orthogonality All updates remain exactly orthogonal to v^harm\hat{v}_{\text{harm}}v^harm. 2. Algebraic Robustness The projection is idempotent, self-adjoint, and non-expansive. 3. Optimal Convergence The standard rate is achieved directly on the safety hyperplane without penalty trade-offs. 4. Manifold Invariance All iterates remain confined to the safe subspace. 5. Topological Preservation (No-Unbinding Theorem) Safety bundle invariants are maintained along continuous trajectories. NSGP is implemented as an in-place kernel with zero additional memory overhead, bypassing inference bottlenecks. Reproducible experiments on proxy models (≤ 16 384 parameters) confirm exact orthogonality to machine epsilon (< 10⁻⁷). A concrete roadmap is provided for scaling to 70B+ production LLMs. By replacing probabilistic suppression with geometric enforcement, NSGP provides formal safety certificates where current methods offer only statistical guarantees—laying the foundation for a principled, scalable science of AI alignment. Code is available upon request for institutional proof-of-concept integrations

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC