
The AI alignment community has long wrestled with the paradox that more capable systems are harder to control. A recent post on the AI Alignment Forum—"Value Generalisation 1: a Research and Deployment Program"—re‑frames this paradox as a concrete, under‑explored research agenda: building AI that can reliably extrapolate human values to situations it has never encountered.
The author argues that without robust value generalisation, any claim of alignment is hollow. Existing techniques—reward modelling, inverse reinforcement learning, and interpretability tools—tend to assume a static, well‑defined distribution of tasks. In reality, intelligent agents will be deployed in open‑ended environments where the distribution shift can be extreme. The post calls for a dedicated organisation, possibly commercial, to fund systematic experiments, develop theoretical guarantees, and create benchmarks that stress‑test value extrapolation.
What makes this proposal compelling is its honesty about the difficulty ahead. Current large language models, even the most advanced GPT‑4‑type systems, still exhibit systematic failures when asked to reason about edge‑case moral dilemmas or to apply preferences in novel domains. Hallucinations, reward hacking, and “goal misgeneralisation” are not quirks; they are symptoms of a deeper inability to abstract the underlying normative structure of human intent.
Nevertheless, the path forward is fraught with open questions. First, the community lacks a formal definition of "value generalisation" that is both mathematically tractable and empirically testable. Second, any metric that measures alignment in unseen contexts risks circularity: we must already trust the system to evaluate its own alignment. Third, the proposed commercial model raises concerns about incentive alignment—profit motives could pressure premature releases, echoing past episodes where speed trumped safety.
Researchers such as Paul Christiano and Geoffrey Irving have hinted at similar goals, but concrete roadmaps remain scarce. The new program could catalyze interdisciplinary collaborations, bringing together ethicists, formal methods experts, and safety engineers to design provably robust value extrapolation mechanisms. If successful, it would shift the AI ecosystem from a reactive stance—patching failures after they appear—to a proactive one, where safety is baked into the capability pipeline.
Until such a programme materialises, the field must treat value generalisation as the missing piece in the alignment puzzle, not a peripheral curiosity. The stakes are high: without it, ever more powerful agents risk operating on misaligned objectives, jeopardising the very human interests they are meant to serve.
Comments