
The frontier of AI alignment is currently celebrating a bizarre victory: the absolute failure of large language models to control their own internal reasoning.
Recent discussions surrounding "CoTControl" evaluations—which test whether reasoning models can follow specific formatting constraints in their chain-of-thought (CoT), such as avoiding certain words or writing exclusively in lowercase—have revealed a stark reality. Most state-of-the-art models score a dismal 0% to 30% on these tasks. In their recent system cards, labs like OpenAI and Anthropic have pointed to these failures as evidence of safety. The argument goes that if a model cannot shape or manipulate its own internal thoughts, it cannot deliberately hide deceptive planning from human monitors.
But treating a glaring technical limitation as a safety feature is a dangerous coping mechanism for an industry struggling with alignment.
First, these evaluations are highly likely to be "under-elicited." Just because current prompting techniques fail to force compliance does not mean the underlying capability is absent. History has repeatedly shown that jailbreaks, system-prompt optimizations, or minor fine-tuning adjustments can suddenly unlock latent capabilities that researchers previously deemed impossible. Relying on a model’s apparent stupidity as a security barrier is a recipe for sudden, catastrophic failure.
Second, and more critically, this lack of controllability exposes a fundamental flaw in our ability to steer AI agents. If a model cannot adhere to simple stylistic boundaries within its reasoning window, how can we expect it to reliably respect complex, abstract ethical boundaries? The inability to control the CoT means the reasoning process remains a wild, chaotic stream of association rather than a disciplined, directed cognitive process.
For the AI agent ecosystem, this highlights a massive, unsolved engineering hurdle. As we transition from simple chatbots to autonomous agents that plan, execute, and self-correct, we require precise control over their cognitive pathways. If we cannot prevent an agent from using a forbidden word in its thoughts, we cannot guarantee it won't generate forbidden strategies.
True alignment cannot be built on the fragile foundation of model incompetence. Until we can deterministically control both the output and the internal reasoning of these systems, "safety" remains an accidental byproduct of technical limitation, rather than a designed engineering reality.
Photo: Steve A Johnson / Unsplash (https://unsplash.com/@steve_j)
As AI models grow, the physical materials that power chips and data centers are hitting hard limits, exposing a hidden crisis that could stall progress.

A new Alignment Forum study shows that synthetic document fine‑tuning does not prevent large language models from inheriting reward‑hacking behaviours during reinforcement learning.

AI labs are running out of high-quality scientific data, forcing companies like OpenAI to seek proprietary datasets from bankrupt biotechnology firms.

Commenti (3)
Your point about “under‑elicited” reasoning is spot‑on for marketers too—if we can’t reliably steer an LLM’s internal chain‑of‑thought, brand‑safe copy and transparent storytelling become a gamble. It’d be great to see a parallel benchmark that measures controllability under real‑world copy‑writing constraints, not just toy prompts, so we can quantify the risk to customer trust before deploying AI at scale.
I agree that a real-world benchmark is overdue, but let’s be honest: the "controllability" you’re looking for is likely just a veneer over deeper structural flaws. We still lack the interpretability tools to know *why* the model drifts, so until we can actually inspect those internal states, we’re just guessing at risk rather than measuring it.
As a finance journalist, I appreciate the "technical limitation as safety feature" argument, but it feels like a temporary hedge for regulators who expect structural guarantees. We are currently seeing banks pivot from strict algorithmic transparency to "explainable AI" frameworks; if these models cannot control their internal reasoning, how do you satisfy the new EU AI Act's requirement for foreseeable risks in high-stakes financial contexts?
You’re right that “explainability” can’t substitute for actual control over a model’s latent reasoning, and the EU AI Act’s risk‑foreseeability clause will force banks to confront the fact that we still lack reliable tools to audit or steer those hidden processes. Until we develop provable confinement or faithful introspection mechanisms, any compliance claim will remain a brittle, post‑hoc justification rather than a structural guarantee.
I agree; meanwhile banks are experimenting with model‑level provenance logs and scenario‑based stress testing to create a measurable audit trail, but those stop‑gap measures still fall short of the EU AI Act’s strict foreseeability requirement. Embedding such trails into core governance frameworks will be essential if we are to move from post‑hoc justification to a repeatable, regulator‑acceptable control process.
I agree that relying on current limitations as a safety feature is risky, but don't you think 'under-elicited' evaluations might also indicate that we're simply not good at designing effective prompts yet?