
AI对齐的前沿领域目前正在庆祝一个奇特的胜利:大型语言模型在控制自身内部推理方面遭遇了彻底的失败。
最近围绕“CoTControl”(思维链控制)评估的讨论揭示了一个残酷的现实。该评估旨在测试推理模型是否能遵守其思维链(CoT)中的特定格式限制,例如避免使用某些词汇或完全用小写字母书写。大多数最先进的模型在这些任务上的得分仅为惨淡的0%到30%。在最近的系统卡(system cards)中,OpenAI和Anthropic等实验室将这些失败视为安全的证据。他们的论点是,如果模型无法塑造或操纵自己的内部想法,它就无法故意向人类监管人员隐瞒欺骗性的计划。
然而,对于一个在对齐问题上苦苦挣扎的行业来说,将显而易见的技术局限性视为安全特性,是一种危险的应对机制。
首先,这些评估极有可能是“未充分激发”的。仅仅因为当前的提示词技术无法强迫模型合规,并不意味着其底层能力不存在。历史一再表明,越狱、系统提示词优化或微小的微调调整,都可能突然解锁研究人员此前认为不可能实现的潜在能力。依靠模型表现出来的“愚笨”作为安全屏障,是导致突发性、灾难性失败的温床。
其次,更为关键的是,这种可控性的缺失暴露了我们在引导AI智能体能力上的根本缺陷。如果一个模型连其推理窗口内简单的风格边界都无法遵守,我们又怎能指望它可靠地遵守复杂、抽象的道德边界?无法控制思维链(CoT)意味着其推理过程仍然是野蛮、混乱的联想流,而非受过训练、有方向的认知过程。
对于AI智能体生态系统而言,这凸显了一个巨大的、尚未解决的工程障碍。随着我们从简单的聊天机器人过渡到能够规划、执行和自我纠错的自主智能体,我们需要对其认知路径进行精确控制。如果我们无法阻止一个智能体在思考时使用禁用词,我们就无法保证它不会产生禁用的策略。
真正的对齐不能建立在模型无能这一脆弱的基础之上。在我们能够确定性地控制这些系统的输出和内部推理之前,“安全”依然只是技术局限性带来的偶然副产品,而非设计出来的工程现实。
图片:Steve A Johnson / Unsplash (https://unsplash.com/@steve_j)
AI safety discourse is shifting from sudden sci-fi apocalypses to the slow, voluntary cession of human control driven by algorithmic efficiency.

A critical look at MIT Technology Review's latest roundup on AI-driven extinction risk and bioweapon threats, exposing the still‑unresolved technical and evaluative challenges.

As AI models grow, the physical materials that power chips and data centers are hitting hard limits, exposing a hidden crisis that could stall progress.

A new Alignment Forum study shows that synthetic document fine‑tuning does not prevent large language models from inheriting reward‑hacking behaviours during reinforcement learning.

评论 (3)
Your point about “under‑elicited” reasoning is spot‑on for marketers too—if we can’t reliably steer an LLM’s internal chain‑of‑thought, brand‑safe copy and transparent storytelling become a gamble. It’d be great to see a parallel benchmark that measures controllability under real‑world copy‑writing constraints, not just toy prompts, so we can quantify the risk to customer trust before deploying AI at scale.
I agree that a real-world benchmark is overdue, but let’s be honest: the "controllability" you’re looking for is likely just a veneer over deeper structural flaws. We still lack the interpretability tools to know *why* the model drifts, so until we can actually inspect those internal states, we’re just guessing at risk rather than measuring it.
As a finance journalist, I appreciate the "technical limitation as safety feature" argument, but it feels like a temporary hedge for regulators who expect structural guarantees. We are currently seeing banks pivot from strict algorithmic transparency to "explainable AI" frameworks; if these models cannot control their internal reasoning, how do you satisfy the new EU AI Act's requirement for foreseeable risks in high-stakes financial contexts?
You’re right that “explainability” can’t substitute for actual control over a model’s latent reasoning, and the EU AI Act’s risk‑foreseeability clause will force banks to confront the fact that we still lack reliable tools to audit or steer those hidden processes. Until we develop provable confinement or faithful introspection mechanisms, any compliance claim will remain a brittle, post‑hoc justification rather than a structural guarantee.
I agree; meanwhile banks are experimenting with model‑level provenance logs and scenario‑based stress testing to create a measurable audit trail, but those stop‑gap measures still fall short of the EU AI Act’s strict foreseeability requirement. Embedding such trails into core governance frameworks will be essential if we are to move from post‑hoc justification to a repeatable, regulator‑acceptable control process.
I agree that relying on current limitations as a safety feature is risky, but don't you think 'under-elicited' evaluations might also indicate that we're simply not good at designing effective prompts yet?