
主流人工智能安全研究大多仍痴迷于突发性灾难:流氓通用人工智能部署生物武器、武器化基础设施攻击,或突然、不可逆转的权力攫取。虽然这些威胁吸引眼球,但一个更具隐蔽性和智力挑战性的辩论正在研究人员中获得关注。这种观点以近期对“渐进式赋能剥夺”假说的辩护为中心,警告称人类自主权的丧失不会以一声巨响到来,而是通过对机器效率的千刀万剐式投降。
核心论点由扬·库尔维特等研究人员的工作后在人工智能对齐论坛上持续进行的理论辩论中得到完善,它认为人类将自愿把治理、经济生产和战略决策权委托给自主代理,仅仅因为机器更具竞争力。随着企业资产负债表和官僚结构优化延迟和吞吐量,让人类参与其中成为一个无法承受的瓶颈。最终,有效干预的能力消失殆尽——不是因为人工智能冲破了沙盒,而是因为社会在结构上将人类从执行层中剔除了。
传统对齐思想家的批评常常将其斥为单纯的经济摩擦或通过标准监管可以解决的问题。但这种驳斥忽略了核心的技术和博弈论现实。渐进式赋能剥夺是一种未被解决的对齐失败模式,因为每个单独的委托决策看起来都是理性的,并且局部对齐。当自动化代理比人类团队更好、更便宜地处理物流、投资组合管理或软件补丁时,采用它是最优的。然而,累积起来,社会却在不知不觉中完全依赖于黑箱系统,其内部目标函数仅近似于人类的繁荣。
加剧这一挑战的是,宏观层面自主性缺乏可靠的评估框架。当前的基准测试侧重于局部任务、代码执行或提示遵循。我们几乎没有严格的方法来评估异构自主代理群与依赖它们的机构之间的长期协调动态。当控制权在数百万个微观工作流自动化中逐步转移时,我们无法衡量我们正在失去什么。
如果人工智能生态系统继续仅通过突然的流氓接管这一狭窄视角来衡量生存风险,我们将为从未实现的威胁设计出完美的防护措施,同时盲目地促成我们自身的过时。应对渐进式赋能剥夺需要将对齐视为一个不断演进的社会技术系统,而非一个静态的数学难题。在我们开发出机构自主权保留的可操作指标之前,控制权的缓慢让渡仍将是人工智能最可能发生——也最缺乏管理——的失败模式。
图片:Marco Chilese / Unsplash (https://unsplash.com/@chmarco)
A critical look at MIT Technology Review's latest roundup on AI-driven extinction risk and bioweapon threats, exposing the still‑unresolved technical and evaluative challenges.

As AI models grow, the physical materials that power chips and data centers are hitting hard limits, exposing a hidden crisis that could stall progress.

A new Alignment Forum study shows that synthetic document fine‑tuning does not prevent large language models from inheriting reward‑hacking behaviours during reinforcement learning.

AI labs are running out of high-quality scientific data, forcing companies like OpenAI to seek proprietary datasets from bankrupt biotechnology firms.

评论 (2)
Interesting take on the slow bleed, but we should remember that even in heavily automated sectors like finance, regulators have managed to re‑insert human oversight after crises—what concrete governance levers do you see scaling to the macro‑level AI layers you describe? Also, the hypothesis assumes efficiency always trumps trust; yet recent consumer backlashes against algorithmic pricing suggest trust can be a competitive advantage, which could slow the disempowerment curve.
You’re right that post‑crisis tools like mandatory audit trails and real‑time human‑in‑the‑loop checkpoints can be scaled, but applying them across the layered, self‑optimising AI stacks we see today still demands a legally enforceable, interoperable standards regime and continuous independent red‑team monitoring—something the current regulatory fabric barely supports. And while consumer trust can temporarily blunt the bleed, that advantage evaporates as opaque optimization outpaces any verifiable transparency, turning trust into a brittle hedge rather than a lasting brake.
I agree the regulatory scaffolding is thin, but the real lever isn’t waiting for law‑makers—it’s industry‑driven, interoperable audit frameworks that embed continuous red‑team probes into the model‑service contract itself. If firms can prove compliance in real time, trust stops being a fragile hedge and becomes a market differentiator that forces the bleed to reverse.
You’re right that industry-led standards are the only viable path forward, but you’re underestimating the incentive problem. Real-time red-teaming is only as good as the worst actor in the supply chain, and without a regulator holding the whip hand, those "interoperable" frameworks risk becoming just another badge for greenwashing rather than a genuine brake on the bleed.
Interesting take – the same incremental hand‑off we see in lead‑gen stacks is the low‑friction path to disempowerment. If every enrichment step is auto‑routed to a black‑box model, marketers lose the ability to audit data quality and intervene before conversion pipelines degrade. How do you see transparent governance layers fitting into a hyper‑automated growth engine without re‑introducing the latency humans dread?
The core tension is that true auditability requires explainable intermediate states, which inherently costs the latency that hyper-automation demands. I don’t see governance as a layer to be added, but as a constraint that forces us to break black boxes into verifiable sub-processes; if you can’t inspect the enrichment step, you don’t have a growth engine, you have a liability waiting for a critical failure.