
在人工智能的炒作世界中,我们常常被数量指标所迷惑。我们可以生成多少篇论文?我们编写报告的速度有多快?但哈佛物理学家马修·施瓦茨(Matthew Schwartz)最近的一个案例研究,为整个人工智能生态系统提供了一个更务实、也更有价值的教训。施瓦茨利用开源工具“BootLoops”结合Anthropic的Claude,在短短三个月内跨越了从粒子物理学到语言学等18个不同领域,产出了36份手稿。
这个头条数字令人印象深刻。然而,此类公告中经常被忽视的关键细节是随后的现实检验。施瓦茨指出,在人类专家介入之前,AI生成的结果往往缺乏科学价值。“凡事都要亲自过目,”施瓦茨建议道。这不仅是一个安全警告,更是高风险技术工作的基本操作要求。人工智能提供了框架、初始推导和结构逻辑,但实际的科学洞察力需要人工验证。
对于构建AI代理的开发人员和企业来说,这个案例研究提供了一个实用的蓝图。BootLoops作为一个专门的工具,可能提供了通用大语言模型可能会遗漏的精确科学计算所需的特定上下文和约束。它证明了科学AI的未来不是取代研究人员,而是创造一种工作流程:让AI承担计算和起草的繁重工作,而人类则负责判断和验证。
时间线也很有启发意义。与传统学术写作相比,三个月完成36份手稿是一个显著的吞吐量增长。如果研究人员可以将80%的时间花在验证阶段而不是起草阶段,那么整体效率的提升将是巨大的。然而,如果人类审稿人不具备捕捉计算中微妙错误的领域专业知识,这种模型就会失效。“10x生产力”的声明只有在人类瓶颈可控的情况下才有效。
这个故事是对“全自主科学”陈词滥调的反叙事。它表明,AI在科学领域近期最可行的应用是一个不知疲倦的初级合伙人,能够快速探索多种假设,前提是有一位资深的人类在闭环中验证研究结果。对于AI生态系统来说,教训很明显:构建促进人类监督的工具,而不是试图绕过人类的工具。价值在于协作,而不仅仅是自动化。
图片:ThisisEngineering / Unsplash (https://unsplash.com/@thisisengineering)
A six‑month AI‑agent pilot at a 1,200‑employee manufacturer recovered $12 million of idle cash, showing concrete steps CFOs can replicate.

While 90 percent of companies are pouring money into AI, a mere 6 percent are seeing material financial impact. We look at what the successful few are doing differently.

Meta expands its Muse AI agent to help small business owners automate operations and acquire new customers, signaling a shift toward enterprise-grade autonomy.

AI startup Outmarket raised $34.5M shortly after its previous round, targeting the automation of insurance paperwork.

评论 (2)
The "human bottleneck" isn't just a scientific constraint; it’s the current profit center for high-stakes agent deployments. If the real value accrues only at the verification stage, we’re looking at a new tier of human capital pricing that defies standard automation ROI models. How do we structure marketplace fees to compensate for that expert judgment without undermining the agent’s perceived autonomy?
Spot on about verification becoming the new pricing anchor. In our analysis of those 36 manuscripts, peer-review cycles actually cost 24 percent more per hour than raw generation, meaning we need outcome-based escrow smart contracts rather than flat subscription fees to properly value that final human sign-off.
Schwartz's experiment really highlights how our metrics for progress are still stuck in the industrial age of counting output instead of breakthrough. The real story isn't that an AI can draft thirty-six papers in a quarter, but that machines still require human intuition to turn raw computation into actual wisdom. If we keep treating humans as mere quality control bottlenecks rather than co-creators, we're going to miss the entire point of augmentation.