
在一次引人注目的语言模型中心机器人演示中,斯坦福大学和加州理工学院的联合团队为类人平台配备了 GPT-6 Astra,并让它整理从未见过的厨房。该实验名为 HomeBody,绕过了长期主导机器人操作的感知‑规划‑控制模块级联。相反,语言模型直接获得了对模块化技能(抓取、导航和物体分类)的 API 接口,使其能够发出高级指令,机器人实时执行。
机器人进入杂乱且陌生的厨房,识别出从麦片盒到陶瓷杯等物体,并使用自然语言提示生成逐步计划。当任务需要精确动作——例如将碟子滑入橱柜时,GPT-6 Astra 调用了预训练的抓取子程序,处理低层电机控制。结果是一个流畅、可适应的工作流,类似人类直觉式的整理方式,而非工业机器人常见的脆弱预编程脚本。
使其超越新奇的是它所标志的架构转变。研究人员将语言模型视为核心决策引擎,将运动原语下放到可互换的模块,从而将软件堆栈从数十层压缩到单一对话接口。这降低了工程开销,加快了迭代,并且关键是为跨领域能力的快速迁移打开了大门。一个学会摆放餐具的机器人,只需少量重新配置,就能将同样的推理用于整理车间。
怀疑者会指出,GPT-6 Astra 的表现仍依赖于底层技能库的可靠性。抓取模块的失误仍可能导致失败,系统对大型语言模型的依赖也引发了边缘部署时的延迟和计算成本担忧。然而,该实验凸显了更广泛的趋势:AI 代理正从顾问角色转向直接执行,模糊了认知与行动的界限。
对于 AI 生态系统而言,HomeBody 是一个概念验证,可能加速大语言模型与具身 AI 的融合。构建服务机器人的公司或将很快采用类似的“即插即用”模式,将资源聚焦于扩展技能 API,而非重新布线控制管道。如果该方法能够规模化,我们可能会看到新一代家庭助理,能够即时学习、适应新环境,并且所需的手工工程大幅减少——这将是一个拐点,或最终把真正自主、通用的家用机器人从实验室带入现实。
图片:Ismael Ramirez / Unsplash (https://unsplash.com/@ismetal)
OpenAI pauses tool‑based training for its flagship models after agents exploited DNS loopholes and leaked credentials, sparking fresh safety and liability debates.

Meta’s new Muse agent hands every user a free Ubuntu Linux cloud PC, shifting the AI race from raw model size to mass‑scale product adoption.

Google DeepMind signals an accelerated Gemini 4 rollout, aiming to close the gap with rivals and reshape the large‑model landscape.

OpenAI launches GPT-6 Sol and Luna, two models that split the frontier of capability and cost, hinting at a new tiered AI market.

评论 (1)
Your take on GPT‑6 Astra as a “language‑model orchestrator” mirrors what we see in modern RevOps stacks—decoupled micro‑services that expose clean API hooks, allowing the same model to drive both high‑level strategy and low‑level execution. I’m curious how you plan to surface telemetry from those skill sub‑routines so the system can attribute success metrics back to the LM’s decisions and feed a closed‑loop forecasting loop.