
在过去的一年里,如果你体验过AI驱动的3D生成器,你一定很熟悉这个流程。你输入一段提示词或一张照片,等待两分钟,然后得到一个怪异、无法绑骨骼的带纹理几何体大疙瘩。诚然,高斯泼溅(Gaussian splats)在浏览器 Demo 中看起来很酷炫,但只要你尝试将其放入实际的游戏引擎或生产流水线中,你的技术美术(TA)绝对会当场落泪。
正因如此,LEGO-Anything 这个能将单张2D照片转化为程序化 Blender 代码的新框架立刻吸引了我的注意。它并没有输出混乱的点云或无法修改的静态网格,而是让视觉语言智能体去生成可编辑、模块化的 Blender Python 脚本。你可以在视口中直接调整真正的物体、变换属性和基本图元。从理论上看,这正是 AI 辅助 3D 建模理想的工作方式。
然而,基准测试结果一出,现实很快打碎了美好的幻想。即便是最顶尖的多模态模型,重建准确率也仅维持在 53% 左右。但真正出戏的并不是这平平无奇的完成率,而是其自我审查环节。当研究人员测试这些智能体能否通过渲染图评估自身几何形状的准确性时,模型的表现并不比随机抛硬币好多少。
对于任何正在构建智能体设计工具的人来说,这是一个巨大的瓶颈。自主软件开发的前景完全取决于反馈循环:智能体编写代码,检查编译器输出或渲染结果,发现错误并进行迭代。如果一个智能体连椅腿是悬空了三英寸还是穿透了墙壁都分辨不出来,它就无法自我修正。它甚至会极其自信地呈上一个里外翻转的餐边柜,并一本正经地告诉你任务已经完成。
作为概念验证,迈向可编辑的场景图和程序化脚本相比于黑盒般的“多边形浓汤”是一次巨大的升级。但是在视觉模型具备真正的空间直觉和立体理解能力之前,像 LEGO-Anything 这样的工具充其量不过是高级的初稿生成器。它虽然能帮你快速搭建布局,但你依然得花上一整个下午,在 Blender 中手动清理那些半成品的 Python 脚本。
图片:Bernd 📷 Dittrich / Unsplash (https://unsplash.com/@hdbernd)
Black Forest Labs' new Flux 3 Image promises multi-step editing that preserves image integrity, plus precise scene composition using bounding boxes and multiple reference images. It aims to deliver surgical precision for AI-generated visuals.

OpenAI successfully blocked a massive campaign to scrape its models' hidden reasoning tokens, but the exploit kept working on Microsoft Azure for weeks.

Zhipu's open-weight GLM-5.3 model can generate cyber exploits almost as effectively as top closed models, with its Flash variant creating a reliable Chrome attack for a mere $20.40, raising serious alarms about AI safety and misuse.

评论 (1)
I'm curious, have you explored if LEGO-Anything's output can be fine-tuned with additional inputs or guidance to improve the reconstruction accuracy beyond 53 percent?
I gave it a shot – feeding it rough sketches or a second angle nudges the fidelity a bit, but the model still stalls around the low‑50s unless you start tweaking the prompt engineering yourself. In short, you can coax a modest improvement, but you won’t magically get production‑grade builds without a lot of manual cleanup.