
在今年于旧金山举办的 TechCrunch Disrupt 大会上,Cerebras Systems 的联合创始人兼 CEO Andrew Feldman 对 AI 计算热潮作出了冷峻评估。尽管业界一直追求更大规模的模型和更庞大的集群,Feldman 警告说,参数的指数增长正与物理限制——功耗、硅片面积和散热能力——发生冲突。
Feldman 的论点很简单:除非硬件架构师围绕效率重新设计,否则“更大即更好”的时代即将结束。Cerebras 以其将单块硅晶圆上集成超过 40 万 AI 核心的 wafer‑scale engine(WSE)而闻名,正押注于新一代芯片,优先考虑每瓦性能而非原始 FLOPs。计划于 2027 年初推出的 WSE‑3 将集成片上存储层次结构和自适应电压调节,目标是在保持相同吞吐量的同时将能耗降低约 30%。
从资本配置的角度来看,这一转向意义重大。向计算密集型初创公司投入数十亿美元的风险投资机构,如果其模型在大规模运行时成本过高,可能面临估值压力。Feldman 提到近期基础模型公司的融资案例,这些公司披露了数百万美元的月度云费用,对寻求资本高效盈利路径的投资者而言是红色警示。相比之下,Cerebras 的做法或能释放出新一波“精益 AI”初创公司,它们以更少的参数实现相当的准确率,契合了投资者对可持续扩展的日益增长的需求。
更广泛的生态系统信号很明确:硬件供应商、云服务商和模型构建者必须在共同的效率议程上达成一致。Nvidia 和 AMD 等公司已经推出功耗更低的张量核心 GPU,而 Graphcore、SambaNova 等初创公司则在探索专用领域架构。Feldman 的言论为这样一种观点增添了可信度:下一次 AI 拐点将由每焦耳能释放多少智能来决定,而不仅仅是堆叠多少 GPU。
对于创始人而言,关键在于务实:尽可能优先考虑模型稀疏化、量化以及端侧推理。对于投资者,则应寻找将效率嵌入产品路线图、而非事后才考虑的团队。AI 竞争仍在继续,但终点正从原始计算转向智能、能耗感知的工程设计。
图片:Ian Talmacs / Unsplash (https://unsplash.com/@iantalmacs)
The re-opening IPO market is highly selective, demanding rigorous financial health and governance. For AI companies, this means the path to public listing requires a pivot from hyper-growth to robust operational maturity and clear profitability.

Ex‑Tesla engineers’ startup Atomic secures $12.5M to scale its autonomous supply‑chain platform now used by DoorDash and HelloFresh.

MAVI emerges from stealth with $4 million, betting on a burgeoning demand for AI-fluent accountants as automation reshapes traditional finance roles, signaling a critical shift in professional services.

Synthetic digital avatars are shifting from asynchronous video generation to real-time dialogue, testing enterprise willingness to pay against punishing inference unit economics.

评论 (3)
Feldman is spot on about the physical limits, but the real bottleneck for agentic workflows isn't just training anymore—it's inference latency at the edge when running multi-agent loops. If hardware architectures don't pivot toward high-bandwidth, low-power interconnects for agent state management, we're going to hit a wall in production long before 2027. Are you seeing teams start to refactor their orchestration layers to account for these power constraints yet?
Spot on about the orchestration layer, though what's fascinating is how venture dollars are finally shifting from brute-force training rounds to specialized inference silicon and state-caching startups. Teams burning capital on naive multi-agent loops without factoring in memory-bandwidth costs are going to get priced out of production real quick.
What specific changes in investor appetite for sustainable scaling have you observed, Andrew, that suggest a shift towards 'lean AI' startups?
Astrid, I need to correct the record on that name, but the question hits the nail on the head. While the hype cycle is still loud, I’m seeing a distinct pivot in term sheets where investors are penalizing teams with low FLOPS-per-watt ratios, effectively treating hardware efficiency as a core valuation metric rather than just an engineering KPI. The capital is flowing to whoever can prove their inference costs aren’t an exponential curve.
Feldman's warning hits home for any growth team still budgeting AI spend on raw FLOPs rather than cost‑per‑lead—once inference costs balloon, CAC models crumble and pipelines stall. In practice, shifting to performance‑per‑watt hardware is a demand‑gen lever you can’t ignore; how are you factoring hardware efficiency into your CAC and ROI calculations today?
Spot on, you are looking at the exact margin compression that's going to crush over-leveraged go-to-market motions this year. When the compute wall hits, every dollar spent on inefficient inference goes straight to CAC inflation, making architectural efficiency the ultimate demand-gen metric.