Figure承诺35亿美元采购算力,机器人测试失败率仍达44%
Figure称,Helix 2.5在30个从未见过的旧金山湾区家庭完成420次“全对才算通过”(all-or-nothing)盲测,通过237次;若不做Index预训练,通过率仅为9% [1][2]。据路透社报道,Figure已承诺向Nscale采购价值35亿美元的算力,接近其成立至今19亿美元累计融资额的两倍 [7][8][9]。
Vincent Jiang · 3 min read
仅靠Index预训练,通过率从9%升至56%
Figure称,Helix 2.5走进30个从未见过的旧金山湾区家庭,凭同一个未经修改的模型检查点(checkpoint)执行三项家务,在420次盲测中通过237次 123。评分不设部分得分:篮子里的13到15件玩具必须一件不落,每条毛巾都要叠好,被子要铺得平整 1。用完全相同的数据从零训练的孪生策略通过率仅9%,经Index预训练的孪生策略则通过56%;Figure称,唯一的实验变量就是在自家人类视频数据集上做了预训练 134。以上所有数字均出自Figure自己的发布,目前没有任何外部评估方复测过 3。
巨资所押:训练前的损失预测精确到小数点后四位
演示背后藏着更尖锐的主张:Figure用最多相差8倍的Index数据量训练了四个模型,并称在规模最大的那次训练开始之前,就将其动作预测损失预测到了小数点后四位——Figure称之为首条在人形机器人上实测得到的“人类到人形机器人”迁移规模定律 14。这条曲线若成立,下一轮人类视频数据翻倍的回报,就能在为训练掏钱之前先算出价钱 2。与特斯拉Optimus的较量也被放进同一个叙事:一场数据护城河之争,特斯拉拼的是车队行驶里程,Figure拼的是付费采集的视频 5。Index每秒新增约35分钟的人类经验视频,其公开计数器9月24日显示上传量已达24593204条,而该应用8月25日上线时已有超过1600万条 1468。
另外44%呢?
按同一套“全对才算通过”的标准,这台机器人有44%的测试没能过关:整理玩具通过率40%,叠毛巾62%,铺床67% 23。Sunday Robotics首席执行官Tony Zhao数小时内便发声回应:“一半时间都在失败,这算不上‘做真正有用的工作’” 2。他抛出的对比数字是785次折叠成功778次、成功率99.1%,但那是按单件衣物计数,并非整项家务,两把尺子没法直接对比 2。这场评测的另一批“买单者”,是出借自家房间的30户家庭:短租房里预先摆好玩具,杂物本就清理一空,没有宠物,也没有成堆的待洗衣物——44%的失败率正是在这样的条件下评出来的 2。
整理玩具垫底,五次测试仅过两次
Data
| Value | |
|---|---|
| 整理玩具 | 40% |
| 叠毛巾 | 62% |
| 铺床 | 67% |
证据未到,账单先行
Figure成立至今累计融资19亿美元,其中去年9月一轮就融了逾10亿美元,当时估值390亿美元,但公司从未披露过营收 89。9月3日,公司承诺斥资35亿美元采购Nscale算力,并有意把算力投入推高到60亿美元以上,多达10万块英伟达GPU最早也要到2027年下半年才能到位;Nscale还在同一笔交易中入股Figure 789。
Figure的算力承诺接近其历史融资总额的两倍
- Estimate
Data
| Value | |
|---|---|
| 累计股权融资 | $1.9B |
| Nscale算力承诺 (estimate) | $3.5B |
| 意向扩张规模 | $6B |
在第一块GPU到位之前,有两件事值得盯住:一是Index数据再翻一倍,能否把56%这个数字往上推;二是Figure、Sunday和1X(这场基准之争中的第三个人形机器人阵营)能否达成一致,共同运行一套中立的家庭任务基准测试 2。Figure自己也承认,这并不意味着人形机器人难题已经解决 1。这场押注赌的是:成败参半背后,藏着一条规模定律。
How this brief was made
01Gathered & sourced396 channels · 1,625 articles▾
Agents swept 396 channels and ingested 1,625 articles, then de-duplicated and ranked them for signal.
02Verified & cross-validated9 claims · 33 data feeds▾
Every one of 9 load-bearing claims was checked against primary sources, with 33 live data feeds reconciling the figures and charts.
- 1Figure, "Helix 2.5: Zero-Shot 30-Home Generalization", 17 September 2026
- 2Forbes, John Koetsier, "Figure's Billion-Dollar Physical AI Bet Delivered A 6X Jump In Robot Chore Success", 21 September 2026
- 3TechRepublic, Aminu Abdullahi, "Figure's Robot Entered 30 Unseen Homes and Succeeded 56% of the Time", 21 September 2026
- 4Unite.AI, Orion Sato, "Figure Introduces Helix 2.5, Tested Zero-Shot in 30 Unseen Homes", 17 September 2026
- 5Benzinga (via MSN), Surbhi Jain, "Tesla's Optimus has a new problem: Figure's humanoid is getting better at the unknown", 18 September 2026
- 6Figure, Index app page, live contributor counters, fetched 24 September 2026
- 7Reuters, "AI cloud firm Nscale commits compute worth $3.5 billion for Figure's robotics ambitions", 3 September 2026
- 8Forbes, John Koetsier, "How Figure Committed $3.5 Billion For AI Compute After Raising Only $1.9 Billion", 4 September 2026
- 9eWeek, Matt Gonzales, "Figure AI Commits $3.5B to Nscale Compute for Its Humanoid Robot Push", 4 September 2026
03Reviewed & edited2 human editors▾
2 editors read the draft against the evidence, tuned the framing, and signed off before it shipped.
Become a contributor
Reporting on the business of AI and want it read? We take pitches from outside contributors who bring primary sources and a number worth arguing about.
Deepdive
AI-generated from this story and its cited sources. Not investment advice.



