Notes
← All posts
2026-09-12 ·Experiment log·NVIDIA GB10 / DGX Spark·MuJoCo Warp · ASE · PPO

Two fighters learned to flank. Ten fighters learned nothing. 两个人学会了包夹,十个人什么都没学到

Sword-and-shield humanoids in a GPU physics engine, trained by competitive self-play from 1 v 1 up to 5 v 5, to see whether teamwork appears on its own. It did, exactly once, in the one place where nobody should be surprised. 在 DGX Spark 上用竞争式自博弈训练物理仿真的剑盾人形,从 1 打 1 一路推到 5 打 5,看会不会自己长出没人写过的配合。配合只出现了一次,而且出现在最不该让人意外的地方。

2.1m

How far apart the pair ended up. Two fighters against one started the run shoulder to shoulder, 0.9 m apart. By epoch 1,500 they stood 2.1 m apart on opposite sides of their target and won 88% of the time. Nothing in the reward says anything about where to stand. 二打一的两个人最后站开了多远。 开局他们并肩站着,间距 0.9 米。到第 1500 轮,他们分站目标两侧,间距 2.1 米,胜率 88%。奖励里没有任何一项提到该站在哪。

01 — The question问题

Does teamwork fall out of a fight?打架打多了,会自己长出配合吗?

The question was easy to state. Put a red team and a blue team of physics-simulated humanoids in an arena. Let every fighter decide for itself whom to attack, from which side, with what. Train both sides against each other for a long time. Then look at whether anything appears that nobody wrote down — a formation, a pincer, focus fire, a fighter that hangs back while another draws the enemy out.

The famous results in this genre all say yes. DeepMind’s Capture the Flag agents learned to camp the base and to follow teammates; OpenAI’s hide-and-seek agents built forts and ramps; DeepMind’s 2 v 2 humanoid football learned passing. All three used either an abstract game world or a compute budget that does not fit under a desk. On the other side, the physics-based fighting papers — Won et al. 2021, NCP, SMPLOlympics — are all one against one. Nobody I could find had pushed a physics fight to many against many and asked the emergence question.

I wanted the physics layer specifically. A flank in a grid world is a coordinate. A flank with a body means walking around a person who is swinging a sword at you, on legs that can be knocked over. If tactics appear there, they mean something.

问题本身很好说。把红蓝两队物理仿真的人形小人放进一个场地,让每个人自己决定打谁、从哪边打、用什么动作。让两边互相打很久。然后看有没有长出没人写过的东西——阵型、包夹、集火、一个人拖住敌人另一个人绕后。

这个方向上的著名结果都说会。DeepMind 的夺旗智能体学会了蹲家和跟随队友;OpenAI 的捉迷藏智能体学会了搭堡垒和推坡道;DeepMind 的 2 v 2 人形足球学会了传球。这三个要么用的是抽象的游戏世界,要么用的算力放不进一张书桌底下。另一边,物理格斗的论文——Won 等 2021、NCP、SMPLOlympics——全是一对一。我没有找到任何人把物理格斗推到多对多,然后问涌现这个问题。

我特意想要物理层。格子世界里的包抄是一个坐标,有身体的包抄意味着绕着一个正在朝你挥剑的人走过去,而你的腿是会被撞倒的。战术如果出现在这里,才说明点什么。

Four days, six scenarios, roughly a billion agent-steps, one machine. The answer: coordination emerges clearly in the one scenario where it obviously pays, weakly and slowly in two more, and not at all in the biggest one. The rest of this post is how I know that.四天、六个场景、约十亿个 agent 步、一台机器。答案:配合在那个明显划算的场景里清楚地涌现了,在另外两个场景里弱而慢,在最大的那个场景里完全没有。下面是我怎么知道的。
02 — The stack技术栈

Three layers on one GPU一块 GPU 上的三层

The machine is a DGX Spark: a GB10, aarch64, 128 GB of unified memory, CUDA 13. That rules out most of the usual toolkit. Isaac Gym ships x86 binaries only. Isaac Lab’s PhysX backend does not support this GPU yet. What does work is MuJoCo Warp, which has a CUDA 13 aarch64 wheel, on top of NVIDIA’s NGC PyTorch container, the only CUDA 13 PyTorch on ARM. Everything ran inside Docker; the host has nothing installed. The one thing worth persisting is Warp’s kernel cache, because JIT for sm_121 is slow the first time.

The architecture is the standard one from the 1 v 1 papers, extended to N fighters per world:

机器是 DGX Spark:GB10,aarch64,128 GB 统一内存,CUDA 13。这把常见工具排除了一大半。Isaac Gym 只有 x86 二进制。Isaac Lab 的 PhysX 后端还不支持这块 GPU。能用的是 MuJoCo Warp,它有 CUDA 13 的 aarch64 wheel,跑在 NVIDIA 的 NGC PyTorch 容器上——ARM 上唯一的 CUDA 13 PyTorch。所有东西都在 Docker 里,宿主机什么都不装。唯一值得持久化的是 Warp 的内核缓存,因为 sm_121 第一次 JIT 很慢。

架构就是 1 v 1 论文里那套标准的三层,扩展到每个 world 里 N 个人:

Tactical policy      one shared network, one copy per fighter, trained by self-play
                     in:  own body state + every other fighter's relative state + health
                     out: a 64-d skill latent  (roughly: whom, from where, which move)
                     every 5 control steps = 6 Hz
        |
Motion prior         NVIDIA's pretrained ASE controller, frozen
                     in:  body state + latent      out: PD joint targets, 30 Hz
        |
Physics              MuJoCo Warp, 1/240 s, N humanoids per world, ~1000 worlds
                     hits = sword-on-body contact force, measured every substep
Physics throughput · MuJoCo Warp · CUDA graph plain 21-DoF humanoid · 5 ms step · worst-case contacts
SceneHumanoids / worldWorldsenv-steps / shumanoid-steps / s
single140961,685,0001,685,000
1 v 122048653,0001,305,000
5 v 51025625,600256,000
Ten humanoids in one world is 65× slower per world than one: the constraint solve is dense and cubic in the number of degrees of freedom.

With the RL loop on top, the rates that matter are high-level decisions per second — each one is 5 control steps and 40 physics steps: about 6,500 for 1 v 1 and 3 v 2, 5,400 for 3 v 3, 1,800 for 5 v 5. That last number is why the 5 v 5 run took eleven and a half hours for 1,500 epochs and is the shortest run in this post by sample count.

加上 RL 循环之后,真正要紧的速率是每秒多少个高层决策——每个决策对应 5 个控制步、40 个物理步:1 v 1 和 3 v 2 大约 6500,3 v 3 大约 5400,5 v 5 大约 1800。最后这个数字就是为什么 5 v 5 跑 1500 轮花了十一个半小时,而按样本数算它却是本文里最短的一次训练。

03 — Borrowed legs借来的腿

Do not train a walker from scratch别从零训练走路

Teaching a humanoid to walk, run and swing a sword from motion capture takes days of GPU time before the first fight can happen. My first plan was exactly that — an AMP pipeline with a 15-body humanoid and a handful of clips — and it was working, slowly, when I remembered that NVIDIA published ASE checkpoints for precisely this character: a sword-and-shield humanoid trained on 87 Reallusion clips, with a 64-dimensional skill space. Those were trained in Isaac Gym. The question was whether a controller trained in one physics engine would stand up in another.

The network port was right the first time. The physics port needed three fixes, and every one of them looked like a policy problem until it was not:

  • Timestep. ASE’s PD gains go up to 1,000. At 1/120 s the character collapses in under a second and some worlds go NaN. At 1/240 s, with the controller at 30 Hz, it stands.
  • Constraint buffers. MuJoCo Warp preallocates constraint rows per world. The default overflowed silently for a character with a sword, dropping floor contacts, so the humanoid fell through the ground at step 21 regardless of what the policy did. It needs about 360 rows per fighter.
  • Collision set. Self-collision off, as in Isaac Gym; the shield’s cylinder replaced by an equal-mass box, because the cylinder collider threw warnings.

After that, random skill latents keep the character on its feet for ten seconds 76% of the time, and the pretrained high-level “strike” controller walks up to a target and swings at it, zero-shot. The observation definition was checked against ASE’s own normalisation statistics: mean ≈ 0, std ≈ 1 on reference poses.

从动捕教一个人形走、跑、挥剑,要烧掉好几天 GPU 时间,第一场架才打得起来。我最初的计划正是这样——一条 AMP 管线、一个 15 个身体部件的人形、一小撮片段——而且它确实在慢慢地跑通,直到我想起 NVIDIA 发布过恰好是这个角色的 ASE 检查点:一个剑盾人形,用 87 段 Reallusion 动捕训练,技能空间 64 维。那些是在 Isaac Gym 里训练的。问题是一个在某个物理引擎里训练出来的控制器,换到另一个引擎里能不能站得住。

网络移植一次就对了。物理移植需要三处修正,而每一处在修好之前看起来都像策略的问题:

  • 步长。 ASE 的 PD 增益高到 1000。在 1/120 s 下角色不到一秒就瘫倒,有些 world 还会出 NaN。在 1/240 s 下、控制器 30 Hz,它站住了。
  • 约束缓冲。 MuJoCo Warp 按 world 预分配约束行数。对一个拿着剑的角色,默认值会静默溢出,丢掉的恰好是地面接触,于是不管策略做什么,人形都会在第 21 步穿过地板。每个人大约需要 360 行。
  • 碰撞集合。 关掉自碰撞,和 Isaac Gym 一致;盾的圆柱碰撞体换成等质量的薄盒,因为圆柱碰撞体会报警告。

修完之后,随机技能隐变量能让角色 76% 的概率站满十秒,预训练的高层「strike」控制器会零样本地走向目标并挥剑。观测定义对照 ASE 自己的归一化统计核过:在参考姿态上均值 ≈ 0、标准差 ≈ 1。

When everything falls over regardless of the policy, check buffer overflows and the timestep before suspecting the network. I lost most of a day to the floor-contact one because the symptom — "falls at step 21 every time" — looks exactly like a broken controller.当所有人不管策略如何都会倒,先查缓冲溢出和步长,再怀疑网络。地面接触那个问题吃掉了我大半天,因为症状——「每次都在第 21 步倒」——看起来和控制器坏了一模一样。
04 — Two fighters两个人

Self-play settles into a truce自博弈会自动停战

The first 1 v 1 run had the obvious reward: damage dealt +1, damage taken −1, knock-out +5, falling over −5. By epoch 600 it had converged — to nothing. No falls, no hits, two fighters holding their shields up two metres apart for the entire episode. The hit reward is sparse and needs a forceful slash to trigger; the fall penalty is dense and triggers whenever you try one. The gradient points at standing still.

Shaping made it worse in the usual way. Rewards for approaching, for pointing the sword tip at the enemy’s torso, for sword-tip speed: the policy learned to collect those and damage per episode fell, from 0.066 to 0.026, while the reward went up. A curriculum against a standing dummy with a tripled hit reward produced 0.007 damage after 100 epochs. Random exploration in a 64-dimensional skill space does not find a slash that lands with force.

What worked was borrowing again. NVIDIA’s pretrained strike controller — a high-level policy that aims at the nearest enemy’s torso — lands 0.14 damage per episode zero-shot, clashes shields thirty times, and also falls over 0.9 times per episode. I appended its 15-dimensional target block to the observation, copied its weights into the actor as a warm start, gave the critic twenty epochs to catch up, and turned on self-play with an opponent pool. By epoch 500, 47 minutes in, both fighters were dealing about 0.32 health per episode with hundreds of blocks, and the video shows two humanoids in a continuous close-quarters fight rather than a stand-off.

第一次 1 v 1 用的是最直白的奖励:造成伤害 +1,受到伤害 −1,击倒 +5,自己倒地 −5。到第 600 轮它收敛了——收敛到什么都不做。零倒地、零命中,两个人举着盾隔两米对峙一整个回合。命中奖励是稀疏的,需要一次有力的劈砍才会触发;倒地惩罚是稠密的,你一尝试劈砍它就触发。梯度指向站着不动。

塑形让事情以最常见的方式变得更糟。给接近、剑尖指向对方躯干、剑尖速度加奖励:策略学会了薅这些分,而每回合伤害下降了,从 0.066 掉到 0.026,奖励却在涨。对着一个站桩对手、命中奖励翻三倍的课程,100 轮后伤害是 0.007。在 64 维技能空间里随机探索,找不到一次有力的劈砍。

管用的办法还是借。NVIDIA 预训练的 strike 控制器——一个瞄准最近敌人躯干的高层策略——零样本每回合造成 0.14 伤害、碰盾三十次,同时自己每回合倒 0.9 次。我把它 15 维的目标块接到观测末尾,把它的权重拷进 actor 做热启动,给 critic 二十轮时间追上来,然后打开带对手池的自博弈。到第 500 轮、47 分钟时,双方每回合各造成约 0.32 血、格挡数百次,录像里是两个人形持续贴身缠斗,而不是对峙。

1 v 1, epoch 800. Every swing, block and stumble is the physics engine; the policy only picks a skill latent six times a second.1 v 1,第 800 轮。每一次挥剑、格挡、踉跄都是物理引擎算出来的;策略只是每秒六次挑一个技能隐变量。

One detail matters for everything downstream. A hit is a real contact: the normal force between the attacker’s sword geometry and the victim’s body geometries, accumulated every physics substep by a small Warp kernel that runs inside the captured CUDA graph. Sword on shield or sword on sword is a block. When I first checked forces at the control rate instead of the substep, I was missing most of them; measured at the substep, the median sword-on-body force is 54 N and the 90th percentile 238 N, so damage ramps from 20 N to full at 120 N, and a block lets 30% through so that a shield wall is not an unbeatable equilibrium.

有一个细节影响后面所有东西。命中是真实的接触:攻击方的剑几何体与受击方身体几何体之间的法向力,由一个跑在 CUDA graph 里的小 Warp 内核在每个物理子步累积。剑碰盾或剑碰剑算格挡。我最初在控制频率而不是子步上查接触力,漏掉了大多数命中;在子步上测,剑碰身体的力中位数是 54 N、第 90 百分位 238 N,所以伤害从 20 N 开始线性上升、到 120 N 满额,格挡透 30% 伤害,这样盾墙不会成为一个打不破的均衡。

05 — Measuring it怎么量

Define the metric before you train先定义指标,再开始训练

“It looks like they are cooperating” is not evidence. Before the first team run I fixed a set of per-episode numbers, all computed from the simulator’s own positions and hit matrix:

  • Focus fire. The share of a team’s damage in an episode that lands on its single most-hit enemy. 1/(number of enemies) means evenly spread; 1 means everyone on one target.
  • Local numerical superiority. For every living fighter that has an enemy within 2 m: living teammates minus living enemies inside that radius, averaged. A lone duel scores −1 (one enemy, no teammate). A pair on a single scores 0 from the pair’s side and −2 from the victim’s. So in a symmetric fight, “everyone duels” sits at −1 and “everyone piles into one cluster” at 0. Rising means fighters are ending up in situations where they have the numbers.
  • Attacker superiority. The same quantity, but only at the moments a hit lands. This separates “I gang up” from “I run away when outnumbered”: both raise the first metric, only the first raises this one.
  • Nearest teammate distance and team spread, for formation.
  • Attack angle. Where the attacker is relative to the victim’s facing: frontal, flank, or from behind.
  • Win rate against a frozen early checkpoint of the same run, as the individual-skill control.

Every 50 or 100 epochs each checkpoint plays 256 worlds for two episodes against that frozen opponent. That is the protocol behind every number and curve below.

「看起来他们在配合」不是证据。第一次团队训练之前,我先定了一组按回合算的数字,全部从仿真器自己的位置和命中矩阵里算出来:

  • 集火。 一队在一个回合里的总伤害中,落在它打得最多的那一个敌人身上的比例。1/敌人数 表示平均分散,1 表示所有人打同一个目标。
  • 局部人数优势。 对每个 2 米内有敌人的存活者:这个半径内的存活队友数减去存活敌人数,再取平均。单挑是 −1(一个敌人,没有队友)。二打一时,两人那边是 0,被打的那边是 −2。所以在对称的对局里,「人人单挑」在 −1,「所有人挤成一团」在 0。它上升,说明大家越来越多地处在自己人多的局面里。
  • 出手时的人数优势。 同一个量,但只在命中发生的那一刻统计。它把「我去围攻」和「人少时我跑」分开:两者都会抬高上一个指标,只有前者会抬高这个。
  • 最近队友距离和队形离散度,看阵型。
  • 攻击角度。 攻击者相对受击者朝向的位置:正面、侧面、还是背后。
  • 对同一次训练的冻结早期检查点的胜率,作为个人战力的对照。

每 50 或 100 轮,每个检查点在 256 个 world 里对那个冻结对手打两个回合。下面所有数字和曲线都是这个流程算出来的。

06 — Results结果

It emerged once只涌现了一次

Three panels: win rate against a frozen early checkpoint, local numerical superiority, and nearest teammate distance versus training epoch for 2v1, 3v3 mixed spawn, 3v3 two lines, 3v2 and 5v5
Five runs on the same axes. Only 2 v 1 (orange) moves on all three: it wins more, it gets the numbers, and the pair spreads out. 5 v 5 (black) does not move on any.五次训练画在同一组坐标上。只有 2 v 1(橙色)三张图都在动:赢得更多、占到人数优势、两人拉开。5 v 5(黑色)哪张都没动。
Six scenarios · first vs last analysed checkpoint win rate vs frozen early checkpoint · superiority = local numerical superiority · single seed each
ScenarioEpochsWin rateSuperiorityFocus fireTeammate dist.
2 v 1, mixed spawn1,5000.60 → 0.88−0.13 → −0.060.88 → 0.950.93 → 2.09 m
3 v 3, two lines (5 reward variants)~1,500rises within each run−0.20 → −0.18flat1.45 → 1.90 m
3 v 3, mixed spawn1,3500.44 → 0.74−0.48 → −0.290.58 → 0.612.66 → 2.77 m
3 v 3, set-pooling policy5000.47 → 0.56−0.24 → −0.240.57 → 0.592.76 → 2.77 m
3 v 2, mixed spawn6000.76 → 0.80−0.13 → −0.100.68 → 0.763.19 → 3.56 m
5 v 5, mixed spawn1,5000.51 → 0.47−0.25 → −0.250.41 → 0.432.95 → 3.32 m
Focus fire baselines differ by scenario: even spread is 0.33 with three enemies, 0.5 with two, 0.2 with five. Superiority in 2 v 1 and 3 v 2 is for the larger team.

Two against one: the pincer二打一:包夹

This was meant as a positive control — a scenario where coordination so obviously pays that if it did not appear, the training setup was broken. It appeared. Over 1,500 epochs the pair’s win rate against its own early checkpoint went from 0.60 to 0.88 monotonically. The two fighters started each episode 0.9 m apart and, by the end, stood 2.1 m apart: on opposite sides of the lone enemy, who cannot face both. Focus went from 0.88 to 0.95. Total damage dealt fell, from 0.90 to 0.64 per episode, because the fights ended sooner.

No reward mentions position, spacing, or teammates. The damage reward is shared as a team mean and a wipe pays a bonus; that is all. The pincer is the policy’s own answer to “how do two of us beat one of them.”

这本来是个阳性对照——一个配合明显划算到如果它不出现,就说明训练流程坏了的场景。它出现了。1500 轮里,两人一方对自己早期检查点的胜率从 0.60 单调升到 0.88。两个人开局间距 0.9 米,到最后站到了 2.1 米:分在那个落单敌人的两侧,而他没法同时面对两个人。集火从 0.88 升到 0.95。总伤害反而下降了,从每回合 0.90 降到 0.64,因为架结束得更早。

没有任何奖励提到位置、间距或队友。伤害奖励按全队均分,全歼有额外奖励,仅此而已。包夹是策略自己对「我们两个怎么打赢他们一个」给出的答案。

2 v 1 at epoch 1,500. Watch the two reds split around the blue instead of queueing behind each other.2 v 1,第 1500 轮。注意两个红方是分开绕到蓝方两侧,而不是一前一后排队。

Three against three: real, slow, and it stops三打三:真实、缓慢、然后停住

The first 3 v 3 runs spawned the teams in two facing lines. Across five reward variants — team damage share, a wipe bonus, halved health, no penalty for taking damage, chip damage through blocks — individual skill improved every time and no cooperation metric moved. Two things did move, in the wrong direction: when I raised lethality so that fights actually ended, focus dropped from 0.63 to 0.58 and hits from behind fell from 10% to 6%. Deadlier fighters became more individual, not less.

Two lines is a bad scenario for emergence: everyone starts paired off with the fighter opposite, and the fastest reward is to fight that one. Mixed spawn — six fighters placed on random ring slots, red and blue interleaved — is where the metric finally moved. Local numerical superiority rose monotonically from −0.48 to −0.29 over 1,350 epochs. But attacker superiority, the version that only counts moments of landing a hit, went from −0.85 to −0.72: most of the gain is fighters avoiding 1-on-2s, and only a little of it is joining a 2-on-1. Then it flattened. Switching the policy network to a permutation-invariant set-pooling architecture over teammates and enemies, warm-started exactly from the plateaued weights, gave nothing more in 500 epochs.

最初的几次 3 v 3 把两队排成面对面的两排出生。五种奖励变体——团队伤害分成、全歼奖励、血量减半、取消受伤惩罚、格挡透伤——每一次个人战力都在涨,而没有一个配合指标在动。有两样东西动了,方向是反的:当我把致死性调高让架真的能打完时,集火从 0.63 掉到 0.58,背击从 10% 掉到 6%。更能打的人变得更各打各的,而不是更配合。

两排对冲对涌现是个坏场景:每个人开局就和正对面那个人配好了对,最快的奖励就是打他。混杂出生——六个人随机放在环上的槽位里,红蓝穿插——才是指标终于动起来的地方。局部人数优势在 1350 轮里从 −0.48 单调升到 −0.29。但出手时的人数优势——只在命中那一刻统计的版本——是从 −0.85 到 −0.72:大部分涨幅来自躲开一对二,只有一小部分来自加入二对一。然后它平了。把策略网络换成对队友和敌人做置换不变集合池化的结构,从平台期的权重精确热启动,500 轮里没有再多出任何东西。

Three against two, and five against five三打二,和五打五

Three against two was the obvious next probe: closer to 2 v 1 in its arithmetic, since one side always has a spare fighter. It behaved like 3 v 3, not like 2 v 1. Focus rose from 0.68 to 0.76 and the larger team’s superiority from −0.13 to −0.10 over 600 epochs; the three learned to win more cheaply, losing 1.58 fighters per episode at the start and 1.18 at the end. But attacker superiority stayed at −0.74, and the video shows three reds each finding their own opponent, with an occasional accidental 2-on-1.

Five against five, 1,500 epochs, 122 million agent-steps: nothing. Not the cooperation metrics — those I expected to be hard — but even the win rate against the epoch-50 checkpoint stayed at 0.5. Ten fighters in one arena learned nothing measurable in eleven and a half hours.

三打二是很自然的下一个探针:它的算术更接近 2 v 1,因为总有一边多出一个人。它的表现像 3 v 3,不像 2 v 1。600 轮里集火从 0.68 升到 0.76,多的那队的人数优势从 −0.13 升到 −0.10;三人一方学会了赢得更省,开局每回合损失 1.58 人,最后是 1.18 人。但出手时的人数优势停在 −0.74,录像里是三个红方各自找对手,偶尔碰巧凑成一次二对一。

五打五,1500 轮,1.22 亿 agent 步:什么都没有。不只是配合指标——那些我本来就预期很难——连对第 50 轮检查点的胜率都停在 0.5。十个人在一个场地里,十一个半小时,没学到任何可测量的东西。

5 v 5 at epoch 1,500. Ten fighters, ten duels. The team that wins is the one whose duelists happened to win.5 v 5,第 1500 轮。十个人,十场单挑。赢的那队,是碰巧单挑赢得多的那队。
Four stills from the final policies: 1 v 1 close fight, 2 v 1 pincer with the blue fighter between two reds, 3 v 2 with fighters paired off, 5 v 5 spread across the arena
Final policies of four scenarios, same renderer. Top right is the pincer. Bottom right is ten fighters who have each found someone to duel.四个场景的最终策略,同一渲染器。右上是包夹。右下是十个各自找到了单挑对象的人。
07 — Why为什么

Coordination has to pay, and has to be findable配合得划算,还得找得到

I do not have a controlled explanation, only a consistent one. Five things line up.

The marginal value of helping shrinks with the team. In 2 v 1, moving to the other side of the target converts a fair fight into a sure win, and the pair collects that reward within the same episode. In 3 v 3, a fighter who leaves its duel to help a teammate gives up its own damage reward now, for a shared team reward that is diluted over three agents and 150 steps and that also depends on what the other two do. The signal is there; it is buried.

The prior is a duelist. Every policy in this post is warm-started from NVIDIA’s strike controller, which aims at the nearest enemy. That is what made 1 v 1 work at all, and it means the starting point of every team run is “fight whoever is closest.” Coordination is not something to learn on top of that; it is something that has to be unlearned first.

Exploration happens in skill space, not tactic space. The policy’s action is a 64-dimensional latent with Gaussian noise of std 0.1 to 0.2. That explores around the current swing, not around the current plan. Wider noise, which I tried, mostly produces stumbling.

Symmetric self-play converges to symmetric play. Both teams improve at the same thing — dueling — and the equilibrium stays where it started. The two runs that moved, 2 v 1 and 3 v 2, are the asymmetric ones.

The budget is small. A billion agent-steps across six scenarios on one desktop is a rounding error next to the population-based runs that produced the famous emergence results. I make no claim that this setup cannot get there with more of everything; I claim only that it did not get there with this much, and that the amount it did get is proportional to how obviously cooperation pays.

我没有受控的解释,只有一个自洽的。五件事对得上。

帮忙的边际价值随队伍变大而缩小。 二打一里,绕到目标另一侧能把一场势均力敌的架变成必胜,而且两人在同一个回合里就能收到这笔奖励。三打三里,一个人离开自己的单挑去帮队友,是用眼下确定的个人伤害奖励,去换一份摊在三个人、150 步上、而且还取决于另外两个人怎么做的团队奖励。信号在,只是埋得深。

先验是个单挑手。 本文所有策略都从 NVIDIA 的 strike 控制器热启动,它瞄准的是最近的敌人。这是 1 v 1 能打起来的原因,也意味着每一次团队训练的起点都是「打离我最近的那个」。配合不是在这之上再学一样东西,而是得先把它忘掉。

探索发生在技能空间,不在战术空间。 策略的动作是一个 64 维隐变量,加标准差 0.1 到 0.2 的高斯噪声。这是在当前这一剑的附近探索,不是在当前这个计划的附近探索。更大的噪声我试过,得到的主要是踉跄。

对称的自博弈收敛到对称的打法。 两队在同一件事上变强——单挑——均衡就停在起点。动了的两次训练,2 v 1 和 3 v 2,恰好是不对称的那两个。

预算很小。 一台桌面机器上六个场景共十亿个 agent 步,放在那些产出著名涌现结果的种群式训练旁边,是个舍入误差。我不主张这套设置加上更多的一切也到不了那里;我只主张它在这个量下没到,而它到了的那一点,和配合有多明显地划算成正比。

One more honest caveat: the metrics count what I defined. Hits from behind stayed at 2–5% in every scenario, because the strike prior attacks frontally and nothing ever trained it not to. A tactic I did not think to measure could be sitting in those videos.再补一句实话:指标只数我定义了的东西。背击在每个场景里都停在 2% 到 5%,因为 strike 先验是正面进攻的,而从没有什么训练过它不这么做。某种我没想到要量的战术,可能就躺在那些录像里。
08 — Next接下来

Three ways forward, none of them free三条路,没有一条免费

  • Pay for it explicitly. An assist reward — a teammate within 2 m attacking the same target, or landing a hit within a second of a teammate’s — would move the focus and superiority metrics within hours. It would also make them meaningless as evidence of emergence: you get exactly the tactic you paid for.
  • Curriculum from the one place it worked. Take the 2 v 1 pincer policy and transfer it into 3 v 2 and 3 v 3. The set-pooling policy already handles any team sizes through a same-team flag, so the transfer is mechanical. The question becomes whether a learned tactic survives contact with a symmetric fight.
  • Change the game so cooperation pays by construction. Capture the Flag did not get teamwork from a deathmatch; it got it from an objective that one agent cannot complete alone. A point to hold, a VIP to protect, a gate that needs two fighters.

I am going to try the second before the first. If the pincer survives the transfer, the emergence claim gets stronger; if it dissolves back into dueling, that is an answer too.

  • 明码标价。 加一个助攻奖励——队友在 2 米内攻击同一目标、或在队友命中后一秒内也命中——几小时内集火和人数优势指标就会动。它同时也会让这些指标作为涌现证据失去意义:你得到的恰好是你付了钱的那个战术。
  • 从唯一成功的地方做课程。 拿 2 v 1 的包夹策略迁移到 3 v 2 和 3 v 3。集合池化的策略网络已经通过同队标记支持任意队伍规模,迁移本身是机械的。问题变成:一个学会的战术,碰上对称的对局还能不能活下来。
  • 改游戏,让配合在构造上就划算。 夺旗的团队协作不是从死亡竞赛里长出来的,是从一个单人完不成的目标里长出来的。一个要守的点、一个要保护的人、一扇需要两个人才能推开的门。

我会先试第二条。如果包夹在迁移中活下来了,涌现这个说法就更硬;如果它退化回单挑,那也是一个答案。

09 — Reproduce复现

Configuration配置

Machine
NVIDIA GB10 · 128 GB unified
Driver / CUDA
580.95 · 13.0 · sm_121
Container
NGC PyTorch 25.11 (aarch64)
Physics
MuJoCo 3.11 · MuJoCo Warp 3.11 · Warp 1.17 cu13
Motion prior
ASE sword-shield LLC (NVIDIA), frozen
Tactical policy
PPO, one network shared by all fighters

Numbers that had to be right必须对的数字

physics dt        1/240 s, decimation 8  (controller at 30 Hz)
hl decision       every 5 control steps  (6 Hz), 64-d latent, tanh
constraints       njmax 360 / fighter, nconmax 64 / fighter
hit               sword-on-body normal force per substep, 20 N -> full at 120 N
                  block = sword on shield/sword, 30% chip damage
                  damage cap 0.6 / control step, health 0.4
reward            damage dealt +3 (team mean), KO +5, fall -2,
                  team wipe +10, timeout draw -3, style (ASE disc.) x0.05,
                  approach / aim / swing shaping 0.02 each
ppo               horizon 32, minibatch 16384, 6 mini-epochs, lr 3e-5,
                  gamma 0.99, lambda 0.95, clip 0.2, learnable log-std from -1.6
opponent pool     snapshot every 100 epochs, 30% of worlds vs a random snapshot
eval              every 50-100 epochs, 256 worlds x 2 episodes vs frozen early ckpt

One run, as launched一次训练的启动命令

python scripts/train_fight.py --run curr3v2_a --arch attn \
  --init-attn-attn /runs/team3v3_i/model_000500.pt \
  --envs 1024 --fighters 5 --teams 0,0,0,1,1 \
  --spawn-mode random --spawn-radius 2.0 --arena-radius 6 \
  --episode-length 150 --epochs 600 --horizon 32 --minibatch 16384 \
  --lr 3e-5 --learn-std --init-log-std -1.6 --critic-warmup 5 --pool-prob 0.3 \
  --reward-dealt 3 --reward-taken 0 --shared-dmg --reward-fall 2 --reward-ko 5 \
  --reward-team-wipe 10 --reward-draw 3 --hit-force-full 120 --block-frac 0.3 \
  --damage-per-hit 0.6 --max-health 0.4 \
  --reward-near 0.02 --reward-aim 0.02 --reward-swing 0.02 --style-w 0.05

Everything runs from a docker compose service with the NVIDIA runtime, the repository mounted at /workspace, checkpoints at /runs, and Warp’s kernel cache at /cache/warp. The host machine has nothing installed beyond Docker.

所有东西都通过一个带 NVIDIA runtime 的 docker compose 服务运行,代码挂在 /workspace,检查点挂在 /runs,Warp 的内核缓存挂在 /cache/warp。宿主机除了 Docker 什么都没装。