Does teamwork fall out of a fight?打架打多了,会自己长出配合吗?
The question was easy to state. Put a red team and a blue team of physics-simulated humanoids in an arena. Let every fighter decide for itself whom to attack, from which side, with what. Train both sides against each other for a long time. Then look at whether anything appears that nobody wrote down — a formation, a pincer, focus fire, a fighter that hangs back while another draws the enemy out.
The famous results in this genre all say yes. DeepMind’s Capture the Flag agents learned to camp the base and to follow teammates; OpenAI’s hide-and-seek agents built forts and ramps; DeepMind’s 2 v 2 humanoid football learned passing. All three used either an abstract game world or a compute budget that does not fit under a desk. On the other side, the physics-based fighting papers — Won et al. 2021, NCP, SMPLOlympics — are all one against one. Nobody I could find had pushed a physics fight to many against many and asked the emergence question.
I wanted the physics layer specifically. A flank in a grid world is a coordinate. A flank with a body means walking around a person who is swinging a sword at you, on legs that can be knocked over. If tactics appear there, they mean something.
问题本身很好说。把红蓝两队物理仿真的人形小人放进一个场地,让每个人自己决定打谁、从哪边打、用什么动作。让两边互相打很久。然后看有没有长出没人写过的东西——阵型、包夹、集火、一个人拖住敌人另一个人绕后。
这个方向上的著名结果都说会。DeepMind 的夺旗智能体学会了蹲家和跟随队友;OpenAI 的捉迷藏智能体学会了搭堡垒和推坡道;DeepMind 的 2 v 2 人形足球学会了传球。这三个要么用的是抽象的游戏世界,要么用的算力放不进一张书桌底下。另一边,物理格斗的论文——Won 等 2021、NCP、SMPLOlympics——全是一对一。我没有找到任何人把物理格斗推到多对多,然后问涌现这个问题。
我特意想要物理层。格子世界里的包抄是一个坐标,有身体的包抄意味着绕着一个正在朝你挥剑的人走过去,而你的腿是会被撞倒的。战术如果出现在这里,才说明点什么。
Three layers on one GPU一块 GPU 上的三层
The machine is a DGX Spark: a GB10, aarch64, 128 GB of unified memory, CUDA 13. That rules out most of the usual toolkit. Isaac Gym ships x86 binaries only. Isaac Lab’s PhysX backend does not support this GPU yet. What does work is MuJoCo Warp, which has a CUDA 13 aarch64 wheel, on top of NVIDIA’s NGC PyTorch container, the only CUDA 13 PyTorch on ARM. Everything ran inside Docker; the host has nothing installed. The one thing worth persisting is Warp’s kernel cache, because JIT for sm_121 is slow the first time.
The architecture is the standard one from the 1 v 1 papers, extended to N fighters per world:
机器是 DGX Spark:GB10,aarch64,128 GB 统一内存,CUDA 13。这把常见工具排除了一大半。Isaac Gym 只有 x86 二进制。Isaac Lab 的 PhysX 后端还不支持这块 GPU。能用的是 MuJoCo Warp,它有 CUDA 13 的 aarch64 wheel,跑在 NVIDIA 的 NGC PyTorch 容器上——ARM 上唯一的 CUDA 13 PyTorch。所有东西都在 Docker 里,宿主机什么都不装。唯一值得持久化的是 Warp 的内核缓存,因为 sm_121 第一次 JIT 很慢。
架构就是 1 v 1 论文里那套标准的三层,扩展到每个 world 里 N 个人:
Tactical policy one shared network, one copy per fighter, trained by self-play
in: own body state + every other fighter's relative state + health
out: a 64-d skill latent (roughly: whom, from where, which move)
every 5 control steps = 6 Hz
|
Motion prior NVIDIA's pretrained ASE controller, frozen
in: body state + latent out: PD joint targets, 30 Hz
|
Physics MuJoCo Warp, 1/240 s, N humanoids per world, ~1000 worlds
hits = sword-on-body contact force, measured every substep
| Scene | Humanoids / world | Worlds | env-steps / s | humanoid-steps / s |
|---|---|---|---|---|
| single | 1 | 4096 | 1,685,000 | 1,685,000 |
| 1 v 1 | 2 | 2048 | 653,000 | 1,305,000 |
| 5 v 5 | 10 | 256 | 25,600 | 256,000 |
| Ten humanoids in one world is 65× slower per world than one: the constraint solve is dense and cubic in the number of degrees of freedom. | ||||
With the RL loop on top, the rates that matter are high-level decisions per second — each one is 5 control steps and 40 physics steps: about 6,500 for 1 v 1 and 3 v 2, 5,400 for 3 v 3, 1,800 for 5 v 5. That last number is why the 5 v 5 run took eleven and a half hours for 1,500 epochs and is the shortest run in this post by sample count.
加上 RL 循环之后,真正要紧的速率是每秒多少个高层决策——每个决策对应 5 个控制步、40 个物理步:1 v 1 和 3 v 2 大约 6500,3 v 3 大约 5400,5 v 5 大约 1800。最后这个数字就是为什么 5 v 5 跑 1500 轮花了十一个半小时,而按样本数算它却是本文里最短的一次训练。
Do not train a walker from scratch别从零训练走路
Teaching a humanoid to walk, run and swing a sword from motion capture takes days of GPU time before the first fight can happen. My first plan was exactly that — an AMP pipeline with a 15-body humanoid and a handful of clips — and it was working, slowly, when I remembered that NVIDIA published ASE checkpoints for precisely this character: a sword-and-shield humanoid trained on 87 Reallusion clips, with a 64-dimensional skill space. Those were trained in Isaac Gym. The question was whether a controller trained in one physics engine would stand up in another.
The network port was right the first time. The physics port needed three fixes, and every one of them looked like a policy problem until it was not:
- Timestep. ASE’s PD gains go up to 1,000. At 1/120 s the character collapses in under a second and some worlds go NaN. At 1/240 s, with the controller at 30 Hz, it stands.
- Constraint buffers. MuJoCo Warp preallocates constraint rows per world. The default overflowed silently for a character with a sword, dropping floor contacts, so the humanoid fell through the ground at step 21 regardless of what the policy did. It needs about 360 rows per fighter.
- Collision set. Self-collision off, as in Isaac Gym; the shield’s cylinder replaced by an equal-mass box, because the cylinder collider threw warnings.
After that, random skill latents keep the character on its feet for ten seconds 76% of the time, and the pretrained high-level “strike” controller walks up to a target and swings at it, zero-shot. The observation definition was checked against ASE’s own normalisation statistics: mean ≈ 0, std ≈ 1 on reference poses.
从动捕教一个人形走、跑、挥剑,要烧掉好几天 GPU 时间,第一场架才打得起来。我最初的计划正是这样——一条 AMP 管线、一个 15 个身体部件的人形、一小撮片段——而且它确实在慢慢地跑通,直到我想起 NVIDIA 发布过恰好是这个角色的 ASE 检查点:一个剑盾人形,用 87 段 Reallusion 动捕训练,技能空间 64 维。那些是在 Isaac Gym 里训练的。问题是一个在某个物理引擎里训练出来的控制器,换到另一个引擎里能不能站得住。
网络移植一次就对了。物理移植需要三处修正,而每一处在修好之前看起来都像策略的问题:
- 步长。 ASE 的 PD 增益高到 1000。在 1/120 s 下角色不到一秒就瘫倒,有些 world 还会出 NaN。在 1/240 s 下、控制器 30 Hz,它站住了。
- 约束缓冲。 MuJoCo Warp 按 world 预分配约束行数。对一个拿着剑的角色,默认值会静默溢出,丢掉的恰好是地面接触,于是不管策略做什么,人形都会在第 21 步穿过地板。每个人大约需要 360 行。
- 碰撞集合。 关掉自碰撞,和 Isaac Gym 一致;盾的圆柱碰撞体换成等质量的薄盒,因为圆柱碰撞体会报警告。
修完之后,随机技能隐变量能让角色 76% 的概率站满十秒,预训练的高层「strike」控制器会零样本地走向目标并挥剑。观测定义对照 ASE 自己的归一化统计核过:在参考姿态上均值 ≈ 0、标准差 ≈ 1。
Self-play settles into a truce自博弈会自动停战
The first 1 v 1 run had the obvious reward: damage dealt +1, damage taken −1, knock-out +5, falling over −5. By epoch 600 it had converged — to nothing. No falls, no hits, two fighters holding their shields up two metres apart for the entire episode. The hit reward is sparse and needs a forceful slash to trigger; the fall penalty is dense and triggers whenever you try one. The gradient points at standing still.
Shaping made it worse in the usual way. Rewards for approaching, for pointing the sword tip at the enemy’s torso, for sword-tip speed: the policy learned to collect those and damage per episode fell, from 0.066 to 0.026, while the reward went up. A curriculum against a standing dummy with a tripled hit reward produced 0.007 damage after 100 epochs. Random exploration in a 64-dimensional skill space does not find a slash that lands with force.
What worked was borrowing again. NVIDIA’s pretrained strike controller — a high-level policy that aims at the nearest enemy’s torso — lands 0.14 damage per episode zero-shot, clashes shields thirty times, and also falls over 0.9 times per episode. I appended its 15-dimensional target block to the observation, copied its weights into the actor as a warm start, gave the critic twenty epochs to catch up, and turned on self-play with an opponent pool. By epoch 500, 47 minutes in, both fighters were dealing about 0.32 health per episode with hundreds of blocks, and the video shows two humanoids in a continuous close-quarters fight rather than a stand-off.
第一次 1 v 1 用的是最直白的奖励:造成伤害 +1,受到伤害 −1,击倒 +5,自己倒地 −5。到第 600 轮它收敛了——收敛到什么都不做。零倒地、零命中,两个人举着盾隔两米对峙一整个回合。命中奖励是稀疏的,需要一次有力的劈砍才会触发;倒地惩罚是稠密的,你一尝试劈砍它就触发。梯度指向站着不动。
塑形让事情以最常见的方式变得更糟。给接近、剑尖指向对方躯干、剑尖速度加奖励:策略学会了薅这些分,而每回合伤害下降了,从 0.066 掉到 0.026,奖励却在涨。对着一个站桩对手、命中奖励翻三倍的课程,100 轮后伤害是 0.007。在 64 维技能空间里随机探索,找不到一次有力的劈砍。
管用的办法还是借。NVIDIA 预训练的 strike 控制器——一个瞄准最近敌人躯干的高层策略——零样本每回合造成 0.14 伤害、碰盾三十次,同时自己每回合倒 0.9 次。我把它 15 维的目标块接到观测末尾,把它的权重拷进 actor 做热启动,给 critic 二十轮时间追上来,然后打开带对手池的自博弈。到第 500 轮、47 分钟时,双方每回合各造成约 0.32 血、格挡数百次,录像里是两个人形持续贴身缠斗,而不是对峙。
One detail matters for everything downstream. A hit is a real contact: the normal force between the attacker’s sword geometry and the victim’s body geometries, accumulated every physics substep by a small Warp kernel that runs inside the captured CUDA graph. Sword on shield or sword on sword is a block. When I first checked forces at the control rate instead of the substep, I was missing most of them; measured at the substep, the median sword-on-body force is 54 N and the 90th percentile 238 N, so damage ramps from 20 N to full at 120 N, and a block lets 30% through so that a shield wall is not an unbeatable equilibrium.
有一个细节影响后面所有东西。命中是真实的接触:攻击方的剑几何体与受击方身体几何体之间的法向力,由一个跑在 CUDA graph 里的小 Warp 内核在每个物理子步累积。剑碰盾或剑碰剑算格挡。我最初在控制频率而不是子步上查接触力,漏掉了大多数命中;在子步上测,剑碰身体的力中位数是 54 N、第 90 百分位 238 N,所以伤害从 20 N 开始线性上升、到 120 N 满额,格挡透 30% 伤害,这样盾墙不会成为一个打不破的均衡。
Define the metric before you train先定义指标,再开始训练
“It looks like they are cooperating” is not evidence. Before the first team run I fixed a set of per-episode numbers, all computed from the simulator’s own positions and hit matrix:
- Focus fire. The share of a team’s damage in an episode that lands on its single most-hit enemy. 1/(number of enemies) means evenly spread; 1 means everyone on one target.
- Local numerical superiority. For every living fighter that has an enemy within 2 m: living teammates minus living enemies inside that radius, averaged. A lone duel scores −1 (one enemy, no teammate). A pair on a single scores 0 from the pair’s side and −2 from the victim’s. So in a symmetric fight, “everyone duels” sits at −1 and “everyone piles into one cluster” at 0. Rising means fighters are ending up in situations where they have the numbers.
- Attacker superiority. The same quantity, but only at the moments a hit lands. This separates “I gang up” from “I run away when outnumbered”: both raise the first metric, only the first raises this one.
- Nearest teammate distance and team spread, for formation.
- Attack angle. Where the attacker is relative to the victim’s facing: frontal, flank, or from behind.
- Win rate against a frozen early checkpoint of the same run, as the individual-skill control.
Every 50 or 100 epochs each checkpoint plays 256 worlds for two episodes against that frozen opponent. That is the protocol behind every number and curve below.
「看起来他们在配合」不是证据。第一次团队训练之前,我先定了一组按回合算的数字,全部从仿真器自己的位置和命中矩阵里算出来:
- 集火。 一队在一个回合里的总伤害中,落在它打得最多的那一个敌人身上的比例。1/敌人数 表示平均分散,1 表示所有人打同一个目标。
- 局部人数优势。 对每个 2 米内有敌人的存活者:这个半径内的存活队友数减去存活敌人数,再取平均。单挑是 −1(一个敌人,没有队友)。二打一时,两人那边是 0,被打的那边是 −2。所以在对称的对局里,「人人单挑」在 −1,「所有人挤成一团」在 0。它上升,说明大家越来越多地处在自己人多的局面里。
- 出手时的人数优势。 同一个量,但只在命中发生的那一刻统计。它把「我去围攻」和「人少时我跑」分开:两者都会抬高上一个指标,只有前者会抬高这个。
- 最近队友距离和队形离散度,看阵型。
- 攻击角度。 攻击者相对受击者朝向的位置:正面、侧面、还是背后。
- 对同一次训练的冻结早期检查点的胜率,作为个人战力的对照。
每 50 或 100 轮,每个检查点在 256 个 world 里对那个冻结对手打两个回合。下面所有数字和曲线都是这个流程算出来的。
It emerged once只涌现了一次
| Scenario | Epochs | Win rate | Superiority | Focus fire | Teammate dist. |
|---|---|---|---|---|---|
| 2 v 1, mixed spawn | 1,500 | 0.60 → 0.88 | −0.13 → −0.06 | 0.88 → 0.95 | 0.93 → 2.09 m |
| 3 v 3, two lines (5 reward variants) | ~1,500 | rises within each run | −0.20 → −0.18 | flat | 1.45 → 1.90 m |
| 3 v 3, mixed spawn | 1,350 | 0.44 → 0.74 | −0.48 → −0.29 | 0.58 → 0.61 | 2.66 → 2.77 m |
| 3 v 3, set-pooling policy | 500 | 0.47 → 0.56 | −0.24 → −0.24 | 0.57 → 0.59 | 2.76 → 2.77 m |
| 3 v 2, mixed spawn | 600 | 0.76 → 0.80 | −0.13 → −0.10 | 0.68 → 0.76 | 3.19 → 3.56 m |
| 5 v 5, mixed spawn | 1,500 | 0.51 → 0.47 | −0.25 → −0.25 | 0.41 → 0.43 | 2.95 → 3.32 m |
| Focus fire baselines differ by scenario: even spread is 0.33 with three enemies, 0.5 with two, 0.2 with five. Superiority in 2 v 1 and 3 v 2 is for the larger team. | |||||
Two against one: the pincer二打一:包夹
This was meant as a positive control — a scenario where coordination so obviously pays that if it did not appear, the training setup was broken. It appeared. Over 1,500 epochs the pair’s win rate against its own early checkpoint went from 0.60 to 0.88 monotonically. The two fighters started each episode 0.9 m apart and, by the end, stood 2.1 m apart: on opposite sides of the lone enemy, who cannot face both. Focus went from 0.88 to 0.95. Total damage dealt fell, from 0.90 to 0.64 per episode, because the fights ended sooner.
No reward mentions position, spacing, or teammates. The damage reward is shared as a team mean and a wipe pays a bonus; that is all. The pincer is the policy’s own answer to “how do two of us beat one of them.”
这本来是个阳性对照——一个配合明显划算到如果它不出现,就说明训练流程坏了的场景。它出现了。1500 轮里,两人一方对自己早期检查点的胜率从 0.60 单调升到 0.88。两个人开局间距 0.9 米,到最后站到了 2.1 米:分在那个落单敌人的两侧,而他没法同时面对两个人。集火从 0.88 升到 0.95。总伤害反而下降了,从每回合 0.90 降到 0.64,因为架结束得更早。
没有任何奖励提到位置、间距或队友。伤害奖励按全队均分,全歼有额外奖励,仅此而已。包夹是策略自己对「我们两个怎么打赢他们一个」给出的答案。
Three against three: real, slow, and it stops三打三:真实、缓慢、然后停住
The first 3 v 3 runs spawned the teams in two facing lines. Across five reward variants — team damage share, a wipe bonus, halved health, no penalty for taking damage, chip damage through blocks — individual skill improved every time and no cooperation metric moved. Two things did move, in the wrong direction: when I raised lethality so that fights actually ended, focus dropped from 0.63 to 0.58 and hits from behind fell from 10% to 6%. Deadlier fighters became more individual, not less.
Two lines is a bad scenario for emergence: everyone starts paired off with the fighter opposite, and the fastest reward is to fight that one. Mixed spawn — six fighters placed on random ring slots, red and blue interleaved — is where the metric finally moved. Local numerical superiority rose monotonically from −0.48 to −0.29 over 1,350 epochs. But attacker superiority, the version that only counts moments of landing a hit, went from −0.85 to −0.72: most of the gain is fighters avoiding 1-on-2s, and only a little of it is joining a 2-on-1. Then it flattened. Switching the policy network to a permutation-invariant set-pooling architecture over teammates and enemies, warm-started exactly from the plateaued weights, gave nothing more in 500 epochs.
最初的几次 3 v 3 把两队排成面对面的两排出生。五种奖励变体——团队伤害分成、全歼奖励、血量减半、取消受伤惩罚、格挡透伤——每一次个人战力都在涨,而没有一个配合指标在动。有两样东西动了,方向是反的:当我把致死性调高让架真的能打完时,集火从 0.63 掉到 0.58,背击从 10% 掉到 6%。更能打的人变得更各打各的,而不是更配合。
两排对冲对涌现是个坏场景:每个人开局就和正对面那个人配好了对,最快的奖励就是打他。混杂出生——六个人随机放在环上的槽位里,红蓝穿插——才是指标终于动起来的地方。局部人数优势在 1350 轮里从 −0.48 单调升到 −0.29。但出手时的人数优势——只在命中那一刻统计的版本——是从 −0.85 到 −0.72:大部分涨幅来自躲开一对二,只有一小部分来自加入二对一。然后它平了。把策略网络换成对队友和敌人做置换不变集合池化的结构,从平台期的权重精确热启动,500 轮里没有再多出任何东西。
Three against two, and five against five三打二,和五打五
Three against two was the obvious next probe: closer to 2 v 1 in its arithmetic, since one side always has a spare fighter. It behaved like 3 v 3, not like 2 v 1. Focus rose from 0.68 to 0.76 and the larger team’s superiority from −0.13 to −0.10 over 600 epochs; the three learned to win more cheaply, losing 1.58 fighters per episode at the start and 1.18 at the end. But attacker superiority stayed at −0.74, and the video shows three reds each finding their own opponent, with an occasional accidental 2-on-1.
Five against five, 1,500 epochs, 122 million agent-steps: nothing. Not the cooperation metrics — those I expected to be hard — but even the win rate against the epoch-50 checkpoint stayed at 0.5. Ten fighters in one arena learned nothing measurable in eleven and a half hours.
三打二是很自然的下一个探针:它的算术更接近 2 v 1,因为总有一边多出一个人。它的表现像 3 v 3,不像 2 v 1。600 轮里集火从 0.68 升到 0.76,多的那队的人数优势从 −0.13 升到 −0.10;三人一方学会了赢得更省,开局每回合损失 1.58 人,最后是 1.18 人。但出手时的人数优势停在 −0.74,录像里是三个红方各自找对手,偶尔碰巧凑成一次二对一。
五打五,1500 轮,1.22 亿 agent 步:什么都没有。不只是配合指标——那些我本来就预期很难——连对第 50 轮检查点的胜率都停在 0.5。十个人在一个场地里,十一个半小时,没学到任何可测量的东西。
Coordination has to pay, and has to be findable配合得划算,还得找得到
I do not have a controlled explanation, only a consistent one. Five things line up.
The marginal value of helping shrinks with the team. In 2 v 1, moving to the other side of the target converts a fair fight into a sure win, and the pair collects that reward within the same episode. In 3 v 3, a fighter who leaves its duel to help a teammate gives up its own damage reward now, for a shared team reward that is diluted over three agents and 150 steps and that also depends on what the other two do. The signal is there; it is buried.
The prior is a duelist. Every policy in this post is warm-started from NVIDIA’s strike controller, which aims at the nearest enemy. That is what made 1 v 1 work at all, and it means the starting point of every team run is “fight whoever is closest.” Coordination is not something to learn on top of that; it is something that has to be unlearned first.
Exploration happens in skill space, not tactic space. The policy’s action is a 64-dimensional latent with Gaussian noise of std 0.1 to 0.2. That explores around the current swing, not around the current plan. Wider noise, which I tried, mostly produces stumbling.
Symmetric self-play converges to symmetric play. Both teams improve at the same thing — dueling — and the equilibrium stays where it started. The two runs that moved, 2 v 1 and 3 v 2, are the asymmetric ones.
The budget is small. A billion agent-steps across six scenarios on one desktop is a rounding error next to the population-based runs that produced the famous emergence results. I make no claim that this setup cannot get there with more of everything; I claim only that it did not get there with this much, and that the amount it did get is proportional to how obviously cooperation pays.
我没有受控的解释,只有一个自洽的。五件事对得上。
帮忙的边际价值随队伍变大而缩小。 二打一里,绕到目标另一侧能把一场势均力敌的架变成必胜,而且两人在同一个回合里就能收到这笔奖励。三打三里,一个人离开自己的单挑去帮队友,是用眼下确定的个人伤害奖励,去换一份摊在三个人、150 步上、而且还取决于另外两个人怎么做的团队奖励。信号在,只是埋得深。
先验是个单挑手。 本文所有策略都从 NVIDIA 的 strike 控制器热启动,它瞄准的是最近的敌人。这是 1 v 1 能打起来的原因,也意味着每一次团队训练的起点都是「打离我最近的那个」。配合不是在这之上再学一样东西,而是得先把它忘掉。
探索发生在技能空间,不在战术空间。 策略的动作是一个 64 维隐变量,加标准差 0.1 到 0.2 的高斯噪声。这是在当前这一剑的附近探索,不是在当前这个计划的附近探索。更大的噪声我试过,得到的主要是踉跄。
对称的自博弈收敛到对称的打法。 两队在同一件事上变强——单挑——均衡就停在起点。动了的两次训练,2 v 1 和 3 v 2,恰好是不对称的那两个。
预算很小。 一台桌面机器上六个场景共十亿个 agent 步,放在那些产出著名涌现结果的种群式训练旁边,是个舍入误差。我不主张这套设置加上更多的一切也到不了那里;我只主张它在这个量下没到,而它到了的那一点,和配合有多明显地划算成正比。
Three ways forward, none of them free三条路,没有一条免费
- Pay for it explicitly. An assist reward — a teammate within 2 m attacking the same target, or landing a hit within a second of a teammate’s — would move the focus and superiority metrics within hours. It would also make them meaningless as evidence of emergence: you get exactly the tactic you paid for.
- Curriculum from the one place it worked. Take the 2 v 1 pincer policy and transfer it into 3 v 2 and 3 v 3. The set-pooling policy already handles any team sizes through a same-team flag, so the transfer is mechanical. The question becomes whether a learned tactic survives contact with a symmetric fight.
- Change the game so cooperation pays by construction. Capture the Flag did not get teamwork from a deathmatch; it got it from an objective that one agent cannot complete alone. A point to hold, a VIP to protect, a gate that needs two fighters.
I am going to try the second before the first. If the pincer survives the transfer, the emergence claim gets stronger; if it dissolves back into dueling, that is an answer too.
- 明码标价。 加一个助攻奖励——队友在 2 米内攻击同一目标、或在队友命中后一秒内也命中——几小时内集火和人数优势指标就会动。它同时也会让这些指标作为涌现证据失去意义:你得到的恰好是你付了钱的那个战术。
- 从唯一成功的地方做课程。 拿 2 v 1 的包夹策略迁移到 3 v 2 和 3 v 3。集合池化的策略网络已经通过同队标记支持任意队伍规模,迁移本身是机械的。问题变成:一个学会的战术,碰上对称的对局还能不能活下来。
- 改游戏,让配合在构造上就划算。 夺旗的团队协作不是从死亡竞赛里长出来的,是从一个单人完不成的目标里长出来的。一个要守的点、一个要保护的人、一扇需要两个人才能推开的门。
我会先试第二条。如果包夹在迁移中活下来了,涌现这个说法就更硬;如果它退化回单挑,那也是一个答案。
Configuration配置
Numbers that had to be right必须对的数字
physics dt 1/240 s, decimation 8 (controller at 30 Hz)
hl decision every 5 control steps (6 Hz), 64-d latent, tanh
constraints njmax 360 / fighter, nconmax 64 / fighter
hit sword-on-body normal force per substep, 20 N -> full at 120 N
block = sword on shield/sword, 30% chip damage
damage cap 0.6 / control step, health 0.4
reward damage dealt +3 (team mean), KO +5, fall -2,
team wipe +10, timeout draw -3, style (ASE disc.) x0.05,
approach / aim / swing shaping 0.02 each
ppo horizon 32, minibatch 16384, 6 mini-epochs, lr 3e-5,
gamma 0.99, lambda 0.95, clip 0.2, learnable log-std from -1.6
opponent pool snapshot every 100 epochs, 30% of worlds vs a random snapshot
eval every 50-100 epochs, 256 worlds x 2 episodes vs frozen early ckpt
One run, as launched一次训练的启动命令
python scripts/train_fight.py --run curr3v2_a --arch attn \
--init-attn-attn /runs/team3v3_i/model_000500.pt \
--envs 1024 --fighters 5 --teams 0,0,0,1,1 \
--spawn-mode random --spawn-radius 2.0 --arena-radius 6 \
--episode-length 150 --epochs 600 --horizon 32 --minibatch 16384 \
--lr 3e-5 --learn-std --init-log-std -1.6 --critic-warmup 5 --pool-prob 0.3 \
--reward-dealt 3 --reward-taken 0 --shared-dmg --reward-fall 2 --reward-ko 5 \
--reward-team-wipe 10 --reward-draw 3 --hit-force-full 120 --block-frac 0.3 \
--damage-per-hit 0.6 --max-health 0.4 \
--reward-near 0.02 --reward-aim 0.02 --reward-swing 0.02 --style-w 0.05
Everything runs from a docker compose service with the NVIDIA runtime, the repository mounted at /workspace, checkpoints at /runs, and Warp’s kernel cache at /cache/warp. The host machine has nothing installed beyond Docker.
所有东西都通过一个带 NVIDIA runtime 的 docker compose 服务运行,代码挂在 /workspace,检查点挂在 /runs,Warp 的内核缓存挂在 /cache/warp。宿主机除了 Docker 什么都没装。