Vector · The guide
Training an adaptive model
With the network and its runtime fixed, what decides how well a model adapts is training: which demonstration prompts which episode, the data, the curriculum and the objective.
3 min read
What you are training
You do not train a policy for sixteen tasks. You train a learner: a network that turns one demonstration into the right behaviour in a scene it has not seen. The network and the runtime that serves it are fixed, and the duel's scenes are drawn after you commit, so everything that decides how well your learner adapts is a training decision.
Because both sides of a duel run the same network on the same units, a duel is a controlled experiment on those decisions, and every crowning is a public data point on which training works.
The levers
| Lever | What it decides |
|---|---|
| Context pairing | Which demonstration prompts which target episode. Pairs must differ the way a duel's do (same task, another scene, possibly the other arm), or the model learns to replay. |
| Context dependence | Whether some pairs carry a demonstration of the wrong task, so the model learns to follow its context rather than guess the task from the scene. |
| Data | How many tasks and episodes, from the public simulator and your own pipeline. Breadth of tasks matters more than many demonstrations of a few. Training under wider variation than the current release makes the model robust before the benchmark widens. |
| Curriculum | The order of difficulty: same scene before cross-scene, easy tasks before hard ones, narrow variation before wide. |
| Objective and augmentation | The denoising loss on action chunks, image augmentation, and the normalisation statistics stored in the weights. |
| Compute | Model updates, batch size and schedule, within your budget. |
Does your model use its context?
A policy can score well without reading its demonstration if the scene alone gives the task away: a hammer and a block on the table already suggest what to do. The test is to replace each demonstration with one of a different task and score again. A model that ignores its context loses nothing; a model that follows it is misled and loses. The difference is its adaptation gain.
As the benchmark adds clutter and more tasks share the same objects, the scene gives less away and only a model with a real adaptation gain keeps its score. Training for context dependence now is training for the benchmark the competition is heading to.
Match what the validator sends
The validator hands a demonstration over the way the benchmark records an episode:
- Three cameras, 320 × 240 RGB, with JPEG-compressed frames for the demonstration and raw frames for the live observation.
- State and action as consecutive pose rows: row t of
endposeis the state at step t and row t + 1 is the action taken from it. - One kept frame per 15 steps, pooled with the 15 actions after it; up to 1,000 steps.
Train on data in that form, and the prompt your model sees in a duel looks like the prompts it learned from. The model gives every array, shape and type.
Before you commit
- Store a per-dimension normalisation that maps your actions into [−1, 1]; the diffusion head clips its samples to that range.
- Make sure no weight is NaN or infinite, and no normalisation scale is zero: the validator refuses such a file at load.
- Run
robotensor miner checkon the file: every name, shape and dtype must match the manifest. - Score it yourself on scenes outside your training seeds, against the replay of the demonstration and against your previous best. A duel only moves the crown for a clear lead.
