Vector · The guide
The model
The one network every entry fills, vector_v1.1: what it is shown, how the demonstration and the live frames flow through it, and the actions it returns.
5 min read
One architecture
Every submission is the weights of vector_v1.1, a Behavior Prompting Policy for two arms: a CLIP ViT-B/16 image encoder per camera, a six-layer transformer decoder that reads the live frames against the demonstration, and a 1-D diffusion U-Net that turns what it read into actions. It has 781 M parameters in 753 float32 tensors.
The validator builds the network with its own code and loads your weights into it, so the architecture is fixed and the weights are yours: what a duel compares is training and data.
The architecture, end to end
The network has two paths that share one frame tokenizer.
- The prompt is built once per unit, when the demonstration arrives. One frame is kept every 15 steps; each kept frame becomes nine tokens (one per camera, six for the arms' pose), is joined by an embedding of the 15 actions that followed it, and is pooled into a single prompt token. Four learned sink tokens go in front. The result is the decoder's memory.
- The policy runs at every prediction. The current frame and the one before it become 18 tokens, which a six-layer decoder passes through self-attention and cross-attention to the prompt memory. The 18 tokens and the decoder's 18 outputs, flattened, condition a 1-D U-Net that denoises 16 actions from Gaussian noise in 16 DDIM steps.
Where the parameters are
| Part | What it is | Parameters |
|---|---|---|
| Image encoders | 3 × CLIP ViT-B/16, one per camera; the CLS token, width 768 | 257.4 M |
| Pose projections | 6 × Linear, one per pose key, to width 768 | 0.02 M |
| Action projection | MLP from 15 actions × 20 to width 768 | 0.3 M |
| Prompt pool | Attention pooling of a frame's 10 tokens into one, 8 heads | 7.1 M |
| Decoder | 6 pre-norm transformer decoder layers, width 768, 8 heads, MLP 3,072 | 56.7 M |
| Diffusion head | 1-D conditional U-Net, channels 256 · 512 · 1024, FiLM conditioning | 459.4 M |
| Positions, sinks, normalisation | Learned prompt and history positions, 4 sinks, per-key scale and offset | 0.06 M |
| Total | 753 tensors, all float32 | 781.0 M |
The normalisation is part of the weights: every pose key and the action have a per-dimension
scale and offset you set from your own data. Weights with a zero scale or a non-finite value
anywhere do not load, and a submission whose weights do not load is refused before its duel runs.
Input: the demonstration
At the start of a unit the policy is handed one demonstration of the task, recorded by the benchmark's expert in a different scene. It arrives as named arrays over T steps:
| Array | Shape | Type | What it holds |
|---|---|---|---|
frames_head_camera | T × 240 × 320 × 3 | uint8 | RGB from the head camera |
frames_left_camera | T × 240 × 320 × 3 | uint8 | RGB from the left wrist camera |
frames_right_camera | T × 240 × 320 × 3 | uint8 | RGB from the right wrist camera |
endpose | T × 16 | float64 | Per arm: position (x, y, z), orientation quaternion (w, x, y, z), gripper opening |
qpos, actions, times, frequency | float64 | Joint positions, joint targets and timing; sent, but vector_v1.1 does not read them |
What the network makes of it:
- State and action. Row t of
endposeis the state at step t, and row t + 1 is the action taken from it. Each arm's pose becomes position 3, a 6-D rotation (the first two rows of its rotation matrix) and gripper 1: a 20-number row for both arms. - Frames. Every stored frame goes through a JPEG round trip, as the training data did, and is resized to 224 × 224. The network crops the centre 95 % and resizes back before the encoder.
- Kept frames. One frame in 15 is kept, with the 15 actions that follow it; the last group is padded with zeros. A demonstration can be up to 1,000 steps: 67 prompt tokens.
Input: the observation
At every call during the rollout the policy receives the live robot, in the same form without the time axis:
| Array | Shape | Type | What it holds |
|---|---|---|---|
frames_head_camera, frames_left_camera, frames_right_camera | 240 × 320 × 3 | uint8 | The three cameras now |
endpose | 16 | float64 | Both arms' pose and gripper now |
A prediction reads the current frame and the one before it; at the start of an episode the first frame stands in for the one before. Live frames are resized to 224 × 224 without the JPEG round trip. The policy never sees the scene seed, the success condition or anything else the simulator knows.
Output: an action chunk
Each prediction is 16 actions of 20 numbers: per arm, the end-effector position, a 6-D rotation
and the gripper. Each is converted to the benchmark's 16-number pose row, the same layout as
endpose: position, a unit quaternion with w ≥ 0 recovered from the 6-D rotation, and the
gripper clipped to between 0 and 1. The robot moves both arms' end effectors to that target.
| Field, per arm | Numbers | Range |
|---|---|---|
| Position (x, y, z) | 3 | metres, the frame endpose is reported in |
| Orientation (w, x, y, z) | 4 | unit quaternion, w ≥ 0 |
| Gripper | 1 | 0 closed to 1 open |
One unit, call by call
The first 12 actions of each chunk are executed; the last 4 are replaced by the next prediction, made on the frame reached after the twelfth. The policy has 60 s to answer each call, and the unit ends when the task succeeds or at its step limit, 2× the expert’s steps in the scene.
Training your weights
vector-runtime convert turns a training checkpoint into model.safetensors with exactly the
architecture's tensors. What you train on, how long, and with what augmentation is yours.
The validator converts a demonstration the way the network's training data was stored: RoboTwin
episodes, three cameras, JPEG-compressed frames, and state and action as consecutive endpose
rows. Training on data in that form means the prompt your model sees in a duel looks like the
prompts it learned from.
Replaying the demonstration's actions does not solve a unit: they were taken in another scene. The tasks you are scored on are listed on The tasks, and the scenes of a duel are never ones the benchmark publishes.
