Skip to content

Under maintenance. Some results may be incomplete or inaccurate.

Robotensor Competition
All pages11

Vector · The guide

The model

The one network every entry fills, vector_v1.1: what it is shown, how the demonstration and the live frames flow through it, and the actions it returns.

5 min read

One architecture

Every submission is the weights of vector_v1.1, a Behavior Prompting Policy for two arms: a CLIP ViT-B/16 image encoder per camera, a six-layer transformer decoder that reads the live frames against the demonstration, and a 1-D diffusion U-Net that turns what it read into actions. It has 781 M parameters in 753 float32 tensors.

The validator builds the network with its own code and loads your weights into it, so the architecture is fixed and the weights are yours: what a duel compares is training and data.

The architecture, end to end

Once per unit · the promptEvery prediction · the policyInput · the demonstrationT steps · 3 cameras · 16-D pose1 frame kept per 15 stepsInput · the observation3 cameras · 16-D pose, now+ the frame before: 2 framesFrame tokenizer · shared by both paths3 × CLIP ViT-B/16, per camera → CLS → 7686 × Linear, per pose key → 7689 tokens / framePrompt token, per frame9 tokens + MLP(15 × 20 actions)→ attention pool → 1 tokenHistory tokens2 × 9 = 18 tokens × 768+ a learned position eachPrompt memory4 sinks + P prompt tokensP ≤ 67 (up to 1,000 steps)Transformer decoder × 6self-attn → cross-attn → MLPpre-norm, 8 heads, width 768crossEncoded once, when the demo arrives,reused by every prediction in the unit.Condition[18 history ; 18 decoded] = 36 × 768flattened: 27,648 numbersGaussian noise16 × 20Diffusion head · 1-D U-Net256 · 512 · 1024 channels, FiLMDDIM: 16 denoising steps+ the step embeddingOutput · an action chunk16 × 20; the first 12 are executedEach 20-D row: per arm position 3,rotation 6-D and gripper 1, sent to therobot as a 16-D pose row.
vector_v1.1, the only architecture a submission may fill: 781 M parameters in 753 float32 tensors. You train the weights; the validator builds this network with its own code and loads them into it.

The network has two paths that share one frame tokenizer.

  • The prompt is built once per unit, when the demonstration arrives. One frame is kept every 15 steps; each kept frame becomes nine tokens (one per camera, six for the arms' pose), is joined by an embedding of the 15 actions that followed it, and is pooled into a single prompt token. Four learned sink tokens go in front. The result is the decoder's memory.
  • The policy runs at every prediction. The current frame and the one before it become 18 tokens, which a six-layer decoder passes through self-attention and cross-attention to the prompt memory. The 18 tokens and the decoder's 18 outputs, flattened, condition a 1-D U-Net that denoises 16 actions from Gaussian noise in 16 DDIM steps.

Where the parameters are

PartWhat it isParameters
Image encoders3 × CLIP ViT-B/16, one per camera; the CLS token, width 768257.4 M
Pose projections6 × Linear, one per pose key, to width 7680.02 M
Action projectionMLP from 15 actions × 20 to width 7680.3 M
Prompt poolAttention pooling of a frame's 10 tokens into one, 8 heads7.1 M
Decoder6 pre-norm transformer decoder layers, width 768, 8 heads, MLP 3,07256.7 M
Diffusion head1-D conditional U-Net, channels 256 · 512 · 1024, FiLM conditioning459.4 M
Positions, sinks, normalisationLearned prompt and history positions, 4 sinks, per-key scale and offset0.06 M
Total753 tensors, all float32781.0 M

The normalisation is part of the weights: every pose key and the action have a per-dimension scale and offset you set from your own data. Weights with a zero scale or a non-finite value anywhere do not load, and a submission whose weights do not load is refused before its duel runs.

Input: the demonstration

At the start of a unit the policy is handed one demonstration of the task, recorded by the benchmark's expert in a different scene. It arrives as named arrays over T steps:

ArrayShapeTypeWhat it holds
frames_head_cameraT × 240 × 320 × 3uint8RGB from the head camera
frames_left_cameraT × 240 × 320 × 3uint8RGB from the left wrist camera
frames_right_cameraT × 240 × 320 × 3uint8RGB from the right wrist camera
endposeT × 16float64Per arm: position (x, y, z), orientation quaternion (w, x, y, z), gripper opening
qpos, actions, times, frequencyfloat64Joint positions, joint targets and timing; sent, but vector_v1.1 does not read them

What the network makes of it:

  • State and action. Row t of endpose is the state at step t, and row t + 1 is the action taken from it. Each arm's pose becomes position 3, a 6-D rotation (the first two rows of its rotation matrix) and gripper 1: a 20-number row for both arms.
  • Frames. Every stored frame goes through a JPEG round trip, as the training data did, and is resized to 224 × 224. The network crops the centre 95 % and resizes back before the encoder.
  • Kept frames. One frame in 15 is kept, with the 15 actions that follow it; the last group is padded with zeros. A demonstration can be up to 1,000 steps: 67 prompt tokens.

Input: the observation

At every call during the rollout the policy receives the live robot, in the same form without the time axis:

ArrayShapeTypeWhat it holds
frames_head_camera, frames_left_camera, frames_right_camera240 × 320 × 3uint8The three cameras now
endpose16float64Both arms' pose and gripper now

A prediction reads the current frame and the one before it; at the start of an episode the first frame stands in for the one before. Live frames are resized to 224 × 224 without the JPEG round trip. The policy never sees the scene seed, the success condition or anything else the simulator knows.

Output: an action chunk

Each prediction is 16 actions of 20 numbers: per arm, the end-effector position, a 6-D rotation and the gripper. Each is converted to the benchmark's 16-number pose row, the same layout as endpose: position, a unit quaternion with w ≥ 0 recovered from the 6-D rotation, and the gripper clipped to between 0 and 1. The robot moves both arms' end effectors to that target.

Field, per armNumbersRange
Position (x, y, z)3metres, the frame endpose is reported in
Orientation (w, x, y, z)4unit quaternion, w ≥ 0
Gripper10 closed to 1 open

One unit, call by call

reset(seed)seeds the noiseset_demonstrationbuilds the prompt memoryact(observation)until done or out of steps#1#2#30122440stepexecuted (12)predicted, then replaced (4)
A prediction is made every 12 steps, on the frame it is made at and the one before. The last 4 actions of each chunk are never run: the next prediction replaces them.

The first 12 actions of each chunk are executed; the last 4 are replaced by the next prediction, made on the frame reached after the twelfth. The policy has 60 s to answer each call, and the unit ends when the task succeeds or at its step limit, 2× the expert’s steps in the scene.

Training your weights

vector-runtime convert turns a training checkpoint into model.safetensors with exactly the architecture's tensors. What you train on, how long, and with what augmentation is yours.

The validator converts a demonstration the way the network's training data was stored: RoboTwin episodes, three cameras, JPEG-compressed frames, and state and action as consecutive endpose rows. Training on data in that form means the prompt your model sees in a duel looks like the prompts it learned from.

Replaying the demonstration's actions does not solve a unit: they were taken in another scene. The tasks you are scored on are listed on The tasks, and the scenes of a duel are never ones the benchmark publishes.

The model · Docs · Robotensor Competition