Vector · Operating
Running a validator
One GPU host with three Python environments: the simulator, the policy runtime and the subnet. How to set it up, check it, run it, and what it costs.
2 min read
What a validator does
A validator reads commitments from the chain, keeps the queue, plays each entry's duel against the reigning king on its own GPUs, writes the result to an append-only store, and sets weights for the champions every 360 blocks. Nothing a miner submits is executed: the validator serves every entry with its own runtime, from weights it has checked against the pinned architecture.
The host
One host with an NVIDIA GPU and three Python environments, kept apart because the simulator and the policy runtime pin different stacks:
| Environment | Python | Holds |
|---|---|---|
| Simulator | 3.10 | RoboTwin-Vector (robotensor_bench/scripts/install_robotwin.sh) and vector-protocol |
| Policy | 3.12 | torch 2.8.0 with CUDA 12.8, vector-runtime[model] and vector-protocol |
| Host | 3.12 | robotensor, vector-orchestrator and vector-runtime |
One evaluation unit, a simulator and a policy server together, used up to 34 GB of GPU memory in
our measurements. The number of units each card runs at once is the workers setting, one by
default; an 80 GB card can run two.
Setup
git clone -b vector https://github.com/robotensor/RoboTwin-Vector.git
git clone https://github.com/robotensor/vector-orchestrator.git
git clone https://github.com/robotensor/robotensor-subnet.git
uv venv --python 3.12 .venvs/subnet
uv pip install --python .venvs/subnet/bin/python -e robotensor-subnet -e vector-orchestrator \
-e vector-orchestrator/packages/vector-protocol -e vector-orchestrator/packages/vector-runtime
Point the [vector] table of config/<network>.toml at the two other environments and at your
wallet:
network = "finney"
netuid = 0
wallet = { name = "validator", hotkey = "default" }
[vector]
policy_python = "/abs/.venvs/vector-policy/bin/python"
simulator_python = "/abs/.venvs/robotwin/bin/python"
simulator_root = "/abs/RoboTwin-Vector"
workers = 1 # units per GPU at once
mirror = "" # a Hugging Face dataset to mirror the result store to, if any
On a rented GPU pod, docker/build.sh in robotensor-subnet builds an image with all three
environments, the simulator's assets and the checkouts; run pod-check --selftest on a new pod.
Check, then run
robotensor doctor --config config/<network>.toml
export HF_TOKEN=...
robotensor validator --config config/<network>.toml run
doctor checks the config, the chain, the wallet, the Hugging Face token, the GPU, disk and clock,
the contract, the benchmark and both interpreters before the validator goes live. Once it runs:
| Command | What it does |
|---|---|
robotensor validator … status | The queue, the king and the champions |
robotensor validator … weights --dry-run | The weights the validator would set now, without setting them |
robotensor validator … duel --challenger owner/name@sha | Run one duel by hand |
An interrupted duel resumes from the validator's data directory (var/<network>/), with its seed
block kept.
Cost
A duel plays 160 (10 × 16 tasks) per side and gets at most 24 h of work; units still unplayed when that runs out are void. Every entry is the same network doing the same work per prediction, so a duel's cost does not depend on whose weights it plays. The per-unit limits are on What you submit.
