Skip to content

Under maintenance. Some results may be incomplete or inaccurate.

Robotensor
All pages21

Horizon · The guide

Reading the dashboard

What every part of an epoch’s page is showing, and which rule decided it — the page to read when you are looking at the dashboard and wondering what you are looking at.

4 min read

The dashboard shows what an epoch measured. It names the rule each number was measured under and links to it, and it explains nothing itself: a page that restated a rule beside the numbers would be a second copy of it, free to drift from the one that is maintained. This page is the map between the two — what each part of the dashboard is, and where its rule is written down.

The dashboard

The dashboard is two things and nothing else: the epoch that is running, and a table of every epoch this store has published, newest first.

Epochs are numbered rather than named. The first one the competition published is Epoch 1, and the count goes up from there in publication order. The number is this site's count of what this store holds, so no address is built from it — a mirror missing an epoch would renumber every one after it — and the record's own id, which every mirror agrees on, is on the epoch's own page.

Each row says whether the epoch is open or closed, whether its record verified, when it was published, how many units it ran, and the submission its record ranked first with that submission's score. The table is grouped into score series, and a note marks where one run ends and the next begins. A dry run is in none of it: it is counted nowhere and named on one line under the table, because a rehearsal is not a week the competition ran. An epoch, end to end is both rules: what ends a series, and what a dry run is.

An epoch: one table

An epoch's page is one table and nothing else. Its rows are the models the epoch scored; its columns are the score each one reached on each axis, and the overall score that is the mean of those. A row opens on what is under it, and nothing on the page is a second copy of what the table already says.

The headline

The epoch's id — the record's own name for it — the window it ran in, when its record publishes one, how many units it derived, and the axes it scored.

The table

One row per model, in the rank the record published: best overall first, ties broken by the time the intake stamped each entry with. No row is a reference — the competition evaluates no base model — and no number in the table was worked out here. Scores is how each one is formed. Every column header carries a sentence saying what its cells hold, so nothing on the table has to be guessed at.

A record published under the older rule carries no overall score. Its per-axis scores are shown and the table says so; nothing is worked out in its place.

A row, opened

Three levels, each the one above it broken down, and each the record's own numbers:

  1. The axis — the score that model reached on one benchmark. The axes is what an axis is, and why one of them is not a result on its own.
  2. The task — the success rate over that task's units. An axis score is the mean over its tasks, so a model strong on one task and absent on another reads the same as one that is even; this is where the difference between the two is visible.
  3. The episode — one run: what it came to, and the two clips of it. The demonstration the model was shown, and the rollout it produced. An episode the model has no result for says so rather than reading as a failure, and a clip that is not in the epoch directory is named as missing rather than given a player with nothing behind it. The frozen pool is what a demonstration bundle records, and why every model in the epoch was shown the same one.

The episodes come from the epoch directory, which carries no signature. Nothing in it is shown as this epoch's unless it agrees with the signed record; where it does not — or where there is no such directory on this box — the table says so in place of the episodes. Verification is that check in full.

The page an epoch leads to

A unit has a page of its own, reached from an episode: its demonstration and the rollouts run against it, one per model evaluated on this unit. Each names the bundle it ran against, and one that is not the bundle the pool froze for that unit is marked — a rollout against another demonstration is not a result on this one.

Reading the dashboard · Docs · Robotensor competitions