Horizon · The rules
Scores
A task score is a success rate, an axis score is the mean over its tasks, and a submission’s score is the mean of its axis scores.
3 min read
Task rate to axis score
A task's score is its success rate over the units that produced a score. An axis's score is the mean over its tasks, so a task with many episodes does not outweigh one with few — and so a submission strong on one task and absent on another reads the same as one that is even. The rates the mean was taken over are published with it, which is where that difference shows.
Scores are fractions in the record and are shown as percentages here, on the axis scores and on the one number that is their mean alike.
The score
A submission's score is the mean of its axis scores. That is the whole of the rule. Every submission that ran is scored, every submission that was scored is ranked, and the rank is by that score, highest first.
The mean is over the axes the epoch actually scored, and every submission is scored on the same axes, so every mean is taken over the same set. The base model has a score too — it is the reference, and a reader should be able to see where it lands — but it is not in the rank: it is the organiser's yardstick, not an entrant.
A tie is broken by the submitted_at the intake stamped the entry with, and then by the key, so
the order is total and entering a copy of somebody else's model cannot get ahead of the model it
was copied from.
Nothing is paid. The competition is a leaderboard: a score and a rank, and no reward of any kind is computed, published or owed.
Paired units
Every submission on the shortlist is scored on one set of units: a unit any of them came back harness-void on is dropped for all of them, which keeps the comparison paired, and a runtime void that outlived its retries counts as that submission's own failure.
That distinction is the whole of it. A unit the harness could not run is nobody's fault and is taken away from everybody; a unit your model could not finish is yours.
Voids and a dropped axis
A unit nothing measured is one the axis is short of: it stands in no submission's rate, and an axis short of enough of them is not measuring what it was meant to.
An axis whose voids are over the epoch's max_void_fraction has its void units re-run once, and
is dropped from the epoch's scoring if it is still over or none of them can be re-run; the epoch
page gives the reason. An axis dropped that way is in nobody's mean — it is simply not one of
the epoch's axes — so a dropped axis costs every submission the same, which is nothing.
The record still publishes the voids and the dropped axes either way. How much of an epoch went unmeasured is a fact about the epoch worth knowing, not a scoring rule.
Both the fraction an axis reached and the limit it was compared with are the record's own: the limit is configuration, published in the epoch record and shown on Key numbers, and an epoch's own void counts and fractions, per axis, are on its page.
What an older record was scored under
A record of epoch record version 2 was scored under the old rule and carries no overall score at all. Its epoch's table shows the per-axis scores it does publish, and says plainly that no overall score was published for it.
A record from before version 2 was scored on each submission's own units, a void left out of that submission's score only.
