Progress of the P-program track toward the Hutter Prize
on enwik9. Higher is better — an inverted loss curve.
This page is the entry point for the current P-program track: the
top-level index links here and nothing else from the track, so everything below is
reached from this page. Each instrument states what it is and what it is not; per
#hutter_publication_handoff, none of them is a result — the chart
above is where results go.
The chart carries no p8v2 point, and that is the intended state.
No axis has been decided, so there is no p8v2 to plot; each variant's k is on its own
page, by size class. The registry that feeds the chart
(docs/update-progress.py) says so where the series are defined, so the
absence is not read as an omission.
#f-p8-forward, model header
byte 42) and what a position does with the applications that reach it over
time is axis H (#f-p8-decay, byte 43), so the decay rule can be
moved with the delivery held still for the first time. The six retired axis-A ideas
are restated on the cap-8 delivery and run as v026…v031,
against the baseline v024 and the B2 ground-truth column
v025 — both of which reproduce their generation-4 rows byte for
byte, because the two new digits are written only when a variant names the split.
Same instrument as generation 4 — the f panel, the pattern browser with its
concordances, the support painted over the sample, the probability section —
with a fourth pane, the ladder, carrying every decay rule at e64, e1k and
e10k, computed from the dumps rather than copied from the report.
settled_ok is partly
circular, because a rule whose readout coincides with the predictor omega used to
choose which bytes to record scores near-perfectly without settling doing any work.
Every settled_ok in generations 1–5 is a measure of agreement with
the causal pass — causal in the NLP sense, meaning left to right,
left context only, no lookahead, and not a claim about causation — rather
than of recovery. Do not carry that word onto f. Settling is deliberately not
causal in that sense: rather than going one token at a time as a transformer does, it
puts a set of salient points into a memory trace and works out the intermediate
points through the window of indeterminacy from both sides, so every
position takes a backward application from its right neighbour that the left-context
pass could not use. That direction is the one the metric cannot see.
v024 carries A7, the pair — the k=1 pattern delivered at
min(8, ws) with fall-off 1, one alternative because the
delivery and the fall-off only make sense together — and v025 sets
it against B2 as ground truth, where the stored learned weights are the
delivery. The ground truth cuts both ways: at e64 the stated 8 recovers fourteen
unrecorded positions where the stored truth recovers one, and by e10k the stored
matrix wins. Same instrument as generation 3 — the f panel, the pattern browser
with its concordances, the support painted over the sample, the probability section
— with the constants f uses now read from axes.json's new
constants section rather than held anywhere.
v002 against A4, A5 and A6, the three rules
that set a pattern's fall-off from its own support rather than from a constant.
Carries the panel that is f — for one position and one time step, every
pattern application in SN with its arithmetic, in order, ending in the new activation
vector — and the LATD expansion, which names the input positions a pattern
records. Since 2026-08-15 that expansion is the instrument: selecting a position
paints each pattern's support over the sample itself, one colour per
application and striped where two of them were learned from the same byte. Under the
grid, three panes share its screen — f at the selected position, a browser over
all 303 patterns with a concordance for each, and every settled
prediction traced back through its patterns to the data. Then
what licenses that chain, which renders
the justification where one exists and asks for it where none does. Nothing
here is pinned: every rule accumulates with LSA addition, which draws from a stream
the page cannot replay, and the page says so and compares what it still can.
v002, ten variants.#hutter_metrics. Kept
because its runs were re-run after the k=0 removal and are therefore comparable with
generation 2's.gen1-viz.py, which takes the
generation off the posdir name; they are served under the p8v2-gen1/ path
because that is a published URL, not because they are generation 1's.k=0, which no longer exists.The reports. What the runs say, and the picks they ask for, are
published beside the pages rather than summarised here:
generation 5
(and the goal that
carried the picks in),
generation 4
(and the goal that
carried the picks in),
generation 3
(and the goal it
answers, and the push that opened it),
generation 2
(and the goal it answers),
generation 1,
#variant_protocol for why
agents do not make the picks, and
#hutter_publication_handoff
for the division of labour.
The metric. Score k is defined by S/U = 0.99k, i.e.
k = log(S/U) / log(0.99), where U is the uncompressed size and
S the archive (compressor + decompressor + model), extrapolated to whole
enwik9 (109 bytes). Each +1 in k is a 1% smaller archive, so the
dashed reference lines are, from the bottom: the zip "dumb" baseline, the current
Hutter record L, and the record improved by 1% (the minimum prize claim).
The extrapolation holds the measured small-sample rate constant, so a plotted k is a
conservative floor. Full definition and provenance live in block
#hutter_leaderboard_goal_20260701; data in
progress.json.
The rate r. Each point is measured on a small sample as
r = P/U: U is the sample size in bytes and P the
data-dependent archive payload in bytes — the sections whose size scales with
the input. r < 1 means the sample actually got smaller. Costs that
are the same size for any input (the decompressor binary, model header, fixed tables) are
the fixed cost, paid once; the extrapolation above is
S/U = r + fixed/109.
Predictions. The gold mini-series pinned at
2026-07-06 — P8's date, not when any experiment runs —
is P8's ladder over input scale: one dot per order of magnitude, 104 to
109 bytes. The two smallest scales are measured enwik9-prefix runs,
drawn as flat dashes (a measurement, no model CI); from 106 up each step is
the component-model extrapolation, drawn as a violin — the density in k
implied by the block's σr = 0.045 (95% = ±2σ) —
with the median line going dashed at that discontinuity. Marks are labelled by scale
only; hover one for its k, r, CI and training time. By default the ladder is
collapsed: drawn at the true time scale of the shared x-axis, and since the whole
ladder streams in ≈83 minutes at the pinned 200 kHz input clock, it sits as a
sliver at P8's date. Click the gold cluster, or its gold
10⁴…10⁹ ⊞ label on the x-axis, to expand just that
slice of the x-axis to log training time (0.6 days per decade) — the fragment
#pred-ladder goes into the URL, so the expanded view is linkable — and
click anything gold again to collapse.
The k axis is shared with the measured points, so later experimental numbers plot
directly over the ladder — and when the slice is expanded, measured runs at a
ladder scale align to that scale's x, sitting in their violins. The gold violin is the 95% confidence
interval at enwik9, shaped by the density the prediction's r-uncertainty
implies in k. These are extrapolations from the measured prefix ladder (block
#pprog_p8_words_20260706, 2026-07-06), not measured points: predicted rates r
come from the block, k is computed here from r.
Every measured point, grouped by program — click a chart point to jump to its row. Provenance shared by a whole series is stated once; the table notes only what differs per point, and each row's exact reproduction commands and pins are under its "repro & provenance" toggle.