e64 — stepping f

The 64-byte enwik9 prefix, settled one time step at a time under live settling rules. Select a position; step f; the patterns that reach it light up the data they were learned from. Generation 3, frozen; the current instrument is generation 4. Notes and open questions · what licenses any of this.

THIS IS THE GENERATION 3 RECORD, AND IT IS FROZEN. Generation 4 answered the choice this generation left open — the k=1 strength — and its instrument is /hutter/pprog/p8v2-gen4-e64/: the A7 pair (delivery min(8, ws), fall-off 1) against the B2 ground truth. This page is not rebuilt to follow it: a superseded page is frozen so it can still be compared with. On this page the k=1 strength readings are generation 3's — the query layer of the time reported the literal 1 — and generation 4 extirpated both that literal and the 255 reading it sat beside. See #viz_standard, "Versioning a viewer".
the controls
controlwhat it changes
A1…A6What a position does with the applications that reach it. This is f, and it is the axis generation 3 is working. A1 assignment: the longest k that fires assigns. A2 LSA addition against a decay of 2 per step. A3 addition with renormalisation to constant sum 255. A4 accumulate then decay by the support deficit smax−s. A5 no decay: a pattern of support s is applied only on time steps that are multiples of 2smax−s, so the rate is the fall-off. A6 accumulate then decay by the number of applications that arrived. All applied live. A1 and A3 are pruned and kept as controls; A2 keeps its measurements and lost its justification (notes).
smaxThe reference support A4 and A5 measure a deficit from. Which reading the runs used is an open choice, so it is a control and not a constant.
B1 B2The backward k=1 application. B1 reconstructs it from the forward argmax table, each of the n predecessors paying the ceiling of the base-2 logarithm of n. B2 reads a stored 65536-byte LPP — refused where that would borrow it from a model with a different learned k=1.
C1 C2The AND gate at partial activation. C1 min(wa,wb). C2 sums and thresholds at 128. C2 never binds at any runnable sample: the learned k=2 weights are 1 and 2 and the gate has nothing to bite on.
min / subThe application rule. min(wp,ws) is what every run used and is frozen normatively in #f-p8. wp−(255−ws) was tried and rejected, is implemented in no variant, and every frame under it is unverified. Set it on the baseline and step: every non-absolute application dies one hop from a clamped position and the window collapses. That collapse is the single most useful thing on this control.
D E FWhich learned model is loaded — not settling rules. Sparsification, replay and pruning change what was learned, so each selects a different model file. Only single-axis combinations were run; two at once is refused rather than borrowed.
stepsHow far f runs. An application travels one position per time step, so stopping short of W leaves the far end of the window on its initial state. On the baseline the agreement readout gives 11, 12, 14, 27, 49, 62, 64 at 1, 2, 4, 8, 16, 24 and 32 steps.
presetThe four runs. Each carries its recorded settled_ok (positions whose settled argmax is the sample byte, clamped ones included, so it is bounded below by the 5 recorded positions) and conv (the time step after which the window's argmax sequence stops changing; equal to W means it never did). Anything not on a preset is the viewer's own and is marked as such.
how to read the grid
bytethe sample. Non-printable bytes are ·. Shaded where the selected pattern or application was learned from, in that application's colour; striped where two of them share a position.
argmaxthe settled argmax at this time step. Blue = clamped (the trace records it, held at 255 on its own event, never updated). Green = agrees with the byte, red = disagrees.
w1the weight of that entry, in log support: v stands for a magnitude of about 2v. Over 99 is shown as tens with a subscript zero.
chgshaded where this time step changed the position's argmax. A clear row means the window has stopped changing.
the axes

the operations, in ordinary arithmetic

Every weight here is a log support, which is why the arithmetic is not the arithmetic it looks like. Nothing on this page names an operation without a row here.

Every pattern the model holds inside the window — 256 k=1 rules and the k=2 tokens — each expanding to the input positions that formed it. Selecting one writes its concordance and lights those positions in the grid above. At this sample the model header says n = M = 64, so the training set is the 64 bytes on screen and the expansion is exact, not sampled.

Every settled prediction, the patterns that produced it, and the input positions those patterns record. One row per position, over the full run of f under the current controls. Click a row to drive the instrument to it.

Two things in that table are not defects. A row often lists a pattern out of 0x00: at the first time steps an unclamped neighbour is all zeros, its argmax is byte 0, and the rule out of byte 0 is genuinely what fired. And the backward direction is named as a reconstruction rather than by an atom, because under B1 it is not a stored pattern — the query layer prints no atom for it and this page will not mint one. It has a support set all the same.

Is any of this probability?

The chain the instrument puts on screen is: input positions → a pattern's log support for a joint event → f → the log support values of the settled event space. What licenses each link. Nothing here is this page's own derivation — everything asserted is quoted from a document published beside this page, and where no document says it, this section says that instead.

the linkwhat licenses it
input positions → a pattern. A pattern is about the positions the concordance lists. #f-p8: “the M-1 pattern with one byte and the next byte being input and output spaces … is a sufficient statistic on the frequency of the joint events, i.e. the byte pairs.” So the antecedent–consequent pair names a joint event and the strength is a log count of it. That count is what the concordance enumerates, which is why LATD is not a debugging view but the pattern's own definition read back.
a count → a byte. Why a strength of 1 is not “seen twice”. LSA.md: an LSA value “represents the log value of a count”, the error is uniform in the log domain, two values are required to have uncorrelated error distributions, and the increment gives w+1 with probability 2−w. The stored weight is a stochastic estimate of the count, so the two are not required to agree — and on this sample they do not.
one pattern → a probability, from an absolute antecedent. #f-p8: “In the case of assignment, given the full joint pattern on the ESs in which t_i and t_j participate, we see immediately that softmax on the ES of t_j gives the correct probability distribution given the event e_i (corresponding to t_i = 255) and the information about the frequency of the joint events, the sufficient statistic, captured by the pattern.” With LSA.md: “LSA variables which form an event space are interpretable as probabilities via softmax.” Read the conditions: assignment, not accumulation; t_i = 255; before f has run.
two patterns → one distribution, and the condition under which that is correct. #f-p8 states it with its condition attached: “If another pattern that is derived from a disjoint set of observations is brought in, these two distributions can be combined into the correct (per probability theory) combined distribution, which is not true if this additional parameter is not known (or if the observations are not disjoint, or more generally, are correlated in a way that is unknown).”

Four places the chain is not covered

1. THE ANTECEDENT IS NOT ABSOLUTE. The softmax result is derived for t_i = 255. In f a forward application delivers min(wp, ws) where ws is the source position's current activation — after one time step that is neither 255 nor a count of anything observed, it is the output of the previous time step. min of a log count of a joint event and an activation is not derived anywhere we can find. Axis B is frozen at b1 normatively, but a freeze is a decision, not a derivation.
2. THE PATTERNS THAT MEET AT A POSITION ARE NOT DISJOINT, AND THE GRID SHOWS IT. The condition quoted above is disjointness of the observation sets. The forward k=1, the backward k=1 and the k=2 token that meet at one position were all learned from the same 64 bytes, and their support sets overlap — which is what the striped cells in the grid are. Measured live, for the position selected: Under A2, A4, A5 and A6 those three vectors are combined with LSA addition, which is log2(2a + 2b): addition in the count domain. Adding the counts of two event sets that share observations counts the shared observations twice. Whether that is the intended reading, and what the operation should be when the sets are known to overlap, is not written down in any document published beside this page.
3. ONE PATTERN, THREE STRENGTHS, AND NONE OF THEM IS ITS SUPPORT. The k=1 section of the model is 256 bytes of argmax and stores no weight, so every reader supplies one. The query layer supplies a literal 1. The settling shell supplies 255 — not by writing it, but by omitting the attenuation entirely (fwd[table[a]] = wsrc), and 255 is what min's identity is on a byte. Under B2 the same pattern is applied at its actual learned weight, lsa_min(c1[a×256+c], wsrc), which at this sample is 0, 1 or 2. On a log-support scale those are counts of 2255, 2, and 1–4 of the same joint event, and every operation downstream — min, LSA addition, a support deficit — treats them as the same kind of number. Detail, and the reason the 255 is not simply a bug, in the notes.
4. THE SETTLED VECTOR. Is softmax of a settled activation vector a probability distribution over the next byte? The softmax result is stated for one pattern, from an absolute antecedent, by assignment, before f runs. After W time steps of accumulate and fall off, no document published beside this page makes the claim again. The original derivation had a bridge — a constant sum through time, from which the fall-off rate followed — and #p8v2_gen3 removed it: “There is no constant sum ever in an ES. The only thing constant is total probability = 1 and that is after softmax.” That passage of #f-p8 is marked WITHDRAWN and nothing has replaced it.

These four are a request for a document, not a finding: the derivation is the programmer's, and inventing one here would be the same failure as rendering a rule that was never run. Written up as #p8v2_latd_probability_ask_20260815. When the document exists, this section renders it instead of asking for it.

what the conformance check does and does not pin

which page this is, and what it is not

Generation 3. Generation 2's instrument is /hutter/pprog/p8v2-e64/, frozen as that generation's record. A page is not rebuilt in place when its generation is superseded: the generation-1 page was, and no longer exists to be compared with. The convention is in #viz_standard, “Versioning a viewer”.

Four variants were run — the baseline and generation 3's three new axis-A alternatives — and are marked RUN; every other setting is the viewer's own and is marked not a measurement. There is no pick, no ranking and no compression number here; the picks are the programmer's (#variant_protocol).

sources
models/p8v2/enwik9/64/the retained models; k=1 argmax table at offset 48, backward LPP at 304, then four bytes per token rule. No k=0 section
gen3-pos/generation 3's recorded endpoints, which the conformance check compares against
gen2-pos/generation 2's, kept because v013's stored count matrix is read from that generation
gen3.tsvgeneration 3's measured facts, including the e64 settling diagnostics
axes.jsonthe axes, their alternatives and their definitions — the page reads it, it is not copied here
#f-p8f itself: the single normative statement. The application rule, the window and its tiling, the time steps, initialisation, clamping and the argmax tie-break are all in it
#f-p8-deficit, #f-p8-period, #f-p8-indegreeA4, A5 and A6, each in its own words
LSA.mdwhat an LSA value is: the log of a count, uniform expected error in the log domain, uncorrelated between values by construction
#p8v2_latd_probability_ask_20260815the four questions above, written as a request for a derivation
p8v2-words.mdthe fixture; § Axis A is the settling shell this page implements
the generation 3 reportwhat these runs say, and the picks it asks for
#pprog_p8v2_gen3_goal_20260810the goal generation 3 was run against, and the choices it waits on
#p8v2_gen3the push that opened the generation, with the programmer's answers inline
#p8v2_gen3_viz_goal_20260810the goal this page was built from
#viz_standardwhat goes on anything we publish, including how a viewer is versioned
#hutter_publication_handoffthe division of labour between the two trees
build-p8v2-gen3-e64rebuilds this page and the notes. Runs no compression
p8v2-gen1/the generation panels: every variant of a generation over the same input positions

Published under #viz_standard. Self-contained: no external stylesheet, no script, no font, no image, nothing fetched. Data symlinked from ../cmpr-src/; page regenerated by docs/pprog/build-p8v2-gen3-e64, never hand-edited. The blocks are the authority.