Notes — generation 5 on e64

What the split of f changed, what the eight runs say, and what is still open. The instrument is here, and every claim below that names a cell links into it. This page carries no controls and computes nothing.

Nothing here is a pick. Eight variants were run; the choices are the programmer's (#variant_protocol), and this page is evidence put in front of that judgement. The goal block carried the picks in, so the runs waited on nothing.

THE SPLIT, WHICH IS THE GENERATION

Axis A had been carrying a conflation since generation 1: what a k=1 application delivers and what a position does with the applications that reach it over time moved together. That is why generation 4 had to make “the pair” the unit, and why the support-driven rules carried the k=1 strength gap inside themselves — an alternative that changed the rate could not hold the delivery still. Generation 5 carves f in two:

axiswhat it iswhere it lives
#f-p8-forward G, the delivery. Digit 1, the baseline: the k=1 argmax is delivered at min(8, ws), forward and backward agreeing on wp after the backward direction has paid its predecessor split. Under B2 the stored weights are the delivery and this axis has no effect at all. model header byte 42
#f-p8-decay H, the decay. Digit 1, the baseline: accumulate every application that fired, one complete 256-entry pass at a time in the order forward, backward, token, then a constant fall-off of 1. Digits 2…7 are the six retired axis-A ideas restated on the cap-8 delivery. model header byte 43

Axis A is retired, not deleted. Its seven digits keep their recorded meanings forever, each naming the pair a historical variant ran; no future variant is built on them and they leave the live comparison (#variant_protocol). On the instrument that is why there is no A control and why every vector here begins with a 7.

THE VECTOR IS NOW VARIABLE LENGTH, AND THAT IS LOAD-BEARING. The two new digits are written only when a variant names the split, so v024 is still 711111 and v025 is still 721111 — six characters, the same spec, the same generated C, the same model file. Which is why the claim that generation 5 changed nothing underneath it is checkable rather than asserted: both variants reproduce all six of their generation-4 rows byte-identically, settled_ok, mean_sweeps and settle_apps alike. The six restatements carry eight digits, 71111112 through 71111117. The builder refuses to build if any variant's field disagrees with that.

Why 8, and why it is now on its own axis

8 is the weight that makes the argmax carry half the mass of an event space whose other 255 events sit at 0, and the measured share of the argmax at its contexts expresses as 9 / 8 / 7 at e64 / e1k / e10k — so 8 is the measured value at e1k and within one unit of it at every sample size (#f-p8-forward). Generation 4 could only argue that the fall-off, not the strength, carried the gain, because it moved both at once and reasoned about a grid. Generation 5 measured it. The three runs at e1k are (255, 2) = 165 from generation 4, (8, 2) = 167, and (8, 1) = 220: the fall-off carries essentially the whole gain and the strength carries two positions. At e10k the same corner says something the grid did not — 1177, 1133, 1288 — so the cap costs 44 positions when the fall-off is wrong and gains 111 when it is right. The pair really was a pair; it is now two axes and the interaction is visible.
255 IS PASS-THROUGH, AND THAT IS AN ABSOLUTE PATTERN, NOT A BUG. MJC's correction, which stands over every page here: “255 is pass-through and it is an absolute pattern. It is literally an absolute rule applied to something that represents partial knowledge. It is the highest strength a pattern can have because of the UM max-min forward pass, and because 255 is the identity of min.” Since the split it has a name and a place: it is the pass-through setting of axis G, one of the two points #f-p8-forward says the axis exists for. Neither is built — generation 5 moved the decay axis only — so both are drivable here and every frame under them is the viewer's own. That is the point of driving them: attenuate is the only form in which a strength of 1 can be applied at all, and stepping it shows why — min(1, ws) puts the whole window in a 0-to-1 band where the stochastic add is noise.

The six restated, and what each is for

Hthe rulewhat it is for
H1 accumulate, constant fall-off 1 the baseline. Not a restatement: this is what v024 runs, and what generation 4 shipped as half of the pair.
H2 assignment — the longest k that fires assigns; no accumulation, therefore no decay the axis's degenerate point, and the control that shows what accumulating buys once the delivery is no longer saturating. The only rule here that draws no entropy, which is what makes it exactly conformance-checkable in a viewer — see below for why its score is not a win.
H3 accumulate, constant fall-off 2 the one-constant control, and the reason the axis was worth running. With H1 and generation 4's retained (255, 2) runs it fills in the corner of the square that was never measured, and the attribution stops being an inference from a grid.
H4 renormalise to the arriving support the correction, not the restatement — see below.
H5 decay by the support deficit smax − s of the strongest arrival a fall-off set by the pattern rather than by a constant: in LSA a rate of one in 2d is a subtraction of d.
H6 no decay; a pattern of support s fires only on time steps that are multiples of 2smax−s the alternative that takes the word “rate” at face value — and the one whose promised fix did not land.
H7 decay by the number of applications that arrived the smallest correction that removes the withdrawn constant-sum argument entirely. It needs no support at all, so it is untouched by the k=1 strength gap and by axis B. It differs from a constant exactly where applications are missing: window edges, positions with no token context, and every position whose backward predecessor set is empty under B1 — 218 of the 256 bytes at e64.
WHAT A UNIFORM SUBTRACTION CAN AND CANNOT DO. This is the whole of H4, and it is worth more than its score. Subtracting a constant d from every entry of a position's vector divides every count it represents by 2d. That leaves the softmax distribution exactly unchanged and changes only the total number of observations the position claims to rest on. So a renormalisation is never a statement about which byte a position favours; it is only ever a statement about how much evidence the position represents. Generation 1's A3 renormalised to a sum of 255 and was read as imposing an invariant on the distribution — which is both the thing #p8v2_gen3 ruled out (“there is no constant sum ever in an ES”) and a thing the operation cannot do. H4 renormalises instead to the LSA sum of the supports of the applications that fired, folded with LSA addition rather than taken as a maximum, which is the correct combination when the observations are disjoint. Its measured behaviour is exactly what its block predicted before the run: worst on the board at the small samples, where the 0 floor annihilates, and within 40 of the baseline at e10k, where it barely bites. It is the only alternative on this axis whose vectors mean the same thing at time step 5 and at time step 100, which is why gap 4 names it.

The finding is about the measurement, not about an alternative

FIRST, THE WORD. “The causal pass” is omega’s sparsifier, and causal here means left to right, left context only, no lookahead — the NLP sense, most literally true of a character-level RNN taking one symbol at a time. It is not a claim about causation. Concretely: one left-to-right sweep predicting each byte from the kept k=2 rule for (p−2, p−1), or the k=1 table from p−1 where there is none, recording the byte into the trace exactly when that prediction is wrong. Read causal as left-context throughout.

And then do not carry the word onto f, because settling is deliberately not causal in that sense. That is the architecture, not a caveat. A transformer goes one token at a time and can only look left. The idea here is the opposite: put a set of salient points into a memory trace, then work out the intermediate points through the window of indeterminacy from both sides. Every position takes a forward application from its left neighbour and a backward one from its right, and the second is information the left-context pass structurally could not use. Backward causality from a memory trace is a claim this work makes — in these models and, the argument goes, in the brain. Kept and glossed rather than renamed, per MJC 2026-08-17: short and clear once you know what it means, overloaded enough that it does not travel unannotated (#viz_standard).

settled_ok IS PARTLY CIRCULAR. H2 scores 59/59, 775/833 and 5172/5349 on unrecorded positions — far above everything else, including the ground truth. It is not better. Generation 1's v001 is the same rule at a delivery of 255 and scored 64, 942 and 9828 in raw settled_ok against H2's 64, 942 and 9823 today: four generations apart, a delivery of 255 against 8, and the score is the same to within five positions out of 5349, the residue being entropy-stream drift. Two things follow. First, it confirms generation 3's principle from the other side — a uniform delivery cannot reorder an event space either, so capping at 8 cannot move an argmax that assignment then reads out. Second, and this is the one that matters: assignment's readout is the argmax of the single strongest application, which is very nearly the left-context (“causal”) predictor omega used to decide which bytes to record, and the unrecorded positions are by construction the ones that predictor got right. A rule whose readout coincides with the sparsifier's predictor scores near-perfectly without settling doing any work at all. On the measured set the two readings are literally the same event: at e10k the left-context prediction equals the true byte at 5349 of the 5349 unrecorded positions, because that identity is what “unrecorded” means — so “settled back to the true byte” and “agreed with the left-context prediction” cannot be told apart there. Measured on the dumps, the share of unrecorded positions whose settled argmax matches that prediction is 96.7% for H2 at e10k against 18–26% for every other rule on the axis. The direction settling exists for is the one this metric cannot see: the backward application from the right neighbour. That is why A1 was pruned in generation 1 despite the best score on the board, and the prune was right; what generation 5 adds is the reason. Every settled_ok in generations 1–5 is a measure of agreement with the causal pass, not of recovery. Read the ladder with that in hand.
THE SUPPORT-DRIVEN RATES STILL DO NOT BEAT A CONSTANT. H5, H6 and H7 all sit below H1 at every sample — 174, 164 and 135 against 220 at e1k; 1064, 987 and 972 against 1288 at e10k. The a5 direction has now been tried three ways across two generations on two deliveries and has not once won on this metric. H6 converges far faster than anything else (1.00, 13.50, 8.43 time steps) and spends the least, while recovering the least of the three: the rate idea doing exactly what it says and paying for it.
H6 DOES NOT CARRY THE FIX IT WAS PROMISED. The goal block said the period would decide before each pattern's vector was computed, so settle_apps would finally measure the energy a rate is supposed to save — generation 3 recorded that A5 computed all three vectors and counted them before the period decided, so a skipped application cost what a taken one cost. Moving the decision ahead of the messages requires restructuring the position loop, and the generated settling then stopped honouring the frozen shell's clamping: all 64 positions at e64 diverged from the reference, including recorded positions that must never be updated at all. It was reverted. So H6 is A5's schedule on the cap-8 delivery — a real restatement, and the delivery is the thing being varied — but the vectors are still computed, here and in the runs alike, and settle_apps still measures work done rather than work saved. Realising the saving is an open item against the frozen shell, not against that alternative.
THE FOLD THAT TREATED ZERO AS ABSENT. H4 needs the combined support of the applications that fired, which is an LSA addition of their supports. The generated C wrote it as t ? lsa_add(t, s) : s, which reads a support of 0 as “nothing has contributed yet” — but an LSA 0 means one observation and is a real support, so a first contributor with support 0 was silently dropped. Caught only because the replay conformance gate disagreed. This page's own f carries the fixed form, with an explicit first-contributor flag; the same defect would be invisible here otherwise.

Where the support comes from

H5 and H6 need each pattern's own support: c1[x × 256 + argmax], a value in the 65536-byte learned count matrix that a B1 model does not serialize. Generation 5 does not borrow it. The matrix is in the run set — v025 is this generation's B2 ground-truth column and stores it — so unlike the generation-3 page nothing is reached for outside the generation, and there is exactly one such variant, so nothing is picked between. The builder still compares the k=1 table and the token section of all eight variants byte for byte and refuses to build otherwise; that identity is what makes reading v025's matrix under v024's siblings honest.

The reference support smax at this sample is 2 (2 at e64, 6 at e1k, 9 at e10k), so H5 and H6 have three rates here; which reading of smax should win remains generation 3's open choice B. It stays a control on the instrument rather than a constant.

A PERIOD CAN OUTGROW THE WINDOW, AND e64 HIDES IT. The period is 2smax−s against a window of W time steps. At e64 the periods are 1, 2 and 4 and everything fires; at e1k they reach 64; at e10k they reach 512, and a support-0 pattern never fires inside a window at all — 169 of the 256 bytes at that sample. Whether that is right for weak evidence or a defect is what the generation is for. This page cannot show it, being one sample, and says so rather than leaving it to be discovered.
THE MODEL NAMES THE WRONG SUCCESSOR, AND LATD IS HOW YOU SEE IT. The k=1 pattern from m names e. Expand it: m occurs at positions 1, 12, 29, 44 and 61 of this sample, followed by e twice and l three times. The rule the model kept is the less frequent of the two, in the only data there is. The stored counts are c1[m][e] = 2 and c1[m][l] = 1: the log estimate inverted the order of a 3-count and a 2-count — the uncorrelated-error property LSA is designed around, seen from the data side. Unchanged since generation 3, because learning is unchanged. Open it: the pattern itself, with its concordance.

Standing facts, carried forward

THERE IS NO k=0 BACKGROUND. An unclamped position initialises to 256 zeros, and under B1 a byte with no predecessor sends no backward application at all rather than falling back to a base rate. The model file carries no byte-frequency section: the k=1 argmax table starts at offset 48, and a model built before 2026-08-06 is 256 bytes longer and is rejected by the builder on the size identity.
NEITHER f AXIS CAN REACH THE RATE, AND THAT IS NOT A DEFECT OF EITHER. Under the D1 baseline the trace is decided by the left-context (“causal”) pass and the decoder reconstructs the same way, left to right, so settling never runs at decode time and nothing G or H does can move the compression. The axes are read off settling instead, which is what the instrument draws.
C2 STILL NEVER BINDS. The AND gate is byte-identical to the baseline at every sample that can be run, because the learned k=2 weights are 1 or 2 and the gate has nothing to bite on. Pruned as of generation 3; kept as a control.
conv STILL MEASURES THE WRONG THING FOR B2 AT e64. v025 reports conv = 64, the never-converged sentinel: with the recorded positions clamped and 1-unit weights everywhere else, the argmax sequence keeps oscillating around the stochastic add. Generation 3's caveat on the conv column stands unchanged — a window that keeps flickering and a window that settled instantly can both defeat the measure, in opposite directions.

Where these go on

Published under #viz_standard. Regenerated by docs/pprog/build-p8v2-gen5-e64 from docs/pprog/p8v2-gen5-e64-notes.tpl.html, never hand-edited.