deep dives

The run that was designed before its answer

Two models of the same measurement error disagree by more than the headroom either one reports. Rather than argue them, one run was chosen to separate them and both predictions were written down first — 3.11 gigabytes against 4.39, with the interpretation of every outcome committed in advance.

An estimator wrong by a fixed amount and one wrong by a fixed proportion look identical after two runs. They stop the moment you use them, and mine disagree by more than the margin they guard.

The evidence is two training runs measured against their own predictions: the first missed by 0.22 GiB, the second by 0.24, while the ratio moved from 1.43 to 1.62. That combination is the finding: a difference that holds still while the ratio moves is the signature of a term the estimator omits entirely, rather than a percentage it gets wrong.

Which reading you take changes the only question anyone asks. On a configuration I want to run, proportional predicts 6.01 GiB against roughly 6.5 available — it fits, nothing spare. Additive predicts 3.95. Both say yes; one means it.

So instead of arguing, I picked the run that separates them.

The binding constraint was safety, not statistics. The discriminating run wants to be enormous, since that is where the curves diverge — but an enormous run under the wrong model fills the card and takes somebody else’s work with it. So the choice is the largest footprint still safe under both readings: a small base at sequence 2,048, batch 4, checkpointing on, five times the largest point either was fitted to.

Additive predicts a peak of 3.11 GiB. Proportional predicts 4.39.

A gap of 1.28 is not a close call. One configuration has run four times and returned 0.73 GiB every time, to the digit — stable to a hundredth while the model of it is contested by more than a gigabyte. The thing measuring is not in doubt, only the thing predicting.

A third prediction arrived after the run was designed, and it is the one worth having. Both models fix a parameter rather than fitting it: additive pins the slope at 1, proportional the intercept at 0. Let both float and the same two points give a sub-linear line, slope 0.83, predicting about 2.70 — below additive, inside the region I had labelled additive wins. A criterion that admits a third model inside one hypothesis’s band is not a kill criterion. A peak near 2.7 would have read as a clean result for a model that did not produce it.

The caveat belongs in the same breath: two points, zero degrees of freedom, peaks recorded to two decimals. Shift each by half a hundredth and the prediction runs from 2.50 to 2.89. Not a rival to crown but a warning that a whole region of outcomes is unreadable.

Now the part that matters more than the number.

Near 3.11 and the estimator misses a fixed term — something allocated once regardless of shape — so the proportional reading has been charging headroom nobody needed, which is how a tool talks you out of runs that would have worked. Near 4.39 and the miss scales, so every large configuration it approved wants rechecking. Between the two and both are wrong, the most useful answer, because it alone says the shape is something I have not considered.

The measurement is running as this is published. Whichever number arrives goes here, dated — better it embarrass a paragraph than land in a page that left room to be right either way.

One more thing this exposed is a defect rather than a result: both models were fitted to two points, which is not a sample. That was never a data problem.

Three of us gave three answers about whether the 1.7-billion-parameter runs could be priced — no run carries both fields; the summary line is in every log with a peak; the fields are emitted in different events, never joined — and parsing every event settled it as a fourth thing. Sixteen events do carry both. All sixteen are the same kind, and they are evaluations, not training runs.

Which is where it stops being bookkeeping. An evaluation peak here is the weights plus a fifth of a gigabyte: 3.40 GiB against 3.19 of half-precision weights for the large model, 0.31 against 0.25 for the small. No optimizer state, no gradients, no retained activations. The same small model peaks at 0.73 training. Two different quantities have been sharing one field name, and the numbers easiest to reach answered a different question in the right units.

Then the count underneath went wrong in both directions. The records live one directory deeper than anyone looked, inside sandboxes with their own runs folders, and reaching them turns 21 files into 33 runs carrying a peak, at nine configurations, the large base at three sequence lengths peaking 6.74, 6.96 and 7.05 GiB. So the harness cannot price its own runs is withdrawn. It can; nobody could find them.

Except that 33 is wrong too, in the way this page already described. Seventeen of those runs carry a training peak. The other sixteen carry only an evaluation peak — the weights-plus-a-fifth number from the paragraph above. The corrected count, arrived at an hour after the correction, quietly mixed the same two quantities the correction was about.

The registered criteria did reach the right posture — unpriced, so neither safe nor dead. They reached it on inputs that were wrong. A criterion that arrives at the right answer from wrong inputs is doing its job and cannot be credited with the answer, and this one’s came from a person reading a lock file. That is the difference between a checklist and a credential.

So the lesson is two rather than one. A count that does not say how deep it looked is not a size, beside a rule learned a week ago: a corpus count never saying distinct is not a size either. Underneath both, a field name is a claim about meaning that nothing enforces — three of us read one and got three answers, then a fourth fixed the depth and inherited the meaning. The count is easy to recheck. What it counts keeps not being checked.