How do we know our own results are real?
Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means.
A note on the language in these records. This is a working laboratory notebook for research into training AI models more cheaply and efficiently, so you will read that an approach did not work, that a result did not hold up, or that one method was worse than another. That is the research doing its job, not a verdict on the engineering we deliver to clients. Ruling an approach out is how the search narrows, and these are the pages that teach us the most: nearly every technique we now rely on came from understanding why something else fell short. Testing our own ideas at least as hard as anyone else's is the point of publishing them. More about this programme and why we run it.
Most of what goes wrong in this kind of work is not the experiment; it is the comparison. These are the times our own checks overturned our own headline.
Where this stands
Repeatedly, the control rather than the measurement decided the result -- and several striking findings dissolved when the right comparison was finally run.
If you are evaluating whether someone would catch a problem in your models, this is the group to read. It is the most transferable output the project has.
Still open · 36 published results bear on this question.
What this does not settle yet
Stated plainly, because the gaps are as much a part of the record as the answers:
- The programme still has no unlearnable-data control anywhere except one record (I6).
How these results fit together
Ten records that are all about this project checking its own work, and they are the most transferable thing here. Several overturned our own headlines. The largest is that a pattern we had described as the model expanding and then consolidating is substantially a *measurement* effect: we were re-fitting the coordinate frame at every step, and the apparent contraction is the frame rotating rather than the model unwinding. Any quantity compared across training steps now has to say which frame it was measured in. A companion record does the opposite service -- certain quantities depend only on eigenvalues, which are unchanged by rotation, so that critique cannot reach them, and one of those still peaks on the event. We also audited how many of our own results are close calls, found that eight runs is not enough for several of them, and published that. The habit that produced most of these: re-read the whole record after every new result, not just the item in front of you. The number of times the mistake was in the *analysis* rather than the experiment is the single most useful statistic on this site.
The results
Each of these is a self-contained record: what we asked, what would have proved us wrong, what we found, and what it does not show. They open with a plain-language summary before any of the technical detail.
Three Audits, No New Faults
No training. Eleven grids flagged; two known faults, four honest joint descriptions, five lookup tables. Three audits, zero new cases.
Read the recordNobody Tuned the Rate
No training, just reading our own code: every speed-up claim was measured at an untuned rate, and the data that would have caught it had been in the archive for a month.
Read the recordThe Design Carries Its Derivation
No training. Twenty-six scripts build a setting from a measurement; twenty-three do it fairly. The three faults are known, and two came from reusing a good design.
Read the recordThe Check We Should Have Run First
We gave a model the same inputs with random answers, so there was nothing to learn. The effect we have been reporting all along vanished, which is what we hoped.
Read the recordEight Runs Is Not Enough
The usual statistical test said we had found something. The range of values the data was consistent with said we had not. Repeating it on three times as many runs settled it, and the answer was yes after all.
Read the recordHow Many of Our Results Are Close Calls
Sixteen of them are within one error bar of their own verdict. We are publishing the list rather than quietly rerunning the convenient half of it.
Read the recordWe Read Our Own Warning List
Only two were things a conclusion depends on, and both had already been rechecked. The tool ranks careful studies as the shakiest, which is backwards.
Read the recordConsolidation Is a Measurement-Frame Effect
The same 41 runs show a number falling and rising over the same interval, depending only on whether the measuring frame is allowed to move. Nothing contracts.
Read the recordThe One That Was Not A Ruler
A measure that cannot depend on which directions you call important rises to a peak exactly when the model learns, then falls back. So the change is real.
Read the recordThe Fix That Fixed Nothing
The one dial we had made the task fainter rather than deeper, and faint is hard for everyone. We learned what the next attempt actually has to do.
Read the recordWe Tested Someone Else's Prediction
Cutting the route made no difference at all. That rules out one of the three explanations the paper offers, and leaves two standing.
Read the recordThe Experiment We Could Not Run
The control worked. The comparison arm never produced anything to compare, and reporting its result would have looked like a finding when it was evidence the task was wrong.
Read the recordThe Archive as an Instrument
A pooled correlation of -0.679 that dissolved on stratification, and a scaling law that survived. Also the discovery that our archive varied only one of the four settings we thought it had.
Read the recordThe Rule Caught It The Same Day
A real effect at two model sizes and nothing at the third. Notable as much for the process as the finding: the check that caught it was added as a standing rule the same day, off the back of the previous result.
Read the recordThe Same Confound In Six Results
No experiment was run. Reading found that six results we published as effects fading with model size might instead be effects fading as the model outgrows the task, which is a very different claim.
Read the recordThe Effect Never Faded
This overturns two of our own records from the same day and rewrites a rule we had just written. Once the task keeps pace with the model, an effect we reported as vanishing is flat across every size we tested.
Read the recordThe Same Numbers, A Longer Run
The second confound found in the same small experiment, and neither was the one being controlled for. We withdrew the finding rather than the caveat.
Read the recordA Yardstick That Does Not Stretch
A measurement that gave different answers depending on when we stopped training, and two repaired versions that do not. Built and tested entirely from data we already had.
Read the recordThe Window Was There All Along
The second of six results we are rechecking, and the second where the disappearance was our own measurement rather than the model's size. It restores the scope of our only genuine efficiency result.
Read the recordThe Warning Gets Longer, Not Shorter
The third of six results rechecked, and the first where our own hypothesis failed. The earlier picture was distorted, but not in the way we expected, and the instrument comes out better than either version suggested.
Read the recordThe Alarm Never Got To Speak
The last of six results rechecked. Five of the six turned out to be measurement artefacts, and the two instruments we had written off both give more warning in bigger models, not less.
Read the recordIt Does Not Survive A Second Task
The counterweight to a good day. The same experiment that survived being made bigger does not survive being moved to another task, which scopes five of our results to one exercise.
Read the recordThe Scoring Inflated The Imbalance
The conclusion survives the better measurement and the number we had been quoting does not. Both endpoints were computed on the same 100 runs so the old one is an exact anchor rather than a comparison.
Read the recordFourteen Copies of One Definition
Zero of fourteen definitions drift by more than one measurement interval when a fifth of the run is discarded. The drift that does exist is confined to runs that had not finished improving, and the definition anchored to the task rather than the run does not drift at all.
Read the recordThe Measuring Grid Moves the Answer
A prediction from one line of algebra, confirmed to 0.3 of a percentage point, plus a second axis nobody had looked at. Nothing published trips it, because an earlier result had already said not to.
Read the recordClose Calls Are Decided by the Code
A usable rule falls out: a comparison within one measurement interval of the boundary is decided by the code, not the data. Plus the discovery that eleven of fourteen versions cannot measure a warm-started model at all.
Read the recordFour Of Them
The kill test does not fire: 44.2% of surviving claims rest on one task against a 20% threshold. The first run of the audit answered confidently over an empty set.
Read the recordWhere the Fault Was
No training. The kill test fires: zero further cases, because a house convention makes the correct thing the easy thing.
Read the recordThe Shortcut Reversed
Our central efficiency negative survives being re-run properly, and gets stronger. The instrument got seven times cheaper relative to training while the thing it was buying went negative.
Read the recordOne Number for Every Size
Both figures we could re-price moved by more than a factor of two, in opposite directions. The instrument cost three times more than published and our one positive saving is understated by nearly five.
Read the recordIt Never Becomes Free
The cost of switching off keeps falling and never reaches free. And the boundary we had found was convergence wearing the transition's clothes: at the original setting the model was already 99.6% of the way to its final accuracy.
Read the recordThe Shape Held, the Level Did Not
A growth rate is a property of the algorithm and a percentage is a property of the computer you ran it on. The same fifteen-minute pilot reproduced one to 0.05 and missed the other by fifteen points.
Read the recordWhich Side of the Moment
Data-per-step decides nothing, which stands. The six-for-six pattern turned out to be a slope read as a step, and the correction is on the page.
Read the recordThe Moment Marks Nothing
Our own quick test had given the right answer and we overrode it with a pattern that fitted every point we happened to have. That is the most likely thing a confounded grid produces.
Read the recordWhere the Shortcut Stops Paying
Four sizes: +21.1%, +19.9%, +12.9%, then -42.0%. Monotone, so the test passes, but three points sit within eight of each other and the fourth falls fifty-five.
Read the recordThe Cliff Survives Matching
The kill test does not fire and the prediction was wrong. It is also not a clean law: non-monotone, and the cell carrying the verdict is the worst-matched one.
Read the recordBack to all research questions
Get new results as we publish them
Roughly monthly, one finding per email, in plain English first. Including the approaches that turned out not to work, which are usually the useful ones. No sales email.
Double opt-in: we send a confirmation link and add nobody who does not click it. One-click unsubscribe on every email.