Research question

How do we know our own results are real?

Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means.

A note on the language in these records. This is a working laboratory notebook for research into training AI models more cheaply and efficiently, so you will read that an approach did not work, that a result did not hold up, or that one method was worse than another. That is the research doing its job, not a verdict on the engineering we deliver to clients. Ruling an approach out is how the search narrows, and these are the pages that teach us the most: nearly every technique we now rely on came from understanding why something else fell short. Testing our own ideas at least as hard as anyone else's is the point of publishing them. More about this programme and why we run it.

Most of what goes wrong in this kind of work is not the experiment; it is the comparison. These are the times our own checks overturned our own headline.

Where this stands

Repeatedly, the control rather than the measurement decided the result -- and several striking findings dissolved when the right comparison was finally run.

If you are evaluating whether someone would catch a problem in your models, this is the group to read. It is the most transferable output the project has.

Still open · 36 published results bear on this question.

What this does not settle yet

Stated plainly, because the gaps are as much a part of the record as the answers:

  • The programme still has no unlearnable-data control anywhere except one record (I6).

How these results fit together

Ten records that are all about this project checking its own work, and they are the most transferable thing here. Several overturned our own headlines. The largest is that a pattern we had described as the model expanding and then consolidating is substantially a *measurement* effect: we were re-fitting the coordinate frame at every step, and the apparent contraction is the frame rotating rather than the model unwinding. Any quantity compared across training steps now has to say which frame it was measured in. A companion record does the opposite service -- certain quantities depend only on eigenvalues, which are unchanged by rotation, so that critique cannot reach them, and one of those still peaks on the event. We also audited how many of our own results are close calls, found that eight runs is not enough for several of them, and published that. The habit that produced most of these: re-read the whole record after every new result, not just the item in front of you. The number of times the mistake was in the *analysis* rather than the experiment is the single most useful statistic on this site.

The results

Each of these is a self-contained record: what we asked, what would have proved us wrong, what we found, and what it does not show. They open with a plain-language summary before any of the technical detail.

Three Audits, No New Faults

No training. Eleven grids flagged; two known faults, four honest joint descriptions, five lookup tables. Three audits, zero new cases.

Read the record

Nobody Tuned the Rate

No training, just reading our own code: every speed-up claim was measured at an untuned rate, and the data that would have caught it had been in the archive for a month.

Read the record

The Design Carries Its Derivation

No training. Twenty-six scripts build a setting from a measurement; twenty-three do it fairly. The three faults are known, and two came from reusing a good design.

Read the record

The Check We Should Have Run First

We gave a model the same inputs with random answers, so there was nothing to learn. The effect we have been reporting all along vanished, which is what we hoped.

Read the record

Eight Runs Is Not Enough

The usual statistical test said we had found something. The range of values the data was consistent with said we had not. Repeating it on three times as many runs settled it, and the answer was yes after all.

Read the record

How Many of Our Results Are Close Calls

Sixteen of them are within one error bar of their own verdict. We are publishing the list rather than quietly rerunning the convenient half of it.

Read the record

We Read Our Own Warning List

Only two were things a conclusion depends on, and both had already been rechecked. The tool ranks careful studies as the shakiest, which is backwards.

Read the record

Consolidation Is a Measurement-Frame Effect

The same 41 runs show a number falling and rising over the same interval, depending only on whether the measuring frame is allowed to move. Nothing contracts.

Read the record

The One That Was Not A Ruler

A measure that cannot depend on which directions you call important rises to a peak exactly when the model learns, then falls back. So the change is real.

Read the record

The Fix That Fixed Nothing

The one dial we had made the task fainter rather than deeper, and faint is hard for everyone. We learned what the next attempt actually has to do.

Read the record

We Tested Someone Else's Prediction

Cutting the route made no difference at all. That rules out one of the three explanations the paper offers, and leaves two standing.

Read the record

The Experiment We Could Not Run

The control worked. The comparison arm never produced anything to compare, and reporting its result would have looked like a finding when it was evidence the task was wrong.

Read the record

The Archive as an Instrument

A pooled correlation of -0.679 that dissolved on stratification, and a scaling law that survived. Also the discovery that our archive varied only one of the four settings we thought it had.

Read the record

The Rule Caught It The Same Day

A real effect at two model sizes and nothing at the third. Notable as much for the process as the finding: the check that caught it was added as a standing rule the same day, off the back of the previous result.

Read the record

The Same Confound In Six Results

No experiment was run. Reading found that six results we published as effects fading with model size might instead be effects fading as the model outgrows the task, which is a very different claim.

Read the record

The Effect Never Faded

This overturns two of our own records from the same day and rewrites a rule we had just written. Once the task keeps pace with the model, an effect we reported as vanishing is flat across every size we tested.

Read the record

The Same Numbers, A Longer Run

The second confound found in the same small experiment, and neither was the one being controlled for. We withdrew the finding rather than the caveat.

Read the record

A Yardstick That Does Not Stretch

A measurement that gave different answers depending on when we stopped training, and two repaired versions that do not. Built and tested entirely from data we already had.

Read the record

The Window Was There All Along

The second of six results we are rechecking, and the second where the disappearance was our own measurement rather than the model's size. It restores the scope of our only genuine efficiency result.

Read the record

The Warning Gets Longer, Not Shorter

The third of six results rechecked, and the first where our own hypothesis failed. The earlier picture was distorted, but not in the way we expected, and the instrument comes out better than either version suggested.

Read the record

The Alarm Never Got To Speak

The last of six results rechecked. Five of the six turned out to be measurement artefacts, and the two instruments we had written off both give more warning in bigger models, not less.

Read the record

It Does Not Survive A Second Task

The counterweight to a good day. The same experiment that survived being made bigger does not survive being moved to another task, which scopes five of our results to one exercise.

Read the record

The Scoring Inflated The Imbalance

The conclusion survives the better measurement and the number we had been quoting does not. Both endpoints were computed on the same 100 runs so the old one is an exact anchor rather than a comparison.

Read the record

Fourteen Copies of One Definition

Zero of fourteen definitions drift by more than one measurement interval when a fifth of the run is discarded. The drift that does exist is confined to runs that had not finished improving, and the definition anchored to the task rather than the run does not drift at all.

Read the record

The Measuring Grid Moves the Answer

A prediction from one line of algebra, confirmed to 0.3 of a percentage point, plus a second axis nobody had looked at. Nothing published trips it, because an earlier result had already said not to.

Read the record

Close Calls Are Decided by the Code

A usable rule falls out: a comparison within one measurement interval of the boundary is decided by the code, not the data. Plus the discovery that eleven of fourteen versions cannot measure a warm-started model at all.

Read the record

Four Of Them

The kill test does not fire: 44.2% of surviving claims rest on one task against a 20% threshold. The first run of the audit answered confidently over an empty set.

Read the record

Where the Fault Was

No training. The kill test fires: zero further cases, because a house convention makes the correct thing the easy thing.

Read the record

The Shortcut Reversed

Our central efficiency negative survives being re-run properly, and gets stronger. The instrument got seven times cheaper relative to training while the thing it was buying went negative.

Read the record

One Number for Every Size

Both figures we could re-price moved by more than a factor of two, in opposite directions. The instrument cost three times more than published and our one positive saving is understated by nearly five.

Read the record

It Never Becomes Free

The cost of switching off keeps falling and never reaches free. And the boundary we had found was convergence wearing the transition's clothes: at the original setting the model was already 99.6% of the way to its final accuracy.

Read the record

The Shape Held, the Level Did Not

A growth rate is a property of the algorithm and a percentage is a property of the computer you ran it on. The same fifteen-minute pilot reproduced one to 0.05 and missed the other by fifteen points.

Read the record

Which Side of the Moment

Data-per-step decides nothing, which stands. The six-for-six pattern turned out to be a slope read as a step, and the correction is on the page.

Read the record

The Moment Marks Nothing

Our own quick test had given the right answer and we overrode it with a pattern that fitted every point we happened to have. That is the most likely thing a confounded grid produces.

Read the record

Where the Shortcut Stops Paying

Four sizes: +21.1%, +19.9%, +12.9%, then -42.0%. Monotone, so the test passes, but three points sit within eight of each other and the fourth falls fifty-five.

Read the record

The Cliff Survives Matching

The kill test does not fire and the prediction was wrong. It is also not a clean law: non-monotone, and the cell carrying the verdict is the worst-matched one.

Read the record

Back to all research questions

Want this measured on your data?

We build private models our clients own and run on their own infrastructure, and every engagement proves measured lift on the client's own tasks before we call it done. Start free with a readiness scorecard that tells you whether your data can support it, or book a short call.