What actually happens at the moment a model learns?
Ongoing research. This is an experimental result from active work, not a settled conclusion. The numbers are what we measured and the method is described so you can judge it, but the programme is still running and later experiments may revise what it means.
A note on the language in these records. This is a working laboratory notebook for research into training AI models more cheaply and efficiently, so you will read that an approach did not work, that a result did not hold up, or that one method was worse than another. That is the research doing its job, not a verdict on the engineering we deliver to clients. Ruling an approach out is how the search narrows, and these are the pages that teach us the most: nearly every technique we now rely on came from understanding why something else fell short. Testing our own ideas at least as hard as anyone else's is the point of publishing them. More about this programme and why we run it.
The jump from guessing to competent takes a few dozen steps. What is going on inside the model during those steps?
Where this stands
It builds machinery rather than selecting it, working through the task in a reproducible order and trying a simpler wrong rule on the way. The visible training curve cannot tell you which is happening.
This is where the commercially relevant finding lives: a model handed finished machinery shows the same training curve as one building it from nothing, so the curve alone cannot tell you what you paid for.
Still open · 14 published results bear on this question.
What this does not settle yet
Stated plainly, because the gaps are as much a part of the record as the answers:
- Why is the first-acquired part the most expensive to remove? The mechanism is unexplained.
- Everything here is one architecture family on one task family.
How these results fit together
Models on these tasks do not improve gradually. They sit near chance for a long time and then climb steeply, and these nine records are about what is inside that climb. The most useful finding is that it is not one event: what looks like a single jump is several acquisitions in a reproducible order, roughly sixteen steps apart. That ordering is stable enough across seeds to be a real property of the task rather than noise. Two records complicate it in ways worth knowing. First, some of that apparent separation is an artefact of scoring by accuracy, which is a winner-takes-all measurement -- scored in bits instead, the acquisitions are about half as separated, though the *order* is untouched. Second, a model handed finished internal machinery largely skips the reorganisation, while one handed a perfect external shortcut does not, which points at the climb being the construction of machinery rather than the discovery of an answer. The most striking negative here is that two models can have nearly identical improvement curves while one is building something and the other arrived with it. The curve alone cannot tell you which you are looking at.
The results
Each of these is a self-contained record: what we asked, what would have proved us wrong, what we found, and what it does not show. They open with a plain-language summary before any of the technical detail.
Skills Compete
We predicted a queue and found competition: take one unneeded skill away for a while and the rest learn faster.
Read the recordA Queue, Not a Graph
We tried to find which skill feeds which by holding one back. What came out looked like a queue: everything learned later waited, whether it needed the held-back skill or not.
Read the recordDepth Sets the Order
The longer rerun of the five-skill test: every skill learned, and on all six runs in the order the dependencies predict rather than the order difficulty predicts.
Read the recordDepth Costs Most of the Budget
Five planted skills: learned in the order of what depends on what, not of how hard each is alone. The deepest one mostly was not learned in time, so a longer test is next.
Read the recordOne Jump, or Several?
What looks like one sudden jump is four, about sixteen steps apart, ordered by which element of the pattern is being predicted. The order is reproducible at +0.983 across seeds.
Read the recordThe Training Curve Cannot Tell You
Give a model finished internal machinery and the reorganisation stops, though its accuracy still jumps. The visible curve looks the same either way.
Read the recordIt Tries The Wrong Rule First
The model's answers resemble one simple rule early and a different, correct one later. Its leftover mistakes are systematic rather than random.
Read the recordWe Gave It The Answer Anyway
Handing a model a correct shortcut changes nothing about how much it rearranges inside. Handing it finished internal machinery does. The difference says what the rearranging is.
Read the recordThe Scoring Made The Steps
Scored a second way with no right-or-wrong line in it, the stages of learning are about half as separated. The order they arrive in is unchanged.
Read the recordDoes the Transition Create Features, or Select Them?
A fixed random memory never reaches full accuracy at any width tested, and at a compute-matched budget improves 18.7 times more gradually. The sharp part of learning needs a memory that can change.
Read the recordLearning It Again Is Not Like Learning It
The same amount of learning, six times less internal upheaval. Which means the signature marks learning from scratch, not learning in general.
Read the recordThe Habit It Did Not Need
Removing the habit costs almost nothing, so it is not a stage the model needs. What each removal costs turns out to follow the structure of the task itself.
Read the recordThe Easiest Skill Is Learned Last
A recent study of large models found combined skills appear after their parts and read it as evidence of prerequisites. In ordinary tasks the combined skill is also the harder one, so we built a case where those two explanations disagree.
Read the recordA Dial, Not a Switch
Scramble a quarter of the labels, then half, then all of them. The effect shrinks smoothly with each step, and two of our measurement attempts failed first.
Read the recordBack to all research questions
Get new results as we publish them
Roughly monthly, one finding per email, in plain English first. Including the approaches that turned out not to work, which are usually the useful ones. No sales email.
Double opt-in: we send a confirmation link and add nobody who does not click it. One-click unsubscribe on every email.