Private & owned AI, explained
Plain-English writing for leaders deciding how to bring AI into a business the right way: owned, private, and built on your own data.
We Tested Five Ideas From This Week's AI News
Five training ideas made the news this week. We tested each on small models in a day. Three worked, one lost to a fixed plan, and the best was not the obvious one.
Read the articleA Model Can Ace Its Own Test and Still Slow You Down
One model in our pool scored a perfect 1.000 on the task it was trained for, and building on it made a new model slower than starting from scratch.
Read the articleThe Number Was Right. The Advice Was Wrong.
We tested one of our own cost estimates on a task it had never seen. It held to within ten percent. The rule of thumb we had written underneath it cost double when we followed it.
Read the articleWhat 48 Fine-Tuning Experiments Showed
Six open-weight models, three task types, three random seeds each. Every one of the 48 runs improved. The full numbers, the two predictions we got wrong, and what the measurements do not cover.
Read the articleWhat Is Actually In Your Fine-Tuned Model?
Software has had a bill of materials for years and models have not. A real one from a real training run, field by field, including the field where the fine-tune lost ground to the model it started from.
Read the articleSix Results of Our Own That Did Not Survive Checking
An 88-point improvement that turned out to be our own file format. A 45-point gain that reversed to a 9-point loss. Six measurements we made, believed, and then checked properly, with the check that now catches each one.
Read the articleWhat 510 Real Contracts Taught Us About Measuring Extraction
Why we exclude the contract fields buyers most want reported, why exact match is right for a date and meaningless for a 317-character clause, and two pipeline defects that only appeared on the real corpus.
Read the articleYour Base Model Is Better Than You Think
We sell fine-tuning, and this is the case for measuring your untuned model first. It already scores 78 to 84% on eight-way classification where guessing scores 12.5%. Four situations where we would tell you not to fine-tune.
Read the articleChoosing a Base Model Matters Less Than You Think
After training six models on the same contracts, the gap between best and worst narrowed from 15.0 points to 4.5. What should decide your choice instead, and the counter-example that keeps it from being a law.
Read the articleHow We Prove Your Model Is Actually Better (Not Just Different)
Ask most AI vendors how they know their model is better and you get a demo. Here's the method that produces a number instead, with real results on 510 lawyer-annotated contracts (62% → 85%), how much data it actually took, and the two results we measured, caught, and threw away.
Read the articleOwn, Don't Rent: The Case for Private AI in Regulated Industries
The fastest way to get AI into your business is to rent it, and it may be the most expensive decision you make this decade. Why owning a private, fine-tuned model beats renting a cloud API on privacy, cost, control, and risk.
Read the articleWhat Air-Gapped AI Actually Means (and When You Need It)
"Air-gapped" gets used loosely, and the loose version is exactly what a security team can't accept. A precise read on what true isolation requires, who needs it, and why you can only get it from a model you own.
Read the articleWhere Your AI Comes From: Model Provenance & Country of Origin
For defense, government, and regulated buyers, who built the model, and where, is part of the risk calculus. Why provenance belongs on your procurement checklist, separate from license and from data safety.
Read the articleIs Your Data Ready for AI?
Most organisations have more usable training data than they think, in worse shape than they hope. What "AI-ready" actually means, the exact thresholds we score against, and how to check yours without sending it anywhere.
Read the articleWhen Your Metric Moves With What It Measures
The same forty-one training runs showed a number going down and going up at the same time. The difference was not the data. It was whether the measuring stick was allowed to move, and why that matters for any quality metric you are sold.
Read the articleA Strong Correlation That Meant Nothing
A correlation of -0.68 across forty-one experiments looked like a finding. Splitting the data three ways dissolved it completely, and two of the three pieces pointed the opposite way. The most common way a benchmark misleads.
Read the articleWhen a Benchmark Is Too Easy to Be Useful
Every quality check in a year of our own experiments passed. Then we noticed the test was scored at 99%, which meant nothing could have failed it. What we built to replace it, and how to spot the problem in a vendor's numbers.
Read the articleWe Tried to Kill Our Own Finding
Our main research result had only ever been measured with one training algorithm. We changed it to see whether the result would survive. Half of it did, and half of it turned out to belong to the algorithm.
Read the articleThe Answer Was There Before the Model Could Say It
A model knew the answer about 23 training steps before it could produce one. We nearly missed it, because the first version of the experiment used a measuring tool too weak to see it, and gave a clean wrong answer instead.
Read the articleWhat a Model Has to Build, and What It Gets for Free
We froze the part of a model that does the remembering and trained only the two ends. It still learned, slowly and partially, and the difference turned out to be about shape rather than score.
Read the articleThe Experiment That Cancelled Our Next Six Months
We had a whole line of research planned. One cheap experiment, designed specifically to kill it, killed it in an afternoon. That is the experiment working, not failing.
Read the articleChecking Whether Our Result Was Only True of One Model
Everything we had found came from one model design on one task. So we built four more and ran the whole grid. Two thirds of it held up, and the part that did not is the interesting part.
Read the articleOur Two Findings Turned Out to Be One
We had two separate results about how models learn. Checking whether they held in other designs revealed they were the same result wearing two hats, and both were narrower than we thought.
Read the articleWhat Moves Is Not What Matters
We measured which part of a model changes most while it learns, then held each part still to see which one it actually needs. They were different parts, which is a warning about a whole style of analysis.
Read the articleDoes Randomness Help AI Learn?
There is a popular theory that AI models learn suddenly because random noise knocks them out of a rut. We added noise deliberately to find out. It made things worse, and the experiment that seemed to show otherwise was measuring something else.
Read the articleHow to Tell a Real Early Warning From a Fake One
We found a signal that fired 74 steps before an AI model learned a task, in every single run. It was worthless. A month later we found one worth having. The difference came down to one cheap test that almost nobody runs.
Read the articleAn Average Is Not a Promise
We reported a warning signal that worked in 49 runs out of 49. Two days later we showed that number was about the runs we happened to measure, not about the method. Here is how we caught it, and why the distinction matters outside our lab.
Read the articleAI Models Have Critical Periods Too
One part of a model has a forty-step window in which it has to be working. Interrupt it then and learning takes twice as long. Interrupt it forty steps later and nothing happens at all.
Read the articleThe Price We Quoted for Our Own Measurement
We priced our own monitoring in a unit that cancelled out of its own ratio, then quoted that number in eight results. Recounting moved two figures by more than a factor of two, in opposite directions.
Read the articleFourteen Versions of One Number
We went looking for one bad measurement and found fourteen copies of it, differing in ways nobody had decided. Then we made them all answer the same questions.
Read the articleThe Day Our Results Turned Out to Be Our Own Mistake
We had six results saying the same thing. They agreed with each other because every one of them was produced the same wrong way, and it took reading somebody else's paper to notice.
Read the articleThe Control That Cost Us Our Result
We found that some parts of a model repay training effort sixteen times better than others. Two controls then showed that training had nothing to do with it. Here is what each one caught.
Read the articleThree Buckets That Lied to Us
We grouped our data into before, during and after. The three groups came out cleanly different with error bars that did not overlap. The difference was entirely an artifact of the grouping.
Read the articleSeveral of these are the plain-language versions of studies from our open research programme. The full records, with the designs, the tables of numbers and the limits, are in the research section.
Wondering if your data is AI-ready?
Start with a free readiness scorecard you run yourself. We never see your data.
Get the free scorecard