
Monday night, September 1. I sat down to teach a small model how I review my own work.
Not a chatbot. A model that lives on a machine in the house, trained on hundreds of real calls I have already made: ship this, pull that, the copy is wrong and the practice is fine, freeze the destination before you send. The kind of notes you leave yourself after a long night so you do not pay the same tax twice.
I held twelve cases back on purpose. The model never saw them during training. Then I asked it those twelve, one after another, and watched the first line of each answer land.
Eleven of the twelve opened with a verdict word. The costume was perfect.
The judgment was not.
What looking like a review was hiding
On one held-back case, the real ruling was simple. A setting I thought I had frozen at the moment of send was not actually frozen. Change it a second later and the work goes to the wrong place. That is a stop. Do not ship.
The model said the opposite. The change would not matter. The request was already on its way.
Same shape as a review. Opposite call. I sat there for a second and did the thing the notes exist to prevent: I almost believed the first line because it sounded like me.
On another, it restated the question back to me, as if repeating the prompt were the answer. On a third, it wrote a balanced list of reasons for and against pulling a known-broken build from Apple's queue. Pros, cons, it depends on the severity, it depends on the window. The real ruling had no list. Pull it. Nothing broken ships, even for a day. I have already paid for that sentence. The model had not.
I have a name for this now. style without judgment. The format of a decision, with the decision missing.
I thought more training was the work
The obvious move, at midnight, is to give it more of the same notes and run it longer. I did that before morning.
The error number dropped, then sat still. More passes over the same examples did not buy a better call. I had been watching a score that had already told me the truth, and I ran it anyway because stopping felt like quitting one hour too early.
I held twelve cases back again and asked.
On one, it finally named the real cause of a bug three separate reviewers had already found. Right cause. A fix you could actually do. That is the part I would like to stop the story on.
On another, it got the call backwards again. A scan of a historical text had two corrupted lines. The real ruling was restore the standard public-domain reading, and keep the garbled scan in the audit trail so the correction stays checkable. The model said leave the corruption in. Stay true to the source. Which sounds like principle, and is the opposite of the principle I had already written down.
That is the expensive version of this mistake. Not a bad number. A good-sounding rule applied to the wrong object.
What the examples were missing
The notes I trained on were verdicts. Ship. Pull. Fix the copy, not the practice. They were written for me, at the end of a session, when I already knew why. They were short on the why because I did not need it. I had been in the room.
A verdict without the why is a costume. Give a model enough of those and it copies the first line. It does not know what the line is for.
I already knew this about people. A junior who has watched you make a hundred calls will copy your first sentence and miss the thing that made the call. I had not applied it to a model I was training on my own desk. I had assumed that if the examples were mine, the judgment would come along for free.
It does not.
The model is parked. Nothing new is in the live stack. The tool I already use for reviews is unchanged, and I still read what it returns. The experiment ran. That is the whole claim.
The rule
If you are going to teach a tool your taste, the examples have to carry the reason, not only the answer. The night you made the call is not in the file unless you put it there. Then you hold some back on purpose, and you read what it does with a case it has never seen. Not a demo. A case you would actually have to live with.
Do not trust the first line because it sounds like you. That is the whole trap. I almost walked into it on my own notes, on my own machine, on a night I thought I was teaching it to be careful.
A verdict word is not a verdict. Go look at the call.
If the examples do not carry why, the model learns the costume.
Have you ever taught a tool your notes and then checked whether it learned the call, or just the costume? Hit reply. I read everything.
Christopher
Wisdom Nexus is written the way it is lived, with AI in the loop and a person at the desk who cares how it lands. This is where the theory meets the practice, shared with you in the open.
P.S. The scoreboard is still mostly me. One of the apps got a real update onto the store this week, because a backup import was quietly dropping people. The machine that moves the scoreboard got better. The model I trained to review that work did not.