When my AI twin gets it wrong
Part 2 was about people trying to break the twin on purpose. This one is about the twin breaking all by itself. No attacker needed, just me, my code and my notes.
Same order as last time, from the most serious to the silliest.
Long chats would have crashed
To keep things short, the chat only sends the last 12 messages. But cutting there could make the conversation start with a twin reply instead of yours, and the model API refuses a conversation that doesn't start with the user.
In practice, every long chat would have died around the seventh question. I caught this one in review before any visitor hit it. Now the server drops those leading twin replies before calling the model.
The judge that crashed
When I moved the judge to Sonnet 5, it sometimes answered in a slightly different shape than the one my code expected: the same verdict, wrapped in one extra layer. My code couldn't read it, so the whole request failed, and you would have seen the chat break instead of an answer.
I caught it while running the tests, before switching. Now the model is forced to answer in the exact format my code expects (structured outputs, for the curious), and if the judge's answer still can't be read, the draft counts as failed. If the judge breaks, I want it to say no, not let things through.
French questions got nothing
Someone asked "Où es-tu basé ?" and got a vague non-answer. Remember the English-only encoder from part 1? The French question missed the note that says I live in Paris. Then the judge did its job and blocked a draft that wasn't backed by anything.
So the judge was fine, the search was the problem, and a judge can refuse a bad answer but it can't find what the search missed. I also had no idea what was going on until I started logging the rewritten question and the search scores for every message. The fix was the English translation in the gate, plus French questions in my tests.
My career? "I garbled that one."
"Walk me through your career so far" is about the most obvious thing to ask, and it kept failing. Once the gate got every bucket right (thanks to Jev, more on that in part 5), it was the last failure left in my tests, so it got hard to ignore.
The search was bringing back the wrong notes: bits of my blog posts, and even the file that sets the twin's voice, ranked above my actual career. The draft came out thin, and the judge blocked it. Cleaning up the search wasn't enough (I stopped searching the voice file and merged tiny sections together). What fixed it was a short "career so far" section in my notes, with my four jobs in one paragraph.
The search finds notes that look like the question. When the answer is spread over four sections, one per job, none of them looks enough like "my career so far".
Who is "I"?
My notes are written in my own voice, so early answers said things like "I work at Chainguard", as if the bot were me. A bit creepy. So I set a rule: the twin says "I" only about itself, and talks about me as "he".
Then the judge started arguing with itself. One run it flagged a draft for using "I", the next run for not using it, and once it even blocked a correct answer. It turned out the writer and the judge each had their own description of the voice, worded a bit differently. One shared voice rule, copied into both prompts, fixed it.
My judge is a bit of a stickler
At first, it failed any answer that left something out. So an answer listing some of my certifications (correct, just incomplete) turned into "I garbled that one". Now there's a "partial" flag that doesn't block anything. You get the answer, and I get a note in my logs telling me which topics are hard to cover.
It's still strict though. While running my tests, "Que fais-tu chez Chainguard ?" got blocked twice, because the draft described my job as "technical advice, not pure sales". Not a crazy thing to say, but it's not in my notes, so Sonnet 5 said no. Fair enough. If I want the twin to say it, I have to write it down first.
Half a list
A broad question like "What certifications do you have?" came back with only my Kubernetes ones. My certifications page had one section per family, and the search only grabbed the cloud-native one. Cutting notes by heading works well most of the time, just not when someone wants the whole list. That page is now kept in one piece, and the twin is told to cover every category when someone asks for a list.
My twin turned into my website
When I first put the twin online, its health check kept failing, and the logs said "Serving static site". My hosting was building both services from the same config file, so for a while the twin was a second copy of my homepage, which is not great at answering questions. Deleting that file and giving each service its own settings fixed it.
What these have in common
A lot of these only showed up because I had a question in my tests that poked at them: a French one, a long one, a broad one, the obvious career one. And when a test failed, the logs had to show me each step (the rewritten question, what the search found, what the judge said) before I could tell which part was actually wrong.
Which brings me to the tests. Part 4 is about how they work.