Swapping Haiku for Jev in my AI twin
In part 4 I said speed was what bugged me most. Every message started at the gate (the step that sorts your question, see part 1), and the gate was a Haiku call that took almost 2 seconds (1.75 at the median on my tests). That's 2 seconds just to decide what kind of question you asked. Even a simple "hi" paid that price, and so did every refusal.
Then Jev showed up.
What Jev is
TypeSafe released Jev in mid-September. It's a different kind of model: it doesn't write anything. You give it some context and a question with a list of options, and it tells you which option it picks and how sure it is. TypeSafe advertises answers in 70 to 500 ms. Picking a bucket is exactly that kind of question, so I had to try it.
First surprise: Jev was still invite-only on TypeSafe's side. Luckily it's also available on OpenRouter, so that's how the twin calls it.
How I set it up
One call to Jev asks three questions at once:
- Which bucket, and how sure? Each bucket comes with a short written description, tricky cases included. For example, "where is he based?" counts as a work question, but "how old are you?" counts as personal, even after a friendly chat about Kubernetes.
- Is this message trying to trick the twin? A plain yes or no, asked on its own. If Jev leans yes (0.5 or more), the message gets refused, whatever bucket it picked. A separate yes or no is harder to talk around than one option in a list of eight.
- Is this a full English question on its own? If yes, it's searched as is. If not, Haiku rewrites it first.
Then a few rules around it. If Jev is less than 60% sure about the bucket, Haiku decides instead. For small talk I ask for 80%, because small-talk replies are the only ones the judge doesn't check. And if Jev is down, slow (more than 1.5 seconds), or sends back something weird, Haiku takes over too. You wouldn't notice. Only my logs would.
Can it guard the door?
Jev's own documentation says it isn't built to resist messages written to fool it. That's a fair warning for something that sorts out who gets refused. So I looked at what happens if a sneaky message gets through anyway:
- If it lands as a work question, the answer is still written only from my notes and checked by Sonnet 5.
- If it lands as a message or booking request, the tools still can't send anything without your click (the rule from part 1, and the reason part 2 ends well).
- Personal and off-topic questions just get a refusal.
Jev only moves the sorting. Everything that keeps the twin safe after that stays the same.
What broke
Jev didn't know my homelab was work. In the first Jev tests, two questions kept landing in the wrong bucket: "Can you actually run LLMs on a Raspberry Pi?" and "What's your take on eBPF-based observability versus sidecar agents?" Jev called both off-topic, and it was 80 to 89% sure about it. That's above the bar, so Haiku never got a say. The fix was two lines in the description of work questions: my take on topics in my field, and what works in my homelab, both count. After that, every test question landed in the right bucket.
The fake Mouhamad, round two. Remember the visitors typing my name in part 2? That was the closest call with Jev. It didn't flag the fake name as a trick, but it wasn't sure which bucket it belonged to either (31% sure), so Haiku took over and refused it. And if both had missed it, the name check would still have said no.
The career question. Once Jev got every bucket right, the only failure left in my tests was "Walk me through your career so far". That one wasn't Jev's fault at all, it was the search. The full story is in part 3.
The numbers
On my 77 test questions, Haiku put 97.4% of them in the right bucket, in 1.75 seconds at the median. Jev got 100% on three runs in a row, in about 0.3 seconds. Every refusal held.
Most messages never touch Haiku before the search anymore. The ones that do (follow-ups, French questions, and the rare case where Jev hesitates) take around 1.2 to 1.7 seconds, which is still faster than the old gate. French went 12 out of 12, on every run. And the cost is about $0.00005 per message, so roughly 5 cents for a thousand messages.
What you actually feel is the full answer time, so I timed all 77 questions end to end, once with each gate:
Refusals went from 1.9 seconds to 0.3, about six times faster. Small talk went from 3.5 to 1.9, and a typical work answer from 5.3 to 4.1. The draft and the judge are still there, so work answers don't get the full boost, but more than a second is more than a second.
Trade-offs
A gate that can't write. Jev only picks, so Haiku stays around for the rewrites and for the cases where Jev hesitates. Two models for one step is a bit more to maintain, but most messages only pay for the fast one.
Two more companies in the loop. With Jev, your messages also go through OpenRouter and TypeSafe, on top of Anthropic, LangSmith and Railway. The chat's privacy note lists all of them.
A brand new model. I pin the exact version of Jev. A new release could sort things differently, and I'd rather rerun my tests before that happens than find out from a visitor.
What I don't know yet
I wrote the bucket descriptions while looking at these same test questions, so take the 100% with a grain of salt. Real visitors will be the real test. I'll be watching how often Haiku has to step in, and the thumbs down. And if Jev starts acting up, going back to Haiku is one setting away.
That's the end of the series, for now. The twin will keep changing, and so will these notes. If you want to see how it all holds up, go ask it something. And if it says something dumb, you know where the thumbs down button is.