TL;DR
Most of what we call an agent’s “thinking” is a multiple-choice question: which tool, which team, is this safe, is this a regression. We’ve been answering those with frontier LLMs, which are generative and slow, and I learned the cost in wall-clock seconds, not tokens. Jev from TypeSafe AI is a System 1 model. It takes state in and returns typed, calibrated probabilities, not prose. The idea is right, and the calibration matters more than the speed. But a decision layer you can’t run yourself is a vendor, not infrastructure. For regulated and PII-heavy workloads, the hosted API is a “not yet” until you can self-host an equivalent.
Originally published at portfolio.hagzag.com.
Too Much Noise Not to Try It
Within weeks of TypeSafe coming out of stealth, my feeds were full of Jev: launch threads, “System 1” explainers, and open-source clones fighting over the same name. At some point the noise got loud enough that I had to try it myself, so I opened the playground. My first question wasn’t “how fast is it?” but “where would I actually use this?” The obvious place was my current work in compliance. Much of a control review is really multiple-choice: does this evidence satisfy the control, is this exception in scope, which risk tier does this finding belong to. Typed, calibrated answers to those questions could guide the decisions and tell a human where to look first. Whether an auditor will accept “the model said 0.93” as part of the evidence trail is another matter, and I’m honestly not sure they will. Only time will tell if yes or no, one clear place I thouight I could use this is in a model routert - this is something which is definetelly choice driven.
The Router That Made the App Feel Broken
The router wasn’t expensive. I checked. The token bill for routing was a rounding error next to the actual work.
What it cost was time. Every request went through a frontier model with one job: read the user’s message, pick one of a handful of downstream paths, and return clean JSON. It did that well, in a few seconds. Those few seconds were enough for the app to feel like it had hung. I’m the kind of user who taps the button again after a second and a half, so I didn’t need a UX study to know it was wrong.
Looking back, the mistake was in what kind of question it was, not in the prompt. “Which of these six paths?” isn’t a reasoning problem. It’s a classification problem, and I was paying for a model that writes essays to answer it.

Most of Your Agent’s Thinking Is a Multiple-Choice Question
Jev is TypeSafe AI’s answer to that mismatch. It’s named after the economist William Stanley Jevons and it doesn’t generate text at all. You give it state (text or JSON) and a set of typed questions, and it returns probability distributions. The playground shows the whole surface area in three primitives:
- Choice: pick one of N options (up to 255).
- Score: grade against a rubric with 2–10 levels.
- Noul: how true is this statement, as a single probability. (Not a typo; it’s their name for the yes/no type.)
Playing with the online playground, what stood out wasn’t the speed. It was the examples TypeSafe chose to lead with: LLM guardrails, support agent audit, résumé screening. None of them are about generating anything. They’re all about judging something an LLM or a human produced. That’s the right framing.
The vendor numbers: 70–500 ms per call, $0.042 per million input tokens, and output is free because nothing is generated. Those are TypeSafe’s figures, not mine. But they match the thesis of my AI cost post, where the real lever is deciding which tasks deserve an agent at all. A decision layer that answers in milliseconds for fractions of a cent is that lever.
# Illustrative only: the shape mirrors the playground (state + typed questions).
# Check TypeSafe's docs for exact field names before copying.
decision = jev.decide(
state={"message": user_message, "plan": customer_plan},
questions={
"route": {
"type": "choice",
"question": "Which path should handle this request?",
"options": ["billing", "infra", "security", "product", "smalltalk"],
},
"is_outage": {"type": "noul", "question": "Is the user reporting a production outage?"},
},
)
Look at what’s missing: no system prompt begging for JSON, no Pydantic model, no retry when the model wraps the JSON in a Markdown fence. The schema is enforced by the model’s own interface, not by my parsing code.
A Confident Witness Is Not a Judge
The speed gets the headlines. Calibration is what changes the architecture.
In Probability Pivot, Part 5 I wired an LLM-as-a-Judge into the metrics pipeline: sample traces, send them to an evaluator with a rubric, emit the score as a gauge. It works, but there’s an awkward truth underneath. An LLM that says “confidence: 0.9” is a confident witness. Nothing ties that 0.9 to how often it’s actually right. A 0.9 that’s right 60% of the time is worse than no score, because you’ll build thresholds on it.
Jev’s claim is that its probabilities are calibrated: 85% means right about 85% of the time. TypeSafe trains for that with proper scoring rules like the Brier score rather than human preference ratings. That claim is what makes threshold-based routing safe:
route = decision["route"]
if route.confidence >= 0.95:
dispatch(route.answer) # fast path, no LLM in the loop
elif route.confidence >= 0.70:
dispatch_with_llm_review(route) # System 2 only for the gray zone
else:
escalate_to_human(user_message, decision) # don't guess in the dark
That’s the pattern my slow router should have been: a System 1 reflex that handles most traffic in milliseconds and calls the expensive System 2 model only when the reflex isn’t sure. It’s the same shape as the evaluate step in Loop Engineering. If the evaluator is a frontier model, the loop is slow and expensive by design.
So far there’s one independent data point, and it’s encouraging: on 662 prompt-injection messages, the jev-sec-bench benchmark measured an expected calibration error of 0.0588 (via Victor Dibia). One benchmark isn’t a track record. It is more than most LLM judges have.

Where It Would Have Fit in What I’ve Already Built
Looking back through the last year of posts, the System 1 slot was empty in several places:
- The LLM Gateway: I described the gateway routing simple requests to cheap models “based on the tags or the content”. Something has to read that content. A Choice question (“which model tier does this request need?”) is that something: smarter than tag rules, and much cheaper than an LLM router.
- The Agentic Highway: a Noul check (“is this tool argument a prompt-injection attempt?”) before the tool executes. The returned probability is exactly the kind of field the Black Box audit trail should record next to which agent and which task.
- Probability Pivot: swap the LLM judge for a Score rubric and the quality gauge becomes cheap enough to run on every trace, not a 5% sample.
What Would Break: A Vendor Is Not Infrastructure
Here’s where I get off the hype train.
It’s a hosted API, and for some of my clients that ends the conversation. Jev runs only as TypeSafe’s managed service: no weights, no VPC deployment, waitlisted keys. Every payload leaves your environment. In healthcare, where those payloads are logs and messages that can contain PII or PHI, that’s a “not yet” for me. It’s the same boundary argument I made in FedRAMP, Part 4: a decision layer that crosses your authorization boundary for every request isn’t a building block. It’s a subprocessor.
Self-hosting changes the answer, but read the label. The pattern matters more than the vendor, and open alternatives showed up within weeks. They come in two very different kinds.
- Wrappers like OpenJev keep the interface: the same
state+ typedquestionsrequest, Choice/Score/Noul, and a probability per option. Under the hood they score each option with an ordinary chat LLM you bring (OpenAI, Groq, Anthropic, or any OpenAI-compatible endpoint, including one you host) and normalize the results. That’s how I’m prototyping the pattern with clients: same contract, our own model, our own boundary. But be honest about what you get. Latency is still LLM latency (one call per option, run in parallel), and the confidence is a normalized score, not a trained calibration. - Trained decision models like RSI-Jev go after the model itself: a 4B Qwen-based model with decision heads, open weights, and the project claims TypeSafe’s wire format. That’s the kind of thing you’d put on the hot path, after you’ve measured its calibration on your data.
Either way, run the decision layer inside your own cluster (pointed at a model you serve, not a SaaS endpoint) and the boundary objection goes away. Then you inherit the operational work of any model you serve: GPU capacity, versioning, and drift.
It’s not Redis, and it’s not eBPF. The research I started from compared Jev to in-memory lookups and kernel packet filters. That’s wrong by about three orders of magnitude. 70–500 ms is a network call. The win is relative to a multi-second LLM, not relative to your existing fast path. Don’t put it in your hot path assuming microseconds.
The numbers are mostly the vendor’s. Apart from that one calibration benchmark, the speed and cost multipliers come from TypeSafe. They also say Jev is weak at counting, arithmetic, and anything needing multi-step reasoning, and that’s by design. If your “routing” question quietly needs a calculation, you’re back to System 2.
Conclusion
The lesson from my slow router wasn’t “use a faster model.” It was “stop asking a writer to do a clerk’s job.” Split your agent’s work into reflexes and reasoning, give reflexes a typed, calibrated, millisecond decision layer, and keep the frontier model for the gray zone where it earns its latency.
Then run that layer yourself. The cheapest token is the one you never generate. The safest one never leaves your cluster.
Discussion