AI has dominated tech conversations for years now, and LLMs have had the loudest microphones. So it was a welcome change of pace when something genuinely different joined the hype cycle: Jev, a model from TypeSafe AI.
This post is my attempt to figure out how it actually works. The short version: it’s a decision model — you give it state, ask typed questions about it, and get probabilities back.
It’s not an LLM — it doesn’t generate strings
An LLM takes your prompt and generates a response, one token at a time, as a stream of text. Your code then has to parse that text, hope it matches the shape you wanted, and hope the model didn’t invent a field along the way.
Jev works differently. TypeSafe calls it a System One model: built for fast, gut-check judgments — classify, route, score — the kind a knowledgeable person makes in a few seconds. It evaluates your state and returns typed answers with probabilities, nothing more.
The “System One” label comes from Daniel Kahneman’s Thinking, Fast and Slow, which splits human thinking into two systems. System 1 is automatic and intuitive — snap judgments, little effort, but can be wrong in predictable ways. System 2 is deliberate and analytical — long division, planning a budget, debugging — accurate when you engage it, but effortful and lazy by default.
TypeSafe borrowed the label for their model class. A System One model is the machine version of System 1: fast, focused judgments over state you provide, not slow open-ended chain-of-thought. Kahneman’s System 1 has a reputation for being error-prone; TypeSafe’s counter-claim is that theirs don’t have to be, because every answer ships with calibrated probabilities so your code knows when the gut check is trustworthy.
The deeper difference is who each was built for. LLMs were tuned with RLHF to produce words people want to read. Jev was built to work with machines: TypeSafe calls it machine-native intelligence, “natively used by machines.”
Why “Jev”? It’s named after William Stanley Jevons, the 19th-century English economist behind Jevons paradox: when a resource gets cheaper to use, we end up using more of it. More efficient steam engines didn’t save coal — they burned more of it. TypeSafe expects cheaper machine intelligence to work the same way.
State in, questions out
Every Jev request has exactly two parts:
- State — the material to judge: a string, or structured data like JSON. For a ticket-triage example, that’s the email’s subject and body.
- Questions — the typed judgments you want made about that state. You define the possible answers in advance; Jev only fills them in.
Jev evaluates every question in one parallel pass. Here’s a complete request:
POST /v1/decide
{
"model": "typesafe-ai/jev",
"state": {
"subject": "Production down — 500s on checkout",
"body": "Since 14:02 UTC, checkout returns 500. Blocking all purchases."
},
"questions": {
"department": {
"type": "choice",
"instructions": "Which department handles this message?",
"criteria": {
"billing": "payments, invoices, refunds, pricing",
"technical": "bug, error, outage, crash",
"sales": "plan comparison, purchase intent",
"support": "how-to, account help",
"other": "unclear or none of the above"
}
},
"is_urgent": {
"type": "noul",
"instructions": "Is this urgent, time-sensitive, or blocking?",
"criteria": {
"true": "outage, down, blocked, critical, ASAP",
"false": "routine, no time pressure"
}
}
}
}{
"answers": {
"department": {
"type": "choice",
"choice": "technical",
"probabilities": {
"billing": 0.01,
"technical": 0.91,
"sales": 0.01,
"support": 0.05,
"other": 0.02
},
"confidence": 0.91
},
"is_urgent": {
"type": "noul",
"noul": 0.94
}
},
"usage": {
"input_tokens": 286,
"output_tokens": 64
}
}Because questions are independent and evaluated in parallel, the useful habit is decomposition: ask separately about department, urgency, and violation type, then combine the results in code — rather than asking one giant question and hoping for a clean answer.
Why it’s so fast
LLMs sample sequentially: token 1, then token 2, each conditioned on the last. Jev’s parallel sampler produces all of its outputs in a single query — it never writes a token stream at all.
TypeSafe reports end-to-end response times of roughly 70–500ms, versus multi-second latencies for frontier LLMs on comparable tasks — about 40×–200× faster in their measurements, and dramatically cheaper ($0.042 per million input tokens; output is free, “too cheap to meter”).
Adding more questions barely moves the response time, because they all run in the same pass. You can ask a dozen questions in one round trip for roughly the price of one.
Confidence you can branch on
Reinforcement Learning for Calibrated Decisions (RLCD) is the method TypeSafe uses to train Jev: probabilities are optimized against outcomes, not human rater preference for nice-sounding prose.
| RLHF · preference | RLCD · outcomes | |
|---|---|---|
| Optimizes for | what human raters prefer to read | probabilities that match real outcomes |
| Good at | chat, instruction-following | decisions your code can branch on |
| Weakness | prone to overconfidence, mode-dropping | narrower scope, no generation |
RLHF is the standard way chat models are tuned: humans rank or rate model outputs, and the model is optimized to produce text people prefer — helpful, fluent, well-formatted. It can then sound sure even when it shouldn’t be.
RLCD instead rewards a model when its stated probabilities line up with how often it is actually right — honest odds your code can branch on.
“Calibrated” has a precise meaning: if Jev says 90% confident on 100 similar questions, it should be right about 90 of them. That says nothing certain about any single answer, but over many decisions the numbers are trustworthy.
That precision is the whole point. Every answer comes with a confidence value, and when Jev says it’s 90% confident it should be right about 90% of the time. That’s what makes rules like “if confidence > X, auto-act; else escalate” meaningful.
Worth being precise about one thing: confidence is not the literal probability that the answer is correct. It is derived from the shape of the answer’s probability distribution. In practice, low-confidence answers fall back to “other” or get routed to a human. The model provides the signal; your application code defines the policy.
What “can’t hallucinate” means
Because the answer space is defined in advance, Jev cannot return a value outside your schema — no type errors, no invented enum values, no walls of prose where a field should be. That’s a structural guarantee: TypeSafe’s claim is that schema matching is mathematically impossible to violate, so their type-error rate is 0%.
It is not a claim that every judgment is semantically correct. A well-formed answer can still be wrong, so you still measure accuracy on your own data. “Type-safe” and “always right” are different promises.
What Jev doesn’t do: generate text, write code, fill tool arguments, or explain its reasoning. For those jobs you use an LLM. The intended pairing is complementary — Jev handles the frequent, narrow decisions along the way (classify, route, score, verify), and a language model handles open-ended reasoning and generation.
Jev isn’t the only one
Jev is TypeSafe’s first System One model, but similar decision models exist — for example Laya by ConvAI Innovations, which is open-weights.