Back to all stories
Operational Chaos
🟡 Inspired by Real Events

Jev and the Agentic Stack: A Primitive You Can Actually Test

Typed Noul, Score, and Choice decisions with calibrated confidence product and ops can threshold, audit, and test in production

2026-09-28·6 min read·By Alain Prasquier
Jev and the Agentic Stack: A Primitive You Can Actually Test

In our enterprise software systems, most of the work is not writing paragraphs. It is making the same kinds of decisions, thousands of times a day, with enough confidence that software can act. Agentic systems are no different.

Is this a complaint? How satisfied is this customer? Which team should own this email? Which tool or sub-agent should this be passed to?

Large language models can answer those questions. They answer them in language. Your code then has to interpret the language. That is where flaky agents live: brittle parsers, prompt drift, and unit tests that assert on vibes. They are slow, expensive and we have no control over their reliability.

Jev, TypeSafe AI's first System One model, is built for a different job. It does not return chat text. It returns a typed decision your program can branch on directly — with a calibrated confidence score you can threshold, log, and test.

That score is the reliability and transparency unlock. For years, "how much can we trust this model — can we trust this answer?" was a gut feel. With calibrated confidence, it becomes an explicit number in code: raise the bar, lower it, escalate below it. Finally, a simple way to decide how much to trust the model — not vibes, thresholds.

That semi-deterministic shape is the point. It is how you turn agent logic into components you can actually test.

What Jev is

The "System One" label is not ours. It comes from Daniel Kahneman's Thinking, Fast and Slow (2011): System 1 is fast, intuitive judgment; System 2 is slow, deliberate reasoning. TypeSafe AI brands Jev as a System One model on that metaphor — built for snap decisions software needs in-line (classify, score, gate), not for the fluent System-2-style prose of frontier chat models.

You send context (a ticket, a message, a state object) and a map of questions. Jev evaluates them and returns structured answers. The output shape is fixed by the question type you declared. Your agent does not have to hope the model "remembered" the JSON schema.

Every question is one of three primitives. Each comes back with a decision and a confidence you can act on.

The three outputs

Noul — calibrated yes / no

Noul returns a probability between 0 and 1 for a yes/no question — the confidence is the answer shape.

Example: Is this customer message a complaint or not?

Score — ordered judgment

Score places the input on an ordered scale (a small number of levels you define) and returns a position — often fractional — plus a distribution you can read as confidence across the scale.

Example: Sort these messages by customer satisfaction.

Choice — pick one labelled path

Choice selects one option from a labelled set, with per-option probabilities and an overall confidence.

Example: Classify these emails by type — customer success, billing, and so on.

Together, Noul, Score, and Choice cover most of the closed decisions agents make before they ever need to write anything. The common thread is not the labels — it is that every path carries a number you can trust enough to automate, or distrust enough to escalate.

Why this matters for agentic systems

In production, agent stacks fail less often on the dramatic "write a plan" step and more often on the control-plane edges: filters, routers, escalation gates, and policy checks that have to run on every ticket, every message, every handoff — under SLA, under budget, and under scrutiny.

When those edges are LLM prose, enterprises inherit risk they cannot operationalize:

  • Outputs that look fine until a new phrasing breaks a parser buried in a workflow
  • Regression tests that snapshot English instead of asserting on typed decisions
  • No durable place for product and operations to own a confidence threshold — only prompt tweaks and tribal knowledge
  • Cost and latency that do not scale when the same closed decision fires thousands of times an hour

When those edges are typed primitives with calibrated confidence, you get something closer to a supervised control plane:

  • if noul > 0.8 escalate — a threshold risk and ops can change without rewriting prompts
  • choice == "billing" && confidence > 0.7 route — a labelled path with an explicit trust bar
  • score < 2.0 pull a human in — a gate you can put in a runbook, an audit log, and a golden-set test suite

That is still AI. It is also testable, auditable, and operable. You can hold labeled corpora, assert on expected Choice / Score / Noul bands, catch regressions when models or policies change, and show why a decision was automated or escalated. Transparency stops being a slide for the board and becomes evidence in the log.

Cost and speed close the enterprise case. High-volume closed decisions are exactly where a full chat LLM is often the wrong tool for production: too slow for the critical path, too expensive to call at ticket volume for a yes/no or a labelled Choice. TypeSafe reports Jev answering in milliseconds and at a fraction of frontier-model cost (their published framing is on the order of tens to hundreds of times cheaper/faster than typical chat LLM calls for these jobs). We have not independently re-measured those numbers here; the operating point still holds. In many of those situations, Jev is a no-brainer: typed output, calibrated confidence, and economics that make in-line decisions practical in a real control path — not a luxury reserved for demos.

Semi-deterministic does not mean perfect. It means the contract between model and code is stable enough to put under governance. In enterprise agent stacks — especially regulated or heavily supervised workflows — that contract (typed decision, honest confidence, affordable latency at volume) is the difference between a prototype and a system you can put in production and defend.

What we are doing at Runwaize

We have started rolling Jev into our internal processes and into existing solutions — Supervaize control paths, interview and ops workflows, the places where agents already make closed decisions at volume, and where calling a frontier chat model for every hop was never going to pencil out on latency or cost.

This is not a drop-in replacement for frontier models. We are not ripping out reasoning where reasoning belongs. We are treating Jev as a gradual optimization under scrutiny: swap or wrap the decision edges first — especially the high-volume ones — measure calibration against our labels, keep humans and stronger models in the loop where confidence is low.

In a market this frantic, that discipline matters. It is easy to chase every new capability. It is harder to notice when a primitive quietly changes what you can ship and test — and how clearly you can say why you trusted a decision.

Easy to miss — and hard to ignore

Jev's early traction is a signal that builders who care about delivering — not just demonstrating — have not missed this paradigm shift. Typed decisions with calibrated confidence are exactly the kind of boring foundation agent platforms need if we want control, auditability, and quality floors. The game changer is not that the model is smarter. It is that trust finally has a dial.

The room will not stay empty

I do not expect TypeSafe to own this category alone for long.

Incumbents will not leave this room to Jev. I expect Anthropic and OpenAI to ship their own classifier / System-One-style surfaces soon — constrained decision APIs sitting next to their chat models. Laya and open-source stacks will remain real alternatives for teams that want different tradeoffs on cost, hosting, or control.

That competition is healthy. The category is the point.

Jev's edge, today, is ease of use — plus economics that fit volume. A small set of primitives, a clear contract, a confidence score you can put in a threshold, decisions your code can act on without a parsing layer, at a speed and cost that make sense when you run them all day. For agent builders, that ease is not a nice-to-have. It is how you turn judgment into components — and components into systems you can trust under scrutiny.

The teams that win the next chapter of agentic software will not be the ones with the longest prompts. They will be the ones who put testable decisions — and an explicit trust dial — where chat used to be.

— Alain Prasquier

Founder & CEO, Runwaize

#creditAI: contribution by GrokBot


Sources