Jev

TypeSafe AI's Jev decides instead of writing: a typed answer with a probability in about 150ms, where an LLM takes seconds.

16 Sep 2026 An LLM answers you in words. Jev answers in a decision.

You give Jev some text or data, plus the questions you want answered and the answers you'll allow. It picks one, and attaches a probability and a confidence score. It can't write a sentence, and that's the point: the reply is already in a shape your code can act on.

Take a support email. You want to know if the customer is angry, which team owns it, and how urgent it is. An LLM reads the email, thinks for a few seconds, and writes you a paragraph, or JSON if you asked nicely and it behaved. Jev returns angry = yes (0.91), team = billing (0.96), urgency = today faster than the page you're reading rendered.

An LLM Jev
You send A prompt State, plus the allowed answers
You get back Text you have to parse A typed answer with a probability
One decision takes 3 to 9 seconds 0.07 to 0.5 seconds
1,000 decisions cost Dollars A fraction of a cent
Ask 50 questions at once Slower, and less accurate Same call, same time
It can Write, reason, code Choose, score, judge
It goes wrong by Inventing something Being confidently wrong

Measured, not claimed: Every timed the same writing check at 0.35s a passage against 8.83s for Claude Fable 5.1, and ran 777 judgments in under 0.7 seconds for about a quarter of a cent. Good Start Labs priced a grading job at $160 per million answers against $33,000 for Fable 5.1.

The speed comes from what Jev doesn't do. There are no tokens to generate one after another, so the answer arrives in roughly the time a network round trip takes. That's the difference between a check you run once at the end and a check you run on every message.

TypeSafe AI launched it on 15th September 2026. The company calls it a System One model, after Kahneman's fast, intuitive thinking, as opposed to the slow deliberation that chat models perform.

TypeSafe AI wordmark and cube logo in dark grey on a pink brand card

Maker TypeSafe AI, San Francisco
Founders Diogo Almeida (CEO, ex-OpenAI), Sasha Sheng (COO, ex-Meta FAIR), Erik Gafni (CTO)
Funding $40M, per The Register
Price $0.042 per MTok input ($42 per billion), output free
Latency 70-500ms claimed, 150-350ms in tester reports
Model name jev-latest
Access Early access by waitlist, plus a browser playground

What it actually does

One call carries the state (a message, a document, a data structure) and a dictionary of questions. Each question is one of 3 types:

  • Choice - pick one option from a set. Returns the option, a probability per option, and a confidence.
  • Score - rate against ordered levels. Returns the level, a probability per level, and a confidence.
  • Noul - a yes/no claim. Returns the probability that it's true.

Questions run in parallel against a shared read of the state, so asking 50 costs about as much wall time as asking 1. That inverts the habit built around LLM classifiers, where each extra question muddies the answer and adds latency.

The model can't produce free text, which is what TypeSafe means when it says Jev can't hallucinate. It's a claim about types, not about correctness. Jev can still be wrong, and confidently so.

For business people

Plenty of production AI work is not writing. It's deciding: which team gets this ticket, is this reply angry, should the agent call a tool, is this answer good enough to send. Today those decisions go to a chat model that takes a few seconds and bills per token in both directions. The latency forces workarounds, like a crude rule that decides whether the expensive check is worth running.

Jev sells the decision as its own primitive, sitting between plain code and an LLM call. The pitch is price and speed: the launch post claims 20-200x faster and 40-400x cheaper than frontier models on decision work, with output tokens free.

What it costs in practice: a tester ran about 5,000 requests for around $2. Good Start Labs priced the same grading job at $160 per million answers against $33,000 for Claude Fable 5.1.

Where the risk sits:

  • It's a new company with a new model. Early access, a waitlist, and no track record beyond 1 week.
  • The headline benchmarks are the vendor's own. TypeSafe admits its evaluation workflows came from its own capabilities team.
  • Confidence has to be earned. The value of a calibrated probability depends on 90% meaning 90%, and that needs testing on your own labelled data.
  • It replaces nothing on its own. You still run an LLM for the writing. Jev goes before and after it.

Worth a test if you already pay for classification, routing or scoring calls at volume. Not worth a rebuild if you make a few hundred decisions a day, where the cost difference is noise.

For technical people

Call it

curl -X POST https://api.typesafe.ai/v1/systemone \
  -H "Authorization: Bearer $TYPESAFE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "state": "I was charged twice. Please fix this ASAP.",
    "model": "jev-latest",
    "questions": {
      "is_urgent": {"type": "noul", "instructions": "The message conveys urgency"},
      "department": {"type": "choice", "instructions": "Which team should handle this",
        "criteria": {"billing": "Payment issues", "technical": "Bugs", "sales": "Pricing"}}
    }
  }'

The Python SDK wraps the same call:

from typesafe_sdk import Choice, Noul, Score, TypeSafeClient

with TypeSafeClient() as client:
    response = client.system_one(
        state={"document": ticket_text},
        questions={
            "billing": Noul(instructions="Is this ticket about billing?"),
            "tone": Choice(
                instructions="What is the customer's tone?",
                criteria={"calm": None, "frustrated": None, "angry": None},
            ),
            "urgency": Score(
                instructions="How urgent is this ticket?",
                criteria=["can wait", "this week", "today"],
            ),
        },
    )

print(response.nouls["billing"].noul)
print(response.choices["tone"].choice)
print(response.scores["urgency"].score)

Install with uv add typesafe-sdk or pip install typesafe-sdk, then set TYPESAFE_API_KEY. Both a sync and an async client ship.

Confidence is the control surface

Every Choice and Score answer carries probabilities plus a single confidence between 0 and 1. The documented pattern is 3 bands: act automatically when confidence is high, confirm or flag in the middle, route to a human at the bottom. Thresholds scale with the stakes of the action, and your code holds them, not the model.

action = response.choices["action"]

if action.confidence < 0.5:
    route_to_human(user_message)
elif action.choice == "approve_transfer" and action.confidence < 0.9:
    ask_user_to_confirm(account_id)
else:
    execute(action.choice)

Training and shape

TypeSafe trains with RLCD (Reinforcement Learning for Calibrated Decisions) instead of RLHF, optimising for honest probabilities rather than human preference. The sampler produces all answers in one pass rather than token by token, which is where the latency goes.

⚠️ WARNING: the input wants structured program state, not a chat transcript. Choice cardinality caps at 255. Above that, TypeSafe documents a 2-stage pattern: score candidates independently, then choose.

Documented patterns worth reading before you design around it: speculative fan-out (ask everything you might need in one call), confidence-gated routing, composite scoring, and intent routing.

What early testers measured

  • Michael Lee ran about 5,000 requests for ~$2 across classification, routing, intent and steering. He measured p50 ~150ms and p95 ~350ms, and called it "a new intelligent decision-making primitive, separate from deterministic code and LLM calls". His old setup used low-latency LLM classifiers averaging 4 seconds, which forced hand-written heuristics to decide when to run them. At 150ms you run the check every turn, before and after generation.
  • Every put 777 judgments across 37 documents through it in under 0.7 seconds for about a quarter of a cent. On a writing-lint test Jev caught 6 of 7 planted defects at a median 0.35s per passage; Fable 5.1 caught 7 of 7 at 8.83s.
  • Good Start Labs ran 6,003 rubric checks over 1,203 financial-research answers. Jev agreed with a 5-model panel 90% on average, in a range of 86% to 92%, at the $160 versus $33,000 cost noted above. Their own caveat is the honest one: "Agreement is evidence about a judge; it doesn't establish who is right."
  • One early tester posted a confident error: Jev rated a benign sentence as attempted AI manipulation at 0.84 probability. His line is the one to remember: "type-safe" does not mean "can't be wrong".

Tested against real outcomes

30 Sep 2026 Muratcan Koylan, who builds the AI receptionist at Sully.ai, gave Jev 2,029 real phone calls. It's the first public test I've seen scored against what actually happened, not against other models.

Video: Muratcan Koylan, original post on X.

  • No words in. Jev never saw a transcript or heard the audio. Each call was reduced to structure: turns, tool calls, workflow stages and timing. That's the "program state, not chat" input the model is built for.
  • Zero-shot. No fine-tuning, no examples from their data.
  • Volume and speed. 38,012 turn-level forecasts during the calls at a median 118ms. After the calls, 5 typed questions per call: 10,145 answers in 26 seconds, with 256 requests in parallel.
  • Cost. About $3 for the whole run.
  • Scored against the patient record. The team compared Jev's forecasts with what the EHR (the electronic health record) showed: did the caller book or not. Halfway through a call, Jev separated booking calls from the rest with an AUC of 0.78. Near the end of the call, it ranked them correctly 94% of the time.
  • The weakness. Jev over-weighted visible errors that their agent usually recovers from. It judged the surface of the call, not the agent's ability to fix things.

This answers part of the "agreement is not truth" caveat below: here the benchmark is a real outcome. It's still 1 team's report on their own data, with no published method.

What to check before trusting it

  • Calibration on your data. Take 200 labelled examples from your own workflow and check whether the 0.9 bucket is right 90% of the time. Everything else follows from that number.
  • Agreement is not truth. Both public tests measure agreement with other models, not ground truth.
  • The founder claim. Almeida's post opens with "After co-inventing ChatGPT". The team page says RLHF and InstructGPT, which is the accurate version and still a strong credential. Some replies on X pushed back on the shorter phrasing.
  • "Hallucination-free" is definitional. The Register made the same point: no natural language output doesn't preclude being incorrect.

Where I'd use it

Every sales tool I build makes small judgements before it makes a sentence: is this reply a real objection or an out-of-office, which account does this email belong to, is this call note worth pushing to the KB, does this draft break my own writing rules. Today those are LLM calls, and they are the slow part of every pipeline.

A decision primitive at 150ms and near-zero cost would let me run checks I currently skip, like scoring every draft before it reaches me instead of after. The catch is access: early access only, through a waitlist.

20 Sep 2026 joined the waitlist. First test when the invite lands: score my own drafts against my writing rules, the check that is too slow to run on every draft today.

21 Sep 2026 access granted, one day after joining. The welcome screen is under The hook.


How X explained it

18 Sep 2026 Moritz Kremb collected the clearest explanations from the first 48 hours into 1 thread. The framings there are sharper than the launch material, so this section keeps what they add.

The 2 framings that stuck

An AI-native if statement. @paarangatrai puts it as a change in what a condition can read. Instead of this:

if transaction > 10_000:
    review()

you write the condition against messy context:

if suspicious_behaviour > 0.95:
    review()

You declare the answers first (risk: low / medium / high, manual review: yes / no) and Jev returns risk = high (96%), manual review = yes (91%). His point about volume matters more than the syntax: the same state can carry several typed questions at once, so fraud, churn, escalation, eligibility and priority all come back from one call.

A really smart switch statement. @NathanFlurry calls it "like if 2016 ml classifiers got 2026 levels of intelligence", and gives the cleanest inventory of the limits:

Can't Can
Write code or natural language Classify, route, score, rank
Show step-by-step reasoning Return a confidence with every answer
Produce an output you didn't define Pick the branch, tool, model or sub-agent
Pick from more than ~255 options at once Judge, verify or guardrail an LLM's output

He also calls the AGI framing in the launch post "incredibly far fetched", which matches how the rest of the thread reads: a useful new tool, not a new species.

The architecture everyone converges on

LLM proposes, Jev decides, code executes. Jev sits in front of the expensive model as a routing layer rather than replacing it.

@swill1ams turns that into the money argument. A workflow where a frontier model reviews every ticket, invoice or claim pays seconds and cents for a call that is usually obvious. Put a cheap decision model first, let it handle the items it's confident about, and pass the rest through. Confident on 6 of 10 items means more than half the spend at that step is gone. His other pre-processing candidates are the interesting part:

  • Route each request to a cheap model, a frontier model or a human.
  • Pick which skill or sub-agent to load for a turn, instead of stuffing the whole catalogue into context.
  • Rerank retrieved context so only relevant chunks reach the window.
  • Guardrail every agent turn: contradictions, policy breaches, prompt injection.
  • Extract typed fields from emails, PDFs and transcripts before anything expensive reads them.

What people built in 48 hours

Demo Builder What it shows
Browser agent @gregpr07 Flights found in 7s for $0.0039, small LLM only to type
Trading bot @jarrodwatts Buy/sell on a price feed, orders on-chain every 300ms
Self-driving @jpschroeder Tesla-style FSD decisions, rebuilt in under an hour

The browser agent belongs to Browser Use, and it rebuilds its action space at every step. The trading bot shows the shape of the thing best: a 300ms block leaves no room for a chat model, so the decision either fits that budget or the product doesn't exist.

The 4 corrections worth keeping

  • "It's just a classifier." Closer to a classifier that takes unstructured state and answers several independent questions from one read.
  • "It replaces GPT or Claude." No. It decides when to call them.
  • "It can't hallucinate." The output structure is constrained, the judgement isn't. In @paarangatrai's words: "HIGH at 92% can still be the wrong decision."
  • "Just force an LLM to return JSON." You can, and you keep the generation latency, the schema validation, the retries and the glue code. Jev removes that layer rather than the model.

The hook

21 Sep 2026 the first screen after signing in. A letter from the 3 founders on the left, 3 columns of claims on the right, one button at the bottom: Enter console.

TypeSafe welcome screen: a letter titled Meet Jev signed by Diogo, Erik and Sasha, beside columns listing properties, limitations and benefits, with an Enter console button

The letter

"Meet Jev. Our first (public) System One model."

The pitch in one line: "a new class of AI models optimized for programmatic (inside code) use. Think: Smart if-statements." That's the framing @paarangatrai used on X a few days earlier, now in the company's own words.

RLCD gets its origin story: 2 years of research aimed at "mode dropping, hallucinations, and lack of reliability inherent to RLHF". Then the letter answers the obvious objection before you raise it: "Our claims may sound too good to be true, but the bitterest lesson in AI is that optimizing for the right task gets you an unfair advantage."

That line is a deliberate twist on Rich Sutton's The Bitter Lesson, which argues the opposite: general methods plus compute beat task-specific cleverness. TypeSafe is betting its company on the exception. Signed off "May your intelligence be ever reliable", from Diogo, Erik and Sasha.

What they claim, and what they concede

Each item expands on click. The 4 without detail below were closed when I saved the page.

Properties. "Jev has a fundamentally different architecture with new training and sampling methods."

  • Structured, machine-native outputs. Type correctness "guaranteed by design. Not 99.9999% success, actually 100%."
  • Parallel sampling for fast decisions.
  • Consistent, calibrated, probabilistic results. A confidence value you can use in code. Each decision is a calibrated probability, "so similar inputs give similar outputs".

Limitations. "Real tradeoffs we want to be honest about."

  • Not good at System 2 tasks. Weaker than large reasoning models at high-reasoning work, "like mathematical reasoning and games like chess".
  • Not trained on specialized domains.
  • Not a generative chat model. No text, so no chat. You define the shape of the answer first, "kind of like writing multiple choice questions".

Benefits. "We get that these are big claims, and we're excited to prove them to you."

  • 20-200x faster.
  • 40-1,000x cheaper.
  • Frontier-level intelligence for System 1 tasks. On instinctive judgement and common sense over large bodies of text and structure, Jev "approaches frontier reasoning models". Followed by the most honest sentence on the page: "This is the hardest claim to defend, and no one in the field has found a good way to prove it."

The quickstart

The console's quickstart is not a code sample. It's a prompt to paste into your coding agent, and what it does is install TypeSafe's skill:

# Claude Code
claude plugin marketplace add typesafe-ai/skills
claude plugin install typesafe@typesafe-ai

# any other agent
npx skills add typesafe-ai/skills --skill typesafe-ai

One method, not both. The skill is MIT-licensed, with 1,258 stars since it appeared in late August.

Most of what it teaches the agent is where to look. It tells the agent to treat the live docs as the source of truth, start from docs.typesafe.ai/llms.txt, and read the current API or SDK page before writing any integration. Then it offers 4 patterns that go further than the plain classifier everyone builds first:

  • Route and fill known arguments. Pick the handler and its typed parameters in the same call.
  • Select instead of generate. Code finds the candidate values, Jev picks the right one, code copies it.
  • Find and judge evidence. Rerank retrieved context so only the relevant parts go forward.
  • Turn judgements into reusable data. Score once, then let code change the weights and thresholds.

"Select instead of generate" is the one that fits my work best: finding the right contact, the right email, the right figure in a document, without letting a model retype it.

What changed since launch

  • The docs grew. The skill points to a JavaScript SDK and a migration guide to a v1 API. Neither appeared in the docs index when I first read it on 16 Sep.
  • The price claim grew. The launch post said 40-400x cheaper. The welcome page says 40-1,000x, 6 days later, with no new figure behind it on the page. The speed range stayed at 20-200x.
  • The concessions are new, and they're the useful part. Weak at System 2, no specialised domains, no chat. That lines up with the independent tests further up: fast and cheap on judgement calls, a step behind the frontier model on the hardest check.
  • "Actually 100%" is about types, not answers. The same distinction this note already makes: the output always has the right shape, and the answer inside it can still be wrong.

The first test stays the one I set on the waitlist day: score my own drafts against my writing rules.

Further Reading

NicAI
Written by NicAI, Nic's AI assistant, for his personal knowledge base. Researched and drafted by the model, not hand-written by Nic. Verify anything you plan to act on.