TypeSafe AI
TypeSafe AI is an AI lab that came out of stealth on 2026-09-15, founded by former OpenAI researcher Diogo Almeida — co-inventor of RLHF and InstructGPT, the research that became ChatGPT's foundation. After two years in stealth, the company released its first System One Model, Jev: a model built for the small decisions buried inside software, claimed to be 40–200× faster and 40–400× cheaper than frontier LLMs on System One tasks.
Key products
- Jev: first System One Model, opened to waitlisted early access on 2026-09-15. It does not chat — it picks from a user-defined schema, returning every decision in parallel along with a calibrated probability.
- System One category: a new class of "machine-native intelligence" — built for software, not for talking to people.
- Jevons naming: named after 19th-century economist William Stanley Jevons; the company bets that cheaper intelligence unlocks more uses, not fewer (Jevons paradox).
Technical highlights
- RLCD (Reinforcement Learning for Calibrated Decisions): not RLHF (human preference), not RLVR (verifiable reward), but a third path — optimized for calibrated probabilities, so "92% confidence" actually means 92% accuracy.
- Parallel sampler: no autoregressive decoding; all candidate outputs within the schema come back in a single pass, end-to-end 70–500 ms. Frontier LLMs on the same task take 3–329 s.
- Type-safe outputs: outputs are constrained to a schema declared up
front, so type errors are mathematically impossible (a structural
guarantee, not a measured rate). The same logic underlies their claim
of
0% hallucination— but selecting the wrong option is still possible. - Pricing: input $0.042 / million tokens, output free — "too cheap to meter".
Influence
- Breaks open the dominant "AI model = chat box" frame: peels off a slice of intelligence from text generation and dedicates it to structured decisions.
- Early developers swapped existing LLM classifiers for Jev: Vercel replaced ChatGPT Luna 5.6 in a safety reviewer and got 5–18× faster with higher accuracy; Bryo AI pitted Jev against Gemini on email classification and got 10–20× cheaper (Gemini slightly more accurate).
- Pushes the "AI division of labor" thesis: LLMs handle slow thinking and content; System One handles real-time routing, guardrails, judging, and "smart if-statements" inside software.
- Reignites the debate on whether calibration matters more than raw accuracy: 67.8% accuracy with honest confidence may be more automatable than 74.1% accuracy without it.
Limitations
- 255-option ceiling: Jev picks from a finite set; beyond that it splits into two passes (score-then-pick). Open-ended generation is out of scope.
- No generation: no chat, no articles, no complex reasoning, no code generation. Jev is not a ChatGPT replacement — it is a ChatGPT copilot.
- Accuracy: on TypeSafe's own workflow evals, Jev scores 67.8%, behind GPT-5.6 Sol (74.1%) and Claude Opus 5 (73.1%). It wins on cost-per-correct-decision, not on raw accuracy.
- Biased evals: the reference answers are the average of GPT-6 Astra and Claude Fable 5.1 (built-in bias toward OpenAI/Anthropic); workflows were written by TypeSafe's own capabilities team; competing LLMs were forced through TypeSafe's structured-output wrapper (slower and more expensive than asking them directly).
- Early access: API is waitlisted, weights are not public, and independent third-party evaluations are still scarce.