Structured Output
Structured output is the capability to make an LLM emit values that conform to a predefined schema (usually JSON), rather than free-form text. It turns the LLM from "a text generator that talks" into "a programmable function call" — downstream code no longer has to do brittle string parsing and retries.
Three mainstream implementations
1. JSON Mode (early)
- Model is constrained to emit only valid JSON strings
- Field names / types are not enforced — unexpected keys can still appear
- OpenAI introduced it in 2023, Anthropic and Google followed
2. Structured Outputs (strict schema)
- Caller provides a JSON Schema or similar DSL
- Model is forced to follow the schema; type-mismatched output simply won't generate
- OpenAI shipped it in August 2024 (GPT-4o); Anthropic followed with stricter tool-use in October 2024
3. Function Calling / Tool Use
- "Structured output" repackaged as "tool call"
- Model emits
{name, arguments}where arguments is structured JSON - Native to Anthropic, OpenAI, Google, xAI, Mistral
- Combined with MCP (Model Context Protocol) it became the de-facto agent tool-call standard
Why it matters
- Directly consumable by code: instead of "Is this email a complaint?"
returning a paragraph, the model returns
{category: "complaint", confidence: 0.92}and downstream code just branches - Eliminates parsing failures: no more handling
{"answer": "..."}vsAnswer: ...vs markdown-wrapped JSON - Cuts retry cost: a malformed output used to mean a full retry; a schema mismatch now produces a typed error
- Auditable: every output is a structured value that can land in version control / databases
The extreme form: System One Models
TypeSafe AI Jev pushes structured output to its logical extreme:
- No string-generation phase at all; choices are enumerated in the schema up front
- Every option comes with a calibrated probability
- 0% type errors (a structural guarantee, not a statistical rate)
- End-to-end 70-500 ms — 40-200× faster than LLMs
See typesafe-ai.md.
Evaluation
- Schema compatibility: 100% (TypeSafe Jev) / 99.x% (GPT-4o structured)
- Field accuracy: depends on prompt design and schema complexity
- Hallucination boundary: structured output guarantees shape, not substance — Jev can still pick the wrong option, which is why probabilities matter
Common deployment scenarios
- Data extraction (contracts, invoices, resumes)
- Classification / routing / scoring
- Agent tool calls (MCP)
- Real-time decisioning (structured if-statements)
- Table / database writes