Structured Outputs

Redaktion ·

An LLM answers in free-flowing prose by nature — even when you ask it to “please return only JSON.” Structured Outputs flip that around: instead of a request to the model, a JSON schema gets technically enforced. The model literally cannot emit tokens during generation that would violate the schema. For any pipeline that processes the result by machine — writing it to a database, handing it to another system, parsing it inside an agent loop — that’s the difference between “usually valid” and “guaranteed valid.”

Three tiers: free text, JSON mode, Structured Outputs

The confusion comes from three tiers of very different strictness getting lumped together:

  1. Free text with a request for JSON. The prompt says “respond as a JSON object with fields x and y.” The model mostly complies, but “mostly” isn’t good enough for a production parser. A markdown code fence, an explanatory sentence before it, a missing comma — any of these trips up JSON.parse().
  2. JSON mode. A flag like response_format: {"type": "json_object"} guarantees syntactically valid JSON — brackets match, quotes match. What JSON mode does not guarantee: that the right fields exist, that types are correct, that no extra fields show up. Valid JSON with the wrong shape is still valid JSON.
  3. Structured Outputs. Here, not just JSON syntax is enforced but the entire schema: exact field names, exact types, no additional properties, every required field populated. This is the tier pipelines actually need.

How it works technically: grammar, not a request

The trick is called grammar-constrained decoding. Before generation, the JSON schema gets compiled into a formal grammar that dictates which tokens are even allowed at each position. When the model reaches a spot where the schema calls for a boolean, only the tokens for true and false are permitted at all — everything else is excluded during generation, not discarded afterward by validation.

That also explains the latency quirk shared by all three providers: the first request with a new schema takes a bit longer because the grammar has to be compiled first. After that, the compiled grammar is cached (24 hours at Anthropic) and subsequent requests with the same schema are fast again.

Anthropic: JSON format and strict tool use

Anthropic offers Structured Outputs in two variants that can be combined:

  • JSON output mode via the output_format parameter (Python SDK: client.messages.parse() with a Pydantic model) — for pure data extraction, with no tool involved.
  • Strict tool use via strict: true on the tool definition — here the grammar enforcement applies directly to the arguments the model produces during Function Calling. This is the case most relevant for agents: an agent that fills a tool’s arguments wrong breaks the entire chain.
from pydantic import BaseModel
from anthropic import Anthropic

class Ticket(BaseModel):
    title: str
    priority: str
    assigned_to: str

client = Anthropic()
response = client.messages.parse(
    model="claude-sonnet-5",
    max_tokens=512,
    messages=[{"role": "user", "content": "Create a ticket from: ..."}],
    output_format=Ticket,
)
print(response.parsed_output)

OpenAI: response_format with json_schema and strict

At OpenAI, the same thing runs through response_format: {"type": "json_schema", "json_schema": {...}, "strict": true}. Without strict: true you stay in the weaker JSON mode — the option has to be set explicitly, it doesn’t kick in automatically just because a schema is attached. It has been supported since the gpt-4o-2024-08-06 snapshots and runs consistently across the current GPT-5.x and GPT-6 Astra model family.

Google Gemini: responseSchema and responseMimeType

Gemini enforces the format via two fields in generationConfig: responseMimeType: "application/json" unlocks JSON output in the first place, and responseSchema (a subset of the OpenAPI 3.0 schema) defines the exact structure. Since 2026, the Gemini API also understands regular JSON Schema format, letting libraries like Pydantic or Zod pass their schemas through directly, without translating into the OpenAPI subset.

The reliability gain for pipelines

Without Structured Outputs, every extraction pipeline needs a fallback loop: parse, retry on failure, give up and escalate after three failed attempts. That costs tokens, latency, and code for an edge case that should never have occurred in the first place. With an enforced schema, that loop disappears entirely — the result is either there on the first attempt, or the request itself fails (e.g. because of an invalid schema), never “technically valid JSON in the wrong shape.” That makes Structured Outputs the foundation of any LLM API integration meant to produce data instead of prose — from classification pipelines all the way to tool calls inside an AI agent.

Limits and pitfalls

The grammar enforcement carries similar restrictions across all three providers:

  • Shape only, not truth. An enforced schema guarantees structure, not content. A model can deliver a syntactically perfect JSON object full of made-up values — the problem of hallucination doesn’t go away because of it.
  • Restricted schema vocabulary. In strict mode at Anthropic and OpenAI, refinements like minLength, maxLength, minimum or maximum aren’t allowed, recursive schemas aren’t either, and additionalProperties: false plus a complete required array are mandatory on every object. If you bring an existing, looser schema, you first have to trim it down to this standard.
  • The first call is slower. Grammar compilation on the first request with a new schema costs noticeable time — with frequently changing schemas (e.g. dynamically generated ones), the caching benefit evaporates.
  • No substitute for server-side validation. An object that’s guaranteed schema-conformant can still be wrong on the business level (a priority level that doesn’t exist in the target system). Business rules remain the job of your own application.

FAQ

Is Structured Outputs the same thing as JSON mode?
No. JSON mode only guarantees syntactically valid JSON, regardless of field names or types. Structured Outputs additionally enforces the entire schema — wrong or missing fields simply cannot occur during generation.
Do I need Structured Outputs for chat answers to humans too?
Usually not. The benefit shows up when a machine processes the answer unchecked — for plain chat text, a free-form prompt is simpler and the grammar enforcement is unnecessary overhead.
Does the schema also guarantee factually correct answers?
No. The schema only enforces shape — which fields exist and what type they are. Whether the values inside are correct is a separate question; hallucination is not ruled out by it.
Why is the first request with a schema slower?
The schema gets compiled into a grammar before first use, which restricts the token-sampling process. That compilation costs time once, but is then cached for a certain period.
Does strict tool use also work with several tools at once?
Yes, each tool gets its own strict flag and its own schema. The model still freely chooses which tool to call — what gets enforced is only that the arguments of the chosen tool exactly match its stored schema.
See everything in one place:API Usage