Era 6 · Generative models and systems · 2023 / 2024

55 Structured Output

function calling (OpenAI, 2023) · Efficient Guided Generation (Outlines) · Willard & Louf, 2023 · XGrammar, 2024 · OpenAI Structured Outputs, 2024
🟧 read selectively~2–2.5 horiginal ↗
The gist in 20 seconds. How to make an LLM emit a guaranteed valid structure (JSON matching a schema) rather than a "usually valid" one. Three waves: prompt-JSON (unreliable) → function calling (train the model to emit schema-shaped calls, 2023) → constrained / guided decoding (at every step mask the logits against an automaton, leaving only tokens that keep the grammar intact). Outlines builds an index of "FSM state → allowed tokens"; XGrammar uses a byte-level pushdown automaton. OpenAI Structured Outputs (2024) gets 100% schema validity this way. One subtlety: a strict format can hurt reasoning.

Context

An LLM is a distribution over text, but applications (agents, tool use, API pipelines) need machine-readable output: valid JSON matching a schema exactly. Simply asking it to "reply in JSON" is unreliable: the model drops a comma, adds prose, muddles types. On large schemas prompt-only gets roughly a third of responses valid — unacceptable in production. What you need is a guarantee, not a hope.

The idea and the mechanism

1. Function calling (2023). The model is fine-tuned to emit {"name": …, "arguments": {…}} given a function description and its parameter schema. This made tool use practical — but validity is still probabilistic: a trained model usually hits the schema, but is not obliged to.

2. Constrained / guided decoding (2023–24). The guarantee comes not from training but from the decoding stage. You run an automaton that tracks the already-generated prefix against the grammar; at each step you mask the logits of every token that would make the output invalid, and sample only from what is allowed. The structure is valid by construction. Outlines reduces a regex or grammar to a finite automaton and precomputes an index: state → the set of allowed tokens (otherwise you would have to scan the whole vocabulary at every step). XGrammar takes a context-free grammar, compiles it into a byte-level pushdown automaton and caches the masks. OpenAI Structured Outputs (2024) compiles a JSON schema into a grammar and so guarantees 100% conformance (plus training on schemas).

automata · decoding Masking logits, and why token alignment is the hard part

The core. Let the automaton be in state s, and let \( \mathcal{A}(s) \subseteq V \) be the set of tokens after which a valid completion is still possible. The model gives logits \( z \in \mathbb{R}^{|V|} \); we mask and sample only from the allowed set:

\[ z'_i = \begin{cases} z_i, & i \in \mathcal{A}(s)\\[2pt] -\infty, & i \notin \mathcal{A}(s)\end{cases}\qquad x_t \sim \mathrm{softmax}(z'),\qquad s \leftarrow \delta(s, x_t) \]

Forbidden tokens get \(-\infty\) → zero probability after the softmax; the distribution is renormalized over the valid ones. The chosen token moves the automaton to a new state \(\delta(s,x_t)\). Which kind of automaton you need depends on the grammar:

\[ \text{regex} \Rightarrow \text{DFA } (Q,\Sigma,\delta,q_0,F)\qquad\qquad \text{JSON / CFG} \Rightarrow \text{PDA (with a stack)} \]

A flat format is described by a regular language (a finite DFA). But JSON is recursive — nested, balanced {} and [] do not form a regular language, so you need a pushdown automaton: the stack keeps count of the nesting depth.

The hard part is token alignment. The grammar is defined over characters/bytes \(\Sigma\), while the model generates tokens V (BPE subwords). One token covers several characters, and one string has many tokenizations — there is no direct correspondence. So you cannot simply "allow the right characters": for each state you have to work out the allowed tokens. Hence the precomputed index:

\[ \mathrm{Idx}:\; Q \;\to\; 2^{V},\qquad \mathcal{A}(s) = \mathrm{Idx}(s) \]

The cost. Checking every token for validity naively is \(O(|V|)\) per step (a vocabulary of ~10⁵). Outlines' index makes it \(\approx O(1)\) (plus a cheap vectorized mask application). XGrammar goes further: it splits tokens into context-independent ones (validity is clear from position alone — precomputable, >99% of them) and context-dependent ones (the whole stack matters), caching masks by the top of the PDA stack. The upshot — <50 µs per token, negligible against 10–50 ms of inference.

Python Constrained decoding: masking logits against an automaton
def constrained_decode(model, automaton, idx):   # idx: state -> set(token_id)
    s, out = automaton.start, []
    while not automaton.is_accept(s):
        z = model.logits(out)                     # logits over the vocabulary V
        allowed = idx[s]                          # precomputed index: tokens allowed from s
        z = mask_fill(z, allowed, NEG_INF)        # forbidden -> -inf (zero probability)
        tok = sample(softmax(z))                  # sample ONLY from the valid ones
        out.append(tok)
        s = automaton.step(s, tok)                # automaton transition (regex->DFA, JSON->PDA+stack)
    return detokenize(out)                         # guaranteed valid structure
the decoding loop (every step): logits zover V mask A(s)invalid → −∞ samplesoftmax automaton: state+stackδ(s, tok) the automaton's state sets the next mask three waves: prompt-JSON → function calling (2023) → guided decoding (Outlines/XGrammar) → guaranteed (OpenAI SO, 2024)
At each step the automaton (for JSON, one with a stack) sets the set of allowed tokens; the logits of the forbidden ones are driven to −∞, sampling happens only among the valid, and the chosen token advances the automaton. The structure is valid by construction.
Analogy. A sat-nav that lights up only the legal turns at every junction. The road map is the automaton; the lit turns are the mask of allowed tokens; the PDA stack is its memory of how many brackets are still open. Wherever the model "drives", it physically cannot turn into a forbidden street — so it always arrives at a valid address (one matching the schema). Training (function calling) only makes the driver used to such routes; the mask is what guarantees they do not break the rules.

Why it matters

This is the foundation under all modern tool use, agents and dependable API pipelines: without a guarantee of valid output you cannot build systems where a program parses the LLM's answer. "Structured output", which everyone uses today (OpenAI, Anthropic tool use, vLLM/TGI/llama.cpp with grammars), is exactly this line of work: training on schemas + constrained decoding. The broader lesson: the reliability of LLM systems is often born not in the model itself but in the decoding layer around it.

Connections

← lives in the decoder32. Transformer

Constrained decoding wedges itself in exactly where an autoregressive Transformer turns logits into a token — at the softmax over the vocabulary. Masking logits is possible precisely because generation goes token by token: before each sample you can zero out the forbidden options. Without the step-by-step nature of #32 there would be no such lever.

← the training half44. InstructGPT (RLHF)

Function calling is alignment to a tool: the model is fine-tuned to follow call schemas (a relative of the SFT/RLHF in #44). Training makes the output usually valid and sensible; constrained decoding adds the guarantee. Full structured output is the sum of the two: a trained model plus a mask.

↔ in tension with43. Chain-of-Thought

CoT wants free reasoning out loud; a strict format suppresses it. The study "Let Me Speak Freely?" (2024) showed that a hard JSON mode hurts reasoning — the model is forced to produce the answer before it has "thought". The practical fix is a free-form CoT field first, format afterwards (the order of fields in the schema decides it).

Questions worth asking

If the mask guarantees valid JSON — why not turn constrained decoding on always?

Because the mask distorts the distribution: zeroing out some tokens and renormalizing can push the model into regions it considers unlikely, and impose structure before it has finished thinking. "Let Me Speak Freely?" (2024) showed a measurable drop in reasoning under a strict format — largely because of the reordering: the answer has to be named before the reasoning. Mitigations: a free reasoning field first, strict fields after (the order in the schema), or generate in natural language and convert to the format in a separate step. A format guarantee is not free.

Why is this harder than "allow only the right characters"?

Because of token alignment. The grammar lives over characters/bytes, the model over BPE tokens: one token glues several characters together, and the same string can be tokenized in different ways. So "allow the character {" does not translate directly into "allow this token" — for every automaton state you have to work out which (multi-character) tokens keep the grammar intact. Hence the precomputed index (Outlines) and the byte-level PDA with its split into context-independent tokens (>99%, precomputable) and context-dependent ones (XGrammar). That is where the real engineering depth of the topic sits.

Regex is enough for many formats — why does JSON specifically need a stack (PDA)?

JSON is recursive: an object can contain objects, arrays nest arbitrarily, brackets must balance. The language of balanced brackets is the classic example of a non-regular one (a finite automaton cannot recognize it: you have to "remember" arbitrary depth). Hence a pushdown automaton with a stack, which counts the open brackets and demands they be closed. For flat formats (a date, an enum, a number) regex/DFA is enough — but the moment the structure nests, you cannot do without a stack.

What to read in the original

Read selectively. Outlines (Willard & Louf, arXiv 2307.09702) — the key idea of the "FSM state → allowed tokens" index and the reduction to an automaton. XGrammar (arXiv 2411.15100) — the byte-level PDA, the split of tokens into context-(in)dependent, the mask cache. The OpenAI Structured Outputs post — how a JSON schema is compiled into a grammar for 100% conformance. And do read "Let Me Speak Freely?" (arXiv 2408.02442) — an honest account of what a strict format costs reasoning.