Structured Outputs in 2026: How to Get Reliable JSON from AI Models (OpenAI, Claude, Gemini)

Schema-constrained output fixes JSON syntax, not truncation, refusals or wrong values. How structured outputs work at OpenAI, Anthropic and Google, schema design rules, and a tested Python validation layer.

By Lumis Editorial · 11 min read · October 7, 2026

اقرأ هذا المقال بالعربية

Structured Outputs in 2026: How to Get Reliable JSON from AI Models (OpenAI, Claude, Gemini)

The moment an AI model's answer feeds into code (a database row, an invoice, a CRM record, the next step of an agent), "mostly correct JSON" stops being good enough. One missing bracket or a field named Total instead of total and your pipeline breaks at 3 a.m. In 2026 the major providers all offer structured outputs: you give the model a JSON Schema, and the API constrains generation so the answer follows it. That solves the syntax problem. It doesn't solve everything, and the gaps are exactly where production bugs live. This guide explains how structured outputs work at OpenAI, Anthropic and Google, what they still don't guarantee, how to design a schema that works, and includes a tested Python validation layer you can adapt.

Three levels of "give me JSON"

  • Prompt only. You ask nicely ("reply with JSON in this format"). Modern models comply most of the time, but "most of the time" means you need retries, and failures show up as stray prose, markdown fences, or missing fields.
  • JSON mode. The API guarantees the output parses as JSON, but not that it matches your structure. OpenAI's own announcement points out that JSON mode doesn't guarantee conformance to a particular schema.
  • Schema-constrained output (structured outputs). You send a JSON Schema and the API forces the output to follow it. This is the level to use for anything a program will read.

How constrained decoding works

A language model writes one token at a time, choosing from probabilities over its whole vocabulary. With structured outputs, the provider first converts your schema into a grammar. Then, at every step, it works out which tokens could legally come next and sets the probability of every other token to zero. If the schema says the next thing must be the key "total_aed" followed by a number, the model physically can't write anything else.

OpenAI introduced this in its API in August 2024. On its own evaluation of complex schema following, the then-new GPT-4o with Structured Outputs scored 100%, against less than 40% for the older GPT-4 (June 2023 version). Training alone got the new model to 93%; constrained decoding closed the rest of the gap. That's the core lesson: better models reduce errors, but only enforcement removes them.

Where the providers stand (October 2026)

  • OpenAI: set strict: true, either on a function definition (for tool calls) or on a response_format of type json_schema (for the answer itself). Only a subset of JSON Schema is supported.
  • Anthropic (Claude): two features that can be used separately or together. JSON outputs take a schema in output_config.format; strict tool use adds strict: true to a tool definition so tool names and inputs are validated. The feature launched as a beta in November 2025, and beta headers are no longer required. Anthropic's SDKs can also take a Pydantic (Python) or Zod (TypeScript) model directly.
  • Google (Gemini): set response_mime_type to application/json and pass a schema in response_json_schema. In November 2025 Google added JSON Schema support across its actively supported models, including anyOf, $ref, minimum/maximum and additionalProperties, and output keys now follow the order in your schema for Gemini 2.5 and later.

Each provider supports a slightly different slice of JSON Schema, and the lists change. Check the current documentation for the exact keywords before you rely on one. (For the tool-calling side of this, see our guide to reliable tool calling for AI agents.)

What structured outputs still don't guarantee

Read the fine print in OpenAI's and Anthropic's documentation and the same five gaps appear:

  • Truncation. If the response hits the token limit (max_tokens) before it finishes, the JSON is cut off mid-way. Both providers say the output may then not match the schema. Always check the stop reason before parsing.
  • Refusals. If the model declines a request for safety reasons, the answer may not follow your schema. OpenAI returns a separate refusal field; Anthropic returns stop_reason: "refusal", and you're still billed for the tokens.
  • Wrong values in the right shape. OpenAI says plainly that outputs can still contain mistakes in the values. A schema can force total_aed to be a number; it can't force it to be the right number, or positive, or a real date.
  • Small format surprises. Anthropic's docs note that the capitalization of enum values isn't guaranteed, so compare them case-insensitively. Anthropic also doesn't enforce numeric and length constraints such as minimum or maxLength in the sent schema; its SDKs move them into field descriptions and check them after the response.
  • Latency and limits. The first request with a new schema is slower while the grammar compiles (OpenAI says typical schemas take under 10 seconds, complex ones up to a minute; both providers then cache the result). Very complex schemas can be rejected: Anthropic documents limits such as 20 strict tools and 24 optional parameters per request. OpenAI's strict mode doesn't work with parallel function calls.

Designing a schema that works

  • Keep it flat and small. Every optional field and every union type adds grammar complexity. Extract what you need now, not everything that might be useful one day.
  • Use enums for categories. A fixed list ("office", "travel", "software") beats free text you have to map later.
  • Allow "unknown". If a field can be missing from the source, make it nullable. Otherwise the model is forced to invent a value to satisfy the schema, which is worse than an honest null.
  • Describe each field. Descriptions are part of the prompt. "Total including VAT, in AED, as a number without currency symbols" removes guesswork.
  • Don't put long reasoning inside the schema. Anthropic warns that asking for step-by-step reasoning inside a property can trigger a refusal; ask for a short explanation field instead.
  • Keep the schema stable. Changing its structure invalidates the compiled-grammar cache (and, on Claude, the prompt cache), so version it like an API.

A tested demo: the validation layer you still need

The script below needs no API key. It plays the part of your code receiving five raw responses from an invoice-extraction step, the kind every team sees in production logs. It checks the stop reason, extracts and parses the JSON, checks types, normalizes enum casing, and applies business rules a schema can't express. Each response gets one of three verdicts: OK, RETRY (with a reason you can send back to the model), or ESCALATE to a person.

"""Validate AI model output before your code trusts it (no API key needed).
Constrained decoding fixes the JSON syntax; this layer catches everything else:
truncation, refusals, wrong casing, and values that are valid JSON but wrong.
"""
import json, re
from datetime import date

SCHEMA = {  # what we asked the model to extract from an invoice
    "vendor": str, "invoice_date": str, "total_aed": (int, float),
    "category": str, "vat_included": bool,
}
CATEGORIES = {"office", "travel", "software"}

def right_type(value, expected):
    if expected is bool:
        return isinstance(value, bool)
    return isinstance(value, expected) and not isinstance(value, bool)  # True isn't a number here

def check(stop_reason, text):
    if stop_reason == "refusal":
        return "ESCALATE", "model refused; send to a person"
    if stop_reason == "max_tokens":
        return "RETRY", "output was cut off; raise max_tokens and retry"
    match = re.search(r"\{.*\}", text, re.S)  # tolerate ```json fences or chatter
    try:
        data = json.loads(match.group(0) if match else text)
    except json.JSONDecodeError as e:
        return "RETRY", f"invalid JSON ({e.msg})"
    errors = [f"missing '{k}'" for k in SCHEMA if k not in data]
    errors += [f"'{k}' has the wrong type" for k, t in SCHEMA.items()
               if k in data and not right_type(data[k], t)]
    if errors:
        return "RETRY", "; ".join(errors)
    data["category"] = data["category"].strip().lower()  # casing isn't guaranteed
    if data["category"] not in CATEGORIES:
        errors.append(f"unknown category '{data['category']}'")
    try:
        if date.fromisoformat(data["invoice_date"]) > date(2026, 10, 7):
            errors.append("invoice date is in the future")
    except ValueError:
        errors.append("invoice_date is not a real date")
    if not 0 < data["total_aed"] < 100_000:
        errors.append(f"total {data['total_aed']} AED is outside the allowed range")
    if errors:
        return "RETRY", "; ".join(errors)
    return "OK", data

# Simulated raw responses, the kind every team sees in production logs.
responses = [
    ("end_turn", '{"vendor": "Gulf Office Supplies", "invoice_date": "2026-09-28", '
                 '"total_aed": 1260.5, "category": "Office", "vat_included": true}'),
    ("end_turn", 'Sure! Here is the data:\n```json\n{"vendor": "SkyTravel", '
                 '"invoice_date": "2026-09-30", "total_aed": 2400, "category": "travel", '
                 '"vat_included": "yes"}\n```'),
    ("max_tokens", '{"vendor": "CloudSoft FZ-LLC", "invoice_date": "2026-10-01", "tot'),
    ("end_turn", '{"vendor": "CloudSoft FZ-LLC", "invoice_date": "2026-11-31", '
                 '"total_aed": -399, "category": "software", "vat_included": false}'),
    ("refusal", ""),
]
for i, (stop, text) in enumerate(responses, 1):
    verdict, detail = check(stop, text)
    print(f"{i}. {verdict:8} {detail}")

Output, run on 2026-10-07:

1. OK       {'vendor': 'Gulf Office Supplies', 'invoice_date': '2026-09-28', 'total_aed': 1260.5, 'category': 'office', 'vat_included': True}
2. RETRY    'vat_included' has the wrong type
3. RETRY    output was cut off; raise max_tokens and retry
4. RETRY    invoice_date is not a real date; total -399 AED is outside the allowed range
5. ESCALATE model refused; send to a person

What each case shows:

  • 1. Accepted, after normalizing "Office" to "office". Casing drift is common and harmless if you expect it.
  • 2. Prose, fences and a wrong type. This is what prompt-only JSON looks like. The parser tolerates the chatter, but "yes" isn't a boolean. With strict structured outputs this case shouldn't happen, which is exactly why you turn them on.
  • 3. Truncated. The stop reason says the model ran out of tokens, so the code doesn't even try to parse. No schema can prevent this; only checking can.
  • 4. Perfect JSON, wrong data. November 31 doesn't exist and the total is negative. This would pass any schema-constrained API. Business rules belong in your code.
  • 5. Refusal. Handled as its own path and sent to a person, instead of crashing on an empty string.

When a retry is needed, send the specific error back ("invoice_date is not a real date") rather than repeating the same request. Cap retries at one or two, then escalate. To measure how often each path happens on your real documents, see our guide to AI agent evals.

A checklist before you ship

  • Strict structured outputs are turned on wherever code reads the result.
  • The code checks the stop reason (truncation, refusal) before parsing.
  • Every field the source might not contain is nullable.
  • Enum comparisons are case-insensitive.
  • Business rules (ranges, real dates, totals that add up) are validated in code.
  • Retries send the specific error back, and there's a cap and a human fallback.
  • The schema is versioned, and the first-request latency is acceptable for your use.

The takeaway

Structured outputs turned "please reply in JSON" from a hope into an API guarantee about shape, and every serious integration should use them. But shape isn't truth. Keep a thin validation layer that checks how the response ended, whether the values make sense, and what to do when they don't. It's a few dozen lines of code, and it's the difference between a demo and a system you can leave running overnight. (Writing the instructions around the schema matters too; see our guide to writing system prompts for AI agents.)

Sources: OpenAI, "Introducing Structured Outputs in the API" (August 2024); Anthropic, Claude structured outputs documentation; Google, "Improving Structured Outputs in the Gemini API" (November 2025). Checked October 2026; supported schema features change, so confirm details in each provider's current documentation.

Related articles