Reliable Tool Calling for AI Agents in 2026: Structured Outputs, Validation and Safe Retries

Strict tool use guarantees the shape of a tool call, not its meaning. What structured outputs do and don't cover, how to design tools, and a tested execution layer for validation, idempotency, retries and useful errors.

By Lumis Editorial · 10 min read · October 7, 2026

اقرأ هذا المقال بالعربية

Reliable Tool Calling for AI Agents in 2026: Structured Outputs, Validation and Safe Retries

Most AI agent failures in production aren't the model "being dumb". They happen at the seam between the model and your systems: a tool call with a made-up customer ID, a date that already passed, a retry that books the same appointment twice, or an error message so vague the agent gives up. In 2026, the major model providers can guarantee that a tool call has the right shape. That's a big step, but it solves only half the problem. This guide covers what structured outputs actually guarantee, what they don't, and the small execution layer that closes the gap, with tested Python code you can adapt.

What structured outputs guarantee

Both Anthropic and OpenAI now offer constrained decoding: while the model generates a tool call or a JSON answer, tokens that would break your schema are blocked. The result always parses and always matches the schema's types and required fields.

  • Anthropic offers two generally available features: JSON outputs (output_config.format with a JSON schema) for structured answers, and strict tool use ("strict": true on a tool definition) to guarantee that tool names and inputs follow the schema. The SDKs can also take a Pydantic or Zod model directly.
  • OpenAI introduced Structured Outputs in August 2024, with strict: true for function calling and a json_schema response format. In its own evals, the new model with Structured Outputs scored 100% on schema matching, versus under 40% for an older model without it.

Before these features, teams wrote regex to fix broken JSON and retried on parse errors. That whole class of bug is now avoidable. Use it.

What they don't guarantee

A schema describes types, not truth. Four gaps remain, and the providers' own documentation states them:

  • Valid shape, wrong values. OpenAI's announcement says plainly that the feature doesn't guarantee the accuracy of the values inside the JSON. A perfectly formatted call can still contain a customer ID that doesn't exist.
  • Not every constraint is enforced. Anthropic's documentation lists JSON Schema features that strict mode doesn't support, including minimum, maximum, minLength and maxLength. The SDK helpers move those rules into the field description instead, so the model is asked to follow them but not forced to. A party_size of -4 is still a valid integer.
  • Edge cases. If the model refuses for safety reasons (stop_reason: "refusal") or runs out of tokens ("max_tokens"), the output may not match the schema. Anthropic also notes there's no guarantee on the capitalization of enum values, so compare them case-insensitively.
  • Side effects. Nothing about the schema stops a retried call from creating a second booking, charging a card twice or sending an email again.

So the rule for 2026 is simple: let the model guarantee the shape, and let your code guarantee the meaning.

Design tools the model can use well

Reliability starts before any code runs. Anthropic's tool-use documentation calls the description the most important factor in tool performance, and recommends at least three or four sentences per tool. Practical rules:

  • Say what the tool does, when to use it, and when not to. "Creates a confirmed booking. Use only after the customer has agreed to a specific date and time. Don't use it to check availability; use check_availability for that."
  • Describe every parameter, including its format (ISO date, currency, ID pattern) and where the value should come from ("the id returned by find_customer").
  • Show examples. Anthropic's tools accept input_examples: a few valid sample inputs that show the model which optional fields to include and how nested objects look.
  • Consolidate. One tool with an action parameter is often clearer than five near-identical tools, and fewer tools means fewer wrong choices.
  • Return less, but better. Return the fields the model needs and stable identifiers, not the whole database row. Every extra field is context the model must read.

One 2026 detail: on Anthropic's newest models (Claude Opus 5.5, Sonnet 5.5, Fable 5.1 and Mythos 5.1), tool_choice values that force a tool (any or a named tool) aren't supported and return an error. The documentation recommends auto combined with strict tool use instead. If you're upgrading an older agent that forced tool calls, check this first.

The execution layer: four jobs

Between the model's tool call and your real system, put a thin layer that does four things:

  • 1. Validate meaning. Check the rules a schema can't express: the ID exists, the date is in the future, the amount is within limits, the user is allowed to do this.
  • 2. Make actions idempotent. Give every write a key, so the same request twice produces one result. The model will retry, sometimes after a timeout where the first attempt actually succeeded.
  • 3. Retry only what's transient. Timeouts and 503s deserve a few retries with exponential backoff. Validation errors don't: retrying the same bad input gets the same answer.
  • 4. Return instructive errors. Send failures back as a tool result with is_error: true and a message that says what went wrong and what to do next. Anthropic's docs note that Claude typically retries two or three times with corrections when given a clear error, before telling the user.

A tested demo

The script below implements that layer for a booking tool. It needs no libraries or API key: the model's tool calls are scripted, and all five of them would pass a strict JSON schema (right types, all required fields). A fake calendar API fails half the time, with a fixed random seed so the run is reproducible.

"""A reliable tool-execution layer for an AI agent (no API key needed).
The model's tool calls are scripted below so the run is reproducible.
Structured outputs guarantee the SHAPE of a call; this layer checks the
MEANING, makes retries safe, and turns failures into useful tool_results.
"""
import datetime as dt, hashlib, json, random, time

TODAY = dt.date(2026, 10, 7)
CUSTOMERS = {"C-1001": "Mariam", "C-1002": "Omar"}
BOOKINGS = {}                      # booking_id -> booking
random.seed(7)

class Transient(Exception):        # e.g. timeout or HTTP 503
    pass

def flaky_calendar_api(payload):
    """Pretend calendar service: fails 50% of the time, like a bad day."""
    if random.random() < 0.5:
        raise Transient("calendar service timeout")
    return {"ok": True}

# --- 1. Semantic validation: rules a JSON schema can't express ------------
def validate(args):
    errors = []
    if args["customer_id"] not in CUSTOMERS:
        errors.append(f"customer_id {args['customer_id']} does not exist; "
                      "call find_customer first to get a valid id")
    day = dt.date.fromisoformat(args["date"])
    if day < TODAY:
        errors.append(f"date {args['date']} is in the past; today is {TODAY}")
    if not 1 <= args["party_size"] <= 8:
        errors.append("party_size must be between 1 and 8; "
                      "for larger groups, hand off to a person")
    return errors

# --- 2. Idempotency: the same request never books twice -------------------
def idempotency_key(args):
    raw = json.dumps(args, sort_keys=True).encode()
    return "bk_" + hashlib.sha256(raw).hexdigest()[:10]

# --- 3. Retries with backoff, only for transient errors ------------------
def with_retries(fn, payload, attempts=4, base=0.05):
    for i in range(1, attempts + 1):
        try:
            return fn(payload), i
        except Transient as e:
            if i == attempts:
                raise
            time.sleep(base * 2 ** (i - 1))      # 0.05s, 0.1s, 0.2s ...

def create_booking(args):
    errors = validate(args)
    if errors:
        return {"is_error": True, "content": "Invalid request: " + " | ".join(errors)}
    key = idempotency_key(args)
    if key in BOOKINGS:
        return {"is_error": False, "content": f"Already booked as {key} (no duplicate created)"}
    try:
        _, tries = with_retries(flaky_calendar_api, args)
    except Transient:
        return {"is_error": True, "content": "Calendar unavailable after 4 attempts. "
                "Tell the customer you'll confirm by message; do not retry now."}
    BOOKINGS[key] = args
    return {"is_error": False, "content": f"Booked {key} for {CUSTOMERS[args['customer_id']]} "
            f"on {args['date']} (calendar attempts: {tries})"}

# --- Scripted tool calls (all of them pass a strict JSON schema) ----------
calls = [
    ("valid booking",        {"customer_id": "C-1001", "date": "2026-10-12", "party_size": 2}),
    ("model retries same",   {"customer_id": "C-1001", "date": "2026-10-12", "party_size": 2}),
    ("invented customer id", {"customer_id": "C-9999", "date": "2026-10-12", "party_size": 2}),
    ("date in the past",     {"customer_id": "C-1002", "date": "2026-09-30", "party_size": 3}),
    ("negative party size",  {"customer_id": "C-1002", "date": "2026-10-15", "party_size": -4}),
]
for label, args in calls:
    r = create_booking(args)
    print(f"{label:21} is_error={str(r['is_error']):5} {r['content']}")
print(f"\nbookings stored: {len(BOOKINGS)}")

Output, run on 2026-10-07:

valid booking         is_error=False Booked bk_63413a6eac for Mariam on 2026-10-12 (calendar attempts: 3)
model retries same    is_error=False Already booked as bk_63413a6eac (no duplicate created)
invented customer id  is_error=True  Invalid request: customer_id C-9999 does not exist; call find_customer first to get a valid id
date in the past      is_error=True  Invalid request: date 2026-09-30 is in the past; today is 2026-10-07
negative party size   is_error=True  Invalid request: party_size must be between 1 and 8; for larger groups, hand off to a person

bookings stored: 1

What happened:

  • The valid booking survived a flaky dependency. The calendar timed out twice; backoff and retry got it through on the third attempt, without the model ever seeing the failures.
  • The model's retry didn't double-book. The same request returned the existing booking. Without the idempotency key, there would be two bookings and an annoyed customer.
  • Three schema-valid calls were rejected for meaning: an invented customer ID, a past date, and a negative party size (a range rule strict mode doesn't enforce). Each error tells the model exactly how to recover: look up the customer, use a future date, or hand off to a person.
  • Only one booking was stored, which is the only correct number.

In production, generate the idempotency key from a request ID your application creates once per user action, rather than from the arguments alone, so two genuinely separate bookings with identical details aren't merged. Payment providers like Stripe use the same idea with an idempotency key header.

Errors the model can actually use

Compare two tool results for the same failure:

"Error"

"customer_id C-9999 does not exist; call find_customer first to get a valid id"

The first leaves the model guessing; it may apologize, invent another ID, or give up. The second names the problem and the next step. Good error messages share three traits: they say what failed, why, and what to do now (including "don't retry; tell the user" when that's the right move). Never put stack traces, internal hostnames or secrets in a tool result: everything you return becomes part of the conversation.

Guardrails for tools that change things

  • Separate read and write tools. Reading availability is safe to call often; creating a booking isn't. Different tools make different permissions easy.
  • Human approval for irreversible or expensive actions: payments, refunds, cancellations, messages to many people.
  • Limits in code, not in the prompt. "Never refund more than 500 AED" in a system prompt is a request; the same rule in the execution layer is a guarantee.
  • Treat tool inputs as untrusted. Text the model read from an email or web page can steer it. Permissions and validation in code are what keep a manipulated call harmless. (See AI agent security and prompt injection.)
  • Log every call: arguments, validation result, attempts, final outcome. When something goes wrong, the log is how you find out why.

Test it before you ship

Turn the failure cases into an automated test set: invented IDs, past dates, out-of-range numbers, duplicate requests, a dependency that's down. Run it on every change to prompts, tools or models, and track how often the agent recovers correctly from each error. Our guide on testing AI agents with evals shows how to build that harness, and how to build your first AI agent covers the basic loop this layer plugs into.

The takeaway

Strict tool use and structured outputs have removed the most annoying class of agent bugs: broken JSON and missing fields. Turn them on. Then add the part no provider can do for you: check that the values make sense, make every write safe to repeat, retry only what's temporary, and return errors that tell the model how to recover. That thin layer is the difference between an agent that demos well and one you can leave running.

Sources: Anthropic documentation on structured outputs, strict tool use, tool definitions and handling tool results (checked October 2026); OpenAI, "Introducing Structured Outputs in the API" (August 6, 2024); Stripe API documentation on idempotent requests. The demo is self-contained and uses scripted tool calls.

Related articles