AI Function Calling

Every provider’s function calling docs show you the same thing: define a schema, the model emits a call, you run it, you hand back the result. That part takes an afternoon. Then you put it in front of real traffic and discover the docs covered maybe a third of what you needed.

The loop itself is genuinely simple. You send the model a list of tools. It either answers or it asks for a tool. You run the tool, append the result to the conversation, and send the whole thing back. You keep going until it stops asking. That’s it, and it’s the same shape everywhere.

messages = [{"role": "user", "content": user_input}]

while True:
    response = call_model(messages, tools=TOOLS)
    messages.append(response.message)

    calls = response.tool_calls
    if not calls:
        return response.text

    results = run_in_parallel(calls)
    messages.append(package_results(results))

What follows is everything I had to learn after that loop was already working.

The description is the prompt

The schema tells the model what’s structurally valid. The description tells it when to reach for the tool at all, and that’s the part that decides whether any of this works. A tool called search_database described as “Search internal database” gets called for things it can’t answer and skipped for things it can.

Write the description like you’re onboarding somebody. What’s in this data. What isn’t. When to prefer a different tool. What a good query looks like. It’ll be longer than feels proportionate for a JSON blob, and it does more work than anything else you write here.

The same goes for each parameter. "filters" described as “Optional filters” is a coin flip. Say what keys it accepts and what happens when you omit them.

Return errors, don’t raise them

The instinct from normal code is to let a failing tool throw and handle it in the loop. That throws away the model’s ability to recover.

If a tool fails, hand the failure back as the tool’s result and say what went wrong. Most providers have a flag for this, an is_error marker or equivalent, so the model knows it’s a failure rather than data. Given “no user found with that email”, a model will usually try a different lookup. Given a stack trace in your logs and silence in the conversation, it can’t do anything.

The exception is failures the model can’t route around. An expired API key isn’t something it should retry, and letting it try four times just burns tokens on its way to the same wall.

Return every parallel result, in one message

This is the one the docs are quietest about.

Models can request several tools at once. When you send the results back, all of them go in a single message. Split them across separate messages and most implementations will accept it without complaint, and then the model quietly stops making parallel calls at all. It’s not an error you can catch. It’s a behavioral regression that shows up as latency creeping upward over a week.

Same rule for a call that failed. Return an error result for it. Dropping the result entirely leaves a call with no answer, and behaviour after that is undefined in the useful sense of “you’re on your own”.

Tool results are context, and context is the bill

Every result you append stays in the conversation for the rest of the session, and you resend the whole thing on the next turn. A tool that returns a 40 KB JSON payload has just added 40 KB to every subsequent request in that conversation.

Return what the model needs to answer, not what your API happens to give you. Summarize, truncate, paginate. If the model genuinely might need the full document, put it behind a second tool it can call when it decides to, rather than paying for it on every turn whether it’s used or not.

Providers have started shipping mechanisms for clearing or compacting old tool results out of context automatically. They help, and they’re not a substitute for not putting it there in the first place.

Assume it will be called twice

Retries happen. A timeout you couldn’t distinguish from a slow response, a network blip, a model that asks for the same thing twice because the first result was ambiguous. If your tool sends money or an email, that’s a real problem and not a hypothetical one.

Anything with a side effect needs an idempotency key, and the model shouldn’t be the one generating it. Derive it from the call, or keep a short-lived record of which calls you’ve already executed in this session.

Separately, anything destructive needs a gate that isn’t the model’s judgement. Keep a list of tools that require confirmation and check it before dispatch, not inside the tool.

Constrain the schema where you can

Providers now offer a strict mode that guarantees the arguments you get back actually validate against your schema, rather than merely usually validating. It generally requires you to close the schema, meaning additionalProperties: false and an explicit required list.

Turn it on. Validating hand-parsed JSON in your own code is work you no longer have to do, and the failure mode it removes, a model inventing a plausible parameter name, is one you’d otherwise find in production rather than in tests.

On frameworks

The abstraction layers are genuinely useful once you’re supporting several providers, and they’re a bad place to start. The differences between providers are small and mostly mechanical, roughly the shape of the tool definition and what the result block is called. If you write the loop yourself once, you’ll understand what the framework is doing when it misbehaves, and you’ll be able to tell whether a bug is yours or theirs.

Start direct. Reach for a framework when you have a second provider to support, not before.

I still don’t have a good answer for the loop-termination problem. A hard iteration cap is the standard answer and it’s obviously wrong: it either cuts off legitimate long work or lets a confused model spin ten times before you stop it. Token budgets that the model can actually see are a better idea than a cap it can’t. I haven’t used them enough to say whether they solve it.

Written July 2025 and revised August 2026. The specific model identifiers in the original version had all been retired, so this revision keeps the patterns and drops the version-specific code.