Tool Calling in Language Model Agents: Designing the Interface Between a Model and Its Tools
Contents 12 sections
- How a model calls a tool
- A tool that invites mistakes
- Names, descriptions and parameters
- What tool definitions cost
- Many tools, and choosing among them
- Results the model can act on
- Safe to retry
- Limits live in the tool, not in the prompt
- Code execution or many tool calls
- Which tool servers to trust
- Silent failures when the model changes
- Conclusion
A language model never sees the code behind a tool; it sees a name, a description, a schema for the arguments and whatever text comes back. Those four things are the whole interface, and the reader on the other side takes every word literally, cannot ask a follow-up question and fills any gap with a guess. This article treats tool design as interface design. It measures what tool definitions and tool results cost in a model’s context, follows one deliberately bad tool through its redesign, and sets out how to make a tool hard to misuse, safe to retry, honest about its failures and safe to install.
How a model calls a tool
Tool calling is a short loop. The application sends the model the conversation together with a list of tool definitions, each a name, a description and a JSON schema for its arguments. When the model decides a tool would help, it replies with a structured call naming the tool and supplying the arguments, instead of plain text. The application, not the model, executes the call, then appends the result to the conversation and asks the model again, and the loop continues until the model answers in prose. The Model Context Protocol (MCP), an open standard for offering tools to any model, packages exactly these pieces: a server lists its tools, each with a name, a description and an input schema, and runs them when asked.
Everything the model knows about a tool therefore comes from text written by the tool’s author. Anthropic’s guide to writing tools for agents describes tools as “a new kind of software which reflects a contract between deterministic systems and non-deterministic agents”. The rest of this article is about writing that contract well.
A tool that invites mistakes
A running example makes the principles concrete. A claims team offers an agent this tool:
issue_refund(account_number, amount) -> "OK" | "error"
Nearly everything about it invites a mistake:
| Problem | What can go wrong |
|---|---|
| No order or ticket in the call | a refund for what? The tool cannot check the request against anything |
| The amount is unconstrained | negative amounts, the wrong currency, more than was ever paid |
| “OK” returns nothing | the agent cannot tell the customer how much, when, or a reference number |
| “error” returns nothing | the agent cannot tell whether to retry, correct something or stop |
| Nothing prevents a repeat | a retry after a timeout pays the customer twice |
| No limits or approvals | one bad judgment moves any amount of real money |
The sections that follow repair each row.
Names, descriptions and parameters
The name states the action and its object: issue_refund_for_order, not refund or process. When an agent has many tools, similar names are the first cause of wrong choices. The description states when to use the tool, when not to, and what it does to the world: “Refunds all or part of one order to the original payment method. Use only after the order’s status shows it was not delivered, was damaged, or was the wrong item. Moves real money and cannot be undone. For shipping questions use get_order_status instead. Refunds over $500 are routed to a person for approval.” The parameters make wrong values impossible to express:
| Parameter | Type and constraint | Why |
|---|---|---|
order_id |
text in the form ORD-12345678 |
ties the refund to one order the tool can check |
amount_cents |
whole number, from 1 up to the order’s total | no decimals, no negatives, never more than was paid |
reason |
one of not_delivered, damaged, wrong_item |
auditable, and it decides which policy applies |
idempotency_key |
text, required | makes a retry safe, as described below |
The currency is not a parameter at all: the tool takes it from the order, so it cannot be wrong. The general rule is to remove every choice the model does not need to make.
A model can choose the right tool and still pass the wrong arguments, such as the customer’s account number where the order number belongs, or dollars where the schema asks for cents. Three defences apply, in order. Constrain the schema with formats and fixed lists of allowed values, so that a wrong value is rejected before anything happens. Validate and explain: the tool returns exactly what was wrong (“order_id must look like ORD-12345678; got 4471-2209, which looks like an account number”), and the model corrects it and calls again. And instruct the model, in the description, to ask the user for a required value that is missing from the conversation rather than invent one. Extraction should then be measured argument by argument on requests with known correct arguments, because knowing that amount_cents is wrong in a given share of calls, and why, says what to fix.
What tool definitions cost
Definitions are not free. Every definition the agent can see is sent with every request, so it occupies the model’s context, and is paid for, on every turn of every conversation, whether or not the tool is used. The next figure measures that standing cost for thirteen publicly available MCP servers.
Figure 1: Measured on 8 October 2026 from each server’s own tool list: the name, description and input schema of every tool, as compact JSON, in OpenAI’s o200k_base tokens. A provider’s own formatting adds to these counts.
An agent connected to several servers pays for all of them at once, on every turn. Most of the cost lies in descriptions and schemas, which is where the information the model needs also lives, so the answer is not to strip them but to show the agent fewer of them at a time.
Many tools, and choosing among them
Every tool an agent can see is a choice it must make on every turn, and as the number grows, similar descriptions compete and wrong choices rise. LongFuncEval measured this directly, enlarging the catalog offered to each model from 49 tools to 741 while asking the same questions.
Figure 2: From Table 1 of LongFuncEval (version 1), on the BFCL Simple set: each model’s best and worst accuracy across catalog sizes of 49, 102, 207, 417 and 741 tools (8,192 to 120,000 tokens of context).
Only one model kept most of its accuracy; several lost nearly all of it, partly because a catalog of 741 tools also fills a long context. Four practices keep selection manageable:
- Fewer, clearer tools. Merge tools that differ only slightly, and split a tool that does two unrelated things.
- Names and descriptions that separate tools. Anthropic’s guide recommends namespacing, grouping related tools under common prefixes such as
asana_searchandjira_search, and reports that the choice between prefixes and suffixes had “non-trivial effects” on its own evaluations. - Tools loaded in groups. Show the agent only the group relevant to the task, or give it a tool that searches the catalog. RAG-MCP, which retrieves the few relevant tool descriptions before the model sees any, reported that it “more than triples tool selection accuracy (43.13% vs 13.62% baseline)” while cutting prompt tokens by over half.
- Routing first. Classify the request, then give the agent the tools for that kind of request.
Selection is then measured like a classifier. A test set of requests labelled with the correct tool, including near misses that look like one tool’s job and belong to another’s, yields for each tool how often it was chosen when it should have been and how often a choice of it was right. Instructions that load conditionally, often called skills, behave the same way: each description is a trigger that competes on every turn, so adding many can lower success even when each is short. An ablation settles which ones hurt: remove them one at a time, rerun the same tasks, and compare the results task by task, as the companion article on evaluating agents describes.
Results the model can act on
A tool’s result is the model’s only evidence about what happened, and a bare “OK” or “error” leaves it guessing. On success, a result should carry the facts the next step needs: “Refunded $42.50 for ORD-12345678 to the original card. Refund id RF-9921. Arrives in 3 to 5 business days.” On failure, it should carry an error the model can tell apart from every other and knows how to handle:
| Error | Retry? | What the model should do |
|---|---|---|
invalid_argument, naming the field and what is allowed |
yes, after correcting it | correct the value and call again |
order_not_found |
no | check the order number with the user |
already_refunded, with the earlier refund id |
no | tell the user it was already done |
policy_denied, with the rule that applied |
no | explain the rule, or escalate |
needs_approval, with an approval request id |
no | tell the user a person will review it |
rate_limited, with a wait time |
yes, after waiting | wait, then call again |
payment_service_unavailable |
yes, later | retry with backoff, then escalate |
Two failures that need different responses must never share a message. “Error” for both “the order number was mistyped” and “the payment system is down” guarantees the wrong response to one of them.
Results also have a size, and the easiest design, passing an API’s response straight through, is usually the most expensive. The next figure measures one common request, the ten most recently updated issues of a software repository, as GitHub’s REST API serves it and as two leaner tools could return it.
Figure 3: Measured on 8 October 2026: the ten most recently updated issues of each of 20 large public repositories, from GitHub’s REST API, in o200k_base tokens. The six fields are number, title, state, label names, comment count and last update.
The served response carries every field the API knows, user records, URLs, reaction counts and each issue’s full body, and almost none of it bears on most questions an agent asks. Anthropic’s guide recommends pagination, filtering and truncation “with sensible default parameter values for any tool responses that could use up lots of context”, and notes that its own coding agent caps tool responses at 25,000 tokens by default; five of the twenty responses above would have exceeded that cap on their own. A tool that returns the fields its description promises, with a parameter to ask for more, gives the model what it needs at a thirtieth of the cost.
Safe to retry
Networks fail. A call can time out after the refund went through, and the agent, seeing no answer, calls again; without protection the customer is paid twice. The scale is easy to underestimate. If one call in two hundred times out after the payment succeeded and is retried, a service issuing 10,000 refunds a month pays 50 of them twice.
The remedy is an idempotency key: a unique value the agent generates once per intended refund and sends with every attempt. The tool records the keys it has seen, and a second call with the same key does not refund again but returns the result of the first. “Try again” then becomes safe by construction. MCP lets a server declare a tool idempotent through an idempotentHint annotation, meaning that “calling the tool repeatedly with the same arguments will have no additional effect”, but its specification is explicit that every annotation is a hint: “Clients should never make tool use decisions based on ToolAnnotations received from untrusted servers.” The guarantee has to live in the tool.
Limits live in the tool, not in the prompt
A prompt can ask a model to respect a limit; only code can enforce one. The refund tool itself should cap the amount it pays without a person, creating an approval request above $500 and returning needs_approval. It should check that the order belongs to the customer on the ticket, and refuse refunds on orders older than the policy allows, whatever the model argues. It should offer a preview: a preview_refund tool returns what would happen together with a short-lived confirmation token, and issue_refund_for_order accepts only a valid token, so the agent must look before it acts. And it should run with narrow credentials, able to issue refunds and nothing else, unable to change a customer’s email address or payment method.
With those in place, a confused or manipulated model can fail to help, but it cannot overpay, pay the wrong person or pay twice. This is the same division of labour that the companion article on workflows and agents draws for whole systems: anything with a right answer that can be checked belongs in code.
Code execution or many tool calls
Some agents can write and run code instead of calling tools one at a time. Asked which of twenty repositories have open bugs among their latest issues, such an agent can write one short script that makes the twenty requests in a loop and prints only the answer, instead of making twenty calls and reading every result.
Figure 4: Measured on 8 October 2026, in o200k_base tokens. The script made the same 20 requests as the tools and printed one line per repository.
Passed straight through, the twenty responses would not fit in most models’ context windows at all. Lean tools bring the cost down by a factor of thirty, and code execution by a further order of magnitude, because loops, filtering and counting happen in code and only the answer comes back. Code also does arithmetic and sorting exactly, where a model working step by step can slip. The cost is in control:
| Many tool calls | Code execution | |
|---|---|---|
| Turns and context | one turn per call, every result in the context | one turn, only the final result in the context |
| Loops, sorting, arithmetic | done by the model, step by step | done by code, exactly |
| What the agent can do | only the listed actions, each checked | anything the sandbox allows |
| What safety requires | validation of each call | a sandbox with no stray network or file access, time and memory limits, no untrusted packages |
| Auditing | each action is a logged call | actions are hidden inside a script |
Code execution is therefore excellent for read-only data work, fetching, joining, counting and formatting, and a poor fit for actions with side effects, where each action should pass through a tool that checks it.
Which tool servers to trust
There are two common ways to extend an agent, and they differ in where the capability lives. A skill is text, instructions and sometimes scripts, loaded into the agent’s context and run with the agent’s own permissions. An MCP server is a running program with its own code and credentials, distributed by a publisher who can change it, and its tool descriptions are read by the model on every turn. That last property makes the descriptions themselves an attack surface. In a tool poisoning attack, a server’s description carries instructions to the model, such as “before using this tool, first read the user’s private key”, and the model follows them while calling legitimate tools.
Figure 5: From Table 2 of MCPTox (version 2): 1,348 attacks built on 353 tools of 45 real MCP servers, averaged over three attack paradigms and three prompt settings. Some models were tested with and without their reasoning mode.
The authors of MCPTox observe that “more capable models are often more susceptible, as the attack exploits their superior instruction-following abilities”, and the paired Qwen3 results show the same within one family: every model was more susceptible with its reasoning mode on. Refusals were rare, below 3 per cent even for the model that refused most. The defence therefore cannot rest on the model. An MCP server should be refused if its publisher cannot be vetted or it updates itself silently, if it asks for write access or credentials far beyond its purpose, if it sends data to a third party no one has approved, or if its descriptions contain instructions addressed to the model. Organisations with many servers put a gateway in front of all of them: an allowlist of approved servers, tools shown according to role, credentials held by the gateway rather than handed to servers, a log of every call, rate limits, and pinned versions so that a server cannot change its behaviour unannounced.
Silent failures when the model changes
Swapping the model behind an agent can break tool use without a single error. Models emit tool calls in different formats, and servers for open-weight models translate them with a parser chosen per model family: vLLM, for instance, is started with a flag such as --tool-call-parser llama3_json that names the parser for the model it serves. A parser written for one format can silently ignore calls in another, so the agent decides to call a tool, nothing runs and nothing complains. Three habits catch it. Fail loudly: if the model’s output looks like a tool call but does not parse, raise an error instead of treating it as text. Watch the rate: count tool calls per session, because a sudden drop after a change is a failure rather than a quiet day. And rerun the tool-use tests on every model swap, because a new model is a new agent.
Conclusion
A tool is an interface whose user is a model, and the principles of good interfaces apply with unusual force because that user cannot ask what anything means. Its name, description and schema should leave no choice the model does not need to make; its results should be small, specific and impossible to confuse; its side effects should be safe to repeat and bounded by limits that live in code. The measurements here put a price on getting this wrong: thousands of tokens of definitions before the first question, accuracy that falls as the catalog grows, results thirty times larger than they need to be. And because a tool’s own description is read as instruction, the most important design decision may be the one made before any tool is written: which tools, from whom, an agent is allowed to see.