AI & agents

Why one agent question costs four requests

The mistake

You give an agent a question and it comes back with an answer, so you price it like one completion. The tools run on your own server, so they look free.

The cost is in the asking, not in the tool. A tool call does not pause the request, it ends it. The model replies “run search_docs”, your code runs it, and then the whole conversation goes back to the provider in a brand new request, with your instructions and every tool schema attached again.

The machine

Simulator · the agent loop
cache the prefix

Every count runs on the tested reducer. The loop's rules were read from laravel/ai v1.0.0; the token counts are stated, not measured, and no model was called.

1 provider request, 220 tokens sent, 180 of that the same instructions and tool definitions again. The model asked for search_docs, so the loop goes round again.

Drive it

Send the next request and watch one tool call finish as two.

Switch the run to three tool calls. Count the requests, then look at how much of each bar is the same prefix you sent last time.

Tick both cache boxes. The request count stays at four and the bars shrink.

The mechanism

prompt() runs a loop. In laravel/ai it is literally a for loop, and every pass through it is one HTTP request to the provider:

for ($step = 0; $step < $maxSteps; $step++) {
    // one provider request per pass
}

Each pass builds a fresh request carrying four things: your instructions, your tool definitions, the schema, and the whole message history so far. The first three do not change between steps. They are sent again anyway, because the provider has no memory of the last call.

So one question with three tool calls is four requests. The pattern is n tool calls, n plus 1 requests: each tool call ends a request, and one more is needed to turn the last result into an answer.

Two costs grow, and they grow differently. The prefix, your instructions and tool schemas, is a fixed size paid once per step, so it multiplies by the number of steps. The history is a different shape: every assistant message and every tool result is added to it and then carried by every later request. A 400-token page fetched on step two is sent again on step three and again on step four.

Prompt caching addresses the first cost. Mark the instructions and the tool definitions as cacheable and the provider keeps the prefix between steps, so step one writes it and the rest read it. Nothing addresses the second cost, which is why an agent that fetches large documents gets expensive faster than an agent that makes many small calls.

The loop ends in one of two ways. Either the model stops asking for tools, which is the loop breaking on its own, or the step budget runs out, which is the for condition failing. The first gives you an answer. The second leaves you without one.

There is a budget whether you set one or not, and the documentation does not say what it is. The source does. Left alone, laravel/ai works one out: five steps when the agent declares no tools, and otherwise one and a half times the tool count, rounded, capped at 25. An agent with one tool gets a budget of two, which covers one tool call and an answer, and nothing more. An agent carrying #[RepairToolCalls] gets one extra step on top, and that is the one part of the rule the documentation does spell out.

The last request under that budget behaves differently. On the final step the tool is not run at all, and the loop substitutes a fixed sentence in place of the result:

The agent reached its maximum number of steps without running this tool call.

So a run that hits its ceiling spends its last request, executes nothing, and returns no answer to the caller. Set the budget to two in the simulator and you can watch it happen.

In your code

Set the budget yourself rather than inheriting a derived one, cache the parts that never change, and read the step count back off the response:

use Laravel\Ai\Attributes\{CacheInstructions, CacheToolDefinitions, MaxSteps};

#[MaxSteps(6)]
#[CacheInstructions('1h')]
#[CacheToolDefinitions('1h')]
class ResearchAgent implements Agent, HasTools
{
    use Promptable;
}

$response = (new ResearchAgent)->prompt('Compare the two refund policies.');

count($response->steps); // how many requests that answer actually took
$response->usage;        // and what they cost

If count($response->steps) equals your MaxSteps, treat it as a failure. It means the loop was cut off before it got to an answer.

The fine print

The token counts in the simulator are ours. They are stated in the dataset, not measured from text, because the page is about the shape of the loop and not about any tokeniser. The loop’s rules are the package’s, read from laravel/ai v1.0.0. Two of them, the derived step budget and what the final step does, are not in the documentation at all, which is why the reducer’s spec asserts them against the lines they came from rather than citing a page.

The simulator counts tokens and never money. Cache read and write multipliers differ by provider and by retention window, so any price on this page would be wrong for most providers the day it was written. The Lab enum in v1.0.0 has seventeen cases, and not all of them generate text.

The cache model here is the simple one: the prefix is sent on the first step and read on the rest. Real providers charge a premium to write a cache entry and have a retention window that can expire mid-run, and a provider that does not support caching ignores the attributes rather than failing.

Left out: approvable tool calls, which pause the loop and resume it on your decision, and deserve their own page; streaming and the two frontend protocols; where conversation history is stored; and how the model decides to reach for a tool at all, which is the model’s business rather than the SDK’s.

Further reading

  • Laravel AI SDK documentation is the reference for agents, tools, caching and the response object. Read it for the API, and check the source for the step budget.
  • TextGenerationLoop.php is the loop itself. resolveMaxSteps and resolvedToolResult are the two methods this page rests on, and they are short.
  • Anthropic: prompt caching covers what a cached prefix actually costs to write and to read, which is the half the simulator leaves out.

Spotted a problem, or have a way to make this clearer? Suggest an improvement.