AI, usefully · 4 min read ·

An AI agent needs a budget, not just a goal

Give an AI agent this instruction:

Research the market thoroughly and return the best answer.

When should it stop?

After five searches? Fifty? When two sources agree? When it has checked every plausible explanation? “Be thorough” gives the agent a direction, but no definition of enough.

This matters because an agent can repeatedly search, call tools, ask specialist agents for help and revise its answer. Each step consumes time and computing resources. More activity may improve the result—but it can also repeat the same work, collect weaker evidence and make the final answer harder to audit.

A production agent therefore needs two things:

  1. A goal describing what success looks like.
  2. A budget describing how much work it may do to get there.

The budget is more than money

Cost is the obvious constraint, but it is only one part of an execution budget.

A useful budget may limit:

BudgetWhat it controls
Model callsHow often the agent can ask a model what to do next
Tool callsSearches, database queries and external API requests
TimeHow long the task may run
TokensHow much information models read and generate
Specialist agentsHow widely the supervisor can delegate
RetriesHow many times a failed step may run again
DataWhich records, files or date ranges may be processed

These limits belong in application logic. Telling a model to “avoid unnecessary searches” may influence its behaviour, but it does not enforce a maximum.

The application should count the work and stop or change course when a limit is reached.

Match the effort to the question

Not every task deserves the same architecture.

“What is the delivery status of order 104?” probably needs one lookup.

“Compare five vendors across security, pricing and implementation risk” may benefit from parallel research and several specialist assignments.

The mistake is allowing the second workflow to become the default for the first.

Anthropic reported that its multi-agent research system was valuable for broad, parallel research but used substantially more tokens than ordinary chat interactions. It also found that tasks requiring tightly shared context or many dependencies were poorer candidates for multiple agents.

The lesson is not that multi-agent systems are wasteful. It is that their additional cost should buy something specific: greater coverage, independent investigation or faster parallel work.

If one model call and one tool can answer the question reliably, adding a supervisor and four specialists mostly creates more places for the workflow to fail.

Give the supervisor an effort policy

A supervisor should not have to invent the amount of effort from scratch.

An effort policy might say:

Request typeInitial plan
Direct record lookupOne read-only tool call
Defined comparisonRetrieve the required sources, calculate, validate
Broad research questionDivide into distinct research areas, then synthesize
High-impact actionResearch and prepare only; stop for approval
Ambiguous requestAsk a focused question before using expensive tools

This is still a starting policy, not an automatic guarantee of quality. But it gives the system a sensible default.

The supervisor can then escalate when the evidence justifies it. A conflicting source may warrant another search. A missing identifier should trigger clarification, not ten speculative lookups.

Define what “enough” means

A budget tells the agent when it must stop. A completion rule tells it when it should stop early.

For a research task, completion might require:

Notice what is absent: “use all twenty searches.”

The budget is a ceiling, not a target.

Once the completion rule is satisfied, more calls may add volume without adding value. The supervisor should be able to finish with budget remaining.

It should also be allowed to stop without an answer. If the evidence cannot support the requested conclusion, exhausting the budget should produce a clear account of what is missing—not a confident guess created because the workflow ran out of time.

Retries need their own limit

A failed tool call creates a tempting loop:

Call tool → receive error → try again → receive error → try again

Some failures are temporary. Others will never improve through repetition: access is denied, an identifier is invalid or the requested record does not exist.

Retry policies should distinguish those cases.

A temporary timeout might allow a small number of attempts with increasing delays. A permission failure should stop immediately. Invalid input should return for correction.

The retry count should stay in the task state so a restarted workflow does not quietly begin the same attempts again.

Measure useful work, not visible activity

A highly active agent can look impressive while accomplishing very little.

For each run, record:

Then compare similar tasks.

If a new workflow uses three times as many calls without improving correctness or coverage, the extra orchestration has not earned its place. If parallel specialists cut a complex investigation from an hour to ten minutes while preserving source quality, the expense may be justified.

Cost, latency and reliability should be evaluated together.

Try this

Before your next multi-step AI task, add this to the brief:

Propose the smallest plan likely to answer the question. State the tools required, the maximum number of searches or tool calls, what would justify additional work and the conditions for stopping. If the available evidence cannot support an answer within the budget, return the unresolved gaps instead of guessing.

Then inspect what the assistant proposes.

The goal is not to make AI do less. It is to make every additional step earn its cost.