The build is only part of what an AI agent costs. Once it is live, every task it handles uses a language model, and that usage is billed. Teams are sometimes surprised by the first invoice, because agent costs do not behave like normal software costs. They are still predictable, once you know what drives them.
What you actually pay for
Model providers charge for the text sent to the model and the text it produces, measured in tokens. A token is roughly a short word or part of a word. The important detail is that an agent does not call the model once per task. It calls it once per step, and each call usually includes the conversation so far.
So a task that takes ten steps does not cost ten times a single call. It can cost much more, because the tenth call carries the history of the previous nine.
The three things that drive cost
- Steps per task. An agent that wanders, retries and re-reads uses far more than one that goes straight to the answer.
- Context size. Sending a whole document when one paragraph is needed multiplies the bill on every step.
- Model choice. The most capable models cost many times more per token than smaller ones. Many steps, such as classifying a request or extracting a field, do not need the largest model.
Measure cost per task
The useful number is not the monthly total. It is the cost of completing one task. Log token usage for every run and report the average, and also the expensive tail: the few runs that cost twenty times the typical one. Those usually point to a loop or a tool that returns far too much text, and fixing them is often the cheapest improvement available.
Ways to reduce it
Use the right model for each step. Route simple steps to a small, fast model and keep the capable model for the reasoning that needs it.
Trim what the tools return. A tool that returns an entire customer record when the agent needs three fields wastes tokens on every later step. Return only what is needed.
Cache what repeats. Instructions and reference material that are identical on every call can often be cached by the provider at a lower price. Answers to identical lookups can be cached by you.
Do not use a model where code will do. Date arithmetic, format checks and fixed rules belong in ordinary code. It is cheaper, faster and always gives the same answer.
Set hard limits
Every agent should have a maximum number of steps and a maximum spend per task. When a run hits the limit, it stops and hands over to a person. This turns a possible runaway cost into a visible, bounded event. Add an alert on daily spend as well, so a sudden change is noticed the same day and not at the end of the month.
Compare with the work it replaces
Cost per task only means something next to the alternative. If a person needs ten minutes for the same task, the comparison is clear. Include review time honestly: an agent whose output takes eight minutes to check has not saved much.
Summary
Agent cost depends on steps, context and model, and all three can be measured and limited. Track cost per task, cap every run, and use small models and plain code where they are enough. Cost monitoring is built into the AI agents we deliver. To estimate running costs for a specific use case, contact us.