September 10, 2026 · StartupQuickstart
What an AI agent actually costs to run per month
Token prices from the official pricing pages, a worked example for an 880-run-a-month email agent, and the costs that never show up on the model invoice: retries, evals, and human review.
Founders ask us some version of the same question every month: "what will this agent cost to run?" The honest answer has two parts. The model bill is usually smaller than people fear. The things around the model bill, mainly testing, retries, and the human who checks the output, are usually larger than anyone budgets. This post walks through both with real prices and a worked example you can adapt to your own workflow.
Prices in this post are from the official pricing pages as checked on October 8, 2026. Model prices change several times a year, almost always downward, so treat the arithmetic as the durable part and plug in current numbers before you commit to a budget.
How the meter actually runs
API pricing is per token, quoted per million tokens, with separate rates for what you send (input) and what the model writes back (output). Output is the expensive side, typically five times the input rate. Anthropic's rule of thumb is that one token is roughly four characters or three quarters of an English word, but the same page warns that newer Claude models use a tokenizer that produces about 30% more tokens for the same text. That is a good reason to measure real usage from the API response instead of estimating from word counts.
Three price points cover most small-business work right now:
- Small tier: Claude Haiku 5.5 at $0.10 per million input tokens and $0.50 per million output (for prompts under 100,000 tokens), and OpenAI's gpt-6-luna at the same $0.10 and $0.50.
- Mid tier: Claude Sonnet 5 at $2 input and $10 output, with OpenAI's gpt-6-sol at the same $2 and $10.
- Top tier: Claude Opus 5.5 at $4 input and $20 output, and higher still for the frontier models.
Two discounts matter. Prompt caching lets you reuse a repeated prefix, like your system prompt and tool definitions, at a fraction of the input price; on Claude a cache read costs 10% of the base input rate for most models and a five-minute cache write costs 1.25x. Batch processing gives a 50% discount on input and output for jobs that can wait, and OpenAI's batch prices are also half the standard rate.
A worked example: an inbound email agent
Take a realistic small-business workflow. A services company gets about 40 inbound emails per business day: quote requests, scheduling questions, invoice questions. The agent reads each email, looks the sender up in the CRM, and drafts a reply plus a CRM note for a person to approve. Over 22 business days that is 880 runs a month.
An agent is a loop, and every turn of the loop resends the conversation so far. That is the part people forget when they estimate. Our assumed run has three model calls:
- Read and decide. A 3,000-token system prompt with tool definitions, plus a 1,500-token email thread: 4,500 tokens in, about 300 out (a request to call the CRM tool).
- Look up. Everything from call one, plus the tool call and a 1,000-token CRM record: 5,800 in, 300 out.
- Draft. Everything so far plus a short second lookup: 6,600 in, 500 out (the reply and the note).
That totals 16,900 input tokens and 1,100 output tokens per run. Notice that input is fifteen times output, almost entirely because of the resent context.
On the mid tier at $2 and $10 per million, one run costs $0.034 for input plus $0.011 for output, about $0.045. Multiply by 880 runs and the model bill is roughly $39 a month. The same workload on the small tier comes to about $2 a month. On Opus 5.5 it is about $79.
Caching the 3,000-token prefix on the mid tier saves about a cent per run in this example, because the second and third calls read the prefix from cache instead of paying full price for it. That brings the month to roughly $31. It is real money at scale, and it is not the number that decides whether this project is worth doing.
The costs that do not show up on the model invoice
Retries and failed runs
Production agents hit rate limits, timeouts, malformed tool output, and the occasional run that goes in circles. OpenAI's guidance is to honor the Retry-After header, back off exponentially with jitter, cap both attempts and total retry time, and not retry billing or quota errors. It also notes the official SDKs already retry some rate-limit errors, so your own retry loop can double the attempts if you are not careful. A run that fails on its third call has already paid for the first two. Budget 10-20% on top of the clean-run cost until your logs tell you the real number. In our example that takes $39 to about $45.
Evals, which can cost more than production
Before you trust an agent, and every time you change its prompt, you rerun it against a fixed set of test cases. Anthropic's eval guide pushes toward more test cases with automated grading over fewer hand-graded ones, and its examples run from 50 to 1,000 cases. Say you keep 200 representative emails. One full eval pass on the mid tier costs 200 x $0.045, or about $9, plus whatever a model-based grader costs. During active tuning, when you might run the suite eight or ten times in a month, eval spend can exceed the production bill. Run evals through the batch endpoint, which halves the price, and keep the suite small enough that people actually run it.
Tools that charge per call
Some built-in tools have their own meter. Anthropic's web search, for example, costs $10 per 1,000 searches plus the tokens the results add to the context. An agent that searches twice per run on 880 runs adds about $18 a month before counting those extra tokens, which then get resent on every later turn.
Human review, the biggest line by far
In the example, a person approves every draft. Suppose a good draft takes 30 seconds to check and send. That is 880 x 0.5 minutes, about 7.3 hours a month. The median wage for secretaries and administrative assistants was $23.23 an hour in May 2025, so review alone is roughly $170 a month before benefits and overhead, four times the model bill. If drafts are mediocre and each takes two minutes to fix, review jumps to about $680 a month and the project may not pay for itself at all.
This is the core lesson. Draft quality drives cost far more than model price does. Moving from the small tier to the mid tier costs about $37 a month here. If the better model cuts review time per draft from 90 seconds to 30, it saves several hundred dollars of labor. Pick the model by measuring review time on your eval set, not by reading the price table.
Maintenance
Someone has to read failure logs, update the prompt when your pricing or policies change, and rerun evals. In our experience a single well-scoped agent needs a few hours of attention a month once it is stable, and noticeably more during the first two months. Hosting for the worker that runs the loop is usually small next to all of this.
Putting the month together
For the 880-run email agent on the mid tier, a realistic month looks like this:
- Model usage with caching and retry overhead: about $35-45
- Evals during steady state, through batch: about $10-20
- Human review at 30 seconds per draft: about $170
- Maintenance: a few hours of someone's time
The comparison that matters is against the current cost of the work. If a person spends three minutes per email today, that is 44 hours a month. Getting it to 7 hours of review is a strong return even after every hidden cost above. If the drafts need heavy editing, the same math says stop, fix the drafts, or narrow the agent to the email types it handles well.
How to estimate your own
- Count runs. Volume per day times working days. Be honest about peaks.
- Sketch the loop. How many model calls per run, and what gets added to the context at each step. Remember every call resends everything before it.
- Measure, do not guess. Run 20 real examples and read the token counts from the API response. Your estimate from word counts will be off.
- Price three tiers. Small, mid, and top. The spread is often 40x, and the cheap tier is sometimes good enough for classification steps.
- Time the review. Have the person who will approve outputs check 20 drafts with a stopwatch. This number decides the project.
- Add 10-20% for retries and a line for evals. Then compare the total against what the work costs today.
If the numbers work on paper, the next step is a two-week pilot on one inbox or one queue, with the review timer running from day one. That pilot will tell you more than any pricing page, and it is the same process we use when we scope an agent with a client.
Sources
- Pricing, Anthropic (Claude API docs)
- API pricing, OpenAI
- Rate limits guide, OpenAI
- Define success criteria and build evaluations, Anthropic (Claude API docs)
- Secretaries and Administrative Assistants, Occupational Outlook Handbook, U.S. Bureau of Labor Statistics
Want systems like this built for you?
We build and run data pipelines, websites, and AI automation for startups.
