Two cents looks better than four cents on a pricing spreadsheet. But what happens when the cheaper answer needs another attempt, then a person to fix it? In the hypothetical comparison below, the option with the cheaper AI calls ends up costing more to get the work done.
Compare what it costs to finish a useful task, not just to get an AI response. Include failed attempts, outside tools and the time people spend checking and fixing the result. In our hypothetical example, the 4¢ call produces a lower cost per finished task than the 2¢ call.1, 2

The price tag and the finished job are different numbers
Illustrative USD charges, not current vendor prices.
IllustrativeIllustrative totals after retries, tools and human review. Second figure rounded.
IllustrativeThink of a team using AI to turn incoming leads into usable customer records. Success is not a paragraph saying “done.” The company must be matched correctly, the required fields must be supported, and the record must be ready in time for the next person to use it.
That distinction is central to Anthropic’s guidance on evaluating agents: what an agent says happened is different from the outcome in the actual system. Use that outcome as the unit you are paying for.2
The bill you see is not the whole bill
Anthropic’s pricing separates input, output, cache writes and cache reads. Some server-side tools add their own charges. A single average token price can therefore leave out part of the model bill before anyone even starts reviewing the answer.1
For a useful comparison, make one ledger for a fixed group of tasks. Count every attempt, every outside service charge and every minute spent checking or repairing results. Divide that total by the number of distinct tasks that pass the same acceptance rules.
Decide whether a human-fixed result counts as a success. It can, but then you are comparing assisted workflows. Include the person’s time. For an autonomous-success metric, report the human-rescued cases separately rather than quietly crediting the model.
A worked example: the cheaper call loses
These are invented configurations, not product benchmarks. Both receive the same 1,000 tasks and use the same acceptance rules. The labor assumption is $30 an hour. All figures are in USD, and the attempt counts already include retries.
| Cost or outcome | Option A | Option B |
|---|---|---|
| Total AI attempts | 1,400 | 1,100 |
| Assumed price per attempt | $0.02 | $0.04 |
| AI charges | $28 | $44 |
| Tools, hosting and evaluation | $32 | $26 |
| Review and repair time | 200 minutes | 80 minutes |
| Labor at $30/hour | $100 | $40 |
| Total operating cost | $160 | $110 |
| Accepted tasks | 800 | 920 |
| Cost per accepted task | $0.20 | About $0.12 |
View the underlying values
| Measure | Value (USD / task) |
|---|---|
| Option A | 0.2 |
| Option B | 0.1196 |
Option B spends $16 more on AI, but saves $60 of human work and $6 elsewhere. It also finishes more tasks successfully. That combination takes its cost per accepted result from A’s 20¢ to roughly 12¢. Nothing here proves a pricier model will deliver those gains. It shows what would have to be true.
The break-even number is more useful than the discount
At 20¢ per accepted task, B could spend $184 on its 920 successes and tie A. Its assumed bill is $110, leaving $74 of room for costs we have not included. That is a sharper question to take into a trial: will extra infrastructure or checking consume the $74?
This is also why a retry is not another success. One lead may produce three perfectly formatted responses before becoming one usable record. Count the record once. If the attempt total already includes retries, do not multiply the bill by a retry factor again.
What to measure in your own trial
Keep four things together
- The same mix of tasks and the same deadline for both options.
- The final accepted, rejected and timed-out counts, with each task counted once.
- All model and tool charges, including unsuccessful work.
- Actual checking and correction time, not just the person’s estimated hourly rate.
Use the AI workflow cost calculator for a first estimate. It uses average attempts, tokens, rates and final success. Add review and tool costs under other costs. For caching, tiered pricing or several models in one workflow, keep a separate usage ledger.
The next number worth collecting is not another price per million tokens. It is how long someone spends rescuing each result. In this example, that is where the bargain disappears.
Sources and methodology
Sources checked September 21, 2026. Dates and periods for individual figures are stated beside them.
- Anthropic: Claude API pricing ↗Accessed 2026-09-21
- Anthropic: Demystifying evals for AI agents ↗Accessed 2026-09-21
Scope and assumptions
All workflow counts, per-attempt charges and labor costs in the comparison are illustrative, not vendor quotes or measured outcomes. No model was tested.
Initial setup, tax and the business damage from wrong or late results are not included. Include them when relevant.
AI-assisted research and editing. Our editorial standards.
Continue reading
1,000 Jobs, 3,000 Tasks: Why Software Bills Surprise You →