Two cents looks better than four cents on a pricing spreadsheet. But what happens when the cheaper answer needs another attempt, then a person to fix it? In the hypothetical comparison below, the option with the cheaper AI calls ends up costing more to get the work done.

IN BRIEF

Compare what it costs to finish a useful task, not just to get an AI response. Include failed attempts, outside tools and the time people spend checking and fixing the result. In our hypothetical example, the 4¢ call produces a lower cost per finished task than the 2¢ call.1, 2

Illustrative graphic. The price tag and the finished job are different numbers. Price per AI attempt: 2¢ vs 4¢ — Illustrative USD charges, not current vendor prices.. Cost per accepted task: 20¢ vs 12¢ — Illustrative totals after retries, tools and human review. Second figure rounded.. Values and their context are also available as HTML below.
The price tag and the finished job are different numbers. Illustrative, not measured results. Values and their context are also available as HTML below.

The price tag and the finished job are different numbers

2¢ vs 4¢
Price per AI attempt

Illustrative USD charges, not current vendor prices.

Illustrative
20¢ vs 12¢
Cost per accepted task

Illustrative totals after retries, tools and human review. Second figure rounded.

Illustrative

Think of a team using AI to turn incoming leads into usable customer records. Success is not a paragraph saying “done.” The company must be matched correctly, the required fields must be supported, and the record must be ready in time for the next person to use it.

That distinction is central to Anthropic’s guidance on evaluating agents: what an agent says happened is different from the outcome in the actual system. Use that outcome as the unit you are paying for.2

The bill you see is not the whole bill

Anthropic’s pricing separates input, output, cache writes and cache reads. Some server-side tools add their own charges. A single average token price can therefore leave out part of the model bill before anyone even starts reviewing the answer.1

For a useful comparison, make one ledger for a fixed group of tasks. Count every attempt, every outside service charge and every minute spent checking or repairing results. Divide that total by the number of distinct tasks that pass the same acceptance rules.

Decide whether a human-fixed result counts as a success. It can, but then you are comparing assisted workflows. Include the person’s time. For an autonomous-success metric, report the human-rescued cases separately rather than quietly crediting the model.

A worked example: the cheaper call loses

These are invented configurations, not product benchmarks. Both receive the same 1,000 tasks and use the same acceptance rules. The labor assumption is $30 an hour. All figures are in USD, and the attempt counts already include retries.

Illustrative cost for the same 1,000 submitted tasksIllustrative
Cost or outcomeOption AOption B
Total AI attempts1,4001,100
Assumed price per attempt$0.02$0.04
AI charges$28$44
Tools, hosting and evaluation$32$26
Review and repair time200 minutes80 minutes
Labor at $30/hour$100$40
Total operating cost$160$110
Accepted tasks800920
Cost per accepted task$0.20About $0.12
The cheaper call produces the more expensive finished taskIllustrative inputs
Option A0.2 USD / taskOption B0.12 USD / task
View the underlying values
MeasureValue (USD / task)
Option A0.2
Option B0.1196

Option B spends $16 more on AI, but saves $60 of human work and $6 elsewhere. It also finishes more tasks successfully. That combination takes its cost per accepted result from A’s 20¢ to roughly 12¢. Nothing here proves a pricier model will deliver those gains. It shows what would have to be true.

The break-even number is more useful than the discount

At 20¢ per accepted task, B could spend $184 on its 920 successes and tie A. Its assumed bill is $110, leaving $74 of room for costs we have not included. That is a sharper question to take into a trial: will extra infrastructure or checking consume the $74?

This is also why a retry is not another success. One lead may produce three perfectly formatted responses before becoming one usable record. Count the record once. If the attempt total already includes retries, do not multiply the bill by a retry factor again.

What to measure in your own trial

Keep four things together

  1. The same mix of tasks and the same deadline for both options.
  2. The final accepted, rejected and timed-out counts, with each task counted once.
  3. All model and tool charges, including unsuccessful work.
  4. Actual checking and correction time, not just the person’s estimated hourly rate.

Use the AI workflow cost calculator for a first estimate. It uses average attempts, tokens, rates and final success. Add review and tool costs under other costs. For caching, tiered pricing or several models in one workflow, keep a separate usage ledger.

The next number worth collecting is not another price per million tokens. It is how long someone spends rescuing each result. In this example, that is where the bargain disappears.

Sources and methodology

Sources checked September 21, 2026. Dates and periods for individual figures are stated beside them.

  1. Anthropic: Claude API pricingAccessed 2026-09-21
  2. Anthropic: Demystifying evals for AI agentsAccessed 2026-09-21
Scope and assumptions

All workflow counts, per-attempt charges and labor costs in the comparison are illustrative, not vendor quotes or measured outcomes. No model was tested.

Initial setup, tax and the business damage from wrong or late results are not included. Include them when relevant.

AI-assisted research and editing. Our editorial standards.

Continue reading

1,000 Jobs, 3,000 Tasks: Why Software Bills Surprise You

86% Agreement Can Still Miss Half the AI Failures

90% Cheaper AI Cache Reads Do Not Mean a 90% Smaller Bill