A single model call can make a five-to-one price gap look small in absolute dollars. An agent that plans, searches, calls tools, checks results and retries can make the same gap recur many times inside one task. That is where GPT-6.1 Sol’s pricing becomes a product-design question.1

IN BRIEF

OpenAI lists GPT-6.1 Sol at one-fifth of Astra’s standard API input and output token prices. The distinct business question is repeated agent work: when one task requires planning, tool calls, checking and retries, a lower per-call price can recur across the loop. Token price still is not total task cost.1, 2

The agent-cost frontier. Sol input: $2 — Launch comparison price per 1 million standard input tokens.. Sol output: $10 — Launch comparison price per 1 million standard output tokens.. Price reduction: 80% — Calculated reduction from Astra's $10/$50 launch comparison rates to Sol's $2/$10 rates.. Values and their context are also available as HTML below.
The agent-cost frontier. Values and their context are also available as HTML below.1
Standard API token prices in OpenAI’s launch comparison1
ModelInput / 1MOutput / 1M
GPT-6 Astra$10$50
GPT-6.1 Sol$2$10

The price gap compounds across a loop

At the launch comparison rates, Sol input and output tokens cost one-fifth as much as Astra’s. If two models consumed exactly the same tokens across exactly the same calls, the model-token portion would be 80% lower. Real workloads rarely hold those variables constant.1

The agent-cost frontier

$2
Sol input1

Launch comparison price per 1 million standard input tokens.

$10
Sol output1

Launch comparison price per 1 million standard output tokens.

80%
Price reduction1

Calculated reduction from Astra's $10/$50 launch comparison rates to Sol's $2/$10 rates.

Cheaper calls can support more checking, not just more volume

A lower model price can be spent in different ways: run more tasks, add verification steps, keep more context, or reserve Astra for cases where the more expensive model materially improves completion quality. The economically useful unit is still cost per successful task.1

That is why this story is narrower than our earlier GPT-6 price-compression analysis. The new question is how a near-Astra model at one-fifth of the launch comparison price changes multi-call agent design, not whether AI inference is getting cheaper in general.

Cached context changes the arithmetic again

OpenAI’s current pricing page lists separate cached-input and processing-tier prices. Agent workloads that repeatedly reuse a large stable context can therefore have a different cost structure from a one-shot prompt. A headline input price is only one line in the bill.1

Four costs to keep in the loop

  • Fresh input tokens sent on each model call.
  • Cached or reused context under the applicable pricing tier.
  • Output tokens generated across planning, tool use and verification.
  • Retries, external tools, storage and human review needed to finish the task.

The important change is not that every Astra workload should move to Sol. It is that repeated-agent workflows now have another point on the cost-versus-capability curve, and the five-to-one token-price gap can repeat many times inside one completed job.

Sources and methodology

Sources checked September 29, 2026. Dates and periods for individual figures are stated beside them.

  1. OpenAI API pricing ↗Accessed 2026-09-29
  2. AWS: GPT-6.1 Sol on Amazon Bedrock ↗Accessed 2026-09-29
Scope and assumptions

The one-fifth comparison uses the standard rates presented with the launch and does not include every processing tier or long-context case.

Lower token price does not establish equal task quality, equal token consumption or lower total application cost.

Continue reading

OpenAI Cut Its New Model Prices in Half. What Gets Cheaper When Intelligence Does? →

2¢ vs 4¢: When the Cheaper AI Call Costs More →

90% Cheaper AI Cache Reads Do Not Mean a 90% Smaller Bill →