TL;DR
- GPT-6.1 Sol Ultrafast (rolling out Oct 9, 2026) is OpenAI's speed tier: up to 8x faster than standard Sol, at $12 input / $60 output per million tokens, 6x the standard price.
- Claude Opus 5.5 is $4 / $20. Its fast mode is up to 2.5x faster at $8 / $40, 2x the standard price.
- At standard speed, Opus 5.5 is the faster model: about 95 vs 55 output tokens per second on Artificial Analysis. On the speed tiers it flips. Taking each vendor's ceiling, Ultrafast is up to about 440 t/s vs about 240 for Opus fast mode.
- Ultrafast costs exactly 1.5x Opus fast mode on every line of the bill (input, cached input and output), for up to about 1.8x the speed.
- Nonstop output for an hour costs about $95 on Ultrafast, $35 on Opus fast mode, $7 on Opus 5.5 and $2 on standard Sol.
Watch the speeds
Numbers per second are hard to feel. Below, each setup writes the same 1,500 tokens of code in real time at its speed, with the output cost ticking up as it goes.
Ultrafast finishes in about 3 seconds and standard Sol takes about 27. The cost column is the other half of the story: the fastest lane spends six times as much per token as the slowest.
What launched
On October 9, 2026, OpenAI started rolling out Ultrafast for GPT-6.1 Sol in the API, Codex and ChatGPT Work. OpenAI's pitch is "near-Astra intelligence at up to 8x faster speeds than Sol Standard". In Codex, Ultrafast sits on the $500 Pro plan. In the API it is a service tier with its own price row: $12 per million input tokens, $0.60 cached, $60 output.
The closest thing Anthropic sells is fast mode for Claude Opus 5.5, a research preview on the Claude API: the same model weights on a faster inference setup, up to 2.5x higher output tokens per second, at $8 input and $40 output. Fast mode needs access from Anthropic and is not on Bedrock, Google Cloud or Microsoft Foundry.
The price sheet
Prices are per million tokens, short context, checked on both official pricing pages on October 9, 2026. Standard speeds are Artificial Analysis measurements at max effort. The fast-tier speeds are not measurements: they are the standard speed multiplied by each vendor's own "up to" figure.
| Setup | Input | Cached input | Output | Speed (output t/s) |
|---|---|---|---|---|
| GPT-6.1 Sol | $2.00 | $0.10 | $10 | ~55 |
| Claude Opus 5.5 | $4.00 | $0.20 | $20 | ~95 |
| Opus 5.5 fast mode | $8.00 | $0.40 | $40 | up to ~238 |
| Sol Ultrafast | $12.00 | $0.60 | $60 | up to ~440 |
What you pay for speed
Both labs charge a premium that grows a little slower than the speed it buys. OpenAI's tier is 8x the speed for 6x the price, Anthropic's is 2.5x for 2x. Per dollar of premium the deals are almost identical: 1.33x speed per 1x of price for Ultrafast and 1.25x for Opus fast mode.
The difference is the starting point. Standard Sol is the cheaper and slower base, so Ultrafast stretches further: it ends up about 1.8x faster than Opus fast mode for 1.5x the price. Because every price line scales by the same multiple, that 1.5x holds for input, cached input and output alike.
One 20,000-token job
Say an agent has to write 20,000 tokens of output, for example a large file rewrite. Here is the streaming time and output cost on each setup. Input and fixed wait are left out to keep the comparison clean.
| Setup | Streaming time | Output cost | Seconds saved vs standard Sol, per extra dollar |
|---|---|---|---|
| GPT-6.1 Sol | 6.1 min | $0.20 | baseline |
| Claude Opus 5.5 | 3.5 min | $0.40 | 766 s per $ |
| Opus 5.5 fast mode | 1.4 min | $0.80 | 466 s per $ |
| Sol Ultrafast | 0.8 min | $1.20 | 318 s per $ |
Opus 5.5 at standard speed buys the most time per extra dollar against standard Sol, because it is already faster and only twice the price. Past that point each step costs more per second saved.
An agent request is mostly input
Agent loops read far more than they write. For an illustrative request with 60,000 input tokens, 90% of them cache hits, and a 500-token reply, the output is the small part of the bill:
Ultrafast multiplies the input side by 6x too, so an Ultrafast loop costs about 6x a standard Sol loop no matter how short the replies are. And since a short reply spends most of its wall time on the fixed wait, not on streaming, the speedup you feel in a loop of short turns is smaller than the headline multiplier.
Caveats
- The fast-tier speeds are ceilings. "Up to 8x" and "up to 2.5x" are vendor claims. We multiplied them onto measured standard speeds. Independent measurements of Ultrafast on Sol were not published when we checked.
- Tokens are not the same size. Anthropic says Claude 4.7 and later produce about 30% more tokens for the same text. So 95 Claude tokens per second is closer to about 73 per second in same-text terms. Speed comparisons across vendors are approximate.
- Effort changes everything. The standard speeds are at max effort. Lower reasoning effort writes fewer hidden reasoning tokens, so jobs finish sooner even at the same tokens per second.
- Intelligence is not equal. Artificial Analysis scores Opus 5.5 ahead of GPT-6.1 Sol at max effort (58 vs 52 on its intelligence index). OpenAI describes Ultrafast Sol as near GPT-6 Astra.
- Prices change often. Everything here was checked on October 9, 2026.
Which one to use
- Background agents and batch work: standard Sol or Opus 5.5. Nobody is waiting, so speed is not worth paying for.
- Long interactive output (big rewrites, generated docs, code you watch stream in): this is where Ultrafast's ceiling pays off, if your plan or API tier has it.
- Short-turn agent loops: the fixed wait per request dominates, so measure before paying 6x. Opus 5.5 at standard speed is already faster than standard Sol.
- Mixed fleets: route by job. Keep the expensive speed tier for the step a person is watching and run everything else on standard tiers.
What this means for an AI employee
An AI employee that runs scheduled work overnight does not need Ultrafast. One that sits in a live chat with your customer might. We pick the model and the speed tier per step of the workflow, then check the bill against the logs. Talk to us if you want that routing designed around your workflow.
Sources
- OpenAI Developers: Ultrafast for GPT-6.1 Sol (Oct 9, 2026)
- OpenAI API pricing (Standard and Ultrafast rows for gpt-6.1-sol)
- Anthropic API pricing
- Anthropic fast mode docs (up to 2.5x output tokens per second)
- Artificial Analysis: GPT-6.1 Sol and Claude Opus 5.5 (speed and intelligence)
- VentureBeat on GPT-6.1 Sol and Ultrafast


