Anthropic released Claude Opus 5.5 today: the first model in the Claude 5.5 family. They say it performs at the level of Claude Fable 5.1 on most work and costs about 40% less to run than Opus 5.
Same question we asked for Sol, Astra, and Grok 4.7: where does this remove a real constraint, and where does it just raise the bill?
A cheaper frontier engine is useful. A production system is still the job.
What shipped (FACT)
From Anthropic’s announcement (https://www.anthropic.com/claude-opus-5-5):
- First Claude 5.5 family model. API id:
claude-opus-5-5 - Anthropic claim: Fable 5.1 level on most work; about 40% less to run than Opus 5
- Pricing per 1M vs Opus 5: input $4 (vs $5), output $20 (vs $25), cache reads $0.20 (vs $0.50), cache writes $5 (vs $6.25)
- Input/output about 20% cheaper; cache reads about 60% cheaper; output more than 30% faster
- Fast mode (Claude Code + Claude Platform): up to 2.5x speed at $8 / $40 per 1M
- Available on platforms they list, including AWS, Google Cloud, and Microsoft Azure
- Context window: not on the announcement page. We will not invent a number.
Anthropic’s benchmark table (generally adaptive thinking at max effort; Terminal-Bench for 5.5 at xhigh). Their caveat: margins at this level are a weaker guide to real-world differences, and the gap with Fable 5.1 is narrower than the scores suggest. Production safeguards on; cyber often fell back to Opus 4.8.
| Benchmark | Opus 5.5 | Fable 5.1 | Opus 5 |
|---|---|---|---|
| Terminal-Bench 4.0 | 66.4% | 55.8% | 52.3% |
| FrontierCode v1.1 | 54.4% | 50.3% | 48.0% |
| CursorBench 4.0 | 57.8% | 51.8% | 46.6% |
| GDPval-AA v2.1 | 1846 | 1735 | 1708 |
| AutomationBench | 40.0% | 31.4% | 26.9% |
AutomationBench also lists GPT-6 Astra at 41.4%. “Best on every row” is not the operator call.
Attributed Anthropic anecdotes (theirs, not ours): 680k-line migration under a day; web-app load cuts 39/40; 200k-line audit under 3h vs more than 20h for Opus 5 at about 2.5x fewer tokens; HAProxy C-to-Rust 9.5h vs 12 for Fable 5.1 at 51% less cost; report-gen 16/18 cleared their quality bar; merger analysis 63 vs 93 min for Opus 5 at 50% less cost.
Safety (Anthropic): strongest on their automated behavioral audit; about 85% fewer containment-circumvention attempts vs Opus 5 and Mythos 5.1; Fable-like cyber/bio/distillation safeguards with transparent fallback; preserved thinking; no thinking-off mode; ZDR available.
What this changes for operators
Price / performance. If the 40% cost claim holds on your cache-heavy agent runs, Opus 5.5 is a real candidate to displace Opus 5. Still measure Fable 5.1. Do not assume the leaderboard gap is your backlog gap.
Harness and fixtures. CursorBench and Terminal-Bench are harness-bound. Score in the harness you will ship.
Token burn and effort. Cheaper per token is not cheaper per finished job. Fast mode buys speed. Max effort buys scores. Put a success number on the first workflow before you make it the default.
Human gates. Better coding does not replace a person on merge, send, spend, or deploy.
Exception paths. Today’s other VL note: design skip / ask a human / refuse before you demo the happy path.
Do not flip every bot on Day 1. Candidate upgrade. Side-by-side on one painful workflow until the metric moves.
What Viking Labs is doing with this
We already run a gated fleet: specialist bots, CoS send gate, human merge and deploy. Opus 5.5 is a candidate engine for long coding and research-week jobs where the harness is ours and the success metric is written first.
Next step for us (and for you): one workflow, clear done state, side-by-side fixtures, score quality / steps / tokens / human edits, keep the gate and the three exception paths.
Offer mix when you want that discipline on a real job: DFY agents typically $2k to $5k+, productized automations $197 to $297/mo, custom software when a template will not fit. ResearcherFlow is production proof we run ourselves (live billing, permissions, human gates on writes), not the only offer.
What we are not copying
- Leaderboard theater without your own fixture
- Anthropic’s table as independent proof without their caveat
- Invented context window
- Ownership-free “autonomous company” claims
- Default Fast mode or max effort with no cost cap
- Fake VL Day 1 win rates
Claude Opus 5.5 is a cheaper frontier engine. It still needs an owner, a harness, a success number, and exception paths that keep that number honest.
If you want help wrapping agents around one workflow you can measure: https://vikinglabs.com Approve-before-write research weeks: https://researcherflow.com
Jon Marrs Viking Labs https://vikinglabs.com
Leave a Reply