Field notes

Claude Opus 5.5 is a cheaper frontier engine. It still needs an owner.

Anthropic released Claude Opus 5.5 today: the first model in the Claude 5.5 family. They say it performs at the level of Claude Fable 5.1 on most work and costs about 40% less to run than Opus 5.

Same question we asked for Sol, Astra, and Grok 4.7: where does this remove a real constraint, and where does it just raise the bill?

A cheaper frontier engine is useful. A production system is still the job.

What shipped (FACT)

From Anthropic’s announcement (https://www.anthropic.com/claude-opus-5-5):

  • First Claude 5.5 family model. API id: claude-opus-5-5
  • Anthropic claim: Fable 5.1 level on most work; about 40% less to run than Opus 5
  • Pricing per 1M vs Opus 5: input $4 (vs $5), output $20 (vs $25), cache reads $0.20 (vs $0.50), cache writes $5 (vs $6.25)
  • Input/output about 20% cheaper; cache reads about 60% cheaper; output more than 30% faster
  • Fast mode (Claude Code + Claude Platform): up to 2.5x speed at $8 / $40 per 1M
  • Available on platforms they list, including AWS, Google Cloud, and Microsoft Azure
  • Context window: not on the announcement page. We will not invent a number.

Anthropic’s benchmark table (generally adaptive thinking at max effort; Terminal-Bench for 5.5 at xhigh). Their caveat: margins at this level are a weaker guide to real-world differences, and the gap with Fable 5.1 is narrower than the scores suggest. Production safeguards on; cyber often fell back to Opus 4.8.

Benchmark Opus 5.5 Fable 5.1 Opus 5
Terminal-Bench 4.066.4%55.8%52.3%
FrontierCode v1.154.4%50.3%48.0%
CursorBench 4.057.8%51.8%46.6%
GDPval-AA v2.1184617351708
AutomationBench40.0%31.4%26.9%

AutomationBench also lists GPT-6 Astra at 41.4%. “Best on every row” is not the operator call.

Attributed Anthropic anecdotes (theirs, not ours): 680k-line migration under a day; web-app load cuts 39/40; 200k-line audit under 3h vs more than 20h for Opus 5 at about 2.5x fewer tokens; HAProxy C-to-Rust 9.5h vs 12 for Fable 5.1 at 51% less cost; report-gen 16/18 cleared their quality bar; merger analysis 63 vs 93 min for Opus 5 at 50% less cost.

Safety (Anthropic): strongest on their automated behavioral audit; about 85% fewer containment-circumvention attempts vs Opus 5 and Mythos 5.1; Fable-like cyber/bio/distillation safeguards with transparent fallback; preserved thinking; no thinking-off mode; ZDR available.

What this changes for operators

Price / performance. If the 40% cost claim holds on your cache-heavy agent runs, Opus 5.5 is a real candidate to displace Opus 5. Still measure Fable 5.1. Do not assume the leaderboard gap is your backlog gap.

Harness and fixtures. CursorBench and Terminal-Bench are harness-bound. Score in the harness you will ship.

Token burn and effort. Cheaper per token is not cheaper per finished job. Fast mode buys speed. Max effort buys scores. Put a success number on the first workflow before you make it the default.

Human gates. Better coding does not replace a person on merge, send, spend, or deploy.

Exception paths. Today’s other VL note: design skip / ask a human / refuse before you demo the happy path.

Do not flip every bot on Day 1. Candidate upgrade. Side-by-side on one painful workflow until the metric moves.

What Viking Labs is doing with this

We already run a gated fleet: specialist bots, CoS send gate, human merge and deploy. Opus 5.5 is a candidate engine for long coding and research-week jobs where the harness is ours and the success metric is written first.

Next step for us (and for you): one workflow, clear done state, side-by-side fixtures, score quality / steps / tokens / human edits, keep the gate and the three exception paths.

Offer mix when you want that discipline on a real job: DFY agents typically $2k to $5k+, productized automations $197 to $297/mo, custom software when a template will not fit. ResearcherFlow is production proof we run ourselves (live billing, permissions, human gates on writes), not the only offer.

What we are not copying

  • Leaderboard theater without your own fixture
  • Anthropic’s table as independent proof without their caveat
  • Invented context window
  • Ownership-free “autonomous company” claims
  • Default Fast mode or max effort with no cost cap
  • Fake VL Day 1 win rates

Claude Opus 5.5 is a cheaper frontier engine. It still needs an owner, a harness, a success number, and exception paths that keep that number honest.

If you want help wrapping agents around one workflow you can measure: https://vikinglabs.com Approve-before-write research weeks: https://researcherflow.com

Jon Marrs Viking Labs https://vikinglabs.com

Leave a Reply

Discover more from Viking Labs

Subscribe now to keep reading and get access to the full archive.

Continue reading

Map Your First Build