news

Introducing Claude Opus 5

Anthropic shipped Claude Opus 5 on July 24 priced level with Opus 4.8, pitching near-Fable 5 intelligence at half the price. Its own Frontier-Bench v0.1 figures show more than double the predecessor's score, and cheaper per task.

Published 2026-07-24Source: Anthropic
Anthropic source artwork

Why it matters

Vendor-reported numbers, but the framing is cost per task rather than raw score, which is what finance teams actually budget against. On OSWorld 2.0, Anthropic claims Opus 5 tops Fable 5's best mark while spending a shade over a third as much.

Tokenmaxxing read

The effort setting is the routing story inside one model: dial up intelligence or conserve tokens per call. Anthropic reports roughly 1.5x the next-best pass rate on Zapier's AutomationBench for the same cost per task, and more tasks passed than any rival even at lowest effort.

Source takeaway

Customer notes point the same way: a legal team reports matching quality on 26% fewer tokens than Opus 4.8 at max reasoning, while a trading firm says it burns roughly one-seventh as many reasoning tokens at under half the latency.

Topic links

tokenmaxxingcoding-agentstopicagentstopicscoreboards
Related projects

Tools that match this angle

#4In spirit
Agents

LangGraph

langchain-ai/langgraph

A framework for building resilient stateful agents with explicit graphs, persistence, human-in-the-loop flows, and controllable execution.

39.9K6.7KMIT
agentsstateworkflows
#5Direct
Evaluation

promptfoo

promptfoo/promptfoo

A CLI and CI workflow for testing prompts, agents, and RAG systems across models, with evals and red-team style checks.

24.3K2.2KMIT
prompt-evalscirag
#6In spirit
Evaluation

DSPy

stanfordnlp/dspy

A framework for programming and optimizing language-model pipelines rather than hand-tuning one prompt at a time.

37.3K3.2KMIT
optimizationprogrammingevals
Related feed

More source-linked context

CNX Software - Embedded Systems News source artwork
newsCS
news

Token Monitor - An ESP32-S3 desktop display that tracks AI coding assistant usage (Crowdfunding) - CNX Software

Fractal Manifold is crowdfunding Token Monitor, a EUR 99 ESP32-S3 desk display with a 4-inch touchscreen that shows quota use, session limits, reset timers and estimated token costs for Claude Code, Codex CLI and Antigravity CLI.

tokenmaxxingcoding-agentsagents
Read note
theclimatebrink.com source artwork
newsT
news

The real energy use of agentic AI

Climate scientist Zeke Hausfather metered his own Claude Code habit: 1,138 typed prompts fanned out to more than 14,000 model calls and 3.2 billion tokens in eight weeks, drawing roughly 170 kWh of data-center electricity.

tokenmaxxingcoding-agentsagents
Read note
HackerNoon source artwork
newsH
newsmedium review

Claude Code Was Burning Tokens Until I Put a Gate in Front of It | HackerNoon

Prateek Kapoor wrapped Claude Code in a PreToolUse hook that rejects unbounded greps and whole-file reads, hands the agent a corrected command, and reports 75% to 80% lower token use on heavy editing tasks.

tokenmaxxingcoding-agentsagents
Read note