AI

Anthropic’s Fable 5.1 cuts cache costs 75% — agentic scores more than doubled

Susan Hill

Anthropic updated Claude on September 1 with Fable 5.1, cutting the cost of cached input tokens from $1.00 to $0.25 per million — a 75% reduction that arrives at the moment most production AI developers have discovered that context accumulation, not raw computation, is where their infrastructure budgets actually concentrate. In a live agentic session, a single Claude conversation can hold hundreds of thousands of tokens in cache as context builds across turns. At $0.25 per million, that accumulation becomes a deliberate design choice rather than a cost ceiling a developer works around.

The benchmark results accompanying the price change show performance gains that compound the economic case. On Terminal-Bench-Science 0.1, which measures multi-step scientific reasoning requiring experimental design and iterative analysis, Fable 5.1 scored 52.6% against Fable 5’s 24.7% — more than doubling the previous result. AutomationBench improved from 17.1% to 31.4%, and Terminal-Bench 4.0, a broader agentic capability evaluation, rose from 42.0% to 55.8%. CursorBench 3.2.0 came in at 73.4%. The pattern across all four evaluations is consistent: Fable 5.1 is not incrementally better at the margin, it is substantially more reliable on the categories of tasks that define agentic value.

The combined effect rewrites the cost arithmetic of building with Claude. An agentic application that completes a task in fewer turns — because the model is more reliable — consumes fewer total token cycles on the way to a successful result. Apply the 75% cache cut to the tokens it does use, and the effective per-task cost reduction in high-reuse agentic workflows reaches 45% or more. Anthropic has not published a single headline number for the compound savings, but developers who model their own token economics will arrive at figures in that range for context-heavy workloads.

Fable 5.1 is available immediately through Anthropic’s direct API under the model identifier claude-fable-5-1, and through Amazon Web Services Bedrock, Google Cloud Platform, and Microsoft Azure. Anthropic announced Enterprise Frontier Safeguards as a concurrent capability: customer-controlled data retention, customer-managed encryption keys, and deployment within each customer’s existing cloud infrastructure tenant, included at no additional charge beyond the underlying cloud costs.

Alongside the general release, Anthropic introduced Mythos 5.1, a variant of the same underlying model operating with what the company describes as more permissive safeguards for vetted organizations. Access is restricted to verified cybersecurity research institutions and life-sciences companies. The company cited demonstrated applications including protein binder design and biological model optimization at up to 2.5 times the inference throughput of the standard release. Mythos 5.1 access is not self-service: Anthropic reviews applications individually and has not published the size of the enrolled cohort.

The Fable 5.1 agentic numbers, particularly on Terminal-Bench-Science, put Anthropic’s published results ahead of OpenAI‘s most recently disclosed figures on comparable evaluations. Cross-company benchmark comparisons carry methodological caution — vendors select evaluations that favor their models, and the specific tests Anthropic highlighted were presumably chosen for their favorable outcomes. What is harder to explain away is the internal doubling pattern: Fable 5.1 is not the same model with a price reduction attached. The performance change is large enough that the pricing story and the performance story cannot be separated.

For developers who do not currently use Claude, the cache pricing move is the number worth translating into their own token economics. Any AI provider that does not approach $0.25 per million cached tokens will now face direct cost comparison questions in enterprise sales conversations. Anthropic’s cache price was already lower than several competitors before this reduction; the gap has widened. Performance improvements are harder to generalize across use cases, but in the agentic and scientific-reasoning segment, Fable 5.1 sets a benchmark level that the next major release from any provider will be measured against.

Tags: , , , ,

Discussion

There are 0 comments.