Cache Reads Down 75%: Long-Running Agents Just Got Affordable

Published on: 2026-09-04

Cache Reads Down 75%: Long-Running Agents Just Got Affordable

📖 Glossary

AI Box (also known as Agent Computer / Agent PC), is a dedicated local hardware device that runs AI Agents. Pre-installed with an AI agent management system, plug-and-play, running 24/7. Users can remotely command AI to work via Discord, Slack, Telegram, WhatsApp, and more.

You set an AI agent to watch your project overnight. The next morning you open the bill. The money spent on actual thinking is smaller than the money spent re-reading tens of thousands of lines of code, your system prompt, and the conversation history.

This is the shared pain of anyone running long-horizon tasks. Every step the model takes, it re-reads the context. Cache-read billing quietly eats the largest slice of the invoice.

Figure

The cloud model is responsible for getting smarter. The local base is responsible for staying online. A device like the Kaihe AIBOX, an agent computer, stays on 7x24 the moment you plug it in, running your private agent with memory stored on your own hardware. The cheaper the model gets, the more sense this base makes.

The misconception: stronger models must cost more

Most people assume that when capability goes up, price goes up. This generation of flagship models breaks that logic.

On September 1, Anthropic released Fable 5.1, a broadly available new flagship, alongside Mythos 5.1, a restricted variant on the same base for high-sensitivity scenarios. It did something counterintuitive: capability went up, while the most-used part of the cost went down.

The one-line essence: the cost structure of long tasks changed

Input and output prices stayed flat, at 10 dollars and 50 dollars per million tokens. The cut landed on cache reads: from 1 dollar per million tokens down to 0.25 dollars, a 75% drop.

Anthropic estimates typical workloads cost about 25% less, and heavily agentic tasks save up to 45%. This is not pocket change. It is the recurring cost of re-reading context across long tasks.

Why cache reads matter so much for agents

Prompt caching lets a model reuse context it has already processed instead of recomputing it. For a normal question and answer, you pay once. For an agent that loops through tools, reads files, and checks results, the same code repository, system instructions, and history get sent back on every step. Anthropic says cached material can account for more than half of token use in long tasks. That is exactly where the 75% cut lands, and why the savings compound for always-on agents.

Three dimensions

On performance, the jump is real. Terminal-Bench 4.0, which measures terminal coding, moved from 42.0% to 55.8%. Terminal-Bench-Science 0.1, which tests autonomous research tasks, more than doubled from 24.7% to 52.6%. The knowledge-work benchmark GDPval-AA v2 also rose by 130 points. The model is better at long-chain development across entire codebases, at systematic code review, and at running code, running experiments, and analyzing data on its own. On SWE-bench Verified it reaches 95.0%, and on AutomationBench, a test of real-world multi-step agentic work, it nearly doubled from 17.1% to 31.4%.

On cost, the cut hits the vital point. For enterprises running agents, cache reads often account for more than half of token usage. This single move thins the bill for always-on agent tasks.

On data, there is also relief. The new Enterprise Frontier Safeguards let customers keep monitoring data in their own infrastructure, with zero data retention. No matter how strong the cloud model gets, your customer memory, your conversation history, and your long-term preferences belong to you only when they sit on your own Kaihe AIBOX.

Figure

A note on the restricted sibling

Mythos 5.1 shares the same underlying model but ships lighter guardrails for vetted cybersecurity and life-sciences organizations through a trusted-access program. Alongside it, Anthropic cut safety false positives: cybersecurity filter false positives dropped 60%, and biology filter false positives dropped 85%. For legitimate researchers who kept hitting blocks on reasonable queries, the public model is noticeably more usable. Fable 5.1 also carries invisible watermarks in text and file outputs to meet the EU AI Act.

Compare with the previous generation

Item Fable 5 Fable 5.1
Input price (per million tokens) 10 USD 10 USD
Output price (per million tokens) 50 USD 50 USD
Cache read (per million tokens) 1 USD 0.25 USD
Terminal-Bench 4.0 42.0% 55.8%
Terminal-Bench-Science 0.1 24.7% 52.6%

Prices held. Cache reads cut 75%. Benchmarks up almost across the board. That is the entire point of this update.

Stronger, cheaper models make the base more valuable

Running a long-online private agent used to mean two expensive bills: the model and the base. Now the model side cut cache reads by 75%, and long-task cost drops directly. What remains is a device that stays on.

Figure

The Kaihe AIBOX is that base. Plug it in and it is online. OpenClaw and Hermes come pre-installed. The longer you use it, the more it knows you, and all memory stays in your own hands. The cloud model decides how strong. The local base decides staying online.

Stronger and cheaper at the same time is a rare combination. For anyone who wants to actually use long-running agents, the timing is good.


Related reading: - Models keep racing to do the work, but what you lack is a base that stays online - Voice AI can fool human ears: what you lack is not speech, but an assistant that stays awake - Xiaohongshu open-sources a 280B model: the more models race, the more you need a butler that dispatches for you

KaiheAIBOX #AIFrontier #CacheRead #AgentComputer #AIBOX

Contact: [email protected]

KAIHE AIBOX · 7x24 Agent Computer | AI Frontier

Recommended Products

A1 Home Entry A1 Pro Enhanced A2 Professional A2 Pro Advanced X1 Enterprise G1 Flagship
© KAIHE AI - Agent Computer Specialist