On July 24, Anthropic dropped a quiet bombshell.
No advance notice. No press briefing. Claude Opus 5 just appeared—and my Claude Max default model switched to it instantly.
Then I checked the benchmark table. I had to do a double take.
This isn't "close to flagship" performance. Opus 5 dominated nearly every benchmark: Frontier-Bench coding at 43.3% (Fable 5 scored 33.7%), GDPval-AA knowledge work at 1861 (Fable 5: 1747), SWE-bench Verified at 96.0%. And on ARC-AGI 3—the benchmark designed to test models on completely unseen problems—Opus 5 scored 30.2%, while the runner-up, GPT-5.6 Sol, managed just 7.8%.
That's not a margin. That's a threefold gap.
The price? $5 per million input tokens, $25 per million output—identical to the previous Opus 4.8, and exactly half of Fable 5's $10/$50 pricing.
A model that beats the flagship, sold at the previous generation's price. The "you get what you pay for" model of AI pricing is being demolished—by Anthropic itself.

Why "Half-Price Flagship" Isn't Just a Rebrand
📖 Glossary
AI Box (also known as Agent Computer / Agent PC), is a dedicated local hardware device that runs AI Agents. Pre-installed with an AI agent management system, plug-and-play, running 24/7. Users can remotely command AI to work via Discord, Slack, Telegram, WhatsApp, and more.
The obvious first reaction: isn't this just Fable 5 with a new name and a discount?
No. Opus 5 underwent a fundamental architectural reconstruction. The clearest evidence: Fable 5 never received scaled reinforcement learning training. Researcher Nathan Lambert analyzed that Fable 5's sheer scale makes it currently impractical to run full RL pipelines on it. Opus 5, through architectural optimization, runs complete RL on a more compact model.
This makes Opus 5 not just cheaper, but genuinely "better to use" in many scenarios.
Consider the official demos. Given a mechanical part blueprint and tasked with writing FreeCAD code to reconstruct a 3D model—but explicitly prevented from directly "viewing" the image—Opus 5 built its own computer vision pipeline, extracted geometric shapes from raw pixels, and reconstructed the entire part. Other models failed five out of five attempts.
Or a real open-source package manager bug where the community patch missed an edge case. Opus 5 identified the root cause and fixed the gap; the comparison model patched the surface symptom and declared the bug "resolved."
An intriguing quirk: on the FrontierCode benchmark, Opus 5 scored 53.4% at medium thinking effort. Crank the effort higher, and the score dropped to 48.0%. Overthinking actually made it worse—like an overanxious test-taker.
45 Days to Dethrone Your Own Flagship
Fable 5 launched on June 9 as Anthropic's first publicly available Mythos-class model, priced at $10/$50 per million tokens—the most expensive general-purpose model in history.
Between June 9 and July 24: 45 days. That's all it took for Anthropic to release a model that outperforms Fable 5 at half the cost.
This means AI model pricing no longer follows the "capability tier" system. Previously, choosing a model meant picking from a price menu: Haiku cheapest, Sonnet mid-range, Opus expensive, Fable/Mythos flagship. Price and performance scaled linearly.
Opus 5 broke that linear relationship—drawing a new curve between price and performance, one far more favorable to consumers.
Anthropic's own analogy fits perfectly. Haiku/Sonnet are your V6 family engines. Opus is a high-performance V8. Fable/Mythos are F1 racing engines. Now this V8 is matching F1 lap times while burning fuel at V6 rates.
OpenAI's GPT-5.6 and Google's Gemini 3.1 now face an uncomfortable question: if "half-price flagship" becomes the market standard, how long can your pricing hold?

Cheaper Models Make 24/7 AI Agents More Critical Than Ever
There's a counterintuitive twist here.
People used to say AI is too expensive—a tool for the privileged few. But when Opus 5 halves the price while increasing capability, AI is becoming infrastructure. Like electricity. Like running water.
Which raises a new question: electricity got cheaper, but you still needed a circuit breaker panel. Water got cheaper, but you still needed a faucet.
Models are getting cheaper—but who's going to use them for you, 24 hours a day?
This is where local AI assistants become indispensable.
A Kaihe AIBOX sits at home, plugged in and running on minimal power, no computer required. The Hermes Agent on it operates 24/7—processing emails, organizing reports, running scheduled queries, and pushing results to you while you sleep. You send a WeChat message to your agent, it fires up Opus 5 (or whichever model you prefer), and gets to work. You wake up to results.
All interaction history, preference data, and work output stays on local storage—nothing uploaded to the cloud. Your AI gets smarter the more you use it, and all that "knowledge about you" belongs to you alone.
Someone might argue that "cloud AI is more convenient." But when model pricing drops to a few dollars per million tokens, the real bottleneck isn't affordability. It's whether you have a platform that automatically uses these cheap models on your behalf, 24/7.
Open laptop, launch browser, log in, type prompt—fine a few times a day. What about dozens of times? And who's doing it while you sleep?
"No computer needed, 24/7 online, private local storage, persistent memory, ultra-low power"—these five characteristics together form an AI system that actually saves you time.
The Pricing Model Is Crumbling—What Comes Next
Opus 5 isn't the endgame. It's a signal.
When Anthropic's own mid-tier model outperforms its flagship, it means AI capability improvement has outpaced the pricing system's ability to keep up. Fable 6 and Mythos 6 are almost certainly in preparation—but how long can those flagship price tags survive?
The open-source front is equally restless. Kimi K3's full weights just went public, letting any developer run a near-flagship model locally. OpenAI slashed GPT-5.6 Luna pricing by 80%, only to be outflanked by Opus 5 delivering more for less.
The curve points in one direction: downward.
For the average person, the good news is you'll never have to agonize over "can I afford the best model." The bad news: if you don't have a platform that automatically runs these increasingly cheap models for you around the clock, the "cheap" part doesn't really matter.
What you need isn't a cheaper API key. It's a box with your personal AI inside, running forever at home, remembering everything about you, ready at a moment's notice.
The difference between a few dollars per token matters far less than the difference between "nobody using AI for you" and "someone using AI for you."
Further Reading: - How to Keep AI Working for You After Shutting Down Your Computer - Send WeChat Files to Your Home AI for Automatic Voice Briefings While Commuting - Codex vs Hermes vs WorkBuddy vs Kaihe AIBOX Explained
Contact: [email protected]