Deepseek Cache Price Cut 60 Local Scheduler Kaihe Aibox

Published on: 2026-09-09

When models get cheaper by the day, KAIHE AIBOX makes one thing clear: you need a local scheduler, not a price table.

DeepSeek Just Cut Cache Pricing by 60%: The Cheaper Models Get, the More You Need a Local Scheduler

📖 Glossary

AI Box (also known as Agent Computer / Agent PC), is a dedicated local hardware device that runs AI Agents. Pre-installed with an AI agent management system, plug-and-play, running 24/7. Users can remotely command AI to work via Discord, Slack, Telegram, WhatsApp, and more.

A friend who runs a customer-service system sent me a screenshot yesterday. His monthly API bill dropped nearly 40% compared with the month before. I assumed his business had shrunk. It hadn't. DeepSeek had simply cut the cache pricing on its flash series. On September 9, DeepSeek announced on its open platform that starting at noon on September 10, during off-peak hours the price of cached input hits would drop from 0.05 yuan to 0.02 yuan per million tokens — a 60% cut. Uncached input dropped from 1.5 yuan to 1 yuan (down 33%), and output from 4.5 yuan to 4 yuan (down 11%). On the surface this looks like another vendor price war. But it carries a very practical message for ordinary users: the cheaper and more frequently models switch, the more you need a local scheduling butler.

First, where exactly does that 60% get saved

A cache hit, in plain terms, means you ask the AI the same question repeatedly, or feed it large amounts of repeated content, and that portion of the cost drops sharply. Scenarios with high hit rates — a support bot answering FAQs, document summarization reusing fixed prompts, system instructions repeated across a long conversation — see the steepest bill reductions. One developer estimated overall API costs could fall around 40%.

But here is the catch that needs saying out loud: how much you save depends entirely on your cache-hit rate and the input/output token mix. A workflow with a low hit rate barely benefits. And this cut is really a partial rollback of the August price hike — output at 4 yuan is still double the pre-hike 2 yuan. So don't let the word "discount" spin your head. It only pulled back part of what was raised.

配图

The second catch: peak hours cost double

DeepSeek keeps its peak/off-peak split. Weekdays 9–12 and 14–18 are peak, priced at double. Everything else is off-peak, and weekends are entirely off-peak. The intent is obvious: nudge you to push non-urgent work to evenings or weekends.

That is fine for developers, but a headache for ordinary people. The moment you actually want AI to help during the day is exactly when it costs the most. If you have to remember "what's cheap when," using AI becomes more work than not using it.

Model prices are becoming a moving target

Step back and look at the wider picture. In a single month DeepSeek raised prices, then cut them. GPT and Claude have done similar things. Model pricing behaves like vegetable prices at a market — cheap today, expensive tomorrow. And there is a worse trend: models refresh faster and faster. A new flash arrives, stronger and quicker, but it also burns tokens more aggressively.

This surfaces a question an ordinary person simply cannot compute: for the same task, which model should you dispatch right now? The cheapest flash, or the most reliable Pro? Run it in the cheap off-peak window, or get the result now?

KAIHE AIBOX: keep the scheduling at home

This is exactly the job of a "scheduling butler" that KAIHE AIBOX does. It does not run large-model inference itself. It acts as a local Agent hub, treating every cloud model as a tool it can call on demand. You hand it a task; it picks the most suitable and most cost-effective model based on task type, current prices, and busy/idle windows. A model raises its price — it switches. A new model appears — it plugs in. Off-peak is cheap — it queues non-urgent work for the night.

You don't stare at a price table doing arithmetic. The scheduling logic runs inside that little box in your home.

Why "local" is the step you can't skip

People ask: can't a cloud scheduling tool do the same? Not quite. Two things set it apart.

First, data. Your conversation memory, business materials, and accumulated workflows live locally on KAIHE AIBOX, never uploaded to the cloud. Models are only borrowed for compute; your assets stay with you. Second, always-on. KAIHE AIBOX draws very little power. Even with your computer off and phone locked, it keeps running in the background — scheduled tasks, nightly batches, remote commands, none of it depends on you sitting at a desk.

A cloud scheduler can help you pick a model, but it cannot give you "data never leaves home" and "runs even when powered down." That is why the cheaper and more frequently models switch, the more valuable this local base becomes — it wraps both the freedom of "which model to use" and the bottom line of "who owns the data" in one package.

配图

A concrete picture

Picture this: before bed you drop a task on KAIHE AIBOX — compile the week's customer inquiries into a report. It decides the task isn't urgent, queues it for the cheap early-morning off-peak window, runs it on whatever flash model is cheapest at that moment, and the report is already in your WeChat when you wake up. Which model it switched to, how much it cost — you don't care, and your memory and data never left that box in your home.

Model prices change daily. Your work doesn't get any lighter. What KAIHE AIBOX does is make those fluctuations irrelevant to you.

Further Reading

Want a scheduling butler that is always online at home? Learn about KAIHE AIBOX-A1, or visit the official store for the full lineup.

Learn More about KAIHE AIBOX, search 【KAIHE AIBOX】.

Email: [email protected]

Recommended Products

A1 Home Entry A1 Pro Enhanced A2 Professional A2 Pro Advanced X1 Enterprise G1 Flagship
© KAIHE AI - Agent Computer Specialist