Deepseek V4 Flash Vision Kaihe Aibox

Published on: 2026-08-24

On August 21, DeepSeek quietly listed a new model on its API platform: deepseek-v4-flash-vision-exp. The long name means one thing — this is DeepSeek's first model that can actually "look at images." Where V4-Flash was text-only, this one opens its eyes: screenshots, charts, and UI layouts can now be fed in and understood directly.

For users of a local AI box like KAIHE AIBOX, this update is worth watching. Not because the model runs in your home, but because it pushes the door of "agents that can see the world" open a little wider — and who can keep that door local and running 24/7 is exactly the question we care about.

What It Actually Gained

📖 Glossary

AI Box (also known as Agent Computer / Agent PC), is a dedicated local hardware device that runs AI Agents. Pre-installed with an AI agent management system, plug-and-play, running 24/7. Users can remotely command AI to work via Discord, Slack, Telegram, WhatsApp, and more.

First, the positioning: V4-Flash-Vision-Exp is an experimental build (the "Exp" stands for Experimental). It adds a vision layer on top of the text-only V4-Flash. DeepSeek states that its text capabilities — agents, reasoning, world knowledge — remain exactly on par with the original V4-Flash, with nothing added or lost, just a new pair of eyes.

What can those eyes do? It supports JPEG, PNG, GIF, and WebP, up to 600 images per request, with a max edge of 8192 pixels. Capabilities include image recognition, chart analysis, and text extraction from screenshots. In the official demo, it identifies a British Shorthair cat in seconds and even lists the coat-color distribution and paw-pad color as evidence.

More importantly, there is a leap in multimodal agent ability. DeepSeek claims that on vision-dependent agent benchmarks it far surpasses the text-only V4-Flash, with multimodal agent performance now approaching Claude Opus 4.8. In plain terms: agents doing computer-use tasks used to be "blind"; now they can read screenshots and see interfaces, and GUI automation and screen-operation scenarios suddenly click into place.

配图

Vision Added, but No Price Premium — the Sharpest Move

Multimodal models usually bill separately, and image tokens often cost several times more than the same length of text. DeepSeek folded vision straight into the existing billing system — identical pricing to the text-only V4-Flash.

How it is counted: an image is converted to tokens, capped at 384 tokens per image, billed at the input-token rate. Pricing follows the V4-Flash tier — cache-hit input at 0.05 yuan per million tokens off-peak and 0.10 yuan peak; cache-miss input at 1.5 yuan (off-peak) / 3 yuan (peak) per million, and output at 4.5 yuan (off-peak) / 9 yuan (peak) per million tokens.

Some outlets ran the numbers: for the same-size image, leading models burn 258 to 1334 tokens per image, pushing the cost of a thousand images as high as ~29 yuan; DeepSeek caps at 384 tokens per image, so at the base 3-yuan-per-million input rate, a thousand images cost only about 1.15 yuan — under 0.6 yuan at off-peak half price. The gap can reach 50x.

配图

Also launched: a free Files API. Upload an image once, get a file_id, and reference it in later requests without re-uploading. For visual agents doing batch screenshot analysis, long-document checks, or chart comprehension, this is a thoughtful touch.

What It Means for Developers

The most comfortable part is compatibility. It supports OpenAI's Chat Completions, Anthropic's Messages, and the Responses API — existing agents and tools need no interface rewrite; just change the model name to plug vision in. The same-day DeepSeek Harness 0.1.1 also supports the new model out of the box.

It keeps V4-Flash's 1M context window and supports thinking modes (high/max); complex multimodal tasks are recommended at max. The mobile app also shipped an "image recognition mode" so ordinary users can try image interaction directly.

But set the boundaries straight: this is an API-only experimental build, with no open weights, and you cannot download it to run on your own machine. All image understanding happens on DeepSeek's cloud. The benchmarks are DeepSeek's own runs — across 11 comparisons against Opus 4.8, it leads on 3 and trails on 8, with Terminal Bench 2.1 at 83.9 vs 85.0. So "approaching Opus 4.8" should be heard with a discount.

The Other Layer KAIHE AIBOX Sees

Now that models can see, what does a local AI box get out of it? Simply put, multimodal agent playbooks multiply: screenshot-to-code, UI recognition then auto-form-filling, chart reading for data rollups, long-document figure checks — work that used to need a human staring at a screen can now flow through an agent pipeline.

But the model is in the cloud; the eyes live on DeepSeek's servers. If you ask an agent to analyze a screenshot containing client info, that image has to travel to DeepSeek's cloud first. That is the exact opposite of KAIHE AIBOX's "data never leaves home" principle.

So the real division of labor is this: send the visible, temporary, collaborative visual tasks to cloud multimodal models; keep the sensitive, long-running, always-on tasks on the local base. KAIHE AIBOX's value is not competing with DeepSeek on who sees sharper, but being the ever-plugged-in, 24/7 scheduling hub — while you sleep, it still runs orchestrated flows that call multiple model APIs including V4-Flash-Vision, chaining "seeing" and "doing" end to end.

To learn how a local AI box schedules multiple models while keeping your data in your own hands, see Why an AI Box Beats Installing on a PC, or go straight to the product page and the store.

The age of multimodal agents is just lifting the curtain. DeepSeek's move is a reminder to every local base: models keep getting more capable, but whoever schedules them well and guards data sovereignty is the one still at the table. The giants built the "model that can see" — what we should ask is, on whose body those eyes should sit.


Further Reading

DeepSeek #V4FlashVision #Multimodal #AIFrontier #KAIHEAIBOX

Learn More: Search 【KAIHE AIBOX】

Contact: [email protected]

—— KAIHE Intelligence · Your 24/7 Personal AI Assistant | AI Frontier

Recommended Products

A1 Home Entry A1 Pro Enhanced A2 Professional A2 Pro Advanced X1 Enterprise G1 Flagship
© KAIHE AI - Agent Computer Specialist