DeepSeek V4.1 Flash is available on Novita AI as a serverless API. It’s the best fit for developers who want stronger agentic coding results than the previous DeepSeek V4 Flash at the same 1M-token context, plus native image input, and are willing to pay a higher per-token rate for it. Pricing on Novita AI is $0.3 per 1M input tokens and $1.2 per 1M output tokens, and it’s reachable through the existing OpenAI-compatible chat completions endpoint.
Key Takeaways
- DeepSeek V4.1 Flash is live on Novita AI, serverless, OpenAI-compatible access (see the specs table for the exact model ID).
- It adds multimodal input. Unlike the text-only DeepSeek V4 Flash, V4.1 Flash accepts text and image input and returns text output.
- Agentic coding scores jump over V4 Flash. DeepSeek’s published results show Terminal-Bench 2.1 at 90.6 (vs. 82.7 for V4 Flash) and DeepSWE v1.1 at 74.2 (vs. 54.4).
- Context stays at 1M tokens, with a smaller KV cache and fewer active parameters per token. The architecture activates only 8B parameters during prefill and 16B during decode, versus 13B flat for V4 Flash — a DeepSeek-side efficiency gain, not a Novita price cut.
- Pricing is higher than base V4 Flash but still a mid-tier option. Novita lists $0.3/$1.2 per 1M input/output tokens for V4.1 Flash, between V4 Flash’s $0.14/$0.28 and the V4 Flash 0731 revision’s $0.44/$1.32.
What Is DeepSeek V4.1 Flash?
DeepSeek V4.1 Flash is DeepSeek’s successor to DeepSeek V4 Flash, released September 10, 2026 under the MIT license. It’s a multimodal Mixture-of-Experts model with 552B backbone parameters plus 196B parameters in an “Engram” conditional-memory module, for roughly 748B total parameters. The model natively processes text and images and generates text autoregressively.
The headline change is architectural, not just a parameter bump. V4.1 Flash uses a Causal Encoder-Decoder design: a 40-layer transformer split into a 20-layer causal encoder and a 20-layer decoder, where the decoder’s KV cache is projected from the encoder’s final hidden states instead of being computed fresh at every decoder layer. Combined with FP4 KV caching and a technique DeepSeek calls Compressed Sparse Attention 2, this cuts the global KV cache footprint to roughly a quarter of DeepSeek V4 Flash’s, and keeps active parameters per token to 8B during prefill and 16B during decode. That’s a DeepSeek-side architectural efficiency gain in how the model is served, not a statement about what you pay on Novita AI; the per-token price you see is set by Novita’s pricing, not by the model’s internal compute footprint.
DeepSeek also supports a continuously controllable reasoning effort setting (an integer from 1 to 100) that trades inference cost for accuracy on a single model ID, rather than requiring separate “fast” and “thinking” model variants.
Explore DeepSeek V4.1 Flash on Novita AI
DeepSeek V4.1 Flash Specs and Pricing Summary
| Field | Details |
|---|---|
| Display name | DeepSeek V4.1 Flash |
| Model ID | deepseek/deepseek-v4.1-flash |
| Base URL | https://api.novita.ai/v3/openai |
| Endpoint family | Chat completions, Anthropic-compatible, Responses |
| Input modalities | Text, image |
| Output modality | Text |
| Context window | 1,048,576 tokens |
| Max output tokens | 393,216 tokens |
| Input price | $0.3 per 1M tokens |
| Cache-read price | $0.006 per 1M tokens |
| Output price | $1.2 per 1M tokens |
| Best fit | High-throughput agentic coding and tool-use workloads that also need occasional image input |
Specs and pricing from the Novita AI model page, checked October 10, 2026.
DeepSeek V4.1 Flash API Access on Novita AI
Novita AI hosts DeepSeek V4.1 Flash as a serverless model with the model ID from the specs table above. It’s available through the OpenAI-compatible chat completions endpoint, as well as Anthropic-compatible and Responses endpoint families, so existing integrations built against other Novita-hosted models can point at the new model ID with minimal changes.
For request fields, streaming behavior, image-content formatting, and tool-call syntax, use the current Novita AI Create Chat Completion reference rather than copying request shapes from an older DeepSeek post.
Key Capabilities for Developers
Stronger Agentic Coding Than V4 Flash
DeepSeek’s published benchmarks at maximum reasoning effort show meaningful gains over V4 Flash on agent-style tasks:
| Signal | What it shows | Developer takeaway |
|---|---|---|
| Terminal-Bench 2.1 (Pass@1) | 90.6 for V4.1 Flash vs. 82.7 for V4 Flash | Better at multi-step shell/agent tasks in a sandboxed terminal |
| DeepSWE v1.1 (Resolved) | 74.2 for V4.1 Flash vs. 54.4 for V4 Flash | Higher real-world issue-resolution rate on repo-scale coding tasks |
| AutomationBench (Pass@1) | 54.8 for V4.1 Flash vs. 37.7 for V4 Flash | Stronger at multi-step browser/tool automation |
Data from DeepSeek’s DeepSeek-V4.1-Flash model card, checked October 10, 2026. These are DeepSeek’s own evaluation numbers; validate against your own agent harness and task set before treating them as a production guarantee.
The practical read: if your current V4 Flash integration drives a coding agent, a terminal-automation workflow, or a multi-step tool-use pipeline, V4.1 Flash is worth re-testing on the same prompts before you decide whether the pricing step-up pays for itself.
Multimodal Input Without a Separate Vision Model
V4.1 Flash accepts image input alongside text, so you can send a screenshot, chart, or document page together with an instruction instead of running a separate OCR or vision step first. That’s a meaningful change from the original DeepSeek V4 Flash, which is text-only; previously, image-aware DeepSeek workloads needed the (now-superseded) experimental Vision variant.
Smaller KV Cache and Active-Parameter Footprint at 1M Context
The 8B/16B active-parameter split (prefill/decode) is lower than V4 Flash’s flat 13B, and DeepSeek reports the KV cache footprint at roughly a quarter of V4 Flash’s at the same 1M-token context. That reduces the compute and memory pressure DeepSeek’s own serving infrastructure carries at long context — it is not a claim about your bill. On Novita AI, V4.1 Flash’s listed per-token price is higher than base V4 Flash’s, so your actual cost for a given workload still depends on token volume and Novita’s current pricing, not on the model’s internal efficiency gains. Test your own representative context length and request volume before assuming either model is cheaper for your use case.
When to Use DeepSeek V4.1 Flash
- Coding agents and terminal-automation workflows where Terminal-Bench- and DeepSWE-style tasks are representative of your workload
- Multi-step tool-use and browser-automation agents that benefit from the AutomationBench gains
- Workflows that need to mix a screenshot, chart, or document image with a text instruction in the same request
- Long-context agent sessions (large codebases, long tool-call histories) where you want to test whether the smaller KV cache architecture translates into better throughput or latency for your traffic
- Teams currently on DeepSeek V4 Flash who want to re-evaluate quality on agentic tasks before committing to a model swap
When Not to Use DeepSeek V4.1 Flash
Don’t default to V4.1 Flash for simple, high-volume, latency-sensitive text requests where DeepSeek’s own base-model benchmarks show it roughly matching or trailing V4 Flash (for example, MGSM and MATH). If your workload is short-prompt classification or chat completion without agentic or multimodal demands, the lower per-token price of base DeepSeek V4 Flash may be the better default — test both on your own prompts rather than assuming the newer model always wins.
Also hold off if you need a model with a longer production track record; V4.1 Flash is a recent release, so validate your exact tool-call schema, image-input path, and failure modes before routing production traffic.
How DeepSeek V4.1 Flash Fits Your API Workflow
An existing Novita AI integration needs one change to try V4.1 Flash: swap in the DeepSeek V4.1 Flash model ID (specs table above) on your existing base URL and API key. If your workload sends images, add multimodal message content per the current API reference; if it doesn’t, your text-only request body should work unchanged.
Start with a smoke test on a handful of representative prompts, especially any agentic or tool-use flows, before widening rollout. Keep the model ID, context limit, and pricing in configuration rather than hardcoded, and recheck the Novita AI model page before any cost-sensitive production decision — DeepSeek Flash pricing and model status have changed more than once this year.
Conclusion
DeepSeek V4.1 Flash on Novita AI is the model to evaluate next if your current DeepSeek V4 Flash workload touches coding agents, tool use, or terminal automation, or if you need occasional image input without standing up a separate vision model. The 1M-token context carries over from V4 Flash, agentic benchmark scores move up meaningfully, and the architecture carries a smaller KV cache and fewer active parameters per token — though on Novita AI, V4.1 Flash’s listed per-token price is still higher than base V4 Flash’s.
Test it against your own agent and coding tasks, compare the real cost at your typical context length rather than only the headline price, and keep the model ID and limits in configuration so you can roll back to V4 Flash if a specific workload doesn’t benefit.
FAQ
What is the model ID for DeepSeek V4.1 Flash on Novita AI?
See the specs table above for the exact model ID to use in Novita AI API requests.
How is DeepSeek V4.1 Flash different from DeepSeek V4 Flash?
V4.1 Flash adds image input, uses a new Causal Encoder-Decoder architecture with a smaller KV cache and fewer active parameters per token, and scores meaningfully higher on agentic coding and tool-use benchmarks. Context stays at 1,048,576 tokens on both models, but Novita’s listed price for V4.1 Flash is higher than base V4 Flash.
How much does DeepSeek V4.1 Flash cost on Novita AI?
As checked October 10, 2026, Novita AI lists $0.3 per 1M input tokens, $0.006 per 1M cache-read tokens, and $1.2 per 1M output tokens. Confirm current pricing on the model page before production budgeting.
Does DeepSeek V4.1 Flash support image input?
Yes. Novita AI lists text and image input modalities for DeepSeek V4.1 Flash, with text output.
What is the context window for DeepSeek V4.1 Flash?
Novita AI lists a 1,048,576-token (1M) context window and a 393,216-token maximum output.
Is DeepSeek V4.1 Flash a good fit for simple chat or classification tasks?
Not necessarily as the default. DeepSeek’s own benchmarks show V4.1 Flash roughly matching or trailing V4 Flash on some non-agentic tasks. Test both models on your actual prompts before choosing, especially for high-volume, latency-sensitive, text-only workloads.
Does DeepSeek V4.1 Flash support function calling?
Yes. Novita AI lists function calling, structured outputs, and reasoning among the hosted features for DeepSeek V4.1 Flash.