AI Mindset · Model Cheatsheets
DeepSeek

The High-Efficiency Technical Specialist

DeepSeek has promised a model it has not shipped. Its September 10 release article says every deepseek-v4-pro request routes to V4.1-Flash from September 14, and that this continues until V4.1-Pro launches. V4.1-Pro has not launched. Pro is still priced, still served, still the model a claude-opus string maps to. So the phase-out that reads like an ending is a bridge with no far end. Meanwhile the lineup is two models, deepseek-flash and deepseek-v4-pro, prices were cut on September 10, image input is native, and the weights are MIT. And the price moved twice in under a month, once up and once down. Any number off an old screenshot is wrong in one direction or the other.

Verified October 1, 2026V4.1-Flash · V4-Pro · V4.1-Pro unshipped1M context · 384K outputPrices cut September 10, 2026MIT weights · dual API formats
1 / Meet DeepSeek

Strong capability per dollar, with a different operating model

Judge DeepSeek as infrastructure, not as a cheaper consumer chatbot. Its value rides on task routing, deployment choice and model evaluation. And on one human question: can your organization manage the legal and operational context that comes attached? Three working defaults follow. Re-price the workload against the September 10, 2026 rates, not any earlier reading. Default to deepseek-flash, because the vendor no longer claims Pro performs better. And remember that peak applies Monday through Friday, minus Chinese public holidays, which DeepSeek does not list.

Personality

Technical, efficient and configurable

It’s built for teams that are comfortable choosing modes, APIs, weights and deployment architecture. Not for teams that want one polished end-user product.

  • Thinking and non-thinking modes
  • Agent and coding orientation
  • MIT-licensed weights and a hosted API
Deploy it for

High-volume technical work

Coding, tool use, document and image processing, long-context analysis, and any workload where token economics change the business case.

  • Route almost everything to deepseek-flash
  • Raise thinking effort before changing model
  • Exploit caching, concurrency and the off-peak clock
Choose DeepSeek when

Control and economics justify ownership

Use it when your team can evaluate quality, design safeguards and run the hosted or self-hosted path responsibly. And when processing in the People’s Republic of China under PRC law is acceptable for the data involved.

  • Benchmark on your own tasks
  • Resolve data and jurisdiction questions first
  • Budget for model operations

Is DeepSeek the right strategic choice?

Choose the requirement that is driving the decision.

API economics
Re-price against the September 10 rates, in both directions

deepseek-flash is the cheapest way into the lineup. It’s also got the highest listed concurrency at 2,500, and it is now the only model with image input. Those rates took effect at 04:00 UTC on September 10, 2026 and sit below DeepSeek’s August rates across the board. So anyone still costing against the pre-August flat rates is under-budgeting. Anyone costing against the August peak and off-peak rates is over-budgeting. Both are wrong, in opposite directions. The numbers to budget against: output at $0.60 per million tokens off-peak and $1.20 at peak, cache-miss input at $0.15 and $0.30, cache hit at $0.003 and $0.006. One catch on the off-peak half. Peak now excludes Chinese public holidays, and DeepSeek publishes no list of which days those are. So your off-peak share is a range, not a figure.

  • Model quality on your data, not the headline rate
  • Cache-hit and off-peak scheduling assumptions
  • Peak is weekdays only, and Chinese public holidays are off-peak in full
  • Total workflow cost, not token price alone
2 / What’s Current

A new default model, a price cut, and two retirements

Everything about the lineup changed on September 10, 2026. DeepSeek shipped DeepSeek-V4.1-Flash, retired two model IDs, and reduced prices with effect from 04:00 UTC that day. The two retired names are still accepted. But they are, in DeepSeek’s own wording, temporarily routed to V4.1-Flash and billed at the Flash price. So anything that pins them is running on a compatibility shim with no stated end date. Image input is now native to the default model rather than a separate experimental build. Pro was supposed to stop taking traffic on September 14. It did not. Seventeen days later it is still priced, still rate-limited, still monitored and still the model a claude-opus string buys. The release article conditions the routing on a successor, V4.1-Pro, which has not shipped. Both current models carry 1M context and 384K maximum output, and neither of the July-retired aliases appears on the pricing page.

Half price
Off-peak rate. Peak is 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday. Chinese public holidays are off-peak
$0.60 / $1.20
deepseek-flash output per million tokens, off-peak and peak
$1.98 / $3.96
deepseek-v4-pro output per million tokens, off-peak and peak
04:00 UTC
The hour the current price list took effect, September 10, 2026

DeepSeek-V4.1-Flash New · September 10, 2026

The new default, and the only model in the lineup that accepts images. The architecture lines below are vendor claims, not measurements.
· New: deepseek-flash. Retired: deepseek-v4-flash and deepseek-v4-flash-vision-exp. Continuing: deepseek-v4-pro. Promised and unshipped: V4.1-Pro.
· DeepSeek describes a 552B-parameter mixture-of-experts model on a new causal encoder and decoder architecture. It splits a 40-layer Transformer into a 20-layer causal encoder and a 20-layer decoder, with just 8B active parameters for input and 16B for output.
· It claims the KV cache needs a quarter of the HBM and an eighth of the SSD storage of the previous generation, and that multimodality is native through DeepSeek-ViT, trained from scratch.
· The weights channel lists a 552B backbone and 763B total parameters under the MIT license.

  • deepseek-flash
  • 1M context, 384K maximum output
  • Image input supported, 2,500 listed concurrency
  • MIT-licensed weights

DeepSeek V4-Pro Still serving, past its stated date

The 0813 build reached general availability on August 13, 2026. It is still fully served on October 1, which you can check in four places at once rather than taking one page’s word for it. What it no longer is, on DeepSeek’s own account, is the stronger model. The September 10 article says several parties measured V4.1-Flash ahead of it on performance, cost, speed and total runtime. Pro costs 3.3 times Flash on output, 4.4 times on cache-miss input and 7.3 times on cache-hit input. It carries 500 listed concurrency against Flash’s 2,500, and the pricing page marks image input as not supported for it. That missing vision is now its only distinguishing feature.
· Checked October 1, 2026: a full price row on the pricing page, 500 concurrency on the Rate Limit page, a named model on the docs homepage, and 99.92% uptime on the status page.

  • deepseek-v4-pro
  • No image input, 500 listed concurrency
  • 3.3x Flash on output, 4.4x on cache-miss input, 7.3x on cache-hit input
  • Read the September 14 card before you depend on it

September 14 came and went Not executed

DeepSeek set a hard date for moving Pro traffic away, and then nothing happened on it. Seventeen days past that date, every Pro request still reaches Pro and bills at the Pro rates. The one sentence anywhere that says service continues with billing unchanged sits on the change log, a page DeepSeek rewrites in place. No date on it, no history behind it, and it is not repeated on the pricing page. So it is a reading, not a commitment.
· September 10 release article: starting at 04:00 UTC on September 14, 2026 all deepseek-v4-pro requests will route to V4.1-Flash at V4.1-Flash rates, and this will continue until V4.1-Pro launches. The same article says Pro is being phased out.
· Change log, checked October 1, 2026: in response to user demand DeepSeek has decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged.

  • The routing did not happen on the stated date
  • The continuation sentence lives only on a rewritable page
  • The release article has not been withdrawn or annotated
  • A Pro dependency needs written confirmation, not a page reading

V4.1-Pro is named and unshipped Watch this

DeepSeek tied Pro’s ending to a model it has not released. The routing runs until V4.1-Pro launches, in the article’s own words. So the phase-out that reads like a shutdown is a hand-off waiting on a product, and the product is not there. No price row. No docs entry. No change-log entry. No published weights. The planning consequence is the awkward one. You cannot move onto the successor, and you cannot put a date on the end of the model it replaces.
· Checked October 1, 2026: no V4.1-Pro on the pricing page, the docs homepage, the change log whose newest entry is September 10, or DeepSeek’s weights channel.

  • Announced in the September 10 release article
  • No pricing, docs, change-log or weights entry
  • Pro’s routing is conditioned on it launching
  • Plan a migration you cannot yet schedule

Retired: V4-Flash and Vision-Exp Executed

Two model IDs went out on September 10, 2026. DeepSeek states that the corresponding models have been retired, and that their requests are served by the DeepSeek-V4.1-Flash model and billed at the Flash price. The change log says the two names are temporarily routed for compatibility. Temporarily is DeepSeek’s word, and it comes with no end date. The July retirement still stands separately.
· Retired September 10, 2026: deepseek-v4-flash and deepseek-v4-flash-vision-exp.
· Dead since 2026/07/24 15:59 UTC: deepseek-chat and deepseek-reasoner. Neither appears on the pricing page.

  • Pin deepseek-flash or deepseek-v4-pro
  • A temporary route is not a supported model ID
  • Legacy deepseek-chat and deepseek-reasoner calls fail outright
  • Update tests, configuration and cost attribution

Vision is native, and the limits are hard

The old routing decision has disappeared. You no longer send images to a separate experimental build. The token cap gets quoted everywhere, and it is the least likely thing to break you. The size and count ceilings are what stop a document pipeline cold, usually on the page nobody tested. Count your pages and weigh your payload before you design the batch. One cost question is still open: no DeepSeek page states a price for the Files API, so treat it as unconfirmed. Not zero.
· Vision guide, checked October 1, 2026: 1024 tokens per image after automatic resizing. JPEG, PNG, GIF and WebP. 32 MiB per image by base64 or URL, 64 MiB through the Files API. 600 images per request. 8,192 characters maximum for a URL. 48 MiB for the whole request body.
· Pricing page, checked October 1, 2026: image input supported for deepseek-flash, not supported for deepseek-v4-pro.

  • 1024 tokens per image, 600 images per request
  • 32 MiB per image, 64 MiB through the Files API
  • 48 MiB total request body, so batch size is bounded twice
  • Files API pricing is not published anywhere

Dual API formats

Two request formats, one service. Auto-mapping is convenient, and it also means a Claude model string sitting in your configuration quietly decides which DeepSeek model you pay for.
· OpenAI-compatible endpoint at https://api.deepseek.com. Anthropic-compatible endpoint at https://api.deepseek.com/anthropic.
· Anthropic guide, checked October 1, 2026: model names starting with claude-haiku or claude-sonnet are mapped to deepseek-flash. claude-opus maps to deepseek-v4-pro and is billed at the V4 Pro price.

  • Lower migration friction
  • Explicit base URLs
  • claude-haiku and claude-sonnet land on deepseek-flash
  • claude-opus lands on Pro and bills at the Pro price

Thinking control, and the trap in it

Here’s the part that breaks code copied between formats. Three request formats, three different spellings, and an effort value that turns thinking off exists in only one of them. Write the setting per format, and set it explicitly. The default effort level is unconfirmed, so don’t lean on it. The pricing page lists no per-effort rate, so effort moves your token count, not your rate.
· Thinking Mode guide, checked October 1, 2026. Anthropic-compatible format: reasoning.effort takes none, low, high or max, and none disables thinking.
· OpenAI-compatible format, which is the default base URL: reasoning_effort takes low, high or max, and you disable thinking with a separate thinking object set to disabled.
· Responses API: output_config.effort takes low, high or max. There is no none on that surface at all.

  • Default effort level unconfirmed, so set it explicitly
  • none exists only in the Anthropic-compatible format
  • The OpenAI format needs a separate thinking toggle
  • The Responses API has no off value for effort

Responses API coverage

Narrower than DeepSeek’s own release material implies. Standardized on the Responses API and also depend on Pro? That’s a gap to confirm with DeepSeek, not one to assume away. The August 13 release article separately claims native OpenAI Responses API support optimized for Codex with one-click setup. That’s a vendor claim to test. It is also the one surface with no way to turn thinking off, so a high-volume extraction job standardized here pays for reasoning it never asked for.
· Responses API guide, checked October 1, 2026: deepseek-flash is named as supported, deepseek-v4-pro does not appear anywhere on that page, and output_config.effort takes low, high or max only.

  • Only deepseek-flash is named as supported
  • Pro is absent from that page, which is not the same as refused
  • No effort value turns thinking off on this surface
  • Codex setup is a vendor claim, so test it

Peak and off-peak pricing Holiday carve-out added

Thirty-five hours a week is the ceiling now, not the number. Peak used to be every weekday hour inside the two windows. It now excludes Chinese public holidays, and DeepSeek neither lists those days nor links a calendar for them. So you can bound your peak exposure from above, and you cannot compute it from first-party material. A batch job moved to a Saturday is off-peak for the full twenty-four hours, and so is one that lands on a holiday nobody on your team had marked. Both price regimes do carry a published effective hour, which is the only reason either can be dated: the pricing page itself has no last-updated stamp.
· Pricing page, checked October 1, 2026: off-peak rates are half of the peak rates. Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday, excluding Chinese public holidays. All other hours are off-peak, including weekends and Chinese public holidays in full.
· August 13 release article: the peak and off-peak regime took effect at 16:00 UTC on August 16, 2026. September 10 release article: the current rates took effect at 04:00 UTC on September 10, 2026, which is 24 days and 12 hours later.

  • Thirty-five hours a week is now an upper bound
  • Weekends are off-peak end to end
  • Chinese public holidays are off-peak and are not listed
  • Both regimes have a published effective hour

Open weights, MIT licensed Released

DeepSeek publishes its V4 weight collections and technical material on its own weights channel under the MIT license. That permits commercial use, modification and redistribution with attribution, and it is much more permissive than the bespoke community licenses common elsewhere. Here’s one oddity worth knowing. The retired DeepSeek-V4-Flash-Vision-Exp repository is still published on October 1, 2026, 21 days after the API retired the model on September 10. And it shows activity since the retirement, with nothing visible about what changed. So the weights channel and the API don’t move together. The channel is also wider than the API: six V4 repositories against two live model IDs.
· Weights channel, read October 1, 2026: V4.1-Flash at 763B total, V4-Pro-0813 at 1.7T, V4-Flash-Vision-Exp at 305B, V4-Flash-0731 at 304B, V4-Flash-DSpark at 165B and V4-Pro-DSpark at 1.7T.
· MIT confirmed on the four repositories opened that day: V4.1-Flash, V4-Pro-0813, V4-Flash-Vision-Exp and V4-Flash-0731. The two DSpark repositories were not opened, so their license is unconfirmed.

  • MIT confirmed on four of the six V4 repositories
  • Hosting control, model operations responsibility
  • Hardware and quantization planning
  • A retired API model can still be self-hosted

Service status

DeepSeek publishes a status page with per-service uptime. Those are the vendor’s own figures, on a live page that is rewritten continuously. So they’re a snapshot, not a commitment. There is no published SLA attached to them. And the service list itself moves: Chat currently appears twice, with two different numbers, which tells you the page is a dashboard rather than a register.
· Checked October 1, 2026: all systems operating as expected, no incidents listed. Published uptimes of 99.92% for the V4 Pro API, 99.69% for the V4.1 Flash API, 99.84% and 99.68% for two separate Chat entries, 100.00% for File Upload and 99.91% for Search.

  • Green across all services on October 1, 2026
  • Vendor-published uptime, not a contractual SLA
  • The service list changes shape, so do not diff it by hand
  • Design fallback for a provider with no SLA
Model and rate windowInput: cache hitInput: cache missOutputConcurrency
deepseek-flash (DeepSeek-V4.1-Flash) · off-peak$0.003 / 1M$0.15 / 1M$0.60 / 1M2,500
deepseek-flash (DeepSeek-V4.1-Flash) · peak$0.006 / 1M$0.30 / 1M$1.20 / 1M2,500
deepseek-v4-pro (DeepSeek-V4-Pro-0813) · off-peak$0.022 / 1M$0.66 / 1M$1.98 / 1M500
deepseek-v4-pro (DeepSeek-V4-Pro-0813) · peak$0.044 / 1M$1.32 / 1M$3.96 / 1M500
deepseek-v4-flash · retired, legacy aliasBilled at the Flash rate aboveBilled at the Flash rate aboveBilled at the Flash rate aboveServed by V4.1-Flash
deepseek-v4-flash-vision-exp · retired, legacy aliasBilled at the Flash rate aboveBilled at the Flash rate aboveBilled at the Flash rate aboveServed by V4.1-Flash
Model IDBuild namedContextMax outputImage inputThinkingConcurrency
deepseek-flashDeepSeek-V4.1-Flash1M384KSupportedSupported2,500
deepseek-v4-proDeepSeek-V4-Pro-08131M384KNot supportedSupported500
3 / Model Router

Use the least expensive mode that meets the task’s failure cost

The routing map got simpler on September 10. Images no longer route anywhere special, because the default model has vision. Hard work no longer routes to Pro, because the vendor no longer claims Pro performs better. What is left is one model, an effort dial and a clock. Run deepseek-flash at the right effort, scheduled against the weekday peak windows. Keep Pro for the narrow cases that really call for it. And check which request format you are on before you write the effort value, because the off switch does not exist on all three.

Route the workload

Select the dominant task property.

Everyday agent task
deepseek-flash · effort high

The right economic default. Flash with thinking at high covers planning, tool calling and coding where the model needs a deliberate plan. Its output costs under a third of Pro’s, and it has five times the listed concurrency. Set high explicitly, because the vendor default level is unconfirmed.

  • Strong economic default
  • Set the effort value explicitly
  • Set maximum steps
  • Image input works here too, capped at 1024 tokens per image

Reasoning effort versus operational cost

Three active thinking levels, low, high and max, plus an off state that only two of the three request formats have. The default level is unconfirmed, so set it explicitly. The Anthropic-compatible format sets effort through reasoning.effort and accepts none. The OpenAI-compatible format uses reasoning_effort and turns thinking off through a separate thinking object. The Responses API uses output_config.effort and has no off value. Raise effort only when the task actually benefits from deeper search or tool use.

ThroughputMaximum reasoning
Effort high
Everyday complex work

High effort fits daily agent workflows: analysis, tool calling and coding where the model needs a deliberate plan. Set it explicitly. The vendor default level is unconfirmed.

  • The working default for agents
  • Good balance of cost and depth
  • Set completion criteria
  • Check whether your workload actually needs it
Routing evaluation
Run this task suite on deepseek-flash with thinking disabled, then at effort low, high and max, and on deepseek-v4-pro at high. State which request format you used and how you disabled thinking in it. Compare exact-task success, latency, output tokens, retries and human correction time. Recommend a routing policy and state the cost per successful task at both peak and off-peak rates.
Agent boundary
Complete this repository task with a maximum of 12 tool steps. Stop before any external network or deployment action. Report tests, files changed and unresolved risk.
4 / Deployment

Hosted API and open weights solve different problems

Open weights improve control. They do not remove cost. They swap per-token pricing for infrastructure, model serving, security, monitoring and lifecycle work. What DeepSeek does remove is the licensing obstacle. Every V4 weight repository opened on October 1, 2026 was MIT.

Choose the deployment path

Start with the organizational requirement, not the ideology.

Fastest path to production
Hosted DeepSeek API

Use the official API when speed of integration, provider-operated serving and listed token economics are the priority. Accept that content is processed and stored in the People’s Republic of China. And accept that the same Open Platform terms cover individual and enterprise developers.

  • OpenAI or Anthropic request format
  • Provider data and jurisdiction review
  • Monitor rate and pricing changes, which moved twice in a month
DimensionHosted APISelf-managed weights
Time to startFastInfrastructure project
Variable costPublished token pricingCompute, staffing and utilization
Data controlProcessed and stored in the PRC under PRC lawOrganization controls environment
LicensingOpen Platform terms, no enterprise variant. Training other models on the outputs, including distillation, is explicitly permittedMIT on four of the six V4 weight repositories as read October 1, 2026; the two DSpark repositories unconfirmed
ScalingProvider-managed within listed concurrencyOrganization provisions capacity
Model updatesProvider-led, and a model can retire on the day it is announcedOrganization tests and deploys
ObservabilityAPI-level signalsFull stack if implemented
Availability promiseNone. The terms disclaim availability by jurisdiction and reserve broad suspension rightsOrganization owns uptime
Operational burdenLowerHigher

Context caching

Caching on disk is enabled by default for all users, with no code change. And the discount is not a published percentage. It’s the cache-hit column of the price list. A cache hit costs 2 percent of the cache-miss rate on deepseek-flash and 3.3 percent on deepseek-v4-pro. So prompt architecture is part of the economics, not a micro-optimization.

  • Keep stable instructions stable
  • Measure the real hit rate through prompt_cache_hit_tokens
  • Do not assume every request qualifies

MIT-licensed weights

Permissive licensing removes one common blocker to a self-hosted path. And the Vision-Exp repo shows a second use. It is still published 21 days after the API retired the model, so a model DeepSeek has pulled from the API is still available for you to run.
· Six V4 repositories on DeepSeek’s weights channel, read October 1, 2026: V4.1-Flash, V4-Pro-0813, V4-Flash-Vision-Exp, V4-Flash-0731, V4-Flash-DSpark and V4-Pro-DSpark.
· MIT confirmed on the first four, which were opened. The two DSpark repositories were not opened.

  • Commercial use, modification and redistribution with attribution
  • A retirement escape hatch for a model you depend on
  • The weights channel and the API move separately, so verify against both

User isolation

The user identifier is not just a label. On an account with expanded quota, each user_id gets its own concurrency allowance, so the field is a capacity lever as well as a safety and cache boundary. Leave it empty and the empty value becomes its own user_id, which means every anonymous caller shares one bucket.
· Rate Limit page, checked October 1, 2026: the identifier must match the regex [a-zA-Z0-9\-_]+, with a maximum length of 512.

  • Do not place personal data in the identifier
  • Use stable internal pseudonyms, 512 characters maximum
  • An empty identifier is itself one shared bucket
  • Expanded-quota accounts get the limit per user_id

Concurrency planning

Published account-level limits differ sharply: 2,500 for deepseek-flash and 500 for deepseek-v4-pro. Flash is now cheaper, and on DeepSeek’s own account it measures better too. So the concurrency gap is one more reason the migration runs toward Flash rather than away from it.

  • Model peak demand
  • Handle 429 responses
  • Request capacity expansion when justified

Off-peak scheduling

The clock is a cost lever, and it just got one notch less predictable. Off-peak rates are half of the peak rates. Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday, now excluding Chinese public holidays, which are off-peak in full. Batch, backfill and evaluation runs that do not need to be interactive belong in the other seventeen hours of a weekday. Or anywhere at all on a Saturday or Sunday. The catch for a forward budget is that DeepSeek does not publish the holiday list, so your peak share has a ceiling and no exact value. Build the forecast on the ceiling and treat the holidays as upside.
· Pricing page, checked October 1, 2026, for the window text and the holiday carve-out. No calendar is listed or linked.

  • Queue deferrable work for off-peak
  • A weekend run is off-peak for the full day
  • Track the UTC windows, not local time
  • Budget peak at the ceiling, because holidays are unlisted

Service status and fallback

DeepSeek publishes per-service uptime on a live status page. Those are vendor figures on a page that is rewritten continuously, and DeepSeek publishes no SLA alongside them. The terms go further than silence: they disclaim liability for inaccessibility, disruption, delay and suspension, and reserve the right to suspend service. So the fallback provider is not a nice-to-have.
· Checked October 1, 2026: all systems operating as expected, no incidents listed, uptimes between 99.68% and 100.00% across the V4 Pro API, the V4.1 Flash API, two Chat entries, File Upload and Search.

  • Subscribe to the feed rather than checking manually
  • Uptime disclosure is not an SLA
  • The terms disclaim liability for disruption and suspension
  • Design an approved fallback provider anyway
5 / Agentic Work

Thinking inside tool use is useful, and needs hard boundaries

Cheap tokens make a bad loop cheap to run, right? That’s the whole risk here. DeepSeek’s docs list agent integrations from Claude Code and Codex to OpenCode. And its September 10 article says OpenCode and WorkBuddy fully support V4.1-Flash. The operational risk has not changed, and it is slightly worse now that tokens are cheaper again. An efficient model can take a lot of inexpensive wrong steps before anyone notices.

The bounded DeepSeek agent loop

Make the stop conditions part of the task.

Brief
Define outcome and maximum authority

State the finish line, the available tools, the prohibited actions, the step limit and the validation method.

  • Maximum turns or budget
  • No implicit external side effects
  • Named fallback
Tool-using analysis
Research this technical decision using only the supplied documentation and approved web domains. Maximum 10 tool calls. Cite every conclusion and stop if the sources conflict.
Agentic coding
Implement the change, run the tests and inspect the diff. Maximum 15 tool steps. Do not modify deployment, credentials or network configuration. If blocked, report the exact evidence.
Structured verification
Return the result as the required JSON schema, then run an independent validation pass that checks completeness, allowed values and source coverage.
Commit rule
After two verification passes, commit to the strongest supported answer. Do not continue rechecking unless new evidence appears.
6 / Behavior Playbook

Nine habits for extracting the economic advantage safely

DeepSeek rewards engineering discipline. Explicit routing. Structured outputs. Hard agent limits. Scheduling against the weekday peak windows. And evaluation on your organization’s real tasks, not on a price list that changed twice in one month.

01

Pin a current model ID

Send deepseek-flash or deepseek-v4-pro. deepseek-v4-flash and deepseek-v4-flash-vision-exp were retired on September 10, 2026 and are only temporarily routed to V4.1-Flash. deepseek-chat and deepseek-reasoner died at 2026/07/24 15:59 UTC and now fail outright.

02

Route by effort, not by model

Flash covers the range. Escalate low, high, max before you consider a different model. And write the off switch for the request format you are actually on, because the three formats spell it three ways.

03

Turn thinking off deliberately

Set it explicitly, because the default effort level is unconfirmed. Do not pay reasoning latency for transformations a schema can validate. Effort none works only in the Anthropic-compatible format. The OpenAI format needs a separate thinking object. The Responses API has no off switch at all.

04

Exploit stable prefixes

A cache hit costs 2 percent of the cache-miss rate on Flash. Design repeated context to benefit from caching without hiding changing instructions.

05

Set hard agent limits

Maximum steps, budget, tools and stop conditions belong in the task contract.

06

Evaluate cultural and political domains

Test behavior on topics that matter to your users and markets, not only on code benchmarks.

07

Separate hosted and self-managed risk

The same model can create different legal and security profiles by deployment. Hosted means PRC processing, and inputs that can be used to improve the models unless someone turns that setting off. MIT weights mean self-hosting is a real option for you.

08

Measure correction cost

Include retries and human review when you compare providers.

09

Schedule deferrable work off-peak

Off-peak rates are half of the peak rates. Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday, excluding Chinese public holidays. Everything else is off-peak, weekends and those holidays included. Move batch and evaluation runs out of the windows, and budget peak at the ceiling, because the holiday list is not published.

7 / Watch Outs

The risk profile belongs in the architecture decision

Capability answers none of the questions that actually stall a deal. Jurisdiction. Hosted-service data handling. Political behavior. Licensing. Operational responsibility. DeepSeek states the jurisdiction facts plainly in its own policies, so you can decide on them rather than guess at them.

Data residency and governing law

Stated plainly in DeepSeek’s own policies. So decide whether it is acceptable before the technical evaluation, not after.
· Privacy Policy, last updated February 10, 2026: DeepSeek directly collects, processes and stores your personal data in the People’s Republic of China.
· Open Platform Terms of Service, effective April 29, 2026: the agreement is governed by the laws of the People’s Republic of China in the mainland. Suits are filed at the location of the registered office of Hangzhou DeepSeek Artificial Intelligence Co., Ltd.

No promise the service reaches you

Before DeepSeek goes on a critical path, read the availability clause. The terms make no warranty that the service is available, or stays available, in any given jurisdiction. They disclaim liability for inaccessibility, disruption, delay and suspension from causes outside reasonable control. And they reserve broad rights to suspend. Put that beside the missing SLA and you have a provider that has written down in advance that it owes you nothing on uptime or on access. That is not a reason to walk away. It is a reason to make the fallback real and tested rather than named on a slide.
· Open Platform Terms of Service, effective April 29, 2026, sections 1.4, 7.2 and 7.5.

Retention is open-ended

The Privacy Policy commits to keeping personal data for as long as necessary to provide the services. No fixed deletion window is published. It also gives users the right to opt out of the use of their personal data for training DeepSeek’s models or optimizing its technologies. That is a control you have to exercise, not one that applies by default.
· There is no API-specific privacy policy, and the URL that would hold one returns a 404. The Open Platform Terms of Service link out to exactly three documents: the general Terms of Use, the general Privacy Policy and a Model Algorithm Disclosure. So the general policy is the applicable one by DeepSeek’s own linking, not by inference.

Your inputs can train the model

The clause is not in the API terms. It is in the general Terms of Use the API terms link to, and that document names APIs explicitly. DeepSeek reserves the right to use inputs and outputs, to a minimal extent, to provide, maintain, operate, develop or improve the services and the technologies behind them, subject to encryption, de-identification and irreversibility. There is an opt-out and it is named: a setting called Improve the model for everyone. Which reads like a consumer app switch. Have your legal review ask DeepSeek in writing where a developer account turns it off.
· Terms of Use, last updated March 27, 2026.

Your end users are not covered

Build a product on this API and DeepSeek’s privacy policy stops at you. It says the processing rules for personal data collected from end users of downstream systems built by developers on the open platform fall outside that policy. So you write your own disclosures. You answer your own data-subject requests. And you cannot hand a customer a link to DeepSeek’s page as the answer. Teams usually find this during a security review, which is the expensive place to find it.
· Privacy Policy, last updated February 10, 2026.

No separate enterprise terms

The same Open Platform Terms of Service cover individual and enterprise developers. There is no enterprise agreement, no separate data-processing addendum on any first-party page, and no published SLA behind the status page uptimes. So if your procurement process assumes a negotiated contract exists to fall back on, check that assumption early.

Political sensitivity

Evaluate behavior on politically sensitive and culturally specific topics that matter in your markets. This is an evaluation task on your own prompts. No benchmark answers it for you.

Provider concentration

Do not build a critical system without model abstraction, fallback and exit planning. A provider that retired two model IDs on the day it announced their replacement is a provider whose lineup you should be able to leave quickly.

Open-weight operations

Self-hosting transfers patching, serving, scaling, safety and monitoring to the organization. The MIT license makes that legally simple. Legally simple isn’t the same as operationally cheap.

Long-context confidence

A 1M window does not guarantee full coverage, reliable retrieval or balanced attention. DeepSeek’s efficiency claims for the new KV cache design are vendor claims about resource use, not about retrieval quality.

Agent loops

Cheap tokens can make a runaway verification or tool loop inexpensive and still operationally harmful. And tokens just got cheaper again.

Pro outlived its own shutdown date

A dated article gave Pro a September 14 ending. Seventeen days on, Pro is still fully served. That sounds like good news and it is really a planning problem, because nothing dated says the reprieve is permanent. The article conditions the routing on V4.1-Pro launching, and V4.1-Pro has not launched. The only sentence saying service continues with billing unchanged lives on the change log, a page rewritten in place with no version history, and the pricing page does not repeat it. So anyone with a production dependency on Pro should get written confirmation from DeepSeek rather than rely on any page at all. Diarize the date, re-read the source page on the same cadence, and keep a dated copy of the row you depend on. A date written down once and trusted is how you miss a deadline that moved.
· September 10 release article: all deepseek-v4-pro requests route to V4.1-Flash from 04:00 UTC on September 14, 2026, and this continues until V4.1-Pro launches.
· Change log, checked October 1, 2026: Pro continues after that date with unchanged billing.

The price moved twice in a month

This is the one to say out loud to a client, right? The peak and off-peak policy went live at 16:00 UTC on August 16, 2026 and raised every billing item. Twenty-four days and twelve hours later, at 04:00 UTC on September 10, the whole list was replaced again and the volume rates came down. Output about 9 percent lower, cache-miss input about 32 percent lower and cache-hit input about 57 percent lower than the August V4-Flash rates. Those three comparisons cannot be checked, incidentally, because DeepSeek published the August rates as a chart image and superseded rates leave the pricing page. Two full repricings in under a month, one up and one down. That’s the pattern to plan for. A vendor that can do that twice can do it again, in either direction. So budget with headroom, keep an abstraction layer, and re-read the pricing page before any commitment that rests on a published number.

Weights can move under a fixed model ID

Credit where it is due: DeepSeek names builds where buyers can see them. The pricing page identifies DeepSeek-V4.1-Flash and DeepSeek-V4-Pro-0813, matching the weights channel. The caution survives the improvement, though. A model ID is still an alias over weights that can be replaced. Naming a build on a live page is not a versioned endpoint you can pin to. And deepseek-flash is a generational name rather than a dated one, so the next replacement may not change the string you send at all. This is not hypothetical. The V4.1-Flash repository has been modified since release and the model card does not say what moved. So re-run your evaluation suite on a schedule, not only on a version bump.

Retired in the API, still on the weights channel

The DeepSeek-V4-Flash-Vision-Exp repository is still published on DeepSeek’s weights channel on October 1, 2026, 21 days after the API retired the model on September 10. It has even shown activity since. Read that two ways. It’s a real escape hatch, because an MIT-licensed model you can host does not disappear when the vendor withdraws the endpoint. It is also evidence that DeepSeek’s channels do not update together. So a live weights page is not proof that a model is still served.

The benchmark scores exist, on the wrong kind of page

The September 10 release article publishes its benchmark comparison and its price chart as images, so its benchmark and price figures can’t be quoted from it. The scores do exist as text. They are on the change log, which DeepSeek rewrites in place. A page you can read and cannot cite, and that distinction turns real the moment a procurement document needs a source line. Same trap on the August 21 article, whose benchmark figures are chart-only too. So screenshot anything you intend to rely on, treat the claim that V4.1-Flash outperforms V4-Pro as a vendor claim, and measure it on your own tasks.
· Change log, checked October 1, 2026, publishing V4.1-Flash scores as text. Among them: GPQA Diamond 90.9, Codeforces rating 3471, Terminal-Bench 2.1 at 90.6, DeepSWE v1.1 at 74.2 and HLE at 36.8. Vendor figures, on a page that cannot anchor a dated claim.

Risk decisionHosted API questionSelf-managed question
DataContent is processed and stored in the PRC. Is that acceptable for this data class?Who can access model inputs, logs and infrastructure?
SecurityWhat provider and network controls apply?How are serving stack and weights protected?
SafetyWhat provider policies and isolation apply?What filters, evaluation and abuse controls will run, and who operates them?
ReliabilityListed concurrency and a status page, no published SLA, and terms that disclaim availability by jurisdiction. What is the fallback?How does capacity scale and recover?
LegalPRC governing law, Hangzhou jurisdiction, one set of terms for everyone, and distillation from outputs expressly allowed. Who signs off?MIT license, plus internal acceptable-use and export review
PrivacyInputs can be used to improve the models unless the setting is turned off, and your own end users sit outside DeepSeek’s policy. Who writes the downstream disclosures?The organization holds the inputs and still owes its users a policy
LifecycleTwo models retired on the day of announcement. How fast can a migration run?Who owns model testing and rollout?
8 / Sources

Where these facts come from

Every source listed below is a dated first-party DeepSeek article or policy document, read on October 1, 2026. Several load-bearing facts live only on pages DeepSeek rewrites in place. Those can’t anchor a dated claim, so they’re attributed inline instead. Confirm each against the live page before you rely on it, and keep a timestamped copy of anything you’ll need to cite later. One trap worth knowing on DeepSeek’s documentation host: it serves its documentation homepage for unknown paths rather than a 404. So a page that renders is not proof the path exists.

DeepSeek-V4.1-Flash: Smarter, Faster, More EfficientDeepSeek · September 10, 2026 · the current release article · V4.1-Flash architecture and parameter claims, the retirement of V4-Flash and V4-Flash-Vision-Exp, the 04:00 UTC September 10 pricing effective hour, and the still-live sentence routing V4-Pro to V4.1-Flash from September 14 and continuing until V4.1-Pro launchesDeepSeek-V4-Pro GA ReleaseDeepSeek · August 13, 2026 · V4-Pro general availability, the unchanged calling method, and the 16:00 UTC August 16, 2026 effective hour for the peak and off-peak regime · its rate figures are published as a chart image and were superseded on September 10, 2026DeepSeek-V4-Flash-Vision-Exp Release: Multimodal API Now LiveDeepSeek · August 21, 2026 · history only · the model it announces was retired on September 10, 2026, and its benchmark figures are published as a chart image rather than as textDeepSeek V4 Preview ReleaseDeepSeek · April 24, 2026 · history only · preview parameter counts and the alias retirement it still describes in the future tenseDeepSeek Open Platform Terms of ServiceDeepSeek · effective April 29, 2026 · PRC mainland governing law, suits at the registered office of Hangzhou DeepSeek Artificial Intelligence Co., Ltd., one set of terms covering individual and enterprise developers, no warranty of availability in any given jurisdiction, broad suspension rights, and express permission to train other models on the outputs including model distillationDeepSeek Privacy PolicyDeepSeek · last updated February 10, 2026 · personal data collected, processed and stored in the People’s Republic of China, open-ended retention, the opt-out from model training, and the carve-out placing end users of developer-built downstream systems outside this policy · no API-specific privacy policy exists; the Open Platform terms link hereDeepSeek Terms of UseDeepSeek · last updated March 27, 2026 · one of the three documents the Open Platform Terms of Service link out to · it names application programming interfaces, reserves minimal-extent use of inputs and outputs to provide, maintain, operate, develop or improve the services, and names the opt-out control as a setting called Improve the model for everyone

AI Mindset

Explore the model cheatsheets