Concurrency limit specified for V4-Flash-Vision-Exp at 2500
Developers using this model now have a defined concurrency limit of 2500 for rate limiting purposes
Model APIsplatform.deepseek.com10 published changes
Developers using this model now have a defined concurrency limit of 2500 for rate limiting purposes
Users can now see transparent pricing for the new model: $0.007/$0.014 per 1M input tokens (cache hit) and $0.22/$0.44 per 1M input tokens (cache miss), with higher rates for other token types
Developers can now use the new deepseek-v4-flash-vision-exp model to process images in addition to text, enabling multimodal AI capabilities.
api-docs.deepseek.com · also seen at 1 more source
All DeepSeek API users will face new pricing structure effective August 16, 2026. Peak hours (01:00-04:00 and 06:00-10:00 UTC) charge double the off-peak rates. Off-peak rates are half of peak rates.
API users need to update their model references from 'deepseek-v4-pro' to 'deepseek-v4-pro-0813' for correct model targeting.
Developers using DeepSeek-V4-Pro can now use the Responses API, previously limited to deepseek-v4-flash model.
Developers can now access an updated DeepSeek-V4-Pro model (version 0813) by calling deepseek-v4-pro. The calling method remains unchanged.
API pricing will increase to 2x during peak hours (9:00–12:00 and 14:00–18:00 Beijing Time, UTC+8) daily. This affects all billing items. Effective date pending official announcement.
Developers using deepseek-v4-flash will automatically receive the latest model version (0731) without needing to change their API calls or code.
api-docs.deepseek.com · also seen at 1 more source
API users using these legacy model names must migrate to deepseek-v4-flash before the deprecation date to avoid service disruption.
api-docs.deepseek.com · also seen at 1 more source