Tencent Cloud has officially announced the upcoming integration of the DeepSeek-V4 large language model into its specialized service platforms. Scheduled for release in mid-July 2026, the "direct supply from the manufacturer" model will be available via the TokenHub service platform and the provider's dedicated agent development environment. This integration marks a significant step in the expansion of high-performance AI tools within the cloud infrastructure that often supports decentralized application (dApp) development and blockchain data processing.
New Pricing Structure and Peak Hour Mechanisms
To manage the high demand for computational resources, Tencent Cloud is introducing a sophisticated peak and off-peak billing mechanism for the DeepSeek-V4 series. This approach allows developers to optimize costs based on network congestion. The peak hours are defined as 09:00-12:00 and 14:00-18:00 daily (UTC+8). Under this system, the costs for the high-performance DeepSeek-V4-Pro and the efficiency-focused DeepSeek-V4-Flash versions will vary as follows:
- DeepSeek-V4-Pro (Normal): 3 yuan per million tokens for inference input and 6 yuan for output.
- DeepSeek-V4-Pro (Peak): 6 yuan per million tokens for inference input and 12 yuan for output.
- DeepSeek-V4-Flash (Normal): 1 yuan per million tokens for inference input and 2 yuan for output.
- DeepSeek-V4-Flash (Peak): 2 yuan per million tokens for inference input and 4 yuan for output.
Cache hit charges are also scaled by time, with the Pro version costing 0.025 yuan during normal hours and 0.05 yuan during peak periods per million tokens.
Implications for the AI and Web3 Ecosystem
The deployment of DeepSeek-V4 on a major cloud infrastructure like Tencent Cloud provides essential tools for Web3 developers who utilize AI for smart contract auditing, automated trading bots, and on-chain data analytics. Enterprises using the Token Plan Enterprise Edition will be able to apply existing credits to offset these costs, facilitating a smoother transition for large-scale operations. As the convergence of Artificial Intelligence and Blockchain continues to accelerate, the availability of "manufacturer direct" models ensures higher reliability and lower latency for decentralized protocols requiring real-time AI inference.
The introduction of tiered pricing reflects a growing trend in the infrastructure industry to balance load across distributed networks. By providing predictable pricing for both the Pro and Flash variants, Tencent Cloud aims to accommodate a diverse range of use cases, from complex analytical tasks to high-speed automated responses within the digital asset ecosystem. This strategic rollout in mid-July is expected to enhance the accessibility of advanced LLM capabilities for the global developer community.
Frequently Asked Questions
Quick answers to the most common questions about this topic.