
Over 7 Trillion! DeepSeek V4 Flash Tops Global Weekly Usage
Data from OpenRouter shows that DeepSeek V4 Flash topped the global rankings with a weekly usage volume of 7.22 trillion tokens. Chinese large AI models have surpassed US models in weekly usage for 14 consecutive weeks, firmly holding the top spot. Xiaomi's MiMo-V2.5 ranked second, but its traffic declined week-on-week
The weekly global large AI model usage ranking released by the multi-model aggregation platform OpenRouter (July 27 to August 2) shows that DeepSeek V4 Flash ranked first with a usage volume of 7.22 trillion tokens.
According to a post by the open-source project team OpenCode, the usage volume of DeepSeek V4 Flash on its platform surged significantly. On August 1 alone, the model processed 8 trillion tokens in a single day. Of this, 5 trillion tokens were consumed under the free trial allowance, and 3 trillion tokens were paid for by developers using the OpenCode platform.
OpenRouter data shows that the total global usage volume of large AI models last week was 56.8 trillion tokens, a week-on-week decline of 2.07%. Among the listed large AI models, the weekly usage volume of Chinese large AI models reached 28.13 trillion tokens, a week-on-week decline of 14.76%. During the same period, the weekly usage volume of US large AI models was 4.38 trillion tokens, a week-on-week increase of 87.18%. Chinese large models' weekly usage volume has exceeded that of the US for fourteen consecutive weeks, firmly remaining at the top globally.
On August 5, Jiemian News checked OpenRouter's current weekly ranking and found that, as of press time, DeepSeek V4 Flash 0423 still ranked first with a usage volume of 6.92 trillion tokens. This model was first launched in April this year, focusing on balancing inference speed, long-context capabilities, and usage costs. It primarily provides high-throughput inference services for developers and commercial enterprises.

OpenRouter Weekly Usage Ranking (as of press time)
Ranking second is MiMo-V2.5, launched by Xiaomi, with a weekly usage volume of 5.1 trillion tokens, a 52% decline from the previous week. This indicates a significant pullback in traffic and a cooling of market enthusiasm.
Xiaomi's model officially entered public beta on April 23 this year and completed the open-sourcing of its entire series by the end of April. It adopts a Mixture of Experts (MoE) architecture, with total parameters exceeding the trillion level and 42 billion activated parameters. It comes standard with an ultra-long context window of 1 million tokens and covers multimodal interactions including text, voice, and images. Its low inference costs and stable agent operation capabilities are its core advantages in attracting batches of overseas developers.
In the week from July 20 to July 26, MiMo-V2.5 once topped the OpenRouter weekly usage ranking, with a single-week usage volume reaching 10.5 trillion tokens, a 12% week-on-week increase.
Currently in third place is Tencent's self-developed large model Hunyuan Hy3, with a usage scale of 5.01 trillion tokens. Hy3 is Tencent's self-developed general-purpose large model. The traffic for this model remained flat compared to the previous period, with a growth rate of 0%. Hy3 was the model with the strongest growth momentum in the ranking for the week ending July 26, when its weekly usage volume was 3.94 trillion tokens, representing a week-on-week increase of over 999%.
Hy3 was officially open-sourced on July 6, with total parameters of 295 billion and only 21 billion parameters activated per inference. It adopts a fast-and-slow thinking fusion architecture, supporting a maximum context of 256K, with significant iterations in code generation and intelligent interaction capabilities. Tencent Hunyuan previously stated that as of July 15, the total usage volume of Hy3 had increased by more than 68 times compared to the previous generation model, Hy2.
Fourth place is also held by a DeepSeek model, V4 Flash 0731. This is a newly launched version of the model, with weekly traffic of 3.45 trillion tokens. Fifth place is GPT-5.6 Luna, launched by OpenAI, with a token usage volume of 2.99 trillion this week, achieving an explosive growth of 738%.
Among the top five spots on the list, domestic large models occupied the top four positions, demonstrating extremely strong market usage enthusiasm. However, the strong catch-up by OpenAI's newly iterated model means that competition among global leading vendors has not slowed down.
Looking at the top ten, DeepSeek occupies three seats. In addition to the two V4 Flash versions, DeepSeek V4 Pro ranked sixth with 2.97 trillion tokens. Models from Anthropic, Google, NVIDIA, MiniMax, StepFun, Zhipu AI, and other companies also made the list. Free models are capturing an increasing share of traffic. Free models such as NVIDIA Nemotron 3 Ultra, Poolside Laguna S 2.1, and Inclusion AI Ling-3.0-flash achieved rapid growth, with many free models seeing week-on-week increases exceeding 50%.
Notably, the Kimi K3 model, which previously attracted significant attention from Elon Musk, temporarily fell out of the top ten this week and is currently ranked 12th. This model has total parameters of 2.8 trillion and is equipped with an ultra-long context window of 1 million tokens, making it the largest open-source large model in the world by parameter size.
Unlike the past competition model of exchanging low prices for traffic, current domestic models rely on continuously iterated inference performance and stable service capabilities to attract overseas developers. However, challenges remain. OpenAI and Anthropic continue to launch new model versions, constantly narrowing the performance gap. Multiple models in Google's Gemini series steadily occupy the top positions on the list, indicating that overseas giants still have a solid base.
Differentiation within the sector is also worth noting. Some domestic models experienced a significant decline in usage volume during this period. Xiaomi MiMo-V2.5, MiniMax M3, and Zhipu GLM5.2 all saw varying degrees of traffic pullback. As the speed of popularity rotation accelerates, if the pace of product iteration slows down, it is easy to quickly lose developer traffic.
An AI industry analyst pointed out that the overseas independent developer market has entered a phase of intense competition. Relying solely on a single version update is difficult to maintain an advantage. How to continuously iterate and efficiently control inference costs will determine whether large models can retain users in the long term.
Risk Warning and Disclaimer
The market involves risks, and investment should be approached with caution. This article does not constitute personal investment advice, nor does it take into account the specific investment objectives, financial status, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article align with their specific circumstances. Investors bear full responsibility for their own decisions.
