China's AI model weekly call volume tops US for 22 straight weeks; anonymous 'Space Bunny' linked to MiniMax
Editorial responsibility
- No named human review is recorded for this page.
- Source reporting is collected, normalized, translated or condensed automatically when needed.
- Automatically published source-backed update
China's AI large model weekly call volume has surpassed the United States for 22 consecutive weeks, according to OpenRouter data from September 21-27. Global total call volume reached 146 trillion tokens, up 13.18% week-over-week. China's weekly call volume was 62.22 trillion tokens, down 7.77%, while the US had 14.2 trillion tokens, down 0.07%. Among the top five models by call volume, three were Chinese: DeepSeek V4.1 Flash (first, 19.6 trillion tokens), GLM 5.3 Flash (second, 16.3 trillion), and Tencent Hy4 preview (fourth, 9.64 trillion). An anonymous model named Space Bunny (also called 'Yutu') ranked third with 13.9 trillion tokens. Developers have linked Space Bunny to MiniMax's M3.1-Flash-Preview via tokenizer tests, though MiniMax has not confirmed. MiniMax CEO Yan Junjie stated in an earnings call that text model unit compute throughput has tripled in two months, with the goal of reducing M3.1 inference costs to one-third of M3's launch cost. He expects text business to drive gross margin improvement in the second half of the year and into 2025, and that billing models may shift from per-token to per-task or value-based pricing.
Source report
September 28 — Shanghai-based AI giant MiniMax has released its latest large language model, the M3.1-Flash-Preview, now available for public beta testing, as Chinese AI models maintain their global lead in total token consumption for the 22nd consecutive week.
M3.1-Flash-Preview: Key Features
According to MiniMax, the new model supports:
- Native multimodal capabilities
- Million-token context windows
- Reliable performance for everyday development tasks, including bug fixes and full feature development
The model is designed to participate in the full development lifecycle — from issue identification and code implementation to testing and validation. It can handle edge-case logic, improve regression testing, and verify the impact of changes on existing functionality, forming a closed-loop development process from problem diagnosis to deliverable output.
The "Space Bunny" Connection
Following the release of M3.1-Flash-Preview, some developers have drawn comparisons to the recently popular anonymous model Space Bunny (also referred to as "Yutu" or "Jade Rabbit").
According to the latest rankings from global model aggregation platform OpenRouter, between September 21 and September 27:
- Total global AI model token usage reached 146 trillion tokens, up 13.18% week-over-week
- Chinese AI models accounted for 62.22 trillion tokens, down 7.77% from the previous week
- U.S. AI models accounted for 14.2 trillion tokens, down 0.07%
This marks the 22nd consecutive week that Chinese models have surpassed U.S. models in total weekly token usage.
Top 5 Global Models by Weekly Token Usage
| Rank | Model | Weekly Token Usage | Change | |------|-------|-------------------|--------| | 1 | DeepSeek V4.1 Flash | 19.6 trillion | +24% | | 2 | GLM 5.3 Flash (Zhipu) | 16.3 trillion | +16% | | 3 | Space Bunny (anonymous) | 13.9 trillion | — | | 4 | Hy4 preview (Tencent) | 9.64 trillion | -23% | | 5 | — | — | — |
Three of the top five models are Chinese. The anonymous Space Bunny model surged onto the榜单 during the Mid-Autumn Festival, ranking third.
Speculation on Space Bunny's Identity
Tokenizer tests have become a key clue in community discussions. Tests revealed that Space Bunny's token counting characteristics are consistent with MiniMax models. Some Reddit developers have further speculated that Space Bunny may be a preview version of MiniMax's M3.1 Flash. However, MiniMax has not officially confirmed this connection.
MiniMax H3: Open-Source Success
MiniMax recently open-sourced its next-generation multimodal generation model, MiniMax H3, which has achieved:
- Global No. 1 on the Artificial Analysis video editing benchmark
- Global No. 1 on the Arena text-to-video leaderboard
- Recognition from industry figures including Stable Diffusion founder Emad Mostaque and a16z partner Justine Moore
- Top popularity on Hugging Face, surpassing DeepSeek V4 Flash
CEO Commentary on Business Outlook
In a post-earnings conference call, MiniMax CEO Yan Junjie shared the following insights:
- Over the past two months, text model throughput per unit of compute has tripled
- The M3.1 model aims to reduce inference costs to roughly one-third of the M3's initial launch cost
- Cost savings will be split between passing savings to customers (expanding token volume) and improving gross margins
Yan noted that multimodal and voice models currently have relatively higher gross margins. As text model revenue grows rapidly and inference costs decline, he expects text business to become a key driver of overall margin improvement. He projected continued margin improvement in the second half of the year, with further upside in 2025.
He also predicted that as model capabilities and reliability improve, the business model may gradually shift from per-token billing to billing based on task outcomes and professional value.
Source
新浪财经Eastern
Part of this Story
Chinese AI models lead global token usage for 22nd consecutive week; anonymous model enters top three