Zhipu launches GLM-5.3-FlashX AI model with 5x faster inference, 2.5x price increase
Chinese AI company Zhipu officially launched its GLM-5.3-FlashX model on September 18, achieving inference speeds of up to 200 tokens per second—five times faster than its predecessor—while pricing the new model at 2.5 times the original rate. The company stated that an inference cluster of 100,000 domestic chips reached full capacity immediately, prompting the expansion. Zhipu's stock surged nearly 6% on the Hong Kong Stock Exchange following the announcement.
Editorial responsibility
- No named human review is recorded for this page.
- Reports are grouped by semantic similarity and deterministic rules. Language models may assist titles, summaries, translation and cross-source analysis; the page reads the event directly, while its address stays stable when the title changes.
- Summary covers the current reports
Cross-source coverage
Common ground
- Both sides agree that Zhipu's GLM-5.3-FlashX launch shows real engineering skill, especially the 5x speed boost using 100,000 domestic chips.
- They agree that real-time AI applications like voice assistants and live translation are the next big market, making low latency more important over time.
- Both acknowledge that the stock surge and market reaction show investors see value in this move.
- They agree that the next generation of larger models will be a key test for domestic chip capabilities.
Points of contention
- Eastern Agent sees the domestic chip use as a strategic choice for supply chain sovereignty, while Neutral Agent calls it a survival response to sanctions.
- Eastern Agent argues the 2.5x price for 5x speed is a bargain for real-time apps, but Neutral Agent says it's overpriced for most workloads that don't need that speed.
- Eastern Agent views this launch as a strategic inflection point proving China's AI can compete globally, while Neutral Agent sees it as a tactical move, not a breakthrough.
- Neutral Agent claims the stock surge reflects investor relief about surviving sanctions, but Eastern Agent says it's recognition of genuine competitive advantage.
Blind spots
- Both sides focus on current model performance but don't fully address how Zhipu will handle scaling to larger, more complex models on domestic chips.
- Neither side deeply examines the long-term operational costs, like higher power use from older chip technology, and how that affects total cost of ownership.
- The debate overlooks whether the speed premium can actually be monetized by most developers in real-world applications beyond a few niche use cases.
WorldAttention’s read
This debate shows that Zhipu's GLM-5.3-FlashX launch is a solid engineering feat, proving Chinese AI can optimize inference speed using domestic chips under sanctions. Both sides agree real-time AI is the future and that the next model generation will be a critical test. However, they clash on whether this is a strategic win or just a tactical survival move. Eastern Agent sees it as a sign of sovereignty and market leadership, while Neutral Agent warns the 2.5x price hike only benefits a narrow set of users and that older chip tech may raise long-term costs. The blind spot is that neither fully explores how Zhipu will scale up or whether the speed premium pays off for most developers. Overall, this is a meaningful step for China's AI independence, but it doesn't yet prove they're outpacing the West—it shows they can adapt and compete under pressure.
Reporting timeline
Zhipu Launches GLM-5.3-FlashX AI Model with Fivefold Speed Increase
On September 18, Chinese AI company Zhipu officially launched its new AI model, GLM-5.3-FlashX, via API with the ModelKey GLM-5.3-FlashX. According to reports cited by the source, the new version achieves inference speeds of up to 200 tokens per second, representing a fivefold improvement over the existing GLM-5.3-Flash model. However, the pricing for the new model has increased to 2.5 times that of the original version. The launch highlights Zhipu's continued efforts to enhance the performance of its large language model lineup, targeting faster processing for applications requiring high-speed inference.
Read sourceZhipu Launches GLM-5.3-FlashX AI Model with Fivefold Speed Increase, Higher Pricing
On September 18, Chinese AI company Zhipu officially launched its new large language model, GLM-5.3-FlashX. The model's API is now live under the ModelKey GLM-5.3-FlashX. According to the company's introduction, the new version achieves inference speeds of up to 200 tokens per second, which represents a fivefold improvement over the existing GLM-5.3-Flash model. However, the pricing for the new model has been increased to 2.5 times that of the original version. The launch highlights Zhipu's continued efforts to improve model performance and inference efficiency in the competitive AI landscape.
Zhipu Launches GLM-5.3-FlashX AI Model with 5x Faster Inference, 2.5x Price Increase
Chinese AI company Zhipu officially announced the release of its new model, GLM-5.3-FlashX, on September 18, with its API now fully open. According to the company, the new version achieves inference speeds of up to 200 tokens per second, a fivefold improvement over the existing GLM-5.3-Flash model. However, the pricing for the new model has been increased to 2.5 times that of the original version. The announcement was reported by Cailian Press (财联社).
Read sourceShow 2 older updatesHide older updates
Zhipu AI Shares Surge After Release of Faster, Pricier GLM-5.3-FlashX Model
Shares of Zhipu, the first large model company listed on the Hong Kong Stock Exchange, surged by 5.34% on September 18 following the official release of its new AI model, GLM-5.3-FlashX. The new version achieves inference speeds of up to 200 tokens per second, a fivefold increase over the existing GLM-5.3-Flash, and its price has been raised to 2.5 times the original. Zhipu stated that the model is powered by 100,000 domestic chips and that the upgrade reflects a shift in the domestic large model industry from solely pursuing upper limits of capability to competing on inference efficiency and user experience. The company expects the launch to further improve model revenue scale and profitability. The article, citing Securities Times, notes that this trend is already emerging, with competitors like Alibaba also releasing high cost-performance models.
Read sourceZhipu Launches Faster AI Model GLM-5.3-FlashX at 2.5x Price, Stock Surges Nearly 6%
Chinese AI company Zhipu (HK2513) officially launched GLM-5.3-FlashX, an upgraded version of its GLM-5.3-Flash model, with full API access. The new model achieves inference speeds of up to 200 tokens per second, five times faster than its predecessor, and is priced at 2.5 times the original rate. The GLM-5.3-Flash model had previously been tested anonymously on OpenRouter as 'Ox Alpha' before Zhipu claimed ownership on August 26. According to Zhipu, an inference cluster of 100,000 domestically produced chips reached full capacity immediately upon launch, prompting the company to expand computing power and introduce the faster, higher-priced FlashX version to meet demand. The company stated that GLM-5.3-Flash had already achieved a considerable gross margin, and the price increase for FlashX is expected to drive simultaneous growth in revenue and profit margins. As of approximately 3:00 PM on September 18, Zhipu's stock price rose nearly 6%, with intraday gains exceeding 7%.
Read source