Wire flash
Zhipu launches GLM-5.3-FlashX AI model with 5x speed boost, pricing raised 2.5x
Editorial responsibility
- No named human review is recorded for this page.
- Source reporting is collected, normalized, translated or condensed automatically when needed.
- Automatically published source-backed update
On September 18, Chinese AI company Zhipu officially launched its new large language model, GLM-5.3-FlashX. The model's API is now live under the ModelKey GLM-5.3-FlashX. According to the company's introduction, the new version achieves inference speeds of up to 200 tokens per second, which represents a fivefold improvement over the existing GLM-5.3-Flash model. However, the pricing for the new model has been increased to 2.5 times that of the original version. The launch highlights Zhipu's continued efforts to improve model performance and inference efficiency in the competitive AI landscape.
Source report
Jin10 Data, September 18 — Zhipu has officially launched the GLM-5.3-FlashX model, capable of reaching inference speeds of up to 200 tokens per second.
The GLM-5.3-FlashX API is now available, with the model key designated as GLM-5.3-FlashX.
Key Highlights
- Inference speed: Up to 200 tokens/s — a fivefold improvement over the existing GLM-5.3-Flash model.
- Pricing: Increased to 2.5 times the cost of the original version.
Source
金十数据Eastern
Part of this Story
Zhipu launches GLM-5.3-FlashX AI model with 5x faster inference, 2.5x price increase