Wire flash
TechZhipu open-sources GLM-5.3-Flash multimodal model, matching Claude Opus 4.8 performance
Editorial responsibility
- No named human review is recorded for this page.
- Source reporting is collected, normalized, translated or condensed automatically when needed.
- Automatically published source-backed update
Chinese AI company Zhipu has launched and open-sourced GLM-5.3-Flash (320B-A18B), the first native multimodal model in its GLM-5 series. The model achieved an AA comprehensive intelligence index score of 57, matching the performance of Claude Opus 4.8. Its pricing is set at 1/10 of GLM-5.3, and with a limited-time discount, it costs only 1/40 of Opus 4.8. The model has been integrated into platforms like ZCode for open API access. It employs a hybrid architecture combining sparse attention and linear attention, with inference services already running on domestic chip clusters. The article suggests this price-performance ratio could make long-context and coding tasks routine by removing previous cost barriers.
Source report
Zhipu has released and open-sourced GLM-5.3-Flash (320B-A18B), the first native multimodal model in the GLM-5 series. The model achieves an AA comprehensive intelligence index score of 57, matching the performance of Claude Opus 4.8.
Key Highlights
- Pricing: GLM-5.3-Flash is priced at 1/10 of GLM-5.3. During a limited-time discount, it is available at 1/40 of the cost of Opus 4.8.
- Access: The model has been integrated into platforms such as ZCode for open API access.
- Architecture: It adopts a hybrid architecture combining sparse attention and linear attention.
- Deployment: Inference services are already running on domestic chip clusters.
Implications
With performance comparable to Claude Opus 4.8 but at only 1/40 of its price, long-context and coding tasks that were previously constrained by high inference costs may now become routine usage.
Source
aihotEastern
Part of this Story
Zhipu AI Open-Sources GLM-5.3-Flash, a Cost-Efficient Multimodal Model