Alibaba Cloud launches next-gen CPFS storage, cutting AI training costs by 69%
Editorial responsibility
- No named human review is recorded for this page.
- Source reporting is collected, normalized, translated or condensed automatically when needed.
- Automatically published source-backed update
At the 2026 Yunqi Conference on September 22, Alibaba Cloud announced the launch of its next-generation high-performance storage system, CPFS (Cloud Parallel File Storage), designed for AI training clusters. The new system offers up to hundreds of TB/s throughput and billions of IOPS, with a single file system scaling to 100 PiB, a five-fold increase. It aims to accelerate training for million-card models and handle large-scale multimodal data. In practical training scenarios, it reduces average model startup time by 50%, increases peak compute utilization by 30%, lowers AI storage costs by 69%, and doubles business capacity. The system eliminates the need to split or migrate data across training phases. Alibaba Cloud also introduced KVCacheStore, a caching engine for inference that sits between GPU memory, host memory, and remote storage, achieving a 20%+ cache hit rate improvement. Alibaba Cloud storage product head Jiang Jiangwei stated the company will continue to innovate across compute, networking, and storage to provide a high-performance, scalable, and cost-optimized storage foundation for large model training and AI applications.
Source report
Shanghai, September 22 (Reported by Yang Xiangfei) — At the 2026 Apsara Conference, Alibaba Cloud officially launched its next-generation high-performance storage solution, CPFS (Cloud Parallel File Storage), designed for next-generation AI training clusters.
The new CPFS delivers up to hundreds of TB/s throughput and hundreds of millions of IOPS, with a single file system capacity expanded fivefold to 100 PiB. This enables accelerated model training on million-card clusters and supports large-scale model training and multimodal data processing.
In real-world training scenarios, the new CPFS achieves:
- 50% reduction in average model startup time
- 30% improvement in peak compute utilization
- 69% reduction in AI storage costs
- Doubled business capacity
Key Performance Improvements
The next-generation CPFS file system features:
- 5x increase in capacity scale
- 10x improvement in metadata performance
- 100x increase in file count scale
- Scalability to 100 PiB per file system, supporting trillions of files
Enterprises can now manage training corpora, model parameters, checkpoints, and experimental results within a unified file system—without needing to repeatedly split or migrate data across different training phases. The system scales capacity and throughput in sync with GPU clusters.
New KVCache Acceleration Engine for Inference
For large model inference, Alibaba Cloud also introduced the KVCacheStore, a KVCache acceleration engine. Positioned between GPU memory, host memory, and remote shared storage, it forms a new G3.5 storage layer built on near-compute deployment, high-performance networking, and elastic shared storage.
KVCacheStore uses a "storage for compute" approach to handle large caches generated by long-context, multi-turn conversations, and complex Agent tasks. Key specifications include:
- 40 GB/s throughput per compute node
- Millions of QPS via batch interfaces
- Support for hundreds of billions of KV storage per instance
- Over 20% improvement in cache hit rate in real-world tests
Industry Perspective
"Every model leap is the best opportunity for cloud storage to reinvent itself," said Jiang Jiangwei, Head of Alibaba Cloud Storage Products. "Alibaba Cloud will continue to drive coordinated innovation in compute, networking, and storage—enabling data to enter compute faster, making fuller use of GPUs, and providing a high-performance, scalable, and cost-optimized storage foundation for large model training and AI applications."
Source
上海证券报Eastern