Wire flash
Alibaba Cloud showcases chip-model-cloud synergy, plans mass production of Zhenwu V900 chip in Q1 2027
Editorial responsibility
- No named human review is recorded for this page.
- Source reporting is collected, normalized, translated or condensed automatically when needed.
- Automatically published source-backed update
At the 2026 Cloud Summit, Alibaba demonstrated its 'chip-model-cloud' collaborative optimization, claiming to be one of the few companies with full-stack capabilities in self-developed chips, large models, and AI cloud. The company detailed a 'synergy flywheel' where model and agent needs drive chip design, cloud platforms deliver hardware as services, and models optimize inference software and chip design. Key highlights include: the Zhenwu M890-based Lingjun supernode instance stably running Qwen3.8-Max and Kimi K3 models with over 2 trillion parameters; Qwen3.8-Max autonomously designing a chip module, achieving a 42% physical area reduction; and the model self-evolving for over a month with zero human intervention, improving its Artificial Analysis score by 12.5%. On the inference side, Tair KVCM reduced per-token cost by 50% and KVCacheStore cut first-token response time by 54%. Alibaba also announced next-gen chips: Yitian 730/720/750 for AI head nodes and agentic racks, and the Zhenwu V900 AI chip with 216GB HBM and 1200GB/s inter-chip bandwidth, planned for mass production in Q1 2027. Alibaba Cloud plans to add new service nodes in Q4 2026 and aims to operate over 20GW of global data center capacity by 2032, citing strong customer AI demand and supply chain constraints.
Source report
September 22, 2026 – At the ongoing 2026 Apsara Conference, Alibaba has for the first time demonstrated the collaborative optimization results across its "chip-model-cloud" stack. As one of the few global technology companies with in-house capabilities spanning self-developed chips, large language models, and full-stack AI cloud services, Alibaba has achieved coordinated optimization across model self-evolution, chip design, and cloud-based computing services. This integration translates full-stack technical capabilities into tangible performance gains, cost reductions, and large-scale deployment.
A Bidirectional Feedback Loop
The core of this showcase is a bidirectional feedback mechanism among three technical systems:
- Real-world demands from models and agents drive chip and system design.
- The cloud platform organizes hardware capabilities into scalable, deliverable services.
- Models, in turn, optimize inference software and chip design.
As a result, chips, models, and the cloud continuously iterate together around task effectiveness, response speed, throughput, and cost—forming a "synergy flywheel."
Chip-Model Synergy: Chips Empower Models; Models Optimize Hardware and Software
This synergy is already evident in the operation and delivery of cutting-edge models. According to Alibaba, the Lingjun supernode instances built on the Zhenwu M890 chip have been stably running models with over 2 trillion parameters, including Qwen3.8-Max and Kimi K3, and are now available at scale. This is enabled by joint optimization across chips, servers, high-speed interconnects, model software, and cloud resource scheduling. As of June 2026, the Zhenwu series has served over 650 enterprise clients.
Models are also beginning to autonomously design chips to improve infrastructure efficiency. Based on a single real-world chip module specification, Qwen3.8-Max can independently run for over 60 hours, invoking EDA tools more than 10,000 times to complete the full design flow—from functional architecture and verification to physical implementation—achieving a 42% reduction in physical area. This growing model capability makes chip design an iterative, accumulative autonomous process, while better chip hardware accelerates further breakthroughs in model intelligence, creating a self-sustaining loop between chips and models.
These cases reveal another dimension of the synergy flywheel: chips support model capability improvements, which in turn unlock hardware performance and enhance chip implementation efficiency. Progress in model capabilities is now feeding back into the technical infrastructure that supports them.
Model-Cloud Synergy: Bridging Training and Inference to Improve Compute Utilization
The cloud platform translates optimizations across all stages into overall operational efficiency. On the training side, Alibaba Cloud reduces communication latency through high-speed cluster interconnects, improves data supply via high-performance parallel file storage (CPFS), and uses the PAI asynchronous Agentic RL framework to chain together trajectory sampling, reward accumulation, model training, and parameter updates—supporting continuous model iteration. Improvements in networking, storage, and training frameworks enable large-scale compute resources to be more consistently applied to effective computation.
Leveraging this powerful AI cloud infrastructure, Qwen is gradually moving toward self-evolution during training. At the Apsara Conference, Alibaba reported that Qwen3.8-Max autonomously built its training pipeline, constructed training data, designed experiments, and identified defects—completing 33 effective iterations over more than one month with zero human intervention. Through joint optimization of RSI, post-training, and other techniques, the new version of Qwen3.8-Max achieved a 12.5% score increase on Artificial Analysis, placing it in the top tier alongside Claude and GPT's strongest models, and ahead of all domestic models including GLM5.3 and Kimi K3.
On the inference side, joint optimization focuses on long-context and complex agent task state management. As KV Cache continues to grow, model response speed and service costs are increasingly affected by cache scheduling and data movement. According to Alibaba Cloud, in relevant scenarios, Tair KVCM reduces per-token costs by 50% through unified orchestration of memory pools, KVCacheStore, and remote shared storage, while KVCacheStore cuts first-token response time by 54%.
Autonomous inference optimization initiated by the model is also accelerating. On the T-Head self-developed Zhenwu M890, Qwen3.8-Max "adapted from scratch" to the new Qwen3.8-Flash-Next model, delivering a fully operational inference stack. Compared to the first usable version, the optimized model achieved a 47% reduction in long-text first-token latency and a 60% reduction in per-output-token latency, while single-instance throughput in daily conversation scenarios improved by 96%.
Cloud-Chip Synergy: Real Business Workloads Drive Chip Design
Demands from cloud-based workloads are also shaping the product roadmap for next-generation chips. Agent operations require AI chips for model inference as well as CPUs for tool invocation, code sandboxes, retrieval, and multi-agent orchestration. In response to these workloads:
- The planned Yitian 730 is designed for AI head nodes, with enhanced single-core performance.
- The Yitian 720 features 192 cores to support large-scale concurrent execution in agentic racks.
- The future Yitian 750 will connect directly to Zhenwu AI chips via the ICN bus to improve collaboration efficiency.
The next-generation training-inference integrated AI chip, Zhenwu V900, is further upgraded for larger models and longer contexts. It features 216GB of HBM, 1200GB/s inter-chip interconnect bandwidth, and native support for FP8 and FP4 low-precision computing. Mass production is planned for Q1 2027, with large-scale deployment in Alibaba Cloud data centers. Combined with supernode architecture and cloud platform optimization, these hardware capabilities will support future model training, inference, and agent services.
Expanding Compute Supply to Scale Synergy Outcomes
Beyond performance and cost optimization, expanding supply is another key focus for the continued development of this synergy system. Alibaba Cloud plans to add new service nodes in Q4 2026 to further increase the supply of cloud-based supernodes, bringing the benefits of chip, model, and system optimization into more real-world business scenarios.
Alibaba Group CEO Eddie Wu stated that customer AI demand remains strong, but the global shortage in AI data center-related supply chains continues to constrain compute growth. Alibaba will invest jointly with partners in AI infrastructure, with a target of operating over 20GW of global data center capacity by 2032.
The Apsara Conference highlights the interconnectedness of Alibaba's full-stack technology investments: model requirements feed into chip design, hardware and software optimizations are integrated into cloud services, and models themselves participate in the next round of technical improvement. As this feedback mechanism continues to operate, technological advances in individual components are amplified across the entire system, further improving the operational efficiency and large-scale service capabilities of complex AI tasks.
Source
C114通信网Eastern
Part of this Story
Alibaba unveils Zhenwu V900 AI chip, targets 20GW data center capacity by 2032