Huawei launches OceanStor M900 AI memory storage, KV Cache bypasses protocol conversion to cut latency to 60 microseconds
Editorial responsibility
- No named human review is recorded for this page.
- Source reporting is collected, normalized, translated or condensed automatically when needed.
- Automatically published source-backed update
At the Huawei Connect 2026 conference on September 17, Huawei officially launched the OceanStor M900 AI memory storage, targeting AI inference in ultra-large data centers. The product features a 'three-chip-in-one' architecture (CPU, network, storage controller) enabling KV Cache to bypass traditional protocol conversions and go directly from NPU to SSD, reducing access latency from milliseconds to 60 microseconds (a 90% reduction) and providing 40TB/s aggregate bandwidth per cluster. Huawei executives, including distributed storage president Yang Wendao and data storage VP Wu Junjie, argued that the AI industry has moved from training to inference, where the core challenge is speed, accuracy, and cost. They noted that traditional memory and storage solutions cannot handle the growing KV Cache demands of trillion-parameter models and long-context windows. The M900 is positioned as a '3.5-layer' storage tier between memory and local SSDs. Huawei also introduced a '3+1' AI data platform integrating knowledge bases, KV Cache libraries, and memory libraries, which has improved AI-assisted diagnosis report adoption from 50% to 90% in a hospital case and reduced first-token latency by 80% in AI coding scenarios. The company outlined three future innovation directions: building data ontologies, native KV semantic direct access, and agent data resilience.
Source report
At the HUAWEI CONNECT 2026 held on September 17, Huawei officially unveiled the OceanStor M900 AI Memory Storage, designed for AI inference scenarios in ultra-large-scale data centers. The new product strengthens the infrastructure for the Agentic AI era, driving the AI foundation from a computing-centric model toward a new phase of coordinated computing, networking, and storage.
AI’s Second Half: From “Training Success” to “Fast, Accurate, and Cost-Effective Inference”
“The second half of AI lies in data.” This was a recurring theme emphasized by Huawei’s storage team at the conference.
According to Yang Wendao, President of Huawei’s Distributed Storage Domain, the industry has moved past the concentrated construction phase of large model training and is now accelerating toward the large-scale deployment of inference.
“After training a model, we must ultimately answer one question: What is it useful for, and what benefits does it bring to our lives? AI programming is the best example, and it can be widely applied. When enterprises start actually purchasing AI models and using agents to complete specific tasks, the core challenge shifts from ‘Can we train it?’ to ‘Is the inference fast enough, accurate enough, and cheap enough?’”
This assessment is based on several trends:
- Large model parameters are approaching the 10 trillion level.
- AI agents are entering production environments such as healthcare, finance, and software development.
- Context windows have surpassed one million tokens.
- The scale of KV Cache generated during inference continues to grow, making traditional GPU memory and DRAM solutions unsustainable in terms of both capacity and cost.
Building a multi-tier storage system that coordinates GPU memory, DRAM, and SSD — creating a large-capacity, fully shared memory space — has become an industry consensus.
Wu Junjie, Vice President of Huawei’s Data Storage Product Line, noted that different industries face different bottlenecks in AI deployment:
- Some enterprises lack sufficient raw data and high-quality training corpora.
- Some institutions have computing resources but suffer from low accuracy in knowledge retrieval during inference — for example, some medical scenarios achieve only 50–60% accuracy, failing to meet clinical production standards.
- AI programming and financial agent applications are constrained by insufficient KV Cache capacity, degrading inference throughput and response experience.
“Many customers are accustomed to patchwork-style construction — expanding storage when performance is insufficient, adding hardware when latency is high. But AI infrastructure requires holistic planning; piecemeal expansion cannot solve systemic problems.”
Wu emphasized that as large model hardware becomes increasingly similar and the gap in available computing power narrows, the ability to build a high-quality data foundation tailored for AI has become the key differentiator for enterprise AI deployment. Huawei’s full-chain AI data infrastructure solution covers data preparation, model adaptation, inference execution, agent development, and data protection. Enterprises need to plan their complete architecture based on their own business needs, rather than stacking equipment in a fragmented manner.
M900 Breakthrough: “Three-in-One” Architecture Enables Direct KV Cache Access to SSD
It is against this backdrop that Huawei launched the OceanStor M900 AI Memory Storage, targeting AI inference in ultra-large-scale data centers. The product features the industry’s first “three-in-one” architecture integrating CPU, networking, and storage controller. Leveraging Huawei’s Lingqu high-speed interconnect network, it achieves “one-hop direct access” from NPU to SSD for KV Cache, eliminating the complex protocol conversions and CPU forwarding required in traditional solutions.
Performance highlights:
- Access latency reduced from milliseconds to 60 microseconds — a 90% reduction.
- Single cluster provides 40 TB/s aggregate access bandwidth — 1.5 times improvement over industry alternatives.
- In typical AI programming scenarios, inference cluster token throughput doubles, and first-token latency is halved.
In the engineering practice of large model inference, the industry has gradually formed a consensus on a four-tier hierarchical KV Cache storage system. However, the OceanStor M900 targets what is referred to as the “3.5th tier.”
Yang Wendao explained:
“GPU memory reads at nanosecond speeds with TB-level bandwidth; DRAM is still nanosecond-level but bandwidth drops to GB-level; local SSDs read at millisecond speeds with TB-level capacity. Are these three tiers enough? No — because KV Cache is too large.”
He stressed that the “3.5th tier” is not a Huawei invention but a natural technological evolution driven by the industry’s capacity and cost bottlenecks in inference. Huawei is simply the first to launch a dedicated product for this tier.
Regarding compatibility with heterogeneous computing, Yang clarified:
“The M900 will not create new ecosystem barriers. It works with the Ascend SuperNode and Lingqu network as part of an integrated supernode, making it transparent to upper-layer applications.”
For non-Ascend or non-SuperNode scenarios, Huawei’s already-launched OceanStor M800 supports multiple heterogeneous computing platforms through standard networking and hardware adaptation.
AI Data Platform: “3+1” Enables Synergy Among Knowledge, Cache, and Memory
While the OceanStor M900 targets inference in ultra-large-scale data centers, Huawei’s newly introduced “3+1” AI Data Platform Solution addresses broader enterprise-level AI deployment needs.
According to Wu Junjie, the solution is built on Huawei’s OceanStor AI Semantic Storage and integrates three core components:
- Knowledge Base
- KV Cache Library
- Memory Library
Supported by UCM (Unified Cache Management) technology, the solution has already delivered significant results in over 20 key business scenarios, including AI-assisted diagnosis and AI programming.
Medical Scenario Example:
“Some medical clients achieved only 50–60% accuracy in knowledge retrieval for assisted diagnosis, making the data unsuitable for clinical use.”
After adopting Huawei’s AI data platform, a leading Chinese tertiary hospital transformed massive volumes of medical literature, disease codes, and patient histories into a high-precision knowledge base. Results:
- AI-assisted diagnosis report adoption rate rose from 50% to 90%.
- Daily agent calls increased from 2,000 to 16,000.
AI Programming Scenario:
An enterprise using the KV Cache Library and UCM technology adopted a “compute by lookup” approach, achieving:
- First-token latency reduced by 80% to under 10 seconds.
- Token throughput increased by 2.2 times.
“The value of the memory library lies in enabling AI to accumulate memory and experience from past actions and responses over long-term use, making it increasingly intelligent. It can even support cross-enterprise open-source memory libraries for import and cold start,” Wu said.
Three Innovation Directions: Building a Coordinated Computing-Storage-Networking AI Foundation
Looking ahead, at the “Huawei Data Storage Summit” held during HUAWEI CONNECT 2026 on September 17, Yuan Yuan, Huawei Vice President and President of the Data Storage Product Line, outlined three major innovation directions for the AI data platform:
- Build Data Ontology: Improve agent task accuracy to over 90%.
- Native KV Semantic Direct Access: Deliver TB-level storage bandwidth per chassis, reduce I/O latency from 350 microseconds to 55 microseconds, and double KV Cache loading speed.
- Agent Data Resilience: Achieve second-level risk interception and minute-level consistency recovery through an intent recognition engine, with 99% interception of high-risk operations and 99% ransomware detection.
Yuan Yuan previously stated that the second half of AI lies in data. Looking at AI development trends, future AI infrastructure will no longer rely on breakthroughs in computing power alone, but will require a systematic engineering approach integrating computing power, storage capacity, and network bandwidth. Storage is no longer a mere accessory to computing — it is the core foundation that determines how far and how fast Agentic AI can go. In this race, Huawei is positioning itself as both a definer and co-builder of the storage foundation for the AI inference era, with an open ecosystem approach.
Source
环球网Eastern
Part of this Story
Huawei launches OceanStor M900 AI memory storage to tackle inference bottlenecks at Connect 2026