Huawei launches OceanStor M900 AI memory storage to tackle inference bottlenecks at Connect 2026
At Huawei Connect 2026 on September 17, Huawei launched the OceanStor M900 AI memory storage, designed for large-scale AI inference. The product uses a multi-layer KV cache architecture and Lingqu UnifiedBus to enable direct NPU-to-SSD access, reducing latency by 90% to 60 microseconds and extending SSD lifespan by 16 times. Huawei executives stated the AI industry has shifted from training to inference, with the M900 targeting TB-level KV cache demands from long-sequence agent applications.
Reference imageEditorial responsibility
- No named human review is recorded for this page.
- Reports are grouped by semantic similarity and deterministic rules. Language models may assist titles, summaries, translation and cross-source analysis; the page reads the event directly, while its address stays stable when the title changes.
- Summary covers the current reports
Cross-source coverage
Reporting timeline
Huawei launches OceanStor M900 AI memory storage for inference at Connect 2026
At Huawei Connect 2026 on September 17, Huawei副董事长、轮值董事长汪涛 announced the OceanStor M900 AI memory storage, designed to match the 昇腾超节点 (Ascend supernode) for large-scale data center inference. Huawei executives stated that the AI industry has entered a data-critical phase where inference demand surpasses training. The M900 addresses challenges from Agent applications requiring TB-level KV cache per card due to long sequences (4K to 1M tokens, 250x increase) and high-frequency interactions. The product uses a 3.5-layer KV cache architecture with 灵衢总线 (UnifiedBus) enabling NPU-to-SSD access in 60μs (90% latency reduction) and 40TB/s aggregate bandwidth, reducing data movement from 5 passes to 1. Mixed media and optimized Retention algorithms extend SSD lifespan 16x. Huawei aims to deliver a complete stack of 昇腾超节点 + 灵衢网络 + AI memory storage to improve token hit rate by 2x and reduce first-token latency, targeting all AI customers with particular value in inference scenarios.
Read sourceHuawei Launches OceanStor M900 AI Memory Storage for Inference Era
At the Huawei Connect 2026 conference, Huawei announced the OceanStor M900 AI memory storage, designed to match its Ascend super nodes for large-scale data centers. The product addresses challenges in the AI inference era, where agent applications require handling long sequences (from 4K to 1M tokens) and high-frequency interactions, leading to TB-level KV cache per card. Huawei's distributed storage president Yang Wendao stated that the M900 uses a multi-layer KV cache architecture, focusing on the L3.5 layer via the Lingqu bus, reducing data transfers from five to one. It employs hybrid media and optimized retention algorithms to extend SSD lifespan by 16 times. The M900 aims to double token hit rates and reduce first-token latency, improving user experience and return on investment. Huawei positions the product as part of a complete stack including Ascend super nodes and Lingqu network, offering an optimized option for AI infrastructure.
Read sourceHuawei Launches OceanStor M900 AI Memory Storage to Unlock Inference Bottlenecks
At the Huawei Connect 2026 conference on September 17, Huawei officially launched the OceanStor M900 AI memory storage, targeting AI inference in ultra-large data centers. The product features a 'three-chip-in-one' architecture (CPU, network, storage controller) enabling KV Cache to bypass traditional protocol conversions and go directly from NPU to SSD, reducing access latency from milliseconds to 60 microseconds (a 90% reduction) and providing 40TB/s aggregate bandwidth per cluster. Huawei executives, including distributed storage president Yang Wendao and data storage VP Wu Junjie, argued that the AI industry has moved from training to inference, where the core challenge is speed, accuracy, and cost. They noted that traditional memory and storage solutions cannot handle the growing KV Cache demands of trillion-parameter models and long-context windows. The M900 is positioned as a '3.5-layer' storage tier between memory and local SSDs. Huawei also introduced a '3+1' AI data platform integrating knowledge bases, KV Cache libraries, and memory libraries, which has improved AI-assisted diagnosis report adoption from 50% to 90% in a hospital case and reduced first-token latency by 80% in AI coding scenarios. The company outlined three future innovation directions: building data ontologies, native KV semantic direct access, and agent data resilience.
Show 2 older updatesHide older updates
Huawei Launches OceanStor M900 AI Memory Storage at Connect 2026
At Huawei Connect 2026 on September 17, Huawei launched the OceanStor M900 AI memory storage solution built on LingQu. The product is designed for agent inference systems and long-sequence workloads, which require a multi-tier KV Cache architecture. Huawei proposed a multi-layer KV cache architecture and introduced the OceanStor M900 cluster, a PB-scale L3.5-layer KV cache based on LingQu UnifiedBus that supports 'one-hop direct connection.' The M900 employs mixed media combined with an optimized Retention algorithm, extending SSD read/write lifespan by 16 times to ensure high KV hit rates and long-term reliability at the underlying level.
Huawei Launches OceanStor M900 AI Memory Storage at Connect 2026 Conference
At the Huawei Connect 2026 conference, Huawei unveiled the OceanStor M900, an AI memory storage solution built on Lingqu. The product targets agent inference systems requiring multi-tier KV Cache architectures. Huawei proposed a multi-layer KV cache architecture and introduced the OceanStor M900 cluster, a PB-level L3.5 KV cache based on Lingqu UnifiedBus that supports 'one-hop direct connection.' The M900 employs hybrid media combined with optimized retention algorithms, extending SSD read/write lifespan by 16 times and ensuring high KV hit rates and long-term reliability from the underlying infrastructure. The report was published by the STAR Market Daily on the 17th.
Read source