Sugon launches FN Neo all-flash storage, cutting AI inference latency by 73%
Chinese IT firm Sugon launched the FN Neo, a centralized all-flash storage system for AI inference, at the 2026 CCF National Conference on Information Storage Technology. The product uses a pooled shared architecture for cluster-level KV Cache scheduling, reducing first-token response time by 73% and boosting token throughput by 150% versus local NVMe. It offers 160GB/s bandwidth, 99.999% reliability, and supports SAN/NAS/KV protocols.
Editorial responsibility
- No named human review is recorded for this page.
- Reports are grouped by semantic similarity and deterministic rules. Language models may assist titles, summaries, translation and cross-source analysis; the page reads the event directly, while its address stays stable when the title changes.
- Summary covers the current reports
Cross-source coverage
Reporting timeline
Sugon Launches FNNeo, a Centralized All-Flash Storage for AI Inference
At the 2026 CCF National Conference on Information Storage Technology, Sugon (Sugon) officially launched the FNNeo, an upgraded AI inference-native storage solution. The system uses a pooled shared architecture to enable cluster-level KVCache resource scheduling, aiming to provide a storage foundation for large-scale AI inference deployment. As the era of intelligent agents accelerates, AI inference workloads are expanding rapidly, with KVCache becoming critical production data. FNNeo addresses the fragmentation of traditional node-local caching by adopting a centralized all-flash architecture, using a unified storage pool to handle various inference data demands. Test data shows the FNNeo achieves 160GB/s bandwidth per array with sub-millisecond latency, delivering approximately 2x improvements over traditional local solutions in latency, throughput, and concurrent requests. The solution reduces model first-token response time by 73% and increases overall token throughput by 150%. For production-grade AI, FNNeo offers 99.999% reliability with slow-disk isolation and fault-tolerant reconstruction, ensuring 7x24 operation. It features a SAN/NAS/KV triple-protocol ecosystem for seamless integration with mainstream inference frameworks without major environment modifications, and through storage resource pooling, it reduces redundant storage overhead and improves GPU utilization.
Sugon Launches AI Inference Storage FN Neo, Claims 150% Token Throughput Boost
Chinese IT firm Sugon (中科曙光) announced the launch of its new AI inference-native storage product, FN Neo, at the 2026 CCF National Information Storage Technology Academic Conference on September 21. The product is designed to address the inefficiency of traditional node-local caching for large language model (LLM) inference, where KV Cache data is fragmented and cannot be shared across sessions or nodes. FN Neo uses a centralized all-flash storage architecture with a pooled shared design, enabling cluster-level KV Cache resource scheduling. According to Sugon, this allows different model sessions and compute nodes to flexibly read and reuse cache data, reducing redundant GPU computation. The company reports that the single-array bandwidth reaches 160 GB/s with sub-millisecond latency, and that the solution can reduce first-token response time by 73% and improve overall Token throughput by 150% compared to traditional local NVMe solutions. Guo Zhaobin, General Manager of Sugon's Centralized Storage Product Department, stated that the product meets the growing demands of AI agent applications requiring continuous context. FN Neo supports SAN, NAS, and KV triple-protocol integration and is compatible with mainstream inference frameworks such as vLLM, LMCache, and Mooncake. The base model is a 2U 25-bay dual-controller array with domestic high-frequency CPUs, supporting up to 1024 controllers for data-center-scale deployment.
Read sourceSugon Launches FN NEO Storage for AI Inference, Targeting KV Cache Bottlenecks
Chinese tech firm Sugon (中科曙光) has launched the FN NEO, a centralized all-flash storage system purpose-built for AI inference workloads, aiming to address the growing challenge of managing KV Cache data. As AI inference becomes more complex with multi-step, long-context interactions, KV Cache has evolved from temporary cache into a critical data asset that impacts token cost, inference efficiency, and operational costs. The FN NEO introduces a 'shared KV Cache pool' concept, enabling multiple compute nodes to share a single storage system, reducing time-to-first-token (TTFT) by 73% and increasing token throughput by 150%. The product offers 160GB/s single-array bandwidth, 99.999% financial-grade reliability, and compatibility with frameworks like vLLM and SGLang. Sugon's storage product general manager Guo Zhaobin stated that improving token output efficiency under equal compute configurations is the key direction for storage. The launch reflects a broader industry shift from distributed to centralized storage for AI inference, as the latter provides lower tail latency and better resource pooling for demanding AI workloads.
Read sourceShow 2 older updatesHide older updates
Sugon Releases Upgraded AI Inference Native Storage FN Neo at Storage Conference
According to a report from the CCF National Information Storage Technology Academic Conference, as covered by the Science and Technology Innovation Board Daily and republished by Blue Whale Finance, Sugon (中科曙光) has released a fully upgraded version of its AI inference native storage product, FN Neo. The product inherits Sugon's centralized storage high-performance advantages and, by upgrading its 'pooled sharing' capability, challenges the industry's conventional belief that centralized storage is unsuitable for AI inference. FN Neo unifies and shares inference caches across different computing nodes, addressing pressure from long-text and multi-user concurrent access. According to the report, measured data shows FN Neo reduces model time-to-first-token (TTFT) by 73%, outperforming local NVMe, and significantly improves latency, throughput, and multi-request concurrency compared to traditional local NVMe storage.
Read sourceSugon Releases Upgraded AI Inference Native Storage FN Neo for Large Model Deployment
According to a report from the Star Market Daily (科创板日报) on November 21, during the 2026 CCF National Conference on Information Storage Technology, Sugon (中科曙光) announced the release of its upgraded AI inference native storage product, FN Neo. The product inherits Sugon's centralized storage high-performance advantages and, by upgrading its 'pooled sharing' capability, aims to break the industry perception that centralized storage is unsuitable for AI inference. FN Neo unifies and shares inference caches across different computing nodes, addressing pressure from long-text and multi-user concurrent access, thereby improving utilization of existing computing results. Test data shows FN Neo reduces time-to-first-token (TTFT) by 73%, outperforming local NVMe, and delivers significant improvements in latency, throughput, and multi-request concurrency compared to traditional local NVMe storage.
Read source