Wire flash
Sugon launches FNNeo centralized all-flash storage for AI inference workloads
Editorial responsibility
- No named human review is recorded for this page.
- Source reporting is collected, normalized, translated or condensed automatically when needed.
- Automatically published source-backed update
At the 2026 CCF National Conference on Information Storage Technology, Sugon (Sugon) officially launched the FNNeo, an upgraded AI inference-native storage solution. The system uses a pooled shared architecture to enable cluster-level KVCache resource scheduling, aiming to provide a storage foundation for large-scale AI inference deployment. As the era of intelligent agents accelerates, AI inference workloads are expanding rapidly, with KVCache becoming critical production data. FNNeo addresses the fragmentation of traditional node-local caching by adopting a centralized all-flash architecture, using a unified storage pool to handle various inference data demands. Test data shows the FNNeo achieves 160GB/s bandwidth per array with sub-millisecond latency, delivering approximately 2x improvements over traditional local solutions in latency, throughput, and concurrent requests. The solution reduces model first-token response time by 73% and increases overall token throughput by 150%. For production-grade AI, FNNeo offers 99.999% reliability with slow-disk isolation and fault-tolerant reconstruction, ensuring 7x24 operation. It features a SAN/NAS/KV triple-protocol ecosystem for seamless integration with mainstream inference frameworks without major environment modifications, and through storage resource pooling, it reduces redundant storage overhead and improves GPU utilization.
Source report
At the 2026 CCF National Conference on Information Storage Technology, Sugon officially unveiled the fully upgraded FNNeo, an AI inference-native storage solution. Built on a pooled shared architecture, the system enables cluster-level KVCache resource scheduling, providing a robust storage foundation for large-scale AI inference deployment.
Addressing the Challenges of Agent Era Inference
As the era of intelligent agents accelerates, AI inference workloads are expanding rapidly, with inference accounting for an increasing share of total computing demand. In large model and agent-based scenarios, KVCache has become a core production asset during the inference process. Traditional node-local caching models suffer from resource fragmentation, making cross-session and cross-node reuse difficult—ultimately constraining the overall performance of large model inference.
FNNeo is designed to address these industry pain points. Based on a centralized all-flash architecture, it moves beyond the limitations of single-node hardware, using a unified storage pool to handle diverse data requirements in inference workloads.
Performance Gains and Measured Results
According to real-world test data, a single FNNeo array delivers up to 160 GB/s bandwidth with sub-millisecond latency. Compared to traditional local storage solutions, FNNeo achieves approximately 2x improvements in latency, throughput, and concurrent multi-request handling.
As a dedicated centralized all-flash storage solution for AI inference, FNNeo reduces first-token response time by 73% and increases overall token throughput by 150%, significantly enhancing user interaction experience.
Enterprise-Grade Reliability and Ecosystem Integration
Designed for production-level AI workloads, FNNeo offers 99.999% high-reliability operation assurance. It ensures 7×24 uninterrupted inference service through mechanisms such as automatic slow-disk isolation and fault-free reconstruction.
The product also features a triple-protocol unified open ecosystem supporting SAN, NAS, and KV protocols, enabling seamless integration with mainstream inference frameworks. Enterprises can adopt FNNeo without large-scale modifications to their existing inference environments.
By pooling storage resources, the solution reduces redundant storage overhead across multiple nodes and improves the overall utilization of computing hardware such as GPUs.
(Proofread by Li Zhengcao)
Source
爱集微Eastern
Part of this Story
Sugon launches FN Neo all-flash storage, cutting AI inference latency by 73%