Wire flash
Qualcomm, partners deploy 30B-parameter MoE AI model on Snapdragon 8 Super Elite
Editorial responsibility
- No named human review is recorded for this page.
- Source reporting is collected, normalized, translated or condensed automatically when needed.
- Automatically published source-backed update
During the 2026 Snapdragon Summit on September 22, Qualcomm announced a four-party partnership with StepFun, Wulianghuo, and Longsys to deploy the StepEdge-Omni 30B-MoE model on the 6th Gen Snapdragon 8 Super Elite mobile platform. The collaboration aims to create an on-device AI agent assistant with autonomous service and personalization capabilities. The MoE (Mixture of Experts) model architecture activates only necessary expert modules per task, reducing computational load and memory bandwidth compared to dense models. The new platform features Qualcomm Oryon CPU (up to 5GHz), Adreno GPU with Matrix Cores, and Hexagon NPU with Element Accelerator. Wulianghuo's Ember inference engine reduced memory requirements by over 50% without model modification or accuracy loss. Longsys provided on-device AI storage solutions for efficient model weight loading. The demonstration showed a complete autonomous workflow from email understanding to travel planning, calendar syncing, and email drafting, all processed locally. Performance metrics include prefill throughput exceeding 330 tokens/s and decode throughput exceeding 28 tokens/s. Qualcomm's Chris Patrick stated the results enable fast, privacy-focused AI experiences on smartphones. StepFun's Yu Gang noted the platform's on-device computing power unlocks low-latency, high-privacy applications. Wulianghuo's Wang Yukun emphasized the memory reduction allows 30B models to run on phones. Longsys's Huang Qiang highlighted the importance of storage-compute synergy for scalable on-device AI.
Source report
Beijing, September 22 — During the 2026 Snapdragon Summit, Qualcomm announced a four-party collaboration with StepFun (Jieyue), Wulianghuo, and Longsys (Jiangbolang) to deploy a 30-billion-parameter Mixture-of-Experts (MoE) large language model entirely on-device, eliminating the need for cloud-based inference. Based on the Snapdragon 8 Super至尊 Edition mobile platform, the partnership deeply adapted and optimized the StepEdge-Omni 30B-MoE model for on-device inference, creating a personalized, autonomous on-device AI agent. A live demonstration on Qualcomm's reference design validated the deployment and scalability of complex AI agent workflows on smartphones.
MoE Architecture Enables On-Device AI
Unlike dense models that activate all parameters for every inference, MoE models activate only the necessary expert modules per task. This approach reduces computational load and memory bandwidth requirements while preserving large model capacity and capability. As a result, MoE is emerging as a key technical pathway for migrating cloud-scale large models to terminal devices, laying the groundwork for agent-based AI experiences on smartphones, PCs, and other endpoints.
Next-Generation Mobile Platform Provides Hardware Foundation
The Snapdragon 8 Super至尊 Edition mobile platform features:
- Qualcomm Oryon CPU with a peak frequency of 5 GHz and scalable cache architecture (Oryon FlexCache) for multitasking and system planning
- Next-gen Adreno GPU with Adreno Matrix Cores for AI workloads and Adreno Neural Fusion for combining AI and graphics processing
- Redesigned Hexagon NPU with Element Accelerator for Transformer-based AI agent workloads, plus a 50% increase in shared memory capacity for improved on-device inference performance
The platform is also equipped with LPDDR6 memory and high-speed storage channels to support agent-based AI experiences.
Quad-Party Optimization Reduces Memory Requirements by Over 50%
Key optimization outcomes:
- Wulianghuo's Ember inference engine and Ember Kernel leverage the Snapdragon platform's AI performance and software stack. Using XCompute heterogeneous scheduling, the engine achieves unified high-performance orchestration across CPU, GPU, and NPU. Without modifying the model structure or sacrificing inference accuracy, the StepFun 30B-MoE model's runtime memory requirements are reduced by over 50% compared to industry-standard approaches.
- Longsys' on-device AI storage solution enables efficient loading of model weights from flash storage. By coordinating storage and computation, the solution reduces the continuous memory footprint of large models and improves loading and inference efficiency.
The resulting on-device AI agent can autonomously perform a complete workflow—from email understanding and itinerary planning to calendar synchronization, flight and hotel recommendations, itinerary sharing, and email drafting—entirely on-device without any cloud calls.
Performance Metrics
- Prefill throughput: Over 330 tokens/s (NPU + GPU dual-engine collaboration), more than 30% higher than NPU-only solutions
- Decode throughput: Over 28 tokens/s, with sustained stable operation
- Time to first token: Millisecond-level, several to dozens of times faster than cloud API calls
Industry Perspectives
Chris Patrick, Senior Vice President and General Manager of Mobile Business at Qualcomm Technologies, said:
"This collaboration demonstrates how advances in hardware, software, memory, and AI models can drive complex AI workloads to run directly on mobile devices. By working with StepFun, Wulianghuo, and Longsys, we optimized the StepFun 30B-MoE model on the Snapdragon 8 Super至尊 Edition platform, improving throughput, reducing memory requirements, and achieving stable on-device operation. These results help bring responsive, privacy-focused AI experiences directly to smartphones."
Yu Gang, Vice President and Head of On-Device Models at StepFun, said:
"The StepFun 30B-MoE model has been successfully validated on the Snapdragon 8 Super至尊 Edition platform. The platform's on-device computing power unleashes the model's inference capabilities, expanding the application space for low-latency, high-privacy on-device AI agents."
Wang Yukun, CEO of Wulianghuo, said:
"Based on the Snapdragon 8 Super至尊 Edition platform, Wulianghuo's self-developed Ember inference engine reduces large model memory requirements by over 50%, enabling smartphones to run 30B-scale MoE models. We look forward to deepening collaboration with Qualcomm and ecosystem partners to turn smartphones into personal intelligent terminals, delivering low-latency, high-privacy on-device AI experiences."
Huang Qiang, Vice President and General Manager of the Embedded Storage Business Unit at Longsys, said:
"The large-scale deployment of on-device AI requires coordination between computing power and storage. Longsys has adapted its iSA storage agent for the Snapdragon 8 Super至尊 Edition platform, leveraging storage-computing synergy to unlock the platform's computing value and meet the storage demands of on-device large model inference. We look forward to deepening ecosystem collaboration with Qualcomm to create on-device AI solutions that deliver stable, reliable local AI experiences for end users."
Outlook
From email processing to itinerary planning, a 30-billion-parameter AI agent now performs all inference locally on a smartphone—data never leaves the device, and responses are measured in milliseconds. This demonstration shows that large-parameter MoE models have moved from architectural concepts to verifiable smartphone platform practices, and are now approaching the technical conditions for large-scale deployment.
The complete on-device deployment of a 30B MoE model on a flagship mobile platform indicates that, through model architecture and hardware-software co-optimization, smartphones are gaining the technical foundation to run larger models and handle more complex agent tasks. As hardware architectures, inference engines, and storage solutions continue to be optimized together, more sophisticated AI capabilities are expected to run stably on-device, making personalized, low-latency, privacy-focused AI agent experiences a daily reality for smartphone users.
Source
环球网Eastern
Part of this Story
Qualcomm and partners deploy 30-billion-parameter AI model on smartphones for on-device agents