Wire flash
Lenovo China's Huang Shan: Hybrid Token Factory can cut AI token cost from 10 yuan to 1 yuan
Editorial responsibility
- No named human review is recorded for this page.
- Source reporting is collected, normalized, translated or condensed automatically when needed.
- Automatically published source-backed update
This article reports on Lenovo's strategy to reduce the cost of AI token generation through its 'Hybrid Token Factory' solution, as explained by Lenovo China Infrastructure Business Group Strategic Technology Director Huang Shan. Huang argues that enterprises' primary anxiety is not AI technology itself but the inability to clearly account for token costs, which are exploding due to the 2025 agent boom. He introduces a 'nine-layer' cost structure where compute optimization accounts for one-third to one-half of costs, with potential to reduce token cost from ten yuan to one yuan. The Hybrid Token Factory integrates xCloud, smart computing centers, and Wanquan heterogeneous computing to standardize AI production. Huang forecasts that 'super nodes' will become the new form of public inference services, with explosive growth over the next three years, and notes that inference-specific chips (TPU, NPU, LPU) are a key trend. The article positions Lenovo as a system integrator enabling AI affordability for small and medium enterprises and traditional industries, aligned with China's 'AI+' and 'East Data West Computing' national strategies.
Source report
Computing power is becoming as fundamental as water and electricity—a transformation that signals the deepening and practical realization of universal access to computing resources. As China's "AI+" initiative and the "East-to-West Computing" strategy continue to advance, the democratization, convenience, and standardization of computing power have become clear priorities. According to data from the National Data Administration, the average daily token call volume in the Chinese market exceeded 500 trillion in June 2026 and is projected to reach 19,306 trillion by 2030.
Yet this growth has brought with it a new source of anxiety. With the explosion of AI agents in 2025, token consumption has grown exponentially. Enterprises that once said "let's try it" are now asking, "how do we account for this?"—when a single employee might use 100 million tokens in a day, the math becomes critical. In a recent media interview, Huang Shan, Strategic Technology Director of Lenovo's China Infrastructure Business Group, offered a sharp observation: what truly worries enterprises is not the technology itself, but the inability to clearly calculate the costs. This insight cuts to the heart of the industry's transition from "technology exploration" to "value realization."
Token Factory: An Industrial Model for Universal Computing
Huang Shan emphasized that a Token Factory is not simply an intelligent computing center. Rather, it is a complete system encompassing "top-level planning—data readiness—intelligent computing center—model services—workflow alignment." A Token Factory is a standardized, scalable, traceable, predictable, and billable way to produce intelligence.
Lenovo has innovatively built a Hybrid Token Factory solution, integrating products and capabilities including xCloud (Lenovo Intelligent Cloud), intelligent computing centers, Wanquan heterogeneous computing solutions, servers, storage, and data networking—forming a complete business loop from cloud platform to infrastructure. "Lenovo has the ability to integrate all these technical elements into a comprehensive solution," Huang said.
The standardization and billability of the Hybrid Token Factory enable computing power to be used "on demand, like water and electricity," lowering the barrier for AI adoption across industries and aligning with the strategic goal of universal computing under the "East-to-West Computing" initiative.
Cost Breakdown: The "Nine-Layer Tower" and Chip-Model Adaptation
The largest variable in token cost is computation, with computing efficiency optimization accounting for one-third to one-half of total costs. Huang broke this down into a "Nine-Layer Tower":
- Computing layer (7 layers of optimization): Operator libraries, communication libraries, memory semantics, inference framework parallel strategies, etc.—each layer offering roughly a 2x improvement.
- Hardware system layer: Efficiency differences between model-chip combinations can reach 30–40%, and GPU configuration differences can double that.
- Operations layer: Model gateway routing, agent orchestration, Token Hub, FinOps, anti-time-fragmentation scheduling.
- Infrastructure layer: IDC construction, electricity options, liquid cooling, etc.
"The cost of a single token can drop from ten yuan to one yuan," Huang said. "Our slogan is: unleash every bit of computing power's efficiency."
Chip-model adaptation is not a simple selection process but a systems engineering challenge—designing computing systems to run models at peak efficiency. Lenovo's large model ecosystem helps users select models and adapt chips based on specific scenarios. Cost reduction is essential for small and medium-sized enterprises and traditional industries to afford AI, enabling the "AI+" initiative to take root in the real economy. "The supply side must bring costs down before applications can flourish," Huang noted.
Serving Industries: From Top-Level Planning to On-the-Ground Deployment
Lenovo's consulting team has worked in over a dozen prefecture-level and above cities, conducting complete planning—from application to factory construction—based on each city's distinctive industries, and successfully implementing these plans. This includes everything from top-level design to full system construction, from AI agents to chip-level solutions, with long-term iterative testing and production process re-engineering.
Huang believes that AI adoption will follow the trajectory of previous technologies—not every scenario is suited for AI-native approaches. Systematic planning and industry-specific know-how are essential. "AI algorithms cannot replace all algorithms. They only play a transformative role in specific scenarios."
Future Variables: Supernodes and Specialized Chips
Supernodes are emerging as the single biggest variable in computing infrastructure.
Huang predicts that supernodes will become a new form of public inference service, replacing traditional cluster-based computing. Citing analysis from Huatai Securities, he noted that supernode revenue in 2025 remains limited, but as inference scenarios begin adopting supernodes, the market will see year-over-year explosive growth.
The fundamental driver is that Scaling Law continues to hold, making models smarter. Huang emphasized that supernodes themselves are still evolving rapidly: "Supernodes will become the new form of public inference service. For at least the next three years, supernodes will continue to evolve."
Over a three-year horizon, the trajectory of inference-specific chips is equally worth watching. TPUs, NPUs, and LPUs each have their strengths and weaknesses. Autonomous driving chips are already a clear赛道 (track), and specialized chips may emerge for scientific research and other application domains.
Clear Accounting, A New Era Unlocked
The Token Economy is not a concept—it is a bill that must be settled. From the industrial form of the Hybrid Token Factory, to the technical breakdown of the cost "Nine-Layer Tower," to the practical deployment across industries, Huang Shan has articulated Lenovo's clear role in the march toward universal computing: a comprehensive solution provider and a systematic capability integrator.
Against the backdrop of the "AI+" and "East-to-West Computing" national strategies, computing power is transitioning from a scarce resource to a universal infrastructure. But universal access requires someone to bring down costs, make the numbers clear, and deliver solutions to industry. That is precisely what Lenovo is doing: supporting universal computing with systematic technical capabilities, and turning every unit of computing power into tangible productivity.
"Make costs clear. Make ROI transparent. Unleash every bit of computing power's efficiency. Turn every unit of computing power into real productivity." This is both Huang Shan's assessment and Lenovo's most practical commitment in the Token Economy transformation.
Source
驱动中国Eastern
Part of this Story
Lenovo unveils Hybrid Token Factory to slash AI token costs from 10 yuan to 1 yuan