Wire flash
Lenovo's Huang Shan: Token cost must be cut from 10 yuan to 1 yuan to drive AI adoption
Editorial responsibility
- No named human review is recorded for this page.
- Source reporting is collected, normalized, translated or condensed automatically when needed.
- Automatically published source-backed update
In an interview with DoNews, Lenovo China Infrastructure Business Group Strategic Technology Director Huang Shan outlined the company's strategy to reduce token costs and make AI accessible to enterprises. He argued that businesses are not anxious about AI technology itself but about the inability to calculate costs clearly. Huang introduced Lenovo's 'Hybrid Token Factory' solution, which integrates cloud, computing, storage, and networking to standardize AI production. He broke down token cost optimization into a 'nine-layer tower' covering computing, hardware, operations, and infrastructure, claiming costs can be reduced from ten yuan to one yuan per token. Huang noted that chip-model adaptation is a system engineering challenge, with efficiency varying by 30-40% across different combinations. He predicted that 'super nodes' will become the new form of public inference services, with explosive growth expected as inference scenarios adopt them. Huang also highlighted that Lenovo's consulting team has implemented AI plans in over a dozen cities based on local industries. The article frames Lenovo's role as a system integrator enabling the national 'AI+' and 'East-West Computing' strategies.
Source report
Computing power is becoming as fundamental as water and electricity—a clear sign that the democratization of computing resources is moving from concept to reality. As China's "AI+" initiative and the "East-to-West Computing Resources Transfer" strategy continue to advance, the standardization, accessibility, and affordability of computing power have become explicit policy goals.
According to data from the National Data Administration, daily token calls in the Chinese market surpassed 500 trillion in June 2026, and are projected to reach 19,306 trillion by 2030.
The Growing Anxiety
With the explosion of AI agents in 2025, token consumption has grown exponentially. Enterprises have shifted from "let's try it" to "let's think twice"—when a single employee can consume 100 million tokens in a day, how do you calculate the cost?
Huang Shan, Strategic Technology Director of Lenovo's China Infrastructure Business Group, recently offered a sharp observation in an interview: What truly worries enterprises is not the technology itself, but the inability to clearly account for costs. This insight cuts to the heart of the industry's transition from "technology exploration" to "value realization."
Token Factory: The Industrial Form of Computing Democratization
Huang Shan defines a Token Factory not as a simple intelligent computing center, but as a complete system encompassing:
- Top-level planning
- Data readiness
- Intelligent computing centers
- Model services
- Workflow alignment
A Token Factory is a standardized, scalable, traceable, predictable, and billable way to produce intelligence.
Lenovo has innovatively built a Hybrid Token Factory Solution, integrating:
- xCloud (Lenovo Intelligent Cloud)
- Intelligent computing centers
- Wanquan heterogeneous computing solutions
- Servers, storage, and data networking
This creates a complete business闭环 from cloud platform to infrastructure. "Lenovo has the capability to integrate all these technical elements into a holistic solution," Huang said.
The standardization and billability of the Hybrid Token Factory enable computing power to be used "on demand, like water and electricity," lowering the barrier for AI adoption across industries and aligning with the strategic goals of the "East-to-West Computing Resources Transfer."
Technical Breakdown: The "Nine-Layer Tower" of Costs and Chip-Model Adaptation
The biggest variable in token cost is computation, with computing efficiency optimization accounting for one-third to one-half of total costs. Huang Shan breaks this down into a "Nine-Layer Tower" :
- Computing Layer (7 layers of optimization): Operator libraries, communication libraries, memory semantics, inference framework parallel strategies, etc.—each layer can deliver roughly a 2x improvement.
- Hardware System Layer: Efficiency differences between model-chip combinations can reach 30–40%, and GPU configuration differences can double that.
- Operations Layer: Model gateway routing, agent orchestration, Token Hub, FinOps, anti-time-fragmentation scheduling.
- Infrastructure Layer: IDC construction, electricity sourcing, liquid cooling, etc.
"The cost of a single token can drop from ten yuan to one yuan," Huang said. "Our slogan is: Unleash every ounce of computing power's efficiency."
Chip-model adaptation is not a simple selection process—it is a systems engineering challenge: designing computing systems that allow models to run at peak efficiency. Lenovo's large model ecosystem helps users select models and adapt chips based on specific scenarios. Only by reducing costs can small and medium-sized enterprises and traditional industries afford AI, enabling the "AI+" initiative to truly take root in the real economy.
"The supply side must bring costs down—only then can applications flourish," Huang emphasized.
Serving Every Industry: From Top-Level Planning to On-the-Ground Implementation
Lenovo's consulting team has completed full-cycle planning and implementation—from application to factory construction—in over a dozen prefecture-level and above cities, tailored to each city's distinctive industries. This includes top-level design, overall system construction, agent deployment, and chip-level solutions, all through long-term trial and error and production process reengineering.
Huang believes that AI adoption still follows the trajectory of previous technologies. Not every scenario is suitable for AI-native approaches; systematic planning and industry know-how are essential. "AI algorithms cannot replace all algorithms. They only play a transformative role in specific scenarios."
Future Variables: Supernodes and Specialized Chips
Supernodes are emerging as the biggest variable in computing infrastructure.
Huang predicts that supernodes will become a new form of public inference service, replacing traditional cluster computing. Citing analysis from Huatai Securities, he noted that supernode revenue in 2025 remains limited, but as inference scenarios begin adopting supernodes, the market will see year-over-year explosive growth.
The fundamental driver is that Scaling Law continues to make models smarter. Huang emphasized that supernodes themselves are still evolving rapidly: "Supernodes will become the new form of public inference service, and they will continue to evolve for at least the next three years."
Over a three-year horizon, the trajectory of inference-specific chips is equally worth watching. TPUs, NPUs, and LPUs each have their strengths and weaknesses. Autonomous driving chips are already a clear赛道, and specialized chips may emerge for scientific research and other application domains.
Clear Accounting, A New Era
The Token Economy is not a concept—it is a bill that must be settled. From the industrial form of the Hybrid Token Factory, to the technical breakdown of the "Nine-Layer Tower" of costs, to the practical implementation across industries, Huang Shan presents Lenovo's clear role in the democratization of computing power: a holistic solution provider and a systematic capability integrator.
Against the backdrop of China's "AI+" and "East-to-West Computing Resources Transfer" strategies, computing power is transitioning from a scarce resource to a universal infrastructure. The prerequisite for universal access is that someone brings down costs, makes the accounting clear, and delivers solutions that work in the real economy.
This is precisely what Lenovo is doing: supporting the democratization of computing power with systematic technical capabilities, and turning every unit of computing power into tangible productivity.
"Make costs clear. Make ROI calculable. Unleash every ounce of computing power's efficiency, and turn every unit of computing power into real productivity." This is both Huang Shan's judgment and Lenovo's most practical commitment in the Token Economy transformation.
Source
DoNewsEastern
Part of this Story
Lenovo unveils Hybrid Token Factory to slash AI token costs from 10 yuan to 1 yuan