Wire flash
Stanford & Together AI paper: Hybrid local-cloud routing cuts energy, compute, cost by 60%-80%
Editorial responsibility
- No named human review is recorded for this page.
- Source reporting is collected, normalized, translated or condensed automatically when needed.
- Automatically published source-backed update
A new paper from Stanford University and Together AI, titled 'Intelligence per Watt: Measuring Intelligence Efficiency of Local AI,' finds that hybrid local-cloud routing reduces energy, compute, and cost by 60% to 80% compared to a batched-cloud baseline. The study reports that from 2023 to 2025, local AI's intelligence-per-watt improved by 5.3 times, while locally serviceable query coverage jumped from 23.2% to 71.3%. An iPhone 16 Pro achieved roughly 7 times higher intelligence-per-watt than workstation GPUs on the same model and precision, demonstrating mobile AI's power efficiency for lightweight queries. A diverse pool of more than 20 local models outperformed three frontier cloud models on three of four benchmarks when each query was routed to the best model, highlighting the impact of model diversity. Reducing precision from FP16 to FP4 cut inference energy by 3 to 3.5 times, at a cost of approximately 2.5 percentage points in accuracy per precision step. The paper notes a remaining weakness on the hardest reasoning tasks, where about 95% of problems were still unsolved by local models.
Source report
A new paper from Stanford University and Together AI presents a compelling analysis of local AI efficiency, revealing significant gains in performance and cost savings.
Key Findings
- Hybrid local-cloud routing reduced energy, compute, and cost by 60% to 80% compared to the batched-cloud baseline.
- From 2023 to 2025, local AI's intelligence-per-watt improved by 5.3×, while locally serviceable query coverage jumped from 23.2% to 71.3%.
- An iPhone 16 Pro achieved roughly 7× higher intelligence-per-watt than workstation GPUs on the same model and precision, demonstrating the power efficiency of mobile AI for lightweight queries.
- A diverse pool of more than 20 local models outperformed the paper's three frontier cloud models on three of four benchmarks when each query was routed to the best model, highlighting the significant impact of model diversity.
- Reducing precision from FP16 to FP4 cut inference energy by 3×–3.5×, at a cost of approximately 2.5 percentage points in accuracy per precision step. In one test, a larger FP4 model even surpassed a smaller FP16 model.
Remaining Weakness
The most significant remaining weakness is concentrated at the hard end: on the paper's most difficult reasoning slice, about 95% of problems were still unsolved by local models, even as easier and medium-difficulty tasks improved rapidly.
Source
rohanpaul_aiNeutral / independent