Techno Time

AI Efficiency Rises 18-Fold in 16 Months as Models and Hardware Improve

Monday 17 August 2026 11:42
AI efficiency
AI efficiency

The efficiency of running artificial intelligence models locally has increased sharply, with “Intelligence per Joule” rising about 18-fold over a 16-month period, according to a research study that analyzed more than one million real-world queries, over 20 language models and eight AI accelerators between 2023 and 2025.

The study, conducted by researchers including Jon Saad-Falcon and Avanika Narayan as part of a research team associated with Stanford Hazy Research, reflects a broader shift in how AI progress is measured, with energy efficiency becoming increasingly important alongside model quality and response speed.

Intelligence per Joule emerges as a new efficiency metric

Intelligence per Joule measures the amount of performance or accuracy delivered relative to the energy consumed to produce an answer.

Unlike intelligence-per-watt measurements, the metric evaluates efficiency across the full task, taking into account both energy consumption and the time required to complete the task. A system can therefore become more efficient either by using less energy or by completing the task faster.

The metric is becoming increasingly relevant as more AI inference workloads move from data centers to personal devices.

Model improvements drive higher efficiency

The study found that advances in local models played a significant role in improving AI efficiency.

Between 2023 and 2025, the ability of local models to handle queries increased from 23.2% to 71.3% relative to advanced models, representing an improvement of about 3.1 times.

The gains were linked to improvements in model architecture, training and post-training processes, as well as techniques such as Mixture of Experts and more efficient parameter utilization. These developments have enabled greater capabilities from models that can run outside large-scale data centers.

Hardware delivers an even larger efficiency boost

Model improvements were not the only factor behind the gains.

The study found that advances in AI accelerators contributed about 5.9 times to the improvement in Intelligence per Joule, compared with roughly 3.1 times from model improvements.

Hardware affects both major components of the metric: the amount of energy consumed and the time required to complete a task. Faster inference can therefore improve not only response times but also overall energy efficiency.

An 18-fold gain in 16 months

Overall, Intelligence per Joule increased by about 18 times over 16 months, according to the data cited in the study.

The improvement resulted from the combined progress of models and accelerators. When their contributions were examined separately, model advances accounted for about 3.1 times the improvement, while accelerator advances contributed 5.9 times.

The figures indicate that improving AI efficiency is not simply a matter of building larger models. It depends on optimizing the broader system that runs them.

Local AI is expanding its capabilities

Local models are increasingly capable of handling tasks that previously required advanced cloud-based models.

Researchers found that local models successfully handled 88.7% of individual conversational and reasoning queries in the dataset used for the study, although performance varied across fields.

Local models performed better on some types of creative tasks, while coverage was lower in specialized technical areas.

The findings suggest that on-device AI is no longer limited to basic applications, as the range of tasks that can be handled locally continues to expand.

The cloud still maintains an efficiency advantage in some cases

Despite the rapid gains, the study does not suggest that local devices have become more efficient than cloud infrastructure in every scenario.

When running the same model, data-center accelerators such as NVIDIA B200 achieved higher efficiency than some local accelerators, with the gap becoming more pronounced when energy consumption and response time were considered together.

Some cloud accelerators recorded an advantage of between 1.6 and 2.3 times in Intelligence per Joule compared with the Apple M4 Max in the tests cited by the study.

The findings suggest that the future of AI inference may not involve moving all workloads to local devices. Instead, workloads could be directed to whichever environment is best suited to handle them efficiently.

A hybrid approach between local devices and the cloud

The study proposes a hybrid model that combines local devices with cloud data centers, directing each query to the environment most appropriate for the task.

Simulation results indicated that intelligent routing between local and cloud models could significantly reduce energy consumption, computing requirements and costs compared with relying exclusively on the cloud, while maintaining an appropriate level of answer quality.

As local Intelligence per Joule improves, a larger share of workloads can be handled on personal devices, while more demanding tasks can continue to be processed in data centers with greater computing resources.

Energy efficiency becomes a competitive factor

The findings put energy efficiency at the center of the next stage of competition between AI model developers and chipmakers.

The key question is no longer simply which model is the most capable, but also how much energy and computing power are required to deliver that capability.

This is likely to drive the development of more efficient models, specialized accelerators, compression and quantization techniques, as well as better software for running AI systems.

The study points to Quantization as one potential avenue for reducing energy consumption. Tests showed that moving some models from FP16 to FP4 reduced energy consumption by roughly 3 to 3.5 times, with a relatively limited impact on accuracy in the tests conducted.

Toward AI that runs directly on devices

The rise in Intelligence per Joule reflects a broader move toward making advanced AI available across a wider range of devices, from personal computers to smartphones and wearables.

As models and accelerators become more efficient, running advanced AI capabilities locally becomes increasingly practical, reducing the need to send every query to a remote data center.

The shift could also reduce pressure on cloud infrastructure as demand for inference continues to grow, while allowing data centers to focus on workloads that require greater computing resources.

The 18-fold figure does not mean AI became 18 times smarter

The reported 18-fold increase in Intelligence per Joule does not mean that AI models became directly 18 times “smarter.”

Instead, the figure represents an increase in the amount of measured performance that can be obtained per unit of energy under the conditions of the study.

The metric combines improvements in models, hardware efficiency and response time. Continued progress across these areas could make increasingly capable AI models practical for a growing number of applications on personal devices, while cloud infrastructure continues to handle workloads requiring greater resources.