Software can reduce the energy used for a computing task or shift flexible work to a different time or place. Those are distinct levers: one targets energy per task, while the other changes when or where electricity is used. Neither guarantees that a data center’s total consumption will fall.
What software can—and cannot—change
Software can influence how much computing a task requires. Teams can choose a smaller model for routine requests, cache repeated prompts, improve batching and compilers, limit prompt and output tokens, or right-size cloud instances to the work at hand. These approaches aim to avoid unnecessary computation; their results depend on the task and the system running it.
Scheduling works on a different part of the problem. A flexible batch job can be delayed until a later time, or moved to a region with spare capacity and lower-carbon electricity. That changes the timing or location of demand, rather than necessarily reducing the total amount of energy the job needs.
Two efficiency results, two different measures
An ACM paper reports that Perseus reduced energy use by up to 30% in its evaluation of large-model training, without throughput loss or hardware modification. Perseus schedules computation over time to address energy bloat—energy spent beyond what is needed to complete the training work. The result applies to the paper’s evaluation, not every model or data center.
NVIDIA describes a separate approach: power profiles for Blackwell GPUs, released with the Blackwell B200 and supporting AI training, AI inference, and high-performance computing (HPC). The profiles configure GPU power and performance for different workloads.
NVIDIA says its phase-one profiles can save up to 15% energy while keeping performance above 97% for critical applications. In a separate, power-constrained facility scenario, NVIDIA says using the savings to provision more GPUs can raise overall throughput by up to 13%. Energy saved on a GPU and work completed across a facility are different measures; neither figure is a universal outcome.
Scheduling flexible work around the grid
Day-ahead planning and real-time scheduling can help data centers respond to grid conditions by moving flexible batch work to a more favorable time or a region with available capacity. If that region also has lower-carbon electricity, shifting the workload can change its power mix as well.
Moving work across regions has practical limits. Data-sovereignty rules may restrict where information can be processed, and large datasets take effort to transfer. Scheduling can help manage a local capacity crunch, but it does not make those constraints disappear.
Why efficiency is only part of the answer
The International Energy Agency estimates that servers account for around 60% of electricity demand in modern data centers, though the share varies by facility type. Software that reduces unnecessary computing can therefore address a significant part of a facility’s demand. It remains one part of a broader response alongside more efficient hardware and electricity infrastructure.
There is also a rebound risk. The Jevons paradox describes how efficiency gains can make a resource-intensive activity cheaper, encouraging more use. In AI, lower energy per task could make it economical to generate more tokens or run more jobs. A more efficient workload does not, by itself, establish lower total electricity consumption.
What the IEA’s 2030 number covers
The IEA estimates that data centers used about 415 TWh of electricity worldwide in 2024, roughly 1.5% of global electricity consumption. Its Base Case projects about 945 TWh of total data-center electricity use in 2030. That forecast covers all data centers; the agency discusses electricity demand from accelerated servers separately, with AI adoption a main driver of its growth.
For more context on the wider infrastructure pressures behind AI, see NeoTeo’s coverage of AI infrastructure constraints.