Back to Research

United States

Google Puts a Number on a Gemini Query: 0.24 Watt-Hours

For two years, estimates of how much electricity a single AI query uses have ranged widely, because the companies running the models would not say. On August 21 Google became the first major AI developer to publish its own figure. In a technical paper and accompanying blog post, it estimated that the median Gemini Apps text request uses 0.24 watt-hours of energy, emits 0.03 grams of carbon dioxide equivalent, and consumes 0.26 milliliters of water, about five drops. Google compares the energy to watching television for less than nine seconds.

The headline number is small. The more useful parts of the disclosure are the method behind it, the rate at which it has fallen, and what it still does not tell us.

What Google measured

Google says many published estimates count only the energy used by AI chips during active computation, which gives a theoretical rather than operational picture. Its method includes four additional elements: the actual utilization achieved by chips at production scale, which can be well below theoretical maximums; idle machines kept ready for traffic spikes or failover; the host CPU and memory that support the accelerators; and data center overhead such as cooling and power conversion, captured by power usage effectiveness. Water use is included as well.

The difference is substantial. Using only active accelerator consumption, Google estimates the median Gemini text request at 0.10 Wh, 0.02 grams of CO2 equivalent and 0.12 milliliters of water. Its fuller method more than doubles the energy figure.

MIT Technology Review, which interviewed Google chief scientist Jeff Dean about the report, gave the breakdown. Google's custom TPU chips account for 58% of the 0.24 Wh. The host machine's CPU and memory account for 25%. Idle backup machines take 10%, and data center overhead the remaining 8%. Google says its fleet operates at an average power usage effectiveness of 1.09, meaning overhead adds about 9% to the energy drawn by IT equipment.

Outside researchers welcomed the level of detail. Jae-Won Chung, a PhD candidate at the University of Michigan who helps run the ML.Energy leaderboard, told MIT Technology Review it was the most comprehensive analysis so far. Estimates of this kind are generally something only operators can produce, because they run at a scale researchers cannot replicate and can see the full stack of hardware in production.

A 33-fold drop

The paper reports that over a recent 12-month period, the energy of the median Gemini text request fell by a factor of 33 and its total carbon footprint by a factor of 44, while response quality improved. MIT Technology Review reported the comparison as May 2024 against May 2025.

Google attributes the gains to several sources. Mixture-of-experts models activate only part of a large model for each request, which Google says reduces computation and data transfer by a factor of 10 to 100. Speculative decoding lets a small model draft responses that a larger one verifies. Distillation produces smaller serving models such as Gemini Flash and Flash-Lite. On hardware, Google says its latest TPU, Ironwood, is 30 times more energy efficient than its first publicly available TPU, and its serving stack moves models dynamically to reduce idle chips.

How the carbon number is built

The emissions figure deserves a closer look. Google's footnote says emissions per request were estimated by applying its 2024 average fleetwide grid carbon intensity to energy per request. MIT Technology Review explained that this is a market-based figure, which accounts for Google's clean energy purchases. The company has signed agreements for more than 22 GW of power from solar, wind, geothermal and advanced nuclear projects since 2010, and on that basis its emissions per unit of electricity are roughly one-third of the average on the grids where it operates.

That is a legitimate accounting method, but it means the 0.03 gram figure reflects Google's procurement as much as the physics of the query. The same request served from a data center on an average grid would carry about three times the emissions. Water use was estimated the same way, using Google's 2024 fleetwide average water usage effectiveness.

What is missing

The disclosure answers the per-request question but not the system question. Hugging Face researcher Sasha Luccioni told MIT Technology Review that the major missing figure is the total number of Gemini queries each day, which would allow an estimate of total energy demand. Google did not provide it.

Without that number, the per-query figure can be read in very different ways. At one billion text requests a day, 0.24 Wh each would add up to 240 MWh a day, or about 88 GWh a year, equivalent to a constant load of about 10 MW. That would be a modest share of a single large data center. But the figure covers only median text requests. MIT Technology Review notes that image and video generation can require much more energy, and the median does not reflect long or complex requests at the top of the range. Nor does it include the energy used to train the models, which Google did not report in this disclosure.

Luccioni also said the report is not a substitute for a standardized AI energy score comparable to the Energy Star rating for appliances. As long as each company chooses its own method and timing, comparisons across models will be hard.

Why it matters for power planning

The disclosure lands as utilities and grid operators build forecasts around AI demand. For them, per-query efficiency cuts two ways. A 33-fold drop in energy per request in a year is evidence that the cost of serving AI can fall faster than many forecasts assume. But cheaper responses tend to encourage more use, and Google itself says AI demand is growing and it is investing heavily to reduce the power and water required per request.

The arithmetic of data center load depends on the product of two numbers: energy per unit of work and the amount of work. Google has now published the first. The power sector will be watching for the second, and for whether efficiency gains keep pace with adoption.

Sources

  • Google Cloud, Measuring the environmental impact of AI inference, August 21, 2025 cloud.google.com
  • Google, Measuring the environmental impact of delivering AI at Google Scale, technical paper, August 2025 arxiv.org
  • MIT Technology Review, article on Google's first release of Gemini energy data, August 21, 2025 technologyreview.com

Newsletter

Get The TEI Briefing

A weekly read on energy markets, policy and our latest research. Free, and you can unsubscribe at any time.