TY - JOUR KW - Affordable and Clean Energy KW - AI training KW - Artificial Intelligence KW - Benchmark testing KW - Bioengineering KW - Computational modeling KW - Engineering KW - GPU power measurements KW - Graphics processing units KW - Hardware KW - Information and Computing Sciences KW - Large Language Models KW - Networking and Information Technology R&D KW - Power demand KW - Power measurement KW - Stress KW - sustainable computing KW - Technology KW - Training KW - Visualization AU - Imran Latif AU - Alex C Newkirk AU - Matthew R Carbone AU - Arslan Munir AU - Yuewei Lin AU - Jonathan G Koomey AU - Xi Yu AU - Zhihua Dong AB -
The expansion of artificial intelligence (AI) applications has driven substantial investment in computational infrastructure, especially by cloud computing providers. Quantifying the energy footprint of this infrastructure requires models parameterized by the power demand of AI hardware during training. In this work, we measured the instantaneous power draw of an 8-GPU NVIDIA H100 HGX node during the training of open-source image classifier (ResNet) and large-language models (Llama2-13b). We characterize power demand for a single node configuration, providing foundational data for future multi-node studies. The maximum observed power draw was approximately 8.4 kW, 18% lower than the manufacturer-rated 10.2 kW, even with GPUs near full utilization. Holding model architecture constant, increasing batch size from 512 to 4096 images for ResNet reduced total training energy consumption by a factor of 4. These findings can inform capacity planning for data center operators and energy use estimates by researchers. Future work will investigate the impact of cooling technology and carbon-aware scheduling on AI workload energy consumption.
BT - IEEE Access DA - 26/03/2025 DO - 10.1109/access.2025.3554728 N2 -The expansion of artificial intelligence (AI) applications has driven substantial investment in computational infrastructure, especially by cloud computing providers. Quantifying the energy footprint of this infrastructure requires models parameterized by the power demand of AI hardware during training. In this work, we measured the instantaneous power draw of an 8-GPU NVIDIA H100 HGX node during the training of open-source image classifier (ResNet) and large-language models (Llama2-13b). We characterize power demand for a single node configuration, providing foundational data for future multi-node studies. The maximum observed power draw was approximately 8.4 kW, 18% lower than the manufacturer-rated 10.2 kW, even with GPUs near full utilization. Holding model architecture constant, increasing batch size from 512 to 4096 images for ResNet reduced total training energy consumption by a factor of 4. These findings can inform capacity planning for data center operators and energy use estimates by researchers. Future work will investigate the impact of cooling technology and carbon-aware scheduling on AI workload energy consumption.
PB - Institute of Electrical and Electronics Engineers (IEEE) PY - 2025 SP - 61740 EP - 61747 T2 - IEEE Access TI - Single-Node Power Demand During AI Training: Measurements on an 8-GPU NVIDIA H100 System UR - https://doi.org/10.1109/access.2025.3554728 VL - 13 SN - 2169-3536 ER -