Lightweight Reinforcement Learning for Cost and Energy Aware Workload Optimization in Geo-Distributed Data Centers

Authors

  • Rafia Abdul Sattar Department of Computer Science, University of Engineering & Technology Peshawar 25000, Pakistan
  • Muhammad Imran Khan Khalil Department of Computer Science, University of Engineering & Technology Peshawar 25000, Pakistan
  • Arshia Noor Qazi Department of Computer Science, University of Engineering & Technology Peshawar 25000, Pakistan
  • Mujtaba Hasan Department of Computer Science, University of Engineering & Technology Peshawar 25000, Pakistan
  • Amer Taj Department of Computer Science, University of Engineering & Technology Peshawar 25000, Pakistan
  • Imran Rasheed Department of Computer Science, University of Engineering & Technology Peshawar 25000, Pakistan
  • Alauddin Department of Computer Science, University of Engineering & Technology Peshawar 25000, Pakistan

Keywords:

Energy Efficiency, Resource Optimization, Geo-Distributed Data Centers, Geographical Load Balancing (GLB), Q-Learning, Double Q-Learning, Reinforcement Learning, Multi-Objective Optimization

Abstract

The rapid expansion of the hyperscale computing infrastruc-ture has pushed the data center electricity consumption to 1–2% of global demand, a trajectory projected to exceed 1,065 TWh by 2030. The current geographical load balancing (GLB) strategies depend on static, single-metric rules that cannot simultaneously manage the competing pressures of reducing power consumption and cutting electricity costs across markets where both vary hour by hour. This paper proposes an enhanced tabular Q-Learning framework for multi-objective work-load allocation across the geographically distributed data centers. The standard Q-Learning is enhanced with the help of two stabilization mechanisms which is Double Q-Learning, that removes the maximization bias from the value estimation, and Experience Replay, which separates the sequential training samples. Where as an adaptive epsilon-decay schedule progressively transitions the agent from broad exploration to refined exploitation over the training period. A weighted reward function jointly optimizes the normalized power consumption (α = 0.6) and the electricity cost (β = 0.4). This entire framework requires only ∼0.08 kWh of training energy on commodity CPU hardware with no GPU requirement which is over 3,500× less than the GPU-based deep RL systems. The training and evaluation uses real Wikipedia traffic traces (1,464 hourly time steps) paired with market-rate electricity price records from three data centers spanning three countries. Tested against nine benchmark algorithms across three tiers which includes naïve stateless baselines, single-metric heuristics, and multi-metric strategies and the proposed agent outperforms every naïve and single-metric competitor by more than 400 cumulative reward points. The learned policy allocates workloads across data centers in direct proportion to the α/β weighting, providing measurable evidence of genuine multi-objective optimization. These results establish a rigorous, reproducible reference point for lightweight reinforcement learning in data center optimization. This identifies the operational conditions under which adaptive policies hold a practical advantage over cost-focused heuristics.

Downloads

Published

2026-01-25

How to Cite

Rafia Abdul Sattar, Muhammad Imran Khan Khalil, Arshia Noor Qazi, Mujtaba Hasan, Amer Taj, Imran Rasheed, & Alauddin. (2026). Lightweight Reinforcement Learning for Cost and Energy Aware Workload Optimization in Geo-Distributed Data Centers. Spectrum of Engineering Sciences, 4(1), 1593–1606. Retrieved from https://thesesjournal.com.medicalsciencereview.com/index.php/1/article/view/3941