TY - GEN
T1 - Challenges in Designing Robust RL-Based Autoscalers
AU - Asadi, Navidreza
AU - Ali, Dalal
AU - Ursu, Razvan Mihai
AU - Kellerer, Wolfgang
N1 - Publisher Copyright:
© 2025 Copyright held by the owner/author(s)
PY - 2025/10/13
Y1 - 2025/10/13
N2 - Reinforcement learning (RL) offers a promising, adaptive alternative to heuristic-based autoscaling, yet its practical adoption in production environments remains negligible. In this paper, we argue that this gap between promise and practice is caused by three systemic challenges that violate fundamental RL assumptions: (i) generalization failures under workload and system drift; (ii) orchestration interference that obscures causality; and (iii) unreliable, delayed metric feedback. We substantiate these claims through an empirical study of two PPO-based autoscalers on real-world and synthetic workloads, demonstrating how these factors lead to policy instability and performance degradation. Our findings reveal that these challenges collectively frame autoscaling as a Partially Observable Markov Decision Process. We conclude that robust RL-based autoscaling requires a paradigm shift from purely algorithmic solutions toward systems-aware designs that model the partial observability and non-stationarity inherent in service autoscaling.
AB - Reinforcement learning (RL) offers a promising, adaptive alternative to heuristic-based autoscaling, yet its practical adoption in production environments remains negligible. In this paper, we argue that this gap between promise and practice is caused by three systemic challenges that violate fundamental RL assumptions: (i) generalization failures under workload and system drift; (ii) orchestration interference that obscures causality; and (iii) unreliable, delayed metric feedback. We substantiate these claims through an empirical study of two PPO-based autoscalers on real-world and synthetic workloads, demonstrating how these factors lead to policy instability and performance degradation. Our findings reveal that these challenges collectively frame autoscaling as a Partially Observable Markov Decision Process. We conclude that robust RL-based autoscaling requires a paradigm shift from purely algorithmic solutions toward systems-aware designs that model the partial observability and non-stationarity inherent in service autoscaling.
KW - Autoscaling
KW - Delayed Feedback
KW - Distribution Drift
KW - Infrastructure Drift
KW - Kubernetes
KW - Non-Stationary Environments
KW - Partial Observability
KW - Reinforcement Learning
UR - https://www.scopus.com/pages/publications/105020837738
U2 - 10.1145/3766882.3767176
DO - 10.1145/3766882.3767176
M3 - Conference contribution
AN - SCOPUS:105020837738
T3 - PACMI 2025 - Proceedings of the 4th Workshop on Practical Adoption Challenges of ML for Systems
SP - 44
EP - 49
BT - PACMI 2025 - Proceedings of the 4th Workshop on Practical Adoption Challenges of ML for Systems
PB - Association for Computing Machinery, Inc
T2 - 4th Workshop on Practical Adoption Challenges of ML for Systems, PACMI 2025
Y2 - 13 October 2025 through 16 October 2025
ER -