Skip to main navigation Skip to search Skip to main content

Challenges in Designing Robust RL-Based Autoscalers

  • Technical University of Munich

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Reinforcement learning (RL) offers a promising, adaptive alternative to heuristic-based autoscaling, yet its practical adoption in production environments remains negligible. In this paper, we argue that this gap between promise and practice is caused by three systemic challenges that violate fundamental RL assumptions: (i) generalization failures under workload and system drift; (ii) orchestration interference that obscures causality; and (iii) unreliable, delayed metric feedback. We substantiate these claims through an empirical study of two PPO-based autoscalers on real-world and synthetic workloads, demonstrating how these factors lead to policy instability and performance degradation. Our findings reveal that these challenges collectively frame autoscaling as a Partially Observable Markov Decision Process. We conclude that robust RL-based autoscaling requires a paradigm shift from purely algorithmic solutions toward systems-aware designs that model the partial observability and non-stationarity inherent in service autoscaling.

Original languageEnglish
Title of host publicationPACMI 2025 - Proceedings of the 4th Workshop on Practical Adoption Challenges of ML for Systems
PublisherAssociation for Computing Machinery, Inc
Pages44-49
Number of pages6
ISBN (Electronic)9798400722059
DOIs
StatePublished - 13 Oct 2025
Event4th Workshop on Practical Adoption Challenges of ML for Systems, PACMI 2025 - Seoul, Korea, Republic of
Duration: 13 Oct 202516 Oct 2025

Publication series

NamePACMI 2025 - Proceedings of the 4th Workshop on Practical Adoption Challenges of ML for Systems

Conference

Conference4th Workshop on Practical Adoption Challenges of ML for Systems, PACMI 2025
Country/TerritoryKorea, Republic of
CitySeoul
Period13/10/2516/10/25

Keywords

  • Autoscaling
  • Delayed Feedback
  • Distribution Drift
  • Infrastructure Drift
  • Kubernetes
  • Non-Stationary Environments
  • Partial Observability
  • Reinforcement Learning

Fingerprint

Dive into the research topics of 'Challenges in Designing Robust RL-Based Autoscalers'. Together they form a unique fingerprint.

Cite this