
Proactive resource scaling in container orchestration platforms like Kubernetes is essential to maintain application responsiveness under fluctuating workloads. Our study aims to explore predictive auto-scaling by forecasting CPU utilization from service request volumes using statistical and neural-networkbased time series models. We applied 29 forecasting techniques to production traces from a FinTech system, evaluating each model's accuracy using nine metrics: MAE, RMSE, SMAPE, MASE, RMSSE, R², NMAE, and NRMSE, and assessing the actual prediction distance. Experimental results show that neural network models, particularly Transformer and GRU, consistently exhibited strong accuracy performance with high explanatory power (R² values above 0.88) and relatively low mean absolute errors (MAE). However, several statistical models, specifically AutoTheta, FFT, and Exponential Smoothing, achieved higher accuracy than neural approaches, and in particular Exponential Smoothing, was the best performing model scoring the lowest MAE (332.824) and highest R² (0.961). Our findings demonstrate the viability of lightweight forecasting-driven scaling and suggest practical improvements for auto-scaling reliability in Kubernetes environments. The source code we used in this study is available as open-source on github.
Authors: Jonathan Wisborg Fog, Jens Jacob Torvin Møller, Thomas Møller Jensen, Davide Taibi, Michele Albano
Contributing partner: Aalborg University