Topology-Aware Autoscaling of Stateful Microservices Using Messaging Lag and Dependency Criticality

Authors

  • Srinivas Nune Independent Researcher Author

DOI:

https://doi.org/10.71238/snnst.v1i01.180

Abstract

As stateful microservice architectures have spread, they have started to rely on asynchronous messaging middleware to get their distributed state right, but traditional autoscaling controllers still rely on coarse metrics like central-processing-unit occupancy or memory usage. These metrics don't align well with the actual message driven, stateful operator saturation and result in less-than-ideal over-provisioning when the load spike and under-provisioning when the backlog is building. This paper brings together the results of container-orchestration surveys, proactive and coordinated autoscaling, service-dependency-graph analysis, and stateful stream-processing research to suggest a topology-aware autoscaling model consisting of two complementary signals: message-queue consumption lag and dependency-criticality score, which can be based on the service call graph. The proposed framework collects the per-service lag measurements, assigns them weights, based on the criticality coefficient calculated from fan-in, fan-out and latency-sensitivity attributes of the nodes in the dependency graph, and reaches scaling decisions with a decision engine that is connected to a state migration planner for stateful operators. Based on the analysis of the comparative synthesis across benchmark suites applied to microservice performance debugging, criticality-aware and lag-driven scaling can mitigate tail-latency degradation and scaling-action churn compared to scaling based on thresholds, and be as efficient as hybrid proactive scaling approaches on resource-efficiency metrics. The evaluated evaluation techniques are synthesized to show efficiency gains of around 25 to 30 percent with respect to effective central-processing-unit utilization, and reduce the rate of scaling-action oscillation by more than fifty percent. Operators’ considerations regarding cluster operators, compliance to SLOs, future plans for integration with model-driven autoscalers and machine-learning based autoscalers are also covered, as are limitations of checkpoint-consistency overhead when migrating state.

Downloads

Download data is not yet available.

References

D. Balla, C. Simon, and M. Maliosz, “Adaptive Scaling of Kubernetes Pods,” in NOMS 2020 - 2020 IEEE/IFIP Network Operations and Management Symposium, 2020. doi: 10.1109/NOMS47738.2020.9110428.

A. Bauer, N. Herbst, S. Spinner, A. Ali-Eldin, and S. Kounev, “Chameleon: A Hybrid, Proactive Auto-Scaling Mechanism on a Level-Playing Field,” IEEE Trans. Parallel Distrib. Syst., 2019, doi: 10.1109/TPDS.2018.2870389.

A. Bauer, V. Lesch, L. Versluis, A. Ilyushkin, N. Herbst, and S. Kounev, “Chamulteon: Coordinated Auto-Scaling of Micro-Services,” in 2019 IEEE 39th International Conference on Distributed Computing Systems (ICDCS), 2019. doi: 10.1109/ICDCS.2019.00199.

P. Carbone, S. Ewen, G. Fóra, S. Haridi, S. Richter, and K. Tzoumas, “State Management in Apache Flink: Consistent Stateful Distributed Stream Processing,” Proc. VLDB Endow., 2017, doi: 10.14778/3137765.3137777.

E. Casalicchio, “Container Orchestration: A Survey,” in Systems Modeling: Methodologies and Tools, 2019. doi: 10.1007/978-3-319-92378-9_14.

E. Casalicchio, “A Study on Performance Measures for Auto-Scaling CPU-Intensive Containerized Applications,” Cluster Comput., 2019, doi: 10.1007/s10586-018-02890-1.

E. Casalicchio and V. Perciballi, “Auto-Scaling of Containers: The Impact of Relative and Absolute Metrics,” in 2017 IEEE 2nd International Workshops on Foundations and Applications of Self* Systems (FAS*W), 2017. doi: 10.1109/FAS-W.2017.149.

Y. Gan et al., “An Open-Source Benchmark Suite for Microservices and Their Hardware-Software Implications for Cloud & Edge Systems,” in Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems, 2019. doi: 10.1145/3297858.3304013.

Y. Gan et al., “Seer: Leveraging Big Data to Navigate the Complexity of Performance Debugging in Cloud Microservices,” in Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems, 2019. doi: 10.1145/3297858.3304004.

A. U. Gias, G. Casale, and M. Woodside, “ATOM: Model-Driven Autoscaling for Microservices,” in 2019 IEEE 39th International Conference on Distributed Computing Systems (ICDCS), 2019. doi: 10.1109/ICDCS.2019.00197.

M. Imdoukh, I. Ahmad, and M. G. Alfailakawi, “Machine Learning-Based Auto-Scaling for Containerized Applications,” Neural Comput. Appl., 2020, doi: 10.1007/s00521-019-04507-z.

S.-P. Ma, C.-Y. Fan, Y. Chuang, W.-T. Lee, S.-J. Lee, and N.-L. Hsueh, “Using Service Dependency Graph to Analyze and Test Microservices,” in 2018 IEEE 42nd Annual Computer Software and Applications Conference (COMPSAC), 2018. doi: 10.1109/COMPSAC.2018.10207.

M. Mudassar, Y. Zhai, and L. Liao, “Efficient State Management for Scaling Out Stateful Operators in Stream Processing Systems,” Big Data, 2019, doi: 10.1089/big.2018.0093.

T.-T. Nguyen, Y.-J. Yeom, T. Kim, D.-H. Park, and S. Kim, “Horizontal Pod Autoscaling in Kubernetes for Elastic Container Orchestration,” Sensors, 2020, doi: 10.3390/s20164621.

C. Pahl, A. Brogi, J. Soldani, and P. Jamshidi, “Cloud Container Technologies: A State-of-the-Art Review,” IEEE Trans. Cloud Comput., 2019, doi: 10.1109/TCC.2017.2702586.

C. Qu, R. N. Calheiros, and R. Buyya, “Auto-Scaling Web Applications in Clouds: A Taxonomy and Survey,” ACM Comput. Surv., 2018, doi: 10.1145/3148149.

K. Rzadca et al., “Autopilot: Workload Autoscaling at Google,” in Proceedings of the Fifteenth European Conference on Computer Systems (EuroSys 20), 2020. doi: 10.1145/3342195.3387524.

Downloads

Published

2024-01-31

How to Cite

Topology-Aware Autoscaling of Stateful Microservices Using Messaging Lag and Dependency Criticality. (2024). Sciences Du Nord Nature Science and Technology, 1(01), 37-50. https://doi.org/10.71238/snnst.v1i01.180