Runtime Causal Fingerprinting for Root-Cause-Aware Autoscaling and Adaptive Resource Governance in Kubernetes Service Meshes
DOI:
https://doi.org/10.63282/3050-9416.IJAIBDCMS-V6I3P116Keywords:
Kubernetes, Service Mesh, Autoscaling, Causal Inference, Root-Cause Analysis, Distributed Tracing, Microservices, Adaptive Resource Governance, SLO-Aware Resource Management, Cloud-Native SystemsAbstract
Kubernetes-orchestrated microservice systems increasingly operate under volatile workload patterns, strict latency service-level objectives, and complex inter-service dependencies introduced by service mesh communication layers. Although Kubernetes provides established autoscaling mechanisms, conventional reactive autoscaling policies frequently interpret symptoms such as CPU saturation, request latency, or queue growth as direct scaling triggers without distinguishing whether the observed degradation is caused by workload demand, downstream dependency propagation, resource contention, noisy-neighbor effects, configuration drift, or release-induced regression. This limitation often leads to unnecessary scaling, delayed remediation, over-provisioning, and repeated service-level objective violations. This paper proposes a runtime causal fingerprinting framework for root-cause-aware autoscaling and adaptive resource governance in Kubernetes service meshes. The framework constructs continuously updated causal fingerprints from distributed traces, service mesh telemetry, container metrics, event streams, and workload signals. These fingerprints encode not only anomaly signatures but also causal propagation patterns across services, pods, nodes, and resource dimensions. The proposed approach integrates causal graph inference, temporal fingerprint matching, SLO-risk estimation, and policy-constrained scaling actions within a closed-loop governance architecture. Unlike conventional autoscalers that scale based primarily on local resource utilization, the proposed framework differentiates demand-driven saturation from causally propagated degradation and selects remediation strategies accordingly. The paper develops the conceptual model, methodological pipeline, evaluation criteria, and analytical discussion necessary for implementing root-cause-aware autoscaling in production-grade Kubernetes environments. The study contributes a technically grounded framework for improving autoscaling precision, reducing false-positive scaling, mitigating cascading failures, and strengthening operational trust in cloud-native platforms.
References
1. M. Lorido-Botran, J. Miguel-Alonso, and J. A. Lozano, “A review of auto-scaling techniques for elastic applications in cloud environments,” Journal of Grid Computing, vol. 12, no. 4, pp. 559–592, 2014, doi: 10.1007/s10723-014-9314-7.
2. S. K. Gunda, “A Deep Dive into Software Fault Prediction: Evaluating CNN and RNN Models,” in 2024 International Conference on Electronic Systems and Intelligent Computing (ICESIC), Chennai, India, 2024, pp. 224–228, doi: 10.1109/ICESIC61777.2024.10846549.
3. H. Qiu, S. S. Banerjee, S. Jha, Z. T. Kalbarczyk, and R. K. Iyer, “FIRM: An intelligent fine-grained resource management framework for SLO-oriented microservices,” in Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI ’20), 2020, pp. 805–825.
4. P. Jamshidi, C. Pahl, N. C. Mendonça, J. Lewis, and S. Tilkov, “Microservices: The journey so far and challenges ahead,” IEEE Software, vol. 35, no. 3, pp. 24–35, 2018, doi: 10.1109/MS.2018.2141039.
5. Spatharakis, D., Dimolitsas, I., Vlahakis, E., Dechouniotis, D., Athanasopoulos, N., & Papavassiliou, S. (2022). Distributed resource autoscaling in Kubernetes edge clusters. In Proceedings of the 18th International Conference on Network and Service Management (CNSM 2022) (pp. 163–169). IFIP. https://doi.org/10.23919/CNSM55787.2022.996505
6. G. Yu, P. Chen, H. Chen, Z. Guan, Z. Huang, L. Jing, T. Weng, X. Sun, and X. Li, “MicroRank: End-to-end latency issue localization with extended spectrum analysis in microservice environments,” in Proceedings of The Web Conference 2021 (WWW ’21), Ljubljana, Slovenia, 2021, pp. 3087–3098, doi: 10.1145/3442381.3449905.
7. A. Verma, L. Pedrosa, M. Korupolu, D. Oppenheimer, E. Tune, and J. Wilkes, “Large-scale cluster management at Google with Borg,” in Proceedings of the Tenth European Conference on Computer Systems (EuroSys ’15), Bordeaux, France, 2015, pp. 1–17, doi: 10.1145/2741948.2741964.
8. S. K. Gunda, “Machine Learning Approaches for Software Fault Diagnosis: Evaluating Decision Tree and KNN Models,” in 2024 Global Conference on Communications and Information Technologies (GCCIT), Bangalore, India, 2024, pp. 1–5, doi: 10.1109/GCCIT63234.2024.10861953.
9. C. Meng, S. Song, H. Tong, M. Pan, and Y. Yu, “DeepScaler: Holistic autoscaling for microservices based on spatiotemporal GNN with adaptive graph learning,” in Proceedings of the 38th IEEE/ACM International Conference on Automated Software Engineering (ASE 2023), Luxembourg, 2023, pp. 569–580, doi: 10.1109/ASE56229.2023.00038.
10. J. Lin, P. Chen, and Z. Zheng, “Microscope: Pinpoint performance issues with causal graphs in micro-service environments,” in Service-Oriented Computing: 16th International Conference, ICSOC 2018, Hangzhou, China, 2018, pp. 3–20, doi: 10.1007/978-3-030-03596-9_1.
11. J. Soldani and A. Brogi, “Anomaly detection and failure root cause analysis in (micro)service-based cloud applications: A survey,” ACM Computing Surveys, vol. 55, no. 3, Article 59, pp. 1–39, 2022, doi: 10.1145/3501297.
12. S. K. Gunda, “Software Defect Prediction Using Advanced Ensemble Techniques: A Focus on Boosting and Voting Method,” in 2024 International Conference on Electronic Systems and Intelligent Computing (ICESIC), Chennai, India, 2024, pp. 157–161, doi: 10.1109/ICESIC61777.2024.10846550.
13. Y. Gan, Y. Zhang, D. Cheng, A. Shetty, P. Rathi, N. Katarki, A. Bruno, J. Hu, B. Ritchken, K. Hu, M. Pancholi, Y. He, B. Clancy, C. Colen, F. Wen, C. Leung, S. Wang, L. Zaruvinsky, M. Espinosa, R. Lin, Z. Liu, J. Padilla, and C. Delimitrou, “An open-source benchmark suite for microservices and their hardware-software implications for cloud and edge systems,” in Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS ’19), Providence, RI, USA, 2019, pp. 3–18, doi: 10.1145/3297858.3304013.
14. M. Li, Z. Li, K. Yin, X. Nie, W. Zhang, K. Sui, and D. Pei, “Causal inference-based root cause analysis for online service systems with intervention recognition,” in Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD ’22), Washington, DC, USA, 2022, pp. 3230–3240, doi: 10.1145/3534678.3539041.
15. C. Qu, R. N. Calheiros, and R. Buyya, “Auto-scaling web applications in clouds: A taxonomy and survey,” ACM Computing Surveys, vol. 51, no. 4, Article 73, pp. 1–33, 2018, doi: 10.1145/3148149.
16. Y. Gan, M. Liang, S. Dev, D. Lo, and C. Delimitrou, “Sage: Practical and scalable ML-driven performance debugging in microservices,” in Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS ’21), Virtual Event, 2021, pp. 135–151, doi: 10.1145/3445814.3446700.
17. P. Wang, J. Xu, M. Ma, W. Lin, D. Pan, Y. Wang, and P. Chen, “CloudRanger: Root cause identification for cloud native systems,” in Proceedings of the 18th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGrid ’18), Washington, DC, USA, 2018, pp. 492–502, doi: 10.1109/CCGRID.2018.00076.
18. J. Park, B. Choi, C. Lee, and D. Han, “Graph neural network-based SLO-aware proactive resource autoscaling framework for microservices,” IEEE/ACM Transactions on Networking, vol. 32, no. 4, pp. 3331–3346, 2024, doi: 10.1109/TNET.2024.3393427.
19. L. Wu, J. Tordsson, E. Elmroth, and O. Kao, “MicroRCA: Root cause localization of performance issues in microservices,” in NOMS 2020—2020 IEEE/IFIP Network Operations and Management Symposium, Budapest, Hungary, 2020, pp. 1–9, doi: 10.1109/NOMS47738.2020.9110353.
20. H. Ahmad, C. Treude, M. Wagner, and C. Szabo, “Smart HPA: A resource-efficient horizontal pod auto-scaler for microservice architectures,” in Proceedings of the IEEE 21st International Conference on Software Architecture (ICSA 2024), Hyderabad, India, 2024, pp. 46–57, doi: 10.1109/ICSA59870.2024.00013.