Retrieval-Augmented Operational Intelligence for Self-Healing Cloud-Native Microservices Using Traces, Logs, and Events

Authors

  • Pranayachandanreddy Gottimukkala Independent Researcher Author
  • Srinivas Nune Independent Researcher Author

DOI:

https://doi.org/10.71238/snnst.v1i02.183

Abstract

The cloud-native microservice architecture approach breaks application functionality into many small, loosely coupled, independently deployable services, which increases scalability and delivery speed and greatly increases the number, speed and variety of operational telemetry. Conventional AI-enabled systems for IT operations have been built with a series of discrete anomaly detection or root-cause-localization modules, each of which is applied to a single telemetry modality, which means they lack full diagnostic capabilities and require longer to resolve incidents. Based on twenty-five peer-reviewed and archival studies published until 2023, this paper investigates an operational-intelligence framework for retrieval-augmented (RA) fusion of distributed traces, structured and unstructured logs, and discrete system events, which are then coupled with a retrieval-augmented generation layer for self-healing remediation in cloud-native environments. Alongside trace-based localization systems with reported top one accuracy between forty-eight and ninety percent with mean-average-precision values ranging from ninety to ninety-seven percent, the synthesis includes log-parsing and anomaly-detection techniques such as the documented F-measure of 0.96 on the Hadoop Distributed File System benchmark. Large-language-model based incident assistants are analyzed on the basis of over 40,000 production incidents on two performance measures: their ability to retrieve incidents in the context of verified telemetry and their ability to suggest mitigation steps. The results are evaluated across six benchmark studies, and multi-modal, retrieval-grounded pipelines outperform single-modality baselines by 6 to 24 points on typical localization metrics, in a consistent performance improvement. The paper concludes that retrieval-augmented operational-intelligence architectures are a compelling evolutionary trajectory towards verifiable, low-hallucination automated remediation, from the original self-healing promise of autonomic computing, and that key challenges for moving these architectures to production include data-fusion latency, retrieval-corpus staleness and validation of generated mitigation steps

Downloads

Download data is not yet available.

References

T. Ahmed, S. Ghosh, C. Bansal, T. Zimmermann, X. Zhang, and S. Rajmohan, “Recommending Root-Cause and Mitigation Steps for Cloud Incidents Using Large Language Models,” in 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE), 2023, pp. 1737–1749. doi: 10.1109/ICSE48619.2023.00149.

A. Basiri et al., “Chaos Engineering,” IEEE Softw., vol. 33, no. 3, pp. 35–41, 2016, doi: 10.1109/MS.2016.60.

V. Chandola, A. Banerjee, and V. Kumar, “Anomaly Detection: A Survey,” ACM Comput. Surv., vol. 41, no. 3, pp. 1–58, 2009, doi: 10.1145/1541880.1541882.

Y. Chen et al., “Automatic Root Cause Analysis via Large Language Models for Cloud Incidents,” in Proceedings of the 19th European Conference on Computer Systems (EuroSys ’24), 2024, pp. 674–688. doi: 10.1145/3627703.3629553.

Y. Chen, M. Yan, D. Yang, X. Zhang, and Z. Wang, “Deep Attentive Anomaly Detection for Microservice Systems with Multimodal Time-Series Data,” in 2022 IEEE International Conference on Web Services (ICWS), 2022, pp. 373–383. doi: 10.1109/ICWS55610.2022.00062.

Y. Dang, Q. Lin, and P. Huang, “AIOps: Real-World Challenges and Research Innovations,” in 2019 IEEE/ACM 41st International Conference on Software Engineering: Companion Proceedings (ICSE-Companion), 2019, pp. 4–5. doi: 10.1109/ICSE-Companion.2019.00023.

M. Du, F. Li, G. Zheng, and V. Srikumar, “DeepLog: Anomaly Detection and Diagnosis from System Logs through Deep Learning,” in Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security, 2017, pp. 1285–1298. doi: 10.1145/3133956.3134015.

Y. Gan, M. Liang, S. Dev, D. Lo, and C. Delimitrou, “Sage: Practical and Scalable ML-Driven Performance Debugging in Microservices,” in Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, 2021, pp. 135–151. doi: 10.1145/3445814.3446700.

Y. Gan et al., “An Open-Source Benchmark Suite for Microservices and Their Hardware-Software Implications for Cloud & Edge Systems,” in Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems, 2019. doi: 10.1145/3297858.3304013.

Y. Gan et al., “Seer: Leveraging Big Data to Navigate the Complexity of Performance Debugging in Cloud Microservices,” in Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems, 2019. doi: 10.1145/3297858.3304004.

P. He, J. Zhu, Z. Zheng, and M. R. Lyu, “Drain: An Online Log Parsing Approach with Fixed Depth Tree,” in 2017 IEEE International Conference on Web Services (ICWS), 2017, pp. 33–40. doi: 10.1109/ICWS.2017.13.

S. He, P. He, Z. Chen, T. Yang, Y. Su, and M. R. Lyu, “A Survey on Automated Log Analysis for Reliability Engineering,” ACM Comput. Surv., vol. 54, no. 6, 2021, doi: 10.1145/3460345.

J. O. Kephart and D. M. Chess, “The Vision of Autonomic Computing,” Computer (Long. Beach. Calif)., vol. 36, no. 1, pp. 41–50, 2003, doi: 10.1109/MC.2003.1160055.

P. Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” in Advances in Neural Information Processing Systems 33, 2020. doi: 10.48550/arXiv.2005.11401.

Z. Li, “Practical Root Cause Localization for Microservice Systems via Trace Analysis,” in 2021 IEEE/ACM 29th International Symposium on Quality of Service (IWQOS), 2021, pp. 1–10. doi: 10.1109/IWQOS52092.2021.9521340.

J. Lin, P. Chen, and Z. Zheng, “Microscope: Pinpoint Performance Issues with Causal Graphs in Micro-Service Environments,” in Service-Oriented Computing: ICSOC 2018, vol. 11236, 2018, pp. 3–20. doi: 10.1007/978-3-030-03596-9_1.

D. Liu, “MicroHECL: High-Efficient Root Cause Localization in Large-Scale Microservice Systems,” in 2021 IEEE/ACM 43rd International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP), 2021, pp. 338–347. doi: 10.1109/ICSE-SEIP52600.2021.00043.

S. Luo, “Characterizing Microservice Dependency and Performance: Alibaba Trace Analysis,” in Proceedings of the ACM Symposium on Cloud Computing (SoCC ’21), 2021, pp. 412–426. doi: 10.1145/3472883.3487003.

L. Wu, J. Tordsson, E. Elmroth, and O. Kao, “MicroRCA: Root Cause Localization of Performance Issues in Microservices,” in NOMS 2020 - 2020 IEEE/IFIP Network Operations and Management Symposium, 2020, pp. 1–9. doi: 10.1109/NOMS47738.2020.9110353.

Y. Zhang, “CloudRCA: A Root Cause Analysis Framework for Cloud Computing Platforms,” in Proceedings of the 30th ACM International Conference on Information & Knowledge Management (CIKM ’21), 2021, pp. 4373–4382. doi: 10.1145/3459637.3481903.

R. Xin, P. Chen, and Z. Zhao, “CausalRCA: Causal Inference Based Precise Fine-Grained Root Cause Localization for Microservice Applications,” J. Syst. Softw., vol. 203, 2023, doi: 10.1016/j.jss.2023.111724.

G. Yu, “MicroRank: End-to-End Latency Issue Localization with Extended Spectrum Analysis in Microservice Environments,” in Proceedings of The Web Conference 2021 (WWW ’21), 2021, pp. 3087–3098. doi: 10.1145/3442381.3449905.

G. Yu, P. Chen, Y. Li, H. Chen, X. Li, and Z. Zheng, “Nezha: Interpretable Fine-Grained Root Causes Analysis for Microservices on Multi-Modal Observability Data,” in Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE 2023), 2023, pp. 553–565. doi: 10.1145/3611643.3616249.

S. Zhang et al., “Robust Failure Diagnosis of Microservice System Through Multimodal Data,” IEEE Trans. Serv. Comput., vol. 16, no. 6, pp. 3851–3864, 2023, doi: 10.1109/TSC.2023.3290018.

X. Zhang et al., “Robust Log-Based Anomaly Detection on Unstable Log Data,” in Proceedings of the 2019 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 2019, pp. 807–817. doi: 10.1145/3338906.3338931.

Downloads

Published

2024-12-31

How to Cite

Retrieval-Augmented Operational Intelligence for Self-Healing Cloud-Native Microservices Using Traces, Logs, and Events. (2024). Sciences Du Nord Nature Science and Technology, 1(02), 100-112. https://doi.org/10.71238/snnst.v1i02.183