.
ISSN : 2583-2646

Autonomous Multi-Cloud Data Pipeline Orchestration Using AI-Driven Observability and Self-Healing ETL Frameworks

ESP Journal of Engineering & Technology Advancements
© 2026 by ESP JETA
Volume 6  Issue 3
Year of Publication : 2026
Author : Sreenivasa Reddy Vemareddy

Citation:

Sreenivasa Reddy Vemareddy , 2026. Autonomous Multi-Cloud Data Pipeline Orchestration Using AI-Driven Observability and Self-Healing ETL Frameworks,  Volume 6 Issue 3: 52-59.

Abstract:

As new and multiple cloud platforms emerge, each providing different features and capabilities, the data journey is becoming more complex and more important, resulting in an autonomous multi-cloud data pipeline orchestration problem such as container platforms, serverless services, streaming engines, data lakes, and disparate cloud providers host many analytical workloads. This review covers the impact of AI-powered observability and selfhealing ETL solutions in adaptive control in these distributed environments. A survey of the existing literature reveals that there is no single established research stream that is sufficient for handling the autonomous multi-cloud ETL orchestration problem, but it is a cross-section of the following streams of research: cloud-native orchestration, distributed stream processing, data quality monitoring, root-cause analysis, anomaly detection and self-adaptive systems. Several reported studies have been conducted in the fields of resource scheduling, pipeline elasticity, container portability, log-based anomaly detection, and quality measurement, but much remains to be done to break all of the restrictions. The major limitations identified include insufficient cross-cloud empirical validation, weak causal reasoning, limited evidence of successful automated remediation, heavy reliance on labelled incident data, and inadequate integration between data observability and orchestration policies. The article concludes by calling for future research to explore the operability of ETL pipelines as adaptive socio-technical systems that are closed in an operational loop with observability, policy, lineage and remediation.

References:

[1] Varghese, B., & Buyya, R. (2018). Next generation cloud computing: New trends and research directions. Future Generation Computer Systems, 79, 849–861.

[2] Kratzke, N., & Quint, P.-C. (2017). Understanding cloud-native applications after 10 years of cloud computing—A systematic mapping study. Journal of Systems and Software, 126, 1–16.

[3] Pahl, C., Brogi, A., Soldani, J., & Jamshidi, P. (2019). Cloud container technologies: A state-of-the-art review. IEEE Transactions on Cloud Computing, 7(3), 677–692.

[4] Hassan, H. B., Barakat, S. A., & Sarhan, Q. I. (2021). Survey on serverless computing. Journal of Cloud Computing, 10(1), Article 39.

[5] Singh, S., & Chana, I. (2016). A survey on resource scheduling in cloud computing: Issues and challenges. Journal of Grid Computing, 14(2), 217–264.

[6] Isah, H., Abughofa, T., Mahfuz, S., Ajerla, D., Zulkernine, F., & Khan, S. (2019). A survey of distributed data stream processing frameworks. IEEE Access, 7, 154300–154316.

[7] Giebler, C., Gröger, C., Hoos, E., Schwarz, H., & Mitschang, B. (2019). Leveraging the data lake: Current state and challenges. Big Data Research, 16, 1–12.

[8] Firmani, D., Scannapieco, M., Schirru, R., & Tosco, L. (2016). On the meaningfulness of “big data quality”. Data Science and Engineering, 1(1), 6–20.

[9] Cai, L., & Zhu, Y. (2015). The challenges of data quality and data quality assessment in the big data era. Data Science Journal, 14, Article 2.

[10] Ehrlinger, L., & Wöß, W. (2022). A survey of data quality measurement and monitoring tools. Frontiers in Big Data, 5, Article 850611.

[11] Pang, G., Shen, C., Cao, L., & van den Hengel, A. (2021). Deep learning for anomaly detection: A review. ACM Computing Surveys, 54(2), Article 38.

[12] Blázquez-García, A., Conde, A., Mori, U., & Lozano, J. A. (2021). A review on outlier/anomaly detection in time series data. ACM Computing Surveys, 54(3), Article 56.

[13] Soldani, J., & Brogi, A. (2022). Anomaly detection and failure root cause analysis in cloud applications: A survey. ACM Computing Surveys, 55(3), Article 61.

[14] Leitner, P., Cito, J., & Stöckli, E. (2016). Modelling and managing deployment costs of microservice-based cloud applications. Journal of Internet Services and Applications, 7, Article 11.

[15] Di Francesco, P., Lago, P., & Malavolta, I. (2019). Architecting with microservices: A systematic mapping study. Journal of Systems and Software, 150, 77–97.

[16] Weyns, D., Iftikhar, M. U., Malek, S., & Andersson, J. (2016). Claims and supporting evidence for self-adaptive systems: A literature study. ACM Transactions on Autonomous and Adaptive Systems, 11(4), Article 24.

Keywords:

AI Observability. Autonomous Orchestration, Data Pipelines, Multi-Cloud, Self-Healing ETL .