Cortex-Integrated SLO Management Framework: A Dependency-Aware, Closed-Loop Reliability Optimization System

Authors

  • Vaidyanathan Sivakumaran

Keywords:

Service Level Objectives, Site Reliability Engineering, Azure Data Explorer, dependency-aware observability, closed-loop reliability management, consumer impact analysis, cloud-native microservices, telemetry correlation

Abstract

Service Level Objectives (SLOs) are foundational to Site Reliability Engineering (SRE), yet existing implementations lack continuous enforcement, dependency awareness, and automated improvement mechanisms. This paper presents a Cortex-Integrated SLO Management Framework that enables end-to-end lifecycle management of SLOs—from definition and ingestion to evaluation, dependency correlation, and automated recommendations. The system leverages a workflow-driven ingestion mechanism, Azure-based processing pipelines, and a scalable analytics backend using Azure Data Explorer (ADX). By incorporating dependency-aware evaluation and consumer impact propagation, the framework provides actionable insights that significantly improve reliability outcomes. The architecture integrates centralized SLO governance, telemetry aggregation, service correlation, consumer impact analysis, and automated recommendation generation into a unified operational workflow. Controlled experimental evaluation conducted within a cloud-native microservices environment demonstrated substantial improvements in incident detection responsiveness, dependency attribution accuracy, and operational resolution efficiency. Specifically, the framework achieved approximately five-to-eight times faster Mean Time to Detect (MTTD), a fifty to sixty percent reduction in Mean Time to Resolve (MTTR), and root-cause attribution accuracy exceeding ninety percent compared with traditional threshold-based monitoring approaches. These findings demonstrate the value of integrating SLO governance with dependency-aware observability and automated recommendation generation for modern distributed cloud environments.

Downloads

Download data is not yet available.

References

Google, Site Reliability Engineering: How Google Runs Production Systems. O'Reilly Media, 2016.

Google, The Site Reliability Workbook: Practical Ways to Implement SRE. O'Reilly Media, 2018.

Datadog, "Service level objectives and error budgets," Datadog Documentation, 2023. [Online]. Available: https://docs.datadoghq.com

OpenTelemetry, "OpenTelemetry documentation," 2023. [Online]. Available: https://opentelemetry.io

Microsoft, "Azure Monitor documentation," Microsoft Learn, 2024. [Online]. Available: https://learn.microsoft.com

Microsoft, "Azure Workbooks overview," Microsoft Learn, 2024. [Online]. Available: https://learn.microsoft.com

Microsoft, "Azure Data Explorer documentation," Microsoft Learn, 2024. [Online]. Available: https://learn.microsoft.com

Cortex, "What is an internal developer portal?" Cortex.io, 2023. [Online]. Available: https://www.cortex.io

Cortex, "Building reliable services: A guide to setting SLOs," Cortex.io, 2023. [Online]. Available: https://www.cortex.io

Downloads

Published

15.07.2026

How to Cite

Vaidyanathan Sivakumaran. (2026). Cortex-Integrated SLO Management Framework: A Dependency-Aware, Closed-Loop Reliability Optimization System. International Journal of Intelligent Systems and Applications in Engineering, 14(1s), 2050–2062. Retrieved from https://www.ijisae.org/index.php/IJISAE/article/view/8464

Issue

Section

Research Article