An SLO-Driven Approach to Enhance Kubernetes Cluster Reliability

Explore an SLO-driven approach to enhance Kubernetes cluster reliability in this conference talk from KubeCon + CloudNativeCon Europe 2021. Delve into the challenges of defining reliability for large-scale Kubernetes clusters and learn how Service Level Objectives (SLOs) can be effectively implemented. Discover the philosophy behind SLO-driven reliability engineering and gain insights from Ant Financial's experience with one of the world's largest Kubernetes clusters. Examine concrete cases and lessons learned in building SLO frameworks, covering aspects such as monitoring, alerting, and tracing. Understand the complexities of defining SLOs for Kubernetes services compared to classic web services, and explore topics including fleet management, EZE SLO design, fine-grained and component SLOs, alerting philosophy, and SLO management.

Syllabus

Thank You to Our Session Recording Sponsor
Outline
Motivation
Fleet Management
General Approach
SLO Approach
SLO Recap
What SRE cares on K8S?
EZE SLO Design
Fine-grained SLO
Component SLO
Overall SLO Graph
Why RatioRate is bad?
Alerting Philosophy
SLO Management