Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

YouTube

How to Not Destroy Your Production Kubernetes Clusters

USENIX via YouTube

Overview

Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
This course analyzes real production incidents from managing large Kubernetes clusters, including failures caused by routine operational changes. It presents investigation, restoration, mitigation, and lessons for maintaining cluster availability.

Syllabus

Intro
Background
Postmodern Database
Automation
User escalation
Initial investigation
Restoring service objects
Collecting service definitions
The impact of the incident
The reason for the failure
Fixing the webhooks
Why the operator went rogue
Kubernetes label selector package
Test engineer accidentally created app load balancer
What can we learn
Paradoxical Finalizer
Paging Storm
Mitigation
Kubernetes Platform
Manual Operations
Lessons Learned
User Complaints
Monitoring Dashboard
Victim Cluster
Security Context Change
Learnings
Recap
Key takeaways

Taught by

USENIX

Reviews

Start your review of How to Not Destroy Your Production Kubernetes Clusters

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.