Start speaking a new language. It’s just 3 weeks away.
Get 20% off all career paths from fullstack to AI
Overview
Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
This talk follows a real failure in a distributed system and the production troubleshooting process. It examines failure modes, interpreting metrics in context, centralized logging, lost alerts, and rollback.
Syllabus
Introduction
Anatomy of a System
System Definition
Failure Modes
Failure Walkthrough
Signs of Trouble
Shutting It Down
What Changed
Roll It Back
What We Missed
A sensible default
What we learned
Metrics need context
Centralized logging
Losing alerts
Recap
Questions
Taught by
GOTO Conferences