Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

YouTube

Practicalities of Productionizing Distributed Systems

GOTO Conferences via YouTube

Overview

Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
This talk presents tactics for productionizing distributed systems, with a focus on diagnosing slowness and failures, managing overload, and rolling out infrastructure changes safely. It also discusses tracing, metrics, and operating systems with multiple versions in use.

Syllabus

Intro
Why you should listen to me
Quick foundation
What makes distributed systems different
A subset of failures
Clients stuck to an overloaded process
Partial failure
"It's slow" is the hardest problem you'll ever debug
Create partial availability
"Who to Follow" in the monorail
Knowing what the system has done
Percentiles, not averages
Tracing
On profiling
Releases should change a metric
Free-form logs are liars
Common "problems" are overlogged
Uncommon problems
Avoid coordination
Backpressure
Dropping new messages on the floor
Returning "overload" error responses
Timeouts and exponential back-offs
Roll out infrastructure with feature flags
if (Decider.available..)
Multiple versions are the norm
Datacenter schedulers are worth it
Collaboration is politics
No time-traveling stalkers
moral necessity
Data minimization is a

Taught by

GOTO Conferences

Reviews

Start your review of Practicalities of Productionizing Distributed Systems

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.