Courses from 1000+ universities
AI got cheap enough that Duolingo’s most expensive plan may not survive it. I read the earnings call transcript and opened the app to see what is actually changing for learners.
600 Free Google Certifications
Artificial Intelligence
IT & Networking
Software Engineering
Supporting Victims of Domestic Violence
Know Thyself - The Value and Limits of Self-Knowledge: The Examined Life
Understanding Dementia
Organize and share your learning with Class Central Lists.
View our Lists Showcase
This talk examines how SRE teams can shift observability from service health and business metrics toward customer experience, usability, accessibility, and feedback.
Use Jupyter notebooks to investigate and remediate a cache slowdown during an SRE incident, then document the response for retrospective.
Use incident analysis, richer metrics, and chaos experiments to reveal progress toward safer, more adaptive operations.
Research findings on how SRE teams coordinate under uncertainty during outages, and how tools, dependencies, and incident command shape cognitive workload.
Deliberate node failures, observability, and error budgets expose hidden dependencies and improve the reliability of stateful production systems.
Use SLI, SLO, and error budgets to negotiate application and business stress without turning metrics into punitive thresholds.
Breaks PXE bootstrapping into stages—from downloading a bootloader to loading an operating system—and shows how to automate provisioning.
Automate detection and remediation of distributed performance anti-patterns with trace analysis and SLI/SLO quality gates in CI/CD pipelines.
Explore Paxos through Skinny, a distributed lock service, and work through the practical challenges of implementing consensus under failures.
Examines how pervasive automation can harm SRE teams and complex socio-technical systems, with strategies for safer automation and coordination.
Learn how a very small SRE team can amplify its organizational impact through leverage, data, tools, communication, architecture reviews, and game days.
Lessons from two decades of engineering failures for designing and operating reliable large-scale online services.
An empirical approach to surfacing inherent chaos in distributed systems and building confidence in their resilience through structured Chaos Engineering practices.
A practical tour of Linux performance analysis and tuning using observability tools, benchmarking, profiling, tracing, and system-performance recipes.
Learn how the Incident Command System evolved from emergency response into a flexible framework for managing production incidents.
Get personalized course recommendations, track subjects and courses with reminders, and more.