Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

YouTube

Building Fault-Tolerant Massive Ray Clusters on Anyscale

Anyscale via YouTube

Overview

Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
This session explains how Ray is engineered to run reliably on clusters exceeding 10,000 nodes despite network flakiness, spot preemptions, hardware failures, and resource contention. It covers fault tolerance, state management, recovery, elasticity, and workload-aware scheduling for large-scale AI workloads.

Syllabus

Building Fault-Tolerant Massive Ray Clusters on Anyscale | Ray Summit 2025

Taught by

Anyscale

Reviews

Start your review of Building Fault-Tolerant Massive Ray Clusters on Anyscale

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.