ML Training Acceleration with Heterogeneous Resources in ByteDance
CNCF [Cloud Native Computing Foundation] via YouTube
AI, Data Science & Cloud Certificates from Google, IBM & Meta
Learn Generative AI, Prompt Engineering, and LLMs for Free
Overview
Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
This talk examines system-level methods for accelerating large-scale machine learning training on heterogeneous CPU and GPU resources. It covers GPU sharing, NUMA-aware allocation of CPU, memory, GPU, and network resources, and high-throughput communication using RDMA CNI and Intel SR-IOV.
Syllabus
Intro
GPU Offline Training (Network)
GPU Offline Training (Scheduling).
GPU Online Serving
GPU Unified Scheduling
Future Work
Taught by
CNCF [Cloud Native Computing Foundation]