Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

CNCF [Cloud Native Computing Foundation]

ML Training Acceleration with Heterogeneous Resources in ByteDance

CNCF [Cloud Native Computing Foundation] via YouTube

Overview

Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
This talk examines system-level methods for accelerating large-scale machine learning training on heterogeneous CPU and GPU resources. It covers GPU sharing, NUMA-aware allocation of CPU, memory, GPU, and network resources, and high-throughput communication using RDMA CNI and Intel SR-IOV.

Syllabus

Intro
GPU Offline Training (Network)
GPU Offline Training (Scheduling).
GPU Online Serving
GPU Unified Scheduling
Future Work

Taught by

CNCF [Cloud Native Computing Foundation]

Reviews

Start your review of ML Training Acceleration with Heterogeneous Resources in ByteDance

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.