- Technology
- Computer Science
- Distributed Systems
- High Performance Computing
- Parallel Computing
- GPU Programming
- CUDA
- Technology
- Generative AI
- Large Language Models (LLMs)
- AI Models
- Meta
- LLaMA (Large Language Model Meta AI)
- Technology
- Computer Science
- Distributed Systems
- High Performance Computing
- Parallel Computing
- GPU Computing
Heterogeneous Hybrid Distributed Training for Large-Scale Language Models
OpenInfra Foundation via YouTube
Launch Your Cybersecurity Career in 6 Months
Learn AI, Data Science & Business — Earn Certificates That Get You Hired
Overview
Google, IBM & Meta Certificates – 40% Off
One plan covers every Professional Certificate on Coursera.
Unlock All Certificates
Learn about the technical challenges and solutions in heterogeneous distributed training for Large Language Models (LLMs) in this 11-minute conference talk. Explore how integrating different computing resources for distributed parallel acceleration can support the development of LLMs with hundreds of billions of parameters. Discover the research conducted by China Mobile and industry partners to overcome challenges related to GPU architecture differences, memory constraints, and vendor hardware incompatibilities. Gain insights into the core functional components of a training system designed to enable heterogeneous GPUs to work together effectively, contributing to the advancement of the intelligent computing ecosystem.
Syllabus
Heterogeneous Hybrid Distributed Training Helps the Development of Large-Scale Language Model
Taught by
OpenInfra Foundation