Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

Microsoft

Distributed Training & Advanced Application

Microsoft via Coursera

Overview

Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
Enterprise-scale deep learning demands more than a single GPU; it requires distributed infrastructure, optimized training pipelines, and the ability to build advanced multi-domain applications. This course builds the engineering skills to scale deep learning workflows and deploy production-grade models on Azure ML. You'll implement PyTorch Distributed Data Parallel (DDP) and Fully Sharded Data Parallel (FSDP) pipelines, profiling communication overhead and comparing throughput across single- and multi-node configurations. You'll configure DeepSpeed ZeRO optimization stages, including CPU offloading, and apply Microsoft's full training acceleration stack using Microsoft Olive and Azure Container for PyTorch. You'll build object detection and vision-language pipelines using Florence-2 and CLIP, and engineer enterprise NLP workflows for classification, NER, and extractive QA with Hugging Face Transformers and Azure AI Language. You'll also design multimodal fusion architectures combining vision, text, and audio using Microsoft Phi-4 and OpenAI Whisper. By the end of this course, you'll be able to implement distributed training at scale, accelerate large model training with Microsoft's optimization stack, and build production-grade pipelines across computer vision, NLP, and multimodal domains. This course is designed for advanced machine learning engineers scaling infrastructure for massive datasets and multi-modal application development.

Syllabus

  • Project Module: Distributed Training Scenario
    • Synthesize your distributed training and model acceleration skills to scale a massive transformer model across a multi-GPU, multi-node cloud cluster. You will write a Python training script that implements Microsoft DeepSpeed ZeRO-3 parameter sharding and activation checkpointing, paired with a DeepSpeed configuration file implementing ZeRO-3 CPU offloading. You will then write the Azure ML SDK v2 code to submit this job to a remote compute cluster using the curated ACP environment and optimized NCCL environment variables.

Taught by

Microsoft

Reviews

Start your review of Distributed Training & Advanced Application

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.