- Technology
- Computer Science
- Distributed Systems
- High Performance Computing
- Parallel Computing
- GPU Programming
- CUDA
- Technology
- IT & Networking
- Computer Networking
- Network Operations & Management
- Network Performance Analysis
- Technology
- Computer Science
- Distributed Systems
- High Performance Computing
- Parallel Computing
- GPU Computing
New Approaches to Network Telemetry for AI Performance Optimization
Open Compute Project via YouTube
Learn AI, Data Science & Business — Earn Certificates That Get You Hired
Live Online Classes in Design, Coding & AI — Small Classes, Free Retakes
Overview
Google, IBM & Meta Certificates – 40% Off
One plan covers every Professional Certificate on Coursera.
Unlock All Certificates
Learn how to optimize large GPU clusters for machine learning workloads in this 11-minute conference talk from Nvidia's Principal Software Research Architect. Explore why traditional data center telemetry approaches fall short for massive ML models and discover new methods for extracting meaningful metrics from large-scale clusters. Examine how ML workloads create unique patterns of similarity and synchronicity across adaptive-routed, rail-optimized, fat-tree topologies, and understand the specialized abstractions developed to identify performance optimization opportunities in ML-focused infrastructure.
Syllabus
New approaches to network telemetry Essential for AI performance
Taught by
Open Compute Project