AI Infrastructure: Deployment, Networking, and Storage
Google Cloud via Coursera Specialization
Overview
Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
In this Specialization, you’ll learn how to plan, deploy, and optimize AI infrastructure on Google Cloud. You’ll explore how AI and high-performance computing workloads depend on the right deployment model, network design, storage architecture, and accelerator choice.
You’ll compare options for GPU-accelerated clusters, including Google Compute Engine and Google Kubernetes Engine. You’ll also examine how GKE supports inference workflows through containerization, networking configurations, distributed training, GPU sharing, and model-level optimization.
Across the networking and storage courses, you’ll connect each part of the AI pipeline from data ingestion to training, inference, serving, and archiving. You’ll explore Cross-Cloud Network, Cloud Interconnect, Jumbo Frames, RDMA, Titanium offload, GKE Inference Gateway, IAM, Cloud Storage, Anywhere Cache, Dataflux Dataset, Cloud Storage FUSE, Managed Lustre, Hyperdisk ML, GPUs, and TPUs.
By the end of this Specialization, you'll be able to:
Select deployment options for AI workloads using GCE, GKE, GPU clusters, and inference workflows. Design networking and storage choices for AI data ingestion, training, serving, and archiving. Compare GPUs and TPUs and apply optimization strategies for performance, efficiency, and flexibility.
Syllabus
- Course 1: AI Infrastructure: Deployment Types
- Course 2: AI Infrastructure: Networking Techniques
- Course 3: AI Infrastructure: Storage Options
- Course 4: AI Infrastructure: Cloud GPUs
- Course 5: AI Infrastructure: Cloud TPUs
Courses
-
Curious about the powerful hardware behind AI? This module breaks down performance-optimized AI computers, showing you why they're so important. We'll explore how CPUs, GPUs, and TPUs make AI tasks super fast, what makes each one unique, and how AI software gets the most out of them. By the end, you'll know exactly how to pick the right compute for your AI projects, helping you make smart choices for your AI workkoads.
-
Welcome to the Cloud TPUs course. We'll explore the advantages and disadvantages of TPUs in various scenarios and compare different TPU accelerators to help you choose the right fit. You'll learn strategies to maximize performance and efficiency for your AI models and understand the significance of GPU/TPU interoperability for flexible machine learning workflows. Through engaging content and practical demos, we'll guide you step-by-step in leveraging TPUs effectively.
-
This course provides a comprehensive guide to deploying, managing, and optimizing AI and high-performance computing (HPC) workloads on Google Cloud. Through a series of lessons and practical demonstrations, you’ll explore diverse deployment strategies, ranging from highly customizable environments using Google Compute Engine (GCE) to managed solutions like Google Kubernetes Engine (GKE). Specifically, you’ll learn how to create clusters and deploy GKE for inference.
-
Welcome to the ""AI Infrastructure: Networking Techniques"" course. While AI Hypercomputer is renowned for its massive computational power using GPUs and TPUs, the secret to unlocking its full potential lies within the network. High-performance computing and large-scale model training demand incredibly fast, low-latency connections to continuously feed processors with data. In this course, you will learn to leverage Google Cloud's high-bandwidth, low-latency infrastructure to optimize data transfer and communication between all the components of your AI system. By the end, you will grasp the critical role networking plays across the entire AI pipeline from data ingestion and training to inference and be able to apply best practices to ensure your workloads run at maximum speed.
-
In this course, you’ll take a comprehensive journey through the storage solutions available on Google Cloud, specifically tailored for AI and high-performance computing (HPC) workloads. You’ll learn how to choose the right storage for each stage of the ML lifecycle. You’ll explore how to optimize for I/O performance during training, manage massive datasets for data preparation, and serve model artifacts with low latency. Through practical examples and demonstrations, you’ll gain the expertise to design robust storage solutions that accelerate your AI innovation.
Taught by
Google Cloud Training