Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

Microsoft

Data Centre Engineering for AI

Microsoft via edX

Overview

Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates

Modern artificial intelligence workloads demand unprecedented compute density, pushing data center facilities to their physical and operational limits. This course bridges the gap between hardware engineering and cloud-based monitoring, equipping participants to design, sustain, and optimize infrastructure for GPU-intensive AI clusters.

The curriculum begins with power distribution and high-availability engineering. Participants learn to calculate and optimize Power Usage Effectiveness (PUE) using telemetry data from Microsoft Excel and Azure Monitor. Learners implement N+1 and 2N hardware redundancy topologies using Azure Resource Manager (ARM) templates to prevent downtime within strict power envelopes.

From power, the course transitions into precision thermal dynamics. Participants explore airflow physics, hot/cold aisle containment, and localized heat map generation in Excel using Azure IoT Central telemetry streams. The course examines GPU thermal throttling mechanics and teaches proactive mitigation strategies, including fan curve tuning and liquid-to-chip cooling loops, to maintain peak training performance.

Finally, participants will evaluate strategic infrastructure trade-offs and deploy automated protection networks. The curriculum compares on-premise liquid cooling against elastic cloud bursting, modeling Total Cost of Ownership (CapEx vs. OpEx) and data gravity constraints. Participants finish by deploying comprehensive sensor grids in Azure IoT Central and building automated incident response workflows with Azure Monitor, Logic Apps, and Microsoft Teams to preserve high-value AI hardware.

Syllabus

  • Calculate and optimize Power Usage Effectiveness (PUE)
  • Implement hardware redundancy strategies (N+1/2N)
  • Design precision thermal management plans
  • Engineer mitigation strategies to prevent thermal throttling
  • Conduct a comparative analysis of on-premise liquid cooling
  • Architect scalable, hybrid AI infrastructure
  • Deploy comprehensive environmental sensor networks
  • Configure automated incident response

Reviews

Start your review of Data Centre Engineering for AI

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.