This program is designed for DevOps professionals, software engineers, and data scientists who want to master the principles and practices of AIOps. Learners will build a comprehensive skill set in operationalizing the entire lifecycle of AI systems, starting with traditional machine learning (MLOps), advancing to Large Language Model applications (LLMOps), and culminating in the development and management of AI agents (AgentOps). The curriculum focuses on hands-on application of industry-standard tools and cloud-native technologies to build robust, scalable, and maintainable AI workflows.
Overview
Syllabus
- MLOps: Automated Pipelines and Model Monitoring
- This course provides the practical skills needed to deploy and manage machine learning models in real-world production environments. You will learn to design robust data pipelines that ingest, version, and validate data, track experiments, and implement automated continuous training and deployment (CI/CD) pipelines. The course also covers containerizing model APIs, building scalable serving endpoints on AWS, and adding crucial components like feature stores, model registries, monitoring, data drift detection, explainability, and fairness assessments. By the end, you will be able to build, deploy, and maintain reliable and scalable machine learning systems effectively.
- LLMOps: RAG Pipelines, Evaluation, and Cost Control
- Operate LLM applications in production, not just prototype them. You will version prompts as code, stand up a Chroma vector store, and compose a retrieval-augmented generation pipeline on the OpenAI SDK, then make that pipeline observable, measurable, and affordable: tracing with Arize Phoenix, an automated RAGAS evaluation suite gated against a golden set, token-level cost monitoring, semantic caching, a FastAPI gateway, layered guardrails, prompt A/B tests, automated ingestion with a blue/green index swap, and end-to-end latency optimization. A graded project ties the whole stack into a retrieval-augmented FAQ service.
- AgentOps: Deploying, Guarding, and Monitoring AI Agents
- AgentOps teaches you to take an AI agent from prototype to a production system you can deploy, secure, and operate. Using RealThor, a real-estate analysis agent, you will version its prompts, tools, and configuration for reproducibility, build evaluation suites with ship/no-ship gates, and deploy it as a containerized CLI.
You will then make it safe and observable: enforce guardrails and role-based access, sandbox agent-generated code, add human-in-the-loop approval for high-risk actions, and instrument it with structured logging, dashboards, and reasoning traces. You will also debug failed runs via state replay, isolate state across concurrent users, and connect the agent to external tools over MCP. The capstone: operationalize a SalesOps agent for a fictional B2B company.
Taught by
Amal Feriani, Jeff Chen, and Henrique Santana