AgentOps teaches you to take an AI agent from prototype to a production system you can deploy, secure, and operate. Using RealThor, a real-estate analysis agent, you will version its prompts, tools, and configuration for reproducibility, build evaluation suites with ship/no-ship gates, and deploy it as a containerized CLI.
You will then make it safe and observable: enforce guardrails and role-based access, sandbox agent-generated code, add human-in-the-loop approval for high-risk actions, and instrument it with structured logging, dashboards, and reasoning traces. You will also debug failed runs via state replay, isolate state across concurrent users, and connect the agent to external tools over MCP. The capstone: operationalize a SalesOps agent for a fictional B2B company.
Overview
Syllabus
- Welcome to Operationalizing AI Agents
- An orientation to the AgentOps course: the RealThor agent you will operate, the four chapters of the course, and why agent operations differs from general LLMOps.
- Understand Agent System Versioning
- Learn why agent behavior is defined by more than code and how to version prompts, tools, and configurations for reproducibility.
- Version an Agentic System with Git and uv
- Apply structured project layout and environment management to an agent so any past behavior can be reproduced and any change can be traced.
- Understand Evaluation Frameworks for Agentic Systems
- Learn how to design reproducible evaluation suites that establish trust in agent behavior and detect regressions across versions.
- Create an Automated Evaluation Suite for Agents
- Define a benchmark task set with success criteria, build an automated test harness, and produce a regression report comparing agent versions.
- Understand Agent Deployment Principles
- Learn how agents run under real conditions (on-demand, triggered, or scheduled) and how CI/CD pipelines enforce quality before code reaches production.
- Package and Deploy an Agentic System as a Containerized CLI
- Package an agent as a containerized CLI with Docker Compose and a click interface, using an ops.sh script that runs the evaluation suite as a quality gate before the image is built.
- Understand Agent Guardrails and Governance Principles
- Learn the control layer between an agent and the external world that constrains what the agent can do, regardless of what the model reasons.
- Operationalize Agent Guardrails with LangChain
- Build guardrails as LangChain AgentMiddleware that intercept every agent step on the RealThor agent: enforcing role-based access control, regional data access rules, and output redaction.
- Understand Sandboxed Code Execution for Agents
- Learn why agents that execute code require isolated environments and how container-based sandboxing separates the reasoning layer from execution.
- Implement an In-Process Sandboxed Code Execution Tool
- Build an in-process code execution tool that validates agent-generated Python with the ast module, runs it against restricted globals over pandas DataFrames, and enforces a thread-based timeout.
- Understand Human-in-the-Loop Patterns for Agents
- Learn the design principles for systems that pause at high-risk decisions, treating human oversight as a designed control mechanism, not a fallback.
- Implement a Human-in-the-Loop Gate with LangGraph
- Build a LangGraph agent that pauses before a high-risk action, notifies a reviewer, and routes execution correctly based on approve or reject signals.
- Understand Agent Monitoring and Operational KPIs
- Learn the signals that reveal when something is wrong in a deployed agent and how dashboards turn run data into operational visibility.
- Build an Agent Monitoring Dashboard from Structured Logs
- Add structured run-level logging to a deployed agent and build a pure-Python monitor that generates a self-contained report detecting failures, cost overruns, and latency degradation.
- Understand State Persistence and Time-Travel Debugging for Agents
- Learn how to reconstruct the exact conditions of a failed agent run using state checkpoints so failures that cannot be reproduced can be diagnosed.
- Debug a Failed Agent Run via State Replay with LangGraph
- Configure a durable SqliteSaver checkpointer, then reconstruct the context of a failed agent run from its newest persisted snapshot and its runs.jsonl log to diagnose what happened.
- Understand Agent Observability and Reasoning Traces
- Learn how tracing turns debugging from guesswork into understanding by revealing the full sequence of thoughts, tool calls, and intermediate results.
- Implement Agent Tracing with Langfuse
- Instrument a LangGraph agent with Langfuse via a LangChain callback handler and use the resulting span tree to inspect the agent's reasoning and tool calls.
- Understand Concurrent State Management for Multi-User Agents
- Learn how context scoping and session isolation prevent state leakage and loss when a single agent system serves many users simultaneously.
- Manage Concurrent Multi-User Agent State with LangGraph
- Build a multi-tenant LangGraph agent that isolates each user's state with a per-(user, session) thread key backed by a SQLite checkpointer, verified with thread-key and ownership tests.
- Understand Model Context Protocol (MCP) for Agents
- Learn how MCP standardizes connections between AI agents and external tools, enabling structured, scalable communication across systems and other agents.
- Operationalize MCP Servers for Agentic Systems
- Connect an AI agent to a provided FastMCP server that exposes custom tools by loading them through langchain-mcp-adapters, running as an unauthenticated development service.
- AgentOps Course Review
- A capstone reflection on the 12 AgentOps skills and how versioning, deployment, control, observability, and scale combine in production agentic systems.
- Project: Operationalize a SalesOps Agent for Production at UdaCenture
- In this project, you will operationalize a SalesOps agentic workflow for a fictional B2B company. You will start from a partially implemented prototype and transform it into a production-ready system.
Taught by
Henrique Santana