Operate LLM applications in production, not just prototype them. You will version prompts as code, stand up a Chroma vector store, and compose a retrieval-augmented generation pipeline on the OpenAI SDK, then make that pipeline observable, measurable, and affordable: tracing with Arize Phoenix, an automated RAGAS evaluation suite gated against a golden set, token-level cost monitoring, semantic caching, a FastAPI gateway, layered guardrails, prompt A/B tests, automated ingestion with a blue/green index swap, and end-to-end latency optimization. A graded project ties the whole stack into a retrieval-augmented FAQ service.
Overview
Syllabus
- Welcome to LLM Ops
- An overview of the LLM Ops course, covering key topics, tools, and what you will build by the end.
- Understand Prompt Versioning Systems
- Learn why treating prompts as code matters and how version control, templating, and A/B principles apply to prompt management.
- Implement a Prompt Versioning System with Git and Jinja2
- Build a prompt versioning system using Git branches and Jinja2 templates, then run a basic A/B experiment comparing two prompt versions.
- Understand Vector Databases and Similarity Search
- Learn how vector embeddings, ANN algorithms, and vector database architectures enable efficient semantic search over high-dimensional data.
- Configure a Vector Database with Chroma and sentence-transformers
- Configure a local in-process Chroma vector database, ingest chunks embedded with OpenAI text-embedding-3-small, measure recall@5, then compare a local sentence-transformers embedder.
- Understand Retrieval-Augmented Generation (RAG) Architecture
- Learn the end-to-end RAG workflow (from query embedding to context-augmented generation) and the role of each component in reducing hallucinations.
- Operationalize a Basic RAG Pipeline with the OpenAI SDK and Chroma
- Compose a functional RAG application with the raw OpenAI SDK and Chroma that retrieves relevant documents from a vector store and generates context-aware answers via an LLM API.
- Understand LLM Observability and Distributed Tracing
- Learn how distributed tracing concepts apply to LLM applications and what key metrics (latency, token counts, and span hierarchies) reveal about application health.
- Implement LLM Call Tracing with Arize Phoenix
- Instrument a RAG application to send full execution traces to Arize Phoenix, then analyze latency and token usage across each pipeline step.
- Understand LLM Evaluation Methods
- Learn the metrics and methodologies for evaluating LLM applications: reference-based metrics, judge models, golden sets, and regression detection.
- Build an Automated Evaluation Suite with RAGAS
- Build an automated evaluation suite using RAGAS to measure retrieval relevance, answer faithfulness, and answer correctness against a golden test set, with a regression-detection gate.
- Understand LLM Pricing and Token-Based Cost Management
- Learn how LLM providers price API calls by token, how to count tokens programmatically, and the strategies used to aggregate and monitor costs at scale.
- Build a Token Usage and Cost Monitoring System with tiktoken
- Instrument LLM API calls to capture token usage, log per-call cost data, and build a simple dashboard showing cost trends by day, model, and query type.
- Understand Semantic Caching for LLM Applications
- Learn how semantic caching uses vector similarity to serve cached LLM responses for semantically equivalent queries, reducing cost and latency.
- Implement Semantic Caching with Chroma
- Build a semantic caching layer that intercepts LLM calls, checks a vector store for similar cached queries, and stores new results: measurably reducing API calls.
- Understand High-Performance LLM Inference Servers
- Learn why specialized inference servers like vLLM outperform simple API wrappers and how techniques like PagedAttention and continuous batching maximize GPU throughput.
- Understand LLM Gateway Architecture and API Routing Patterns
- Learn how an LLM gateway centralizes model access, enforces organizational policies, secures API keys, and provides load balancing across multiple model providers.
- Build an LLM Gateway with FastAPI
- Build a FastAPI-based LLM gateway that routes queries across model tiers, adds retry with exponential backoff, and abstracts a second provider behind a stubbed adapter.
- Understand LLM Safety and Guardrails
- Learn the safety risks specific to LLM applications (prompt injection, PII exposure, jailbreaks, and unsafe outputs) and the architectural patterns used to mitigate them.
- Implement Input/Output Guardrails with LLM Guard
- Build input and output guardrails using LLM Guard to detect prompt injection, redact PII, and validate LLM output for safety and format compliance.
- Understand A/B Testing Principles for LLM Prompt Optimization
- Learn how controlled experiments, statistical significance, and feature flagging are applied to evaluate competing prompt variants on real production traffic.
- Conduct Prompt A/B Tests with Python Feature Flags
- Implement a traffic-splitting system that routes live requests to two prompt variants, collects performance metrics, and applies a chi-squared test to determine the winner.
- Understand RAGOps and Automated Data Pipeline Architecture
- Learn how event-driven pipeline design and orchestration tools keep a RAG knowledge base current through automated, idempotent ingestion workflows.
- Automate the RAG Data Pipeline (local watcher + blue/green index swap)
- Build an event-driven RAG pipeline with a filesystem watcher and an atomic blue/green index swap that chunks, embeds, and upserts new documents, with S3 + Lambda as the production analogy.
- Understand RAG Latency Bottlenecks and Optimization Strategies
- Learn how to profile each component of the RAG pipeline, identify latency bottlenecks, and apply optimization strategies including token streaming and hardware acceleration.
- Optimize End-to-End RAG Latency with Streaming and Profiling
- Profile a RAG pipeline to locate latency bottlenecks, measure token streaming against a blocking endpoint, and run a vector-search tuning sweep, measuring improvement at each step.
- LLM Ops Course Review
- Recap of all LLM Ops skills covered, key takeaways, and guidance on next steps in the AI Ops Engineer program.
- Project: Production LLM FAQ Service
- Build a retrieval-augmented FAQ service for a fictional e-commerce company, covering the full LLM Ops stack across nine graded deliverables.
Taught by
Jeff Chen