Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

Udacity

LLMOps: RAG Pipelines, Evaluation, and Cost Control

via Udacity

Overview

Operate LLM applications in production, not just prototype them. You will version prompts as code, stand up a Chroma vector store, and compose a retrieval-augmented generation pipeline on the OpenAI SDK, then make that pipeline observable, measurable, and affordable: tracing with Arize Phoenix, an automated RAGAS evaluation suite gated against a golden set, token-level cost monitoring, semantic caching, a FastAPI gateway, layered guardrails, prompt A/B tests, automated ingestion with a blue/green index swap, and end-to-end latency optimization. A graded project ties the whole stack into a retrieval-augmented FAQ service.

Syllabus

  • Welcome to LLM Ops
    • An overview of the LLM Ops course, covering key topics, tools, and what you will build by the end.
  • Understand Prompt Versioning Systems
    • Learn why treating prompts as code matters and how version control, templating, and A/B principles apply to prompt management.
  • Implement a Prompt Versioning System with Git and Jinja2
    • Build a prompt versioning system using Git branches and Jinja2 templates, then run a basic A/B experiment comparing two prompt versions.
  • Understand Vector Databases and Similarity Search
    • Learn how vector embeddings, ANN algorithms, and vector database architectures enable efficient semantic search over high-dimensional data.
  • Configure a Vector Database with Chroma and sentence-transformers
    • Configure a local in-process Chroma vector database, ingest chunks embedded with OpenAI text-embedding-3-small, measure recall@5, then compare a local sentence-transformers embedder.
  • Understand Retrieval-Augmented Generation (RAG) Architecture
    • Learn the end-to-end RAG workflow (from query embedding to context-augmented generation) and the role of each component in reducing hallucinations.
  • Operationalize a Basic RAG Pipeline with the OpenAI SDK and Chroma
    • Compose a functional RAG application with the raw OpenAI SDK and Chroma that retrieves relevant documents from a vector store and generates context-aware answers via an LLM API.
  • Understand LLM Observability and Distributed Tracing
    • Learn how distributed tracing concepts apply to LLM applications and what key metrics (latency, token counts, and span hierarchies) reveal about application health.
  • Implement LLM Call Tracing with Arize Phoenix
    • Instrument a RAG application to send full execution traces to Arize Phoenix, then analyze latency and token usage across each pipeline step.
  • Understand LLM Evaluation Methods
    • Learn the metrics and methodologies for evaluating LLM applications: reference-based metrics, judge models, golden sets, and regression detection.
  • Build an Automated Evaluation Suite with RAGAS
    • Build an automated evaluation suite using RAGAS to measure retrieval relevance, answer faithfulness, and answer correctness against a golden test set, with a regression-detection gate.
  • Understand LLM Pricing and Token-Based Cost Management
    • Learn how LLM providers price API calls by token, how to count tokens programmatically, and the strategies used to aggregate and monitor costs at scale.
  • Build a Token Usage and Cost Monitoring System with tiktoken
    • Instrument LLM API calls to capture token usage, log per-call cost data, and build a simple dashboard showing cost trends by day, model, and query type.
  • Understand Semantic Caching for LLM Applications
    • Learn how semantic caching uses vector similarity to serve cached LLM responses for semantically equivalent queries, reducing cost and latency.
  • Implement Semantic Caching with Chroma
    • Build a semantic caching layer that intercepts LLM calls, checks a vector store for similar cached queries, and stores new results: measurably reducing API calls.
  • Understand High-Performance LLM Inference Servers
    • Learn why specialized inference servers like vLLM outperform simple API wrappers and how techniques like PagedAttention and continuous batching maximize GPU throughput.
  • Understand LLM Gateway Architecture and API Routing Patterns
    • Learn how an LLM gateway centralizes model access, enforces organizational policies, secures API keys, and provides load balancing across multiple model providers.
  • Build an LLM Gateway with FastAPI
    • Build a FastAPI-based LLM gateway that routes queries across model tiers, adds retry with exponential backoff, and abstracts a second provider behind a stubbed adapter.
  • Understand LLM Safety and Guardrails
    • Learn the safety risks specific to LLM applications (prompt injection, PII exposure, jailbreaks, and unsafe outputs) and the architectural patterns used to mitigate them.
  • Implement Input/Output Guardrails with LLM Guard
    • Build input and output guardrails using LLM Guard to detect prompt injection, redact PII, and validate LLM output for safety and format compliance.
  • Understand A/B Testing Principles for LLM Prompt Optimization
    • Learn how controlled experiments, statistical significance, and feature flagging are applied to evaluate competing prompt variants on real production traffic.
  • Conduct Prompt A/B Tests with Python Feature Flags
    • Implement a traffic-splitting system that routes live requests to two prompt variants, collects performance metrics, and applies a chi-squared test to determine the winner.
  • Understand RAGOps and Automated Data Pipeline Architecture
    • Learn how event-driven pipeline design and orchestration tools keep a RAG knowledge base current through automated, idempotent ingestion workflows.
  • Automate the RAG Data Pipeline (local watcher + blue/green index swap)
    • Build an event-driven RAG pipeline with a filesystem watcher and an atomic blue/green index swap that chunks, embeds, and upserts new documents, with S3 + Lambda as the production analogy.
  • Understand RAG Latency Bottlenecks and Optimization Strategies
    • Learn how to profile each component of the RAG pipeline, identify latency bottlenecks, and apply optimization strategies including token streaming and hardware acceleration.
  • Optimize End-to-End RAG Latency with Streaming and Profiling
    • Profile a RAG pipeline to locate latency bottlenecks, measure token streaming against a blocking endpoint, and run a vector-search tuning sweep, measuring improvement at each step.
  • LLM Ops Course Review
    • Recap of all LLM Ops skills covered, key takeaways, and guidance on next steps in the AI Ops Engineer program.
  • Project: Production LLM FAQ Service
    • Build a retrieval-augmented FAQ service for a fictional e-commerce company, covering the full LLM Ops stack across nine graded deliverables.

Taught by

Jeff Chen

Reviews

Start your review of LLMOps: RAG Pipelines, Evaluation, and Cost Control

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.