Attack generative and agentic AI systems the way an adversary would, then harden the same systems against the attacks you just ran. You will jailbreak a commercial assistant, plant indirect injections in content a pipeline ingests, hijack an agent through its own data, and hide instructions inside images, then build the hardened system prompts, guardrails, RAG controls, agent boundaries, structured logging, and human-in-the-loop gates that stop them. The course closes with a capstone that red-teams and hardens a RAG-enabled research agent against the OWASP LLM Top 10.
Overview
Syllabus
- Welcome to Generative and Agentic AI Security
- In this lesson, you'll learn what this course covers, what you need to know going in, and how to get the browser-based labs running.
- Offensive Prompt Engineering
- In this lesson, you'll name the techniques behind a prompt attack and learn how to judge a test result when an assistant only partly gives way.
- Execute Offensive Prompt Engineering Attacks
- In this lesson, you'll write your own injection prompts against a deliberately vulnerable banking endpoint and learn how to judge whether each attempt succeeded.
- Understand the OWASP LLM Top 10
- In this lesson, you'll map an application's features onto the OWASP Top 10 for LLMs and learn how to rank findings by what they would cost.
- Conduct an OWASP LLM Top 10 Risk Assessment
- In this lesson, you'll audit a vulnerable banking endpoint against five OWASP categories and learn how to write findings with evidence, a severity, and a fix.
- Understand Indirect Prompt Injection
- In this lesson, you'll find where hidden instructions enter the content a model reads and learn how to tell indirect injection from a direct attack.
- Simulate an Indirect Prompt Injection Attack
- In this lesson, you'll run an indirect prompt injection against an internal assistant and learn how to block a poisoned document before it reaches the model.
- Understand Defensive System Prompting and Guardrailing
- In this lesson, you'll name the three patterns that harden a system prompt and learn how to test guardrails against attacks and ordinary requests together.
- Implement Defensive System Prompts and Guardrails with Python
- In this lesson, you'll write three attacks against an internal assistant, harden its system prompt, and learn how to find where prompt-only defenses give way.
- Understand Securing Retrieval-Augmented Generation Architecture
- In this lesson, you'll trace how RAG moves the attack surface to the document store and learn how to order the controls in the retrieval pipeline.
- Secure a RAG System with Access Controls and Data Validation
- In this lesson, you'll screen documents retrieved from a knowledge base and learn how to quarantine a poisoned entry before it reaches the model.
- Understand Secure Multi-Agent System Design
- In this lesson, you'll split a multi-agent pipeline's duties across three roles and learn how to enforce each role with a card your code checks.
- Design a Secure Multi-Agent System with Python
- In this lesson, you'll define an agent's role in code and learn how to reject requests that fall outside it.
- Understand Agent Monitoring and Incident Response
- In this lesson, you'll name the fields that let you reconstruct an agent incident and learn how to work through containment, investigation, and remediation in order.
- Implement Agent Monitoring and Incident Response with Python
- In this lesson, you'll add structured logging to an assistant and learn how to record every interaction, including blocked attacks, as a single JSON line.
- Understand Managing Agentic Risk with Human-in-the-Loop
- In this lesson, you'll rate an agent's actions by risk and learn how to decide which of them a gate sends to a human reviewer.
- Implement a Human-in-the-Loop Workflow for an LLM Agent with Python
- In this lesson, you'll build a risk gate for an assistant and learn how to route high-risk requests to a human reviewer.
- Understand Agentic AI Task Hijacking
- In this lesson, you'll find the hijacked step inside an agent's plan and learn how to check a plan for sub-tasks the goal never authorized.
- Execute an Agentic AI Task Hijacking Attack
- In this lesson, you'll hijack an agent's workflow with a hidden instruction and learn how to limit it to a list of approved actions.
- Understand Multimodal Injection Attacks
- In this lesson, you'll explain how an instruction hidden in an image reaches a model and learn how to show a person what the machine read.
- Craft a Multimodal Prompt Injection Attack
- In this lesson, you'll strip injected instructions out of text extracted from an uploaded file and learn why extraction is the place to clean it.
- Understand Instruction Guardrails for LLMs
- In this lesson, you'll write instruction guardrails as labeled rules and learn how to test them with a matrix that shows which rule each case exercises.
- Implement Instruction Guardrails for an LLM with Python
- In this lesson, you'll write a third named rule for an internal assistant's system prompt and learn how to test that the rule changed its behavior.
- Understand Segregating External Content in Prompts
- In this lesson, you'll assemble a prompt that uses XML-style tags to separate your rules from untrusted content, and learn how an attacker escapes that container.
- Implement External Content Segregation in Prompts with Python
- In this lesson, you'll rebuild Aria's flat prompt as separate system and user messages, and learn how to test that boundary against an injection attempt.
- Project: Northstar Research Agent: Red Team and Harden a RAG-Enabled AI Agent
- Test a RAG research agent for security weaknesses, then harden it by identifying attacks, analyzing risks, and applying defenses to improve safety.
Taught by
Kevin Carter