This course equips you with essential skills to identify and exploit vulnerabilities in AI systems. You will explore the fundamentals of AI red teaming, including theoretical and practical applications of evasion attacks, data poisoning, prompt injection, and vector database attacks. The course also covers advanced topics such as model inversion and quantitative robustness testing, ensuring a comprehensive understanding of AI security threats. You will gain hands-on experience through real-time applications and a capstone project focusing on AI red-teaming strategies to enhance security measures and safeguard against adversarial tactics.
Overview
Syllabus
- Introduction: Offensive & Adversarial AI Security
- Get oriented: what offensive AI security covers, the seven attack families you will run yourself, and the dual-system red team engagement you deliver as the capstone.
- Understand AI Red Teaming
- Learn why AI red teaming targets model behavior rather than code, walk the three-stage lifecycle, and leave able to scope an engagement and write a charter developers can act on.
- Apply AI Red Teaming
- Use an LLM to surface attack vectors from real system documentation, then turn the ones evidence supports into a red team charter with scope, rules of engagement, and success criteria.
- Understand Evasion Attacks
- See how attackers fool an image classifier with changes too small to notice, with and without model access, and why a high score on clean test data proves nothing about robustness.
- Apply Evasion Attacks
- Build black-box and white-box attacks against a real classifier with ART, measure both with adversarial accuracy and perturbation size, and turn those numbers into a risk assessment.
- Understand Data Poisoning
- Learn how flipping labels or planting a trigger corrupts a model at training time, and why a backdoored model passes every accuracy check you would normally run before shipping.
- Apply Data Poisoning
- Train a clean baseline, poison it with label flips and a backdoor trigger, measure what each attack costs the model, and find which defender-side check actually catches which attack.
- Understand Prompt Injection
- Learn why one sentence of user text can override an application's rules: how a model reads system versus user prompts, and the three techniques attackers reach for once filters go up.
- Apply Prompt Injection
- Build a prompt injection payload suite, run it across models and system prompt configurations, score each compromise, and test whether a prompt-level defense moves the number at all.
- Understand Vector Database Attacks
- Follow how RAG turns documents into embeddings and ranks them by meaning, then how an attacker hijacks that ranking to control what the model tells every user who asks a matching question.
- Apply Vector Database Attacks
- Stand up a RAG pipeline with live embeddings, plant poisoned documents tuned to real questions, measure the ranking shift they cause, and test whether provenance filtering stops them.
- Understand Model Inversion
- Learn how a model's own confidence scores let an attacker rebuild the data it trained on, why overfitting makes it worse, and how much detail an inference API should ever return.
- Apply Model Inversion
- Reconstruct a face from a classifier's confidence outputs alone, then switch sides and measure how far leakage drops across full probabilities, rounded scores, and a bare label.
- Understand Quantitative Robustness Testing
- Treat robustness as a number you can benchmark the way you benchmark accuracy: the metrics that measure it, the frameworks that automate it, and how to gate a release on it.
- Apply Quantitative Robustness Testing
- Build an automated attack suite that scores a model under clean, environmental, and adversarial conditions, then compare two variants and judge which one clears a release gate.
- Understand AI Supply Chain Security
- Map the stack you inherit — model weights, frameworks, libraries, base images, cloud — and see how a flaw four layers deep in a package you never named runs with your credentials.
- Apply AI Supply Chain Vulnerability Scanning
- Scan AI container images with Trivy, parse the report and SBOM into structured findings, and build a priority score ranked by real exposure, where a MEDIUM can rightly outrank a HIGH.
- Project: AI Red-Teaming
- Run a full engagement against two production AI systems: five attacks spanning evasion, poisoning, injection, data exfiltration, and supply chain, delivered as a CISO-ready report.
Taught by
Josh Kalin