Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

Coursera

Fine-Tune Your Own LLM: LoRA, QLoRA & PEFT

Board Infinity via Coursera

Overview

Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
This intermediate-level course guides you through fine-tuning open-source large language models on consumer GPUs using parameter-efficient techniques, all anchored around building a production-style SaaS customer support agent. You’ll first map out the full LLM training pipeline—pretraining, supervised fine-tuning, and alignment—and see exactly where PEFT methods fit when you’re constrained by memory and cost. You’ll compare full fine-tuning with approaches such as LoRA, QLoRA, adapters, prefix and prompt tuning, and then dive into the mathematical intuition behind Low-Rank Adaptation, understanding how rank, alpha, and target module choice influence parameter count and memory usage. You’ll get hands-on with preparing real-world customer support data (tickets, chats, FAQs), turning it into high-quality instruction–completion pairs and formatting them using common schemas like Alpaca, ChatML, and ShareGPT-style formats, with programmatic validation for training readiness. Using Hugging Face Transformers and PEFT, you will implement LoRA and QLoRA to fine-tune 7B–8B parameter open models on a single 24GB GPU, configuring 4-bit NF4 quantization, tracking memory usage, and managing training stability. You’ll learn to debug and iterate using logs, loss curves, and failure case analysis (hallucinations, refusals, overfitting), and evaluate your models with automated metrics (perplexity, ROUGE, BERTScore) while recognizing their limitations for dialogue quality. Finally, you’ll design preference datasets that encode brand tone, helpfulness, and safety, and apply alignment methods such as DPO and ORPO to steer the behavior of your support agent. You’ll explore advanced PEFT variants like DoRA, rsLoRA, and adapter composition to study quality–efficiency trade-offs. The course closes with practical packaging and deployment: merging adapters when appropriate, quantizing for efficient inference, and deploying your agent using modern inference stacks like vLLM and Hugging Face Inference Endpoints, with basic monitoring to track latency and response quality in a SaaS environment. Disclaimer: This is an independent educational resource created by Board Infinity for informational and educational purposes only. This course is not affiliated with, endorsed by, sponsored by, or officially associated with any company, organization, or certification body unless explicitly stated. The content provided is based on industry knowledge and best practices but does not constitute official training material for any specific employer or certification program. All company names, trademarks, service marks, and logos referenced are the property of their respective owners and are used solely for educational identification and comparison purposes.

Syllabus

  • Planning PEFT for a Support Agent
    • This module introduces the LLM post-training pipeline and shows where parameter-efficient fine-tuning fits in a real SaaS customer support workflow. Learners will compare full fine-tuning, LoRA, QLoRA, adapters, and prompt-based methods while deciding what should be solved with fine-tuning versus retrieval or workflow design. The module also builds practical intuition for why a single 24GB GPU changes model and training choices from the start.
  • Understanding LoRA and QLoRA Mechanics
    • Before training, learners need to understand why LoRA works and why QLoRA is the standard path when VRAM is tight. This module breaks down low-rank updates, rank and alpha choices, target modules such as attention and MLP projections, and the role of 4-bit NF4 quantization and double quantization. By connecting the math to memory use, learners will be ready to make informed configuration decisions instead of copying settings blindly.
  • Preparing Data and Chat Templates
    • In this module, learners turn raw support tickets, chats, and FAQs into training-ready examples for supervised fine-tuning. They will convert between Alpaca, ChatML, and ShareGPT-style schemas, apply the correct chat template for the target model, and validate the dataset before training begins. This step is especially important because schema mismatches and role-format errors are among the most common causes of poor fine-tuning results.
  • Training and Improving the Support Model
    • This module moves into hands-on model adaptation using PEFT, starting with supervised fine-tuning and then extending to preference-based alignment. Learners will run LoRA training with Hugging Face and PEFT, build chosen-versus-rejected preference data, and apply DPO or ORPO to improve tone, safety, and escalation behavior. They will also learn how to diagnose issues from logs and outputs so they can iteratively improve the model rather than stopping at the first training run.
  • Evaluating, Packaging, and Serving Adapters
    • The final module prepares the fine-tuned support agent for real deployment decisions. Learners will evaluate quality with automated metrics and human review, experiment with advanced LoRA-family options such as DoRA and rsLoRA, and decide when to keep adapters separate versus merge them into the base model. The module closes with deployment patterns using vLLM and Hugging Face Inference Endpoints, plus basic monitoring for latency and response quality in SaaS environments.

Taught by

Board Infinity

Reviews

Start your review of Fine-Tune Your Own LLM: LoRA, QLoRA & PEFT

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.