This beginner-level path introduces practical methods for benchmarking large language models on multiple-choice question answering, summarization, and other text generation tasks. You will compare model outputs using fuzzy matching, ROUGE, and embedding-based semantic similarity. You will also examine internal scoring signals such as token log probabilities and perplexity to assess output likelihood and fluency. Guided experiments explore token efficiency, temperature sensitivity, output consistency, prompting strategies, and hallucination detection. This path is designed for learners who want a hands-on foundation in evaluating and comparing language models beyond basic accuracy scores.
Learn Excel and Financial Modeling the Way Finance Teams Actually Use Them
The Most Addictive Python and SQL Courses
Overview
Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
Syllabus
- Benchmark language models on question answering and text generation tasks
- Evaluate generated text with fuzzy matching, ROUGE, and semantic similarity
- Interpret token log probabilities and perplexity scores
- Compare model performance across prompts, tasks, and model versions
- Analyze temperature sensitivity, token efficiency, and output consistency
- Detect hallucinations and other problematic model behaviors