Courses from 1000+ universities
AI got cheap enough that Duolingo’s most expensive plan may not survive it. I read the earnings call transcript and opened the app to see what is actually changing for learners.
600 Free Google Certifications
Artificial Intelligence
Data Science
Cybersecurity
L'Italiano nel mondo
Introduction to HTML5
Umano Digitale
Organize and share your learning with Class Central Lists.
View our Lists Showcase
Learn LLM Evaluation, earn certificates with paid and free online courses from Northeastern University and other top universities around the world. Read reviews to decide if a class is right for you.
Practice benchmarking LLMs on question answering and text generation with automated metrics, logprobs, perplexity, and behavioral tests.
Evaluate and optimize large language models: run statistical significance tests, diagnose hallucinations from logs, track experiments with DVC and W&B, and cut LLM operational costs.
Evaluate large language models with ROUGE, GLUE, SuperGLUE, and BIG-bench, and build summarization, chatbot, and sentiment analysis demos using LangChain and ChromaDB.
Learn practical techniques for evaluating large language models across question answering and text generation tasks. Compare model behavior using reference-based metrics, semantic similarity, log prob
Discover how to build structured, data-driven frameworks for evaluating LLMs in educational products using synthetic data generation and validation pipelines.
Explore Japan's largest LLM evaluation platform and learn how W&B's Nejumi Leaderboard evolved from initial version to v4, sharing operational insights and future prospects.
Benchmark large language models with QA tasks, prompting strategies, fuzzy matching, and smart scoring to compare model performance.
Evaluate LLM fluency and sentence likelihood using token log probabilities and perplexity.
Use lightweight API experiments to investigate LLM token efficiency, temperature sensitivity, output consistency, and hallucination detection.
Test AI Agents, Chatbots & RAG Apps with DeepEval using Metrics, Tracing, Goldens, Safety Evals & G-Eval Custom Metrics
Evaluate LLM outputs with foundational metrics, Vertex AI Automatic Metrics and AutoSxS, and human judgment, and anticipate evaluation trends across text, image, and audio generation.
Bring scientific rigor to AI: score LLM outputs with BLEU, ROUGE-L and cosine similarity, then compare models through A/B tests, p-values and confidence intervals.
Build an adversarial safety test suite for LLM outputs with pytest, then harden it using mutation testing with mutmut to expose gaps in your safety net.
Evaluate LLM outputs with BLEU, ROUGE-L and cosine similarity, trace hallucinations through retrieval logs, prove gains with A/B tests, and tune SQL and vector search.
Build evaluation systems for generative AI that measure relevance and accuracy, compare models, and support continuous quality improvement.
Get personalized course recommendations, track subjects and courses with reminders, and more.