Overview
Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
In this Specialization, you’ll learn to build and deploy generative AI applications with Hugging Face. You’ll work with models, datasets, and Spaces; run inference with the Pipeline API; prepare text with AutoTokenizer; load models with AutoModel; evaluate model cards; and apply model selection and responsible-use checks.
You’ll then turn models into interactive applications with Gradio. You’ll load and preprocess datasets, fine-tune transformer models with the Trainer API, evaluate results, publish model cards, build interfaces with gr.Interface and gr.Blocks, create streaming multi-turn chatbots with gr.ChatInterface, and deploy applications to Hugging Face Spaces. You’ll also configure hardware and secrets, consider cost and performance, and query deployed apps with the Gradio Python client.
The final course extends these skills to multimodal and agentic AI. You’ll use CLIP and vision-language models for visual question answering, image captioning, and document understanding, then work with Whisper, Diffusers, LoRA, multimodal RAG, smolagents, and MCP. You’ll also apply safety filtering, adversarial testing, and failure-mode documentation for responsible deployment.
Syllabus
- Course 1: Building AI Apps with Hugging Face Spaces and Gradio
- Course 2: Getting Started with Hugging Face Transformers
- Course 3: Introduction to Multimodal AI with Hugging Face
Courses
-
By the end of this course, you will be able to: • Explain the role of models, datasets, and Spaces in the HF ecosystem and use the pipeline API to run inference across text, vision, and audio tasks. • Tokenize and encode text inputs using AutoTokenizer, handle padding and truncation, and apply chat templates for LLM-compatible formatting. • Load pre-trained models using the appropriate AutoModel class, inspect model configuration, run manual inference, and load models in reduced precision with device_map="auto". • Evaluate model cards to assess intended use, limitations, bias disclosures, and license compatibility before recommending a model for deployment. Go from zero to confident model evaluation in four hours. All you need is basic Python — no machine learning or Hugging Face experience required. The course opens with a realistic challenge: your VP needs an AI feasibility assessment by Thursday, and the Hugging Face Hub has over 2 million models to choose from. You'll build a systematic approach to navigating that ecosystem, using filters, model cards, and task categories to find the right model instead of guessing. Run inference across text, vision, and audio tasks with the pipeline API, then go deeper: learn how tokenizers convert raw text into the numerical inputs models actually process, debug why a classifier fails silently on long messages, and discover how chat templates turn a language model into a conversation partner. Load models manually with AutoModel classes, inspect their configuration, and manage memory with reduced precision. The course closes with a hands-on model selection challenge: three candidate models, one task, and you have to decide which one ships — backed by model card evidence, not gut instinct.
-
By the end of this course, learners will be able to: • Load and preprocess HF Hub datasets, fine-tune a pre-trained model with the Trainer API, compute evaluation metrics, and push the result to the Hub with a model card. • Build interactive AI applications using gr.Interface and gr.Blocks with multi-component layouts, conditional visibility, session state, and event listeners • Build a streaming multi-turn chatbot using gr.ChatInterface with an LLM backend and extend it with real-time inference workflows. • Deploy a Gradio app to HF Spaces, configure hardware and secrets, evaluate cost vs. performance trade-offs, and query the deployed app programmatically using the Gradio Python client. A model stuck in a notebook is a model nobody uses. Some familiarity with the HF Transformers library and pipeline API will help you hit the ground running. The course starts where most tutorials stop — with the data. Work through a realistic fine-tuning scenario: the off-the-shelf classifier isn’t cutting it for your domain, so you’ll load a dataset from the Hub, preprocess it with the right tokenization strategy, configure the Trainer API, evaluate with real metrics, and publish the result with a model card. Once you have a model that works for your domain, the next question is: how do people use it? Wrap it in a Gradio app, graduate from quick prototypes with gr.Interface to structured applications with gr.Blocks, and add streaming chatbot behavior with gr.ChatInterface — including diagnosing why a chatbot demo feels broken the day before a client presentation. Deploy everything to Hugging Face Spaces, configure ZeroGPU when the budget won’t cover dedicated hardware, and turn your app into a programmable API endpoint. By the end, you’ll have a fine-tuned model and a live, deployed application that other systems can call.
-
By the end of this course, you will be able to: • Explain how CLIP aligns image and text in a shared embedding space, use VLMs to perform visual question answering, image captioning, and document understanding, and navigate the Hub for multimodal models. • Build a pipeline that transcribes audio with Whisper and generates images with Diffusers, and describe how LoRA fine-tuning and multimodal RAG extend VLM capabilities. • Build an agentic workflow using smolagents with VLM support and MCP tool integration to automate multi-step tasks requiring vision and reasoning. • Apply ShieldGemma 2 to filter inputs and outputs of a VLM pipeline, test against adversarial inputs, and document failure modes for responsible deployment. AI that can only read text is already behind. This intermediate course assumes you're comfortable with the HF Transformers library and basic Gradio development. It opens with a practical challenge: 2,000 products with photos but no descriptions, and a stack of invoice PDFs that need structured data extraction. You’ll learn how CLIP aligned images and text in a shared space, then use modern vision-language models to caption products, answer questions about charts, and pull fields from invoices. Go wider: transcribe customer calls with Whisper, generate images from text briefs with Diffusers, and learn when to fine-tune a model versus when to give it better context through retrieval. Build agent workflows that can see screenshots, reason about what’s on screen, and connect to external tools through the Model Context Protocol (MCP) to act on what they find. The course closes with a deployment readiness review: your CTO wants to launch the AI pipeline next week, and you need to decide whether it’s safe to ship — with safety filtering, adversarial testing, and documented failure modes backing your recommendation.
Taught by
Hugging Face