Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

Udemy

Mastering Generative Voice AI: From Tokens to Agentic TTS

via Udemy

Overview

Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
Master SpeechLMs, neural audio codecs, diffusion & flow matching to build real-time agentic voice AI systems

What you'll learn:
  • Explain the physics, phonetics, and acoustic features that underlie human speech production
  • Compare traditional cascade TTS pipelines with modern Speech Language Model (SpeechLM) architectures
  • Build and apply neural audio codecs and semantic tokenization (EnCodec, HuBERT, wav2vec 2.0, RVQ) (
  • Implement autoregressive codec-based TTS with multi-stream token decoding strategies
  • Design unified speech-text models with cross-modal alignment and paralinguistic control
  • Apply latent diffusion and conditional flow matching to generate high-quality mel-spectrograms
  • Evaluate flow matching vs. diffusion trade-offs for speed, quality, and controllability
  • Deploy low-latency, streaming agentic TTS systems with real-time interruption handling

Generative voice AI has moved far beyond simple text-to-speech — and this course takes you from the physics of sound all the way to building production-grade, agentic voice systems.

Most TTS courses stop at basic vocoders or off-the-shelf APIs. This one goes deeper. You'll start with the fundamentals of human speech — acoustics, phonetics, and prosody — before diving into the architectures actually powering today's state-of-the-art voice models: self-supervised representation learning (wav2vec 2.0, HuBERT), neural audio codecs (EnCodec, SoundStream, DAC), and the tokenization strategies that let LLMs "speak."

From there, you'll master the two dominant modern paradigms — autoregressive codec-based TTS and latent diffusion / conditional flow matching — understanding exactly when and why each is used in real systems. You'll also explore unified speech-text models, paralinguistic modeling (laughter, breathing, affect), and zero-shot voice cloning.

By the final module, you'll understand how to build low-latency, streaming, agentic voice pipelines — the same techniques behind real-time conversational AI agents — covering chunked inference, speculative decoding, WebSocket streaming, and turn-taking.

What you'll learn:

  • The science of speech production and acoustic feature extraction

  • How neural audio codecs and semantic tokenization work

  • Autoregressive and diffusion/flow-based TTS architectures

  • Cross-modal speech-text alignment techniques

  • Building low-latency, interruption-aware conversational voice agents

Whether you're an ML engineer, researcher, or voice-tech founder, this course gives you the complete architectural picture — from tokens to agents.

Taught by

Vinit Singh

Reviews

4.8 rating at Udemy based on 47 ratings

Start your review of Mastering Generative Voice AI: From Tokens to Agentic TTS

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.