What you'll learn:
- Explain the physics, phonetics, and acoustic features that underlie human speech production
- Compare traditional cascade TTS pipelines with modern Speech Language Model (SpeechLM) architectures
- Build and apply neural audio codecs and semantic tokenization (EnCodec, HuBERT, wav2vec 2.0, RVQ) (
- Implement autoregressive codec-based TTS with multi-stream token decoding strategies
- Design unified speech-text models with cross-modal alignment and paralinguistic control
- Apply latent diffusion and conditional flow matching to generate high-quality mel-spectrograms
- Evaluate flow matching vs. diffusion trade-offs for speed, quality, and controllability
- Deploy low-latency, streaming agentic TTS systems with real-time interruption handling
Generative voice AI has moved far beyond simple text-to-speech — and this course takes you from the physics of sound all the way to building production-grade, agentic voice systems.
Most TTS courses stop at basic vocoders or off-the-shelf APIs. This one goes deeper. You'll start with the fundamentals of human speech — acoustics, phonetics, and prosody — before diving into the architectures actually powering today's state-of-the-art voice models: self-supervised representation learning (wav2vec 2.0, HuBERT), neural audio codecs (EnCodec, SoundStream, DAC), and the tokenization strategies that let LLMs "speak."
From there, you'll master the two dominant modern paradigms — autoregressive codec-based TTS and latent diffusion / conditional flow matching — understanding exactly when and why each is used in real systems. You'll also explore unified speech-text models, paralinguistic modeling (laughter, breathing, affect), and zero-shot voice cloning.
By the final module, you'll understand how to build low-latency, streaming, agentic voice pipelines — the same techniques behind real-time conversational AI agents — covering chunked inference, speculative decoding, WebSocket streaming, and turn-taking.
What you'll learn:
The science of speech production and acoustic feature extraction
How neural audio codecs and semantic tokenization work
Autoregressive and diffusion/flow-based TTS architectures
Cross-modal speech-text alignment techniques
Building low-latency, interruption-aware conversational voice agents
Whether you're an ML engineer, researcher, or voice-tech founder, this course gives you the complete architectural picture — from tokens to agents.