- Technology
- Computer Science
- Distributed Systems
- High Performance Computing
- Parallel Computing
- GPU Programming
- CUDA
- Technology
- Cloud Computing
- Amazon Web Services (AWS)
- AWS Networking & Content Delivery
- Amazon Elastic Load Balancer
- Technology
- Computer Science
- Distributed Systems
- High Performance Computing
- Parallel Computing
- GPU Computing
Serving Voice AI at $1/hr - Open-source, LoRAs, Latency, Load Balancing
AI Engineer via YouTube
Learn AI, Data Science & Business — Earn Certificates That Get You Hired
Start speaking a new language. It’s just 3 weeks away.
Overview
Syllabus
00:00 Introduction to Gabber and Real-Time AI
02:15 Gabber's Mission for Consumer AI
04:17 The Orpheus Voice Model
05:43 Challenges in Voice Cloning
07:44 Latency Management and "Head of Line Silence"
11:07 Infrastructure for Batch Inference
11:36 Leveraging vLLM and Dynamic Quantization
13:21 Load Balancing with a Consistent Hash Ring
14:17 System Architecture Overview
15:07 Conclusion and Open Source Shout-outs
Taught by
AI Engineer