How Fast Can Your Model Composition Run in Serverless Inference?
CNCF [Cloud Native Computing Foundation] via YouTube
Build the Finance Skills That Lead to Promotions, Not Just Certificates
Start speaking a new language. It’s just 3 weeks away.
Overview
Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
The session presents a RAG application that combines an LLM, an embedding model, and OCR for inference in serverless Kubernetes. It explains how BentoML, Dragonfly’s peer-to-peer network, and other open-source technologies support model packaging, distribution, and rapid deployment.
Syllabus
How Fast Can Your Model Composition Run in Serverless Inference? - Fog Dong, BentoML & Wenbo Qi
Taught by
CNCF [Cloud Native Computing Foundation]