- Technology
- Computer Science
- Distributed Systems
- High Performance Computing
- Parallel Computing
- GPU Programming
- CUDA
- Technology
- Cloud Computing
- Amazon Web Services (AWS)
- AWS Containers
- Amazon Elastic Kubernetes Service (EKS)
- Technology
- Computer Science
- Distributed Systems
- High Performance Computing
- Parallel Computing
- GPU Computing
Scalable LLM Inference on Kubernetes With NVIDIA NIMS, LangChain, Milvus and FluxCD
Master Production-Ready Machine Learning, Step by Step
You’re only 3 weeks away from a new language
Overview
Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
Explore architecting and implementing a scalable LLM inference service on Amazon EKS in this 33-minute conference talk from the Linux Foundation. Dive deep into workload orchestration using Kubernetes as the foundation while integrating NVIDIA NIMS for optimal GPU utilization, LangChain for flexible LLM operations, and Milvus for efficient vector storage. Learn how to leverage FluxCD for GitOps-driven deployments, implement Karpenter for horizontal scaling, and establish comprehensive observability with Prometheus and Grafana. Discover best practices for building production-ready large language model inference systems that can scale effectively in cloud-native environments, combining cutting-edge AI technologies with robust Kubernetes orchestration patterns.
Syllabus
Scalable LLM Inference on Kubernetes With NVIDIA NIMS, LangChain, Milvus and Flu... Riccardo Freschi
Taught by
Linux Foundation