- Technology
- Computer Science
- Distributed Systems
- High Performance Computing
- Parallel Computing
- GPU Programming
- CUDA
- Technology
- Cloud Computing
- Amazon Web Services (AWS)
- AWS Containers
- Amazon Elastic Kubernetes Service (EKS)
- Technology
- Generative AI
- Large Language Models (LLMs)
- AI Models
- Meta
- LLaMA (Large Language Model Meta AI)
- Technology
- Computer Science
- Distributed Systems
- High Performance Computing
- Parallel Computing
- GPU Computing
Deploying and Scaling Large Language Models with NVIDIA NIM on Amazon EKS
AWS Events via YouTube
Learn Backend Development Part-Time, Online
Learn AI, Data Science & Business — Earn Certificates That Get You Hired
Overview
Google, IBM & Meta Certificates – 40% Off
One plan covers every Professional Certificate on Coursera.
Unlock All Certificates
Learn to deploy and scale large language models like Llama3/Mistral7b on Kubernetes through this 19-minute technical video demonstrating NVIDIA Inference Microservices (NIM) implementation on Amazon EKS. Master essential skills including GPU-enabled EKS cluster configuration, Kubernetes scaling strategies, and efficient deployment using NVIDIA's NIM Helm chart. Explore real-time benchmarking techniques with GenAIPerf while gaining practical insights into monitoring costs and performance metrics. Designed for ML engineers and cloud architects, gain hands-on experience through a live demonstration showcasing best practices for cost-effective LLM deployment in production environments on AWS infrastructure.
Syllabus
Why and how to run NVIDIA NIM on Amazon EKS
Taught by
AWS Events