Earn Your Business Degree, Tuition-Free, 100% Online!
Learn AI, Data Science & Business — Earn Certificates That Get You Hired
Overview
AI, Data Science & Cloud Certificates from Google, IBM & Meta — 40% Off
One plan covers every Professional Certificate on Coursera. 40% off Coursera Plus Annual.
Unlock All Certificates
Explore cutting-edge open-source optimizations for large language model inference in this 28-minute conference talk from the Linux Foundation. Discover Arctic Inference, an innovative vLLM plugin developed by Snowflake AI Research that delivers comprehensive performance enhancements for LLM inference workloads. Learn about groundbreaking techniques including SwiftKV, Ulysses, Shift Parallelism, Suffix Decoding, and Arctic Speculator that collectively achieve remarkable performance improvements: up to 3.4x faster time-to-first-token (TTFT), 1.7x higher throughput, and 1.75x faster time-per-output-token (TPOT) compared to existing open-source solutions. Understand the technical innovations behind these optimizations and how the pluggable design enables seamless integration into existing vLLM deployments. Gain insights into practical implementation strategies and discover how these powerful tools can enhance inference efficiency and resource utilization in real-world applications, empowering the open-source community with advanced LLM inference capabilities.
Syllabus
Arctic Inference: Open-Source LLM Inference Optimizations at Snowflake - Aurick Qiao, Snowflake
Taught by
Linux Foundation