A product discussed on AI Engineer.

Deep dive on LLM Inference at Scale — Harshul Jain, Audible & Tanmay Sah, Independent AI Researcher
Sep 8, 2026 · 1:28:12
Harshul Jain of Audible and independent AI researcher Tanmay Sah teach a two-hour workshop building LLM inference optimization from first principles on Mistral 7B. They show KV cache costs 131 KB per token, so 16,000 tokens across 80 concurrent users needs 42 GB of GPU memory, driving the pain points of memory, TTFT, and throughput. Sah covers quantization, multi-head, multi-query, grouped-query, and latent attention, plus flash attention, using his 'ostrich' and 'world cup' teaching algorithms. Jain covers paged attention, continuous batching, prefix caching, and KV quantization, benchmarked against a Hugging Face baseline on vLLM. Their benchmarking found vLLM and SGLang statistically tied on standard workloads, but SGLang three to four times faster once agentic branching enters.

Introduction to LLM serving with SGLang - Philip Kiely and Yineng Zhang, Baseten
Jul 26, 2025 · 43:42
Philip Kiely and Yineng Zhang introduce SGLang, an open-source fast serving framework for LLMs and VLMs, arguing it offers production-ready performance with day-zero support for new models like DeepSeek and Qwen, and a customizable codebase for contributions. Yineng, a core maintainer, traces SGLang's rapid growth from a December 2023 paper to nearly 15,000 GitHub stars and adoption by xAI, AMD, and Meituan. The workshop demonstrates deploying a first model via Baseten's trunk packaging, then tuning the CUDA Graph max batch size flag on an L4 GPU to maintain CUDA Graph-enabled decoding during higher concurrency, boosting generation throughput. They also cover Eagle 3 speculative decoding, where a draft model derived from the target model speculates tokens; users can benchmark different step and top-k configurations on representative prompts to find optimal settings for production. Finally, they invite contributions via GitHub's 'good first issue' tags and highlight Baseten's job openings.
Powered by PodHood