A product discussed on AI Engineer.

Deep dive on LLM Inference at Scale — Harshul Jain, Audible & Tanmay Sah, Independent AI Researcher
Sep 8, 2026 · 1:28:12
Harshul Jain of Audible and independent AI researcher Tanmay Sah teach a two-hour workshop building LLM inference optimization from first principles on Mistral 7B. They show KV cache costs 131 KB per token, so 16,000 tokens across 80 concurrent users needs 42 GB of GPU memory, driving the pain points of memory, TTFT, and throughput. Sah covers quantization, multi-head, multi-query, grouped-query, and latent attention, plus flash attention, using his 'ostrich' and 'world cup' teaching algorithms. Jain covers paged attention, continuous batching, prefix caching, and KV quantization, benchmarked against a Hugging Face baseline on vLLM. Their benchmarking found vLLM and SGLang statistically tied on standard workloads, but SGLang three to four times faster once agentic branching enters.

KV Cache-Aware Routing and P/D Disaggregation on Kubernetes — Yuchen Fama & Ashish Kamra, Red Hat
Aug 27, 2026 · 21:48
Yuchen Fama and Ashish Kamra of Red Hat argue that public inference benchmarks hide the chaotic reality of agentic workloads, and show how KV cache-aware routing plus prefill/decode disaggregation in the open source llm-d framework tackles it. Red Hat's traces show agentic sessions running from a few turns to 3,000, cache hit rates over 90%, and input-output token ratios past 100:1, making a 10x cost gap between cached and uncached tokens. A live demo shows routing reusing cache on the same pod cutting time from 3 seconds to 1, while a fresh system prompt pays the full 3 again. For prefill/decode disaggregation, they report P99 inter-token latency dropping from roughly 900 milliseconds to about 100 across 16 H100s serving gpt-oss, but note it only wins in the middle concurrency band and requires RDMA or RoCE to move KV caches. Their closing case study runs GLM 5.2 on H200s with three prefill workers to one decode, achieving 4x faster time to first token and 60% more requests.

The Desktop Frontier — Ahmad Osman, Osmantic
Jul 21, 2026 · 18:02
Ahmad Osman, founder of Osmantic, argues that within roughly 18 months (by late 2027) a single RTX 5090 will run intelligence equivalent to GLM 5.2, driven by the Densing Law of increasing impact per parameter. He shows this trend through concrete examples: a 27B-parameter Qwen 3.5 now beats the 405B LLaMA 3, and the same eight RTX 3090s that once struggled with LLaMA 2 can now run 15 parallel Qwen 3.5 agents. Osman presents the Densing Law—every 3.5 months, 50% fewer parameters achieve the same capability—as a systematic pattern, not coincidence. He advocates for sovereign AI: owning your own hardware (like a DGX Station or RTX 5090) gives you control, avoids cloud limitations, and sees hardware appreciate in utility as models become more efficient. He asks why fund cloud data centers when local hardware can run frontier intelligence and grow more valuable over time.

Efficient Reinforcement Learning – Rhythm Garg & Linden Li, Applied Compute
Dec 9, 2025 · 20:19
Rhythm Garg and Linden Li, co-founders of Applied Compute, describe how their company uses efficient reinforcement learning (RL) to specialize large language models for enterprise tasks. They explain that synchronous RL wastes GPU time waiting on straggler samples—99% of arithmetic problems complete in ~40 seconds, but the tail takes 80 more seconds—so they adopt asynchronous pipeline RL. This method dedicates fixed GPUs to sampling and training, allowing continuous inference but introducing stale tokens (up to a tolerated staleness threshold) that require importance ratio corrections. To balance speed and stability, they model the system mathematically: using a roofline-based latency curve for sampling, per-GPU training throughput, and constraints on staleness and KV cache memory. Their simulations, parameterized by response length distributions, reveal an optimal GPU allocation that yields ~60% speedup over synchronous RL while keeping staleness within ML limits. This modeling lets them predict runtime and configure runs without expensive trial-and-error.
Powered by PodHood