A product discussed on AI Engineer.

Deep dive on LLM Inference at Scale — Harshul Jain, Audible & Tanmay Sah, Independent AI Researcher
Sep 8, 2026 · 1:28:12
Harshul Jain of Audible and independent AI researcher Tanmay Sah teach a two-hour workshop building LLM inference optimization from first principles on Mistral 7B. They show KV cache costs 131 KB per token, so 16,000 tokens across 80 concurrent users needs 42 GB of GPU memory, driving the pain points of memory, TTFT, and throughput. Sah covers quantization, multi-head, multi-query, grouped-query, and latent attention, plus flash attention, using his 'ostrich' and 'world cup' teaching algorithms. Jain covers paged attention, continuous batching, prefix caching, and KV quantization, benchmarked against a Hugging Face baseline on vLLM. Their benchmarking found vLLM and SGLang statistically tied on standard workloads, but SGLang three to four times faster once agentic branching enters.

State Space Models for Realtime Multimodal Intelligence: Karan Goel
Oct 29, 2024 · 14:26
Karan Goel, founder of Cartesia, argues that state space models (SSMs) are key to real-time multimodal intelligence, offering cheaper, faster, and higher-quality alternatives to transformers for streaming applications like conversational voice and on-device assistants. He contrasts batch intelligence (cloud APIs) with streaming intelligence needed for low-latency tasks, emphasizing that SSMs compress information linearly rather than storing all tokens, enabling efficient long-context processing. Cartesia’s voice generation model achieves instant latency in the data center and is being optimized for on-device Mac and desktop deployment. Goel asserts that compression helps long-context tasks (e.g., 24-hour security footage analysis) more than retrieval, and that SSMs now match transformer quality while scaling better for multimodal data.
Powered by PodHood