Context Engineering in 2026 — Louis-François Bouchard, Omar Solano & Samridhi Vaid, Towards AI
Aug 17, 2026 · 1:03:26
Louis-François Bouchard, Omar Solano, and Samridhi Vaid of Towards AI test context engineering on their AI tutor: keeping full history beat every compaction technique on recall, cost, and latency because 97% of tokens were served from cache (up to 50x cheaper), making summarization a trap unless it shrinks context by more than 50x. Full history recovered specific details 95% vs 32% after summarizing, and distinctive facts survived 800k tokens. Local hardware changes it: a 32k window can't keep everything, and larger models don't widen context. Dense retrieval hit 0% recall at 400k tokens where BM25 got 100%, so they use hybrid retrieval. Rule: name the constraint before compacting.