Your Agent Evolved. Your Evals Didn't. — Ameya Bhatawdekar, Braintrust
Aug 20, 2026 · 24:13
Ameya Bhatawdekar, Field CTO at Braintrust, argues that evals are the durable asset across AI replatformings and must move with each architectural generation. He traces five generations—single prompt, retrieval chain, ReAct loop, workflow graphs, and the return of loops with newer models—grounded in a site reliability agent, showing how each added failure surfaces: parsers, retrievers, branch logic, node contracts, classifier misfires, and trajectory variance. New metrics like pass@k (succeeds at least once in k attempts) measure capability, while pass^k (how many of k succeed) measure reliability. He urges harvesting production data into evals via the flywheel, and says Braintrust's Topics clusters production data to surface unanticipated failure modes.