A company discussed on AI Engineer.

When Will The Benchmaxxing Plague End? — Nick Heiner, Surge AI
Aug 2, 2026 · 17:25
Nick Heiner of Surge AI argues that benchmaxxing — labs gaming benchmarks rather than improving real-world value — is driven by benchmark misalignment and poor methodology, and can end with rigorous human evaluation. He identifies key antipatterns: broken tasks, contamination, reward hacking, and mismatched prompts and verifiers. He cites IF eval’s impossible prompts like 'repeat this verbatim' plus 'translate into Hindi', and LMArena being gameable via watermarked crowd voters. He shows evidence that Anthropic’s Opus 4.8 memorized much of SWE-bench verified without disclosing it. Surge’s Hemingway Bench uses thousands of professional writers for blind model comparisons, since LLM judges lack taste, and he urges benchmark makers and labs to adopt QC, private holdouts, and aligned verifiers.

State of Data — Sean Cai, Independent / State of Data
Jul 26, 2026 · 18:22
Sean Cai argues data markets, not compute, are now the binding constraint for turning generalist AI into expert systems, with the supply chain unbundling from vertical giants into specialist vendors. He distinguishes Type I (real workflow capture) from contrived Type II data, noting the industry sells Type II as Type I. He introduces Verifier's Law—ease of training proportional to task verifiability—to predict domain maturity: code first, then biology, security, finance. He exposes benchmark psychosis: a single number under one scaffold is a noisy sample, requiring cross-harness differencing. Labs like Anthropic's data spending forecasts product launches (e.g., cybersecurity data in January led to Claude Cyber in March). Data companies like Mercor pivot to enterprise, and Cai builds Antikythera mechanisms to monetize real-world workflows and provide RL-as-a-service.
Powered by PodHood