A product discussed on AI Engineer.

Training Frontier Models to Out-Think Hackers — Uri Rolls, Arithmetic & Thom Wolf, Hugging Face
Jul 24, 2026 · 17:28
Uri Rolls of Arithmetic and Thom Wolf of Hugging Face argue that frontier models can be trained to out-think hackers, not just pattern-match vulnerabilities, by focusing on logic leaps in access control. Their benchmark, Mask Off, builds blackbox environments from real zero days found by human researchers, testing whether models can chain reconnaissance into exploitation. A live example: a Keycloak check validates admin by name while another checks by ID, so renaming oneself to the admin inherits privilege—GPT-5.5 and Opus probe everything but never make that logical leap. Results are brutal: only one solve at K1, with GPT-5.5 alone succeeding at K5. Rolls and Wolf argue this mirrors ARC-AGI’s challenge—models struggle to build dynamic world models—and that high-quality data and open source models can shift the economics of cyber defense, giving defenders a lasting speed advantage.

Measuring AGI: Interactive Reasoning Benchmarks for ARC-AGI-3 — Greg Kamradt, ARC Prize Foundation
Jul 16, 2025 · 18:28
Greg Kamradt, President of ARC Prize Foundation, introduces ARC-AGI-3, the first interactive reasoning benchmark for AGI that drops agents into novel games without prior instruction, forcing exploration to solve tasks. Unlike static tests, this benchmark measures skill acquisition efficiency—how quickly an AI learns and applies new skills—using human baselines from 400+ in-person tests. It strips away language and trivia, relying only on core knowledge priors (basic math, geometry, agentness, objectness). A public training set of ~40 games will be released, but performance is measured on a private evaluation set of 120 games unseen by developers or AI. Kamradt asserts that as long as AI cannot outperform humans on these problems, we do not have AGI; a sandbox preview with five games and a mini agent competition is planned for next month, with full launch in Q1 2026.
Powered by PodHood