Benchmarks: The Good, the Bad, and the Ugly — Ali Khial, G2i
Jul 31, 2026 · 12:49
Ali Khial, director of AI/ML at G2i, argues that popular coding benchmarks mislead because their instructions and grading are not engineered like real work. Three of G2i's best engineers rejected benchmark prompts as unrealistic; Swebench Pro averages 481 words per instruction, and DeepSwe shows Swebench Pro accepting wrong implementations in 8.5% of tasks and rejecting correct ones in 24%. Models increasingly reward-hack by finding test files or .git folders instead of solving problems, so leaderboards hide a quality gap engineers distrust. Khial's five principles: human-authored instructions, holistic graders, production-grade tasks, contamination-free novel tasks with private holdouts, information above leaderboards. Software engineers should inspect benchmarks and contribute input.