Scaling to Long Horizons — Ross Taylor & Chengxi Taylor, General Reasoning
Jul 31, 2026 · 18:07
Ross Taylor and Chengxi Taylor, co-founders of London-based RL company General Reasoning, argue that scaling to long horizons demands a mindset shift and better simulation, not just bigger context windows. Ross recounts how Galactica's 2022 thinking tokens were prescient, but its base-model demo backfired, while RLHF made LLMs products; his Meta team's PPO with verifiable rewards worked, but DeepSeek-R1 later showed better base models were key. Chengxi details long-horizon obstacles: sparse rewards, credit assignment, and 1M token limits, solved by value models that reduce variance and enable bootstrapping. Kelly Bench gave frontier models $100K to trade Premier League matches; all lost money, exposing how little current environments simulate real competition. Pipeline RL trades off off-policy staleness (8 steps okay) against GPU utilization, and they point to openreward.ai, with 350+ environments, for long-horizon RL.