Scaling up Continual Learning — Ronak Malde, Trajectory
Aug 12, 2026 · 23:03
Ronak Malde, founder of Trajectory and former Windsurf research lead, explains how on-policy self-distillation (OPSD) scales continual learning beyond GRPO's limits. He argues GRPO requires parallel rollouts and collapses feedback into one sequence-level score, like being handed 87 out of 100 on an essay. OPSD instead matches per-token log probs between a student and a teacher given privileged hints, optimizing the entire vocabulary and reducing tokens to solve tasks. Scaling to 120B models with 100+ tool calls reveals the 'but wait' problem—models hedge into 'maybe'—solved via step-level KL weighting, and hint leakage, countered with residual guidance. Trajectory applies this to production agent traces for Harvey, Decagon, and Rogo.