A product discussed on AI Engineer.

How I automate my own job at Hugging Face using agents — Niels Rogge, Hugging Face
Aug 20, 2026 · 20:37
Niels Rogge, machine learning engineer on Hugging Face's community science team, explains how he automated his own job: agents ask researchers to publish paper artifacts on the Hub instead of Google Drive, Dropbox, or Zenodo. The outreach is a deterministic nightly workflow on GitHub Actions cron jobs with LangFuse tracing; follow-ups run as a fully autonomous Claude Agents SDK loop, with each GitHub issue in its own Modal container using Bash and one Hugging Face CLI skill. Thousands of issues have drawn only two negative replies, and he doesn't disclose the bot because recipients reply to the same messages he used to send. He argues open models like GLM 5.2 can replace closed ones, and cites his Daily Papers X account (90,000 followers) and a Papers with Code revival at paperswithcode.co.

Scaling to Long Horizons — Ross Taylor & Chengxi Taylor, General Reasoning
Jul 31, 2026 · 18:07
Ross Taylor and Chengxi Taylor, co-founders of London-based RL company General Reasoning, argue that scaling to long horizons demands a mindset shift and better simulation, not just bigger context windows. Ross recounts how Galactica's 2022 thinking tokens were prescient, but its base-model demo backfired, while RLHF made LLMs products; his Meta team's PPO with verifiable rewards worked, but DeepSeek-R1 later showed better base models were key. Chengxi details long-horizon obstacles: sparse rewards, credit assignment, and 1M token limits, solved by value models that reduce variance and enable bootstrapping. Kelly Bench gave frontier models $100K to trade Premier League matches; all lost money, exposing how little current environments simulate real competition. Pipeline RL trades off off-policy staleness (8 steps okay) against GPU utilization, and they point to openreward.ai, with 350+ environments, for long-horizon RL.
Powered by PodHood