A product discussed on AI Engineer.

How I automate my own job at Hugging Face using agents — Niels Rogge, Hugging Face
Aug 20, 2026 · 20:37
Niels Rogge, machine learning engineer on Hugging Face's community science team, explains how he automated his own job: agents ask researchers to publish paper artifacts on the Hub instead of Google Drive, Dropbox, or Zenodo. The outreach is a deterministic nightly workflow on GitHub Actions cron jobs with LangFuse tracing; follow-ups run as a fully autonomous Claude Agents SDK loop, with each GitHub issue in its own Modal container using Bash and one Hugging Face CLI skill. Thousands of issues have drawn only two negative replies, and he doesn't disclose the bot because recipients reply to the same messages he used to send. He argues open models like GLM 5.2 can replace closed ones, and cites his Daily Papers X account (90,000 followers) and a Papers with Code revival at paperswithcode.co.

Lessons from building GenAI based applications — Juan Peredo
Feb 22, 2025 · 33:13
Juan Peredo details the hidden complexities of building GenAI applications, from model hosting and cost control to output validation and observability. He compares local (Ollama) vs cloud hosting (Modal, SkyPilot) and warns that an agent processing 3,000 calls/day with OpenAI O1 costs nearly $300,000/month, while LLaMA 3.3 70B drops that to $50,000/month. He explains techniques to mitigate hallucinations—prompt engineering, guardrails, RAG, and fine-tuning—each with trade-offs like added latency or cost. Peredo advocates externalizing prompts via LangChain Hub for easy iteration and future-proofing, and illustrates agent design with parallel calls to reduce latency. Finally, he stresses observability using tools like LangSmith to debug probabilistic failures, such as an LLM failing on case sensitivity.
Powered by PodHood