A product discussed on AI Engineer.

How to Get Your Org to Adopt Coding Agents (Without Shipping Garbage) — Eyal Blum, Figma
Aug 28, 2026 · 17:43
Eyal Blum, a software engineer at Figma, explains why the engineers slowest to adopt coding agents are often the best ones, and how his org is making agent adoption safe without shipping garbage. He outlines a three-act adoption process and notes adoption is uneven, forcing AI-forward and skeptical teams to coexist. Blum argues the highest-value investment is verification — using Playwright MCP, moving deterministic checks down the testing pyramid, and having agents write tests first in TDD style. He advocates detailed planning over prompting: a week-long plan yields 20 small PRs, turning six weeks of coding into one week (a 5x speedup). He also addresses reduced developer agency, says AI-written communication needs explicit labeling because human attention is scarce, and advises treating skeptics' complaints as the roadmap to make agents safer.

The Missing Layer in Agentic AI — Giedrius Šteimantas, Oxylabs
Aug 26, 2026 · 15:04
Giedrius Šteimantas of Oxylabs argues the missing layer in agentic AI is web scraping infrastructure, applying ten years of scraping rules to make agents cheaper and more reliable. His friend's shopping agent used a browser for everything, hit CAPTCHAs, and wasted tokens. Rebuilding it, he replaces discovery's browser and fixed retailer list with Oxylabs' fast search API (under 2,000 tokens, under 700 milliseconds), and the decision stage with a scraper API that returns markdown, fails loudly, runs hundreds of parallel requests, and bills only for successful results. Checkout needs a browser, so Playwright MCP connects to Oxylabs' headless browser with stealth, residential proxy, and geolocation. Cost matters: use a browser only when necessary, validate content—HTTP 200 does not mean valid.

Anthropic Workshop: Build Agents That Run for Hours — Ash Prabaker & Andrew Wilson
May 18, 2026 · 1:15:40
Anthropic's Ash Prabaker and Andrew Wilson detail how to build agents that run for hours by replacing self-evaluation with adversarial evaluator agents that use Playwright to test live apps and grade subjective output via rubrics. They explain that context compaction doesn't cure coherence drift, so structured handoffs between fresh context windows are essential. The generator and evaluator negotiate testable sprint contracts before building, and the planner provides high-level specs without overspecifying technical details. They show that a solo Claude Code session built a retro game maker that looked complete but failed in play mode, while the adversarial harness produced a fully functional app with live physics and AI features. Key takeaways include reading traces as the primary debug loop, deleting harness components as models improve, and using file-system state for long-running agents.
Powered by PodHood