A product discussed on AI Engineer.

Building GTM AI Agents: Lessons from Deploying to 6,000 Users — Sait Izmit, Snowflake
Aug 26, 2026 · 20:39
Sait Izmit of Snowflake says winning over 6,000 go-to-market users comes down to quality over coverage: he wrote 150 sales questions before testing, accepted 50% first accuracy, and chose 50 questions at 95% over 100 at 70%. The Snowflake Cowork agent, live since September, has answered over a million questions (~40,000 weekly); 60% of data arrived post-launch, and it now spans 15 semantic views, 85 tables, 3,000 columns, MCPs, and 20 skills. After pilot and a 600-user 10% beta with 70% retention, GA showed only 20% of the org tried it, so Izmit spends 60-70% of his time on demos and sales meetings. He warns the wow factor collapses quickly, so teams must move from data chat to workflow automation, build fast with today's stack, accept rearchitecture, and mine logs for feedback loops.

Prototyping as Leadership: How a CTO Ships with AI Agents — Hursh Agrawal, The Browser Company
Aug 20, 2026 · 18:18
Hursh Agrawal, CTO and co-founder of The Browser Company, argues AI coding agents let leaders keep building: despite 15+ meetings, 7 direct reports, and a toddler, he ships 2–10 PRs/week. He says frontier models turn over every three months, so hands-on building is the only way to judge them and show engineers working prototypes. His method is an overnight loop: a coworker agent gathers Slack/Jira/Notion context into a prompt at 5pm, coding agent runs for hours, and a morning hour reviews tests, CI, and AI code review. He details three overnight uses: building features, hill-climbing evals from feedback JSONs, and training custom models like a PII classifier on AWS. He warns leaders to avoid critical path work and to rely on trustworthy CI, feature flags, a prototype branch, and readable PRs.

Benchmarking Coding Agents on New vs Legacy Codebases — Denys Linkov, Wisedocs
Aug 8, 2026 · 18:08
Denys Linkov of Wisedocs, whose medical-claims ML pipeline ran across ten legacy repos, argues his team's six-month monorepo refactor was worthwhile even as coding agents improve fast. A refactor task that took o3 three hours and ten major mistakes now takes about one-fifth the time: Sonnet 4.6 needed one extra iteration, Opus 4.8 nearly one-shot it. Yet GPT 5.5 extra high 'completed' the job in 10 minutes 22 seconds, writing 2,000 lines of scaffolding with models missing and admitting no deployment or bootstrap command. So Linkov reads METR's task-length curve at 80-90% success, not 50%; an hour-long agent run at coin-flip odds wastes the hour and your attention. The payoff was social too: commit velocity never flattened, months-long features ship in under a week, and developers now volunteer across the monorepo.

AI System Design: From Idea to Production - Apoorva Joshi, MongoDB
Jun 28, 2026 · 28:53
Apoorva Joshi presents a structured framework for designing AI systems from idea to production, arguing that thorough product specification and evaluation are now more critical than coding itself. Using a health insurance claims review system for MDB Health as an example, she walks through four phases: defining business problems with measurable success metrics (e.g., reducing urgent claim processing from 2 days to 1 hour within 90 days), designing data strategy and architecture using patterns like RAG, control flow, and human-in-the-loop, establishing guardrails and evaluation metrics such as faithfulness and cost per recommendation, and optimizing for accuracy, cost, latency, and reliability before shipping. She emphasizes building evaluation in from the start and iterating from the simplest system.

Demand-Driven Context: A Methodology for Coherent Knowledge Bases Through Agent Failure
May 5, 2026 · 1:08:15
Raj Navakoti, a staff software engineer at IKEA, presents a demand-driven context methodology for building coherent knowledge bases by letting AI agents fail on real problems and surfacing missing institutional knowledge. He argues that enterprises should shift from pushing monolithic documentation to a pull approach where agents reveal undocumented tribal knowledge through repeated failures on incidents and Jira tickets. Using a framework with skills, rules, and hooks, he demonstrates how agents can gradually improve confidence scores (from 1.4 to 4.4 over 14 incidents) by documenting discovered context blocks. Navakoti introduces a context gap scanner that automatically analyzes work items against existing documentation to identify critical gaps, outdated information, and duplications. He advocates storing curated knowledge in GitHub for version control and PR-based collaboration, and emphasizes that this approach helps teams know the unknown, enabling agents to manage knowledge rather than just consume it.

Mergeable by default: Building the context engine to save time and tokens — Peter Werry, Unblocked
May 3, 2026 · 1:41:25
Peter Werry of Unblocked argues that context engines—systems that supply AI agents with only the relevant organizational context—are critical to avoid agent doom loops and wasted tokens. He debunks three myths: naive RAG, connecting MCP servers, and bigger context windows do not solve the context problem. Werry describes building a social engineering graph to identify experts and distill team best practices, and shares hard lessons including hiding conflicts and caching answers. In a benchmark task, Unblocked's context engine reduced a 2.5-hour, 21-million-token task to 25 minutes and 10 million tokens. The talk offers a practitioner's guide to building context engines with conflict resolution, personalization, and access control.

Ship Production Software in Minutes, Not Months — Eno Reyes, Factory
Jul 25, 2025 · 16:06
Eno Reyes, cofounder and CTO of Factory, argues that AI agents can orchestrate the entire software development lifecycle, moving beyond vibe coding to agent-native development where enterprises delegate planning, coding, testing, and incident response to autonomous droids. He explains that AI tools are only as good as the context they receive—missing context from meetings, whiteboards, or Slack is the primary cause of failure, not LLM quality. Factory's droids search codebases, leverage organizational memory, and question unclear tasks before executing, from generating PRDs and tickets to creating runbooks and RCAs from sentry alerts. Reyes demonstrates how an agent can convert user transcripts and ad-hoc notes into a full feature plan, then break it into parallel tickets for multiple code droids. For incident response, droids pull logs, historical runbooks, and team discussions to produce mitigation plans in minutes, cutting response times in half and shifting from reactive to predictive operations. He emphasizes that the future belongs to engineers who manage agents—thinking clearly and communicating effectively—rather than those writing every line of code.
Powered by PodHood