A company discussed on AI Engineer.

Building uReview, Uber’s Multi-Agent Code Review Engine — Will Bond & Ameya Ketkar, Uber
Aug 28, 2026 · 15:07
Will Bond and Ameya Ketkar explain how Uber built uReview, its multi-agent code review engine, to fight a growing review bottleneck: first-review wait times at Uber grew from 3 hours in 2024 to 9 hours in 2026. They describe why Uber built rather than bought (most vendors don't support Phabricator) and how they made the system work by measuring reply sentiment, addressal rate, and agent trajectory, since 'the model never knows that it's wrong.' Letting hundreds of teams write their own reviewers was easy to author but hard to run cheaply at scale. Results: uReview posts about 25,000 comments a week, roughly 67% get addressed, costs fell 60% versus the naive first build, and quality and accuracy rose 70%. They close by arguing the agentic shift is expanding rather than killing the human outer loop, moving engineers from implementation details to architecture and product thinking.

Building Blocks for Uber’s Software Factory— Uday Kiran Medisetty & Adam Huda, Uber
Aug 21, 2026 · 18:26
Uber's Uday Kiran Medisetty and Adam Huda explain the six building blocks of their agentic software factory, which now generates more than 70% of pull requests and has doubled lines of code per engineer year over year. A model gateway handles 100 million requests a day across 800 projects with redaction of 20+ PII types and safety models under 100 milliseconds; an MCP gateway cuts token use over 40%, a skills marketplace runs 20,000 executions daily across 2,500 skills, and a context graph of 40 million entries replaces 20-30 systems. Adam shows one feature end-to-end, from Slack idea to draft PR that stops before CI, validating via simulator screenshots against Figma in an inner loop. The bottleneck, they argue, has moved to deciding whether a thing should be built at all.

The Next Medium: Why Real-Time Interactive Video Changes Everything — Ahmed Ahres, Reactor
Aug 18, 2026 · 17:30
Ahmed Ahres, head of go-to-market at Reactor, argues that world models are real-time interactive video, a change of medium rather than a speedup, because video becomes programmable like software. GPS enabling Uber and viewfinders enabling Instagram and TikTok show how real time unlocks new applications. Reactor's platform serves infinite interactive video (Helios from ByteDance), controllable worlds (Lingbot from Alibaba, LongLive 2 from Nvidia), and live avatars he admits are still uncracked. Users build interactive live streams, medical and cooking simulations, and video-to-video editing. Real-time infrastructure means streaming pixels, live sessions with memory, and sub-100-millisecond latency; he offers promo code AIE2026 for $75 credits.

Building Closed-Loop Evals for a Multimodal Agent at Scale — Soumya Gupta & Jai Chopra, Uber
Jul 24, 2026 · 21:39
Soumya Gupta and Jai Chopra from Uber detail how they built closed-loop evals for their multimodal food photo enhancement agent, which edits images for Uber Eats merchants while preserving authenticity and avoiding homogenization. They describe a routing agent using a recall guardrail to decide whether to enhance or skip an image, and a pass at K metric for iterative enhancement with QA gates that check faithfulness, completeness, and realism. Examples include reward hacking where the agent overcorrected to a generic plate and failures like hallucinating extra chicken wings. They explain multiple feedback loops: a model loop for drift detection using human labels, internal dogfooding with thumbs up/down, and production metrics like conversion rates, all fed into a diagnoser that auto-tunes agents via a reflect-and-synthesize prompt optimizer, ensuring the system evolves without human intervention.

The Future of Evals: From LLM as a Judge to Agent as a Judge — Aparna Dhinakaran, Arize AI
Jul 24, 2026 · 6:06
Aparna Dhinakaran, co-founder of Arize AI, argues that as agents evolved from simple prompts to complex systems with tool calls, reasoning, and long-horizon tasks, evals must evolve too—from deterministic checks to LLM as a judge, and now to agent as a judge. She reveals that the top teams run over 3,800 different evaluators, yet classical LLM-as-judge evals fail to catch subtle failures in agents that generate unique trajectories per user. Arize's new tool, Signal, is a long-running agent that reads traces, discovers patterns like inefficient tool loops, and can even open a PR to fix issues. The episode traces this arc from static checks to adaptive analysis, emphasizing that the future of evals requires all three layers to handle the complexity of modern agents.

How AI is changing Software Engineering: A Conversation with Gergely Orosz, @pragmaticengineer
Apr 21, 2026 · 26:42
Gergely Orosz explains how AI is reshaping software engineering, from token maxing at Meta, Microsoft, and Salesforce—where engineers use AI tools to inflate token counts out of fear of layoffs and performance reviews—to the broader shift in the engineer's role toward orchestrating agents rather than managing people. He notes that big tech companies like Uber, Airbnb, and Shopify are building custom internal AI infra (MCP gateways, coding agents) to stay ahead, even tolerating high churn for a six-month competitive edge. Shopify secured GitHub Copilot a year early by offering feedback from 3,000 engineers. Gergely also shares how The Pragmatic Engineer hit product-market fit: 100 paid subscribers before publishing, reaching 1,000 in six weeks, and becoming the #1 paid tech newsletter by focusing on two deep-dive articles per week for two years.

Taste & Craft: A Conversation with Tuomas Artman, CTO Linear & Gergely Orosz, @pragmaticengineer
Apr 21, 2026 · 29:17
Tuomas Artman, CTO of Linear, warns that AI's ability to instantly ship features risks creating convoluted, low-quality software, arguing that taste and design must guide development. He explains Linear's culture of deliberate product decisions, including a 'zero bug policy' where bugs are fixed within hours, and 'Quality Wednesdays' where engineers find and fix one minute detail each week—resulting in over 2,500 quality fixes. Artman notes that AI lacks human 'taste' and cannot feel user experience nuances like animation timing. He predicts all software engineers will become product engineers, needing to focus on customer needs and UX, and advises aspiring product engineers to build for themselves, talk to customers, and study Apple's Human Interface Guidelines.

Shipping AI That Works: An Evaluation Framework for PMs – Aman Khan, Arize
Dec 26, 2025 · 1:26:16
Aman Khan, AI PM at Arize, presents a framework for product managers to evaluate LLM-powered products beyond gut-feel 'vibe checks.' He demonstrates building an AI trip planner with multi-agent LangGraph, then using Arize's tracing and prompt playground to iterate on prompts. Khan shows how to create datasets from production traces, run A/B experiments on prompts, and use LLM-as-a-judge evals for friendliness and discount offers, comparing against human labels to refine evaluators. He argues evals are the new requirements docs, enabling PMs to own the product experience by writing acceptance criteria as eval datasets. The talk covers building eval teams, handling variance with temperature settings, and continuously improving golden datasets with hard examples, citing real-world analogies from self-driving cars at Cruise.

Insights on Building AI Teams — Heath Black, SignalFire
Apr 15, 2025 · 20:30
Heath Black, Managing Director of Product at SignalFire, uses Beacon platform data to guide AI team building, arguing that credentialism is declining—only 7% of AI engineers had PhDs in 2023 versus 16% in 2015—and that work experience now outweighs education. He shows AI talent concentrates in San Francisco (35% of AI engineers), Seattle (22%), and New York (10%), and that tracking retention rates (e.g., Anthropic at 66% four-year retention vs. Perplexity at 43%) helps time outreach. Black advises hiring based on a candidate's body of work, removing academic requirements from postings, and understanding generational job-hopping (27% of Gen Z left jobs in 2023). He warns against relying solely on salary and equity, as AI engineers command 5% salary and 10–20% equity premiums, and recommends narratives centered on mission, speed, and collaborative teams. The talk emphasizes using data to filter, locate, time, and close hires effectively.
Powered by PodHood