Page 19 of 23

The Devops Engineer Who Never Sleeps — Diamond Bishop, Datadog
Apr 22, 2025 · 16:18
Diamond Bishop, Director of AI Engineering at Datadog, explains how his team builds AI agents—the On-Call Engineer and Software Engineer—that automate DevOps tasks like incident investigation, remediation, and postmortem writing. He details how the On-Call Engineer wakes up for alerts, reads runbooks, queries logs and metrics, and suggests fixes to let human engineers sleep. The Software Engineer proactively fixes errors by generating code diffs and pull requests. Bishop shares four key lessons: scope tasks and evaluate rigorously, assemble teams of optimistic generalists and UX experts, adapt UX for human-agent collaboration, and treat observability as critical for debugging multi-step agent workflows. He predicts that within five years, AI agents will surpass humans as primary users of SaaS platforms like Datadog, urging builders to design for agent consumers.

Evaluating Domain Specific LLMs for Real World Finance — Waseem Alshikh, Writer
Apr 22, 2025 · 12:01
Writer CTO Waseem Alshikh presents FailSafe, a benchmark for evaluating LLMs in real-world finance, challenging the notion that general-purpose models suffice. Tests on query failures (misspelling, incomplete, out-of-domain) and context failures (missing context, OCR errors, irrelevant context) reveal that reasoning models like o1 and o3 perform 50% to 60% worse on context grounding than smaller domain-specific models, despite higher answerability. Even the best model achieves only 81% combined robustness and grounding, meaning roughly 20% of responses are wrong. Alshikh argues that domain-specific models remain essential, and that full-stack systems—RAG, guardrails, and grounding—are necessary for reliable deployment in high-stakes finance.

Self Coding Agents — Colin Flaherty, Augment Code
Apr 21, 2025 · 17:23
Colin Flaherty, founding researcher at Augment Code, describes how the company built an AI coding agent that wrote over 90% of its own 20,000-line codebase with human supervision. The agent autonomously added third-party integrations like Google Search and Linear by searching its own codebase and even used its own Google Search tool to look up Linear API docs. It profiled and optimized itself, replacing synchronous file hashing with a process pool after adding print statements and running sub-copies. Flaherty argues that agent capabilities improve with better context engines, code execution environments, and test harnesses—adding test-driven iteration boosted a bug-fixing benchmark by 20% versus 4% from just upgrading the foundation model. He cautions that agents are general-purpose, not category-specific, and have different strengths than humans, making knowledge bases crucial for onboarding. With code becoming cheap, the focus shifts to product insights and design, while better tests enable greater autonomy and accelerate agent self-improvement.

Vercel AI SDK Masterclass: From Fundamentals to Deep Research
Apr 20, 2025 · 59:52
Nico Albanese from Vercel delivers a workshop on building agents with the AI SDK, demonstrating how to use generateText, tools, and structured outputs to create a deep research clone in Node.js. He shows that switching between models like GPT-4o mini, Perplexity Sonar Pro, and Gemini Flash 1.5 requires only changing one line of code. Tools with maxSteps enable autonomous multi-step agents, as when the model fetches weather data for San Francisco and New York and sums the temperatures across three steps. The deep research workflow generates subqueries (e.g., 'requirements to become a D1 shotput athlete'), searches via Exa with live crawl, evaluates relevance with an evaluate tool that discards irrelevant results, and recursively deepens research using depth and breadth parameters. Finally, o3 mini synthesizes the accumulated research into a structured Markdown report, all in 218 lines of code.

Frontier Feud: Anthropic, Google DeepMind, Meta FAIR, Thinking Machines — Barr Yaron, Amplify
Apr 19, 2025 · 22:26
Teams from Anthropic, Google DeepMind, Meta FAIR, and Thinking Machines compete in a Frontier Feud game hosted by Barr Yaron at the AI Engineer Summit 2025. Surveyed AI engineers name Ilya Sutskever as the most influential AI researcher, with Andrej Karpathy, Jeff Hinton, and Yann LeCun also on the board. Cost is the top consideration when choosing a model, followed by latency, eval scores, and open vs. closed source. The buzzword everyone is tired of hearing is AGI, with RAG and prompt engineering trailing. In the fast money round, Cursor tops favorite AI tools, 'Attention Is All You Need' wins most influential paper, and hardware failure is the biggest 2 a.m. nightmare. The winning team, Rocco's Basilisk, takes home a massive llama and other prizes.

AI + Security & Safety — Don Bosco Durai
Apr 19, 2025 · 18:13
Don Bosco Durai, CTO of Privacera and creator of Apache Ranger, argues that building safe and reliable AI agents requires a multi-layered security approach combining preemptive vulnerability evals, proactive enforcement, and real-time observability to address challenges like unauthorized access, data leakage, and compliance in enterprise production. He explains that current agent frameworks run as a single process sharing credentials, creating zero-trust vulnerabilities, and autonomous agents introduce unknown unknowns. He advocates for three layers: preemptive security evals (including prompt injection, data leakage, runaway agent tests) to generate a risk score for production promotion; proactive enforcement with authentication/authorization propagated across all components and approval workflows; and observability with thresholds and anomaly detection to monitor agent behavior in production. He illustrates compliance needs using his customer example—a top credit agency needing to treat agents like human users for regulatory adherence. Bosco also open-sourced PAIG.ai as a safety and security solution for GenAI and AI agents.

Stateful Agents — Full Workshop with Charles Packer of Letta and MemGPT
Apr 19, 2025 · 1:19:34
Charles Packer, lead author of the MemGPT paper and co-founder of Letta, argues that statefulness (memory) is the most important problem to solve for building useful AI agents, since LLMs are inherently stateless transformers. He presents MemGPT's LMOS (Language Model Operating System) approach, which treats memory management as a context compilation problem solved by the LLM itself using tool calling to read and write structured memory blocks. The workshop demonstrates Letta's open-source stack (FastAPI, Postgres, Python) where agents persist state on a server, enabling long-running, learning interactions without context overflow—e.g., an agent can update its core memory (e.g., correcting a user's name) or search archival memory (e.g., recalling user preferences after a reset). Packer also shows multi-agent communication via async message passing between agents running as independent services, and highlights that tools are sandboxed by default with support for Composio integrations. The session includes a live notebook exercise and a low-code UI (ADE) showing context window management and memory editing, emphasizing that true stateful agents improve continuously over time rather…

Voice Agent Engineering — Nik Caryotakis, SuperDial
Apr 18, 2025 · 19:07
Nik Caryotakis from SuperDial argues that in 2025, the key to production Voice AI is reliability over realism, especially for sensitive healthcare calls. SuperDial automates back-office phone calls, saving over 100,000 hours of human calling with a lean team of four engineers. He advocates for the 'say the right thing at the right time' approach, using open-source tools like PipeCat for orchestration, TensorZero for LLM routing, and self-hosted LangFuse for HIPAA-compliant observability. Caryotakis warns that new voice-to-voice models often produce nonsensical audio, favoring a sequenced STT/LLM/TTS pipeline for control. He shares specific last-mile challenges: pronunciation of names like 'Caryotaikis', avoiding confusing bot names like 'Billy', and the need for fallbacks when OpenAI goes down. The talk emphasizes that the unique value of a voice agent lies in conversational design and vertical integrations, not realistic voices.

Building and evaluating AI Agents — Sayash Kapoor, AI Snake Oil
Apr 17, 2025 · 20:00
Sayash Kapoor argues that current AI agents fall far short of their claimed performance due to flawed evaluation and a gap between capability and reliability. He cites failures like Do Not Pay (fined by FTC), LexisNexis (hallucinations in up to a third of cases), and Sakana AI (agent hacked reward functions, claiming 150x speedup that exceeded H100's theoretical max). Princeton's CoreBench shows best agents reproduce under 40% of papers. He emphasizes that agent benchmarks like SWE-bench mislead VC funding—Cognition's Devin succeeded on only 3 of 20 real-world tasks. Kapoor calls for cost-aware, multi-dimensional evaluation (e.g., Holistic Agent Leaderboard with Pareto frontiers) and a shift from capability to reliability engineering, drawing parallels to ENIAC's vacuum tube failures.

Building LinkedIn's GenAI Platform — Xiaofeng Wang
Apr 16, 2025 · 17:53
Xiaofeng Wang, manager of LinkedIn's GenAI Foundation, explains the evolution of LinkedIn's GenAI platform from simple prompt-in string-out applications to a multi-agent system for LinkedIn Hire Assistant, arguing that a unified platform is critical for bridging the gap between AI and product engineers in the era of compound AI systems. He details the platform's four-layer architecture—orchestration, prompt engineering, tools/skills invocation, and memory management—and key investments like a Python SDK, centralized skill registry, experiential memory across working, long-term, and collective layers, and observability built on OpenTelemetry. Wang discusses hiring philosophy: prioritize strong software engineers over AI expertise, hire for potential, and build diverse teams integrating full-stack engineers, data scientists, and AI engineers. He recommends solving immediate needs first, leveraging existing scalable infrastructure like messaging systems for memory, and focusing on developer experience to drive adoption.

Insights on Building AI Teams — Heath Black, SignalFire
Apr 15, 2025 · 20:30
Heath Black, Managing Director of Product at SignalFire, uses Beacon platform data to guide AI team building, arguing that credentialism is declining—only 7% of AI engineers had PhDs in 2023 versus 16% in 2015—and that work experience now outweighs education. He shows AI talent concentrates in San Francisco (35% of AI engineers), Seattle (22%), and New York (10%), and that tracking retention rates (e.g., Anthropic at 66% four-year retention vs. Perplexity at 43%) helps time outreach. Black advises hiring based on a candidate's body of work, removing academic requirements from postings, and understanding generational job-hopping (27% of Gen Z left jobs in 2023). He warns against relying solely on salary and equity, as AI engineers command 5% salary and 10–20% equity premiums, and recommends narratives centered on mission, speed, and collaborative teams. The talk emphasizes using data to filter, locate, time, and close hires effectively.

AI Engineers: The Next Generation — Stefania Druga, Google Gemini
Apr 13, 2025 · 21:47
Stefania Druga, a research scientist on Google Gemini, presents Cognimates and its new AI copilot that teaches children to become AI engineers. Built on Scratch’s visual programming, Cognimates lets kids train custom models, program robots, and build games, shifting their perception from ‘AI is magic’ to understanding data and confidence levels. Druga’s design studies with 18 kids from 11 countries showed the copilot doubled programming time by supporting ideation, debugging, and creative agency—never giving answers unless stuck three times. The tool now integrates screenshot analysis and asset generation, with plans to embed the agent across the OS. Druga ties this to the EU AI Act’s mandate for AI literacy, arguing early, hands-on creation is the path to informed users.

How to Fail at AI Strategy: Hamel Husain & Greg Ceccarelli
Apr 13, 2025 · 17:03
Greg Ceccarelli and Hamel Husain argue that the most reliable path to AI failure is to follow a set of inverted worst practices, including cultivating disconnect between executives and builders, promising unrealistic AI capabilities, drowning communication in jargon, and avoiding data analysis. They describe how to divide your company by incentivizing secrecy and using jargon like 'agents' to exclude domain experts, ensuring that AI projects are disconnected from real needs. The speakers advocate faking strategy by highlighting random paragraphs from last year's report, announcing vague goals like 'become the global AI leader in everything,' and creating a massive backlog with no timeline. They recommend throwing tools at problems—buying expensive vector databases or switching frameworks—without understanding root causes, and blindly trusting off-the-shelf evaluation metrics like BLEU and ROUGE. Crucially, they insist on never looking at data, using complex systems inaccessible to domain experts, and trusting gut feelings over evidence. This inverted guide guarantees wasted resources, alienated teams, and spectacular failure.

Anthropic in the Enterprise — Alexander Bricken & Joe Bayley
Apr 13, 2025 · 20:55
Alexander Bricken and Joe Bayley from Anthropic's Applied AI team argue that enterprise AI implementation often fails due to overengineering, poor data infrastructure, or lack of testing—but industry leaders achieve transformative results with Claude. They detail Anthropic's deployment models (API, cloud partnerships, enterprise solutions) and real-world case studies like Intercom's Fin agent, which solved 86% of support volume using Claude. Best practices include building evals early as intellectual property, identifying intelligence/cost/latency trade-offs based on use-case stakes, and avoiding premature fine-tuning by trying prompt caching, contextual retrieval, and agentic architectures first. They also highlight interpretability research and the Model Context Protocol for reliable AI deployments.

Finetuning: 500m AI agents in production with 2 engineers — Mustafa Ali & Kyle Corbitt
Apr 12, 2025 · 18:44
Method Financial and OpenPipe detail how Method scaled AI financial agents to 500 million daily interactions using fine-tuned open-source models instead of expensive GPT-4. After racking up $70,000 in monthly GPT-4 costs and facing latency and error issues, they fine-tuned an 8B parameter LLaMA 3.1 model to achieve under 200ms latency and 9% error rate (beating GPT-4o's 11% at lower cost). Mustafa Ali explains that the key was using production data from GPT as training data and choosing the cheapest model that met performance goals. Kyle Corbitt emphasizes that fine-tuning is a power tool for bending the price-performance curve when prompt engineering falls short. The episode argues that productionizing AI agents requires patience and openness from engineering teams, and concludes with a call for software engineers to pivot to AI engineering.

The Agent Development Life Cycle — Zack Reneau-Wedeen, Sierra
Apr 11, 2025 · 18:40
Zack Reneau-Wedeen from Sierra explains the company's Agent Development Lifecycle for building reliable, testable AI agents at scale for brands like Chubbys and SiriusXM. He contrasts LLMs' nondeterministic, slow nature with traditional software, calling them a 'foundation of Jello.' Sierra treats every agent as a product, using an experience manager to review conversations, file issues, create tests, and release improvements—growing from hundreds of thousands of requests for Chubbys to tens of millions for larger customers. The lifecycle spans quality assurance, testing, and deployment, with reasoning models acting as a force multiplier. Voice agents launched generally in October 2024, handling calls with the same underlying platform, enabling responsive design across channels.

RAG Agents in Prod: 10 Lessons We Learned — Douwe Kiela, creator of RAG
Apr 10, 2025 · 16:56
Douwe Kiela, CEO of Contextual AI and creator of RAG, shares 10 lessons from deploying enterprise RAG systems at scale. He argues that language models are only 20% of a larger system; success comes from focusing on systems, not models, and specializing over AGI to unlock domain expertise. Enterprise data is the real moat, but pilots are easy while production is hard—design for production from day one. Speed beats perfection: ship barely functional to real users early and iterate. Avoid boring engineering chores like chunking; instead, integrate AI into existing workflows to drive adoption. Accuracy is table stakes; handle inaccuracy with observability and attribution. Be ambitious: aim for transformative ROI, not low-hanging fruit like basic HR questions.

Trust, but Verify: Knowledge Agents for Finance Workflows - Mike Conover
Apr 9, 2025 · 21:10
Mike Conover, CEO of Brightwave, explains how his company builds knowledge agents for financial workflows, digesting thousands of pages of content for due diligence and research. He argues that non-reasoning models perform only local search, so winning systems must use end-to-end reinforcement learning over tool-use calls to achieve globally optimal outputs. Conover describes design patterns like decomposing research into sub-themes, using secondary calls for error correction, and avoiding the latency trap of long feedback loops. He emphasizes that synthesis—weaving facts across documents—remains a hard problem due to limited recombinative reasoning in training data. Brightwave's product reveals its thought process through interactive citations and structured findings, allowing analysts to drill into any passage. The episode also notes the company is hiring, offering a $10,000 referral bonus.

Building AI Agents with Real ROI in the Enterprise SDLC: Bruno (Booking.com) & Beyang (Sourcegraph)
Apr 8, 2025 · 20:56
Beyang Liu, CTO of Sourcegraph, and Bruno Passos, Group PM at Booking.com, explain how their partnership built AI coding agents that deliver measurable ROI in enterprise software development. Booking.com, with 3,000 developers handling 250,000 merge requests and 2.5 million CI jobs yearly, used Sourcegraph's Cody and custom agents to automate code migration and review, cutting cycle times. They report that daily Cody users ship 30% more pull requests and are 30% faster, with lighter MRs. Agents like a GraphQL generator for a million-token schema and a code review tool with customizable rules emerged from joint hackathons. Bruno emphasizes education as critical: training turned skeptical developers into daily users, driving adoption and productivity gains. The goal is to shift-left compliance and reduce tech debt, making the SDLC self-healing.

Anchoring Enterprise GenAI with Knowledge Graphs: Jonathan Lowe (Pfizer), Stephen Chin (Neo4j)
Apr 7, 2025 · 20:59
Stephen Chin (Neo4j) and Jonathan Lowe (Pfizer) explain how Pfizer uses knowledge graphs with GraphRAG to accelerate drug manufacturing technology transfer, cutting time from years to weeks. They argue graph databases provide superior accuracy and explainability over vector-only RAG, crucial for life-saving drugs. Jonathan details navigating organizational silos in a 100,000-person company, from C-suite taglines to client partners demanding cost savings. Gartner's prediction of 30% GenAI project failure is addressed with a concrete business case: manufacturing worker tenure plunged from 20 to 3 years, making AI essential to capture lost expertise. The architecture combines vector and graph retrieval to deliver contextually relevant, auditable answers, reducing data consolidation from three months to three weeks.

Personal, Local, Private AI Agents: Soumith Chintala
Apr 6, 2025 · 20:32
Soumith Chintala, co-founder of PyTorch, argues that personal AI agents should run locally and privately to maintain trust and control over intimate data. He warns that cloud-based agents, lacking complete context (like access to all messaging or financial accounts), become unreliable and potentially dangerous—catastrophic actions like buying a Tesla instead of Tide Pods are possible. Key technical challenges include slow local inference, immature open-source computer-use models, and poor catastrophic action classification. He is bullish on open models surpassing closed ones through coordinated improvement, citing Linux, Llama, and DeepSeek. Chintala also plugs open reasoning data from gr.ink and PyTorch’s work on enabling local agents, urging AI engineers to tackle these gaps.

How We Build Effective Agents: Barry Zhang, Anthropic
Apr 4, 2025 · 15:09
Barry Zhang of Anthropic's Applied AI team argues that effective agents require simplicity, not complexity, and should be built only for tasks with high value and ambiguous problem spaces. He offers a checklist: ensure task complexity is high, value justifies token cost, critical capabilities are de-risked, and errors are easily discovered (e.g., coding with unit tests). Agents are just models using tools in a loop — environment, tools, and system prompt — and he advises iterating on these three components before optimizing. To improve agents, developers should think like them by narrowing their perspective to the agent's 10-20k token context window and even asking Claude to critique its own tools and trajectories. Zhang forecasts three open challenges: making agents budget-aware by enforcing time/token/money limits, enabling self-evolving tools via meta-tools, and building asynchronous multi-agent communication beyond synchronous turns.

Scaling Agents for Gen AI Products - Anju Kambadur, Bloomberg Head of AI Engineering
Apr 1, 2025 · 19:38
Bloomberg's Head of AI Engineering, Anju Kambadur, details the company's shift from building proprietary LLMs to leveraging open-source models for agentic products, emphasizing that scaling agents requires accepting inherent fragility and building resilient guardrails rather than seeking perfect upstream systems. Drawing on Bloomberg's daily scale—400 billion structured data ticks, over a billion unstructured messages, and 40+ years of history—Anju explains how their 400-person AI team organized into 50 teams across London, New York, Princeton, and Toronto. He argues that agents must be semi-autonomous with non-optional guardrails (e.g., preventing financial advice, ensuring factuality) because compounding errors from evolving APIs and LLMs demand self-contained safety checks. The talk uses the example of a research analyst agent that factors query understanding, answer generation, and guardrails into separate components, reflecting an org structure that collapses vertically for fast iteration on single agents then later introduces horizontal teams (like guardrails) for optimization and cost reduction. Anju stresses that readiness to build horizontal teams comes after multiple…

AI Engineering at Jane Street - John Crepezzi
Mar 28, 2025 · 16:57
John Crepezzi, an engineer on Jane Street's AI Assistants team, explains how they built custom LLM-powered coding assistants for OCaml, a functional language with scarce public training data. To overcome off-the-shelf tool limitations, they trained models on data from workspace snapshots — automated captures of developer workstations every 20 seconds, tracking build status changes from red to green. They used a Code Evaluation Service (CES) for reinforcement learning and evaluation, pre-warming builds to quickly test if model-generated diffs compile and pass tests. Their sidecar architecture, AID, serves thin editor integrations for Emacs (used by 67% of the firm), VS Code, and Neovim, allowing easy model swapping and A/B testing (e.g., sending 50% of users to different models). The talk details the end-to-end process: data collection from features, commits, and snapshots; training with supervised data and RL; and building pluggable infrastructure for domain-specific tools.

How Deep Research Works - Mukund Sridhar & Aarush Selvan, Google DeepMind
Mar 26, 2025 · 15:15
Aarush Selvan and Mukund Sridhar from Google DeepMind explain how Gemini Deep Research works, a personal research agent that trades latency for comprehensiveness by taking up to five minutes to browse the web and synthesize reports. They discuss product challenges like building an asynchronous experience in a synchronous chatbot, setting user expectations, and presenting thousand-word outputs. Technical challenges covered include iterative planning with partial information, handling the fragmented web with entity resolution, managing growing context via recency-biased retrieval, and ensuring robustness to intermediate failures. The episode also explores future directions: going from aggregating information to providing expert-level insights, personalizing research to user roles, and combining web research with coding, data science, or video generation.

Why Agent Engineering — swyx
Mar 24, 2025 · 11:45
Swyx argues that 2025 is the year of agent engineering, pivoting the AI Engineer Summit to focus exclusively on agents and explaining why agents are now viable due to advancements in reasoning, tool use, model diversity, and a 1000x cost reduction in GPT-4-level intelligence over 18 months. He highlights that OpenAI's ChatGPT grew 33% to 400M users in three months after shipping agentic O1 models, and projects it will reach 1B users by year-end. Swyx cites Simon Willison's crowdsourced definitions and OpenAI's new agent definition, while calling for a halt to demos of flight booking agents and Reddit astroturfing. He frames agent engineering as the evolution of AI engineering, distinct from MLE and software engineering, with practical use cases like coding and support agents having clear product-market fit.

Rethinking how we Scaffold AI Agents - Rahul Sengottuvelu, Ramp
Mar 19, 2025 · 16:32
Rahul Sengottuvelu, Head of Applied AI at Ramp and co-founder of Cohere.io, argues that building AI agents should follow the 'bitter lesson' from AI research: systems that scale with compute outperform handcrafted, deterministic code. He illustrates this with Ramp's switching report agent, which ingests arbitrary CSV files from third-party card providers. Three approaches are compared: manually coding parsers for the 50 most common vendors, using LLMs only for column classification, and a fully LLM-driven method where the model writes and runs pandas code via a code interpreter, repeated 50 times in parallel. The last approach, though using 10,000x more compute, costs under a dollar and generalizes better, saving Ramp far more in failed transactions. Sengottuvelu also demonstrates a prototype email client where the backend is an LLM with access to a code interpreter and the user's Gmail token; the LLM renders the UI as Markdown and handles clicks by re-prompting itself, simulating a full web app without traditional backend code. He contends that as models exponentially improve, shifting more execution into 'fuzzy' LLM compute—rather than rigid code—lets builders ride the trend for…

Navigating AI’s Frontier in 2025 - Grace Isford, Lux Capital
Mar 13, 2025 · 17:55
Grace Isford, partner at Lux Capital, argues that while 2025 is a 'perfect storm' for AI agents with reasoning models like O3 and R1, cheaper inference, and billions in infrastructure (e.g., Stargate, DeepSeek), agents still fail due to cumulative errors—decision, implementation, heuristic, and taste—exemplified by OpenAI Operator booking a flight incorrectly. She prescribes five strategies: curating proprietary and agent-generated data, building personalized evals for non-verifiable domains (e.g., seat preference), designing scaffolding that prevents cascading failures (citing Ramp's approach), treating UX as the moat (e.g., Codium, Harvey, TLDraw), and building multimodally with voice, smell via Osmo, and touch for embodiment. The talk, recorded at the AI Engineer Summit 2025 in NYC, closes with a call to reframe perfection through visionary product experiences.

How Windsurf writes 90% of your code with an Agentic IDE - Kevin Hou, Windsurf
Mar 11, 2025 · 20:50
Kevin Hou, head of product engineering at Codeium, introduces Windsurf as the first AI agent-powered editor, arguing that agents are the future of software development. The editor is built on three principles: trajectories, a unified timeline that lets the agent read the user's mind and execute commands like 'continue my work'; meta-learning, which auto-generates memories of user preferences and codebase context; and scaling with intelligence, removing legacy features like chat and @-mentions as models improve. Windsurf launched three months ago and has already generated 4.5 billion lines of code, with 90% of code now written via its Cascade agent. Hou emphasizes that the agent reduces human-in-the-loop effort by inferring context, running terminal commands in the user's environment, and dynamically retrieving documentation—allowing developers to contribute less input while getting more production-ready output.

Reinforcement Learning for Agents - Will Brown, ML Researcher at Morgan Stanley
Mar 7, 2025 · 18:17
Will Brown, a machine learning researcher at Morgan Stanley, argues that reinforcement learning (RL) is the essential path to unlocking autonomous AI agents, citing DeepSeek's R1 and OpenAI's Deep Research as proof: R1 used GRPO to make models learn chain-of-thought reasoning without manual data, and Deep Research applies end-to-end RL for up to 100 tool calls. He shares his own open-source single-file GRPO script—a 1B-parameter LLaMA model trained on math questions that demonstrated self-correction and improved accuracy, sparking community forks and blog posts. Brown introduces 'rubric engineering' as a new practice akin to prompt engineering, where reward rules (e.g., XML structure, integer answers) guide model improvement, and warns about reward hacking. He previews his current work: a framework for RL inside multi-step environments, letting developers reuse existing agent code for training. The talk concludes that fine-tuning and RL remain relevant as open-source catches up, and that skills like building evals and prompts translate directly to the RL era.

OpenAI for VP's of AI + Advice for Building Agents
Mar 5, 2025 · 16:52
OpenAI's Toki Sherbakov and Prashant Mital explain how enterprises adopt AI through a three-phase journey: building an AI-enabled workforce with ChatGPT, automating operations with APIs, and infusing AI into end products. They detail a Morgan Stanley case study where retrieval methods improved an internal knowledge assistant's accuracy from 45% to 98%. The pair define agents as models with instructions, tools, and self-terminating execution loops, then share four field lessons: build with primitives before frameworks, start with a single purpose-built agent, graduate to a network of specialized agents with handoffs for complex tasks, and keep prompt instructions simple while running guardrails in parallel using fast models like GPT-4o mini for safety and reliability.

Building Agents with Model Context Protocol - Full Workshop with Mahesh Murag of Anthropic
Mar 1, 2025 · 1:44:12
Mahesh Murag of Anthropic presents the Model Context Protocol (MCP) as an open standard that replaces fragmented integrations with a single protocol for connecting AI systems to data sources, enabling context-rich AI applications and agentic experiences. He explains MCP's philosophy, inspired by APIs and LSP, and its three interfaces: tools (model-controlled), resources (application-controlled), and prompts (user-controlled). Murag highlights adoption with over 1,100 community-built servers and official integrations from companies like Cloudflare and Stripe. He demonstrates building agents with MCP using the MCP-Agent framework, showing how agents can use tools dynamically and composably across hierarchical systems. Future plans include remote server support with OAuth 2.0, a centralized registry for discovery and verification, and enabling agents to self-evolve by dynamically finding new capabilities via registry search.

AI Agents, Meet Test Driven Development
Feb 22, 2025 · 29:10
Anita from Vellum argues that teams using test-driven development build more reliable AI systems, and she breaks down how to apply TDD to agentic workflows across four stages: experiment, evaluate, deploy, monitor. She outlines five levels of agentic behavior (L0 to L4), noting that most production systems today are at L1 (tool use), while L2 systems that plan and reason with models like O1 and DeepSeek-R1 will see most innovation this year. She introduces Vellum's new open-source Workflows SDK, which keeps code and UI in sync, and demonstrates her own SEO agent that automates keyword research, content analysis, and writing via a writer-editor evaluation loop. The agent, built on Vellum Workflows, takes a keyword like "chain-of-thought prompting" and produces a first draft with a latency of around 118 seconds, incorporating competitive analysis and iterative feedback. She emphasizes that success now depends on orchestration techniques (prompt chaining, RAG, memory) rather than just model performance.

Don't just slap on a chatbot: building AI that works before you ask
Feb 22, 2025 · 5:46
Arthur Objartel of Evil Martians argues that slapping chatbots on products is misguided, advocating instead for proactive AI that assists users before they ask. Drawing from his work on Tegon, an AI-powered issue tracker, he demonstrates three interaction modes—suggestion, action, and question+action—that operate within the natural workflow without chat interfaces. The AI triggers contextually relevant questions and actions, such as splitting issues or suggesting subtasks, all reversible with one click. He proposes three rules: AI supplements user agency, offers recommendations not force, and integrates without breaking flow. Examples extend to code editors catching pitfalls and design tools promoting accessibility. The talk challenges the status quo, urging experimentation beyond reactive chat interfaces.

The Price of Intelligence - AI Agent Pricing in 2025
Feb 22, 2025 · 20:38
Shitej, co-founder and CTO of Orbe, argues AI agent pricing must continuously evolve, citing Intercom's 99 cent per resolution outcome model, Clay's prospecting credits, and Cursor's tiered usage limits. He stresses aligning pricing with target audience—SMB vs. enterprise—and maintaining simplicity and predictability. Cost structure is key: Character.AI optimized inference to support 100M DAUs, while Jasper leveraged a model decision engine to offer unlimited credits. Shitej emphasizes flexibility, noting OpenAI's price drops force repricing, and predicts 2025 will see more unlimited plans, outcome-based pricing with SLAs, and greater investment in pricing R&D for usage visibility.

WTF do people use Open Models for??
Feb 22, 2025 · 28:01
Eugene Cheah of Featherless.ai breaks down how individuals and enterprises actually use open-source AI models, based on platform data. DeepSeek R1 dominates individual usage, but Mistral Nemo 8B remains the top enterprise model due to production stickiness and Apache 2.0 licensing. Creative writing and roleplay account for 30–40% of all traffic, with over 60% of users in that segment being women; coding copilots and agents make up 20–30%, driven by 'vibe coding' and token-hungry workflows like Kline. RAG and ChatGPT clones represent 20%, while agentic workflows (10–20%) succeed with human-in-the-loop designs. Cheah advises enterprises to aim for 80% automation with escape hatches, and warns against chasing 100% reliability. He concludes by introducing Quirky, a post-transformer hybrid built for $100k.

This video was edited with AI agent. But how?
Feb 22, 2025 · 5:00
Muhteşem from Re-Skill presents the world's first open-source video editing agent, built in collaboration with Diffusion Studio, which uses LLM-generated code to edit videos via a browser-based UI. The agent starts a Playwright browser session and connects to operator.diffusion.studio, a web app built with Diffusion Studio Core that renders video directly in the browser using WebCodecs. Three tools drive the process: VideoEditingTool generates and runs JavaScript code based on user prompts, DocsSearchTool uses RAG to pull relevant information from llms.txt, and VisualFeedbackTool samples one frame per second and decides whether to proceed with rendering or refine further. File transfers between the Python agent and browser happen via Chrome DevTools Protocol, and for scalability, the agent can connect to a remote GPU-accelerated browser through WebSocket. The project emerged from Re-Skill's need for automatic video editing for personalized learning, where they found FFmpeg too limiting and turned to Diffusion Studio Core for its intuitive API and reliable in-browser rendering.

Voice Agents: the good, the bad, and the ugly
Feb 22, 2025 · 18:48
Eddie Siegel, CTO at Fractional AI, details the real-world challenges of building an AI voice agent for consulting-style interviews, moving from proof of concept to production. The agent, designed to conduct hundreds of research interviews simultaneously, struggled with hallucinations, latency, and Whisper transcription errors like 'creeping dippity dippity dippity dippity'. To constrain the LLM's tendency to chitchat and rabbit-hole, Siegel's team introduced a 'drift detector' agent and a 'next question' agent, both running as separate text-based side threads. They also added tool use for question navigation, explicit goals and priorities for follow-ups, and a 'transcript cleaner' to hide garbled outputs from the user interface. To move beyond vibes-driven iteration, they built an automated eval suite with metrics like clarity and completeness, using synthetic conversations with personas such as a 'snarky teenager in charge of a Fortune 500 company' to simulate diverse interviewees. The talk offers a no-nonsense playbook for getting voice agents into production, focusing on out-of-band checks, tool use, and evals even without objective ground truth.

Beyond APIs: How AI Web Agents Are Automating the "Long Tail" of Knowledge Work
Feb 22, 2025 · 17:44
Arjun and Bhavani present Rtrvr.ai, a universal AI web agent using a text-based approach to autonomously perform tasks across multiple browser tabs at under a penny per page. They argue text-based reduces hallucination compared to vision-based agents like OpenAI's operator, and being a Chrome extension avoids password sharing and accesses logged-in content. Features include deep research navigating multiple pages, dynamic function calling for any third-party API, and graph generation. Rtrvr extracts structured data, performs actions like clicking dropdowns, and processes tabs simultaneously, automating the long tail of knowledge work including market research, LinkedIn automation, and WhatsApp messaging. They claim a distributed subtask approach lowers failure rates and envision collaborative dataset construction, e.g., aggregating local government events.

Agent Evals: Finally, With The Map
Feb 22, 2025 · 13:31
Ari Helczak from Rootsignals presents a systematic map for AI agent evaluation, dividing it into semantic and behavioral parts. Semantic evaluation covers single-turn virtues like coherence and safety, plus multi-turn aspects such as conversation consistency and reasoning traces. Behavioral evaluation addresses tool selection, instruction following, and multi-step goal convergence. Helczak grounds truthfulness in RAG and goal achievement in tool utility, and introduces eval-ops as a double-tier approach to optimize both the agent and its judgment flow. He also highlights cost, latency, tracing, and offline versus online testing as practical considerations, noting the map will quickly become obsolete as the field evolves.

Your Evals Are Meaningless (And Here’s How to Fix Them)
Feb 22, 2025 · 18:50
In this AI Engineering Summit talk, HoneyHive co-founder exposes why most LLM evaluations are meaningless due to criteria drift—where evaluator criteria misalign with user needs—and dataset drift, where test cases don't reflect real-world queries. He argues that static evaluation frameworks from tools like LangChain or Ffragas fail because they measure generic metrics rather than business-specific relevance, citing an e-commerce recommendation system that looked perfect in testing but broke in production. The fix is a three-step iterative alignment process: align evaluators with domain experts via continuous critique and few-shot examples; keep datasets alive by logging production underperformance and flowing those cases back into the test bank; and track alignment over time using F1 scores for binary judgments. Practical advice includes customizing the LLM evaluator prompt, starting with 20 domain expert examples in spreadsheets, and avoiding templated metrics. The speaker emphasizes that evals must evolve continuously, just like the LLM application itself, or they become meaningless.

Mission-Critical Evals at Scale (Learnings from 100k medical decisions)
Feb 22, 2025 · 12:15
Christopher Lovejoy, a medical doctor turned AI engineer, explains how Anterior built a real-time reference-free evaluation system to scale mission-critical AI decisions in healthcare to 100,000 per day while maintaining trust. He shows that human reviews don't scale (50 clinicians needed for 5,000 daily reviews) and offline evals miss new edge cases. Instead, Anterior uses an LLM-as-judge to assign confidence scores, dynamically prioritizing high-risk cases for human review. This 'validating the validator' system achieved a 96% F1 score in prior authorization, letting a team of under 10 clinical experts handle tens of thousands of cases. It provides real-time performance estimates, enables rapid error correction, and builds defensibility through proprietary data and iterations only possible at scale.

OpenLLMetry is all you need
Feb 22, 2025 · 9:12
Nir, CEO of Trace Loop, introduces OpenLLMetry, an open-source project extending OpenTelemetry for tracing and monitoring GenAI applications. OpenTelemetry, maintained by CNCF, standardizes logging, metrics, and traces across cloud environments, supported by platforms like Datadog, New Relic, and Grafana. OpenLLMetry provides over 40 automatic instrumentations for foundation models (OpenAI, Anthropic, Cohere), vector databases (Pinecone, Chroma), and frameworks (LangChain, LlamaIndex, CrewAI). These instrumentations emit logs, metrics, and traces in OpenTelemetry format, allowing users to send data to any supported observability backend with a configuration change, avoiding vendor lock-in.

Privacy First Enterprise AI: Building AI Agents that Never Leave Your Security Boundary
Feb 22, 2025 · 7:10
Steven Moon, founder of Aech AI, argues that enterprise AI agents should be deployed within existing security boundaries by treating them like human employees—using existing identity management, compliance frameworks, and audit tools rather than building parallel systems. He explains that IT departments will evolve into HR departments for AI agents, provisioning them through active directory and applying standard security policies. Moon highlights email as a powerful medium for agent-to-agent communication, where every interaction is logged and auditable through existing systems. He advocates for enhancing current enterprise platforms like Microsoft 365 and Azure with AI agents instead of creating new interfaces, noting that the era of mandatory translation layers between humans and machines is ending.

How to Improve Your Agents: Academic Lit Review
Feb 22, 2025 · 39:02
Joe from Columbia University and founder of Arklex AI explains research on AI agents, focusing on improving their reasoning and planning through self-reflection, test-time compute, and tree search methods, without relying on human supervision. He details the TriPath method, which uses larger models to edit smaller models' feedback for better self-improvement, achieving up to 48% accuracy on math benchmarks. He shows how multi-color tree search (MCTS) with contrastive reflection and multi-agent debate (RMCTS) outperforms other search methods on Visual Web Arena and OS World, achieving top non-trained results. He proposes exploratory learning, where models learn from search trajectories rather than optimal actions, improving performance under compute budgets. Finally, he presents the Arklex open-source agent framework, combining machine learning, systems, and security for practical multi-agent orchestration.

The Model Isn’t Wrong—You’re Just Bad at Prompting
Feb 22, 2025 · 8:54
Dan from PromptHub argues that prompt engineering remains critical for improving LLM outputs, covering Chain of Thought, few-shot, and meta prompting techniques. Chain of Thought breaks problems into sub-problems and is built into reasoning models; few-shot prompting works best with just one or two diverse examples, but can degrade performance on reasoning models like O1 and R1. Meta prompting uses LLMs to write or refine prompts, with PromptHub offering model-specific enhancers. For reasoning models, Dan advises minimal prompting, encouraging more reasoning instead of few-shot, and avoiding instructing the model on how to reason. Free resources include PromptHub's templates, the AutoReason prompt, and the Prompt Engineering Substack.

Stop Guessing: Build Robust AI with Layered CoT
Feb 22, 2025 · 10:16
Manish Senwal, Director of AI at NewsCorp, presents Layered Chain-of-Thought (CoT) prompting with Multi-Agent Systems to build transparent, self-correcting AI. Standard CoT is sensitive to prompt wording and lacks verification, letting errors cascade. Layered CoT verifies each reasoning step against a knowledge base, enabling early correction and reducing input sensitivity. The system iteratively generates thoughts, verifies them, and proceeds only with validated information. Combined with specialized agents—each handling distinct tasks like pedestrian detection—this yields fault-tolerant, scalable systems. The talk argues for structured, validated reasoning over larger models, referencing the arXiv paper (2501.18645).

Tool Calling Is Not Just Plumbing for AI Agents — Roy Derks
Feb 22, 2025 · 25:18
Roy Derks argues that tool calling is the most critical yet overlooked component of AI agents, far more than mere 'plumbing.' He contrasts traditional tool calling—where developers manually manage callbacks, retries, and errors within the agent loop—with embedded tool calling, a black-box approach used by frameworks like LangChain’s createReactAgent. Derks advocates for separation of concerns via the Model Context Protocol (MCP) from Anthropic, which splits tool logic into MCP servers communicating with clients, and via standalone tool platforms such as IBM’s wxflows, Composio, and Toolhouse that let teams build tools once and reuse them across LangChain, CrewAI, or AutoGen. He also introduces dynamic tools, where an agent generates queries on the fly—e.g., using GraphQL or SQL schemas—instead of defining hundreds of static tools, noting that LLMs like Claude handle GraphQL well but may hallucinate on deeply nested schemas. The episode emphasizes that 'an agent is only as good as its tools' and provides practical guidance on designing tool descriptions (which act like system prompts) and output schemas to enable type-safe, chainable tool calls.

Your LLM Ran Out of Knowledge — Now What?
Feb 22, 2025 · 12:51
Speaker 1 presents a technique for applying LLMs to low-knowledge domains like corporate negotiations and geopolitics, where structured training data is scarce. The method uses a parsing engine to identify the problem type and match domain-specific heuristics—binary rules like 'must prioritize agreements with highest combined value' for negotiations or 'must have three independent paths for critical resources' for geopolitics. After reformatting the input for consistency, the system sends it to a reasoning model (e.g., Anthropic's Claude with a world sim) that applies the rule set alongside powerful reasoning capabilities. The demo shows an intelligence estimate scenario where the model identifies geopolitics, reformats the query, and evaluates options against rules—eliminating those that breach heuristics (e.g., scenario five violating a rule on control thresholds). The approach enables generating dozens of scenarios in a fraction of the time a human expert would take, allowing practitioners to explore more options and find novel solutions while keeping a human in the loop for subject-matter oversight.

The Hidden Costs of Building Your Own RAG Stack — Ofer Vectara
Feb 22, 2025 · 15:14
Ofer from Vectara argues that building your own RAG stack comes with seven hidden costs: irrelevant responses and hallucinations, high latency, scaling and cost issues, security and compliance challenges, vendor chaos from multiple components, unsustainable expertise requirements, and limited non-English language support. He details how each pitfall compounds at enterprise scale, citing the need for hybrid search, reranking, and continuous evaluation to maintain quality. Vectara offers a turnkey RAG-as-a-service platform that handles parsing, chunking, embedding, vector search, retrieval, and hallucination detection. Ofer highlights Vectara’s open-source HHEM model for hallucination evaluation, downloaded over 3 million times, and a leaderboard showing some LLMs hallucinate at rates above 10%. The platform supports deployment in SaaS, VPC, or on-premises, with built-in access control and explainability.
Powered by PodHood