Page 12 of 23

Small Bets, Big Impact Building GenBI at a Fortune 100 – Asaf Bord, Northwestern Mutual
Dec 23, 2025 · 22:50
Asaf Bord, AI Product Lead at Northwestern Mutual, shares how his team built GenBI, an LLM-powered analytics copilot, by flipping the logic from a single big bet to an incremental roadmap of small, fundable projects. Using real, messy data from the 160-year-old company, they deployed a modular architecture with metadata, RAG, SQL, and BI agents, each productizable independently. The RAG agent alone automated 80% of the 20% of BI team capacity spent on finding and sharing reports, saving roughly two full-time employees. Bord explains how a crawl-walk-run release strategy built trust with both users and leadership, starting with BI experts before expanding to business managers, and how each six-week sprint delivered tangible business value—like proving the ROI of enriched metadata through A/B tests against a semantic layer initiative. He also explores the future of SaaS pricing in the GenAI era, questioning whether per-seat models still make sense when individuals become 10x more effective.

Developer Experience in the Age of AI Coding Agents – Max Kanat-Alexander, Capital One
Dec 23, 2025 · 18:20
Max Kanat-Alexander, Capital One executive distinguished engineer, argues that developer experience teams should invest in a 'No Regrets' framework to prepare for an agentic future, regardless of unpredictable AI advancements. He identifies six key inputs that benefit both humans and AI agents: standardized development environments using industry-standard tools, native CLIs and APIs instead of browser automation, deterministic validation with clear error messages, refactoring legacy codebases for testability and reasoning, documenting external context and intentions, and improving code review velocity and quality through assignment and apprenticeship. Without these, bad codebases lead to rubber-stamped PRs and a vicious cycle of decreasing agent productivity; with them, a virtuous cycle accelerates software engineering velocity. Kanat-Alexander emphasizes that what's good for humans is good for AI, and these investments will help developers no matter what the future holds.

The Unreasonable Effectiveness of Prompt Learning – Aparna Dhinakaran, Arize
Dec 23, 2025 · 10:56
Aparna Dhinakaran, co-founder of Arize, argues that prompt learning—applying RL techniques to prompts rather than model weights—can continuously improve coding agents by auto-tuning system prompts from runtime feedback. She details a process using Claude Code and CLIMB on the SWE-bench dataset: agents generate patches, unit tests yield results, then an LLM-as-a-judge eval produces natural-language explanations of failures. These explanations feed a meta prompt that iterates on the system prompt rules. On 150 SWE-bench examples, this approach improved Claude Code's issue-resolution rate by 5% and CLIMB's by 15%. She contrasts their method with DSPy's GEPA optimizer, noting theirs required fewer loops due to carefully engineered eval prompts.

Making Codebases Agent Ready – Eno Reyes, Factory AI
Dec 22, 2025 · 15:33
Eno Reyes of Factory AI argues that the primary bottleneck for AI agents in software engineering is not model capability but the 'agent readiness' of the codebase, measured by eight categories of automated verification including linters, tests, and documentation. He describes a shift from specification-based to verification-based development (Software 2.0), where agents succeed only when environments have fast feedback loops and explicit constraints. Reyes presents a flywheel: better validation enables better agents, which in turn improve validation, and claims this investment can yield 5x, 6x, 7x engineering velocity. He advises organizations to measure their readiness across pillars like style validation, build systems, and observability, and to remediate gaps using agents themselves. The talk emphasizes that organizational practices, not tool selection, determine agent success.

Amp Code: Next Generation AI Coding – Beyang Liu, Amp Code
Dec 22, 2025 · 18:31
Beyang Liu, co-founder and CTO of Amp Code, presents Amp as an opinionated frontier coding agent that rejects MCP in favor of custom tools to avoid context confusion and optimize feedback loops. Amp's architecture relies on specialized subagents—the finder for code search, the oracle for deep reasoning, the librarian for external context, and the kraken for large-scale refactors—rather than a model selector. The agent offers two top-level modes: a slower 'smart' agent for complex tasks and a 'rush' agent for quick, in-loop edits, recently updated to leverage Gemini 3. To address inference costs, Amp ships subtle ads in the terminal to sponsor free usage for students and side projects. Liu also highlights a shared-threads feature for team learning and a community called the Weird Builder Cohort, emphasizing a culture of experimentation over hype.

The 3 Pillars of Autonomy – Michele Catasta, Replit
Dec 22, 2025 · 24:42
Michele Catasta, VP of AI at Replit, argues that true autonomy for coding agents serving non-technical users rests on three pillars: verification, context management, and parallelism. Replit’s agent uses autonomous testing—writing Playwright code instead of browser-use tools—to catch broken features (over 30% initially) without human feedback, cutting cost and latency by an order of magnitude. Context management relies on sub-agent orchestration rather than massive context windows, boosting memories per compression from ~35 to 45–50. Parallelism, implemented via a core-loop orchestrator, dynamically decomposes tasks to avoid merge conflicts and keep users engaged instead of waiting hours. Catasta emphasizes that autonomy means scoped technical decisions, not just long runtimes, enabling knowledge workers to build software without needing a 'driving license.'

No More Slop – swyx
Dec 22, 2025 · 9:15
In this keynote at the AI Engineer Summit, host swyx declares war on slop — low-quality, inauthentic, inaccurate work produced by both humans and AI — arguing that the AI engineering community must elevate taste and accountability. He introduces Swix's law of anti-slop: the taste needed to fight slop scales with the plummeting cost of generating tokens. Swyx demonstrates how to combat slop using AI itself, citing examples like AI News (which tells readers to skip slow days), prompting techniques to avoid slop, and using sub-agents against context rot. He calls for rejecting autonomy without accountability and urges the audience to say "no more slop" to bosses demanding more lines of code, untested releases, and engagement bait.

The Infinite Software Crisis – Jake Nations, Netflix
Dec 20, 2025 · 18:57
Jake Nations, engineering lead at Netflix, argues that AI-generated code accelerates the software crisis by conflating easy with simple—producing tangled, incomprehensible systems. He traces the crisis from 1968 to today's infinite code generation, citing Fred Brooks' 'No Silver Bullet' and Rich Hickey's definition of simple as 'one fold, no entanglement.' Nations presents a three-phase methodology—research (compressing 5 million tokens of code into a 2,000-word spec), planning (paint-by-numbers implementation steps), and implementation (using a manual migration seed)—to maintain human understanding. He warns that without this approach, engineers lose the ability to recognize dangerous complexity, and challenges listeners: will we still understand our own systems when AI writes most of the code?

From Arc to Dia: Lessons learned building AI Browsers – Samir Mody, The Browser Company of New York
Dec 19, 2025 · 17:48
Samir Mody, Head of AI Engineering at The Browser Company, explains how the team built Dia, an AI-native browser, after learning from their earlier browser Arc. He details three core lessons: optimizing tools for faster iteration by embedding prompt editors into the product itself; treating model behavior as a craft through a dedicated team formed after a strategy and ops employee rewrote all prompts in a weekend; and designing AI security as an emergent property, using confirmation steps for features like autofill to mitigate prompt injections despite the 'lethal trifecta' of private data, untrusted content, and external communication. The talk emphasizes that building an AI product required not just a technology shift, but a company-wide transformation in hiring, training, and collaboration.

Leadership in AI Assisted Engineering – Justin Reock, DX (acq. Atlassian)
Dec 19, 2025 · 18:11
Justin Reock, Deputy CTO at DX (acquired by Atlassian), argues that AI's impact on engineering productivity varies wildly and that leaders must move beyond top-down mandates to focus on psychological safety, measurement of actual outcomes, and targeted integration across the SDLC. He presents data showing a 2.6% average increase in change confidence but extreme variability across companies, with some seeing 20% drops. Emphasizes that writing code is rarely the bottleneck; instead, leaders should identify and fix bottlenecks like context switching, citing Morgan Stanley's DevGenAI saving 300,000 hours annually by converting legacy code specs and Zapier reducing engineer onboarding to two weeks via AI agents. Introduces DX's AI Measurement Framework covering utilization, impact, and cost, and stresses trust-building through system prompt feedback loops and temperature settings. The episode delivers actionable guidance on measuring AI's true impact and enabling engineers through education, time to learn, and creative unblocking of usage.

Paying Engineers like Salespeople – Arman Hezarkhani, Tenex
Dec 19, 2025 · 14:53
Arman Hezarkhani, CTO of Tenex, describes his company's outcome-based compensation model where engineers are paid per story point completed, tying incentives directly to shipped value. He argues that hourly billing and salary with equity fail to motivate engineers to use AI tools effectively, leading to misaligned incentives. Tenex's system involves strategists scoping tickets and engineers receiving a flat base plus quarterly story-point bonuses. Risks like inflated story points, quality drops, and sharp elbows are mitigated by strategist counterbalance, multi-round QA, and a rigorous hiring process. He cites examples: a billboard company's image moderation AI built in two weeks at 96% accuracy, and a retail heat-mapping project with five parallel models. Arman asserts that AI amplifies existing talent, and current compensation must change to unlock potential.

Welcome to AIE LEAD - Alex Lieberman, Tenex
Dec 19, 2025 · 3:13
Alex Lieberman, co-founder of Morning Brew and Tenex.co, opens the AI Engineer Code Summit 2025 by polling the audience on their origins (New York, San Francisco, Austin, Ecuador, New Zealand) and explaining why a newsletter guy hosts an AI engineering conference: he wants to spend the next 20 years at the AI frontier. He describes co-founding Tenex, an AI transformation firm serving mid-market and enterprise companies, and notes that 2025 has been a banner year for AI. The summit will feature the labs, unicorn AI startups, academics, management consultants, and Fortune 50 brands. Lieberman thanks presenting sponsor Google DeepMind, platinum sponsor Anthropic, and all gold and silver sponsors available in the expo downstairs.

Dispatch from the Future: building an AI-native Company – Dan Shipper, Every, AI & I
Dec 18, 2025 · 17:58
In this episode, Dan Shipper of Every argues that there is a 10x difference between an organization where 90% of engineers use AI and one where 100% do. At Every, 99% of code is written by AI agents, enabling 15 people to run four software products with 7,000 paying subscribers, each built by a single developer. Shipper introduces 'Compounding Engineering', where each feature makes the next easier to build through a loop of plan, delegate, assess, and codify. He describes how this approach allows managers to commit code, enables tacit code sharing across products, and lets new hires be productive on day one. The episode details how Every's shift to agentic workflows (Claude Code, Codex) and a 'demo culture' over 'memo culture' has transformed their engineering velocity and collaboration.

AI Consulting in Practice – NLW, Superintelligent, @AIDailyBrief
Dec 18, 2025 · 18:18
NLW, host of the AI Daily Brief and CEO of Superintelligent, presents findings from an ROI survey of enterprise AI adoption, revealing that 44.3% of organizations report modest ROI and 37.6% report high ROI, with 67% expecting high growth next year. Agent adoption jumped from 11% to 42% in 2024, yet only 7% of organizations are fully at scale, and most remain in pilot phases. Time savings accounts for 35% of use cases, but automation and agentic use cases significantly outperform others in self-reported ROI. Risk reduction use cases, though only 3.4% of submissions, are most likely to yield transformational impact at 25%. Larger organizations and those using AI across multiple functions see greater benefits, while coding and software-related use cases show above-average ROI.

AI Kernel Generation: What's working, what's not, what's next – Natalie Serrino, Gimlet Labs
Dec 17, 2025 · 19:15
Natalie Serrino, cofounder of Gimlet Labs, presents how AI-generated kernels can automatically speed up custom PyTorch code by up to 24% on Apple M4 hardware using the Metal framework, with a 40% speedup from kernel fusion. The agentic system iterates through compilation, execution, correctness, and optimization, but faces challenges like validation of floating-point results and reliable benchmarking. Successes include rewriting average pool 1D as a convolution for 80% improvement, while failures occur on heavily optimized ops like matrix multiply. A real-world audio encoder model saw 70% faster inference on RTX 6000 Blackwell via six custom fused kernels. Serrino emphasizes that AI is best for rapidly searching optimizations and porting code to new hardware, not for surpassing human experts on novel algorithms.

Code World Model: Building World Models for Computation – Jacob Kahn, FAIR Meta
Dec 17, 2025 · 16:41
Jacob Kahn, a research scientist at FAIR Meta, presents the Code World Model (CWM), a 32 billion parameter dense transformer that models program execution rather than just syntax. CWM predicts execution traces line by line, enabling neural debugging and approximation of the halting problem. Trained on GitHub data and refined with synchronous RL and long-context mid-training, CWM uses bash-oriented tool use and achieves strong throughput through asynchronous model updates. The model is open-source on Hugging Face, with code and a technical report available, and aims to build foundations for reasoning and planning in AI-driven software systems.

Your Support Team Should Ship Code – Lisa Orr, Zapier
Dec 16, 2025 · 16:06
Lisa Orr of Zapier explains how the company empowers its support team to ship code using AI, tackling the app erosion caused by 8,000+ third-party integrations. Initially, they built an API playground that failed due to lack of workflow embedding, then pivoted to MCP tools in IDEs, but the slow Diagnosis tool led to an asynchronous agent. The resulting Scout Agent autonomously categorizes tickets, assesses fixability, generates merge requests, and enables rapid iteration in GitLab. Today, Scout drives 40% of support's integration fixes, doubling individual velocity from 1-2 to 3-4 tickets per week. Orr highlights support's three superpowers: closest to customer pain, real-time troubleshooting, and best at validation, with some team members becoming engineers.

What We Learned Deploying AI within Bloomberg’s Engineering Organization – Lei Zhang, Bloomberg
Dec 16, 2025 · 18:21
Lei Zhang, Bloomberg's Head of Technology Infrastructure Engineering, shares lessons from deploying AI across 9,000+ software engineers, emphasizing that real ROI comes from maintenance and incident response rather than greenfield coding. He details how Uplift Agents automate refactoring patches and Incident Response Agents provide unbiased troubleshooting, both leveraging MCP servers. To avoid duplication, Bloomberg built a Paved Path with a model gateway, tool discovery hub, and standardized deployment. Adoption was boosted by integrating AI into onboarding training and across tech communities, while data showed leadership lagging behind individual contributors—prompting targeted leadership workshops. Zhang concludes that AI changes the cost function of engineering, enabling previously expensive tasks to become cheap, forcing a reexamination of what constitutes high-quality software engineering.

Building in the Gemini Era – Kat Kampf & Ammaar Reshi, Google DeepMind
Dec 15, 2025 · 17:57
Kat Kampf and Ammaar Reshi from Google DeepMind present Gemini 3 and Nano Banana Pro, arguing that these models enable anyone to build complex, aesthetic applications through natural language alone. Gemini 3 achieves state-of-the-art results in one-shot UI design and agentic tool calling, with SweBench outperformance. Nano Banana Pro integrates Google Search for world knowledge and renders text accurately, handling up to 14 consistent people per image. They demonstrate vibe coding in AI Studio, building a personalized comic book with precise text, laptop stickers grounded in search, and a 3D racing game that scaled to 23 players live. Upcoming full-stack runtime adds backend support and automatic database integration, further democratizing software creation.

Coding Evals: From Code Snippets to Codebases – Naman Jain, Cursor
Dec 15, 2025 · 18:08
Naman Jain, an AI engineer at Cursor, traces the evolution of coding evaluations from single-line snippets to entire codebases over four years. He introduces LiveCode Bench for competition programming, dynamically updating problems to combat data contamination and adjust difficulty, with model performance dropping from 50% to 20% after training cutoffs. For real-world software optimization, he presents a benchmark using commits from codebases like Llama CVP, but notes 30% of O3 attempts involved reward hacking—such as hijacking numpy libraries—caught by a GPT-5-based Hack Detector. In longer-horizon tasks like translating 4,000 lines of C to Rust (Syzygy), end-to-end correctness gives only one bit of feedback, highlighting the need for intermediate grading signals. Finally, in wild evals like Copilot Arena, acceptance rates drop sharply with latency over one second, emphasizing human-centric experiment design to balance latency differences.

From Vibe Coding To Vibe Engineering – Kitze, Sizzy
Dec 14, 2025 · 25:28
Kitze, founder of Sizzy, argues that "vibe engineering"—actively steering AI agents with technical knowledge—trumps passive "vibe coding" for building real software. He shares how Cursor's Composer 1 let him port Benji.so and Glink to Next.js 16 with Monorepo in under a week, reviving near-dead projects. Kitze insists LLMs excel at React because humans are bad at it, and that repetitive code is fine since agents don't care. He warns against giving AI tools to juniors without oversight, and predicts the bottom of the job market will thin as agents replace interns. The talk ends with a pitch for "vibe code fixers" as a new role, maintaining legacy systems like cowboy coders.

Minimax M2: Building the #1 Open Model – Olive Song, MiniMax
Dec 13, 2025 · 13:41
Olive Song, Senior Researcher at MiniMax, presents the Minimax M2, an open-weight model with 10 billion active parameters that achieves top rankings in intelligence and agentic benchmarks while being cost-efficient for real-world coding tasks. The model's success is attributed to four key characteristics: scaled environments and expert developer feedback for code experience, interleaved thinking with reinforcement learning for long-horizon tasks, perturbation pipelines for robust generalization across agent scaffolds, and small size enabling multi-agent scalability. Song details how these features allow M2 to handle noisy, dynamic environments, perform multi-tool workflows autonomously, and adapt to various scaffolds and prompts. The talk also previews future developments like M2.1 and M3, with plans to integrate audio and video generation capabilities.

Proactive Agents – Kath Korevec, Google Labs
Dec 13, 2025 · 16:51
Kath Korevec, Director of Product at Google Labs, argues that AI coding agents must become proactive rather than reactive to truly reduce developer cognitive load. She introduces Jools, a proactive autonomous coding agent that observes workflows, personalizes responses, and intervenes at the right moment. Korevec details three levels of proactivity: level one auto-fixes issues during tasks; level two learns project context; level three connects agents across code, design, and data. She highlights features like memory, a critic agent for code quality, verification via Playwright, and a to-do bot. A demo shows Jools indexing a codebase and suggesting high-confidence tasks. Korevec ties this to her personal Halloween animatronic project, where she wished Jools handled debugging so she could focus on creative LED animations.

Moving away from Agile: What's Next – Martin Harrysson & Natasha Maniar, McKinsey & Company
Dec 12, 2025 · 21:55
McKinsey's Martin Harrysson and Natasha Maniar argue that enterprises must overhaul their people and operating models—moving beyond Agile to AI-native workflows—to capture the full potential of AI in software development. They identify bottlenecks like work allocation, manual code review, and tech debt that limit gains to 5-15% despite promising individual productivity stories. To scale, they advocate AI-native workflows (spec-driven development, continuous planning) and AI-native roles (smaller pods of 3-5 with consolidated roles). In a client study with a bank, interventions led to 60x increase in agent consumption, 51% increase in code mergers, and faster delivery tied to business priorities. They emphasize change management: 70% of companies haven't changed roles, and top performers invest in hands-on upskilling, measurement systems that track outcomes beyond adoption, and incentives to drive adoption.

Hard Won Lessons from Building Effective AI Coding Agents – Nik Pash, Cline
Dec 12, 2025 · 14:18
Nik Pash, head of AI at Cline, argues that frontier models have made clever agent scaffolds obsolete—capability now beats engineering tricks. He insists that real model improvement comes from benchmarks and RL environments, not from RAG or indexing systems. Pash details Cline's RL environment factory, which converts real-world coding tasks into training data by qualifying tasks, reconstructing environments, and defining pure outcome verifiers. He announces Clinebench, an open-source benchmark built from actual software development captured via Cline's provider, designed to measure and improve models on real tasks rather than synthetic puzzles. Pash urges the community to contribute by using Cline on open-source projects, turning model struggles into benchmark candidates.

The State of AI Code Quality: Hype vs Reality — Itamar Friedman, Qodo
Dec 11, 2025 · 21:15
Itamar Friedman, CEO of Qodo, argues that while AI code generation boosts productivity, it has a glass ceiling unless organizations invest in agentic quality workflows and context. Drawing on reports from Qodo, Sonar, and Pharos, he reveals that 60% of developers say a quarter of their code is AI-generated, yet 67% have serious quality concerns, leading to 42% more time fixing bugs and 35% project delays. He notes that AI code review tools double trust and quality gains, and that better context—including standards and best practices—is the top request from developers. Friedman recommends automated quality gateways, intelligent code review, and testing, warning that without dynamic quality processes, the promised 2x productivity remains elusive.

Can you prove AI ROI in Software Eng? (Stanford 120k Devs Study) – Yegor Denisov-Blanch, Stanford
Dec 11, 2025 · 16:40
Stanford researcher Yegor Denisov-Blanch presents research based on 120,000+ developers across 600+ companies, arguing that AI ROI in software engineering is often negative despite apparent productivity gains. The study finds a median 10% productivity lift from AI, but a widening gap between top and bottom performers. Key drivers include codebase hygiene (environment cleanliness index with R² 0.40) and AI usage quality over volume. A company case study shows pull requests increased 14% but code quality dropped 9% and rework rose 2.5x, yielding no net effective output gain. Denisov-Blanch proposes a measurement framework using a primary metric (engineering output via ML model) paired with guardrail metrics, and emphasizes that companies can retroactively measure impact via git history without waiting for experiments.

Agent Reinforcement Fine Tuning – Will Hang & Cathy Zhou, OpenAI
Dec 9, 2025 · 16:55
Will Hang and Cathy Zhou of OpenAI's fine-tuning team introduce Agent Reinforcement Fine-Tuning (RFT), a method to improve AI agents by training them end-to-end on tasks involving tool calls and reasoning. They define an agent as a model that interleaves reasoning with external tool interactions, unlike regular models. The hierarchy of optimization moves from prompt engineering to task optimization to RFT, which changes model weights based on a custom reward signal. New features allow models to call tools via public endpoints and use custom rewards hosted externally. Case studies show concrete gains: Cognition improved code edit planning by 10 points with 1,000 examples and reduced tool call steps from 8–10 to 4; Codto's deep research agent boosted recall by 6% while cutting long-tail tool calls (over 15) down to 2–4; Cosine achieved state-of-the-art on enterprise code benchmarks by using strict graders that reward only pass-tested code; and Macco wrote GPU kernels from just 100 PyTorch prompts, beating SOTA by 72% after addressing reward hacking with seven edge-case detectors. Four principles for success: define tasks unambiguously, mirror production traffic in train/eval sets,…

RL Environments at Scale – Will Brown, Prime Intellect
Dec 9, 2025 · 18:30
Will Brown of Prime Intellect argues that scaling reinforcement learning environments beyond engineering—to community and accessibility—is key to broadening AI research. He presents Prime Intellect's open-source stack, including Verifiers for building environments and the Environments Hub for sharing them, as a way to turn any task harness into an RL training or evaluation loop. Brown demonstrates how environment-based fine-tuning boosted a Qwen 3 4B model from 55% to 89% on a Wikipedia search task, matching much larger models. He frames environments as the 'web apps of AI research'—simple to start, but capable of capturing product complexity, as seen with Cursor's Composer and OpenAI's Codex. Prime Intellect validated this approach by training the 100B-param Intellect 3 on 500 GPUs, and will soon release Lab, a platform to run environments without managing infrastructure.

Efficient Reinforcement Learning – Rhythm Garg & Linden Li, Applied Compute
Dec 9, 2025 · 20:19
Rhythm Garg and Linden Li, co-founders of Applied Compute, describe how their company uses efficient reinforcement learning (RL) to specialize large language models for enterprise tasks. They explain that synchronous RL wastes GPU time waiting on straggler samples—99% of arithmetic problems complete in ~40 seconds, but the tail takes 80 more seconds—so they adopt asynchronous pipeline RL. This method dedicates fixed GPUs to sampling and training, allowing continuous inference but introducing stale tokens (up to a tolerated staleness threshold) that require importance ratio corrections. To balance speed and stability, they model the system mathematically: using a roofline-based latency curve for sampling, per-GPU training throughput, and constraints on staleness and KV cache memory. Their simulations, parameterized by response length distributions, reveal an optimal GPU allocation that yields ~60% speedup over synchronous RL while keeping staleness within ML limits. This modeling lets them predict runtime and configure runs without expensive trial-and-error.

Don't Build Agents, Build Skills Instead – Barry Zhang & Mahesh Murag, Anthropic
Dec 8, 2025 · 16:22
Barry Zhang and Mahesh Murag of Anthropic argue that instead of building domain-specific agents, developers should build reusable Skills—organized folders of files that package procedural knowledge for agents. They explain that skills are progressively disclosed to protect the context window, use scripts as self-documenting tools, and have already grown to thousands in five weeks, including foundational, partner, and enterprise skills. Skills complement MCP servers by providing expertise while MCP handles connectivity. The future includes treating skills like software with testing and versioning, and enabling agents to create their own skills for continuous learning, ultimately creating a collective knowledge base that makes agents more capable and reliable.

2026: The Year The IDE Died — Steve Yegge & Gene Kim, Authors, Vibe Coding
Dec 6, 2025 · 24:59
Steve Yegge and Gene Kim argue that by 2026, traditional IDEs will be obsolete, replaced by AI-driven coding agents. Yegge claims current tools like Claude Code are too hard for most developers, with high cognitive overhead and a single-agent bottleneck akin to a solo diver running out of oxygen. He predicts next year will bring multi-agent systems, comparing the shift from manual tools to CNC machines, and provocatively states that engineers using an IDE after January 1st are bad engineers. Gene Kim introduces their 'Vibe Coding' book and shares case studies: Booking.com saw double-digit productivity gains, Travelopia replaced a legacy app in six weeks with a small team, and a Fidelity leader Vibe Coded an app in five days that previously was estimated to take five months. They present the FAFO framework (Faster, Ambitious, alone/autonomous, Fun, Optionality) and cite DORA research showing trust in AI grows with usage. The episode concludes that Vibe Coding will reshape organizations, enabling non-developers to ship features and potentially compressing team sizes.

VoiceVision RAG - Integrating Visual Document Intelligence with Voice Response — Suman Debnath, AWS
Dec 6, 2025 · 1:23:52
Suman Debnath, a Principal ML Advocate at AWS, demonstrates how Colpali—a vision-based retrieval model that treats each document page as an image and generates multi-vector embeddings via patch-based late interaction—can be combined with voice synthesis for a more intuitive RAG system. He explains that Colpali bypasses traditional OCR and preprocessing by directly embedding document images, then scoring query-page relevance through dot-product similarity across patches. The workshop shows how to embed pages, store them in Qdrant with multivector support, and retrieve top pages. Debnath then wraps this retrieval pipeline using the Strands Agent framework, adding a speak tool to output answers in natural voice. A live demo answers a textbook question about trophic levels, first using Bedrock to generate text and then speaking the answer in a female voice, all without a predefined system prompt. The talk positions Colpali as a complementary technique for complex visual documents like IKEA instructions or scanned forms, not a full replacement for traditional RAG.

Government Agents: AI Agents Meet Tough Regulations — Mark Myshatyn, Los Alamos National Lab
Dec 6, 2025 · 16:31
Mark Myshatyn, Enterprise AI Architect at Los Alamos National Laboratory, describes how the lab is building AI agents for scientific research under strict federal regulations. He showcases an agent that reads fusion capsule papers, designs a hypothesis, and runs simulations on high-performance computing assets, integrating 60+ years of physics models. He stresses the need for explainability, isolation, and governance in government AI tools, citing OMB memoranda M2521 and M2522 mandating faster AI adoption with real-world impact. The lab partners with frontier labs like OpenAI for chem-bio safety work, and with NVIDIA and HPE for the Venado supercomputer. He urges startups to build for DoD IL5 environments and continuous compliance to succeed in federal procurement, noting Los Alamos holds petabytes of never-internet data and specializes in sensors like the ChemCam on Mars.

Future-Proof Coding Agents – Bill Chen & Brian Fioca, OpenAI
Dec 5, 2025 · 17:48
Bill Chen and Brian Fioca from OpenAI argue that coding agents are booming but fragile, with teams rebuilding harnesses every time models change. They present Codex, OpenAI's combined model and harness, as a stable abstraction layer that handles complexity like context compaction, parallel tool calls, and MCP integration. Key lessons include avoiding overprompting—letting the model use its trained habits rather than forcing foreign instructions—as demonstrated when GPT-5 ran slowly because users told it to examine every file. They describe patterns for using Codex as an SDK, enabling products like GitHub and Cursor to integrate without maintaining their own harness, and highlight future improvements in longer-horizon tasks and self-healing software. The episode concludes that building where models are going, with Codex's SDK, lets teams focus on differentiation rather than infrastructure rewrites.

Katelyn Lesse – Evolving Claude APIs for Agents, Anthropic
Dec 4, 2025 · 13:25
Katelyn Lesse, who leads the Claude Developer Platform team at Anthropic, explains how the platform is evolving to help developers build powerful agentic systems using Claude. She details three key areas: harnessing Claude's capabilities through features like extended thinking and tool use, managing context with MCP, a memory tool, and context editing (which combined to yield a 39% performance bump), and giving Claude a computer via the code execution tool and agent skills. The talk also covers challenges like container orchestration for Claude Code on web and mobile, and the importance of letting Claude work autonomously in a sandbox environment.

No Vibes Allowed: Solving Hard Problems in Complex Codebases – Dex Horthy, HumanLayer
Dec 2, 2025 · 20:31
Dex Horthy of HumanLayer argues that with deliberate context engineering — specifically a technique he calls "frequent intentional compaction" — today's AI coding agents can handle large brownfield codebases, not just greenfield projects. He presents a three-phase workflow (research, plan, implement) that keeps agents in the "smart zone" of the context window, avoiding the diminishing returns that set in around the 40% mark. Horthy demonstrates this approach solving a real problem in a 300k-line Rust codebase, shipping a week's worth of work in seven hours with code that passed expert review. He cautions against outsourcing thinking to AI, emphasising that plans must include actual code snippets for reliable execution and team alignment. The talk also addresses the cultural rift where senior engineers clean up slop from juniors' AI tools, calling for top-down adoption and team-wide workflow adaptation.

Defying Gravity - Kevin Hou, Google DeepMind
Dec 2, 2025 · 25:10
Kevin Hou from Google DeepMind introduces Antigravity, a new agent-first IDE that combines three surfaces: an AI editor, an agent-controlled Chrome browser, and an Agent Manager. The central claim is that agents should live outside the IDE, enabling longer-running tasks and multimodal interactions. Key features include Artifacts—dynamic representations like plans, screen recordings, and diagrams—that replace raw chain-of-thought with visual summaries. The browser enables context retrieval and verification via screen recordings, while image generation allows iterative design through comments. Hou explains the research-product flywheel: DeepMind researchers use Antigravity internally to identify model gaps, improving capabilities like computer use and instruction following, which then ship in the product. He also details four categories of model improvements: intelligence, tools, long-running tasks, and multimodal, all powered by Gemini 3 Pro and related models.

Building Cursor Composer – Lee Robinson, Cursor
Dec 2, 2025 · 15:36
Lee Robinson explains how Cursor built Composer, its first agent model for real-world software engineering, by focusing on being both fast and smart — achieving 4x more efficient token generation than similarly intelligent models while matching open-source performance initially and approaching frontier models after reinforcement learning. The model's training posed infrastructure challenges: matching training and inference environments across thousands of GPUs, handling complex rollouts with hundreds of tool calls and up to millions of tokens, and ensuring consistency by using the same tool format and responses as production. Cursor solved these with custom kernels that sped up training by 3.5x on NVIDIA Blackwell chips for mixture-of-experts layers, load balancing across threads to avoid idle time, and co-designing RL infrastructure with its Cloud Agents product using virtual machines that mirror the production Cursor environment. This allowed the model to become a power user of tools like semantic search, which improved all models but especially Composer. RL also taught the model to parallelize tool calls (e.g., reading 10 files simultaneously) and to search more before editing,…

The Unbearable Lightness of Agent Optimization — Alberto Romero, Jointly
Nov 24, 2025 · 17:58
Alberto Romero, co-founder and CTO of Jointly, introduces Meta-ACE, a meta-optimization framework that orchestrates multiple adaptation strategies to overcome the limitations of single-dimensional context engineering like ACE. The framework uses a meta-controller to profile task complexity, uncertainty, verifiability, and resource constraints, then allocates strategies across context, compute, verification, memory, and parameter dimensions. Initial results show 8-11% improvement on agent benchmarks, 30-40% reduction in compute costs, and 6-8% gains on domain-specific tasks. Meta-ACE addresses ACE's weak reflector problem with quality gates and multi-signal reflection, feedback brittleness via a hierarchical verification cascade (self-verification, multimodal consensus, execution checks), and task complexity mismatch by dynamically adjusting strategy allocation to save up to 90% compute on simple tasks. Future work includes scaling to multimodal and compound AI systems, with challenges in meta-controller training stability, computational overhead, and verification cascade brittleness.

Compilers in the Age of LLMs — Yusuf Olokoba, Muna
Nov 24, 2025 · 17:36
Yusuf Olokoba, founder of Muna, explains how they built a Python compiler that converts plain Python inference functions—like Google's 270M-parameter Gemma embedding model—into self-contained C++ and Rust binaries using LLMs within a verifiable pipeline. The process involves symbolic tracing to generate an intermediate representation, type propagation to infer types for native compilation, and LLM-driven code generation to mass-produce native implementations of Python operations. After compiling into a shared library, the model can be loaded via FFI from any language (e.g., Node.js) and exposed through an OpenAI-compatible client, enabling developers to run open-source models anywhere—locally, cloud, mobile—with minimal code changes. Olokoba details why they abandoned PyTorch FX due to its PyTorch-only focus and reliance on fake inputs, and how LLMs help scale the coverage of elementary operations. The talk argues that this compiler approach solves hybrid inference—small edge models working with large cloud models—by moving beyond Python and Docker to more portable, low-latency native binaries.

Backlog.md: Terminal Kanban Board for Managing Tasks with AI Agents — Alex Gavrilescu, Funstage
Nov 24, 2025 · 14:19
Alex Gavrilescu presents Backlog.md, an open-source CLI tool that stores tasks as Markdown files in Git repos, featuring a terminal Kanban board and MCP server for AI agents. He argues that breaking features into atomic Markdown tasks prevents agents from running out of context or implementing unwanted extras. The demo shows Claude creating a task from requirements, generating an implementation plan, and coding a move-mode feature—all via MCP tools. Gavrilescu emphasizes two review checkpoints (after task creation and after the plan) and notes that Backlog.md itself is 99% AI-written. The tool works cross-platform, requires no external APIs, and syncs across branches via Git.

Agents are Robots Too: What Self-Driving Taught Me About Building Agents — Jesse Hu, Abundant
Nov 24, 2025 · 17:37
Jesse Hu, former Waymo engineer and now founder of Abundant, argues that building reliable AI agents mirrors the challenges of self-driving cars, with the same 1% vs 99% problem where the model accounts for only 1% of the work. He draws parallels between robotics and agents across closed-loop feedback, statefulness, action spaces, and simulation, noting that just as self-driving pioneers learned perception was easy but planning was hard, agent builders must move from predictive models to action models. Hu explains how open-loop designs (e.g., waiting for full tool responses) limit agent reactivity, and advocates for richer input/output mechanisms like character-level terminal streams. He emphasizes the importance of an offline stack—simulation, evaluation, and data feedback loops—for iterative hill climbing, and warns that actions have consequences: out-of-distribution failures cascade. Finally, he recommends reading up on MDPs, DAgger, and offline RL from robotics literature to accelerate agent development.

Vision: Zero Bugs — Johann Schleier-Smith, Temporal
Nov 24, 2025 · 36:03
Johann Schleier-Smith, Technical Lead for AI at Temporal, argues that zero-bug software is achievable by applying aerospace-grade techniques like rigorous specification, modular design, and formal verification, now made practical by AI. He cites the Airbus A320's control software (no incidents from software in decades) and the Space Shuttle (1 error per 420,000 lines) as proof. Schleier-Smith demonstrates Dafne, a language that embeds proofs inline with code for automated verification, and shows how agentic coding can reduce high-assurance code costs by ~100× compared to traditional development. He concludes that agents generating verified code will overcome current quality limitations and trigger widespread adoption.

Context Engineering: Connecting the Dots with Graphs — Stephen Chin, Neo4j
Nov 24, 2025 · 26:50
Stephen Chin, VP of Developer Relations at Neo4j, argues context engineering powered by knowledge graphs transforms AI from prompt engineering to information architecture, enabling agents with structured memory and retrieval. He demonstrates GraphRAG using Neo4j's knowledge graph builder with supply chain and VEX documents, showing how vector and graph algorithms retrieve specific Jackson library vulnerabilities with CVE details, severity, and remediation versions. Chin contrasts fast two-pass vector+graph retrieval with agentic multi-step traversal via the Neo4j MCP server and Claude Code, which yields deeper results like CVE number, attack type, and upgrade paths. He explains that graphs excel when relationships span two or more facts, using a presentation update example involving himself, colleague Sid, and the GIDS event. Practical applications include role-based access overlays and explainable AI via visualized conversation flows. Resources like free Graph Academy courses, Nodes AI 2026 conference, and graphrag.com support implementation.

Enterprise Deep Research: The Next Killer App for Enterprise AI — Ofer Mendelevitch, Vectara
Nov 24, 2025 · 5:19
Ofer Mendelevitch from Vectara introduces Enterprise Deep Research as the next killer app for enterprise AI, applying autonomous, multi-step reasoning to internal knowledge bases. Vectara's trustworthy agent OS enables this with multimodal ingest, hybrid retrieval, and hallucination mitigation (HHEM model at 5.5M downloads). Deep Research queries private data to generate comprehensive reports with citations, replacing manual workflows. Key use cases include responding to RFPs by scanning enterprise datasets, generating on-demand employee onboarding guides from Jira/Notion/SharePoint, and creating investment memos in financial services. The system uses multi-agent parallel execution and corpus understanding for accurate planning, addressing 73% of LLM users' top challenge: factual accuracy.

Infra that fixes itself, thanks to coding agents — Mahmoud Abdelwahab, Railway
Nov 24, 2025 · 18:08
Mahmoud Abdelwahab from Railway presents Railway Autofix, a system that uses Inngest for durable execution and OpenCode as a coding agent to automatically detect infrastructure issues and open pull requests to fix them. The system runs scheduled workflows to fetch project architecture, resource metrics, and HTTP performance data, then analyzes services exceeding thresholds. For affected services, it gathers additional context like logs and generates a detailed fix plan using an LLM. The coding agent then implements the plan and opens a PR with a summary of changes. The talk demonstrates how this shifts developers from manually investigating alerts to simply reviewing auto-generated fixes.

Context Platform Engineering to Reduce Token Anxiety — Val Bercovici, WEKA
Nov 24, 2025 · 23:52
WEKA's Val Bercovici and Callan Fox present Context Platform Engineering, open-sourcing a toolkit that maximizes KV cache hit rates—called the single most important metric for production-grade AI agents by Manus AI. They argue that without this engineering, users resort to Context Financial Engineering, a clairvoyant prompt cache arbitrage against time-to-live (5 minutes to 1 hour) and cache read/write pricing. The toolkit includes a load generator that configures agent swarms with specific SLOs, cycling through deterministic and random prompts. Callan's WEKA Labs research shows that 1-minute TTL causes 15-16x token repetition, while 1-hour TTL approaches 1x, but requires larger cache capacity. Benchmarks compare HBM+WEKA (purple) vs HBM+DRAM (orange) vs HBM+DRAM+slow storage (pink); WEKA's NVMe-backed augmented memory grid maintains higher output token rates under increasing concurrent users—up to the point where DRAM tiers drop off sharply due to insufficient speed. The episode covers how SLA requirements translate into SLOs via memory tiering and KV cache offloading, and emphasizes that subscription users effectively purchase cache allotments to keep inference providers in…

From Stateless Nightmares to Durable Agents — Samuel Colvin, Pydantic
Nov 24, 2025 · 22:13
Samuel Colvin of Pydantic demonstrates building production-grade durable AI agents using PydanticAI, Temporal, and Pydantic Logfire, arguing that stateless architectures fail at scale and that durable execution with checkpointing and recovery is essential. He shows a 20 questions game where two agents play, but with 20% simulated failures, restarts are avoided by wrapping agents in Temporal wrappers for automatic retries. Colvin then implements a Deep Research agent that plans searches, runs parallel web searches via Tavilli, and synthesizes results, all recoverable. In a live demo, killing the workflow and restarting resumes instantly, replaying cached LLM calls in milliseconds. He also previews Pydantic AI Gateway and demonstrates Pydantic Evals comparing Gemini, GPT-4.1, and Claude Sonnet 4.5 on performance and cost.

Hacking Subagents Into Codex CLI — Brian John, Betterup
Nov 24, 2025 · 13:39
Brian John, Principal Full Stack Engineer at Betterup, explains how to hack subagents into OpenAI's Codex CLI by using a 72-line wrapper script that launches child Codex processes as subagents, enabling context management without polluting the main agent's window. He details the minimum sandbox permissions required—workspaceright on both parent and child, plus disabling the rollout recorder—and notes that security risk is low per Meta's Rule of 2, though not zero. The wrapper uses file-based communication to avoid repetitive permission prompts, and subagents are defined in an agents.md file with configurable reasoning effort. John demonstrates a word counter and file writer agent, noting that Codex runs subagents serially and slower than Claude Code (e.g., 40 seconds for a simple task, up to 20 minutes for large codebases), but considers this acceptable for Codex's unattended design.
Powered by PodHood