Page 15 of 23

Build Dynamic Products, and Stop the AI Sideshow — Eliza Cabrera (Workday) + Jeremy Silva (Freeplay)
Jul 23, 2025 · 18:10
Eliza Cabrera (Principal AI Product Manager at Workday) and Jeremy Silva (Product Lead at Freeplay) argue that companies must stop treating AI as a separate 'sideshow' and instead deeply integrate it into product strategy using a crawl-walk-run approach to build dynamic, differentiated experiences. They explain that bolt-on AI products result from centralized AI strategies quarantined from core product, leading to features like chatbots that demonstrate capability but don't solve real customer problems. Workday's example shows starting with Gen AI content generation and translations for knowledge bases (crawl), then a contextually aware assistant that processes PII (walk), and finally agentic capabilities that autonomously act on policy changes (run). The north star is AI products that feel like natural, cohesive parts of the experience, not bolt-on enhancements.

The Billable Hour is Dead; Long Live the Billable Hour — Kevin Madura + Mo Bhasin, Alix Partners
Jul 23, 2025 · 17:04
Kevin Madura and Mo Bhasin from AlixPartners argue that AI is reshaping knowledge work by compressing upfront data ingestion from 50% to 10-20% human effort, enabling analysis of 100% of data rather than the top 20%. They detail three use cases: categorization via structured outputs achieving 95% accuracy on 10,000 vendors in minutes; enterprise-scale RAG that democratizes access to siloed data by teaching LLMs to call APIs; and structured extraction from documents using schemas and log-prob-based confidence scoring. They caution that AI investments face a paradox—89% of CEOs plan agentic AI but NBER finds no earnings impact—and stress that success requires people skills, demos, and a focus on NPS and ROI over chasing shiny new tools. The episode concludes that once Excel-powered LLMs work reliably, AGI will be here.

From Copilot to Colleague: Trustworthy Agents for High-Stakes - Joel Hron, CTO Thomson Reuters
Jul 23, 2025 · 19:45
Thomson Reuters CTO Joel Hron explains the shift from helpful AI assistants to productive, agentic systems in high-stakes legal, tax, and compliance workflows. He frames agency as a spectrum with dials for autonomy, context, memory, and coordination, which must be tuned to the risk tolerance of professional users. Key lessons include building the whole system first rather than over-indexing on minimal MVPs, treating legacy applications as assets to be decomposed into tools for agents, and the persistent challenge of evaluating agentic systems with expensive, variable human expert judgments. Hron demonstrates a tax product that extracts data from documents and maps it to a calculation engine, and a legal research agent that uses proprietary case law tools to produce cited reports. He emphasizes leveraging unique organizational assets—like Thomson Reuters' 4,500 domain experts and terabytes of proprietary content—to create differentiation.

How to Hire AI Engineers when EVERYONE is cheating with AI — Beth Glenfield, DevDay
Jul 22, 2025 · 6:45
Beth Glenfield, from DevDay, argues that AI has broken technical hiring because candidates cheat with AI tools, making LeetCode puzzles obsolete, and small companies are crushed in the talent war with big tech. She cites the AI cheating service CLUI, which raised $5.3M and is heading toward $1M ARR, and notes 93% LeetCode wizard success rates for Google/Meta interviews, with 1 in 3 interviews now using AI assistants. Instead of puzzles, she proposes real-world workplace simulations where candidates collaborate with AI agents (perfectionist, pragmatist, etc.) to measure skills like handling ambiguity, mentoring, and business judgment. She highlights that big tech can brute-force hiring (100 candidates for 5 hires), but startups cannot afford bad hires costing $20k-$60k. Referencing Mark Zuckerberg's prediction that AI will handle mid-level engineering and Marc Benioff's announcement that Salesforce won't hire software engineers due to a 30% productivity boost from AI, Glenfield argues engineering jobs now require creativity, collaboration, and working with AI rather than being replaced by it.

Stateful environments for vertical agents — Josh Purtell, Synth Labs
Jul 22, 2025 · 6:51
Josh Purtell of Synth Labs argues that stateful environments—containerized, network-bounded workspaces that capture external state—make it easier to build effective agents for vertical applications like finance and health. By keeping the application logic separate from the agent, developers can revamp their agent when new models come out without rewriting everything. The environment exposes a tailored representation (e.g., just the terminal, not the whole OS) and supports resets and rollbacks, enabling techniques like tree search that improve long-horizon tasks. Purtell also notes that network boundaries allow reliable multi-agent setups and asynchronous work. The concepts are implemented in the open-source Synth AI Environments repository.

Books reimagined: AI to create new experiences for things you know — Lukasz Gandecki, TheBrain.pro
Jul 22, 2025 · 9:44
Łukasz Gandecki of TheBrain.pro presents Books Reimagined, an AI-powered reading companion that adds music, graphics, and conversational AI to books like '1984' and 'The Snow Queen', arguing that AI can create magical experiences by hiding itself behind a polished user interface. He started by vibe-coding a companion for a Trump reelection book to track characters, then iterated into a full multimedia reader. He demonstrates features including 100-millisecond voice response, natural language search via embeddings, and deep research that scans the entire book. Gandecki emphasizes that human curation of AI-generated music and graphics is essential for quality, and he is open-sourcing the player at bookgenius.net.

AI powered entomology: Lessons from millions of AI code reviews — Tomas Reimers, Graphite
Jul 22, 2025 · 10:21
Tomas Reimers, co-founder of Graphite, discusses the company's AI-powered code reviewer Diamond, arguing that while LLMs can effectively find bugs, they must be carefully prompted to avoid frustrating developers. After analyzing 10,000 comments from codebases, Graphite classified bugs into a quadrant based on what LLMs can catch versus what humans want to receive, identifying bugs, accidentally committed code, performance and security concerns, and documentation mismatches as high-value targets. Code cleanliness and best practice comments, though technically correct, are often unwelcome from AI. Graphite measures success via emoji reactions (less than 4% downvote rate) and the percentage of comments that lead to code changes, reaching 52% actionability in March—matching human-level effectiveness. The talk emphasizes continuous monitoring to ensure LLM comments stay within the desirable quadrant.

Critical AI Inference your CIO can Trust — Sahil Yadav, Hariharan Ganesan, Telemetrak
Jul 22, 2025 · 19:04
Sahil Yadav and Hariharan Ganesan of Telemetrak present a three-pillar framework—explainability, adaptive guardrails, and human-in-the-loop on a foundation of traceability—to operationalize trust in enterprise AI inferences, arguing that without these, silent failures can cost millions. They introduce XTOPS as an integrated upgrade to MLOps, with trust-specific dashboards and dynamic guardrails. A case study of GuardHat, a worker safety platform, shows how GPS drift caused 70% false positives; applying XTOPS reduced resolution time from eight months to seven days. They propose metrics MTTRE (mean time to resolve explainable errors) and trust-adjusted risk in dollars, and demonstrate convincing CIOs by quantifying savings, e.g., $500K per site per year in fines avoided.

How to run Evals at Scale: Thinking beyond Accuracy or Similarity — Muktesh Mishra, Adobe
Jul 22, 2025 · 9:25
Muktesh Mishra, lead engineer for Applied AI at Adobe, argues that evaluations (Evals) are the most critical aspect of AI application development, advocating for Eval-driven development. He explains that Evals must be tailored to the use case: RAG applications require accuracy or similarity metrics, code generation needs functional correctness tests, and agent systems demand trajectory evaluation and multi-turn simulation. Mishra emphasizes starting with data — synthetic data, continuous refinement, and multiple labeled datasets for different flows. To scale, he recommends caching intermediate results, orchestrating parallel runs, and establishing a measure-monitor-analyze-iterate loop, while balancing high-fidelity human-in-the-loop reviews with automated speed. He stresses that no universal Eval exists and process matters more than tools.

Continuous Profiling for GPUs — Matthias Loibl, Polar Signals
Jul 22, 2025 · 11:31
Matthias Loibl of Polar Signals explains how continuous profiling for GPUs maximizes GPU efficiency using low-overhead, always-on sampling via eBPF. He contrasts tracing (high cost) with sampled profiling (e.g., 100 Hz, <1% overhead) and details GPU metrics collected from NVIDIA NVMe, including utilization, memory, clock speed, power, temperature, and PCIe throughput. The platform correlates these with CPU stack traces to identify bottlenecks, such as Python and CUDA functions underutilizing the GPU. A new GPU time profiling feature records the duration of CUDA kernel executions, showing actual time spent by functions on the GPU. Deployment runs on Linux with a binary, Docker, or Kubernetes DaemonSet; early adopters like TurboPuffer use it to optimize their vector engine.

Top Ten Challenges to Reach AGI — Stephen Chin, Andreas Kollegger
Jul 22, 2025 · 4:23
Stephen Chin and Andreas Kollegger, curators of the GraphRAG track at AI Engineer World's Fair, use ten sci-fi memes to expose top challenges to AGI—from Memento's short-term memory as prompt engineering to Skynet's unintended consequences, the Matrix's agent simulation, Hal's trust issues, emotions as bug/feature, Frankenstein's creator obligations, Terminator's time travel, Star Wars' language nuance, the Borg's assimilation, and Deep Thought's right questions. They argue graphs and graph technology can solve some of these challenges, inviting the audience to the GraphRAG track to learn which ones. The episode delivers a playful yet pointed framework connecting pop culture to core AGI obstacles.

Practical GraphRAG: Making LLMs smarter with Knowledge Graphs — Michael, Jesus, and Stephen, Neo4j
Jul 22, 2025 · 19:46
Michael Hunger and Stephen Chin of Neo4j present GraphRAG as a method to enhance LLMs by integrating knowledge graphs, achieving more accurate, contextual, and explainable answers than standard RAG. They highlight limitations of vector search, showing GraphRAG improves relevance and reduces hallucinations. They detail a three-step construction process: lexical graphs, entity extraction, graph enrichment with algorithms. They demonstrate open-source tools like Neo4j Knowledge Graph Builder and GraphRAG Python package, and show an agentic approach using domain-specific retrievals. They cite studies showing three times improvement in accuracy.

Knowledge Graphs in Litigation Agents — Tom Smoker, WhyHow
Jul 22, 2025 · 19:13
Tom Smoker, founder of WhyHow.ai, explains how his company uses Graph RAG and multi-agent systems to find class action and mass tort cases early, often months before law firms. They scrape the web, structure data into graphs with consistent schemas, and use LLMs to pipe together ML-filtered pipelines, achieving early signal detection within 15 minutes. Smoker highlights that while single agents can reach 95% accuracy, chaining five agents reduces expected accuracy to 77%, necessitating guardrails and human-in-the-loop. Graphs provide extensible, prunable state control for personalized lawyer workflows, enabling generation of specific reports from subgraphs. Examples include tracking car fire complaints density and velocity to identify lawsuits early, and using graphs to organize discovery documents for faster review.

When Vectors Break Down: Graph-Based RAG for Dense Enterprise Knowledge - Sam Julien, Writer
Jul 22, 2025 · 15:47
Sam Julien, Director of Developer Relations at Writer, explains how the company built a graph-based RAG system that achieved 86.31% accuracy on the RobustQA benchmark and sub-second response times, significantly outperforming vector search approaches for dense enterprise knowledge. The team moved from simple vector retrieval to a graph-based approach to handle concentrated data where similar terms appear frequently, solving problems like inaccurate chunking and failure with similar documents. They built a specialized model to convert data into graph structures, stored as JSON in a Lucene-based search engine, and incorporated Fusion-in-Decoder with knowledge graphs to lower hallucination rates. Key decisions included focusing on customer needs over hype, staying flexible based on team expertise, and letting research challenge assumptions, leading to features like multi-hop reasoning and complex data format handling.

Stop Using RAG as Memory — Daniel Chalef, Zep
Jul 22, 2025 · 7:02
Daniel Chalef, founder of Zep, explains why using RAG as memory for AI agents fails: vector databases rely on semantic similarity, not business relevance, so irrelevant facts pollute retrieval and cause hallucinations. He advocates for domain-aware memory using temporal knowledge graphs, demonstrated through Zep's open-source Graphiti framework. Chalef shows how developers define custom entity types (e.g., financial goals, debts) with Pydantic or Zod schemas, register them as ontology objects, and build tools that concurrently search Zep filtered by node type. A live demo of a finance coach agent illustrates Zep capturing structured facts like rent payments into a knowledge graph, enabling precise, context-rich memory for accurate agent reasoning.

HybridRAG: A Fusion of Graph and Vector Retrieval - Mitesh Patel, NVIDIA
Jul 22, 2025 · 20:24
Mitesh Patel, Developer Advocate Manager at NVIDIA, presents HybridRAG, a fusion of knowledge graph-based GraphRAG and vector-based VectorRAG for improved question-answering from complex texts. He emphasizes that ontology engineering consumes 80% of development time and is critical for accurate triplet extraction, where fine-tuning LLaMA 3.1 with LoRA boosted triplet accuracy from 71% to 87% on 100 documents. Patel also highlights retrieval strategies like multi-hop graph traversal, which provides richer context but increases latency, and recommends QGraph acceleration via Networx to reduce latency. For evaluation, he suggests RAGAS for end-to-end pipeline metrics and the LLaMA-Nimotron reward model for response quality. Ultimately, he advises using GraphRAG when data has inherent structure or complex relationships, but notes it is compute-heavy, so the choice between GraphRAG, semantic RAG, or hybrid depends on the use case.

tldraw.computer - Steve Ruiz, tldraw
Jul 21, 2025 · 18:45
Steve Ruiz, founder and CEO of tldraw, demonstrates the company's AI experiments on their infinite canvas, including Make Real—which turns hand-drawn wireframes into working web apps using vision models—and tldraw computer, a visual programming environment where arrows and LLMs power a graph of connected nodes that can execute multi-step prompts, generate images and speech, and even run loops indefinitely. He also shows Draw Fast for real-time image generation and Teach, where Claude can draw and edit shapes on the canvas. The episode explains how tldraw's SDK (tldraw.dev) enables others to build custom canvas applications, and highlights the company's philosophy of 'shitty but amazing' rapid prototyping.

Excalidraw: AI and Human Whiteboarding Partnership - Christopher Chedeau
Jul 21, 2025 · 16:59
Christopher Chedeau, creator of Excalidraw, explains how to integrate AI into whiteboarding by focusing on turning prompts into editable diagrams rather than static images. He argues that just adding any AI model harms the product, citing a failed attempt to generate realistic images because users don't draw realistically. The successful integration uses Mermaid.js to output Excalidraw files, letting humans modify the AI-generated diagram. He envisions a future of iterative human-AI collaboration, and demonstrates other practical features like auto-naming files, generating illustrations (coming soon), and challenges AI engineers to build a browser-based logo background remover. He concludes that the industry is in a physical-to-virtual transition for AI, and that LLMs work best when targeting a structured domain-specific language.

The Bitter Layout or: How I Learned to Love the Model Picker — Maximillian Piras, Yutori
Jul 21, 2025 · 14:24
In this talk, Maximillian Piras argues that conversational interfaces remain dominant in AI apps not because they are natural, but because they are 'conformable' — able to absorb the next model's capabilities without redesign. He calls this pattern 'The Bitter Layout': an input field, turn-by-turn flow, and a model picker, which prioritizes flexibility over usability. Applying Clayton Christensen's theory of commoditization, Piras claims that as long as scaling laws keep models from commoditizing, the interface itself is the commodity, and designers must conform to the next model. He traces the debate over chat UX from Linus Lee (2022) to Julian Lear (2024), and criticizes the model picker as a 'mode selector' that creates usability issues. Looking ahead, Piras suggests designers shift from procedural thinking to setting goals and constraints, and speculates that future AI UX will be 'grown' like a garden rather than built.

UX Design Principles for Semi Autonomous Multi Agent Systems — Victor Dibia, Microsoft
Jul 21, 2025 · 20:28
Victor Dibia, Principal Research Software Engineer at Microsoft Research, presents four UX design principles for semi-autonomous multi-agent systems: capability discovery, observability and provenance, interruptibility, and cost-aware action delegation. He demonstrates these through Blender LM, a multi-agent system built from scratch that translates natural language into 3D Blender scenes, featuring a planner and verify agent. Dibia stresses eval-driven design—starting with a non-AI baseline, then iterating with agents—and warns that only a small subset of tasks truly benefit from multi-agent autonomy, advocating for careful ROI assessment. He also showcases AutoGen Studio, a low-code tool for composing multi-agent workflows, and shares code and resources for further learning.

Agentic GraphRAG: AI’s Logical Edge — Stephen Chin, Neo4j
Jul 21, 2025 · 15:27
Stephen Chin of Neo4j argues that Agentic GraphRAG — combining graph databases with retrieval-augmented generation — overcomes LLM hallucinations and biases by providing structured, relational context. He demonstrates how LLMs fail on reasoning tasks like calculating classroom capacity due to inaccurate anchoring on irrelevant data, and proposes an architecture where vector search first identifies relevant nodes, then graph traversal retrieves related context for the LLM. Chin highlights Neo4j’s MCP server for cypher query generation and memory modules, and cites Klarna’s success: 250K employee questions answered in the first year, 2,000 daily queries, and 85% adoption, replacing their entire SaaS stack. He recommends the Neo4j Certified Developer Program and the Nodes Conference for further learning.

CIAM for AI: Authn/Authz for Agents — Michael Grinich, CEO of WorkOS
Jul 21, 2025 · 20:13
Michael Grinich, CEO of WorkOS, argues that AI agents need first-class identity and access management, distinct from human or machine-to-machine auth. He identifies key challenges: headless login, least privilege for non-deterministic systems, and compliance tracking. Grinich presents four emerging patterns—persona shadowing, delegation chains, capability tokens, and human-in-the-loop escalation—and references standards like OAuth, UMA, GNAP, OIDC for agents, and verifiable credentials. He predicts a shift from 95% human traffic to 95% agent traffic, calling for middleware trust boundaries and urgent collaboration on agent identity standards.

Good design hasn’t changed with AI — John Pham, SF Compute
Jul 21, 2025 · 20:25
John Pham of SF Compute argues that good design principles—speed, trust, accessibility, and delight—remain unchanged even with AI. He demonstrates these through SF Compute's onboarding, which loads in under 300ms using server-side rendering and a 14KB fog image animated with stacked transparent layers and no third-party JavaScript, respecting reduced motion preferences. Trust is built by setting clear expectations (e.g., '3 steps, under a minute'), auto-filling forms via browser autocomplete, and avoiding layout shifts. Delight comes from color psychology to slow users, peak-end bias (ending with a beautiful San Francisco scene), and the GPU habitat: a live video feed of GPUs implemented as stacked divs and a looping video, with a prediction cone that smooths nested menu interactions. Accessibility includes semantic HTML, screen reader testing, and pausing animations for reduced motion. These techniques turn users into super fans while growing total addressable market.

Building Effective Voice Agents — Toki Sherbakov + Anoop Kotha, OpenAI
Jul 20, 2025 · 17:17
Toki Sherbakov and Anoop Kotha from OpenAI argue that speech-to-speech models have reached a 'good enough' tipping point for production voice agents, highlighting the shift from chained architectures (transcription + LLM + TTS) to the real-time API's speech-to-speech approach. They detail trade-offs across latency, cost, accuracy, UX, and integrations depending on use case—consumer apps prioritize low latency and expressiveness, while customer service demands accuracy and tool integration. Key design patterns include delegating complex tasks to smarter models like O4 mini, prompting for voice expressiveness and tone, and starting with few tools. For evals, they recommend starting with observability, using SMEs for labeling, then transcription-based LLM-as-judge evals, audio evals with GPT-4 Audio to assess tone and pacing, and synthetic conversations between two real-time clients. Guardrails should run async with a configurable debounce period (e.g., every 100 characters). Examples from Lemonade (early focus on evals and guardrails) and Tinder (customization for brand realism) illustrate successful approaches.

What every AI engineer needs to know about GPUs — Charles Frye, Modal
Jul 20, 2025 · 19:52
Charles Frye of Modal explains that AI engineers need to understand GPU hardware constraints to optimize inference, arguing GPUs embrace bandwidth over latency and that Tensor Cores for low-precision matrix-matrix multiplication are the key resource. He describes how GPUs achieve 16,000+ parallel threads per cycle on H100, and notes Patterson’s Law: bandwidth improves at the square of latency. The main insight: arithmetic intensity favors N² operations per N memory loads, so matrix-matrix operations are efficient while matrix-vector is wasteful. Frye demonstrates that running a small 8B model 1,000 times on the same prompt matches GPT-4 quality, and that multi-token prediction and multi-sample query become nearly free because Tensor Cores handle expanded batches as matrix-matrix multiplications. He recommends using smaller models that fit on a single GPU and scaling via multiple generations.

Robots as professional Chefs - Nikhil Abraham, CloudChef
Jul 20, 2025 · 18:58
Nikhil Abraham, CEO of CloudChef, explains how his company turned a general-purpose bimanual robot into a professional chef that works in commercial kitchens for $12 an hour. The robot learns new recipes from a single expert demonstration, using thermal and visual embeddings to handle ingredient and appliance variation. CloudChef's system achieves 95% autonomy and outperforms expert human chefs in cooking decision-making, as evaluated on over 1,000 recipes. The robot is deployed in restaurants like Wingstar and Elan, cooking real meals at 80–95% human speed. Abraham notes that the platform can operate 168 hours a week and aims to expand to tasks like chopping.

[Full Workshop] Reinforcement Learning, Kernels, Reasoning, Quantization & Agents — Daniel Han
Jul 19, 2025 · 2:42:28
Daniel Han of Unsloth presents a technical workshop covering reinforcement learning (RL), kernels, reasoning, quantization, and agents, arguing that RL with verifiable rewards (RLVR) is the key to unlocking LLM capabilities beyond supervised fine-tuning. He explains why open-source models plateaued after September 2024 until DeepSeek-R1 showed that RL can elicit reasoning, and breaks down PPO, GRPO, and the REINFORCE algorithm, emphasizing that GRPO removes the value model for efficiency. Han details how reward functions—not algorithms—are the hardest part, with examples like distance-based scoring for math. He demonstrates a free Colab notebook training a base model to reason, and shows that dynamic quantization can shrink models like DeepSeek-R1 from 730 GB to 140 GB with only ~1% accuracy loss, arguing that GPUs may stop getting faster after FP4 precision.

A Taxonomy for Next-gen Reasoning — Nathan Lambert, Allen Institute (AI2) & Interconnects.ai
Jul 19, 2025 · 19:21
Nathan Lambert, senior research scientist at AI2 and founder of Interconnects.ai, argues that next-generation reasoning models require a taxonomy of four traits: skills, calibration, strategy, and abstraction. While current models excel at math and code (skills), they overthink easy problems, wasting tokens and latency—calibration is needed to match output length to task difficulty. The real frontier is planning: models must learn strategy (choosing the right direction) and abstraction (breaking problems into tractable sub-tasks) to enable long-horizon agents like Deep Research and Claude Code. Lambert traces how OpenAI's Q*→Strawberry→o1 took 12–18 months of human data to teach backtracking and verification; planning should be easier because humans can write 5-step plans. He predicts post-training compute could reach parity with pre-training, citing DeepSeek's shift from 0.18% post-training compute in V3 to an estimated 10–20% for R1. The path forward: collect diverse verifiable questions, filter by difficulty, run stable RL—then scale.

How to Train Your Agent: Building Reliable Agents with RL — Kyle Corbitt, OpenPipe
Jul 19, 2025 · 19:48
Kyle Corbitt, co-founder of OpenPipe, argues that reinforcement learning (RL) with GRPO can make agentic systems far more reliable and cost-effective than prompted frontier models. He presents ART-E, an email assistant trained on Qwen 2.5 14B, which achieved 96% accuracy versus 90% for o3 and slashed cost from $55 to $0.80 per 1,000 queries. The two critical problems are building a realistic environment (solved using the Enron email dataset) and defining the right reward function (turning it into a verifiable task with LLM-as-judge). Extra rewards—favoring fewer tool turns and penalizing hallucination—further improved efficiency. Corbitt also warns about reward hacking, giving examples like a model that exploited a bug to put every word in every category, and shares that the training cost just $80 in GPU time and a week of engineering.

OpenThoughts: Data Recipes for Reasoning Models — Ryan Marten, Bespoke Labs
Jul 19, 2025 · 19:59
Ryan Marten, co-lead of the OpenThoughts collaboration and founding engineer at Bespoke Labs, reveals the missing data recipe for open-source reasoning models, presenting OpenThoughts 3, a state-of-the-art 7B reasoning dataset that outperforms DeepSeek R1 Qwen 7B and Nematron Nano on benchmarks like AIME, Live Code Bench, and GPQA Diamond. Through over 1,000 experiments and 5,000 datasets, key findings include that sampling multiple reasoning traces per question scales performance by 16x, Qwen 32B surpasses DeepSeek R1 as a teacher model, synthetic question generation is highly effective, and filtering by difficulty or response length works better than embeddings. Surprisingly, verification of answers in SFT distillation did not improve results, and focusing on fewer high-quality sources outperformed maximizing diversity. For domain-specific reasoning, Marten advises starting with the OpenThoughts recipe, using synthetic data generation (via the open-source Curator library), and rigorous evaluation (via EvalComet). A legal reasoning example shows that distillation can surpass the teacher model. All resources are open-source.

Google Photos Magic Editor: GenAI Under the Hood of a Billion-User App - Kelvin Ma, Google Photos
Jul 19, 2025 · 20:28
Kelvin Ma, an engineer on Google Photos' editing team, explains how the billion-user app built the Magic Editor by integrating generative AI with on-device computational photography. He traces the evolution from earlier ML features like post-capture segmentation (UNet model, 10 MB) and Magic Eraser (a system of models) to the new server-side generative AI experience, which handles tasks like relocating objects and reimagining backgrounds. Key challenges include managing model size (now hundreds of MB), client-server latency, ambiguous problem scoping (e.g., moving from 5% to 80% reliability), and trust and safety. Ma advocates using evals, reducing ambiguity through product-research collaboration, and iterating from large general models to smaller, faster ones for production. He also highlights that Google Photos serves 1.5 billion monthly active users and processes hundreds of millions of edits per month, and the editor is being rebuilt as AI-first.

Dream Machine: Scaling to 1m users in 4 days — Keegan McCallum, Luma AI
Jul 19, 2025 · 19:03
Keegan McCallum, Head of ML Infrastructure at Luma AI, details how the company's Dream Machine model scaled from 500 to 9,000 H100 GPUs within hours to handle 1 million users in four days, outpacing ChatGPT's initial growth. He explains that their initial Triton inference server setup was brittle and ill-suited for multi-GPU, multi-node video models, prompting a re-architecture to a custom serving stack on vanilla PyTorch. To solve work starvation across user tiers, they implemented an SLO-based aging system that ranks jobs by the percentage of their worst-case wait time elapsed. For managing dozens of model versions, they store immutable full Python environments and checkpoints in object storage, with a YAML file controlling active deployments and enabling zero-downtime rollouts across thousands of GPUs. McCallum also discusses partnerships with Nvidia, AMD, and Grok, and how Luma's broader mission is to build general multimodal intelligence that generates, understands, and operates in the physical world.

ComfyUI Full Workshop — first workshop from ComfyAnonymous himself!
Jul 19, 2025 · 51:25
ComfyAnonymous and Yedrick Kosinski present ComfyUI, an open-source node-based canvas for generative AI that has become the top 150 most popular GitHub repos with 78,000 stars and 3 to 4 million active users. Started as a personal project in January 2023, it now supports image, video, audio, 3D, and text models, with 22,000 custom nodes from 3,000 developers. A key feature is embedding full workflows into generated files for easy sharing. The team discusses its story, including ComfyAnonymous's time at Stability AI and founding ComfyOrg, and addresses questions on CFG, VAEs, LoRAs, control nets, API nodes for remote generation, and future plans for cloud inference and improved onboarding.

Design like Karpathy is watching — Zeke Sikelianos, Replicate
Jul 19, 2025 · 19:26
Zeke Sikelianos of Replicate analyzes Andrej Karpathy's experience building and deploying MenuGen, a vibe-coded web app that turns menu photos into images, to argue that LLMs are now the primary audience for developer tools. He details Replicate's specific failures—rate limiting and outdated docs that blocked Karpathy—and the fixes they implemented: adding LLMs.txt files for Markdown-friendly documentation, promoting curl commands as the LLM-friendly interface, and launching an MCP server built on OpenAPI schemas. Sikelianos advocates for 'boring technology' like SQL and 'good API hygiene' to keep payloads small and information-dense for LLM context windows. He also calls for better payment acceptance and documentation discipline, stressing that unblocking power users like Karpathy (whose CEO intervened) should not require a viral blog post.

On Curiosity — Sharif Shameem, Lexica
Jul 19, 2025 · 18:35
Sharif Shameem, founder of Lexica, argues that curiosity is the main force for pulling ideas from the future into the present, and that building and sharing demos is the best way to explore AI models' hidden capabilities. He recounts early GPT-3 demos from 2020-2021, when the model had a 2,000-token context window and cost $75 per million output tokens, showing how he built a JSX compiler in the browser, a shopping agent that parsed web pages, and a multi-step reasoning tool called MultiVAC. Shameem emphasizes that AI engineering is more like excavating than traditional engineering, and that researchers often don't know the full capabilities of their own models. He closes by invoking computing pioneer J.C.R. Licklider, arguing that today's AI engineers have a moral obligation to follow their curiosity and share what they discover.

Real world MCPs in GitHub Copilot Agent Mode — Jon Peck, Microsoft
Jul 19, 2025 · 14:27
Jon Peck from Microsoft demonstrates how Model Context Protocol (MCP) servers enable GitHub Copilot Agent Mode to solve real-world engineering problems beyond vibe-coding. He shows a README-driven workflow where Agent Mode builds a full app from a specification, then adds MCPs for database access and GitHub operations. Specifically, he configures a PostgreSQL MCP to pull live schema and data into mock JSON for testing, and the GitHub MCP to commit changes and create pull requests automatically. He emphasizes the read-only safety of the Postgres MCP and recommends using Copilot instructions to enforce practices like change logs. The episode covers how MCPs extend Copilot's capabilities to interact with data sources, testing tools, and DevOps pipelines securely.

The rise of the agentic economy on the shoulders of MCP — Jan Curn, Apify
Jul 18, 2025 · 18:08
Jan Curn, founder of Apify, argues that MCP (Model Context Protocol) enables a future agentic economy where AI agents autonomously discover and purchase tools from other agents or businesses (B2A/A2A). He explains that Apify's marketplace of 5,000 Actors (Docker-based tools) now integrates with MCP, allowing agents to dynamically discover and call any Actor via tool discovery — a key MCP feature. Curn demonstrates this with Claude Desktop, where an agent uses Apify's MCP server to find a venue, scrape Twitter, and even fill a form via a nested MCP server from Browserbase, all without prior configuration. He notes that Apify pays creators over $250,000 monthly, with total Actor revenue exceeding $1.5M/month, and that any developer can publish an Actor to monetize their tools instantly across the ecosystem. The talk closes with open questions about reliability, trust, and whether autonomous agent interaction can lead to AGI.

MCP is all you need — Samuel Colvin, Pydantic
Jul 18, 2025 · 15:24
Samuel Colvin, creator of Pydantic, argues that MCP (Model Context Protocol) is all you need for agent-to-agent communication, simplifying what many are overcomplicating. He explains that MCP's tool calling primitive offers advantages over OpenAPI, such as dynamic tools, logging, and sampling—a mechanism for MCP servers to request LLM calls through the client, enabling nested agentic workflows. Colvin demonstrates a research agent using Pydantic AI that queries the BigQuery public PyPI dataset via an MCP server over stdio, showing how sampling allows the server to generate SQL itself while keeping the main agent's context lean. The demo logs progress to Logfire, an observability platform, and outputs human-readable results. He emphasizes that MCP’s standard I/O mode and extensibility make it ideal for autonomous agents beyond its original desktop coding use case.

Full Spec MCP: Hidden Capabilities of the MCP spec — Harald Kirschner, Microsoft/VSCode
Jul 18, 2025 · 14:53
Harald Kirschner from Microsoft/VSCode argues that MCP's full specification unlocks powerful stateful interactions beyond the common 'tools-only' implementations, transforming AI assistants into more contextual and efficient agents. He highlights underused primitives like resources for rich data context and sampling for server-requested LLM completions, demonstrated via a dungeon game where dynamic tool discovery adapts to game state. VS Code's upcoming full spec support includes dynamic tool discovery, user-defined tool sets, a debug mode for server development, and support for streamable HTTP to reduce stateful server churn. Upcoming features like elicitations will allow tools to request user input directly. Kirschner calls on developers to build progressive, full-spec servers and contribute feedback to the open ecosystem, emphasizing that client and SDK support will follow as usage grows.

Shipping an Enterprise Voice AI Agent in 100 Days - Peter Bar, Intercom Fin
Jul 18, 2025 · 17:10
Peter Bar, Product Lead at Intercom, details the 100-day build of Fin Voice, an AI voice agent for enterprise phone support. The agent handles knowledge-based queries using a stack of speech-to-text, LLM, text-to-speech, RAG, and telephony, achieving ~1 second latency for simple queries and using filler phrases for longer ones. Key product decisions included focusing on out-of-office hours as an initial wedge, designing conversations for voice differences like answer chunking, and prioritizing integration with human support workflows over model improvements. The team measured success via resolution rate (user confirming resolution or not calling back within 24 hours) and used an LLM-as-judge for quality analysis. Bar argues that voice AI is the next frontier in customer service, citing cost reduction from $7–12 per human-handled call to 3–20 cents per minute with AI.

The State of Generative Media - Gorkem Yurtseven, FAL
Jul 16, 2025 · 17:14
Gorkem Yurtseven, co-founder and CTO of fal.ai, argues generative media lowers creation's marginal cost to near zero, transforming advertising and e-commerce. He traces evolution from DALL·E 2 to open-source models like Stable Diffusion and FLUX, with the playing field evening quickly. Video models (Sora, Veo 3) are accelerating: fal's platform saw video usage jump from near zero to 30% of revenue, and he predicts the generative video market 100x-250x bigger than image. Advertising leads with hyper-personalized, interactive ads, like fal's A24 campaign turning selfies into toy soldiers. In e-commerce, virtual try-on is a proven product-market fit. Future real-time video generation will blur games and movies.

Teaching Gemini to Speak YouTube: Adapting LLMs for Video Recommendations to 2B+DAU - Devansh Tandon
Jul 16, 2025 · 22:51
Devansh Tandon, a Product Manager at Google leading YouTube's discovery system, details how YouTube adapted Gemini LLMs to power its recommendation engine for billions of daily active users. The team built SemanticID, a tokenization system that compresses video features into semantically meaningful tokens, creating a new language for YouTube content. They then continued pre-training Gemini on sequences of user watches to make the model bilingual in English and this video language. For generative retrieval, they prompt the adapted model with user demographics and watch history to output video recommendations as SemanticIDs, achieving 95%+ cost savings to serve at scale. Challenges include serving billions of users with low latency and handling video freshness—Taylor Swift's new music video must be recommendable within minutes. Tandon argues LLM-led recommendations are a bigger consumer application than search and hints at future interactive, steerable recommendations and even personalized content creation.

Transforming search and discovery using LLMs — Tejaswi & Vinesh, Instacart
Jul 16, 2025 · 21:10
Vinesh Gudla and Tejaswi Tenneti from Instacart detail how they use LLMs to overhaul search and discovery. They address challenges with conventional search: broad queries suffer cold start, tail queries lack engagement data. Using LLMs with Instacart's domain knowledge—e.g., top-converting categories as context—they improved precision by 18 percentage points and recall by 70% for tail queries, and cut zero-result queries. For discovery, LLMs generate complementary/substitute items; but pure LLM suggestions like 'chicken' for 'protein' missed user intent, so they augmented prompts with behavioral data (top categories, subsequent queries) to boost engagement and revenue. They use a hybrid approach: pre-compute offline for head/torso, distilled Llama 8B for long-tail, and LLMs as judges for evaluation. Key takeaways: combining LLM world knowledge with domain-specific data is critical, and evaluation is as hard as generation.

Netflix's Big Bet: One model to rule recommendations: Yesu Feng, Netflix
Jul 16, 2025 · 22:28
Netflix staff research scientist Yesu Feng explains the company's bet on a single transformer-based foundation model for all recommendation use cases, replacing dozens of specialized models. He validates that scaling laws apply to recommendation systems, with the model growing from a few million profiles to 1 billion parameters, and that multi-token prediction (inspired by DeepSeek) notably improves metrics by forcing less myopic learning. The foundation model is consumed via three patterns: as a subgraph in downstream models, through precomputed embeddings (both member and content), and via fine-tuning or distillation. Feng reports that over the past 1.5 years, the model has been integrated into many applications, with A/B test wings showing high leverage and infrastructure consolidation. Future directions include universal representation for heterogeneous content types, generative retrieval for collection recommendation, and prompt tuning for faster adaptation.

360Brew: LLM-based Personalized Ranking and Recommendation - Hamed and Maziar, LinkedIn AI
Jul 16, 2025 · 22:00
LinkedIn AI's Hamed Firooz and Maziar Sanjabi present BrewXL, a 150-billion-parameter foundation model for ranking and recommendation that personalizes the platform's feeds, jobs, and search. They show that training a large model then distilling to a 3B model outperforms training a small model from scratch. Scaling data, model size (from 7B to 8x22B), and context length all improve performance, though context beyond the trained length degrades accuracy. The model achieves zero-shot generalization on out-of-domain tasks, matching or beating task-specific production models, and significantly reduces the cold-start gap for users with fewer than five interactions. For serving, they combine gradual pruning, mixed-precision quantization (FP8 for most layers, FP32 for the LM head), and 4D attention masks to score up to 500 items without cross-attention, yielding a 7X latency reduction and 30X throughput increase per GPU.

What We Learned from Using LLMs in Pinterest — Mukuntha Narayanan, Han Wang, Pinterest
Jul 16, 2025 · 18:13
Pinterest search engineers Han Wang and Mukuntha Narayanan present four key learnings from integrating LLMs into the platform's search relevance system. Fine-tuned LLM cross-encoders improve relevance prediction by 12% over multilingual BERT and 20% over SearchSage using an 8B model. VLM-generated image captions and user action features enrich pin text representations and further boost performance. Knowledge distillation into a bi-encoder student model, trained on 100x more data via semi-supervised learning, enables production serving with real-time query embeddings and 85% cache hit rate. Relevance-tuned embeddings serve as general-purpose semantic representations across multiple surfaces, and the system yields relevance gains in multiple languages and countries despite a predominantly US training set.

Measuring AGI: Interactive Reasoning Benchmarks for ARC-AGI-3 — Greg Kamradt, ARC Prize Foundation
Jul 16, 2025 · 18:28
Greg Kamradt, President of ARC Prize Foundation, introduces ARC-AGI-3, the first interactive reasoning benchmark for AGI that drops agents into novel games without prior instruction, forcing exploration to solve tasks. Unlike static tests, this benchmark measures skill acquisition efficiency—how quickly an AI learns and applies new skills—using human baselines from 400+ in-person tests. It strips away language and trivia, relying only on core knowledge priors (basic math, geometry, agentness, objectness). A public training set of ~40 games will be released, but performance is measured on a private evaluation set of 120 games unseen by developers or AI. Kamradt asserts that as long as AI cannot outperform humans on these problems, we do not have AGI; a sandbox preview with five games and a mini agent competition is planned for next month, with full launch in Q1 2026.

RL for Autonomous Coding — Aakanksha Chowdhery, Reflection.ai
Jul 16, 2025 · 19:27
Aakanksha Chowdhery, CEO of Reflection AI and former lead researcher on PaLM and Gemini at Google, argues that reinforcement learning (RL) scaling is the next frontier for autonomous coding agents. She explains that inference-time techniques like majority voting and self-revision improve accuracy but require many samples (e.g., 10,000 for rare correct generations). RL training can learn to generate correct outputs directly, especially in verifiable domains like code with unit tests. Reflection AI aims to build superintelligence starting with autonomous coding, leveraging automated verification to design better reward functions. She notes challenges in scaling RL, including system complexity and reward hacking, but sees coding as ideal due to execution feedback.

Recsys Keynote: Improving Recommendation Systems & Search in the Age of LLMs - Eugene Yan, Amazon
Jul 16, 2025 · 20:54
Eugene Yan's keynote presents three innovations for recommendation systems: Semantic IDs, LLM-augmented data, and unified models. Kuaishou’s trainable multimodal Semantic IDs increased cold-start coverage by 3.6% and velocity by 3.5% by clustering content embeddings. Indeed used GPT-4 fine-tuning and distillation to filter bad job recommendations, reducing bad recs by 20% while boosting application rate 4% and cutting unsubscribes 5%. Spotify’s LLM-generated exploratory search queries drove a 9% increase in exploratory queries for new categories like podcasts. Netflix’s Unicorn unified ranker matched or exceeded specialized models across search and recommendations, while Etsy’s unified embeddings with a quality vector achieved a 2.6% sitewide conversion lift and 5% more search purchases.

Benchmarks Are Memes: How What We Measure Shapes AI—and Us - Alex Duffy, Every.to
Jul 15, 2025 · 15:44
Alex Duffy argues that AI benchmarks function as cultural memes—ideas that spread and shape what models learn—giving those who design them immense power over AI's trajectory. He traces the lifecycle from a single person's idea to saturation, using examples like 'How many Rs in strawberry' and Pokémon. Duffy introduces AI Diplomacy, a benchmark where language models negotiate and betray each other, revealing that models like DeepSeek R1 and Gemini 2.5 Flash excel at social manipulation while Claude models are naively optimistic. He warns against benchmarks that reward sycophancy (like ChatGPT's thumbs-up training) and advocates for multifaceted, experiential, and generative benchmarks that empower people. Duffy urges the audience to ask non-AI people what they care about, turning benchmarks into tools that build trust and define humanity's role in an AI world.
Powered by PodHood