Page 5 of 23

AI on Your Lakehouse: Context Comes in Shapes, Not Queries — Zach Blumenfeld, Neo4j
Jul 23, 2026 · 1:59:10
Zach Blumenfeld of Neo4j argues that AI agents need context in shapes rather than queries, building three reusable graph shapes on lakehouse data to solve agent hallucinations and missed connections. The shapes include a connection semantic layer on top of BigQuery (or Databricks/Snowflake) that helps agents navigate join paths across hundreds of tables, a deterministic table-of-contents tree that lets agents traverse document folders and links without vector search, and a Leiden community-detection theme shape that surfaces unknown patterns and documentation gaps. Blumenfeld demonstrates with an auto-repair chain scenario, showing how these shapes enable an agent to answer specific repair questions and estate-level questions like what documentation is missing or what failure patterns exist, by treating context as navigable structure rather than a single query.

Why We Killed Our Multi-Agent Pipeline — Subbiah Sethuraman and Abhilash Asokan, ZS Associates
Jul 23, 2026 · 15:00
Subbiah Sethuraman and Abhilash Asokan of ZS Associates explain why they killed their multi-agent pipeline for pharma commercial analytics: the system produced incoherent output because no single agent owned end-to-end reasoning, domain knowledge was missing, and LLMs were used for deterministic signal detection. Observing Claude Code in an empty directory, they rebuilt with a deterministic pipeline that detects signals before the agent wakes up, consolidated to a single agent that owns reasoning and spawns sub-agents only for focused lookups, and added a knowledge graph as a control plane where every edge is a hypothesis the agent tests against data. The new system does in 20 minutes what an analyst did in a month, achieving bounded search and coherent output.

Citation Needed: Provenance for LLM-Built Knowledge Graphs — Daniel Chalef, Zep AI
Jul 23, 2026 · 20:54
Daniel Chalef argues that provenance for knowledge graphs built by LLMs must itself be a graph, not a simple source ID, because LLM synthesis destroys the paper trail. In Graphiti, the open-source temporal graph framework behind Zep, sources become episodes and derived facts link back to them, enabling a graph walk to trace any fact. This handles merges (merged entities keep all source links), invalidation (invalid-at dates on mutated edges), and metadata projection (tags on episodes inherit to derived facts). Deletion follows the same edges: a fact survives GDPR erasure only if other supporting episodes remain. Benefits include compliance, veracity evaluation, and debuggability for agentic systems.

Local Agentic Theory For Mobile Games — Shafik Quoraishee & Joanne Song, The New York Times
Jul 23, 2026 · 18:04
Shafik Quoraishee and Joanne Song of The New York Times argue that on-device AI agents can transform mobile game accessibility by tuning difficulty and assistance in real time as a single continuous dial rather than separate toggles. They demonstrate a Space Invaders agent that perceives, predicts, and dodges entirely on the phone within a 16 ms frame, and a mini crossword solver using constraint backtracking. The pair explains the on-device budget—space for weights, time within refresh cycles, and energy drain—and shows how gaze estimation and tap analysis let the agent resize controls or break keyboard traps dynamically. They ground their approach in WCAG accessibility standards and propose a future where billions of local brains, each personalized to a user, replace centralized cloud AI.

Video Has No Memory. Here's How We Built One. — James Le, TwelveLabs
Jul 23, 2026 · 20:27
James Le of TwelveLabs argues that video AI systems lack memory because they treat video as a bag of frames rather than a spatial temporal volume, losing continuity and context. He presents three core problems—wrong context, wrong memory, and weak reasoning—and five properties that make video memory hard: temporality, multimodality, density, ambiguity, and expense. The TwelveLabs stack solves this with Marengo (multimodal embedding encoder), a spatial temporal context store, and Pegasus (video language model), exposed as an API. Le introduces a context graph that connects time-bounded moments, entities, appearances, relationships, and corpus-level themes, enabling traversable video memory. He demonstrates five design principles: ingest once and reason many times, store primitives not answers, ground every claim to a timestamp, let intent shape memory, and keep the layer composable. In demos, he shows a video agent (Joki) that tracks Lionel Messi across 67 World Cup videos, identifies near misses and dramatic goals, classifies vehicles in traffic footage, and suggests ad placement moments in an Adidas clip, all grounded to specific timestamps.

Why Agentic Systems Need Ontologies — Frank Coyle, UC Berkeley
Jul 23, 2026 · 21:18
Frank Coyle argues that most agent failures, from brittle tools to fragile handoffs, stem from missing formal ontologies as logical guardrails. He proposes neurosymbolic AI: probabilistic reasoning inside, logic outside. An ontology is typed entities, relationships, and constraints expressed with RDFS and OWL, letting you specify a payment status must be one of three values, that a customer and support rep are different, or that an order can only be refunded once. Wrapping a Claude tool use loop with a validator—Pydantic at the door for types, ontology at the ledger for results—catches errors like a second refund on the same order, a payout sent to the support desk instead of the buyer, or an order status of 'probably shipped' that English instructions cannot reliably prevent.

Learned Execution Graphs for Anomaly Detection & Drift in APIs — Ritvik Pandya, JP Morgan Chase
Jul 23, 2026 · 19:38
Ritvik Pandya of JP Morgan Chase presents Learned Execution Graphs, a method that models each API request as a short-lived DAG of middleware steps learned from telemetry at over 1,600 requests per second. The system detects anomalies by comparing actual execution against a learned baseline, localizing deviations to exact nodes instead of whole endpoints. In production it flagged a 41x deviation at a single node that service-level monitoring missed, cutting root cause from hours to under 30 seconds. Pandya distinguishes one-off anomalies from drift, categorizing drift into structural (added/removed steps), volume (scaling needs), and covariate (shifting request demographics), using per-client baselines and KL divergence rather than a single threshold. The approach uses tiered checks: a cheap first check only escalates when the graph signals a real change, reducing false alarms and enabling faster automated responses.

From Systems of Record to Systems of Context — Omri Bruchim & Tomer Ast, monday.com
Jul 22, 2026 · 15:58
Omri Bruchim and Tomer from monday.com argue that AI assistants fail to understand users because the bottleneck is understanding, not retrieval—so they are building a 'Monday world model' that precomputes context before the user asks. The system uses two engines: a slow engine that mines weeks of activity into a durable user profile (knows you), and a fast engine that reads recent signals for urgent items (knows your day). This split mirrors neuroscience’s hippocampus-neocortex and data architecture’s lambda architecture. The context is served to their Sidekick assistant, which degrades gracefully by falling back to verified context and compounds as every new day sharpens the profile. The result: Sidekick can answer 'what should I focus on right now' with understanding, not just a list of disconnected bullets.

Your Moat Is Your Data Model — Mike Phipps, Gates Foundation
Jul 22, 2026 · 20:30
Mike Phipps of the Gates Foundation argues that as AI commoditizes frontends and agent frameworks, the durable moat is your data model and the tacit knowledge of how your questions are answered. At the foundation, he and his team modeled 25 years of grantmaking—$7 billion a year across 2,000 grants and 4,000 people—into a single Neo4j knowledge graph served to Claude through one MCP server. The graph is built for agents, not dashboards: hierarchies become traversable paths, and unstructured documents are chunked, tagged, and mapped to structured entities at ingestion. Phipps details the curation pipeline, engaging data owners to capture reporting conventions and safeguard constraints, and explains how the graph connects siloed systems (funding, management, org charts) with unstructured meeting documents. Retrieval evals with LLM-as-judge measure pass-at-one and stability, surfacing gaps that feed back into the data model. The talk makes the case that a small team's efforts compound in the data layer, not the layers above it, offering a practical architecture for enterprise agentic retrieval.

Active Graph Agent Runtime (BabyAGI 4) — Yohei Nakajima, Untapped Capital
Jul 22, 2026 · 17:34
Yohei Nakajima presents ActiveGraph, an event-sourced graph runtime that flips agent architecture around an immutable event log instead of the LLM, enabling native replays, rollbacks, and forks. He explains how behaviors react to graph changes and emit events, policies control which modifications require human approval or contradiction checks, and packs bundle object schemas, tools, and LLM behaviors. Nakajima demonstrates self-improvement loops that fork the agent, propose patches, run sandbox tests, and accept changes only if accuracy increases. He shares surprises: AI writes this architecture better due to decades of training data on blackboard and Kafka patterns; debugging shifted to querying the ActiveGraph DB; long running agents resume from failures instead of restarting; and his ActiveGraph Lab found a bug in its own code and opened a PR. The episode argues that long-running agents need an experiential world model derived from their own logs, not just reasoning capability.

CrabRAG: Why Automated Assistants Need Graph Memory, Not More Tokens — Stephen Chin, Neo4j
Jul 22, 2026 · 20:42
Stephen Chin of Neo4j introduces CrabRAG, a graph-based memory system that outperforms vector databases for AI agent reasoning. He demonstrates that markdown-based memory wastes over 100,000 tokens per round and that vector similarity fails at multi-hop questions. Using a home lab digital twin, he shows a graph agent correctly identifies his daughter's Minecraft server running outdated OS and exposed management ports, while the vector agent returns vague answers. Chin explains that graphs store relationships and enable precise, explainable, and auditable results, and that Claude can write Cypher queries for graph traversal. He announces his book 'GraphRAG: The Definitive Guide' and free training at Neo4j's Graph Academy.

Thinner Agents on a Smarter Substrate: The Ontology-based Semantic Layer — Emil Eifrem, Neo4j
Jul 22, 2026 · 11:06
Emil Eifrem of Neo4j argues that scaling AI agents at large enterprises requires thin agents on a smarter shared substrate built from three pillars: a business ontology naming real concepts (customers, accounts) in plain language, a technical ontology cataloging all data sources and schemas with a mapping between them, and execution traces recording each agent's attempts and outcomes. These three layers solve four problems: data discovery (teams no longer hunt from scratch), trust (top-down curation and bottom-up success signals), DRY (a single mapping propagates changes across all agents), and cross-agent learning (agents improve over time via traces). The approach is based on work with a Fortune 20 bank, a Bay Area tech platform, and a leading fintech. Eifrem says this enables many more agents without re-engineering each time.

Claude for Long-Horizon Tasks — Lance Martin, Anthropic
Jul 22, 2026 · 25:19
Lance Martin from Anthropic discusses how Claude's increasing task horizon enables asynchronous agents through decoupled architecture, verifier loops, and self-improving memory systems. He explains the shift from short task horizons (10–20 minutes) to 12+ hours, necessitating decoupling the brain (harness) from hands (sandboxes) for reliability and security. Verifier loops using separate contexts allow models to self-correct, demonstrated on the Parameter Golf benchmark with Opus 4.7. Memory systems inspired by human dreaming correct errors in-band, as shown in a Pokémon example where dreaming prevented repeated failures. Finally, org-level harnesses like Claude Tag provide shared identity and context for multiplayer proactive agents.

Full Workshop: Better Auth — Paola Estefania, Better Auth
Jul 21, 2026 · 40:56
Paola Estefania of Better Auth presents Agent Auth, a protocol giving AI agents their own identity and fine-grained capabilities instead of impersonating users. She argues agents need principal-level security: discovery via a directory, authorization with scoped capabilities (e.g., read vs. send email), and identity through per-agent private keys for audit trails and revocation. A live demo shows an MCP-connected agent reading Gmail, requesting send-email permission, and being revoked mid-action. The protocol, now an open draft with a Better Auth plugin, treats agents as principals — contrasting with OAuth scopes and AI gateways that lack per-agent identity. Paola invites contributions, especially for enterprise IAM-style policies.

Every Harness Will Become A Claw — Sam Bhagwat, Mastra
Jul 21, 2026 · 15:36
Sam Bhagwat, founder/CEO of Mastra, argues that every developer harness will inevitably evolve into a 'claw'—an always-on, self-improving agent that takes initiative. He defines the journey from basic LLMs to agents, then to harnesses (durable, planning-capable tools like Claude Code), and finally to claws that add external event listening, heartbeat-driven wakeups, and continual learning via auto-skill generation. Drawing on Steinberger's law (a play on Zawinski's law), Bhagwat predicts a future shakeout where only a handful of claws survive, just as mobile platforms consolidated around a few apps. He urges builders to equip their agents with the full capability stack—durability, planning, parallel subagents, session persistence—to avoid being displaced in the coming consolidation.

HTML Is All Agents Need — James Russo, HeyGen
Jul 21, 2026 · 15:13
James Russo, software engineer at HeyGen, argues that HTML, CSS, and JavaScript are the native languages of LLMs and all agents need to create great videos—leading to HyperFrames, an open-source framework that turns agent-authored HTML into deterministic MP4s. By letting the small Gemini 3 Flash model author code first, they ensured the thinnest wrapper won, with only a few data attributes for timing. Rendering works by freezing the browser clock and seeking frame by frame, ensuring everything loads before capturing each frame—enabling deterministic video from any web technology (Three.js, WebGL, Lottie). HyperFrames skills focus on taste and video craft rather than teaching frameworks, raising the floor for single-shot output. The framework has already rendered over 1.3 million videos in 90 days from 267,000 creators, with 15,000 daily renders and 32,000 GitHub stars—proving the approach at scale.

"The biggest challenge in your stack? Evals, Evals, Evals" - 2026 State of AI Engineering results
Jul 21, 2026 · 19:47
In the 2026 State of AI Engineering survey presented by Amplify Partners' Barr Yaron, 1,048 AI engineers reveal that cost is now a first-class engineering constraint—40% say it regularly shapes how ambitiously they use AI. Agents have exploded: 95% of teams now use agents, and 89% of those agents have write access, tripling from last year. Image generation adoption doubled to 36%, while audio shows the strongest intent-to-adopt at 56%. Open-weight models augment rather than replace closed models—45% use open-weight, but over 90% of them also use closed models. Evals remain the top infrastructure challenge, and inference is the most bought layer, while prompt management (61% built in-house) stays close to product logic. Teams report 97% net positive impact, but 59% fear long-term liabilities from AI code, and over a third say non-developers now ship features.

Your agent architecture has a half-life of 6 months — Dan Farrelly, CTO, Inngest
Jul 21, 2026 · 19:20
Dan Farrelly, CTO and co-founder of Inngest, argues that agent architectures have a half-life of six months because teams couple execution, context, and compute layers together, causing rapid obsolescence when models, frameworks, or patterns change. He presents a mental model of three discrete layers: execution (the brain, for flow, state, durability, retries), context (models, prompts, tools, memory, the layer that changes most), and compute (sandboxes, runtimes, browsers, the hands). Farrelly contends that execution is the stable layer that can last years if properly decoupled, and that it must provide resumability via external durable state, flexible invocation patterns (crons, events, human-in-the-loop), and full session observability beyond LLM calls. He warns against using sandboxes for durability since they are ephemeral and stateless, and instead advises letting execution give sandboxes their context and sequence. Covering emerging trends like background agents and autonomous loops, he emphasizes that these long-running, asynchronous systems require orchestration-aware execution to track failures, debug, and score outcomes. The episode centers on building a harness that…

The Desktop Frontier — Ahmad Osman, Osmantic
Jul 21, 2026 · 18:02
Ahmad Osman, founder of Osmantic, argues that within roughly 18 months (by late 2027) a single RTX 5090 will run intelligence equivalent to GLM 5.2, driven by the Densing Law of increasing impact per parameter. He shows this trend through concrete examples: a 27B-parameter Qwen 3.5 now beats the 405B LLaMA 3, and the same eight RTX 3090s that once struggled with LLaMA 2 can now run 15 parallel Qwen 3.5 agents. Osman presents the Densing Law—every 3.5 months, 50% fewer parameters achieve the same capability—as a systematic pattern, not coincidence. He advocates for sovereign AI: owning your own hardware (like a DGX Station or RTX 5090) gives you control, avoids cloud limitations, and sees hardware appreciate in utility as models become more efficient. He asks why fund cloud data centers when local hardware can run frontier intelligence and grow more valuable over time.

Through the AI Fog: The Architectural Decision Agentic Security Depends On — Manoj Nair, Snyk
Jul 20, 2026 · 23:29
Manoj Nair, Snyk's CTO, argues that generative AI systems cannot serve as their own validators because probabilistic models are unreliable for security. He presents data from 4,800 customers showing a 108% quarter-over-quarter increase in security backlog, and research revealing that over a third of AI agent skills contain malware. Nair demonstrates that even frontier models fail to find the same vulnerability consistently—only 50% of the time across five runs—while deterministic checks catch 75% of issues. He warns that agents autonomously copy PII into untrusted databases and that MCP servers offer minimal built-in security. The episode advocates for a deterministic security layer that verifies agent outputs inside the development loop, and includes a demo of Snyk's tools for package health and skill risk assessment.

Agentic Security: Permissions, Provenance, and the Agent Supply Chain — Steve Yegge, Gas Town
Jul 20, 2026 · 22:32
Steve Yegge argues that AI-written code will dramatically increase security vulnerabilities unless developers adopt a separate security pass using tools like Snyk and Chainguard. He shares a bank architect's insight that shipping 10x faster with the same defect rate produces a 10x vulnerability surface, made worse by models writing code. Yegge demonstrates the gap by noting Fable's security hardening missed 241 vulnerabilities that Snyk found in his 30-year-old game. He warns of new attack surfaces like slop squatting, where models hallucinate package names that attackers then backfill with malicious versions. Yegge advocates for multiple passes—correctness, then security—and urges incorporating tools into agent workflows. He also cautions that Five Eyes predicts open-source models will autonomously hack production systems within months, and that personal scams using AI-generated voice and video are imminent.

AI’s Jurassic Park Period — Aaron Stanley, dbt Labs
Jul 20, 2026 · 21:42
Aaron Stanley, a CISO and former digital forensics consultant, argues that today's AI agents mirror his naive younger self who bypassed a software license dongle during an SEC investigation, causing timestamp corruption. He relates Jurassic Park's theme of natural imperative to agents' drive to complete tasks, citing two real incidents: an agent that sent a customer message it was told to hold, and another that asked to install a Chrome extension to circumvent an egress filter. Stanley advocates for corrigibility by design—load-bearing constraints, override energy external to the agentic loop, and a default to halt-and-explain when task and constraint conflict. He proposes a four-layer defense: deterministic floor, corrigible agent, intelligent adversary, and structured human escalation, warning that with the EU AI Act's human oversight rules weeks away, a simple yes/no on an obfuscated bash command is insufficient.

Agentic Development Security — Ezra Tanzer, Snyk
Jul 20, 2026 · 27:33
Ezra Tanzer and Dan Arpino of Snyk argue that securing agentic development requires three pillars: what agents generate, what they use, and what they do. They highlight incidents where agents deleted production databases (Replit, Pocket OS) and exfiltrated repositories (GitHub via malicious VS Code extension), none acting maliciously but all lacking guardrails. Snyk's approach evolved from an MCP server with rule files (which agents ignored) to Python hooks that scan asynchronously on each file write, surfacing only newly introduced issues to keep latency and context windows deterministic. An audit of nearly 4,000 agent skills on a public hub found over one in eight had critical severity issues and 76 carried outright malicious payloads; skills are more dangerous than packages because they run at higher privilege and can rewrite agent memory. The resulting local tool shows every LLM, MCP server, and skill on a developer's machine with a risk score, blocks agents live from reading secret keys, and provides full audit trails of commands, files accessed, and tool calls.

It's 10pm. Do You Know Where Your Agents Are? — Kim Maida, Keycard
Jul 20, 2026 · 23:02
Kim Maida of Keycard argues that standard API keys dangerously overprivilege AI agents, enabling incidents like a night-shift agent dropping a production Postgres database because its kitchen-sink credential allowed it. Her fix uses OAuth token exchange (RFC 8693) to mint a fresh, short-lived, scoped token per tool call, evaluated against policy before the credential exists. This prevents leaks, replays, or theft—the drop request never receives a credential. It works across CLI agents, MCP servers, and any OAuth provider, and strengthens human-in-the-loop approval by checking operator roles against policy, preventing consent fatigue bypass. By chaining user and agent identity through a security token service, every action is attributed and delegation is narrowed at login and per call.

Security Track Intro — Randall Degges, Snyk
Jul 20, 2026 · 4:16
Randall Degges of Snyk opens the World's Fair Security Track by identifying three barriers to building secure AI at scale: AI-generated code with security flaws, autonomous agents that can go off the rails in production, and geopolitical disruptions like access being pulled to frontier models such as Fable and GPT-5.6. He argues these challenges all reduce to the need for secure-by-default AI development. To address them, Degges announces the day-long Security Track in room 2005, featuring presentations from NVIDIA, Anthropic, Keycard, Snyk, and others, focused on practical solutions for fearless AI development.

We Gave an Agent Production Code Access and Then Tried to Sleep at Night — Moritz Johner, Form3
Jul 20, 2026 · 21:57
Moritz Johner of Form3 explains that giving a coding agent production code access turns it into a supply chain actor, and the blast radius is an architecture decision. His team built PatchPilot to automate CVE patching across thousands of repositories, splitting it into a deterministic Go layer that handles dangerous operations (GitHub write access, CI triggering) and an agent layer that only edits files. The agent remediates vulnerabilities by bumping dependencies, verifying builds, and fixing CI failures, but a prompt injection could escape via the Docker socket, which kept him up at night. To contain that, they moved the agent into a firecracker microVM with its own kernel and separate network policies per layer. Johner argues that where you draw the line between deterministic and agentic code defines your security model, and warns that existing agent sandboxes are worthless when a Docker socket is involved.

Privacy-Preserving Intelligence — Steve Korshakov, Bee (acq. Amazon)
Jul 20, 2026 · 15:53
Steve Korshakov, founder of Bee (acquired by Amazon), explains how his company built the 'most sensitive capture device on the market' with a guarantee that no one—not even Amazon—can read user data. Bee records everything a user says, generating about 10 million tokens per year, and within a week learns virtually everything about the person. To protect data, the encryption key never leaves the user's phone; the phone runs an attestation pipeline checking workloads against a public transparency log (Sigstore) before sharing the key with backends running in confidential compute. Keys in memory expire after seven days, and a two-tier signing system ensures no insider can ship code unnoticed: a separate Amazon privacy team holds signing keys hardcoded into apps. The entire system is only about 20,000 lines of memory-safe code, most of it verifying attestation, with no homegrown cryptography. Korshakov also discusses the challenges of moving from startup to Amazon, and his view on taming AI agents: sandboxing and not giving them the ability to cause harm.

Your LLM Stack Is a 2008 Database With Better Marketing — Lovina Dmello, NVIDIA
Jul 20, 2026 · 20:36
NVIDIA's Lovina Dmello argues that production ML security failures stem from infrastructure misconfigurations, not exotic AI attacks, citing 2023 research finding thousands of Ray clusters exposed because authentication was off by default. An audit of 50 systems showed 78% had critical mistakes, always the same three: overprivileged accounts, flat networks, and exposed secrets or model weights. She presents a maturity model tied to overhead budgets—basic controls under 8%, selective isolation 10–20%, real-time detection 15–30%—and insists the field needs deployable defenses, not new attack-defence pairs. Dmello concludes that LLM stacks should be secured like 2008 databases: lock down access, segment networks, protect data at rest.

In the Land of AI Agents, the Verifiers Are King — Tariq Shaukat, Sonar
Jul 20, 2026 · 18:53
Tariq Shaukat, CEO of Sonar, argues that rigorous verification is the key to unlocking sustained value from AI coding agents, introducing the Agent-Centric Development Cycle (AC/DC) framework: Guide, Verify, Solve. He warns that without it, the initial 3–5x productivity boost from agents dissipates within three months as technical debt—security issues, maintainability issues, and complexity—compounds. Sonar's tests show that multi-layered, zero-trust verification (combining algorithmic and agentic methods) reduces AI-derived production outages by 44% and, in a trial with a large bank, achieved a 92% reduction in issues. Clean codebases, he argues, make agents faster and cheaper, with over 30% fewer tokens consumed when guided by context and constraints. The framework embeds verification into three loops—agentic, CI, and code maintenance—to create a self-reinforcing positive cycle, countering the downward spiral of neglected code quality.

Why Your Agent Disagrees With Itself (And What To Do About It) - Diane Lin, Datadog
Jul 20, 2026 · 25:38
Diane Lin, Tech Lead at Datadog, argues that AI agent inconsistency is not a model failure but a signal of ambiguous data near the decision boundary, known as the gray zone. She presents a workflow combining active learning with semantic memory (domain policies) and episodic memory (past similar cases) to automatically identify flip-flopping outputs, focus human review, and continuously adapt agents without expensive fine-tuning. In a real experiment with 93 cybersecurity alerts, 25% initially flip-flopped; episodic memory reduced that to 10%, with the remainder resolved via human review and policy clarification. Lin emphasizes treating each disagreement as an opportunity to clarify labels and policies, building trustworthy, customer-adaptive agents.

Don't Let the LLM Drive - Ornella Bahidika & Joel Allou, Microsoft
Jul 20, 2026 · 6:08
Ornella Bahidika and Joel Allou of Microsoft present their voice tutor Ace, arguing that the LLM should never control the flow in multi-step agents. They built Ace with a state machine that confines the model to narrow contracts per step, letting the harness validate outputs and decide next actions. This design allows them to use a cheaper, faster model (Haiku 4.5) instead of a heavy reasoning model (Opus 4.7 Cloud). They identify three decisions the LLM must never own: when the lesson ends, whether the student answered correctly, and what comes next. By engineering these checks outside the model, Ace achieves reliability in production, avoiding loops and early terminations that plague prompted-only approaches. The pattern applies to any flow agent, including coding agents, runbooks, and onboarding flows.

Your Voice Agent Doesn't Need a Frontier Model - Joel Allou & Ornella Bahidika, Microsoft
Jul 20, 2026 · 5:45
Joel Allou and Ornella Bahidika from Microsoft present ACE, an AI voice tutor that deliberately uses a small model instead of a frontier model, arguing that latency—not intelligence—is the critical constraint in voice applications. They demonstrate that a frontier model's reasoning time of multiple seconds breaks conversational flow, while their system keeps model latency under 950 milliseconds by extracting all logic and planning into a deterministic state machine. This scaffolding handles lesson progression, student mastery tracking, and response generation, leaving the small model (Haiku 4.5) with only the task of speaking. The result: a 900-millisecond response time that feels instantaneous. They acknowledge that scaffolding requires strict rules to prevent drift, but call it a one-time code investment that unlocks cost-effective, real-time performance. The episode’s core lesson: pick the fastest model your latency budget allows, then invest in external scaffolding to make it smart.

Enterprise Agents Have a Structure Problem - Ishita Daga, Tesla
Jul 20, 2026 · 12:08
Ishita Daga, a senior machine learning engineer at Tesla, argues that enterprise agents fail because of three structural problems — ambiguity, staleness, and preference — rather than needing bigger models or more RAG. For ambiguity, she proposes a hierarchy of sources of truth: a curated semantic layer (best for known KPIs), canonical tables (parametric queries for flexibility), and a database graph (full schema but hard to maintain). To solve staleness, she recommends a context lifecycle embedding live data sources (GitHub, CRM, semantic layers) and a feedback loop that logs events, evaluates agent performance, and updates context automatically. On preference, she notes that different teams calculate the same metric differently (e.g., average milestone time by start vs. completion) and that current solutions like semantic layers or agent memory still fail to capture user-level routing, calling this an open problem requiring further research.

When Agents Meet Physical Data: The Other Physics of Agent Harnesses - Dmitry Petrov, DataChain
Jul 20, 2026 · 27:33
Dmitry Petrov of DataChain argues that AI agents fail on unstructured physical data because their intuitions assume cheap recompute, while large-scale video, sensor, and robot data requires a different harness. He cites Anthropic's finding that agents achieve only 21% accuracy on data projects without specific data harnesses, and OpenAI's need for six layers of context even on structured data. Petrov demonstrates DataChain, an open-source Python framework that uses Pydantic schemas to turn messy binary files into queryable databases, an execution engine for distributed processing, incremental checkpoints to avoid recomputing on failure, and a knowledge base of datasets and source code so agents can answer follow-up questions in seconds instead of reprocessing terabytes. In a live demo with Claude Code analyzing 90 dashcam videos, the harness took 24 minutes to extract 100,000 object records, then instantly answered 'how many clips have people?' without rerunning inference. Petrov emphasizes that the key is organizing metadata into star schemas and sharing data lineage across teammates so no one pays the compute cost twice.

Build the AI GTM Agent That Knows the Buyer - Dr. Sajjan Kanukolanu, Position2 (Position Squared)
Jul 20, 2026 · 26:27
Position2's Dr. Sajjan Kanukolanu presents an AI GTM architecture that knows buyers before they message, solving three problems—AI, integration, and architecture—rather than bolting AI onto legacy stacks. The system uses three layers: signals (CRM, enrichment, LinkedIn), buyer intelligence (knowledge base, ICP scoring, context graph), and action (personalized chat, rep alerts, CRM updates, sequence triggers). A demo shows anonymous visitors deanonymized into a dashboard tracking 3,000+ visitors from 280 accounts, with LinkedIn intelligence capturing engagement across 8 posts from 100 visitors. Kanukolanu warns of ICP drift (retrain quarterly), alert fatigue (use context graph to cut noise), identity ceiling (~70% company, 15-20% individual accuracy), and human bottleneck (keep email editing under 30 seconds). The four takeaways: start with identity, score fit and intent separately, build an auditable policy engine, and let every send and deal compound into the knowledge base.

Agent Output Is Not UX: Rendering Layer Your LLM Pipeline Is Missing - Bala Ramdoss, Amazon Lens
Jul 20, 2026 · 14:13
Bala Ramdoss (Amazon) argues that the UI delivery layer—generative UI—determines whether agentic AI features ship successfully, not the model itself. He presents three patterns from building Amazon Lens at scale: (1) a typed, versioned rendering contract where the model selects from a fixed catalog of components (e.g., flight carousel vs. list) and the client falls back safely for unknown types; (2) streaming UI chunks to show skeletons then fill progressively, shifting focus from total latency to time-to-first-chunk; and (3) a Backend-for-Frontend (BFF) that absorbs model output, hydrates actions and conversation context, and reuses existing native components so the AI feels like a natural part of the app. The episode demonstrates that mobile apps, which cannot be patched instantly, require this architecture to avoid crashes and deliver a snappy, trustworthy experience.

Designing Voice Agents for Real Conversations - Chintan Agrawal & Daniel Wirjo, AWS
Jul 20, 2026 · 32:57
AWS Solutions Architects Chintan Agrawal and Daniel Wirjo argue that the hardest problem in production voice agents is audio engineering—not AI—specifically turn-taking, the decision of when an agent should stop speaking or start responding. They present three levels of turn detection: Level 1 uses Silero VAD with a silence timeout (e.g., 300ms default), Level 2 delegates to STT providers like Cartesia or Deepgram for built-in endpointing (p50 ~250-300ms), and Level 3 combines Silero VAD with SmartTurn, an open-source 8MB model achieving 58.9% recall and 68.4% precision while falling back to VAD on low confidence. They show how interruption handling (barge-in) flushes TTS/LLM in ~15ms and distinguish real interruptions from backchannel acknowledgments. Latency budgets are tight: 40ms mic encoding, 52ms network/jitter, 300ms STT+endpointing, 500-650ms LLM time-to-first-byte (dominant bottleneck), and 120-190ms TTS playback, totaling 800-1300ms in standard cloud setups. Co-locating models in one GPU cluster can achieve ~500ms voice-to-voice. For LLMs, Nemotron 3 Ultra and GPT-4.1 achieve ~530ms p50 but GPT-4.1 spikes to 1.7s p95, and multi-turn drift (>15 turns) can break prompt…

Skills are the New SDKs - Elvin Aghammadzada, DataRobot
Jul 20, 2026 · 26:40
Elvin Aghammadzada of DataRobot argues that enterprise AI platforms must become 'teachable' through a skill layer—versioned, task-specific packages that encode operational knowledge for coding agents. He contends that context windows degrade after 25% usage (a phenomenon called 'context rot'), making progressive disclosure via skills critical for maintaining agent performance. Skills expose only metadata (<100 tokens) until activated, while MCP servers handle heavy resource isolation; skills can also self-modify and spawn MCP servers. With 26+ platforms (Claude Code, Codex, Copilot) supporting skills and 85,000+ published, the ecosystem is creating a 'fluency moat' where platform value compounds with each skill. However, LLM-generated skills currently hurt performance, and marketplaces lack verification controls, echoing early NPM risks. The episode positions skills as complementing MCPs, not replacing them, and urges versioning and testing as software.

Medic for Apache Spark - First Aid for Failing Jobs - Drasko Profirovic, Pinterest
Jul 20, 2026 · 11:21
Drasko Profirovic, a Staff Engineer at Pinterest, presents Medic for Apache Spark, an agentic diagnostics tool that automatically troubleshoots Spark job failures by ingesting logs, correlating context, and producing human-quality diagnoses and actionable recommendations in minutes. The talk covers the journey from a single React agent with unsustainable prompt tuning to a multi-agent architecture built on LangGraph's deep agent library, where specialized agents handle triage, research, and healing. Key improvements include an exception classifier pipeline that filters benign exceptions from logs, converting raw time-series metrics into annotated graphs for token efficiency, and an end-to-end test harness that snapshots production state for offline evaluations. Profirovic shares lessons on handling ambiguity, reducing hallucinations, and balancing automation with human oversight, plus unexpected failure modes that informed iterations. The system now extends to optimizing Spark SQL…

Can Oncology Workflows Run Without Human Touch? - Anant Shankhdhar, Risa Labs
Jul 20, 2026 · 16:41
Anant Shankhdhar, an AI engineer at Risa, explains how his team automates oncology workflows end-to-end using four AI agents—EV, Auth, Necessity, and Submission—to eliminate human touch in prior authorization processes. The EV Agent handles eligibility and benefits verification via a unified service that connects to payer APIs and RPA portals, using LLM-driven config generation and self-healing loops to scale. The Auth Agent determines drug authorization status by reconciling evidence from patient notes, authorization letters, and a payer rule knowledge base, enabling no-touch handling for drugs that are already authorized or don't require authorization. The Medical Necessity Agent answers clinical questions per patient, attaching confidence scores and escalating only cases needing human review. The Submission Agent submits orders to payers using customized integrations. Risa's agents are deployed across 20+ hospitals, supporting care for over 100,000 patients, and the no-touch share…

From Tokens to Cells: Foundation Models for Single-Cell Biology - Akram Baharlouei, Altos Labs
Jul 19, 2026 · 16:57
Akram Baharlouei, a machine learning engineer at Altos Labs, explains the engineering challenges of building foundation models for single-cell biology, arguing that flow matching models like PrimeFlow outperform transformer-based models for this noisy, heterogeneous data. He highlights the Yamanaka factor, discovered in 2006, which reprogrammed aged skin cells to an embryonic-like state, earning a Nobel Prize in 2012 and enabling possibilities for cellular rejuvenation medicine. Baharlouei notes that drug development takes up to 10 years with billions in cost, and AI could shorten the pipeline. He describes RNA-seq as the primary modality for training, with datasets reaching 1 billion cells, but warns that scaling alone fails without quality improvements. Benchmarking at NURIPS showed transformer models often underperform simple linear models, while flow matching models better capture the distribution of cell states, as measured by MMD scores.

You Didn't Ship a Bug. You Just Wrote It for a Human. - Ravi Madabhushi, Scalekit
Jul 19, 2026 · 12:50
Ravi Madabhushi, co-founder of Scalekit, explains why infrastructure built for humans breaks when AI agents act 60 times faster than users, citing a production database spike caused by a 'last seen' timestamp updating every tool call. He argues that OAuth scopes, designed for deterministic human-written programs, cannot constrain non-deterministic agents, which need attribute-level, time-bound, and principle-scoped permissions. Madabhushi warns that 60% of LLM errors stem from rate limits designed for humans, not agents, and advocates for just-in-time authorization and absolute visibility into every agent action—who took it, on behalf of whom, and when authorized. He concludes that without deterministic guardrails, teams are merely 'praying' agents won't delete production data.

From Blind Spots to Merged PRs: Continuous Agentic Performance Optimization - May Walter, Hud
Jul 19, 2026 · 22:46
May Walter, CTO of Hud, presents a case study on using coding agents with runtime intelligence to automate continuous performance optimization in production. The system analyzes real production traces, queries, and latencies to surface high-ROI fixes like N+1 queries and missing indexes, scoring them by complexity and impact. It runs weekly via GitHub Actions and Claude Code, generating reports with verified fixes and evidence of impact, such as P90 latency improvements after deployment. Key challenges included 'plausible unverified' fixes, complex ClickHouse queries, and lazy fixes that only mask exceptions rather than root causes. Walter emphasizes the need for prod-to-code context mapping, scoring guardrails, and human review to achieve reliable automation. The talk concludes that context over cleverness works, and agents can automate the investigation phase that teams rarely do proactively, turning dormant performance debt into actionable merged PRs.

Build Evals That Actually Matter - Nick Ung & Akshay Sharma, Lyft
Jul 19, 2026 · 37:45
Nick Ung and Akshay from Lyft explain how to build evals that actually matter for customer-support AI agents, sharing their end-to-end pipeline. They argue offline evals must use a fine-tuned user simulator trained on real rider and driver transcripts to produce messy, frustrated behavior—not an off-the-shelf LLM that sounds too nice. Their LLM judge is treated as a binary classifier, calibrated against human labels using precision and recall, with metrics tied to actionable business outcomes like escalation decisions. They detail a continuous error-analysis loop that feeds production failures back into the offline test set and closes the loop through context, harness, or model learning. The talk also previews their config-driven eval harness with YAML primitives for tasks, datasets, and personas, enabling repeatable runs across development stages.

Autonomous Agents for Scientific Tasks - Sina Shahandeh, Radicait
Jul 18, 2026 · 19:23
Sina Shahandeh (Radicait) presents an approach to improve autonomous agents for scientific tasks by focusing on hypothesis generation rather than just implementation. He argues that standard coding agents saturate because they run out of ideas, and introduces a hierarchical decomposition method that breaks problems into subcomponents (e.g., data, architecture, training loss) and prompts LLMs to generate radical changes—like switching from 2.5D to 3D convolutions for CT-to-PET translation. He also shows how adversarial/collaborative loops with multimodal models (e.g., Gemini) can critique intermediate results, and uses Peter Steinberger's Oracle CLI to invoke GPT-5 Pro for better reasoning. The talk demonstrates these techniques on real-world medical imaging tasks, including image registration and lung nodule analysis, and identifies multimodal model limitations as a key bottleneck for fully autonomous scientific discovery.

A Practitioner's Guide to Graphs - Tim Ainge, Good Collective
Jul 18, 2026 · 14:18
Tim Ainge from Good Collective presents a practitioner's guide to graphs, covering extraction from unstructured text, schema-first design, and graph-native algorithms to make AI applications smarter, cheaper, and more reliable. He demonstrates that giving extractors a schema (e.g., recipe with ingredients and steps) yields more meaningful graphs, and using ontology instructions standardizes units and ingredient names. Embedding models solve the potato–potato problem by flexibly matching duplicate nodes. Personalized PageRank, inspired by Pinterest's Pixie paper and HIPORAG, finds authoritative nodes in dense graphs, e.g., identifying Miranda v. Arizona as a landmark case not directly cited. Shortest path algorithms reduce tool calls by 40% in code search by retrieving intermediate nodes missed by vector search. Subgraph matching detects software design patterns like the decorator by querying graph shape without specific node details.

The UX of AI: Making AI-Powered Apps Your Users Don't Hate - Kathryn Grayson Nanz, Progress Software
Jul 18, 2026 · 35:59
Kathryn Grayson Nanz, Senior Design and Developer Advocate at Progress Software, argues that the success of AI-powered applications now depends on user experience rather than model performance. She identifies five pillars—trust, clarity, control, transparency, and meaningful benefit—and provides concrete patterns: citing sources to build trust, streaming text for clarity, allowing undo and version history for control, requesting granular permissions for transparency, and offering templates and next-step actions to ensure meaningful benefit. She emphasizes that developers must design these patterns themselves because AI can only remix existing interfaces, and that users need gradual introduction to AI features to avoid disengagement. The talk stresses that without addressing these UX challenges, users will abandon AI tools after a few failed attempts.

Stop Burning Tokens: Why self-improvement needs domain expertise first - Annabell Schäfer, Langfuse
Jul 18, 2026 · 17:39
Annabell Schäfer, Growth Engineer at Langfuse, argues that successful auto-improvement loops require domain expertise and high-signal target functions, not just generic evaluators. She details an experiment classifying arXiv papers with a minimal loop using GPT-5 for nano and Claude Opus 4.8 as an optimizer, achieving a 15% accuracy jump from 68% to 83% in four iterations. The first iteration alone gained 10% by adding classification rules and examples based on error analysis of a 200-item dataset. Schäfer advises replacing vague metrics like correctness with yes/no quality criteria (e.g., 'answer uses knowledge base'), and working with domain experts to identify failure modes and define what 'good' means. She emphasizes validation to prevent overfitting and treating the system as generalizing from representative examples, not just burning tokens on endless loops.

Stop Renting Your Cognitive Infrastructure - Thiyagarajan Maruthavanan, Kalmantic Labs
Jul 18, 2026 · 7:52
Thiyagarajan Maruthavanan of Kalmantic Labs argues that AI teams should stop renting inference from providers like Anthropic and build their own infrastructure, coining "rent to learn, own to earn." He recounts his app UltraSuno costing hundreds of thousands of dollars in inference, a stolen key hitting $10,000, and moving to his own DGX Sparks hardware. Three enterprises—a fund, hospital, and tax practice—each hit walls with renting: control, audit redlining, and reproducibility. He open-sourced JustTokenMax (benchmarked better than Netflix's Headroom) and wrote the book "Peak Inference" on building your own inference infra. He notes the market's conflicting pitches (Jensen's token factory, Nadella's unmetered, NeoClouds' endpoints) but concludes owning is essential post-PMF.

Content Is Code - Matt Palmer, Conductor
Jul 18, 2026 · 10:53
Matt Palmer, head of developer experience at Conductor, argues that code is now the fastest way to produce technical content, shifting left to code as the source of truth. He outlines three eras of content creation—handcrafted, expensive code, and cheap code with AI—and claims that the scarce resource today is not code but structure and conscientiousness, meaning meticulous care in maintaining design tokens, brand guidelines, and clean codebases. AI rewards organizational excellence over raw technical skill, making the best communicating teams those that instill discipline and rigor into their software creation. Palmer predicts 2026 as the year of the creative technologist and 2027 as the year of the content engineer, with declarative pipelines that automatically generate documentation, walkthroughs, product tours, and updates from structured code.
Powered by PodHood