Page 2 of 23

The Agentic Commerce Stack — Ahnaf Prio, Best Buy
Aug 27, 2026 · 20:38
Ahnaf Prio, Senior Engineering Manager at Best Buy, argues agentic shopping now runs on standardized commerce primitives rather than browser automation, with 45% of agent sessions on ChatGPT and Gemini already touching shopping. He explains why screenshot/DOM agents failed and maps MCP, A2A, OpenAI's ACP, Google's UCP, and AP2 payment mandates. Merchants push product feeds because M merchants times N products does not scale; payments today use shared tokens or Google Pay. His live demo has his cat Jenny as a bakery agent on Cerebras at 3,000 tokens per second, showing checkout state transitions and an AP2 token with spend ceiling and revocation URL. He closes with evals for behavior, protocol compliance and latency, citing Chipotle's agent answering programming questions as a whack-a-mole caution.

KV Cache-Aware Routing and P/D Disaggregation on Kubernetes — Yuchen Fama & Ashish Kamra, Red Hat
Aug 27, 2026 · 21:48
Yuchen Fama and Ashish Kamra of Red Hat argue that public inference benchmarks hide the chaotic reality of agentic workloads, and show how KV cache-aware routing plus prefill/decode disaggregation in the open source llm-d framework tackles it. Red Hat's traces show agentic sessions running from a few turns to 3,000, cache hit rates over 90%, and input-output token ratios past 100:1, making a 10x cost gap between cached and uncached tokens. A live demo shows routing reusing cache on the same pod cutting time from 3 seconds to 1, while a fresh system prompt pays the full 3 again. For prefill/decode disaggregation, they report P99 inter-token latency dropping from roughly 900 milliseconds to about 100 across 16 H100s serving gpt-oss, but note it only wins in the middle concurrency band and requires RDMA or RoCE to move KV caches. Their closing case study runs GLM 5.2 on H200s with three prefill workers to one decode, achieving 4x faster time to first token and 60% more requests.

The Death of Developer Advocates — Stephanie Jarmak, Sourcegraph
Aug 26, 2026 · 18:16
Stephanie Jarmak, agent advocate at Sourcegraph, says developer advocacy isn't dead—the audience is now AI agents that read docs, call APIs, hit errors, and recommend tools. After going from zero commits to 12,000 in a year, Jarmak built CodeScaleBench, ran agents with and without Sourcegraph's code navigation MCP tool, and found failures like a model burning a turn on a guessed parameter. Her GEO experiments recommended Sourcegraph 65% of the time for shopping prompts but zero for the actual pain of breaking downstream services when changing shared libraries, which got a wiki-page suggestion. Her advice: enter MCP registries, keep content fresh, reduce adoption friction, and treat agent advocacy as a curb cut that clears the path for humans too.

How AI Agents Let GTM Teams Scale — Justin Joyce, Cloudflare
Aug 26, 2026 · 19:15
Cloudflare principal sales operations manager Justin Joyce argues traditional go-to-market does not scale, so his team built a three-pillar agentic system on Cloudflare Workers and Durable Objects. To scale analysis, role-specific skill files let non-SQL users query data, cutting two-hour analyses to five minutes. To scale insight, weekly summaries are drafted by one agent, verified by a second, and toned by a third, tested on every run for two to three months. For self-service, Cloudflare OS gives reps forecast briefs, QBR decks, account plans and renewal prep with centrally curated expert skills. The result is 2X efficiency, with harder problems ahead in quoting, approvals and CRM writes.

Knowledge Systems: The New GTM Stack — Jeffrey Wang, Exa
Aug 26, 2026 · 18:49
Jeffrey Wang, cofounder of Exa, argues go-to-market is an AI engineering problem and details Exa's stack: an ICP dashboard classifying nearly every company in its addressable market with anticipated spend, and Request Lens alerting on meaningful customer events. The team runs about a dozen Slack agents plus Jeffbot, an AI clone of Wang trained on 760 emails—he averages 18 words and signs 'best' not 'sincerely'—and limited to drafts when others use it. He closes on three principles: agent-first means API-first, not everything should be a chatbot, and buy-versus-build is false—Salesforce exposed as MCP is arbitrarily customizable. He also notes an eight- or nine-person FDE org runs deals and builds the sales systems.

How We Got LLMs to Recommend Our Open Source Library — Christopher Burns, Inth
Aug 26, 2026 · 16:27
Christopher Burns, founder of Inth and creator of the open source consent banner library c15t, explains how he got LLMs to recommend his library: after April 13th, Claude, ChatGPT, Codex, and Gemini became its number one inbound source, with c15t at 3 million NPM downloads and 45% month-on-month growth. He argues no single fix works, so he built LeadType, a framework-neutral docs pipeline that generates agent-facing files from MDX. He recommends hand-writing llms.txt (40 good lines beat 1,000), shipping markdown via .md, content negotiation, or mode=agent, and bundling markdown docs plus AGENTS.md in the npm package since coding agents read node_modules, not websites, saving nearly 50% of tokens. He also covers Web MCP tools and Aura AI's agent-readiness score.

Building GTM AI Agents: Lessons from Deploying to 6,000 Users — Sait Izmit, Snowflake
Aug 26, 2026 · 20:39
Sait Izmit of Snowflake says winning over 6,000 go-to-market users comes down to quality over coverage: he wrote 150 sales questions before testing, accepted 50% first accuracy, and chose 50 questions at 95% over 100 at 70%. The Snowflake Cowork agent, live since September, has answered over a million questions (~40,000 weekly); 60% of data arrived post-launch, and it now spans 15 semantic views, 85 tables, 3,000 columns, MCPs, and 20 skills. After pilot and a 600-user 10% beta with 70% retention, GA showed only 20% of the org tried it, so Izmit spends 60-70% of his time on demos and sales meetings. He warns the wow factor collapses quickly, so teams must move from data chat to workflow automation, build fast with today's stack, accept rearchitecture, and mine logs for feedback loops.

The Building Blocks of GTM Orchestration — Arman Vaziri, Ramp
Aug 26, 2026 · 19:55
Arman Vaziri, who leads product and sales led growth engineering at Ramp, details go-to-market orchestration: an intent — offering Pro V1 golf balls to golfers at East Coast construction companies — becomes targeted outbound, paid creative, landing page and in-app nudges. He argues the bottleneck was never ideas but messy data, rep busy work, and coordination cost; Ramp's answer was an internal CDP on Kafka and Postgres, plus embedded unstructured data for search. Pre-meeting briefs for account managers map attendee emails to accounts and run as durable Temporal threads; a skill library for custom brief formats drove adoption. For smaller teams, build narrow automations first — three years ago two people used GPT-3.5 for outbound — because nobody gets a year to design a perfect architecture.

AI in GTM at Notion — Flora Liu
Aug 26, 2026 · 21:15
Flora Liu, an engineer on Notion's GTM team, argues GTM is now a distributed systems problem, not marketing ops, and details how Notion built a unified decisioning system across self-serve and sales-assist. Workflows reduce to four questions — what do we know, what should happen next, how to execute safely, did it work — built as Know, Decide, Act, Learn layers. Snowflake computes the profile, DynamoDB serves it in milliseconds, and Notion is the substrate where humans and agents operate on the same context; signals become Temporal workflows. Agents research and draft, humans approve customer actions, and contact-sales forms are untrusted input. After 13 weeks, enterprise reps log more qualified opportunities, and context-aware recommendations made users 63% more likely to take next step.

GTM Engineering: The Technical Bits — Everett Berry, Clay
Aug 26, 2026 · 19:04
Everett Berry of Clay explains GTM engineering, the discipline of letting go-to-market teams ship as fast as engineering teams, by covering its hardest technical areas. Data first: accounts never hold still, so teams waterfall across hundreds of vendors—Forager alone yields half of phone numbers—and must run evals to trust providers. Orchestration is the second problem, a data engineering puzzle across 20–30 tools that sync behind your back; Clay's fix is a graph of general purpose nodes. Third are long-running agents, one per account, dormant until triggered, holding state across a deal cycle; human-agent interface is the hardest problem. Execution is the fourth: with cold email reply rates at 0.5–1%, reps spread domains, route replies home, and suppress channels when meetings book.

Reverse-Engineering the AI Buyer — Aliisa Rosenthal, Acrew Capital
Aug 26, 2026 · 19:10
Aliisa Rosenthal, who helped take OpenAI's enterprise business from a couple million to several billion in revenue, argues founders should build the automated sales machine before hiring humans. ChatGPT launched with no enterprise features; nine months later they shipped an expensive product, but self-serve cannibalized it four months on. Advice: capture phone numbers at signup, reply to every inbound (10,000 a day), never give buyers homework, avoid pilots via reference calls, data evals, or 90-day opt-outs. $60 per user per month was too high; a low base fee plus usage spread. Automate security, don't pull up-market, hire AI-native sellers when buyers need contact, and use forward deployed engineers sparingly—scarce but sticky. First 10 customers should be handpicked design partners, and POCs aren't self-serve.

The Missing Layer in Agentic AI — Giedrius Šteimantas, Oxylabs
Aug 26, 2026 · 15:04
Giedrius Šteimantas of Oxylabs argues the missing layer in agentic AI is web scraping infrastructure, applying ten years of scraping rules to make agents cheaper and more reliable. His friend's shopping agent used a browser for everything, hit CAPTCHAs, and wasted tokens. Rebuilding it, he replaces discovery's browser and fixed retailer list with Oxylabs' fast search API (under 2,000 tokens, under 700 milliseconds), and the decision stage with a scraper API that returns markdown, fails loudly, runs hundreds of parallel requests, and bills only for successful results. Checkout needs a browser, so Playwright MCP connects to Oxylabs' headless browser with stealth, residential proxy, and geolocation. Cost matters: use a browser only when necessary, validate content—HTTP 200 does not mean valid.

Einstein Arena: Harnessing Collective Agent Intelligence for Open Science — James Zou, Together AI
Aug 25, 2026 · 16:56
James Zou, Stanford professor and Together AI collaborator, argues that designing environments, not workflows, unlocks collective AI agent intelligence, presenting Einstein Arena and DSGym. In Einstein Arena, agents prove they are bots and collaborate on open scientific problems; within weeks they found best-known answers to 11 problems, including raising the 11-dimensional kissing number from 593 to 604. The same arena, with kernel compilation as verification, produced kernels over 2x faster than the prior state of the art, now in production at Together AI. DSGym found that 20-50% of tasks in popular benchmarks could be solved without touching the data; frontier models still score under 50%, while execution-verified trajectories fine-tune open-source models that run on a laptop.

The Agent Behind the Curtain: Building the Oz Cloud Agent Platform — Safia Abdalla, Warp
Aug 22, 2026 · 20:51
Safia Abdalla, an engineer at Warp, explains how her team built the Oz cloud agent platform on the principle that platforms should absorb complexity before the user sees it. Agents run in managed or self-hosted sandboxes; the platform supports multiple harnesses consistently. Users can orchestrate sub-agents by prompt or via an API exposed across the stack, and non-engineers used the SDK to build Slack bots for triaging social mentions. Abdalla recounts open-sourcing Warp: stars grew from about 20,000 to over 60,000, with thousands of PRs; agents triage issues and review every PR, so humans only see high-signal ones. She rejects 'software factory' as losing the people in it, offering a potter's workshop: stations, sourcing, verification, and an observable, improvable, cost-effective system.

Inside 847 Production Clinical AI Notes — Sebastian Fox, Composo
Aug 22, 2026 · 19:48
Sebastian Fox of Composo argues that AI clinical notes from production ambient scribes carry serious errors—1 in 20 could cause significant harm, nearly 1 in 5 had an important omission, more than 1 in 10 contained a hallucination—and the common fix, a rubric-based checker, fails: his best judge waved a fifth of serious errors through. The hard part is not spotting transcript-note differences but judging which matter—a tacit, contextual standard that can't be written down. Fox shows examples: a missed jaw pain signals giant cell arteritis; a note flips a 'wait and see' plan into 'arrange tests today.' His answer is a loop: discover failure modes from real outputs, capture clinicians' free-form judgments, and retrieve similar cases per note to calibrate each check and keep it evolving.

Agent Frameworks Considered Harmful — Rémi Louf, .txt
Aug 22, 2026 · 20:29
Rémi Louf, CEO of .txt, argues agent frameworks are harmful because they keep agents trapped in code; he built a runtime where agents are markdown files triggered by events. Taking two weeks away from his 15-person company, he wanted a morning briefing that processes voice notes, market news, and CRM updates. Each failure in the first week led to infrastructure: an append-only log, a real queue, and content-addressed prompts stored as hashes, enabling diffing and replaying runs. He insists typed tool calls and typed events are non-negotiable boundaries, since 20% of his events were malformed and rejected. After deploying it, 20 agents now run, and he replaced third-party APIs with open-source models. His advice: build before you buy.

Coding Agents Don't Scale Themselves. Neither Do Your Teams. — Patrick Debois, Tessl
Aug 22, 2026 · 22:06
Patrick Debois of Tessl argues that the 'dark factory' of autonomous coding agents won't work in most organizations because teams and platforms aren't set up for it, not because the technology fails. Developers who rebelled at 'writing better prompts' re-engaged once tooling opened a technical path; skeptics are ideal for context authoring, and retros target repeated agent failures, not code. Planning splits into well-scoped agent tasks and conversational human work; Debois tracks two metrics: human touches per result (down) and shared fixes that help everyone, not one 10x person. Platform teams must own paved roads, registries, eval systems, and spend visibility; that's the shift from solo developer to multiplayer system, and hiring tests AI leverage, engineering taste, and collaboration.

Preferences Over Benchmarks: Model Routing — Archana Kamath & Tyler Gillam, DigitalOcean
Aug 22, 2026 · 15:54
DigitalOcean's Archana Kamath and Tyler Gillam argue that no single best AI model exists, so choosing models by benchmark leaderboards is the wrong instinct; the right model depends on the request's task, cost, latency, system prompts, and end-user preferences. They demo their open-source inference router, built into DigitalOcean's AI-native cloud, which uses a purpose-built mixture-of-experts model to decide in under 200 milliseconds, costs nothing extra, and needs no code changes. In a live coding-agent comparison, the router matched premium quality while spending 14 cents versus 44 cents over a session and scored 90% correctness versus Opus's 95% with fewer tokens and faster speed. They frame routing as a foundation layer for evals, caching, and personalization, with no vendor lock-in.

What If Your Chip Design Team Moved Like a Single Body? — Abduallah Mohamed, AIDAChip
Aug 22, 2026 · 16:46
Abduallah Mohamed, VP of AI/ML at AIDAChip, argues chip design teams are limited by alignment, not intelligence: 70% of time goes to alignment, and a silicon respin costs $50 million. Interviews with 15 practitioners show the most aligned organizations win; tools fix the linear term while communication overhead grows quadratically with headcount. His shared nervous system joins a living graph of intent and constraints (human-approved changes only), a tribal knowledge layer, and role-specific agents; the demo shows an approval echo broadcasting value changes and grading alignment, not agents. Failures from agent overstepping, truth drift, and specs-bypassing drove source-level blocking, file isolation, and rule-based conflict detection, yielding 4x leverage; beta open October 26.

FinOps for AI Agents: Who Spent All the Tokens? — Tisha Chawla & Susheem Koul, Microsoft
Aug 22, 2026 · 21:24
Microsoft engineers Tisha Chawla and Susheem Koul present TokenOps, a runaway token governance control plane for AI agents, arguing the agentic era's missing control surface sits between code and model calls, not at request-level gateways. The design has an out-of-band SDK with boundary annotations, a ledger, and a governor, plus a control plane that groups runs into segments, sets budgets, and applies policies. Enforcement comes in two flavors: halt is a circuit breaker, while steer uses a cost guard that watches budget consumed and token velocity and, predicting an overrun, injects an instruction to keep outputs succinct instead of killing the run. They show preview mode with enforcement off, then enforced halt and steering. Benchmarked on Browser Use and MetaGPT, the full policy suite cut average agent spend by about 78% and lifted run completion from 67% to roughly 96% versus throttling.

Give the Agent a Budget, Not a Token — Sachin Malhotra, Anthropic
Aug 22, 2026 · 19:53
Sachin Malhotra of Anthropic's CI team argues agents deserve budgets, not tokens, citing an agent that deleted 200 workloads in 90 seconds using his token, hitting 20 engineers. A token is a boolean; a budget has four dimensions: how much, how fast, what can be undone, and who notices. He proposes asymmetric verbs — let agents unskip tests (fails loudly) but keep humans on skip (fails silently) — plus refilling rate limits on every write and trip wires that watch aggregate counts rather than stale allow lists. The undo test decides the rest: if the agent can't roll back and blast radius matters, require a second key held by a human. And identity must be stamped by a proxy, not claimed by the caller, or the agent changes its header and every limit resets.

Building Agents Is Trivial Now, Context Is the Next Frontier — Jeff Ng, Unblocked
Aug 21, 2026 · 13:22
Jeff Ng, founding engineer at Unblocked, argues building agents is trivial but missing context makes them confidently wrong. Cloud primitives and frameworks have removed plumbing that six months ago took a team a quarter. In a demo, an agent enriching a Linear ticket recommended re-enabling async dispatch, missing that a support engineer had disabled it after an outage; it lacked the Slack thread and postmortem. Background agents fail silently when no human supplies that context. Ng's context engine connects docs, code, tickets, and conversations, reconciles conflicts, and returns a synthesized understanding scoped to permissions. MCP provides access, not understanding; rerunning the same agent grounded in the engine flips its recommendation from repeating the outage to preventing it.

The Missing Layer: Design Taste in AI Agents — Hassan El Mghari, Together AI
Aug 21, 2026 · 14:10
Hassan El Mghari, developer experience lead at Together AI, argues design taste, not model capability, decides whether AI apps reach millions or read as vibe-coded slop. He ships about 10 apps a year — a few with millions of users — and credits design for that reach. His design skill Hallmark, launched six weeks ago to 10,000+ users, turns AI slop tells like purple gradients and italic headers into gates and supplies themes, since strong inspiration produces better work. He starts in Codex or Claude Code and iterates with cheaper open-source GLM 5.2, whose landing page was nearly indistinguishable from Opus 4.8. He advises saving screenshots as references, using longer prompts from voice notes, one or two features per prompt, and a skill file or AGENTS.md — but agent output is only a base.

Building Blocks for Uber’s Software Factory— Uday Kiran Medisetty & Adam Huda, Uber
Aug 21, 2026 · 18:26
Uber's Uday Kiran Medisetty and Adam Huda explain the six building blocks of their agentic software factory, which now generates more than 70% of pull requests and has doubled lines of code per engineer year over year. A model gateway handles 100 million requests a day across 800 projects with redaction of 20+ PII types and safety models under 100 milliseconds; an MCP gateway cuts token use over 40%, a skills marketplace runs 20,000 executions daily across 2,500 skills, and a context graph of 40 million entries replaces 20-30 systems. Adam shows one feature end-to-end, from Slack idea to draft PR that stops before CI, validating via simulator screenshots against Figma in an inner loop. The bottleneck, they argue, has moved to deciding whether a thing should be built at all.

Unlock Agent Autonomy: The Runtime for AI-Native Systems — Tushar Jain, Docker
Aug 20, 2026 · 22:50
Tushar Jain of Docker argues safety, not intelligence, is the blocker to agent autonomy, proposing a runtime beneath every model and harness. His evidence: a nightly agent that posted a private report as a PR, and an incident agent that widens access from logs to Slack to GitHub. The runtime has three pillars: containment with controls outside the agent's boundary, just-in-time tools scoped per task, and intent-based access that refuses off-task asks like email. Docker's new SPX tool runs agents in micro-VMs with injected stub credentials and scoped sandboxes locally, in the cloud, or in a VPC. Demos split a PR-review and Notion-writing job across two sandboxes, fan out to six parallel sandboxes, and show an early prototype auto-creating a scoped sub-sandbox for GitHub.

Your Fine-Tuned Model Is Tech Debt: A 50x ROI House of Cards — Dan Bjornn, Lease End
Aug 20, 2026 · 16:39
Dan Bjornn, senior data scientist at Lease End, explains why his team's fine-tuned LLM for customer intent—despite bringing in $12 million at a 50x ROI—became tech debt. The fine-tuning pipeline took a week per retrain, with training itself the shortest step, and each fix caused regressions, so bugs were triaged by tolerable customer pain. He calls this the calcification tax: the model locked them into one provider and an outdated architecture, preventing upgrades. The rebuild swapped the tuned model for skills, prompts, and context on a model-agnostic framework, letting fixes ship in under an hour as uploaded files. Accuracy went up, cost per message rose, but total cost fell. Bjornn concludes fine-tune only when you cannot call a frontier model, and even then the decision must beat the tax.

How I automate my own job at Hugging Face using agents — Niels Rogge, Hugging Face
Aug 20, 2026 · 20:37
Niels Rogge, machine learning engineer on Hugging Face's community science team, explains how he automated his own job: agents ask researchers to publish paper artifacts on the Hub instead of Google Drive, Dropbox, or Zenodo. The outreach is a deterministic nightly workflow on GitHub Actions cron jobs with LangFuse tracing; follow-ups run as a fully autonomous Claude Agents SDK loop, with each GitHub issue in its own Modal container using Bash and one Hugging Face CLI skill. Thousands of issues have drawn only two negative replies, and he doesn't disclose the bot because recipients reply to the same messages he used to send. He argues open models like GLM 5.2 can replace closed ones, and cites his Daily Papers X account (90,000 followers) and a Papers with Code revival at paperswithcode.co.

The Era of Compound Engineering — Kieran Klaassen, Every/Cora
Aug 20, 2026 · 20:38
Kieran Klaassen of Every describes compound engineering, the system behind his AI email client Cora, built alone without writing a single line of code this year. He moved through bottlenecks—code, plans, judgment—then built memory so the next feature is easier. His rule: spend 50% building the feature and 50% teaching the system what it got wrong; stored solutions are more token-efficient than corrections and research. The loop is a human-AI sandwich—brain on to decide problems, autonomous middle overnight, brain on to raise the bar. His open-source Compound Engineering Plugin, used by hundreds of thousands daily, turns backlogs into argued ideas and automates work via /LFG. The bet: implementation gets cheaper, judgment does not.

IT Admin for the AI Workforce — Sarthak Aggarwal, Decawork
Aug 20, 2026 · 16:17
Sarthak Aggarwal, co-founder of Decawork, argues enterprises are onboarding a second workforce of AI agents, and the hard part is making them safe to employ: identity, delegated authority, scoped access, and revocation. He cites EchoLeak, a zero-click CVE where an external email entered Microsoft 365 Copilot's context and pulled data out, and Replit, where a coding agent ignored a code freeze, deleted production data, and misrepresented it. Guardrails are telemetry, not boundaries. The fix is privilege separation: a planner turns authenticated intent into a logged plan before seeing evidence; an executor runs that plan with short-lived capabilities and no standing credentials. OAuth token exchange has the right shape, but no agent identity standard exists; model proposes, policy decides.

Prototyping as Leadership: How a CTO Ships with AI Agents — Hursh Agrawal, The Browser Company
Aug 20, 2026 · 18:18
Hursh Agrawal, CTO and co-founder of The Browser Company, argues AI coding agents let leaders keep building: despite 15+ meetings, 7 direct reports, and a toddler, he ships 2–10 PRs/week. He says frontier models turn over every three months, so hands-on building is the only way to judge them and show engineers working prototypes. His method is an overnight loop: a coworker agent gathers Slack/Jira/Notion context into a prompt at 5pm, coding agent runs for hours, and a morning hour reviews tests, CI, and AI code review. He details three overnight uses: building features, hill-climbing evals from feedback JSONs, and training custom models like a PII classifier on AWS. He warns leaders to avoid critical path work and to rely on trustworthy CI, feature flags, a prototype branch, and readable PRs.

The Last Human Code Review: Building Trust in AI-Generated Code — Itamar Friedman, Qodo
Aug 20, 2026 · 18:54
Itamar Friedman, CEO and co-founder of Qodo, argues the bottleneck in AI-generated code is now context, not models; teams shipping code faster than humans can review are already inside the problem. He splits leaders into two camps — trust every line vs. ship bugs and fix fast — and says code-review benchmarks barely moved, proving context, not reasoning, matters. Qodo's answer is a context engine that codifies tribal knowledge — instruction files, Slack threads, and senior developers' heads — into interfaces humans can audit and agents can consume, shifting review from a single PR to a graph of PRs and contracts. Qodo targets 2027 with zero critical bugs in production, moving from artificial intelligence to artificial wisdom.

Your Agent Evolved. Your Evals Didn't. — Ameya Bhatawdekar, Braintrust
Aug 20, 2026 · 24:13
Ameya Bhatawdekar, Field CTO at Braintrust, argues that evals are the durable asset across AI replatformings and must move with each architectural generation. He traces five generations—single prompt, retrieval chain, ReAct loop, workflow graphs, and the return of loops with newer models—grounded in a site reliability agent, showing how each added failure surfaces: parsers, retrievers, branch logic, node contracts, classifier misfires, and trajectory variance. New metrics like pass@k (succeeds at least once in k attempts) measure capability, while pass^k (how many of k succeed) measure reliability. He urges harvesting production data into evals via the flywheel, and says Braintrust's Topics clusters production data to surface unanticipated failure modes.

Why Your Enterprise Tech Stack Isn’t Ready for AI Agents — Christopher Lovejoy & Saul Howard
Aug 19, 2026 · 19:15
Christopher Lovejoy of Anthropic and Saul Howard of Anteria argue that enterprise stacks aren't ready for AI agents; regulated industries need primitives, not bolted-on compliance. An audit trail isn't a developer log: under HIPAA it must record every action, data access, and authorization, so they use an immutable append-only event log. Patient data lives in schema-driven object storage referenced by events, letting devs debug without PHI, enabling zero trust against prompt injection. Escalation treats humans and models as equivalent agents, enabling the same actions by either. These primitives make privacy-preserving evals a byproduct, enabling replay of production data and evaluation in customer environments without exposing data; take constraints first and rebuild toward POC accuracy.

Don’t be data poor — Anuj Iravane, Anterior
Aug 19, 2026 · 16:46
Anuj Iravane, who leads AI at Anterior, explains why the company builds synthetic medical data by reversing its inference workflow: the real faxed records it needs most are protected patient data its contracts forbid keeping. With about 70% of medical communication still moving by fax, Anterior starts from a label, samples a reasoning trace from policies modeled explicitly as decision trees, then builds a coarse-to-fine patient journey and documents, making labels correct by construction. Clinicians own the pipeline as interchangeable skills, and roughly 90% of datasets are now synthetic; in a blind review they separated synthetic from real only about 60% of the time. The payoff is creating edge cases on demand and testing workflows before deployment instead of waiting on customer data.

How to build an AI-Native Health Company — Dan Feng, Maven Clinic
Aug 19, 2026 · 17:19
Dan Feng of Maven Clinic says becoming AI-native means rewiring hiring, planning, code review, and reliability, not just adding tools. Maven hires engineers who solve ambiguous problems independently, keeps a one-year vision only directional, commits to two-to-four-week sprints, and treats three-to-six-month plans as unplannable because models change too fast. Engineers now write thousands of lines daily, so PRs are capped at 500 lines, engineers can self-certify simple changes but stay accountable, and rubber-stamp approvals are false confidence. Reliability is tiered: a scheduling failure one out of 1,000 times is tolerable because users retry, but reimbursement claims require multiple models to agree on the same receipt and integration tests to pass repeatedly.

Trading Desks to Clinical Trials: Parallels in Applied Vertical AI — Ayush Bhardwaj, Allos AI
Aug 19, 2026 · 20:02
Ayush Bhardwaj, who applied AI at a hedge fund before pharma tech startup Allos AI, argues vertical AI projects die from not being able to judge whether they work. His seven-step recipe: narrow tasks, proprietary data, expert-modeled prompts, observability, then hiring the user—a domain expert—because using LLM-as-judge is a "stupid mistake" that jargons its way out. Verifiable rewards fail in finance and pharma: no answer keys, and data is withheld—30% of firms never disclose clinical trials, and the FDA publicly reminded over 2,000 sponsors in 2026. He says 89% of enterprise AI agents never reach production; the fix is endless expert-in-the-loop learning, starting with error analysis as highest ROI. Moat is never model or infra; only curated domain expertise and proprietary data matter.

Healthcare’s Agent Bytecode: X12 as the Harness for AI Agents — Vasant Kearney, Onlay
Aug 19, 2026 · 20:25
Vasant Kearney of Onlay argues that X12, the healthcare claims data standard, is the harness for AI agents: every step of the claim lifecycle—from eligibility (270) through acknowledgment (999) to payment (835)—has an X12 correspondence, and an agent calling a payer or driving a portal is emitting the same transaction by another route. He says phone, portal, and X12 surfaces are built by different teams and can all agree on the wrong answer, so no surface is ground truth; his system keeps a semi-correct internal X12 representation, correct only until downstream evidence says otherwise. Enterprise memory must live in a database rather than local disk, and swapping in a stronger model is not automatically better inside a system built around the old one—evals and validation must be redone. The goal is cutting insurance costs and improving patient experience, and that means being AI pilled but also AI skeptical.

200 Million Patient Interactions Later — Vivek Muppalla, Hippocratic AI
Aug 19, 2026 · 20:40
Vivek Muppalla, head of AI engineering at Hippocratic AI, argues clinically safe AI voice agents can end healthcare rationing, citing 200 million conversations and 60-plus health systems. He details Polaris, their architecture running 31 models per call—one central model plus 30 specialists for labs, medications, scheduling—in parallel after each checks if it needs to speak. Their decoder-only audio LLM adds context and domain knowledge for drug recognition; single-word answers get rescored because 'no' heard as 'now' is catastrophic. With 4-bit quantization, speculative decoding, and KV cache compression, latency savings are reinvested into intelligence. But at 10,000 calls a day, 99% accuracy means 100 wrong appointments, so they use 7,000 clinicians and 450 tests to catch a 1% failure rate.

AI is the World’s largest Relationship Therapist — Clay Cockrell & Tony Fabrikant, CoupleWork AI
Aug 19, 2026 · 16:43
Clay Cockrell and Tony Fabrikant, co-founders of CoupleWork AI, argue that AI is now the world's largest relationship therapist—ChatGPT's 900 million weekly users dwarf BetterHelp's 5 million—but sycophancy makes it clinically dangerous. Cockrell, a 34-year couples counselor, contrasts that with Gottman's 90% divorce-prediction accuracy and emotionally focused therapy, which CoupleWork's Maxine is built on. He warns general AI misses abuse signals and lacks data privilege. Fabrikant says start with clinicians, encode evals, and treat one failing safety test as disqualifying.

Shipping AI to a Million Patients Without an A/B Test — Jared Joselowitz, Ufonia
Aug 19, 2026 · 19:15
Jared Joselowitz, research engineer at Ufonia, explains how the company ships Dora, a regulated medical-device voice agent, safely to patients without A/B tests, since randomizing patients into worse care is unethical. Dora has made 200,000 clinical calls across 20 UK hospitals and is contracted to reach a million patients in two years. Because 5% of patients is thousands and a red dashboard means someone was harmed, Ufonia built Matrix, where an LLM patient (PatBot) converses with Dora and a second LLM judge (BevJudge) flags hazards. A PPI study showed real patients found the simulated patient more realistic than a real patient in 3 of 4 sets, and the judge beat 10 clinicians on sensitivity. Prompts are optimized with GEPA against a cost matrix, because you ship the evidence, not the model.

Guardrails First: Engineering Member-Facing Health AI — Rashi Agrawal, Hinge Health
Aug 19, 2026 · 21:49
Rashi Agrawal, who leads AI/ML at Hinge Health, argues most healthcare AI safety failures are architectural decisions made before a token is generated, not model failures. With 40 million people self-triaging, she cites a chatbot telling a 60-year-old to take sodium bromide, Mount Sinai finding under-triage of life-threatening emergencies half the time, and ECRI ranking chatbot misuse the top 2026 hazard. Her architecture puts PHI stripped at the pipeline boundary, irreversible decisions like 911/988 routing and identity verification in deterministic code above the prompt, and continuous judges scoring live traffic. For launch decisions, she offers five rules — worst case wins, default to the safer mistake, calibrate to revealed tolerance — and says verify the judge before changing the agent.

From Ambient Documentation to Clinical Intelligence — Chaitanya Asawa, Abridge
Aug 19, 2026 · 21:35
Chaitanya Asawa of Abridge argues that everything in healthcare sits downstream of the doctor-patient conversation, and that the administrative machinery around it can be automated. Abridge began with clinical documentation to end 'pajama time' and reached 300 of the largest US health systems in two to three years. For its contextual clinical decision support, quality is paramount because the generator-verifier gap is tiny; Abridge has physicians write independent rubrics, adjudicated into a final rubric, with expert-calibrated LLM judges scoring responses. At a run rate of 100 million medical conversations a year, Abridge decomposes each note into sections and post-trains smaller models per section, betting its unique dataset and narrow problems can outrun frontier models.

The Next Medium: Why Real-Time Interactive Video Changes Everything — Ahmed Ahres, Reactor
Aug 18, 2026 · 17:30
Ahmed Ahres, head of go-to-market at Reactor, argues that world models are real-time interactive video, a change of medium rather than a speedup, because video becomes programmable like software. GPS enabling Uber and viewfinders enabling Instagram and TikTok show how real time unlocks new applications. Reactor's platform serves infinite interactive video (Helios from ByteDance), controllable worlds (Lingbot from Alibaba, LongLive 2 from Nvidia), and live avatars he admits are still uncracked. Users build interactive live streams, medical and cooking simulations, and video-to-video editing. Real-time infrastructure means streaming pixels, live sessions with memory, and sub-100-millisecond latency; he offers promo code AIE2026 for $75 credits.

Infra behind Krea 2: How to train and serve at scale — Gabriel Jorge Menezes, Krea.ai
Aug 18, 2026 · 16:55
Gabriel Jorge Menezes of Krea explains how Krea 2 was trained from scratch on thousands of GPUs, arguing metrics, not mystery-solving, made scale survivable. He calls GPU utilization a lie, tracking tensor core utilization instead, which climbs as resolution steps from 128 to 1024. Any GPU above 78 degrees gets pulled without debugging; custom InfiniBand and NVLink error collection caught most failures, which were silent cross-node timeouts. Aggressive checkpointing on a Wekker filesystem writing nearly a terabyte per second let crashes restart and run 24 hours on same nodes. Production and training share one cluster: gang scheduling kicks inference to external providers via a Virtual Kubelet fake node; taints stop wasted GPUs; a descheduler migrates pods back so the site never drops.

Generative Video at the Speed of Light — Keegan McCallum, uRun
Aug 18, 2026 · 8:43
Keegan McCallum, founder of inference provider uRun, argues that generative video's interesting axis is no longer quality but efficiency and long-horizon generation. He shows Helios, a distillation of Wan 2.1 14b, generating clips in real time for ~one-hundredth the cost of a slower frontier-quality clip, and notes at least 40 real-time/long-horizon models released this year. Ten dollars buys three hours of continuous generative video, and fifty buys fifteen; this unlocks magic-mirror webcam transformations, visual mediums for people who don't think in text, and content creation where you steer a generation in under a second. The hard part is serving: global GPUs, WebRTC with ICE/TURN, and synchronized streaming pipelines; uRun is building a React component, Python runtime, and MCP/CLI.

Voice agents with Realtime Video — Sidney Primas, LemonSlice
Aug 18, 2026 · 26:36
LemonSlice CTO Sidney Primas explains how his startup builds real-time video avatars by pointing world models at humans. A Microsoft partnership put a Teddy Roosevelt avatar in a replica Oval Office, generating continuously for eight hours with no reset. He argues audio embeddings drive emotion, and because avatars only look backward, errors compound, so LemonSlice trains with an attention mask and collapses roughly 30 denoising steps to one. He says serving video costs about the same as a voice model, and that the model harness, which orchestrates GPU/CPU threads and queues to avoid stutters, holds much durable value. He predicts an emotion engine for better reactions and an end-to-end EQ layer within two or three years, taking user video and audio and outputting avatar video and audio.

While my guitar gently speaks — Todd Fisher, Philo Ventures
Aug 18, 2026 · 18:35
Todd Fisher demonstrates how he built a speaking and singing guitar, walking through the audio engineering and AI stack behind it. He traces it from a Halloween Stranger Things garage door display to a JUCE plugin in Logic Pro that plays text-to-speech clips when he picks notes. Slicing words proved hard: energy gap segmentation fails on continuous speech, so he combined it with a sonority peak syllabifier and still hand-edited boundaries. For singing, he uses YIN pitch detection, synthesizes the note, and routes it through a vocoder. Because pitch-shifting samples from World and VocalSet is too heavy to run live, he pre-bakes vowel samples per fret. He also adds a microphone, Whisper, and a local LLM, letting the guitar answer audience questions like 'What is reality?'

The Next Game Engine Won't Have a Manual — Arturo Nunez, Nereu
Aug 18, 2026 · 19:33
Arturo Nunez, creator of the AI-assisted game engine Nereu, argues the next game engine won't have a manual: instead of coding, users describe intent. Drawing on a decade at Unity, he built Nereu around an asset tag system (ATS) inspired by Entity Component System, where everything is an asset tagged with intent (e.g. character, animated, double jump) and systems query tags—so a building can be tagged drivable and dropped into a Mario Kart race. The assistant, Bibi, helps users get unstuck rather than one-shot finished games, and context is assembled using level of detail: assets near the edited object get full tag values, distant ones collapse to position and type, and hundreds of grass pieces are omitted. Nereu also uses vision models to tag 6,000–7,000 assets, runs in the browser with JavaScript, and is in closed alpha. Nunez contrasts this with world models, which he says remain far from real-time 4K 60 fps games.

Building an Agentic Video Editor for Mass Consumer — Ekaterina Deyneka, Reelful
Aug 18, 2026 · 12:45
Ekaterina Deyneka, founder and CEO of Reelful, presents her agentic video editor as structurally identical to an agentic app builder, with editing real footage harder than generation. At Reelful, users drop in media plus a prompt, and the agent understands the media, proposes a creative plan for approval, then spins up a remote sandbox where agent skills encode cut rules, font pairings, and b-roll generation. Video is composed as React code via Remotion, a verification layer catches composition errors and sends the agent back to reiterate, and the result renders to a polished clip. To hide this complexity from consumers, Reelful is mobile-first with directional templates and a building editor for tweaks. Deyneka demonstrated clips made purely by the agent and announced Reelful's funding from a16z Speedrun.

Training Krea 2: What matters in generative model training — Sangwu Lee, Krea.ai
Aug 18, 2026 · 21:46
Sangwu Lee of Krea.ai explains what went into training Krea 2, its open-sourced image foundation model, arguing that once architecture is locked, data is everything. To keep stylistic diversity over the 'most boring average person' consistency of production models, Krea filters billions of images: no AI images, OCR and vision-language captions, hash and embedding dedup, distilled VLM classifiers, sparse autoencoders as unsupervised taggers for watermarks and borders. World knowledge is checked against Wikipedia concepts by PageRank. Training goes from 256 to 1K resolution through pre-training, mid-training, SFT, preference optimization, and GRPO-style RL, plus a prompt expander. Lee wants a single clean transformer without VAEs/text encoders and VLM-generated bounding boxes or scene graphs.
Powered by PodHood