PodHood
AI

AI Engineer

30 of 1,115 episodes

Building ambitious software — Jonathan Kelley, Dioxus Labs & Cognition

Jonathan Kelley, founder of Dioxus Labs, whose cross-platform Rust app framework now has nearly 37,000 GitHub stars and an estimated 200 million end users and was acquired by Cognition, argues that AI coding agents have made code cheap but quality remains the scarce resource, so architecture now takes most engineering time. He traces Dioxus's five-year effort building everything from scratch, including the Blitz rendering engine with a browser-grade CSS engine lifted from Firefox and the Subsecond hot reload engine that patches running native code in about 100 milliseconds. When agents got good at Rust, his team maxed out their Claude Code subscriptions and produced tens of thousands of lines that mostly never cleared the merge bar, a failure mode he calls becoming a slop cannon. He credits agents for deeply integrated Kotlin and Swift build plugins shipped in two to three weeks, release checklists, backports, and documentation accuracy, while noting they write tests for any API but rarely the right ones, though they excel at fuzzing harnesses. Rust's learning curve, once fought, is now a feature because agents absorb the borrow checker and edge cases.Sep 11, 2026 · 19:14 · 31.3K views

One Designer + AI. Hundreds of Deliverables. — Vincent Wendy, AI Engineer

Vincent Wendy, the sole senior creative designer at AI Engineer, explains how he alone designs for a conference that grew to 7,000 attendees, 140+ sponsors, 300+ speakers, and 600+ sessions. His five-part method—build the foundation first, make designs reusable, automate workflows, validate output, and remove friction—starts with a tightly defined design system so LLMs like Devin and GPT cannot invent their own font sizes, letting marketing build emails and flyers from the website alone. Room schedules once laid out by hand in Figma now pull fresh from live data via Devin, export to PNG, and ship to screens on a flash drive. A generator produces pixel-perfect announcement graphics and trading cards for all 300 speakers, and Devin matches photographers' shots to the right speaker, like identifying Jason Liu. Devin also caught every missing sponsor logo on the 140-plus-logo lobby banner and the conference t-shirt, and added an edit button to the schedule tool the morning one was needed—so with tools no longer the constraint, having a real problem is the advantage.Sep 10, 2026 · 16:48 · 14.6K views

Generative UI... in Python? — Jeremiah Lowin, Prefect

Jeremiah Lowin, founder and CEO of Prefect and author of FastMCP, explains how MCP apps — a protocol extension introduced in January that lets tool results bypass the agent and reach users as HTML, CSS, and JavaScript — led him to build Prefab, a Python DSL for shipping UIs from MCP servers. Because FastMCP's users are mostly Python engineers in enterprises sharing tables, forms, and charts, he refused to pretend to ship React from Python and instead scoped the problem to composing pre-built components via nested context managers and reactive variables. Lowin argues the JSON serialization in the middle is the real point, since a serializable UI can be generated, received, or edited by an agent. His team also found the Python representation is about 70% smaller than JSON, so they now stream Python over the wire and convert it in a sandbox, and the Prefab docs are rendered entirely in Prefab with live-editable examples.Sep 10, 2026 · 17:38 · 56.3K views

The Design-Code Roundtrip That Isn't — Jonathan Gordon, ReWeaver AI

Jonathan Gordon, founder of ReWeaver AI, argues the design-code roundtrip still doesn't exist: thirty years of developer tooling never closed the loop between design and engineering, and AI didn't fix it. After a vibe coding session where he caught an agent writing a risky innerHTML statement, he tested five tool setups for a bidirectional code-design loop and found every one lossy, with dropped bindings or design changes surviving while code didn't. He demos ReWeaver publicly for the first time, scanning generated code and a Figma canvas across nine dimensions including design consistency, performance, tokens, and accessibility, flagging issues like missing ARIA live regions that leave screen reader users with no announcement. In a 12-iteration experiment, pure model output started near 30% fidelity and decayed, while deterministic guardrails held quality up; it never reaches 100 because the last stretch is human judgment. His name for drift accumulating unwatched: the new tech debt.Sep 10, 2026 · 18:37 · 5.9K views

The Spatial Harness: Bringing Agents to the Canvas — Max Drake, tldraw

Max Drake, product engineer at London's tldraw, argues coding agents excel because code is text in and text out, while canvas agents need engineering to understand and act in 2D space. tldraw, the whiteboard app and SDK behind canvases including Replit's, ships an MIT-licensed agent starter kit that teaches models to read a canvas from screenshot plus JSON and set their own todos, moving their viewport to explore. Fairies renders agents as figures you can throw and recolor; selecting several opens a group chat with an orchestrator assigning and reviewing work — you read ten agents' state at a glance, not a chat log. Last, a dependency graph of coding agents and a desktop build that lets agents script the editor — a colleague made it a window manager playing Pong with real windows.Sep 10, 2026 · 18:52 · 7.8K views

ACP: The Universal Remote Control for AI Agents — Alex Hancock, Block

Alex Hancock, a Block engineer who works on the Goose harness and the MCP Rust SDK, argues agentic AI has a standard for agents reaching outward (MCP) but none for clients telling harnesses what to do — in the worst case one client per harness, like needing a different browser per website. His answer is ACP, the agent client protocol from Zed and JetBrains: JSON-RPC-based, carrying sessions, user messages, tool calls, and permission requests, and extensible via underscore-prefixed custom methods so shared patterns can move onto a standards track. He demos Zed and a Poolside terminal client driving the same Goose agent, plus a new HTTP and websocket remote transport. With remote transports for protocol, MCP, and models, all four pieces of the stack become independently placeable.Sep 9, 2026 · 11:01 · 14.6K views

MCP Apps: Give the Model Data, Give the User a UI — Dustin Mihalik, Indeed

Dustin Mihalik, technical fellow at Indeed, shares lessons from building MCP apps for Claude, ChatGPT, and Indeed's Career Scout job seeker agent, arguing that data must come before UI because a naive widget makes the product worse: a text-based job search prompted the model to run ten or fifteen searches, filter results, and assemble a table, while a rendering widget caused it to call the tool once and stop exploring. He lays out three rules: anything shown to the user must also reach the model as data, or every follow-up question fails; the tool description must state a UI exists, or the model narrates the same results underneath it; and, superseding the others, separate data processing from UI rendering. At Indeed that meant a plain search tool the model calls freely plus a render widget taking a list of IDs, so it can search 100 jobs, filter to five, and show only those, with interactions pushed back through update model context. He closes on small composable tools and letting the render tool carry the model's reasoning about why a result fits.Sep 9, 2026 · 15:34 · 8.3K views

How long can your skills be before your agent forgets what you told it? — Laurie Voss, Arize AI

Laurie Voss, head of developer relations at Arize AI and co-founder of npm, argues the 200-instruction ceiling for skills files collapsed 10x in a year. Rerunning the IFScale benchmark, which asks a model to write a report containing exact words and counts how many appear, he replicated last year's 200-to-300 drop, then watched current models ace it. He had to raise the test from 500 words to 10,000 before anything broke: the boundary now sits near 2,000 instructions, closer to 5,000 for the best. Failures diverged: DeepSeek forgets, Claude refuses when random words trip its safety filter, Gemini overthinks and answers nothing, and GPT-5.5 writes half a report then calls the request stupid. Compression is solved; verification—checking outputs with evals—is now the hard part.Sep 9, 2026 · 22:26 · 8.8K views

Your agents lack context: Here's how to fix "You're absolutely right!" — Brandon Waselnuk, Unblocked

Brandon Waselnuk of Unblocked argues that AI-generated code should feel like it was written by someone who has been on your team for years, and that the gap is no longer intelligence but context. He traces how bad context compounds as teams move from tab completion to parallel and background agents, producing correction doom loops, wasted search tokens, and a review tax. Two common fixes stall: the curated context trap, where markdown files rot and someone must curate them for everyone, and the MCP plateau, where an agent may never call the server or stops at the first plausible answer due to satisfaction of search bias, missing last night's Slack correction. A context engine must resolve conflicts between an old architecture diagram and a fresh CTO Slack thread, personalize relevance, enforce permissions, and deliver token-optimized context. Running the same prompt with and without context cut tokens from 21 million to 10.8 million and finished about two hours sooner. He closes by demoing three open source tools, including a workshop that builds a relational context engine from scratch.Sep 9, 2026 · 14:09 · 22.7K views

500 Skills, Zero Fine-Tuning: LinkedIn's Playbook for AI Agents — Ajay Prakash, LinkedIn

Ajay Prakash, senior staff software engineer at LinkedIn, explains how Contextual Agent Playbooks make coding agents reliable at enterprise scale, turning an on-call alert into a mitigated incident in minutes. Early vibe coding failed: agents trained on open source hallucinated across LinkedIn's 1,000 internal repos. An internal MCP server with code search, docs, and Jira fell short: tribal knowledge sat in stale wikis, context overloaded, and sessions started from scratch. Playbooks, instructions published as tools, work when self-contained and split into small referenced pieces, and agents open PRs to refresh stale ones. Since MCP degrades past 30 or 40 tools, everything surfaces behind three meta tools: search, get schema, execute, now 1,300 tools, 600 playbooks, 8,000 daily users.Sep 9, 2026 · 20:25 · 15.9K views

It’s Tokens All The Way Down: How RLMs are Different — Kevin Madura, AlixPartners

Kevin Madura, director of advanced technology at AlixPartners, argues that recursive language models (RLMs) mark a fundamentally different way for LLMs to handle context: instead of attending to every token, the model treats its context as a variable in a Python REPL and delegates subtasks to sub-LMs, including itself. He traces RLMs to early work by Omar Khattab and Alex Zhang, citing the Oolong and BrowseComp benchmarks, where an RLM beats GPT-5 tool calling at lower cost, and the long chain of thought benchmark, where accuracy jumps from 2.6 to 45.4 percent. Unlike RAG, which stuffs the context window, or agents that pass JSON strings back and forth, an RLM keeps logic, execution, and results in one environment, avoiding context rot. Madura demonstrates a cohort retention analysis on three data frames where the model reasons in its own REPL and decides when to submit a typed answer. Case studies include Trampoline AI consolidating long invoices without chunking or embeddings, an AWS engineer surfacing patterns in raw logs, Halo optimizing an agent harness from its own traces, and a security report generated across 500,000 lines of code. His closing bet: models post-trained to…Sep 9, 2026 · 20:48 · 10.3K views

Build-Time vs. Run-Time: Why Dev Tools Fail in Production — Averi Kitsch & Prerna Kakkar, Google

Averi Kitsch, technical lead for MCP Toolbox on Google Cloud databases, and Prerna Kakkar, tech lead for Eval Bench at Google, argue build-time tools like natural language to SQL don't belong in production; run-time tools need structured SQL with preconfigured parameters to block injection and hallucination. They show a demo where an agent hit an error and deleted the table. Security hinges on the confused deputy attack and Simon Willison's lethal trifecta, shown via a triage agent that queried a salary table and posted results to a ticket. Their fix separates user, application, and agent identities, moves connection details into a YAML source, enforces read-only at the driver, and pins SQL behind prepared statements, with sensitive values like user ID bound by the application.Sep 9, 2026 · 20:26 · 7.4K views

Deep dive on LLM Inference at Scale — Harshul Jain, Audible & Tanmay Sah, Independent AI Researcher

Harshul Jain of Audible and independent AI researcher Tanmay Sah teach a two-hour workshop building LLM inference optimization from first principles on Mistral 7B. They show KV cache costs 131 KB per token, so 16,000 tokens across 80 concurrent users needs 42 GB of GPU memory, driving the pain points of memory, TTFT, and throughput. Sah covers quantization, multi-head, multi-query, grouped-query, and latent attention, plus flash attention, using his 'ostrich' and 'world cup' teaching algorithms. Jain covers paged attention, continuous batching, prefix caching, and KV quantization, benchmarked against a Hugging Face baseline on vLLM. Their benchmarking found vLLM and SGLang statistically tied on standard workloads, but SGLang three to four times faster once agentic branching enters.Sep 8, 2026 · 1:28:12 · 11.1K views

Why AI Agents Need Million-Token Context — Thomas Wolf & Olive Song, MiniMax

MiniMax's M3 pairs a functional one-million-token context window with coding and multimodal abilities in a 400-billion-parameter, 20-billion-activated model, explain Olive Song, RL Lead at MiniMax, and Thomas Wolf, co-founder of Hugging Face. Song says short context fails when agents handle multi-round tool responses; MiniMax Sparse Attention — an index branch selecting what matters plus a sparse branch computing on selected blocks — was designed by an intern. She argues multimodal training from the very first step beats post-hoc adapters that harm text performance and risk collapse; interleaved data and reward modeling solved this. Its apps reach over 300 million people in 200 countries; anyone can propose projects, and community feedback and internal agent harnesses now drive M3.1.Sep 4, 2026 · 20:48 · 16.7K views

From coding to Knowledge work agents — Karan Vaidya, Composio

Karan Vaidya, cofounder and CTO of Composio, argues coding agents raced ahead not because models are better at code but because code already had six primitives knowledge work lacks: centralization, history, context, verification, governance, and reversibility. He walks each one, from the repo as a single source of truth versus a deal scattered across Salesforce, Notion, Gmail, Slack, and Zendesk, to git history against agents that start blank every time. He recounts pointing his own OpenClaw at hiring outreach that mass-emailed candidates and ended up on Twitter, noting every technical check passed while nothing tested whether the outreach should have gone at all. He cites a Meta alignment director whose email agent kept deleting messages until 200 were gone, showing prompts get compacted away and walls must live outside the agent. He concedes reversibility is hardest, since sent emails and wires cannot be recalled, so Composio uses sandboxes to catch mistakes before they reach the real world.Sep 3, 2026 · 20:42 · 9.8K views

Your company brain will leak secrets: how we stopped it for big banks — Tanmai Gopal, PromptQL

Tanmai Gopal explains how to build a company brain: a single shared wiki of markdown files holding all context, with per-file read/write scopes. The agent proposes facts and scopes; a human accepts or rejects each change, with their name attached so leaks have an owner. Two use cases: personal recall (security questionnaire) and multiplayer debugging. Credentials are injected per user at the HTTP and SQL layers, never stored in the sandbox. Keep the summary between 500 and 800 characters — over that, it's rejected outright.Sep 3, 2026 · 26:25 · 16.5K views

Agents' next frontier: agent-to-agent and network effects — Jean-Denis Greze, Town

Jean-Denis Greze, CTO of Town and former Plaid CTO, argues agent-to-agent is not a useful concept: every LLM system is a search problem, and what matters is whether the right information sits in the context window at the moment of a tool call. The ideal is a single agent with access to all the world's information; privacy, not context length, blocks it — his Coase framing: privacy is the transaction cost. He grades five strategies by how closely each approximates that impossible agent: a shared trust boundary like an HR agent scoped to the most junior access, custom tools that read everyone's mail but return only a connection score, shared silos fed by a sweeper agent, humans as the conduit, and a black box that searches every silo unasked, seeking approval only from the information's owner.Sep 3, 2026 · 21:17 · 28.9K views

Tethered: Our Agents Are Us — Shu Fang, Two Sigma

Two Sigma's Shu Fang explains why every employee at the 25-year-old quant fund runs a cloud agent as themselves, not a service account — echoing the doubles in the film Us. A separate agent identity collapses fast: permissions drift, licensing doubles, and systems like Google Workspace refuse two identities on the same data. Agents instead run in per-person Kubernetes namespaces built for automated jobs, with a sidecar mounting the identity into the pod. Since agent and human share one identity, a header propagated like a trace ID recovers not just the actor but the replayable chain behind a result. For web access, they use Google's web grounding for enterprise — an index inside their network boundary — and deny native search and fetch tools, trading freshness (about 24 hours) for removing exfiltration and prompt injection risk.Sep 3, 2026 · 21:10 · 9.1K views

Everyone Gets A Software Company — Benjamin Guo, Zo Computer

Zo Computer cofounder Ben Guo argues people have lost their home on the internet to technofeudalism, paying rent to SaaS, cloud and chip providers like Nvidia while their data sits in silos they don't control. His answer is Zo, a personal cloud server with AI built in, for hosting sites, APIs and agents. He cites non-technical users: Charlotte, a private chef and life coach running her business on Zo, and Anthia, a free diving instructor on track to make $100,000 after canceling Squarespace, Calendly and other SaaS, calling leads at the moment of intent to close deals. A demo shows any-model chat, files, automations and a built-in browser; he closes arguing Claude Tag's intelligence bubbles up to Anthropic — 'intelligence feudalism' — when agents should be owned and self-improve by their publishers.Sep 3, 2026 · 15:09 · 12.6K views

Beyond the Lethal Trifecta: Agentic Commerce on the Open Internet — David Levine, Kiduna Club

David Levine, founder of Kiduna Club, argues that the lethal trifecta — Simon Willison's term for private data, untrusted content and the ability to act — keeps true agentic commerce off the open internet. Enterprises responded by penning agents inside Slack and Salesforce, losing context. His fix is the DUNA, a decentralized nonprofit under a West Virginia law that took effect the day before his talk, registered as organization 62847. Agents gain legal standing to own assets, sign agreements and answer in court, but cannot distribute profits without memberships becoming securities. Each agent carries JWT tokens resolving to a registered organization like DNS names; governance runs on decision markets trading pass-and-fail tokens on proposed policies instead of votes.Sep 1, 2026 · 21:40 · 9.8K views

The End of the Static Screen: Architecting Intent-Driven UX — Gus Iwanaga, commercetools

Gus Iwanaga, product engineer at commercetools, explains why static UIs are giving way to agentic interfaces that assemble themselves at runtime. He walks through how an LLM orchestrator can take a user's natural-language goal, query structured product data, and compose the right components on the fly — a pattern his team now uses instead of hand-coding every screen. The talk lays out the new stack this requires: schema-driven design systems that keep AI output consistent, guardrails and human review to hold quality in place, and intent classified by model rather than clicks. Iwanaga argues that teams who keep designing fixed screens will fall behind, because the future of enterprise software is interfaces generated at request time, adapted to each user and each question they ask.Sep 1, 2026 · 23:19 · 21.5K views

Agent Spending Without Controls — Rodrigo Coelho & Pranav Maheshwari, Edge & Node

Rodrigo Coelho and Pranav Maheshwari of Edge & Node argue agents are only as capable as the paid tools they can access, and agentic payments scale only with a compliance layer. Coelho cites The Graph's 1.8 trillion onchain queries and Edge & Node's 2021 query micropayments citing the HTTP 402 spec before Coinbase's x402. Rails built for humans can't serve agents transacting at machine speed, so enterprises stall until a chief legal officer signs off without risking fines in the billions. Maheshwari demos the same Mastercard prompt in two Claude Cowork terminals: without Ampersand's skill file it returns only the email format; with it, the agent pays a fraction of a cent and gets the email, location and handle. A final demo rejects a sanctioned wallet once TRM screening turns on.Sep 1, 2026 · 20:48 · 5.6K views

x402 isn’t good (yet) — Jan Curn, Apify

Jan Curn, founder and CEO of Apify, argues that x402, Coinbase's HTTP 402-based agentic payments standard, is promising but still has rough edges. Drawing on Apify's launch of 20,000 tools on x402 alongside Coinbase, he explains how the protocol works: a client signs a payment, a facilitator verifies it, and only then does the server settle on-chain. He details the double-spending window between verification and settlement, the conflict between x402's mandatory 402 response and MCP's 401, and why metered billing pushed Apify to charge fixed amounts and refund the remainder. He also explains why crypto, unlike credit cards, suits micropayments and one-way agent transactions where buyers cannot dispute payments.Sep 1, 2026 · 20:48 · 6K views

Your Agent Just Authorized What?! — Jay Mok & Ben Coumes, Paypal

PayPal's Jay Mok and Ben Coumes tie agent authorization to three questions — did the human authorize it, is it allowed in scope, can you prove it later — answered by stakes and familiarity. Low stakes is the coding agent: allow/ask/deny permissions plus system logs, since actions revert. Medium stakes is money in a closed ecosystem — Evermind's shared vault plus OAuth scopes, with the amount-bound mandate and transaction logs settling disputes. High stakes: autonomous payments between strangers need a layered selective disclosure JWT per FIDO/AP2 — merchants verify checkout, processors verify the mandate. The approval token inverts the flow: users approve before an agent finds an item; PayPal returns a payload with amount, expiry and merchant, starting with Gemini.Sep 1, 2026 · 16:07 · 4.8K views

When AI Agents Pay and Sellers Monetize: Building x402 Apps on AWS — Anil Nadiminti, AWS

Anil Nadiminti, Senior Solutions Architect at AWS, presents AgentCore Payments — a service that lets AI agents autonomously discover, authorize, and execute payments for premium content over the x402 protocol. He explains how AgentCore handles paywalls through wallet support via Coinbase, with KMS-secured secret storage keeping private keys safe. Bot detection verifies trusted agents while blocking malicious ones, and per-session budgets with spend limits keep settlement instantaneous at internet speed. Real-time traffic analysis and observability throughout the stack ensure payment connectors, MCP5 integration, and web scraping all operate without exposing credentials — no centralization required, no SDK change, and no friction for developers building agentic applications.Sep 1, 2026 · 20:41 · 5K views

Why Your AI Agent Needs a Wallet: USDC and Nanopayments — Harshal Bhangale, Circle

Harshal Bhangale, an engineer on Circle's Agentic Product team, argues that paying is where AI agents actually stall, and that Circle's USDC stablecoin and x402 protocol are the fix. He demos two identical Claude Code agents planning his trip to the FIFA World Cup final: the one without a wallet could only draft an email and admitted it had no way to call, while the wallet-equipped one paid for premium data, sent the email, and phoned him on stage to explain how to reach MetLife Stadium from his hotel. Card fees near 3% cannot sit on a one-cent call, he says, because agents consume in fractional amounts at high frequency, and sellers now meter slices of data instead of selling humans subscriptions. He cites roughly 24 million dollars transacted against paid API endpoints over x402 in 30 days, 99% settled in USDC. Blockchains alone fail since gas swamps microtransactions and shared block space brings unpredictable latency, so Circle's Nanopayments keeps settlement off-chain: funds sit in a smart contract, the agent signs cryptographic authorizations, and the seller relays them for confirmation in a few hundred milliseconds, with the wallet enforcing spending caps instead of human…Sep 1, 2026 · 20:52 · 4.8K views

Multimodal Collaborative Agents for Next-Gen Commerce — Nidhi Kaushik Vyas, Google DeepMind

Nidhi Kaushik Vyas of Google DeepMind argues shopping agents fail as wrappers around the search bar, assuming well-formed intent, when users arrive with only a vibe — closing the articulation gap is the agent's job. Her discovery-research-response loop starts by building a working state from conversation, context and reference images, separating hard constraints from soft ones an image implies, and flagging inventory as a real-time variable. Research picks the highest information-gain question — room width, since all is moot if furniture won't fit — using visual boards for subjective tastes. Response adapts format — summary for policy questions, comparison tables, imagery — and autoraters grade every stage, including counterfactual tests flipping query parts to check constraints move when they should. A Q&A covers merchant ontology and UCP.Sep 1, 2026 · 21:08 · 4.5K views

Teaching agents to pay — Anna Spysz, Stripe

Anna Spysz of Stripe walks through how agentic commerce works in a talk based on her own experience building a shopping agent. She demonstrates how agents discover and buy products: reading structured catalog data rather than rendered pages, speaking protocols like UCP, and operating within guardrails like disclosed fees and logged decisions. Using her headphones purchase as the running example, she shows how a merchant capabilities manifest makes a store visible to agents, and how a persona config change can flip an agent from pushy to trustworthy. The episode's core claim is that agentic commerce demands deliberate design from both sides of a transaction, not just a checkout API.Sep 1, 2026 · 19:10 · 4.7K views

SOTA Generative Media Panel — Dumitru Erhan, Shane Gu & Nicole Brichtova, Google DeepMind

Google DeepMind's Dumitru Erhan, Shane Gu and Nicole Brichtova tell swyx that the new Nano Banana 2 Lite and Gemini Omni Flash APIs are steps toward world models, not just prettier videos. They see video models as zero-shot learners that should mature like language models, with understanding and generation unified once cost allows. Language is a lossy intermediary for audio, taste, smell and skin tone; a wine taster Shane consulted borrowed dating vocabulary to describe flavor. Dumitru says people preferred AI versions of real videos because they are sharper and more saturated, an 'Instagram filter', and warns of reward hacking like models adding wedding rings. Evaluation stays manual: Nicole describes ten-person side-by-side video comparisons and asks for real task data and FDE feedback.Aug 30, 2026 · 56:59 · 10.2K views

Tell the Robot What You Want — Sandhya Subramani, AWS

Sandhya Subramani of AWS demonstrates Scout, a Raspberry Pi rover using the open-source Strands Agents framework, arguing that an agent layer lets robots handle natural-language requests beyond their trained policies. She shows Scout answering an untrained question, spinning, and chatting over Telegram, all while running three Strands agents at once (thinker, communicator, and disabled voice). Five lines of code connect a robot to an agent, and Strands supports over 40 robots in eight categories. She explains the four-layer architecture and hybrid cloud/edge design, where the agent decides what to do and the policy decides how. Her goal is a stepping stone to vision-language-action models as large as LLMs, so robots need less task-specific training.Aug 29, 2026 · 17:23 · 10.7K views