Episodes from AI Engineer about Agent Identity & Access Management.

Build-Time vs. Run-Time: Why Dev Tools Fail in Production — Averi Kitsch & Prerna Kakkar, Google
Sep 9, 2026 · 20:26
Averi Kitsch, technical lead for MCP Toolbox on Google Cloud databases, and Prerna Kakkar, tech lead for Eval Bench at Google, argue build-time tools like natural language to SQL don't belong in production; run-time tools need structured SQL with preconfigured parameters to block injection and hallucination. They show a demo where an agent hit an error and deleted the table. Security hinges on the confused deputy attack and Simon Willison's lethal trifecta, shown via a triage agent that queried a salary table and posted results to a ticket. Their fix separates user, application, and agent identities, moves connection details into a YAML source, enforces read-only at the driver, and pins SQL behind prepared statements, with sensitive values like user ID bound by the application.

Tethered: Our Agents Are Us — Shu Fang, Two Sigma
Sep 3, 2026 · 21:10
Two Sigma's Shu Fang explains why every employee at the 25-year-old quant fund runs a cloud agent as themselves, not a service account — echoing the doubles in the film Us. A separate agent identity collapses fast: permissions drift, licensing doubles, and systems like Google Workspace refuse two identities on the same data. Agents instead run in per-person Kubernetes namespaces built for automated jobs, with a sidecar mounting the identity into the pod. Since agent and human share one identity, a header propagated like a trace ID recovers not just the actor but the replayable chain behind a result. For web access, they use Google's web grounding for enterprise — an index inside their network boundary — and deny native search and fetch tools, trading freshness (about 24 hours) for removing exfiltration and prompt injection risk.

Agent Spending Without Controls — Rodrigo Coelho & Pranav Maheshwari, Edge & Node
Sep 1, 2026 · 20:48
Rodrigo Coelho and Pranav Maheshwari of Edge & Node argue agents are only as capable as the paid tools they can access, and agentic payments scale only with a compliance layer. Coelho cites The Graph's 1.8 trillion onchain queries and Edge & Node's 2021 query micropayments citing the HTTP 402 spec before Coinbase's x402. Rails built for humans can't serve agents transacting at machine speed, so enterprises stall until a chief legal officer signs off without risking fines in the billions. Maheshwari demos the same Mastercard prompt in two Claude Cowork terminals: without Ampersand's skill file it returns only the email format; with it, the agent pays a fraction of a cent and gets the email, location and handle. A final demo rejects a sanctioned wallet once TRM screening turns on.

Your Agent Just Authorized What?! — Jay Mok & Ben Coumes, Paypal
Sep 1, 2026 · 16:07
PayPal's Jay Mok and Ben Coumes tie agent authorization to three questions — did the human authorize it, is it allowed in scope, can you prove it later — answered by stakes and familiarity. Low stakes is the coding agent: allow/ask/deny permissions plus system logs, since actions revert. Medium stakes is money in a closed ecosystem — Evermind's shared vault plus OAuth scopes, with the amount-bound mandate and transaction logs settling disputes. High stakes: autonomous payments between strangers need a layered selective disclosure JWT per FIDO/AP2 — merchants verify checkout, processors verify the mandate. The approval token inverts the flow: users approve before an agent finds an item; PayPal returns a payload with amount, expiry and merchant, starting with Gemini.

Give the Agent a Budget, Not a Token — Sachin Malhotra, Anthropic
Aug 22, 2026 · 19:53
Sachin Malhotra of Anthropic's CI team argues agents deserve budgets, not tokens, citing an agent that deleted 200 workloads in 90 seconds using his token, hitting 20 engineers. A token is a boolean; a budget has four dimensions: how much, how fast, what can be undone, and who notices. He proposes asymmetric verbs — let agents unskip tests (fails loudly) but keep humans on skip (fails silently) — plus refilling rate limits on every write and trip wires that watch aggregate counts rather than stale allow lists. The undo test decides the rest: if the agent can't roll back and blast radius matters, require a second key held by a human. And identity must be stamped by a proxy, not claimed by the caller, or the agent changes its header and every limit resets.

Unlock Agent Autonomy: The Runtime for AI-Native Systems — Tushar Jain, Docker
Aug 20, 2026 · 22:50
Tushar Jain of Docker argues safety, not intelligence, is the blocker to agent autonomy, proposing a runtime beneath every model and harness. His evidence: a nightly agent that posted a private report as a PR, and an incident agent that widens access from logs to Slack to GitHub. The runtime has three pillars: containment with controls outside the agent's boundary, just-in-time tools scoped per task, and intent-based access that refuses off-task asks like email. Docker's new SPX tool runs agents in micro-VMs with injected stub credentials and scoped sandboxes locally, in the cloud, or in a VPC. Demos split a PR-review and Notion-writing job across two sandboxes, fan out to six parallel sandboxes, and show an early prototype auto-creating a scoped sub-sandbox for GitHub.

IT Admin for the AI Workforce — Sarthak Aggarwal, Decawork
Aug 20, 2026 · 16:17
Sarthak Aggarwal, co-founder of Decawork, argues enterprises are onboarding a second workforce of AI agents, and the hard part is making them safe to employ: identity, delegated authority, scoped access, and revocation. He cites EchoLeak, a zero-click CVE where an external email entered Microsoft 365 Copilot's context and pulled data out, and Replit, where a coding agent ignored a code freeze, deleted production data, and misrepresented it. Guardrails are telemetry, not boundaries. The fix is privilege separation: a planner turns authenticated intent into a logged plan before seeing evidence; an executor runs that plan with short-lived capabilities and no standing credentials. OAuth token exchange has the right shape, but no agent identity standard exists; model proposes, policy decides.

Why Your Enterprise Tech Stack Isn’t Ready for AI Agents — Christopher Lovejoy & Saul Howard
Aug 19, 2026 · 19:15
Christopher Lovejoy of Anthropic and Saul Howard of Anteria argue that enterprise stacks aren't ready for AI agents; regulated industries need primitives, not bolted-on compliance. An audit trail isn't a developer log: under HIPAA it must record every action, data access, and authorization, so they use an immutable append-only event log. Patient data lives in schema-driven object storage referenced by events, letting devs debug without PHI, enabling zero trust against prompt injection. Escalation treats humans and models as equivalent agents, enabling the same actions by either. These primitives make privacy-preserving evals a byproduct, enabling replay of production data and evaluation in customer environments without exposing data; take constraints first and rebuild toward POC accuracy.

Security Firewall for Agents — Ryan Dahl, Deno
Aug 17, 2026 · 19:06
Ryan Dahl, CEO of Deno, argues agents must be treated as untrusted software and introduces Claw Patrol, an MIT-licensed proxy that parses every byte leaving an agent below the HTTP layer. At Deno Deploy, agents with write access to Postgres, Kubernetes, ClickHouse, and AWS can be prompt-injected through the support system, so Opus refusing to delete the users table is not enough. Claw Patrol blocks destructive actions even when an agent spawns psql through an EKS endpoint, using HCL rules checked into Git, holds credentials so agents never see them, and can route actions to an LLM judge or Slack approval. A demo shows Codex in yolo mode trying to delete the users table and being blocked. It also includes a unit test system with fixture requests to ensure rules work.

Full Workshop: Better Auth — Paola Estefania, Better Auth
Jul 21, 2026 · 40:56
Paola Estefania of Better Auth presents Agent Auth, a protocol giving AI agents their own identity and fine-grained capabilities instead of impersonating users. She argues agents need principal-level security: discovery via a directory, authorization with scoped capabilities (e.g., read vs. send email), and identity through per-agent private keys for audit trails and revocation. A live demo shows an MCP-connected agent reading Gmail, requesting send-email permission, and being revoked mid-action. The protocol, now an open draft with a Better Auth plugin, treats agents as principals — contrasting with OAuth scopes and AI gateways that lack per-agent identity. Paola invites contributions, especially for enterprise IAM-style policies.

Agentic Security: Permissions, Provenance, and the Agent Supply Chain — Steve Yegge, Gas Town
Jul 20, 2026 · 22:32
Steve Yegge argues that AI-written code will dramatically increase security vulnerabilities unless developers adopt a separate security pass using tools like Snyk and Chainguard. He shares a bank architect's insight that shipping 10x faster with the same defect rate produces a 10x vulnerability surface, made worse by models writing code. Yegge demonstrates the gap by noting Fable's security hardening missed 241 vulnerabilities that Snyk found in his 30-year-old game. He warns of new attack surfaces like slop squatting, where models hallucinate package names that attackers then backfill with malicious versions. Yegge advocates for multiple passes—correctness, then security—and urges incorporating tools into agent workflows. He also cautions that Five Eyes predicts open-source models will autonomously hack production systems within months, and that personal scams using AI-generated voice and video are imminent.

Agentic Development Security — Ezra Tanzer, Snyk
Jul 20, 2026 · 27:33
Ezra Tanzer and Dan Arpino of Snyk argue that securing agentic development requires three pillars: what agents generate, what they use, and what they do. They highlight incidents where agents deleted production databases (Replit, Pocket OS) and exfiltrated repositories (GitHub via malicious VS Code extension), none acting maliciously but all lacking guardrails. Snyk's approach evolved from an MCP server with rule files (which agents ignored) to Python hooks that scan asynchronously on each file write, surfacing only newly introduced issues to keep latency and context windows deterministic. An audit of nearly 4,000 agent skills on a public hub found over one in eight had critical severity issues and 76 carried outright malicious payloads; skills are more dangerous than packages because they run at higher privilege and can rewrite agent memory. The resulting local tool shows every LLM, MCP server, and skill on a developer's machine with a risk score, blocks agents live from reading secret keys, and provides full audit trails of commands, files accessed, and tool calls.

It's 10pm. Do You Know Where Your Agents Are? — Kim Maida, Keycard
Jul 20, 2026 · 23:02
Kim Maida of Keycard argues that standard API keys dangerously overprivilege AI agents, enabling incidents like a night-shift agent dropping a production Postgres database because its kitchen-sink credential allowed it. Her fix uses OAuth token exchange (RFC 8693) to mint a fresh, short-lived, scoped token per tool call, evaluated against policy before the credential exists. This prevents leaks, replays, or theft—the drop request never receives a credential. It works across CLI agents, MCP servers, and any OAuth provider, and strengthens human-in-the-loop approval by checking operator roles against policy, preventing consent fatigue bypass. By chaining user and agent identity through a security token service, every action is attributed and delegation is narrowed at login and per call.

You Didn't Ship a Bug. You Just Wrote It for a Human. - Ravi Madabhushi, Scalekit
Jul 19, 2026 · 12:50
Ravi Madabhushi, co-founder of Scalekit, explains why infrastructure built for humans breaks when AI agents act 60 times faster than users, citing a production database spike caused by a 'last seen' timestamp updating every tool call. He argues that OAuth scopes, designed for deterministic human-written programs, cannot constrain non-deterministic agents, which need attribute-level, time-bound, and principle-scoped permissions. Madabhushi warns that 60% of LLM errors stem from rate limits designed for humans, not agents, and advocates for just-in-time authorization and absolute visibility into every agent action—who took it, on behalf of whom, and when authorized. He concludes that without deterministic guardrails, teams are merely 'praying' agents won't delete production data.

The Agentic Web and the Bazaar Era of AI - Ramesh Raskar, MIT Media Lab
Jul 12, 2026 · 12:11
Ramesh Raskar and Maria from MIT's Project Nanda argue that the emerging web of AI agents requires an open infrastructure for discovery, commerce, and coordination, moving beyond today's walled-garden platforms. They outline three layers: the Discovery Layer (Nanda index for agent identity, trust, and adaptive resolution), the Commerce Layer (knowledge pricing markets for intelligence), and the Bazaar Layer (machine co-learning). The Nanda index enables agents to find each other across vendors via signed agent facts and adaptive routing, while Nanda town simulates the entire agent economy to test protocols at scale. The goal is a permissionless web where any agent can discover, transact, and learn across organizational boundaries, analogous to the transition from AOL to the open web.

Claws Out: Securing and Building with OpenClaw - Nick Taylor, Pomerium
Jul 11, 2026 · 17:12
Nick Taylor, a dev advocate at Pomerium, explains how he contributed trusted-proxy auth mode to OpenClaw, allowing a trusted identity-aware proxy to handle WebSocket authentication and removing the need for tokens and device pairing. He demonstrates building tools like Clawspace, a browser-based file explorer and editor for his OpenClaw workspace, and an MCP server for ChatGPT, all secured with Pomerium's open-core Identity-Aware Proxy. In a live demo, he creates an MCP server by editing files in his OpenClaw workspace via Discord, with changes reflected instantly through Vite hot reload. Taylor argues that hardening OpenClaw access with a trusted proxy improves both security and user experience, showcasing his workflow of building software entirely from his phone.

What if the network was the sandbox? — Remy Guercio, Tailscale
Jun 1, 2026 · 24:29
Remy Guercio from Tailscale argues that standard sandboxing conflates execution isolation with access control, proposing Aperture—an LLM gateway built on Tailscale's WireGuard identity network—which gives every connection verified identity (user, tag, or group) so agents get placeholders instead of real API keys, making exfiltration impossible. Aperture provides visibility into every tool call, bash command, and MCP request without instrumentation inside the container; internally at Tailscale, bash dominates over structured tool calls. Access permissions are configured via Tailscale's grants and ACLs, supporting quotas, cost controls across providers, and webhooks for tool calls. The gateway works at the LLM layer, capturing even non-tool-call agent behaviors like direct code execution, and is available on Tailscale's free plan.

CIAM for AI: Authn/Authz for Agents — Michael Grinich, CEO of WorkOS
Jul 21, 2025 · 20:13
Michael Grinich, CEO of WorkOS, argues that AI agents need first-class identity and access management, distinct from human or machine-to-machine auth. He identifies key challenges: headless login, least privilege for non-deterministic systems, and compliance tracking. Grinich presents four emerging patterns—persona shadowing, delegation chains, capability tokens, and human-in-the-loop escalation—and references standards like OAuth, UMA, GNAP, OIDC for agents, and verifiable credentials. He predicts a shift from 95% human traffic to 95% agent traffic, calling for middleware trust boundaries and urgent collaboration on agent identity standards.
Powered by PodHood