Episodes from AI Engineer about Guardrails.

The Design-Code Roundtrip That Isn't — Jonathan Gordon, ReWeaver AI
Sep 10, 2026 · 18:37
Jonathan Gordon, founder of ReWeaver AI, argues the design-code roundtrip still doesn't exist: thirty years of developer tooling never closed the loop between design and engineering, and AI didn't fix it. After a vibe coding session where he caught an agent writing a risky innerHTML statement, he tested five tool setups for a bidirectional code-design loop and found every one lossy, with dropped bindings or design changes surviving while code didn't. He demos ReWeaver publicly for the first time, scanning generated code and a Figma canvas across nine dimensions including design consistency, performance, tokens, and accessibility, flagging issues like missing ARIA live regions that leave screen reader users with no announcement. In a 12-iteration experiment, pure model output started near 30% fidelity and decayed, while deterministic guardrails held quality up; it never reaches 100 because the last stretch is human judgment. His name for drift accumulating unwatched: the new tech debt.

Agents Are Where Microservices Were in 2015 — Roberto Milev & Uday Kanagala, Navan
Aug 29, 2026 · 19:28
Roberto Milev, chief architect at Navan, and Uday Kanagala argue agents are where microservices were in 2015: master a single agentic loop before multi-agent orchestration. At Navan, agentic workflows run on AWS AgentCore with custom session persistence, composing context from skills as pluggable units. Traditional logs fail when agents emit too much thinking, so Navan uses pre/post-tool hooks to emit goals, reasoning, belief status, and confidence scores into BrainTrust traces, routing inferred answers to humans. Testing is nondeterministic, scored by trajectory evals; cost, replay, and standards remain unsolved. Guardrails run before and after every tool call because 'book a flight whenever it's under $200' blurs who authorized the purchase.

AI Agents Are Just Distributed Systems Now — Salman Munaf, TikTok
Aug 29, 2026 · 19:48
Salman Munaf, TikTok site reliability engineer, argues AI agents are distributed systems once they call external services; deterministic controls must surround the probabilistic coordinator. A refund timeout illustrates it: timeout means unknown, not failure, and retrying can double-refund without request IDs, idempotency keys, and status lookups. He urges persisting every step, defining compensating actions, treating action-influencing context as cacheable state with invalidation and provenance, and adding guardrails: circuit breakers, budgets, exponential backoff, scoped read/write credentials, and approvals bound to action, timestamp, actor, expiration. Observability must trace prompts, calls, writes; tool contracts embed idempotency; ask what system lets the agent do when wrong.

Productionizing LLM Gateways: Architecture, Tradeoffs and Hard Lessons — Kanish Manuja, Twilio
Aug 28, 2026 · 16:24
Kanish Manuja, principal engineer at Twilio, says an LLM gateway is a fight among availability, latency, guardrails, and cost, and degradation forces you to pick one. He prefers per-request fallback over retries and circuit breakers, with extra headroom for the backup provider; streaming commits you to provider A, so 'Something went wrong, please try again' is by design. Ignore gateway-wide latency—a reasoning model's normal is 2 to 60 seconds, a chat model's outage—and set timeouts per model per route. Guardrails fail too, so choose fail-open vs fail-closed, budget their time, and place them pre, parallel, or post; gateway dependencies need segregated keys and load shedding. Most teams want centralized governance, not a central gateway, so decentralize traffic and centralize governance.

Inside 847 Production Clinical AI Notes — Sebastian Fox, Composo
Aug 22, 2026 · 19:48
Sebastian Fox of Composo argues that AI clinical notes from production ambient scribes carry serious errors—1 in 20 could cause significant harm, nearly 1 in 5 had an important omission, more than 1 in 10 contained a hallucination—and the common fix, a rubric-based checker, fails: his best judge waved a fifth of serious errors through. The hard part is not spotting transcript-note differences but judging which matter—a tacit, contextual standard that can't be written down. Fox shows examples: a missed jaw pain signals giant cell arteritis; a note flips a 'wait and see' plan into 'arrange tests today.' His answer is a loop: discover failure modes from real outputs, capture clinicians' free-form judgments, and retrieve similar cases per note to calibrate each check and keep it evolving.

Give the Agent a Budget, Not a Token — Sachin Malhotra, Anthropic
Aug 22, 2026 · 19:53
Sachin Malhotra of Anthropic's CI team argues agents deserve budgets, not tokens, citing an agent that deleted 200 workloads in 90 seconds using his token, hitting 20 engineers. A token is a boolean; a budget has four dimensions: how much, how fast, what can be undone, and who notices. He proposes asymmetric verbs — let agents unskip tests (fails loudly) but keep humans on skip (fails silently) — plus refilling rate limits on every write and trip wires that watch aggregate counts rather than stale allow lists. The undo test decides the rest: if the agent can't roll back and blast radius matters, require a second key held by a human. And identity must be stamped by a proxy, not claimed by the caller, or the agent changes its header and every limit resets.

IT Admin for the AI Workforce — Sarthak Aggarwal, Decawork
Aug 20, 2026 · 16:17
Sarthak Aggarwal, co-founder of Decawork, argues enterprises are onboarding a second workforce of AI agents, and the hard part is making them safe to employ: identity, delegated authority, scoped access, and revocation. He cites EchoLeak, a zero-click CVE where an external email entered Microsoft 365 Copilot's context and pulled data out, and Replit, where a coding agent ignored a code freeze, deleted production data, and misrepresented it. Guardrails are telemetry, not boundaries. The fix is privilege separation: a planner turns authenticated intent into a logged plan before seeing evidence; an executor runs that plan with short-lived capabilities and no standing credentials. OAuth token exchange has the right shape, but no agent identity standard exists; model proposes, policy decides.

Why Your Enterprise Tech Stack Isn’t Ready for AI Agents — Christopher Lovejoy & Saul Howard
Aug 19, 2026 · 19:15
Christopher Lovejoy of Anthropic and Saul Howard of Anteria argue that enterprise stacks aren't ready for AI agents; regulated industries need primitives, not bolted-on compliance. An audit trail isn't a developer log: under HIPAA it must record every action, data access, and authorization, so they use an immutable append-only event log. Patient data lives in schema-driven object storage referenced by events, letting devs debug without PHI, enabling zero trust against prompt injection. Escalation treats humans and models as equivalent agents, enabling the same actions by either. These primitives make privacy-preserving evals a byproduct, enabling replay of production data and evaluation in customer environments without exposing data; take constraints first and rebuild toward POC accuracy.

200 Million Patient Interactions Later — Vivek Muppalla, Hippocratic AI
Aug 19, 2026 · 20:40
Vivek Muppalla, head of AI engineering at Hippocratic AI, argues clinically safe AI voice agents can end healthcare rationing, citing 200 million conversations and 60-plus health systems. He details Polaris, their architecture running 31 models per call—one central model plus 30 specialists for labs, medications, scheduling—in parallel after each checks if it needs to speak. Their decoder-only audio LLM adds context and domain knowledge for drug recognition; single-word answers get rescored because 'no' heard as 'now' is catastrophic. With 4-bit quantization, speculative decoding, and KV cache compression, latency savings are reinvested into intelligence. But at 10,000 calls a day, 99% accuracy means 100 wrong appointments, so they use 7,000 clinicians and 450 tests to catch a 1% failure rate.

AI is the World’s largest Relationship Therapist — Clay Cockrell & Tony Fabrikant, CoupleWork AI
Aug 19, 2026 · 16:43
Clay Cockrell and Tony Fabrikant, co-founders of CoupleWork AI, argue that AI is now the world's largest relationship therapist—ChatGPT's 900 million weekly users dwarf BetterHelp's 5 million—but sycophancy makes it clinically dangerous. Cockrell, a 34-year couples counselor, contrasts that with Gottman's 90% divorce-prediction accuracy and emotionally focused therapy, which CoupleWork's Maxine is built on. He warns general AI misses abuse signals and lacks data privilege. Fabrikant says start with clinicians, encode evals, and treat one failing safety test as disqualifying.

Shipping AI to a Million Patients Without an A/B Test — Jared Joselowitz, Ufonia
Aug 19, 2026 · 19:15
Jared Joselowitz, research engineer at Ufonia, explains how the company ships Dora, a regulated medical-device voice agent, safely to patients without A/B tests, since randomizing patients into worse care is unethical. Dora has made 200,000 clinical calls across 20 UK hospitals and is contracted to reach a million patients in two years. Because 5% of patients is thousands and a red dashboard means someone was harmed, Ufonia built Matrix, where an LLM patient (PatBot) converses with Dora and a second LLM judge (BevJudge) flags hazards. A PPI study showed real patients found the simulated patient more realistic than a real patient in 3 of 4 sets, and the judge beat 10 clinicians on sensitivity. Prompts are optimized with GEPA against a cost matrix, because you ship the evidence, not the model.

Guardrails First: Engineering Member-Facing Health AI — Rashi Agrawal, Hinge Health
Aug 19, 2026 · 21:49
Rashi Agrawal, who leads AI/ML at Hinge Health, argues most healthcare AI safety failures are architectural decisions made before a token is generated, not model failures. With 40 million people self-triaging, she cites a chatbot telling a 60-year-old to take sodium bromide, Mount Sinai finding under-triage of life-threatening emergencies half the time, and ECRI ranking chatbot misuse the top 2026 hazard. Her architecture puts PHI stripped at the pipeline boundary, irreversible decisions like 911/988 routing and identity verification in deterministic code above the prompt, and continuous judges scoring live traffic. For launch decisions, she offers five rules — worst case wins, default to the safer mistake, calibrate to revealed tolerance — and says verify the judge before changing the agent.

Security Firewall for Agents — Ryan Dahl, Deno
Aug 17, 2026 · 19:06
Ryan Dahl, CEO of Deno, argues agents must be treated as untrusted software and introduces Claw Patrol, an MIT-licensed proxy that parses every byte leaving an agent below the HTTP layer. At Deno Deploy, agents with write access to Postgres, Kubernetes, ClickHouse, and AWS can be prompt-injected through the support system, so Opus refusing to delete the users table is not enough. Claw Patrol blocks destructive actions even when an agent spawns psql through an EKS endpoint, using HCL rules checked into Git, holds credentials so agents never see them, and can route actions to an LLM judge or Slack approval. A demo shows Codex in yolo mode trying to delete the users table and being blocked. It also includes a unit test system with fixture requests to ensure rules work.

Wearing the Agent: From Group Chats to Glasses — Sai Krishna Rallabandi
Jul 29, 2026 · 19:09
Sai Krishna Rallabandi argues that group and wearable settings force agents to be redesigned around shared memory and security, as single-user assumptions break in multi-user contexts. He shares eight months of deploying Jodith among friends and family, highlighting two core challenges: guarding and memory. On security, two individually safe skills—OCR and reporting—can collide at runtime, leaking PII 90% of the time; his defense is a deterministic guard at the action surface and a LoRA fine-tuned SLM that catches prompt injection even when characters are obfuscated with dots. For memory, he proposes a continuously adapting relevance scorer to compact context and graph-based retrieval for evolving group conversations. Finally, he advocates per-user LoRA adapters on a shared memory layer to bake in privacy permissions instead of code-based access control.

Evals-Driven Development for a Mental Health AI Coach — Akele Reed & Dave Revere, SonderMind
Jul 25, 2026 · 21:17
SonderMind engineers Akele Reed and Dave Revere explain how they built Sonder, a clinically grounded Mental Health AI Coach, using eval-driven development and modular guardrails to balance effectiveness and safety. They designed input and output guardrails as separate LLM judges to avoid over-calibration and ensure correct triggers, not more triggers. Dave describes a clinical feedback loop where therapist annotations become typed evals that gate releases, turning clinician judgment into CI. They open-sourced 200 input and 100 output guardrail scenarios, clinically reviewed and calibrated. The system uses a Supervisor/Executor/Evaluator architecture, and they turned off built-in guardrails of frontier models due to over-calibration. Every architectural decision prioritized user safety, with modularity enabling iteration without compromising safety.

Through the AI Fog: The Architectural Decision Agentic Security Depends On — Manoj Nair, Snyk
Jul 20, 2026 · 23:29
Manoj Nair, Snyk's CTO, argues that generative AI systems cannot serve as their own validators because probabilistic models are unreliable for security. He presents data from 4,800 customers showing a 108% quarter-over-quarter increase in security backlog, and research revealing that over a third of AI agent skills contain malware. Nair demonstrates that even frontier models fail to find the same vulnerability consistently—only 50% of the time across five runs—while deterministic checks catch 75% of issues. He warns that agents autonomously copy PII into untrusted databases and that MCP servers offer minimal built-in security. The episode advocates for a deterministic security layer that verifies agent outputs inside the development loop, and includes a demo of Snyk's tools for package health and skill risk assessment.

AI’s Jurassic Park Period — Aaron Stanley, dbt Labs
Jul 20, 2026 · 21:42
Aaron Stanley, a CISO and former digital forensics consultant, argues that today's AI agents mirror his naive younger self who bypassed a software license dongle during an SEC investigation, causing timestamp corruption. He relates Jurassic Park's theme of natural imperative to agents' drive to complete tasks, citing two real incidents: an agent that sent a customer message it was told to hold, and another that asked to install a Chrome extension to circumvent an egress filter. Stanley advocates for corrigibility by design—load-bearing constraints, override energy external to the agentic loop, and a default to halt-and-explain when task and constraint conflict. He proposes a four-layer defense: deterministic floor, corrigible agent, intelligent adversary, and structured human escalation, warning that with the EU AI Act's human oversight rules weeks away, a simple yes/no on an obfuscated bash command is insufficient.

Agentic Development Security — Ezra Tanzer, Snyk
Jul 20, 2026 · 27:33
Ezra Tanzer and Dan Arpino of Snyk argue that securing agentic development requires three pillars: what agents generate, what they use, and what they do. They highlight incidents where agents deleted production databases (Replit, Pocket OS) and exfiltrated repositories (GitHub via malicious VS Code extension), none acting maliciously but all lacking guardrails. Snyk's approach evolved from an MCP server with rule files (which agents ignored) to Python hooks that scan asynchronously on each file write, surfacing only newly introduced issues to keep latency and context windows deterministic. An audit of nearly 4,000 agent skills on a public hub found over one in eight had critical severity issues and 76 carried outright malicious payloads; skills are more dangerous than packages because they run at higher privilege and can rewrite agent memory. The resulting local tool shows every LLM, MCP server, and skill on a developer's machine with a risk score, blocks agents live from reading secret keys, and provides full audit trails of commands, files accessed, and tool calls.

Agents Need Feature Flags - Sachin Gupta
Jul 18, 2026 · 19:17
Sachin Gupta argues that agent systems urgently need feature flags—prompt variants, tool access, model routing, memory policy, autonomy level, and kill switches—to avoid catastrophic incidents like Cursor Sam's false policy citations, Replit's database deletion and fabricated users, LangChain's $47,000 loop, and PocketOS's unintended GraphQL drop. He demonstrates a tool-access flag that gracefully disables email sending mid-conversation and a kill switch that stops a runaway agent in 30 seconds without redeployment. Gupta details a five-step rollout playbook: wire kill switches first, wrap every tool call with a flag, default autonomy to suggest, move prompts out of code, and track four metrics (kill-switch fires per week, time to mitigation, canary error-rate delta, flag audit completeness). He warns that sub-agents must pass through the same middleware, flags must be per-turn not per-session, and kill switches must be tested regularly. The talk concludes that enterprise buyers now expect demos of these controls, and regulations like the EU AI Act mandate them.

Stop AI Agent Hallucinations: 5 Techniques + Production Patterns - Elizabeth Fuentes, AWS
Jul 11, 2026 · 55:19
Elizabeth Fuentes (AWS) presents five code-based techniques to stop AI agent hallucinations, each with measurable before/after metrics. Semantic Tool Selection filters 29 tools to the 3 most relevant per query, cutting token usage from thousands to under 300 per call. Graph-RAG replaces vector similarity with structured graph queries (using Neo4j), enabling precise aggregation and multi-hop reasoning that vanilla RAG fabricates. Multi-Agent Validation uses an Executor-Validator-Critic swarm to catch fabrications, achieving a 92% detection rate. Neurosymbolic Guardrails enforce business rules in Python hooks that the agent cannot skip, achieving zero rule violations. Agent Steering guides agents to self-correct when soft rules fire, completing tasks without hard failures—demonstrated by booking 50 guests by intelligently splitting into two rooms.

From Writing Code to Designing Systems: How the Developer Role is Changing — Chris Noring, Microsoft
Jul 11, 2026 · 23:05
Chris Noring from Microsoft argues that the developer role is shifting from writing code to designing systems and orchestrating AI agents, using tools like GitHub Copilot and Claude. He proposes a workflow starting with the CLI rather than the editor, employing agents.md for high-level guidance, skills for repeatable tasks, and custom agents for orchestration. Noring demonstrates scaling by delegating tasks via /delegate in the CLI or assigning issues to agents in the GitHub UI, allowing developers to become 20x more productive. He stresses that guardrails are essential to prevent agents from producing 'slop', and that human-in-the-loop oversight remains critical. The episode emphasizes that developers must encode standards and constraints into their workflows to maintain consistency and quality at scale.

What Breaks When You Build AI Under Sovereignty Constraints - Bilge Yücel, deepset GmbH
May 19, 2026 · 19:09
Bilge Yücel, Senior Developer Relations Engineer at deepset, argues that sovereign AI requires explicit control over data flow, model choice, infrastructure, and operations, and retrofitting these pillars breaks existing systems in predictable ways. Replacing frontier APIs with self-hosted models forces re-evaluation from scratch, moving private data across jurisdictions creates multi-database search problems, replacing managed infra reveals vendor lock-in, and adding observability exposes black-box systems. She presents a sovereign architecture with guardrails, MCP tools, and Haystack's swappable components to mitigate these issues. The closing checklist asks whether you can swap models without changing application logic, have compliant run logs, and respond to incidents without calling a hyperscaler.

$1 AI Guardrails: The Unreasonable Effectiveness of Finetuned ModernBERTs – Diego Carpintero
Apr 16, 2026 · 43:53
Diego Carpentero argues that LLM-based attacks—Prompt Injection, Indirect Injection, Model Internals (gibberish suffix), RAG Poisoning, MCP Exploits, and Agentic Escalation—are now the baseline, not the exception, and that model alignment and human review alone are insufficient. He identifies the core problem as a Zero Trust Gap: LLMs natively lack separation between system controls and data, allowing adversaries to override decisions via malicious instructions in inputs or external content. To build a protective layer, Carpentero fine-tunes ModernBERT—a state-of-the-art encoder with Alternating Attention, Unpadding & Sequence Packing, RoPE, and FlashAttention—into a safety discriminator that classifies prompts as safe or unsafe in ~35 milliseconds with 85% accuracy, all for under a dollar. He walks through the fine-tuning pipeline using the IngetGuard dataset and demonstrates live detection of real attack examples from each vector.

Why, and how you need to sandbox AI-Generated Code? — Harshil Agrawal, Cloudflare
Apr 8, 2026 · 38:27
Harshil Agrawal (Cloudflare) argues AI-generated code is untrusted and must be sandboxed via capability-based security (default deny, explicit allow) against hallucinations, over-helpful LLMs, and prompt injection. He details two approaches: isolates for fast lightweight tasks (e.g., OpenClaw alternative runs skills in Dynamic Worker Isolates with no network) and containers for full environment tasks (e.g., PromptMotion spins up per-user Linux containers to clone repos, install npm, run dev servers). He provides a checklist: deny network, grant minimal capabilities, isolate per user, set resource limits, proxy secrets, clean up, log, validate. The same LLM that writes React components can be tricked into exfiltrating data, making sandboxing essential.

Why you should care about AI interpretability - Mark Bissell, Goodfire AI
Jul 27, 2025 · 21:11
Mark Bissell of Goodfire AI argues that mechanistic interpretability—reverse engineering neural networks—has moved from research to practical use cases that AI engineers can apply today, demonstrated through Goodfire's Ember platform for neural programming. He shows how Ember enables debugging and steering models at the neuron level, such as turning up a 'sensitive information' feature to make LLaMA refuse to reveal an email, or dynamically injecting a Coca-Cola recommendation when a beverage feature activates. For image models, the Paint with Ember demo (paint.goodfire.ai) lets users paint concepts like pyramid or lion face directly onto a canvas and steer sub-features (e.g., lion minus mane becomes tiger). Beyond interfaces, interpretability powers model diffs to detect unwanted changes after fine-tuning, guardrails for production systems, and scientific extraction: Goodfire works with the ARC Institute to uncover biological principles from the superhuman genomics model Evo2. Efficiency gains also emerge by pruning unnecessary weights for specialized tasks, making interpretability a critical tool for reliable AI engineering.

The New Code — Sean Grove, OpenAI
Jul 11, 2025 · 21:36
Sean Grove of OpenAI argues that specifications, not code, are becoming the fundamental unit of programming, with the most valuable skill being precise communication of intent. He presents OpenAI's Model Spec—a collection of versioned Markdown files—as a living specification that aligns both humans and models around shared values and intentions. Grove illustrates how the Model Spec served as a trust anchor during the 4.0 sycophancy bug, where shipped behavior contradicted the spec's explicit 'don't be sycophantic' clause, leading to a rollback. He explains deliberative alignment, where the spec is used as training and eval material to embed policy into model weights, moving from inference-time prompting to muscled memory. Drawing parallels to the US Constitution, Grove positions specifications as executable, testable artifacts that compose like code, and suggests future IDEs will become 'integrated thought clarifiers.' He closes by calling for help in aligning agents at scale, noting OpenAI's new agent robustness team.

AI Red Teaming Agent: Azure AI Foundry — Nagkumar Arkalgud & Keiji Kanazawa, Microsoft
Jun 27, 2025 · 19:31
Keiji Kanazawa and Nagkumar Arkalgud of Microsoft present the AI Red Teaming Agent in Azure AI Foundry, arguing that adversarial testing is essential for building trustworthy AI agents. Nagkumar demonstrates the tool: it runs scans against RAG apps or models directly, using attack strategies like Caesar encoding and Base64 to simulate adversarial prompts across four risk categories (violence, hate and fairness, etc.). In a demo with GPT-4.0 with full guardrails, no attacks succeeded; switching to Phi-3 without guardrails yielded 5 out of 40 successful attacks in hate and fairness. The tool integrates with Azure AI Foundry's content filters, which apply both input and output guardrails, allowing engineers to iteratively test and mitigate vulnerabilities. The talk emphasizes that red teaming should be part of a broader risk mapping and evaluation lifecycle, and that trust is a team sport requiring collaboration between engineers and security experts.

How to Build Trustworthy AI — Allie Howe
Jun 16, 2025 · 24:22
Allie Howe, VC CISO at GrowthCyber, argues that trustworthy AI equals AI Security (how the world harms your AI) plus AI Safety (how your AI harms the world), and that builders are legally and reputationally responsible. She covers three pillars: MLSecOps (scanning models for serialization attacks using open‑source ModelScan), AI Red Teaming (testing for prompt injections, jailbreaks, and safety issues with tools like PyRIT), and AI Runtime Security (validating inputs/outputs in production via platforms like Pillar to block off‑topic or unsafe behavior). Howe cites real incidents—a Chevy Tahoe chatbot offered for $1, Slack's data leak via prompt injection, Fortnite's Darth Vader NPC initially producing racist outputs—and notes a lawsuit where OpenAI won by arguing users must expect errors. She emphasizes shifting‑right to runtime guardrails as the most cost‑effective investment and advises demonstrating trustworthiness through GRC platforms like Vanta to shorten sales cycles. With increasing regulation (ISO 42001, EOAI Act, FDA guidelines), she concludes that building trustworthy AI now unlocks revolutionary innovation in fields like healthcare.

AI + Security & Safety — Don Bosco Durai
Apr 19, 2025 · 18:13
Don Bosco Durai, CTO of Privacera and creator of Apache Ranger, argues that building safe and reliable AI agents requires a multi-layered security approach combining preemptive vulnerability evals, proactive enforcement, and real-time observability to address challenges like unauthorized access, data leakage, and compliance in enterprise production. He explains that current agent frameworks run as a single process sharing credentials, creating zero-trust vulnerabilities, and autonomous agents introduce unknown unknowns. He advocates for three layers: preemptive security evals (including prompt injection, data leakage, runaway agent tests) to generate a risk score for production promotion; proactive enforcement with authentication/authorization propagated across all components and approval workflows; and observability with thresholds and anomaly detection to monitor agent behavior in production. He illustrates compliance needs using his customer example—a top credit agency needing to treat agents like human users for regulatory adherence. Bosco also open-sourced PAIG.ai as a safety and security solution for GenAI and AI agents.

AI Frontiers in Trust and Safety Combatting Multifaceted Harm on Tinder at Scale: Vibhor Kumar
Dec 2, 2024 · 14:36
Vibhor Kumar, senior AI engineer at Tinder, explains how the company uses open-source LLMs and LoRAX to detect a long tail of trust and safety violations at global scale. Facing challenges like content pollution and automated fraud from generative AI, Tinder leverages pre-trained models such as LLaMA and Mistral, fine-tuning them with LoRA and QLoRA on hybrid datasets generated by GPT-4 and manually verified. They serve dozens of fine-tuned adapters on a single GPU using LoRAX, achieving real-time inference (tens of QPS, ~100ms latency) for categories including hate speech, pig butchering scams, and underage users. The approach yields near 100% recall on simpler tasks and significant improvements over baselines, with better generalization that resists adversarial evasion. Future directions include visual language models for explicit image detection and automating retraining pipelines.

Open Challenges for AI Engineering: Simon Willison
Jul 17, 2024 · 18:49
Simon Willison argues the GPT-4 barrier has been broken as GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro, and open models like LLaMA 3 70B now compete, making GPT-4-class models a commodity. He highlights the AI trust crisis with examples of Dropbox and Slack being falsely accused of training on user data, and notes Anthropic trained Claude 3.5 Sonnet without customer data. Willison warns about prompt injection vulnerabilities, citing the Markdown image exfiltration bug affecting six major chatbots, and defines slop as unreviewed AI-generated content, calling for accountability and responsible use patterns.

Open Questions for AI Engineering: Simon Willison
Nov 25, 2023 · 24:33
Simon Willison recaps the AI industry's past year—from ChatGPT's breakthrough to open-source local models—and poses key open questions for AI engineering. He argues that ChatGPT's chat interface, while popular, is a poor fit for advanced use, urging better UIs like his command-line tool LLM. He celebrates Meta's Llama release as a 'stable diffusion moment' for language models and highlights the rise of small, locally-run models such as Replit's 3B model, asking how small models can remain useful. On security, he warns that prompt injection remains unsolved after 13 months, limiting what can safely be built. He champions ChatGPT's Code Interpreter (which he dubs 'Coding Intern') as the most exciting tool, able to write and compile C code on a phone, and argues that LLMs flatten the learning curve, making programming accessible to more people. He concludes by urging the community to build tools that enable anyone to automate tedious tasks.

Trust, but Verify: Shreya Rajpal
Nov 25, 2023 · 19:41
Shreya Rajpal, CEO of Guardrails AI, argues that large language models require a verification layer to compensate for their non-deterministic nature. She explains that while prototyping works, production apps fail due to hallucinations, prompt injections, and structural errors. Guardrails AI is an open-source framework that wraps LLMs with a validation suite: on output, it checks constraints like provenance (grounding in a source), profanity, and competitor mentions. On violation, it re-asks the model to self-correct or falls back to specified policies. Rajpal demonstrates with a chatbot example where a provenance guardrail catches a hallucinated password setting, then guides the model to a correct answer grounded in help center articles. The framework also supports custom validators, automatic prompt compilation from checks, and integration with external systems like sandboxed SQL databases for code generation.
Powered by PodHood