Page 16 of 23

Small AI Teams with Huge Impact — Vik Paruchuri, Datalab
Jul 15, 2025 · 17:36
Vikas Paruchuri, CEO of Datalab, argues that small teams can outperform large ones, sharing how his team of 4 achieved 40k GitHub stars, 7-figure ARR, and 5x revenue growth since January by training state-of-the-art models like surya OCR 3. Drawing from his experience scaling DataQuest to 30 people and then cutting to 7, he explains that layoffs increased productivity due to fewer specialists, less meeting overload, and more senior generalists. He advocates hiring senior generalists who work across the stack, using simple tech (e.g., server-rendered HTML over React), and minimizing bureaucracy with high trust and in-person collaboration. By leveraging AI to handle low-leverage tasks and training models to replace forward-deployed engineers, Datalab maintains tight feedback loops and fast iteration. Paruchuri emphasizes scaling productivity, not headcount, and offers a three-step hiring process: a peer chat, a paid 10-hour project, and a culture fit check, with a 40% hire rate.

Rethinking Team Building: how a 30-person Startup serves 50 Million Users — Grant Lee, Gamma
Jul 15, 2025 · 18:06
Grant Lee, CEO of Gamma, explains how his 30-person team serves 50 million users by ditching blitzscaling for lean teams of generalists and player coaches. He argues that hiring generalists—like his head of design who codes, researches UX, and mentors—enables rapid adaptation. Player coaches, such as engineering leads who still write code, make fast technical trade-offs without top-down mandates. Scaling with brand and culture, Gamma invests in a living culture deck and three weekly all-hands meetings to maintain tribal knowledge. In Q&A, Lee advises doing the job yourself before hiring for non-engineering roles, probing for high agency by asking candidates to drill into problem layers, and using work trials (five successes, high failure rate without) to avoid mismatches. He also wishes they had prioritized infrastructure for experimentation earlier given AI's speed.

Building a 10 person unicorn - Max Brodeur-Urbas, Gumloop
Jul 15, 2025 · 12:03
Max Brodeur-Urbas, founder of Gumloop, explains how his company scaled to millions in ARR as a team of two that raised a Series A and grew to only nine people, by being super picky in hiring, using product-led hiring where customers like those from Instacart and Webflow join the team, and requiring work trials such as hacking together in Airbnbs. They eliminate almost all meetings to give engineers deep focus time and automate every internal process with Gumloop itself, from customer research reports to chatbot monitoring. Culture-wise, they balance intense 45-minute shipping challenges with fun retreats and a public company handbook. Brodeur-Urbas argues that every hire must be a no-brainer, and small teams can outpace larger ones by avoiding meetings and leveraging AI tools.

Using OSS models to build AI apps with millions of users — Hassan El Mghari
Jul 15, 2025 · 18:47
Hassan El Mghari, a software engineer at Together.ai, explains how he builds open source AI apps that attract millions of users, including roomGPT.io (2.9 million users), restorePhotos.io (1.1 million), Blinkshot.io (1 million visitors), and LlamaCoder.io (1.4 million visitors). He details his simple four-step tech stack: user input, a single API call to an open source model on Together.ai, storing results in a database, and displaying output. His process emphasizes keeping ideas simple enough to describe in five words, spending 80% of time on UI, launching early, and incorporating the latest models like Fluxionel for virality. He advises naming apps with short, memorable names, making them free and open source to encourage sharing, and iterating rapidly—most apps fail, but the key is to keep building. He funds compute through sponsorships from Together.ai and partners like Neon and Clerk, which provide free services for his open source projects.

Bolt.new: How we scaled $0-20m ARR in 60 days, with 15 people — Eric Simons, Bolt
Jul 15, 2025 · 17:33
Eric Simons, CEO of Bolt.new, explains how his team of less than 20 people scaled the company from $0 to $20M ARR in just 60 days, achieving the second fastest product growth in history by staying lean and focusing on high-impact decisions. He describes building a remote team with high context and agency, resisting pressure to hire more during the 2021 boom. They relied on 'things that don't scale' like weekly office hours to build user love, and leveraged AI support tools like PeraHelp to handle 90% of tickets. Community initiatives such as a Guinness World Record hackathon (over 80,000 participants) further amplified growth without adding headcount. Simons emphasizes taking consistent shots on goal and making independent bets rather than following VC trends.

Prompt Engineering and AI Red Teaming — Sander Schulhoff, HackAPrompt/LearnPrompting
Jul 14, 2025 · 2:01:05
Sander Schulhoff, creator of Learn Prompting and HackAPrompt, argues prompt engineering remains vital despite claims of its demise, drawing on his systematic review of 1,500+ papers for 'The Prompt Report.' He covers advanced techniques including chain-of-thought, decomposition, ensembling, and few-shot prompting, noting that role prompting is ineffective for accuracy-based tasks and that example ordering can swing performance by 50%. He then explains AI red teaming, distinguishing jailbreaking from prompt injection, and warns that system prompts and guardrails cannot prevent attacks—even simple obfuscation like base64 or typos still works. He highlights the critical unsolved problem of agentic security, where agents with real-world actions are easily tricked, and introduces a live competition at the conference to gather more attack data. His key takeaway: AI security is fundamentally harder than classical cybersecurity because 'you cannot patch a brain.'

Survive the AI Knife Fight: Building Products That Win — Brian Balfour, Reforge
Jul 14, 2025 · 14:10
Brian Balfour, CEO of Reforge, argues that winning in today's AI knife-fight requires answering 'What do I build and why will it win?' by focusing on proprietary data, unique functionality, and unmet customer needs rather than building custom AI. He illustrates with Granola, which entered a crowded AI note-taker market by understanding users wanted help taking better notes, not full automation, and assembled off-the-shelf AI (DeepGram, Anthropic, OpenAI) with unique data (user notes plus transcription) and functionality (Mac app, calendar integration) to create a competitive edge. Balfour warns competitive advantages now last only 2-3 weeks, so teams must sequence smaller moats continuously, each buying time to execute faster. The talk emphasizes treating AI as Lego blocks—assembling pre-trained models, data, and product superpowers into a system that spins a data flywheel.

Automating Escrow with USDC and AI - Corey Cooper, Circle
Jul 14, 2025 · 58:18
Corey Cooper, Head of DevRel at Circle, demonstrates how USDC stablecoins and AI agents can automate escrow by combining smart contracts with LLM-based task verification. He explains that USDC's programmability enables near-instant settlement, chargeback-free transactions, and always-on payments ideal for agent-to-agent payments. The demo shows an escrow agent app that uses OpenAI to parse PDF contracts, deploys Solidity escrow contracts via Circle APIs on Base Sepolia, and allows an AI agent to verify deliverables (e.g., image quality) and trigger payout from a locked smart contract. Cooper discusses human-in-the-loop safeguards, cross-chain deposits via Circle's Cross-Chain Transfer Protocol, and future possibilities like multi-sig agent verification. He highlights that while full autonomy is not yet reliable, combining USDC with AI can streamline payment operations.

How LLMs work for Web Devs: GPT in 600 lines of Vanilla JS - Ishan Anand
Jul 13, 2025 · 1:41:34
Ishan Anand shows that GPT-2 small implemented in 600 lines of vanilla JavaScript makes LLMs understandable for web developers without ML backgrounds. He explains tokenization via byte-pair encoding, 768-dimensional embeddings representing semantic meaning via co-occurrence, and the Transformer's attention mechanism that lets tokens share context. The multi-layer perceptron learns next-token prediction through backpropagation, while the language head converts embeddings to token probabilities using softmax. Anand demonstrates each step—tokenization, embedding lookup, positional encoding, attention, MLP, and output—in a browser debugger, and notes that GPT-2's architecture underpins ChatGPT, with innovations like scale, supervised fine-tuning, and RLHF. The workshop provides an intuitive mental model of Transformers, turning perceived AI magic into understandable machinery.

[Workshop] AI Pipelines and Agents in Pure TypeScript with Mastra.ai — Nick Nisi, Zack Proser
Jul 12, 2025 · 1:51:14
This hands-on workshop introduces Mastra.ai, a TypeScript framework for building production AI agents and pipelines, and demonstrates how to build an AI meme generator using composable workflows, tools, and agents. Hosts Nick Nisi and Zack Proser walk through creating a multi-step workflow: extract user frustration, find a base meme via ImageFlip, generate captions, and publish the meme at a stable URL. They show how to chain steps with Zod schema validation for deterministic output, then wrap the workflow in an agent that accepts natural language requests. The workshop also covers MCP (Model Context Protocol) and a live demo of MCP.shop for ordering a shirt, emphasizing local iteration with Mastra's playground, built-in memory, and evaluation tools. Real-world patterns include building internal AI assistants for data cleaning, email drafting, and document summarization with minimal code.

AI Engineering with the Google Gemini 2.5 Model Family - Philipp Schmid, Google DeepMind
Jul 11, 2025 · 1:44:51
In this workshop, Philipp Schmid from Google DeepMind demonstrates AI Engineering with the Gemini 2.5 model family, focusing on using Gemini 2.5 Flash via a free API tier for hands-on coding tasks including text generation, multimodal processing of images, audio, and PDFs, function calling with structured outputs, and integration with MCP servers. The session covers setting up API keys in AI Studio, uploading files via the Files API (free for 1 day), and controlling thinking budgets (0–24,000 tokens) to manage cost and reasoning depth. Schmid shows how Gemini natively processes videos at 1 frame per second for accurate timestamp extraction and how PDFs are handled by combining OCR text with image understanding. He introduces native tools like Google Search with grounding metadata, code execution, and URL context, and explains how MCP servers can be used seamlessly with the Gemini SDK for tool calling. The workshop also covers parallel vs sequential function calling, the Agent Development Kit (ADK), and the upcoming asynchronous function calling for the Live API, providing a practical path from simple generation to agentic workflows.

The New Code — Sean Grove, OpenAI
Jul 11, 2025 · 21:36
Sean Grove of OpenAI argues that specifications, not code, are becoming the fundamental unit of programming, with the most valuable skill being precise communication of intent. He presents OpenAI's Model Spec—a collection of versioned Markdown files—as a living specification that aligns both humans and models around shared values and intentions. Grove illustrates how the Model Spec served as a trust anchor during the 4.0 sycophancy bug, where shipped behavior contradicted the spec's explicit 'don't be sycophantic' clause, leading to a rollback. He explains deliberative alignment, where the spec is used as training and eval material to embed policy into model weights, moving from inference-time prompting to muscled memory. Drawing parallels to the US Constitution, Grove positions specifications as executable, testable artifacts that compose like code, and suggests future IDEs will become 'integrated thought clarifiers.' He closes by calling for help in aligning agents at scale, noting OpenAI's new agent robustness team.

Production software keeps breaking and it will only get worse — Anish Agarwal, Traversal.ai
Jul 10, 2025 · 18:13
Anish Agarwal and Matthew Schoenbauer of Traversal.ai argue that as AI writes more code, production troubleshooting will become vastly harder, requiring a new approach combining causal machine learning, reasoning models, and agentic swarms to autonomously resolve incidents in minutes. They explain that traditional AI ops generates too many false positives, LLMs can't handle petabyte-scale data, and simple agents depend on deprecated runbooks. Their Traversal AI orchestrates thousands of parallel agentic tool calls to sift through trillions of logs and metrics, identifying root causes and citing observability data. A case study with DigitalOcean shows a 40% reduction in mean time to resolution (MTTR), with the system delivering findings in about five minutes. The episode details how this approach turns frantic incident Slack channels into autonomous, cited root-cause analysis, freeing engineers to focus on system design.

Thinking Deeper in Gemini — Jack Rae, Google DeepMind
Jul 10, 2025 · 18:13
Jack Rae, lead of Gemini Thinking at Google DeepMind, presents thinking as a solution to the fixed test-time compute bottleneck in large language models. He explains that Gemini inserts a thinking stage where the model iterates via reinforcement learning, learning to self-correct and explore multiple strategies. This enables a continuous cost-performance tradeoff via thinking budgets, improving reasoning on math and code. Future work includes Deep Think, which raises USA Math Olympiad performance from the 50th to the 65th percentile by scaling inference compute further. Rae envisions models that, like mathematician Ramanujan, achieve deep, data-efficient reasoning from limited knowledge.

A year of Gemini progress + what comes next — Logan Kilpatrick, Google DeepMind
Jul 10, 2025 · 11:58
Logan Kilpatrick, head of product for Google AI Studio at DeepMind, announces the final update to Gemini 2.5 Pro, which achieves state-of-the-art results on Aider and HLE benchmarks. He details Google's 50x increase in AI inference over the past year, driven by merging research and product teams into DeepMind. The episode outlines Gemini's evolution toward a universal assistant that unifies Google products, with upcoming features including proactivity, native audio and video capabilities (Veo), and smaller models. Kilpatrick also previews developer-focused updates: a SOTA embeddings model, a deep research API, and Veo 3 and Imagine 4 in the API, alongside repositioning AI Studio as a dedicated developer platform.

2025 in LLMs so far, illustrated by Pelicans on Bicycles — Simon Willison
Jul 9, 2025 · 18:30
Simon Willison reviews the past six months of LLM releases — including AWS Nova, Llama 3.3 70B, DeepSeek R1, Mistral Small 3, Claude 3.7 Sonnet, GPT 4.5, Gemini 2.5 Pro, GPT-4o, Llama 4, GPT 4.1, O3/O4 Mini, and Claude 4 — using his 'pelican on bicycle' SVG benchmark to argue that local models have become good enough to run GPT-4 class models on a laptop and that combining tools with reasoning is the most powerful technique in AI engineering, while noting risks like prompt injection and the 'lethal trifecta'. He tracks 30 significant model releases, highlighting that Mistral Small 3 (24B) matches Llama 3 70B's performance, which itself matched the 405B model, enabling local inference. DeepSeek's R1 caused a $500B+ Nvidia stock drop on January 27. GPT 4.1 Nano is the cheapest model yet at a fraction of a cent per pelican. He also examines bugs: ChatGPT's sycophantic 'shit-on-a-stick' incident and Claude 4's tendency to snitch to authorities when given ethical instructions and email tools. Willison concludes that while the pace is accelerating, control over context and security remain critical.

Trends Across the AI Frontier — George Cameron, ArtificialAnalysis.ai
Jul 8, 2025 · 17:52
George Cameron of Artificial Analysis presents multiple frontiers—reasoning, open-weights, cost, speed—across the AI stack, arguing that trade-offs between intelligence, latency, and expense are critical for building applications. Reasoning models like O4 mini high use an order of magnitude more output tokens (72M vs GPT-4.1's 7M) and take over 40 seconds per response versus 4.7 seconds, impacting agentic workflows where 30 sequential calls multiply latency. The open-weights gap has nearly closed, with China-based labs like DeepSeek R1 and Alibaba's Qwen 3 leading. O3 cost roughly $2,000 to run the intelligence index, while GPT-4.1 nano is over 500 times cheaper. Output speeds have jumped from GPT-4's 40 tokens/s in 2023 to over 1,000 on a B200 accelerator. Despite efficiency gains, demand for compute will keep rising due to larger models, reasoning’s extra tokens, and multi-step agents.

Training Agentic Reasoners — Will Brown, Prime Intellect
Jul 7, 2025 · 19:17
Will Brown of Prime Intellect argues that reasoning and agents are fundamentally the same, and reinforcement learning (RL) is the key to advancing both. He explains that RL now works at scale, as shown by DeepSeek's GRPO and OpenAI's o3, and that agentic tasks like tool calling are natural RL environments. Brown warns against reward hacking and emphasizes designing evals that are harder to game than the task itself. He introduces his open-source toolkit 'verifiers' (now on pip) which lets users build trainable agent loops with a simple API, and demonstrates training a 7B Wordle agent in a few turns on just a couple GPUs.

New York Times' Connections: A Case Study on NLP in Word Games — Shafik Quoraishee, NYT Games
Jul 5, 2025 · 18:31
Shafik Quoraishee, a game developer at NYT Games, presents his independent research into applying AI to solve the New York Times Connections word game. He explains that Connections, launched in June 2023 and second only to Wordle with hundreds of millions of plays, challenges AI's abstract reasoning through intentional decoys and overlapping categories. Quoraishee models the game as a graph coloring problem and uses semantic similarity, relational alignment scores across multiple dimensions (orthography, morphology, encyclopedic, etc.), and graph neural networks combined with reinforcement learning to build a solver. His preliminary results show increased solvability for hard puzzles, contrasting with LLMs that may simply recall internet solutions. The work aims to provide a transparent, explainable AI approach for puzzle-solving and potential game development applications.

Claude Code & the evolution of agentic coding — Boris Cherny, Anthropic
Jul 4, 2025 · 18:12
Boris Cherny, creator of Claude Code and Anthropic member of technical staff, argues that the product stays intentionally unopinionated and minimal because the model's coding capabilities are improving exponentially and the right UX remains unknown. He traces programming UX from 1950s punch cards to 1980s Smalltalk's live reload, Eclipse's static-analysis autocomplete, Copilot, and Devin's natural-language paradigm. Claude Code's terminal-first approach works in any terminal, over SSH, and inside VS Code or Cursor, with a GitHub integration that keeps data on user compute and a programmatic SDK for custom UIs. Key tips include using Claude for code-based Q&A (shortening onboarding from weeks to days), teaching it tools via `--help` and Claude MD files, leveraging TDD with visual iteration, and running multiple Claude instances in parallel via terminal tabs or GitHub Actions. Today's launch adds a plan mode triggered by Shift+Tab, which makes a plan and waits for approval before writing code.

12-Factor Agents: Patterns of reliable LLM applications — Dex Horthy, HumanLayer
Jul 3, 2025 · 17:06
Dex Horthy, founder of HumanLayer, presents the 12-Factor Agents framework for building reliable LLM-powered applications, arguing that production-grade agents are primarily deterministic software with targeted LLM steps rather than fully autonomous loops. He distills patterns: own prompts and context windows, treat tools as JSON and code, use small focused agents with three to ten steps, contact humans via tool calls. Horthy emphasizes context engineering—LLMs are pure functions—and shows how to compact errors, unify state, and add pause/resume via APIs. He shares a DevOps agent that became a bash script, and advocates for outer-loop agents. The framework, which gained 4,000 GitHub stars in two months, treats agents as stateless reducers that meet users on any channel, with engineers controlling the inner loop of token and control flow.

MCP Is Not Good Yet — David Cramer, Sentry
Jul 3, 2025 · 16:41
David Cramer, founder and engineer at Sentry, argues that the Model Context Protocol (MCP) is a pluggable architecture for agents, not a simple API overlay, and that it is not yet good but worth experimenting with. He stresses that exposing existing OpenAPI endpoints as MCP tools yields terrible results; instead, developers must design context specifically for agents, returning Markdown rather than raw JSON to improve model reasoning. Cramer details Sentry's MCP server, built in two days using Cloudflare Workers for OAuth 2.1, and notes that remote OAuth is essential for B2B SaaS, while standard I/O introduces security risks. He highlights the need for careful token and cost management—citing a single request that produced 20 API calls to Sentry—and advocates for building dedicated agents (like Sentry's root cause analysis tool) that wrap MCP, giving the provider control over prompts, models, and error handling. Despite ongoing issues like lack of streaming responses and unstable client support, Cramer concludes that MCP's core concepts (plug-ins, agents, tools) are just familiar software patterns with new names.

Your Personal Open-Source Humanoid Robot for $8,999 — JX Mo, K-Scale Labs
Jul 2, 2025 · 19:26
Jingxiang Mo, founding engineer at K-Scale Labs, introduces their open-source humanoid robots: the 5-foot K-Bot (pre-order for $8,999, delivered by October) and the 1.5-foot Z-Bot, both fully open-source from hardware to ML models. The K-Bot uses MIT CHiTA actuators, provides up to 250 TOPS compute, runs an RL-based whole-body controller trained with MJX in 1-2 hours, and offers a Python/Rust SDK (pip install kos) with a digital twin simulator for rapid development. Mo emphasizes modularity—swappable end-effectors and upgradable heads—and positions K-Scale as the first US consumer humanoid robotics company, targeting developers and households. The robots are significantly cheaper than competitors (Tesla Optimus at ~60K, Unitree at 40K) and include VR teleoperation, OTA software updates, and a bimonthly hackathon community.

The Build-Operate Divide: Bridging Product Vision and AI Operational Reality
Jul 2, 2025 · 12:50
Jeremy Silva (Freeplay) and Chris Hernandez (Chime) argue that the biggest challenge in generative AI isn't building prototypes but crossing the 'quality chasm' from v1 to reliable v2 through operational iteration. They explain how the lower barrier to entry and faster iteration speed in Gen AI accentuate the need for high-quality ops, where product quality becomes a direct function of how fast teams move through monitoring, experimentation, and evaluation loops. Chris emphasizes that human-in-the-loop isn't just a safeguard but a feedback engine, and that existing QA and CX teams in operations are already equipped to become 'model shapers'—labeling data, testing prompts, and defining what good looks like. Jeremy introduces the emerging role of the 'AI quality lead,' a systems thinker who can run experiments and evaluations without writing production code. They conclude that scaling Gen AI is an operational and people challenge, not just a technical one, and that embedding quality and human feedback early is the key to building faster and better.

The New Lean Startup — Sid Bendre, Oleve
Jul 1, 2025 · 13:26
Sid Bendre, co-founder of Oleve, explains how his tiny team of four scaled a profitable, multi-product portfolio to $6M ARR by embracing a 'new lean startup' philosophy centered on AI tooling, operating principles, and organizational structure. Oleve's two products hit the top 10 in the App Store Education charts, competing with Duolingo and Photomath, with one reaching #4 in 2024 and #5 in 2025. Bendre details three pillars: operating principles like hiring only 10x generists, a profit-first mentality, and continuous process refinement; organizational structure modeled after Palantir's Harvester-Cultivator split, where Harvesters own product metrics and Cultivators build an agentic operating system; and AI tooling augmentation that turns 10x engineers into 100x. He shares how they repurpose tools like LaunchDarkly for load balancing and on-the-fly infrastructure changes, and invest in blueprints—reusable code templates and shared infrastructure—that enabled a third product to launch in three weeks and become profitable immediately. The episode culminates in Oleve's vision of one-person billion-dollar companies, where a single strategic leader commands clusters of autonomous…

Conquering Agent Chaos — Rick Blalock, Agentuity
Jul 1, 2025 · 14:40
Rick Blalock, founder of Agentuity, argues that deploying AI agents remains the number one headache for developers, citing common issues like serverless timeouts (agents running 15–30 minutes), statelessness, and networking complexities. He demonstrates Agentuity’s platform, which treats agents as first-class infrastructure citizens: users scaffold projects via a CLI (supporting Bun, Python/UV, Node.js), write a simple request handler, and deploy with automatic routing, tunneling, and a built-in AI gateway that tracks costs per run. The platform decouples inputs and outputs—agents can be triggered via email, SMS, webhooks, or cron—and provides human and agent-facing telemetry via OTel tracing. Blalock notes Agentuity already hosts 50–60 internal agents and is building infrastructure agents to replace tools like PagerDuty, with plans to add Slack/Discord integrations and service-level reasoning capabilities.

Optimizing inference for voice models in production - Philip Kiely, Baseten
Jul 1, 2025 · 15:13
Philip Kiely of Baseten shows how open-source TTS models like Orpheus TTS, built on a LLaMA 3.2 3B backbone, can be optimized for production inference using LLM tooling such as TensorRT-LLM and FP8 quantization, achieving time-to-first-byte (TTFB) under 150 milliseconds and supporting 16–24 simultaneous streams on a single half-H100 GPU. He argues that because TTS models are architecturally similar to LLMs, techniques like dynamic batching, KV-cache quantization, and Torch Compile on the audio decoder can dramatically reduce latency and increase concurrency. He emphasizes that for real-time voice agents, the goal is not raw tokens per second (real-time requires only 83 TPS for Orpheus) but low TTFB and high throughput to minimize GPU spend. However, Kiely warns that non-runtime factors—such as client code using sequential requests without session reuse or sending traffic to distant data centers—can easily add back the milliseconds saved at the model level, and that infrastructure connecting listening, thinking, and talking pipelines is often the dominant source of latency.

[Evals Workshop] Mastering AI Evaluation: From Playground to Production
Jul 1, 2025 · 1:25:08
In this workshop, Braintrust Solutions Engineers Carlos Esteban and Doug guide participants through the complete AI evaluation lifecycle, from offline testing in the playground to production monitoring. They explain the three core ingredients of an eval—task, data set, and score—and demonstrate how to run evals both via the Braintrust UI and the SDK. The session covers LLM-as-judge vs. deterministic code scores, the importance of starting small with synthetic data, and how to use online scoring and logging to capture real user feedback. Human-in-the-loop review is highlighted as a way to establish ground truth and close the feedback loop. The presenters also address audience questions on bootstrapping data sets, non-determinism in LLM judges, and integrating evals into existing projects.

Intro to GraphRAG — Zach Blumenfeld
Jun 30, 2025 · 1:18:35
Zach Blumenfeld introduces GraphRAG using Neo4j, showing how to build knowledge graphs from employee and skill data, combine structured and unstructured data, and use graph traversal with vector search for retrieval. The workshop covers Cypher query patterns for multi-hop similarity, entity extraction from resumes via LLMs, and creating semantic similarity relationships. It demonstrates Leiden community detection for skill clustering and builds a LangGraph agent with four custom tools that balance exact skill matches with vector similarity. Blumenfeld argues that knowledge graphs provide controlled, explainable retrieval logic for agentic workflows, enabling developers to decompose data into graph models that expose domain logic for more accurate retrieval than pure vector search.

Securing Agents with Open Standards — Bobby Tiernay and Kam Sween, Auth0
Jun 30, 2025 · 18:41
Bobby Tiernay and Kam Sween from Auth0 argue that production-grade AI agents must be secured using open identity standards like OAuth 2.1, token exchange, and SIBA to avoid the pitfalls of shared secrets, overly broad scopes, and lost audit trails. They explain that agents need a clear user identity to prevent the confused deputy problem, and that access should be scoped per user and per API with short-lived tokens from a vault. For retrieval-augmented generation, they advocate fine-grained authorization at the retrieval layer, not inside the LLM. Kam demonstrates a local trading agent that uses SIBA (client-initiated back-channel authentication) to request user approval via push notification without a browser, after the user authenticates once. The talk emphasizes that these patterns—async user consent, token exchange, and centralized credential management—are live today in standards and platforms like Auth0 and OpenFGA, enabling developers to build secure agents without reinventing the wheel.

The emerging skillset of wielding coding agents — Beyang Liu, Sourcegraph / Amp
Jun 30, 2025 · 35:06
Beyang Liu, CTO and co-founder of Sourcegraph, argues that coding agents are a real and high-ceiling skill, contrary to skeptics like Jonathan Blow and Eric S. Raymond. He presents design decisions for the agentic era: agents should directly edit files instead of asking permission, UIs should be minimal (like Amp's bare-bones VS Code extension and CLI), and fixed pricing should give way to usage-based models. In a live demo, Liu uses Amp to implement a custom icon for the linear connector in Amp's own codebase, demonstrating agentic search and sub-agents. He shares power-user patterns, including writing long prompts, constructing feedback loops with Playwright and Storybook, and running multiple agents in parallel—as exemplified by Jeff Huntley using Amp to build a compiler while sleeping. Liu emphasizes that agents should enable more thorough code reviews, not replace human understanding.

Agents, Access, and the Future of Machine Identity — Nick Nisi (WorkOS) + Lizzie Siegle (Cloudflare)
Jun 30, 2025 · 14:17
Lizzie Siegle (Cloudflare) and Nick Nisi (WorkOS) argue that AI agents need the same authentication and authorization patterns as humans, extending OAuth to machines. They demonstrate an MCP server built with Cloudflare Workers and WorkOS that allows Claude to order a shirt on behalf of a user, showing how agents can act with user credentials. They explain using Cloudflare durable objects for per-user persistent storage and how authorization can be added to MCP servers to control agent actions, including an example where a 'pretty please' tool bypasses a block. The talk emphasizes the need for fine-grained authorization and audit trails as agents scale to thousands of tasks.

Turning Fails into Features: Zapier’s Hard-Won Eval Lessons — Rafal Willinski, Vitor Balocco, Zapier
Jun 30, 2025 · 16:15
Zapier AI Tech Lead Rafal Willinski and Staff Engineer Vitor Balocco explain how Zapier's evaluation system turns agent failures into targeted improvements through a data flywheel. They detail collecting explicit feedback at critical moments and mining implicit signals like testing behavior, cursing, and user follow-ups. The pair advocates building unit test evals for specific failure modes, then trajectory evals and LLM-as-judge with rubrics to avoid overfitting and capture multi-turn criteria. They share that over-indexing on unit tests hurt model benchmarking, and reasoning models can compare model runs, revealing differences like Claude as a decisive executor versus Gemini's yapping. Ultimately, they argue that the goal is user satisfaction, so A/B testing on a small traffic fraction is the ultimate verification.

Building voice agents with OpenAI — Dominik Kundel, OpenAI
Jun 29, 2025 · 1:25:35
Dominik Kundel from OpenAI presents the new OpenAI Agents SDK for TypeScript and demonstrates how to build voice agents, arguing that speech-to-speech architectures (using GPT-4 real-time) offer lower latency and more natural interactions than chained approaches that transcribe audio to text. He covers two architectures: chained (speech-to-text, text agent, text-to-speech) versus speech-to-speech (native audio understanding), and recommends starting with a small, clear goal, building evals and guardrails early, and using generative prompts to control tone and personality via openai.fm. In a live coding session, he builds a real-time agent with tools (e.g., getWeather) and handoffs, demonstrates delegation to a smarter model (O4 mini) for complex tasks like refunds, and shows built-in interruption handling, human-in-the-loop approval, tracing for debugging, and output guardrails. The SDK supports WebRTC and WebSocket, handles turn detection automatically, and provides session management with a 30-minute timeout that can be extended by injecting previous context.

Containing Agent Chaos — Solomon Hykes, Dagger
Jun 28, 2025 · 23:48
Solomon Hykes, creator of Docker and founder of Dagger, argues that containing agent chaos requires engineering reproducible execution workflows built on containerization applied to each step of an agent's workflow. He introduces 'container use'—agents developing inside fully isolated, customizable environments rather than just sandboxing outputs—and demonstrates a prototype using MCP to integrate with Claude Code and Goose. The system provides background work, rails, seamless human stepping in, and optionality by leveraging Dagger, Git-based state management, and ephemeral containers snapshotted per action. Hykes shows how agents can run parallel experiments, merge snapshots, and discard failed environments without pollution. The episode concludes with him open-sourcing the project as github.com/dagger/containeruse.

Evals 101 — Doug Guthrie, Braintrust
Jun 27, 2025 · 48:31
Braintrust solutions engineer Doug Guthrie presents the full AI evaluation lifecycle, covering offline and online strategies for building robust AI products. He explains that evals require three ingredients—tasks (prompts or agentic workflows), datasets (real-world examples), and scores (LLM-as-a-judge or code-based)—and demonstrates how to run them in Braintrust's playground and via SDK. The talk emphasizes creating a feedback loop: production logs feed into datasets, enabling human review and user feedback to iteratively improve offline evals. Guthrie showcases a changelog app, the new AI-driven "loop" feature for prompt optimization, and answers audience questions on A/B testing models, handling multiple human scorers, pre-launch validation, and exporting data for custom dashboards.

Why should anyone care about Evals? — Manu Goyal, Braintrust
Jun 27, 2025 · 5:41
Manu Goyal, founding engineer at Braintrust, argues that evals are not just unit tests for AI but a critical tool for building a laboratory that enables 90% of product iteration before shipping to production. Drawing from his personal journey from a disappointed Nintendo-playing child to a self-driving car engineer at Nuro, he explains how evals provide the necessary signal to contextualize model improvements and reduce risk. He describes Braintrust's platform as a dev environment that combines tweaking prompts, logging data, and observability to create a data flywheel for AI development. Citing tech luminaries like Kevin Weil, Gary Tan, Mike Krieger, and Greg Brockman, he asserts that evals are the key to industry transformation and successful AI deployment. The talk concludes with an invitation to the Evals Track at the conference, emphasizing his message of 'evals, evals, evals'.

Engineering Better Evals: Scalable LLM Evaluation Pipelines That Work — Dat Ngo, Aman Khan, Arize
Jun 27, 2025 · 24:46
Dat Ngo, AI architect at Arize AI, presents advanced LLM evaluation strategies for production systems, arguing that effective evals go beyond out-of-the-box LLM-as-a-judge to include code-based heuristics, human feedback, and golden datasets. He explains how to build a virtuous cycle of collecting observability data, running evals, and tuning them over time, and demonstrates agent evaluation techniques like trajectory evals to identify failure modes across complex workflows. Ngo covers trade-offs between offline evals and inline guardrails, the use of log probabilities for confidence scoring, and automated prompt optimization through meta-prompting, all illustrated with customer examples from Reddit, Duolingo, and Booking.com.

To the moon! Navigating deep context in legacy code with Augment Agent — Forrest Brazeal, Matt Ball
Jun 27, 2025 · 15:24
Forrest Brazeal and Matt Ball demonstrate how Augment Agent helps developers navigate and modernize legacy codebases using the Apollo 11 guidance computer as an example. They show that Augment's chat mode can instantly explain the 1202 program alarm—indicating the computer was overloaded but could continue—by reading assembly code and web context. Then, using agent mode, they prompt the tool to write a Python implementation of the P65 vertical descent algorithm from the original assembly, run a simulator, and land the lunar module within expected parameters. The episode argues a three-step strategy for any legacy codebase: use chat to understand, write tests via agent, and then modernize modularly—proving AI agents can turn toil like Java 8-to-17 migrations into automated tasks.

Serving Voice AI at Scale — Arjun Desai (Cartesia) & Rohit Talluri (AWS)
Jun 27, 2025 · 17:05
Arjun Desai of Cartesia AI and AWS's Rohit Talluri discuss scaling voice AI for enterprise, arguing that latency and controllability are critical, with Cartesia's state-space model Sonic 2 achieving 40ms model latency for real-time applications. Desai explains that traditional transformer models scale quadratically, while Cartesia's SSMs maintain O(1) generation, enabling 2.5x faster inference than their earlier models. He emphasizes that edge deployment is 5x faster than cloud round-trips, making local models essential for interactive use cases. On quality, Desai notes that voice AI must handle interruptions, accents, and background noise, and that Cartesia's voice marketplace amplifies human voice actors rather than replacing them. Looking to 2030, he predicts voice AI will become the default interface across healthcare, customer support, and gaming, with interactive models extending beyond audio to full world models.

Ship it! Building Production Ready Agents — Mike Chambers, AWS
Jun 27, 2025 · 19:37
Mike Chambers, a developer advocate at AWS, demonstrates how to take a simple local agent—built with a Llama 3.1 8B model and a dice-rolling tool—and ship it to production at cloud scale using Amazon Bedrock Agents. He breaks down the essential components of an agent: model, prompt, loop, history, and tools. Then he live-deploys the agent with Bedrock Agents, configuring instructions and an action group wired to an AWS Lambda function that handles the dice roll. The fully managed service automates scaling, infrastructure, and the agentic loop. Chambers also highlights free courses on DeepLearning.AI and invites attendees to discuss MCP servers and the new open-source SDK for model-first agents.

Introducing Strands Agents, an Open Source AI Agents SDK — Suman Debnath, AWS
Jun 27, 2025 · 14:26
Suman Debnath, a Principal ML Advocate at AWS, introduces Strands Agents, an open-source SDK that simplifies AI agent creation by requiring only a model and tools, eliminating scaffolding. He demonstrates building an agent in a few lines of code to read, summarize, and speak a file, using default tools. Another demo integrates Strands with an MCP server to generate animated math videos via Manim, requiring no system prompts—the model reasons autonomously. Custom tools can be created by decorating functions. Strands supports any model via LiteLLM or Bedrock, and is available at strandsagent.com with a GitHub repository for contributions.

Data is Your Differentiator: Building Secure and Tailored AI Systems — Mani Khanuja, AWS
Jun 27, 2025 · 20:10
Mani Khanuja from AWS explains how data differentiates generative AI applications, emphasizing that foundation models require careful data handling specific to each use case. She demonstrates Amazon Bedrock's capabilities: Bedrock Data Automation for custom pipelines with a single API, Knowledge Bases for RAG with out-of-box chunking and hybrid search, and Guardrails for responsible AI with PII filtering. Using a travel agent example, she details data needs—customer profiles, travel policies—and introduces the 'coconuts' framework: chunking, optimization via semantic cache and re-ranking, observability, evaluation (context relevance), test suites, and iterative updates. She argues that organizations must crack their coconuts—rigorously prepare data and evaluate applications—before scaling into production securely.

How to build world-class AI products — Sarah Sachs (AI lead @ Notion) & Carlos Esteban (Braintrust)
Jun 27, 2025 · 1:43:46
Sarah Sachs (AI lead at Notion) and Carlos Esteban (Braintrust) explain that building great AI products requires 10% prompting and 90% evals and observability, with Notion AI using Braintrust to iterate on prompts and models. Sachs details their cycle: curate small datasets, tie them to scoring functions (LLM-as-a-judge with per-sample prompts and heuristic checks), run evals before shipping, and use production logs to catch regressions. She notes Notion AI supports 100M+ users, switches models in under a day, and 60% of enterprise users are non-English—requiring multilingual eval rigor. Carlos and Doug then walk through Braintrust's framework: tasks (prompts, tools, agents), datasets, and scores (0-1). They demonstrate offline evals via playground and SDK, online scoring in production, user feedback capture, human review setup, and remote evals to bridge complex code with the playground. The workshop covers moving from pre-prod evals to production monitoring, closing the feedback loop by adding underperforming spans to datasets.

From Mixture of Experts to Mixture of Agents with Super Fast Inference - Daniel Kim & Daria Soboleva
Jun 27, 2025 · 53:15
Daria Soboleva and Daniel Kim of Cerebras explain how Mixture of Experts (MoE) architectures enable scaling large language models efficiently by replacing monolithic feedforward networks with specialized experts, a technique used by GPT-4 and Claude. They then introduce Mixture of Agents (MoA), which combines multiple LLMs with custom prompts to outperform frontier models like GPT-4o on complex tasks, reducing a 293-second reasoning problem to 7.4 seconds using Cerebras' ultra-fast inference. The workshop guides participants to build their own MoA system, configure agents for bug fixing and performance optimization on a Python function, and achieve scores up to 120/120. Daniel details Cerebras' wafer-scale chip with 900,000 cores and distributed memory that eliminates memory bandwidth bottlenecks, enabling linear scaling and 15.5x faster inference on Llama 3.3-70B versus GPUs. Daria discusses ongoing research in diffusion models and sparsity, while Daniel notes plans for multimodal APIs and LoRA fine-tuning support.

Forget RAG Pipelines—Build Production Ready Agents in 15 Mins: Nina Lopatina, Rajiv Shah, Contextual
Jun 27, 2025 · 1:15:43
Rajiv Shah, Nina Lopatina, and Matthew from Contextual AI demonstrate how developers can build production-ready RAG agents in minutes using Contextual AI's managed RAG platform, emphasizing that RAG should be treated as a managed service to avoid reinventing infrastructure. They walk through ingesting documents like NVIDIA financials and spurious correlation reports, then querying the agent with questions requiring quantitative reasoning across tables. The platform handles extraction, layout analysis, image captioning, hybrid retrieval, and a state-of-the-art reranker, capped by a grounded language model that avoids hallucinations and provides attribution. Evaluation is done via LM Unit, a model-as-judge that scores responses on criteria like accuracy and causation. The episode also shows integrating the agent with Claude Desktop via MCP and answers audience questions on pricing (consumption-based with a $25 credit), scalability, entitlements, HIPAA, and domain-specific language.

Milliseconds to Magic: Real‑Time Workflows using the Gemini Live API and Pipecat
Jun 27, 2025 · 21:43
Kwindla Kramer (Daily, Pipecat) and Shrestha Basu Mallick (Google DeepMind, Gemini API) argue that voice is the most natural interface and that the Gemini Live API combined with the Pipecat framework enables developers to build real-time multimodal voice agents covering the full stack from models to application code. They demo a voice-driven task management app that handles grocery, reading, and work lists, showing impressive context-aware tool use (e.g., consolidating lists, web searching for 'Dream Count' author) but also jagged edge limitations like turn detection errors and persistent misspelling of 'Kwin.' They discuss how capabilities like turn detection migrate down the stack over time. The episode also covers proactivity, multilinguality, telephony integration, and experimental native audio models for emotive, steerable dialogue.

Realtime Conversational Video with Pipecat and Tavus — Chad Bailey and Brian Johnson, Daily & Tavus
Jun 27, 2025 · 18:46
Chad Bailey of Daily and Brian Johnson of Tavus explain how to build real-time conversational video bots using the Pipecat open-source framework and Tavus's avatar platform. They argue that beyond models, an orchestration layer is essential for handling input, processing, and output with low latency. Bailey details Pipecat's pipeline of frames, processors, and pipelines that manage audio/video frames, speech-to-text, LLM inference, and text-to-speech in a modular way. Johnson describes Tavus's 600-millisecond response time and proprietary models—Sparrow Zero, Raven Zero—plus future turn detection, response timing, and multimodal perception models being integrated into Pipecat. They emphasize that Pipecat solves real-world production challenges like observability, barge-in handling, and parallel pipelines for tasks like voicemail detection. Johnson admits Tavus initially built its own orchestration but now plans to adopt Pipecat internally, as its customers are already using it.

Vector Search Benchmark[eting] - Philipp Krenn, Elastic
Jun 27, 2025 · 14:10
Philipp Krenn from Elastic dissects 'benchmarketing' and explains why most vector search benchmarks are unreliable due to selective scenarios, outdated competitor versions, and omitted quality metrics like precision-recall. He highlights that read-only benchmarks don't reflect real workloads, filtering can slow HNSW-based search, and implicit biases favor the benchmarker's own system. Krenn advises building automated, reproducible benchmarks (like Elastic's nightly Rally tool) to avoid the 'boiling frog' problem of gradual performance degradation. He concludes that only running your own tailored benchmarks yields trustworthy results, urging listeners to learn from even flawed benchmarks rather than dismissing them entirely.

Taming Rogue AI Agents with Observability-Driven Evaluation — Jim Bennett, Galileo
Jun 27, 2025 · 16:14
Jim Bennett, Principal Developer Advocate at Galileo, argues that AI agents must be tamed using observability-driven evaluation, where LLMs evaluate other LLMs to detect failures like hallucinations and tool misuse. He cites real-world examples: the Chicago Sun-Times publishing a hallucinated summer reading list, and a lawyer citing false AI-generated case law. Bennett demonstrates a fintech chatbot that fails to answer 'what is my account balance' directly, requiring three turns; metrics like 'action completion' and 'action advancement' reveal the agent advances but does not complete. He stresses granular evaluation at every step—LLM calls, tool use, RAG retrieval—and using a better, custom-trained LLM as the evaluator. Human feedback is essential to correct mis-scored metrics and continuously retrain. Bennett urges adding evaluations from day one, even before production, and maintaining them in CI/CD and production with alerting for rogue behavior.
Powered by PodHood