Page 21 of 23

No-code fine-tuning: Mark Hennings
Feb 5, 2025 · 9:27
Mark Hennings, creator of Entrypoint, argues that fine-tuning large language models is now accessible without code, offering faster and cheaper alternatives to prompt engineering: GPT-3.5 fine-tuned runs at 73ms per token vs GPT-4's 196ms, saving 88.6% in cost and cutting prompts by 90%. He explains that fine-tuning reduces prompt injection risks, enables team collaboration via training data, and only requires 20 examples to start. Hennings demonstrates Entrypoint's no-code UI that lets users import CSV data, structure fields with templating, and fine-tune GPT-3.5 Turbo, then iteratively improve models by feeding production feedback back into the dataset. He proposes a dev lifecycle: prototype with prompt engineering, use it to build a dataset, fine-tune, evaluate, deploy, and continuously refine.

BotDojo Launch: Enhancing AI Assistants with Evaluations and Synthetic Data
Feb 5, 2025 · 5:47
BotDojo founder Paul Henry demonstrates how his platform uses synthetic data generation and evaluations to productionize AI assistants. He shows a chatbot template with a low-code editor, node tracing, and JSON schema support. By running batch evaluations, he identifies knowledge gaps in the vector database. Then, he generates synthetic test data from live support tickets by extracting question-answer pairs and writing new documents to the index. Running the evaluations again, the scores improve from having red failures to all green, measurably boosting the chatbot's performance. Henry stresses that while hooking up a vector database to an LLM is easy, getting it production-ready requires systematic evaluation and data augmentation.

Building efficient hybrid context query for LLM grounding: Simrat Hanspal
Feb 5, 2025 · 11:35
Simrat Hanspal from Hasura explains how to build efficient hybrid context queries for LLM grounding using Hasura's GraphQL API, demonstrated via an e-commerce product search use case. The talk covers three query types: semantic search (e.g., 'essential oils for relaxation'), structured search (e.g., 'products less than $500'), and hybrid queries combining both (e.g., 'essential oil diffusers between $500 to $1,000'). Hanspal shows how Hasura unifies relational (Postgres) and vector (Weaviate) databases into a single GraphQL API, auto-vectorizing records via events. A critical security demonstration highlights how role-based permissions (e.g., 'Product Search Bot' role with only select access) prevent malicious insert mutations like adding a fake product with ID 7001. The approach enables secure, dynamic data retrieval for Retrieval-Augmented Generation (RAG) pipelines without building separate APIs.

Best Practices for Evaluating Large Language Model Applications with llmeval: Niklas Nielsen
Feb 5, 2025 · 9:33
Niklas Nielsen, CTO and co-founder of Log10, introduces llmeval, a command-line tool that enables teams to ship reliable LLM applications by evaluating and testing prompts and configurations. The tool initializes with four lines of code, uses Meta's Hydra for configurable test structures, and runs multiple samples (default 5) per test to assess stability. Nielsen demonstrates prompt engineering for a math problem, showing how stripping spaces or adding instructions like 'only return the answer' affects strict pass/fail results across Claude, GPT-4, and GPT-3.5. Advanced use cases include testing tools for Python code generation and model-based evaluation, where a larger model grades outputs on criteria like mermaid diagram quality. He notes pitfalls: models favor their own output and struggle with point scores. Log10 addresses this by bridging model-based and human feedback—collecting prior human ratings to train an 'auto-John' that pre-fills reviews for new completions.

How to evaluate a model for your use case: Emmanuel Turlay
Feb 5, 2025 · 7:32
Emmanuel Turlay, CEO of Sematic, explains why evaluating large language models for specific use cases is more difficult than traditional ML evaluation, as metrics like BLEU and ROUGE and benchmarks like GLUE do not measure task-specific performance. He advocates using another LLM as a grader, describing a workflow where a scoring prompt with grading criteria is fed to a model like GPT-4 (best but costly) or FLAN-T5 (good speed–accuracy trade-off) to numerically score outputs. Turlay demonstrates with a politeness evaluation for email closings and introduces AirTrain, a platform that lets users upload datasets, compare models, and visualize metric distributions to make data-driven LLM selections. The episode argues that teams must build custom evaluation procedures rather than relying on generic benchmarks, treating evaluation as a test suite in the ML development pipeline.

Prompt Engineering Tactics: Dan Cleary
Feb 5, 2025 · 5:12
Dan Cleary, co-founder of PromptHub, presents three research-backed prompt engineering tactics to improve LLM output consistency and reduce hallucinations. Multipersona prompting, from University of Illinois, calls on multiple AI agents to collaborate on complex tasks like writing a book. The "according to" method, from Johns Hopkins, grounds prompts to a specific source (e.g., "according to Wikipedia") and can reduce hallucinations by up to 20%. The motion prompt, from Microsoft, adds emotional stimuli at the end of prompts, improving output accuracy by 8% to 115% depending on the task. These tactics are available as templates on PromptHub. Cleary emphasizes that even small changes in prompts can have outsized effects, crucial for maintaining user trust in AI-integrated products.

Scaling AI in Education: A Khanmigo case study: Shawn Jansepar
Feb 5, 2025 · 22:39
Shawn Jansepar, Khan Academy's Director of Engineering, details building Khanmigo, an AI tutor and teacher assistant on GPT-4 with OpenAI, arguing generative AI can democratize one-on-one tutoring. He describes a rapid prototyping culture that launched Khanmigo in three months via a company-wide hackathon, replacing traditional agile with a prototype-to-beta-to-launch framework. Technical challenges include math accuracy through a math agent and chain-of-thought prompting, refactoring prompts into a component architecture for testing, and managing scale using multiple models and dedicated Azure compute. Jansepar highlights ethical design: a Socratic tutor that avoids answers, teacher moderation, and a writing coach with revision history. Khanmigo now has 200,000 paid users, half in school districts, with teacher tools sponsored by Microsoft free to US teachers. Plans include releasing a math tutoring benchmark and evaluating smaller models like Phi-3.

Cohere for VPs of AI: Vivek Muppalla
Feb 5, 2025 · 16:11
Vivek Muppalla, Director of Engineering at Cohere, details the company's enterprise AI strategy centered on security, customization, and deployment flexibility. He presents Cohere's product line: Command R and R+ for generation, plus advanced retrieval models like embeddings and the ReRanker, which reduces RAG costs by narrowing context. Key claims include a focus on enterprise-specific eval suites (health, HR, finance), out-of-the-box citations, and multilingual performance. Partnerships with Accenture and McKinsey bridge the last-mile gap, while wins often stem from private cloud deployment and data control. In Q&A, he recommends purpose-built classifiers for high-throughput production and notes current 128k context windows.

Mastering LLM Inference Optimization From Theory to Cost Effective Deployment: Mark Moyou
Jan 1, 2025 · 33:39
NVIDIA solutions architect Mark Moyou explains that LLM inference differs fundamentally from standard deep learning deployment, requiring careful management of KV Cache, attention mechanisms, and GPU memory to control cost. He details how tokens are processed: prefill computes attention across the entire prompt, then generation produces one token at a time, with KV Cache storing key-value pairs to avoid recomputation. Llama's 32 attention heads and FP8 quantization (halving memory with near-identical accuracy) are cited as key optimizations. Moyou emphasizes measuring time to first token, inter-token latency, and input/output sequence length distributions to size inference engines. He presents NVIDIA's TRT-LLM (model compilation for LLMs) and Triton inference server as tools to maximize throughput, and discusses how query patterns like long-input-short-output or short-input-long-output impact GPU utilization and deployment cost.

Navigating Challenges and Technical Debt in LLMs Deployment: Ahmed Menshawy
Dec 31, 2024 · 16:15
Ahmed Menshawy, VP of AI Engineering at Mastercard, argues that deploying LLMs requires overcoming massive technical debt, with over 95% of the work outside the ML code. He details how Mastercard boosted fraud detection by 300% using LLMs while adhering to seven responsible AI principles. The episode covers RAG as a solution for hallucination, attribution, and access control, noting that 80% of organizational data is unstructured and 71% of companies struggle to manage it. Menshawy highlights that GPU compute for inference now exceeds training, and that fine-tuning open-source models enables joint training of retriever and generator. He also cautions against AGI hype, quoting Ada Lovelace and a Nature article urging focus on current AI risks.

LLM Safeguards: Security Privacy Compliance Anti Hallucination: Daniel Whitenack
Dec 31, 2024 · 34:10
Daniel Whitenack of Prediction Guard outlines a practical checklist for deploying secure and accurate LLMs in enterprise, addressing hallucination, supply chain vulnerabilities, flaky model servers, data breaches, and prompt injection. He proposes a factual consistency model fine-tuned to detect inconsistencies between AI output and ground truth, rather than relying solely on RAG. For supply chain risks, he recommends a trusted model registry and industry-standard libraries. Data breaches are mitigated via PII detection filters and confidential computing (e.g., Intel SGX) to encrypt server memory. Prompt injection is countered with a custom firewall layer using ensembled classification models. In Q&A, he discusses latency trade-offs using smaller NLP models, pre-production testing for visibility, data access controls via role-based database queries, SIEM integration for monitoring new artifacts like model caches, and additional challenges with agents such as excessive agency, proposing dry-run approvals.

Cooking with fire without burning down the kitchen: Dominik Kundel
Dec 31, 2024 · 20:35
Dominik Kundel, who leads product and design for Twilio's Emerging Tech & Innovation team, explains how the company balances disruptive AI innovation with its existing communications and customer data platform businesses. He distinguishes sustaining innovation (e.g., Apple's AI features) from disruptive innovation (e.g., agents not yet enterprise-ready due to quality and cost), arguing that ignoring disruptive AI is increasingly dangerous because quality improves daily. Kundel shares three key lessons from Twilio's AI journey: first, ship early and often—even rough prototypes—to gather real feedback, which led to the Twilio Alpha sub-brand for setting expectations; second, build a curious, problem-owning team rather than requiring existing AI expertise; third, share learnings internally and externally to avoid operating in a silo and to help customers be thought leaders. He recounts how an initial AI personalization engine was too disruptive for mainstream R&D, so the team iterated on an AI Assistance agent builder, using internal hackathons and dogfooding with low-risk use cases like IT helpdesk to improve quality before broader release.

E-Values Evaluating the Values of AI: Sheila Gulati and Nischal Nadhamuni
Dec 31, 2024 · 28:13
Sheila Gulati (Tola Capital) and Nischal Nadhamuni (Klarity CTO) argue that evaluating AI systems is at a seminal moment, requiring a shift from simplistic benchmarks to multi-faceted, user-centric approaches. Gulati highlights the inadequacy of current leaderboards, citing benchmark hacking and the need for dynamic datasets and value alignment. Nadhamuni shares how Klarity, which raised a $70M Series B, tackles eval challenges in its document-processing product—processing over 500K documents for customers with 15+ LLM use cases. He details practical strategies: front-loading UX risk, thinking backwards from user experience, reducing degrees of freedom by standardizing prompt engineering across customers, and building scrappy future-facing evals. The episode emphasizes that the AI engineer community must innovate benchmarks, bring depth and community sharing to evaluations, and introspect the values instilled in AI systems.

Hiring & Building an AI Engineering Team: Dr. Bryan Bischof
Dec 31, 2024 · 29:07
Dr. Bryan Bischof, Head of AI at Hex, argues that building an AI engineering team requires hiring based on product stage—starting with full-stack engineers and data profiles, then adding designers and later MLEs—while avoiding the 'mythical man month' trap in early AI products. He advocates for data intuition over LeetCode in interviews, using a take-home data exercise to assess candidates' ability to extract meaning from data and give feedback. Key attributes he looks for are curiosity, urgency, and product-mindedness, noting that enthusiasm alone is insufficient. Bischof also recommends working directly with domain experts to model AI behavior and suggests a centralized AI platform team to support multiple product teams, rather than having each team build AI infrastructure independently.

RAG for VPs of AI: Jerry Liu
Dec 31, 2024 · 26:51
Jerry Liu, CEO of LlamaIndex, argues that building production RAG systems requires a new data processing stack distinct from traditional ETL, emphasizing data quality through advanced parsing (Llama Parse) and rigorous eval over chunk size tuning. He explains that while longer context windows may eliminate fine-grained chunking, retrieval from external storage remains vital for multi-doc enterprise systems. Liu addresses data privacy with VPC deployment options, and notes that Llama Cloud processes but does not store data. He also highlights Llama Agents, an open-source multi-agent framework for deploying agentic microservices, and shares that Llama Parse processes tens of millions of pages monthly.

AI Platform Engineering: Patrick Debois
Dec 31, 2024 · 28:18
Patrick Debois, who coined DevOps in 2009, argues scaling GenAI requires an AI platform team, mirroring the cloud and DevOps pattern. He lists shared infrastructure: model access, vector databases, RAG connectors, version control, proxy, observability, monitoring, caching, feedback services, and stresses enablement through prototyping and frameworks. He warns of pitfalls like chasing GenAI without use case, overfocus on fine-tuning, and cost obsession. Citing the Ironies of GenAI Automation, he notes that as engineers shift from producing to reviewing code—copying co-pilot suggestions more on weekends—they risk losing situational awareness. For governance, he advises awareness programs, opt-out training, license checks, and layered guardrails: central rules plus team-specific overlays. He recommends combining cloudops, secops, devx, data platform, and AI platform teams for collaboration.

Real ROI: Lessons from Enterprises that have already succeeded with LLMs at Scale: Raza Habib
Dec 31, 2024 · 20:01
Raza Habib, CEO of Humanloop, shares lessons from enterprises like Duolingo, Filevine, and Ironclad that have achieved real ROI with LLMs at scale. He argues that success hinges on centering domain experts (e.g., Duolingo's linguists do all prompt engineering), needing less ML expertise than expected, and breaking down evaluation into small, testable components rather than chasing a single metric. Habib emphasizes collecting end-user feedback (acceptance rates, edits, thumbs up/down) and building tooling for logging, regression testing, and team collaboration. Filevine doubled revenue by launching six LLM-powered products; Ironclad's open-source Rivet logging tool enabled agents to auto-negotiate 50% of contracts. The talk covers how to optimize the four-component chain (base model, prompt, data selection, function calling) and why simpler systems will win as models improve.

Understanding AI Stakes to Break Production Code: Philip Rathle
Dec 31, 2024 · 23:24
Philip Rathle, CTO of Neo4j, argues that the level of 'stakes' in an AI application determines the production barriers and appropriate solutions, from low-stakes summarization to high-stakes uses requiring knowledge graphs and human-in-the-loop systems. He distinguishes pilot (full autonomy) from co-pilot (human oversight), showing how vector RAG solves moderate stakes but fails for high-stakes needs like tightening bolts on a 737 MAX 9. He promotes Graph RAG for deterministic reasoning and fact retrieval. Attendees share learnings: using LLMs to write code rather than reason directly, adapting to user keyword habits, focusing on few projects with achievable accuracy, building eval pipelines early, and for regulated loans offering three options with explanations instead of a single recommendation.

The ROI of AI: Why you need Eval Framework - Beyang Liu
Dec 31, 2024 · 25:28
Beyang Liu, CTO of Sourcegraph, presents six evaluation frameworks for measuring the ROI of AI coding tools, arguing that no single metric suffices because AI ROI reduces to the intractable problem of developer productivity. He dismisses the 'roles eliminated' framework as irrelevant for engineering, noting that backlogs are never fully addressed. Instead, he highlights customer examples: A/B testing velocity at Palo Alto Networks showed 20-30% acceleration with Cody; time saved as a function of engagement lower-bounds gains (e.g., each code search saves ~2 minutes); and LITOS tracked value-add zones (building features, writing tests, reviewing code) instead of lines of code. For key initiatives like code migrations, ROI is measured by how many months the migration is pulled forward. Liu also advocates developer surveys bounded by a 5-25% productivity tool budget, and warns against the Mythical Man Month fallacy: preferring one 100x human developer over 100 mediocre AI agents.

AI Frontiers in Trust and Safety Combatting Multifaceted Harm on Tinder at Scale: Vibhor Kumar
Dec 2, 2024 · 14:36
Vibhor Kumar, senior AI engineer at Tinder, explains how the company uses open-source LLMs and LoRAX to detect a long tail of trust and safety violations at global scale. Facing challenges like content pollution and automated fraud from generative AI, Tinder leverages pre-trained models such as LLaMA and Mistral, fine-tuning them with LoRA and QLoRA on hybrid datasets generated by GPT-4 and manually verified. They serve dozens of fine-tuned adapters on a single GPU using LoRAX, achieving real-time inference (tens of QPS, ~100ms latency) for categories including hate speech, pig butchering scams, and underage users. The approach yields near 100% recall on simpler tasks and significant improvements over baselines, with better generalization that resists adversarial evasion. Future directions include visual language models for explicit image detection and automating retraining pipelines.

Enhancing Quality and Security in CI: Gunjan Patel
Nov 27, 2024 · 18:27
Gunjan Patel, Director of Engineering at Palo Alto Networks, presents Ghost Pilot, an AI-powered CI pipeline that outsources boring software development tasks to enable self-evolving code. Unlike real-time Copilots, Ghost Pilot operates as a 'slow system' in CI, iteratively improving variable names and code comments, then generating unit tests by first listing edge cases and personalizing them via team context files. For security, it simulates three AI roles—Red Team engineer, developer, and engineering manager—who debate identified vulnerabilities, prioritize fixes by risk and effort, and propose changes with citations. Patel shares a reusable CI template with a Bring-Your-Own-LLM option, demonstrated by catching a logical Kubernetes bug missed by static analysis tools.

Iterating on LLM apps at scale Learnings from Discord: Ian Webster
Nov 22, 2024 · 18:26
Ian Webster, Senior Staff Engineer at Discord and maintainer of Promptfoo, shares how Discord built and scaled Klyde AI, a chatbot for 200 million users, focusing on evaluation and safety. He argues that evals should be treated as simple, deterministic unit tests that run locally, avoiding complex metrics, and that breaking the system into small, testable pieces (e.g., checking for lowercase output to enforce casual tone) achieves 80% of the goal with 1% of the work. Webster details how Discord mitigated risks like the 'grandma jailbreak' (which originated on Discord) by using an attacker LLM to generate adversarial inputs and a judge to refine them, exposing cracks in safeguards. He advocates for pre-deployment red teaming over live filtering, and describes using Promptfoo for risk assessment across brand, legal, and safety categories. The episode also covers prompt management via Git and Retool, routing with occasional GPT-4 responses to correct model drift, and the challenge of closing the feedback loop due to privacy constraints, relying instead on dogfooding and public examples.

Decoding Mistral AI's Large Language Models: Devendra Chaplot
Nov 21, 2024 · 18:16
Devendra Singh Chaplot of Mistral AI details the company's open-source large language models, including Mistral 7B, Mixtral 8x7B, Mixtral 8x22B, and CodeStral 22B, arguing that open models complement rather than compete with profit by serving as branding tools and driving customer acquisition for proprietary upgrades. He explains the three-stage LLM training process—pre-training on trillions of tokens, instruction tuning with prompt-response pairs, and learning from human feedback via preference optimization—emphasizing that more data does not guarantee better performance due to noise. The episode highlights Mistral's focus on optimizing the performance-to-cost ratio, with CodeStral 22B outperforming larger models like Code LLaMA 70B while being smaller and multilingual across 80+ programming languages. Practical guidance is given: prototype with high-end commercial models, then fine-tune open models for specific tasks to balance performance and cost.

The AI emperor has no DAUs why most devs still don't use code AI: Quinn Slack
Nov 20, 2024 · 18:45
Quinn Slack, CEO and cofounder of Sourcegraph, argues that despite massive hype, only about 5% of professional developers actually use Code AI tools, with total recurring revenue from Code AI sitting at roughly $300 million ARR—a fraction of Salesforce's $36 billion. He cites GitHub's 1.3 million paid Copilot subscribers and just 935,000 yearly active users receiving suggestions, revealing the gap between perception and reality. Slack warns that the entire AI ecosystem—foundation models, infra, and applications—risks collapse if usage doesn't grow, and most revenue in AI flows to NVIDIA and chip makers, not software. From building Cody, the number two Code AI product, he shares lessons: hype fools everyone, autocomplete is a freakishly good feature that spoils expectations, while chat and agents are harder to vet and adopt. He advises builders to use their own product daily, ignore customer demands for buzzwords like fine-tuning, and manually build explicit interactions before adding magic. Slack concludes that the industry must collectively dehype and focus on real daily active users to turn the potential into sustained enterprise revenue.

A Practical Guide to Efficient AI: Shelby Heinecke
Nov 18, 2024 · 17:45
Shelby Heinecke, who leads an AI research team at Salesforce, presents five orthogonal dimensions for making AI models efficient: efficient architecture selection, pre-training, fine-tuning, inference, and prompting. She highlights the power of small models like Phi-3 (3.8B parameters outperforming a 7B model), mobile LLM (350M parameters on par with 7B after fine-tuning), and Octopus (2B fine-tuned Gemma exceeding GPT-4 on Android tasks). For efficient inference, she explains post-training quantization, showing 4-bit quantization nearly halves memory usage without performance loss (e.g., LLaMA models), but warns 3-bit can degrade quality. She recommends frameworks like LLaMA CBP and ONNX Runtime for quantization and introduces her team's open-source Mobile AI Bench for evaluating quantized models, including an iOS app to measure latency and battery drain. The central claim is that deploying AI in constrained environments—cloud, on-prem, or edge—demands efficiency, and these practical techniques bridge the gap from demo to production.

Moondream: how does a tiny vision model slap so hard? — Vikhyat Korrapati
Nov 14, 2024 · 19:26
Vikhyat Korrapati built Moondream, a tiny open-source vision language model under 2 billion parameters that matches LLaVA 1.5, a model four times its size, on VQA v2 and GQA benchmarks. He attributes this to focusing on image understanding over world knowledge and investing heavily in synthetic training data—a pipeline that generated 35 million images and used two orders of magnitude more compute on data than training. Key lessons: community engagement was critical, open-source builds trust, and safety guardrails should be application-layer, not baked in. He argues tiny models will dominate production due to cost and privacy advantages, and that prompting will replace custom model training for most vision tasks. Moondream raised a seed round from FullySysAscent and the GitHub Fund.

Navigating RAG Optimization with an Evaluation Driven Compass: Atita Arora and Deanna Emery
Nov 12, 2024 · 18:14
Atita Arora (Qdrant) and Deanna Emery (Quotient AI) demonstrate how evaluation-driven optimization improves Retrieval Augmented Generation (RAG) systems, using Qdrant's vector database and Quotient's evaluation platform. They detail a 10-experiment pipeline on Qdrant documentation, starting with naive RAG and incrementally testing chunk sizes, embedding models, LLMs (Mistral, GPT-3.5), re-rankers (MixedBread, Cohere, ColBERT), and hybrid search. Key metrics—faithfulness (hallucination reduction), context relevance, and chunk relevance—guide decisions: increasing chunk size hurt faithfulness (0.76 to 0.72), while smaller chunks with larger retrieval windows improved it. Switching to GPT-3.5 lifted all metrics, but poor context relevance revealed retrieval issues. Adding Cohere re-ranking boosted faithfulness to 0.82, and hybrid search with re-ranking achieved 0.85, proving domain-specific terminology demands tailored search strategies. The episode stresses avoiding over-engineering without metrics, keeping evaluation datasets current, and using domain understanding for systematic RAG improvement.

How Zapier Builds AI Products and Features with the Help of Braintrust: Ankur Goyal & Olmo Maldonado
Nov 7, 2024 · 14:59
In this talk, Olmo Maldonado (Senior AI Engineer at Zapier) and Ankur Goyal (CEO of Braintrust) explain how Zapier uses Braintrust's evaluation and observability platform to build and improve AI features like AI Zap Builder and Copilot. They moved from 7 manual unit tests to over 800 automated evals, improving accuracy by nearly 300%. Braintrust's tracing capabilities let them dissect Copilot's multi-tool agent performance, leading them to adopt GPT-4 Turbo for the message router despite latency trade-offs. When switching to GPT-4.0 caused a regression below 80% scores, they diagnosed the issue via evals—GPT-4.0 was ignoring system prompts—and fixed it by relaxing prompt engineering and adjusting tool choices, recovering accuracy and cutting response time from 14 to 3 seconds. The episode details how eval-driven development and observability enable rapid iteration and reliable AI products at scale.

What It Actually Takes to Deploy GenAI Applications to Enterprises: Arjun Bansal and Trey Doig
Nov 4, 2024 · 21:30
Trey Doig of Echo AI and Arjun Bansal of Log10 recount Echo AI's journey deploying a GenAI-native conversational intelligence platform for billion-dollar retail brands, focusing on the centrality of accuracy. Echo AI ingests all customer conversations, uses LLMs to surface insights at 100% coverage, but must overcome enterprise trust issues by achieving 95% accuracy within seven days. The platform relies on Log10's auto feedback system, which uses AI-based review to match human accuracy with model speed, yielding a 20 F1 point improvement in one use case. The episode details how Echo AI's solution engineers use Log10 to grade summarizations, catch hallucinations, and track model drift, turning human feedback into curated datasets for fine-tuning. Ultimately, the partnership demonstrates a path to self-improving LLM applications through iterative accuracy measurement and improvement.

Knowledge Graphs & GraphRAG: Techniques for Building Effective GenAI Applications: Zach Blumenthal
Nov 1, 2024 · 1:39:52
In this workshop, Zach Blumenthal of Neo4j demonstrates how to build a GraphRAG application using Neo4j, OpenAI, and LangChain on an H&M fashion dataset. He argues that combining knowledge graphs with vector search and graph embeddings improves retrieval and personalization for LLMs. The session covers creating a Neo4j sandbox, loading data, performing vector search with OpenAI embeddings, adding collaborative filtering via graph traversal patterns, using graph data science (FastRP) to generate graph embeddings for recommendations, and integrating everything into a LangChain chain that generates personalized marketing emails. Attendees build a Gradio app that, given a customer ID and season, outputs a tailored email with product recommendations, demonstrating how structured graph data enhances RAG systems beyond pure vector search.

AI Engineering Without Borders — swyx
Oct 30, 2024 · 10:32
In this talk from the AI Engineer World's Fair, host swyx argues that AI inherently disrespects human-made borders—it is naturally multilingual, multimodal, and indifferent to copyright or ground truth. He challenges the field to define its own laws, distinguishing constants (e.g., humans speak at 80 wpm vs. read at 200 wpm) from contingent facts (e.g., Apple Intelligence's 30 tokens/sec baseline). Reflecting on one year of the 'Rise of the AI Engineer,' he notes that the conference tracks—RAG, code gen, agents, multimodality—are arbitrary constructs we created, not natural categories. He proposes that AI Engineering sits between software engineering and real engineering: it must apply natural sciences for humanity's benefit. The talk concludes with a call to 'disagree more'—with your own conclusions, each other, and the status quo—and to transform the Shoggoth of raw AI into mass transit tools for society.

State Space Models for Realtime Multimodal Intelligence: Karan Goel
Oct 29, 2024 · 14:26
Karan Goel, founder of Cartesia, argues that state space models (SSMs) are key to real-time multimodal intelligence, offering cheaper, faster, and higher-quality alternatives to transformers for streaming applications like conversational voice and on-device assistants. He contrasts batch intelligence (cloud APIs) with streaming intelligence needed for low-latency tasks, emphasizing that SSMs compress information linearly rather than storing all tokens, enabling efficient long-context processing. Cartesia’s voice generation model achieves instant latency in the data center and is being optimized for on-device Mac and desktop deployment. Goel asserts that compression helps long-context tasks (e.g., 24-hour security footage analysis) more than retrieval, and that SSMs now match transformer quality while scaling better for multimodal data.

Second Order Effects of AI: Cheng Lou
Oct 28, 2024 · 21:46
Cheng Lou explores how to anticipate second-order effects of AI by examining who is learning, widening information bandwidth, and extrapolating quantity to extremes. He uses chess and Go as examples where AI initially seemed to end human play but actually improved it, likening this to Conway's Game of Life's emergent behavior. He argues AI can aid human learning in drawing through stroke auto-completions and in music via indirect manipulation of spectrograms, shifting focus from automation to personal skill development. He envisions personalized AR language translation to replace one-size-fits-all text, and critiques current UI designs by proposing machine-learned gesture interpretation that considers full context. Finally, he extrapolates generating thousands of AI-curated UI layouts at design time, then using classification at runtime to deliver dynamic, context-aware interfaces, moving beyond static media queries.

Build an AI Research Agent: Apoorva Joshi
Oct 25, 2024 · 27:33
In this workshop, Apoorva Joshi, an AI Developer Advocate at MongoDB, teaches how to build an AI research agent using MongoDB as the memory provider and knowledge store, open-source LLMs from Fireworks AI (Fire Function V1) as the agent’s brain, and LangChain to orchestrate the workflow. The agent searches for research papers, summarizes them, and answers questions based on past research, using tools like ArXivLoader and a MongoDB vector store. Joshi explains key agent concepts—planning with chain of thought and react patterns, short-term and long-term memory, and tool creation—then guides attendees through hands-on sections to build the agent step by step. Attendees learn to create agent tools, implement reasoning with react, and add short-term memory persisted to MongoDB. The workshop emphasizes that agents enable complex, multi-step tasks through iterative reasoning, tool use, and memory, trading higher cost and latency for improved accuracy.

The Multimodal Future of Education: Stefania Druga
Oct 24, 2024 · 20:05
Stefania Druga, a research scientist on Google's Gemini team, argues that multimodal AI can transform education by making learning more interactive and critical. She presents Cognimates, an open-source platform she built that expands Scratch to let kids train custom AI models and program smart devices, showing a longitudinal study where kids became more skeptical of AI's intelligence after tinkering. Druga also introduces a benchmark for math misconceptions and demonstrates live demos using the Gemini API—a science tutor that responds to drawings, a math coach that avoids giving direct answers, and a curiosity-driven object identifier—all in under 100 lines of code. She emphasizes the need to cultivate AI literacy early, as 70% of generative AI users are Gen Z, and invites engineers to build tinkering tools that balance agency between learners and AI.

Productionizing GenAI Models – Lessons from the world's best AI teams: Lukas Biewald
Oct 23, 2024 · 22:36
Lukas Biewald, founder of Weights and Biases, shares lessons from productionizing GenAI models, emphasizing that while AI is easy to demo, it is hard to productionize. He illustrates this with a personal project building a custom Alexa-like device using LLaMA and Whisper, where accuracy improved from 0% to 98% through prompt engineering, switching to Mistral, and fine-tuning with QLoRA. Biewald argues that tracking experiments (including failures) is critical for reproducibility and collaboration, and that a robust evaluation framework—beyond 'testing by vibes'—is essential for iterating and shipping v2. He notes that 70% of the audience had LLM apps in production, yet many lack solid evaluations, and recommends starting with lightweight prototypes and incorporating end-user feedback.

Code Generation and Maintenance at Scale: Morgante Pell
Oct 17, 2024 · 18:54
Morgante Pell, founder of Grit, argues that AI agents for code generation must be built for modifying large existing codebases rather than generating new apps, and that combining static analysis with AI enables reliable migrations at scale. Grit has merged more PRs than any other company by focusing on supercharging top engineers, using GritQL to precisely find code via syntactic and semantic queries, and relying on compilers like TSC to catch errors that LLMs miss. They overcome slow enterprise builds (10 minutes for type checking) by precomputing in-memory indexes and using Firecracker to snapshot and fork environments for parallel agent execution, achieving PRs with only a few iterations. For editing, they developed a custom 'loose search and replace' format to avoid expensive full-file generation and LLM laziness. This approach allowed one customer to complete a multi-year OpenTelemetry migration in a week with under 100 developer hours. Morgante envisions future UIs that let engineers manage entire codebases like SimCity.

The Hierarchy of Needs for Training Dataset Development: Chang She and Noah Shpak
Oct 15, 2024 · 16:32
Chang She (CEO of LanceDB) and Noah Shpak (AI data platform lead at Character AI) argue that data infrastructure is the critical bottleneck for LLM training, and LanceDB's columnar format solves the 'new cap theorem' for AI: needing fast scans, random access, and handling large multimodal blobs simultaneously. Noah explains how Character AI structures pre-training around wide domain coverage and post-training around granular analytics like token counts and difficulty scores, using synthetic data, quality scoring, and dataset selection to improve models. Chang details how Lance format provides zero-copy schema evolution, time travel, and indexing extensions for vector, scalar, and full-text search, enabling a single table to serve SQL analytics, PyTorch training, and production vector search. The episode emphasizes that speed and iterative dataset management are key to accelerating AI research, with LanceDB facilitating cheap random access and low-infra billion-scale vector search.

No more bad outputs with structured generation: Remi Louf
Oct 14, 2024 · 15:32
Rémi Louf, CEO of .txt and co-maintainer of Outlines, argues that structured generation—guiding LLMs to produce valid regex, JSON, or context-free grammar outputs—eliminates parsing errors and hallucinations while adding negligible overhead. Outlines masks tokens that violate the target structure, enabling Mistral 7B to achieve 99.9% valid JSON (vs. 17% without). It also accelerates inference: a chatty 50-token ChatGPT answer shrinks to 8 tokens, and structured generation boosts open models beyond GPT-4—Phi-3 Medium hits 96.5% on Berkeley function calling (GPT-4 gets 93.5%). With one shot matching eight shots in accuracy, Louf contends that most text is structured and that structured generation should be the default for non-chatbot workflows.

Realtime Data Connectivity for AI: Tanmai Gopal
Oct 11, 2024 · 7:14
Tanmai Gopal from Hasura introduces Pacha DDN, an AI-powered data access layer that lets LLMs securely query live data from multiple sources, arguing that the key is treating all data—structured, unstructured, and APIs—with a unified SQL-based query language. He demonstrates with a Blockbuster example: writing an email to a top customer by querying database transactions and recent rentals via natural language. The system uses an object model for authorization that applies rules based on data schema and session properties, regardless of data origin. To overcome LLM reasoning limitations, Pacha DDN asks the LLM to write Python code to retrieve data instead of reasoning directly. The talk concludes that for AI to be useful, it needs realtime data access, and Pacha DDN provides a secure, explainable query planner to achieve that.

Architecting and Testing Controllable Agents: Lance Martin
Oct 11, 2024 · 2:21:54
Lance Martin presents LangGraph, a graph-based framework for building controllable agents that trade some open-ended flexibility for significantly higher reliability compared to classic React agents, achieving 100% consistent tool-calling trajectories even with an 8B local model. He demonstrates self-corrective RAG patterns like Corrective RAG, SelfRAG, and Adaptive RAG, where the agent grades retrieved documents, checks for hallucinations, and routes to web search when needed. Martin also covers three testing loops: in-app error handling with LangGraph, pre-production evaluation using LangSmith to compare agent answers and tool trajectories against ground-truth datasets, and production monitoring with online evaluators that flag retrieval quality, answer relevance, and hallucinations without reference answers. He shares results from a five-question evaluation showing LangGraph agents achieve 80% answer accuracy and 100% tool-trajectory correctness with Fire Function V2, while React agents with GPT-4o degrade in tool reliability. The talk addresses practical concerns like handling many tools (suggesting RAG for tool selection), multi-turn conversations, and the importance of…

Breaking AI's 1-GHz Barrier: Sunny Madra (Groq)
Oct 10, 2024 · 20:11
Sunny Madra (Groq) argues that LLM inference speed is undergoing a transformation akin to the 1-GHz microprocessor breakthrough, with Groq increasing LLaMA 3 8B speed by over 50% in just two months. He draws parallels to the industrial revolution: just as mass production replaced bespoke manufacturing, AI now enables 1,000 outputs in a minute where previously one designer produced one per day. Specific applications like Globe.Engineer plan a trip in seconds by processing 10,000 input tokens per second. Madra envisions LLMs as future OS cores, enabling instantaneous decision-making, universal natural-language interfaces, and multi-agent collaboration that allows smaller models to compete with larger ones. Edge AI projects like hyperspace.ai distribute unused GPU compute, and personalized education (citing Khan's two-sigma effect) becomes practical with cheap inference. He also highlights automated data science agents (Pioneer) that run endlessly, discovering correlations like productivity decline after certain performance reviews. The episode underscores that speed unlocks predictive analytics, context-aware systems, and enhanced security against AI-powered scams.

Making Open Models 10x faster and better for Modern Application Innovation: Dmytro (Dima) Dzhulgakov
Oct 9, 2024 · 18:55
Dmytro Dzhulgakov, CTO of Fireworks AI, argues that open-source models are the future for GenAI applications because they offer lower latency, lower cost, and domain adaptability compared to proprietary models. He explains that open models can be up to 10x faster for narrow domains and cut costs significantly, using examples like fine-tuned Llama 3 for function calling. Fireworks addresses the challenges of setup, optimization, and production readiness with a custom serving stack that delivers the fastest inference for long prompts and image generation (e.g., SDXL). He highlights FireFunction V2, an open-source model for function calling that combines chat and tool use, and notes that Fireworks serves over 150 billion tokens per day for companies like Quora and Cursor. The talk emphasizes that platforms like Fireworks enable developers to start with serverless inference, fine-tune models, and scale to enterprise-grade deployment with dedicated hardware.

We accidentally made an AI platform: Jamie Turner
Oct 8, 2024 · 5:36
Jamie Turner, co-founder of Convex, explains how his company's backend platform accidentally became an AI platform because its reactive data flow paradigm—extending React's state reactivity to the server with subscribable queries and mutations—perfectly suits generative AI workflows. He describes Convex's architecture as seamlessly syncing state between backend steps and the application, enabling concurrent chains of tasks like automatic speech recognition, summarization, embedding generation, and finding related notes. Turner notes that post-ChatGPT, over 90% of Convex projects are generative AI. To support this shift, Convex added native vector indexes for schema fields and launched a startups program with discounts and exclusive forums. Upcoming high-level components will encapsulate state machines for sophisticated workflows. Turner presents Convex as a way for generative AI engineers to ship quickly and confidently.

Building and Scaling an AI Agent Swarm of low latency real time voice bots: Damien Murphy
Oct 8, 2024 · 1:07:23
Damien Murphy, Senior Applied Engineer at Deepgram, demonstrates building and scaling low-latency real-time voice bots using Deepgram's new voice agent API, which wraps speech-to-text, LLM, and text-to-speech into a single streaming endpoint. He shows a drive-thru ordering demo with function calling (add/remove items) using GPT-4o, achieving sub-second latency by co-locating components. Murphy explains scaling to millions of concurrent calls through regional Kubernetes clusters and multi-agent swarms (routing, booking, support agents) to reduce complexity and cost. He addresses endpointing challenges, VAD-based barge-in, and cost/quality trade-offs with hosted vs. self-hosted models, noting Deepgram offers 20x cheaper TTS than ElevenLabs and 50ms STT latency self-hosted. The talk emphasizes keeping agents simple, using smallest capable LLMs, and composability for reuse.

The era of unbounded products: Designing for Multimodal IO: Ben Hylak
Sep 25, 2024 · 20:32
Ben Hylak, founder of Dawn and former Apple Vision Pro designer, argues that the key to building intuitive AI products in the era of unbounded interfaces is adding structure—highlighting what matters, establishing hierarchy, and leveraging familiarity—lessons from designing VisionOS. He shows how successful AI apps like Dot, Perplexity, and Claude use structure (e.g., Claude pulling code into artifacts) while the Vercel chatbot's inline dynamic UI is an anti-pattern because it disrupts conversation flow. For agents, spreadsheets (like Clay) make unfamiliar multi-step tasks familiar. Looking ahead, Hylak predicts less prompt engineering via sparse autoencoders for millions of ranked, personalized presets, shifting product evaluation from evals to user analytics as apps become increasingly personalized.

Everything you need to know about Fine-tuning and Merging LLMs: Maxime Labonne
Sep 25, 2024 · 17:52
Maxime Labonne from Liquid AI explains the LLM training lifecycle—pre-training, supervised fine-tuning (SFT), and preference alignment—and when to use fine-tuning over prompt engineering. He details SFT dataset creation (accuracy, diversity, complexity) and techniques like full fine-tuning, LoRA, and QLoRA, with key hyperparameters. The core of the talk is model merging: combining weights of fine-tuned models without GPU, using methods like SLERP (spherical linear interpolation for two models), TIES (pruning redundant parameters to merge many models), pass-through (concatenating layers, e.g., Meta LLaMA 3 120B Instruct by repeating layers gives strong creative writing), and Franken-MoE (extracting FFN layers from domain-specific models with a router). Labonne demonstrates these with his NeuralBeagle and Beyonder models, noting merged models dominate the OpenLLM leaderboard and that TIES merging often outperforms more experimental Mixture of Experts approaches.

LLM Scientific Reasoning: How to Make AI Capable of Nobel Prize Discoveries: Hubert Misztela
Sep 23, 2024 · 20:00
Hubert Misztela, an AI researcher at Novartis, argues that LLMs require more than naive RAG to achieve scientific reasoning capable of Nobel-level discoveries. He uses the 1990s petunia flower experiment — where three separate biological phenomena (including RNA interference) went unexplained for eight years until their common cause was found — as a benchmark. Misztela classifies question complexity from one-to-one to multi-needle problems and demonstrates that reasoning before retrieval (e.g., routing, GraphRAG) and after retrieval (e.g., relevance classifiers) improves hypothesis generation. His experiments show that strict prompting and a relevance classifier that evaluates each paper's contribution to advancing a hypothesis can recover the correct RNA/DNA link without post-discovery knowledge. He concludes that harder problems demand reasoning steps and that brute-force checking with LLMs may outperform embedding distances.

How to Construct Domain Specific LLM Evaluation Systems: Hamel Husain and Emil Sedgh
Sep 19, 2024 · 18:45
Emil Sedgh (CTO at Rechat) and Hamel Husain (independent consultant) describe how they built a domain-specific evaluation system for Lucy, an AI assistant for real estate agents, to move beyond vibe checks and achieve production reliability. They argue that systematic evaluation starts with simple unit tests and assertions based on observed failure modes, logging traces for human review with custom tools to remove friction, and using LLMs to synthetically generate test inputs. The presenters emphasize that focusing on process over tools, avoiding generic off-the-shelf evals, and not jumping too early to LLM-as-a-judge are critical to success. This eval framework enabled Rechat to rapidly increase Lucy's success rate and made possible fine-tuning for complex tasks like mixing structured and unstructured outputs, handling multi-step commands that invoke five or six tools, and incorporating user feedback loops—capabilities they could not achieve with few-shot prompting alone.

From model weights to API endpoint with TensorRT LLM: Philip Kiely and Pankaj Gupta
Sep 13, 2024 · 1:40:01
Philip Kiely and Pankaj Gupta of Baseten lead a workshop on TensorRT-LLM, NVIDIA's high-performance inference framework for LLMs, arguing its use delivers best-in-class throughput and latency on NVIDIA GPUs. They explain that TensorRT-LLM optimizes computational graphs via plugin kernels and in-flight batching, achieving 216 tokens per second and 180ms time to first token on Mistral 7B. The workshop demonstrates building an engine for TinyLlama 1.1B, including FP8 quantization that reduced engine size from 2 GB to 1.2 GB with minimal quality loss. They benchmark a deployed model, showing 7,000 total tokens per second at batch size 64 on an A10G. The presenters compare TensorRT-LLM favorably to VLLM for high-throughput production use and introduce Truss, Baseten's open-source packaging tool, alongside their managed platform for automatic scaling and fast cold starts.
Powered by PodHood