Episodes from AI Engineer about Enterprise AI Solutions.

Which AI startups actually land enterprise contracts? — Brian Lewis, Millennium
Aug 29, 2026 · 18:45
Brian Lewis of Millennium, a hedge fund, explains why only ~5% of AI vendor demo calls end in signed contracts: on the buying side, 10-15 startups per pain point become two or three demos, zero or one pilot, and one contract per four pilots. He catalogs failures—a vendor asking customers to self-report gateway telemetry so it could charge margin on traffic it never carried, a 'zero retention' vendor revealing data it shouldn't have, betas engineered around permissive retention, and a core API down for hours on a trading day. His larger claim: 60% of becoming AI-native is unsexy work—entitlements, integration, change management—and agents inherit whatever foundations you have, so fix the boring 60% first. He advises startups to treat Millennium-grade security and ops as the bar.

How AI Agents Let GTM Teams Scale — Justin Joyce, Cloudflare
Aug 26, 2026 · 19:15
Cloudflare principal sales operations manager Justin Joyce argues traditional go-to-market does not scale, so his team built a three-pillar agentic system on Cloudflare Workers and Durable Objects. To scale analysis, role-specific skill files let non-SQL users query data, cutting two-hour analyses to five minutes. To scale insight, weekly summaries are drafted by one agent, verified by a second, and toned by a third, tested on every run for two to three months. For self-service, Cloudflare OS gives reps forecast briefs, QBR decks, account plans and renewal prep with centrally curated expert skills. The result is 2X efficiency, with harder problems ahead in quoting, approvals and CRM writes.

Knowledge Systems: The New GTM Stack — Jeffrey Wang, Exa
Aug 26, 2026 · 18:49
Jeffrey Wang, cofounder of Exa, argues go-to-market is an AI engineering problem and details Exa's stack: an ICP dashboard classifying nearly every company in its addressable market with anticipated spend, and Request Lens alerting on meaningful customer events. The team runs about a dozen Slack agents plus Jeffbot, an AI clone of Wang trained on 760 emails—he averages 18 words and signs 'best' not 'sincerely'—and limited to drafts when others use it. He closes on three principles: agent-first means API-first, not everything should be a chatbot, and buy-versus-build is false—Salesforce exposed as MCP is arbitrarily customizable. He also notes an eight- or nine-person FDE org runs deals and builds the sales systems.

Building GTM AI Agents: Lessons from Deploying to 6,000 Users — Sait Izmit, Snowflake
Aug 26, 2026 · 20:39
Sait Izmit of Snowflake says winning over 6,000 go-to-market users comes down to quality over coverage: he wrote 150 sales questions before testing, accepted 50% first accuracy, and chose 50 questions at 95% over 100 at 70%. The Snowflake Cowork agent, live since September, has answered over a million questions (~40,000 weekly); 60% of data arrived post-launch, and it now spans 15 semantic views, 85 tables, 3,000 columns, MCPs, and 20 skills. After pilot and a 600-user 10% beta with 70% retention, GA showed only 20% of the org tried it, so Izmit spends 60-70% of his time on demos and sales meetings. He warns the wow factor collapses quickly, so teams must move from data chat to workflow automation, build fast with today's stack, accept rearchitecture, and mine logs for feedback loops.

The Building Blocks of GTM Orchestration — Arman Vaziri, Ramp
Aug 26, 2026 · 19:55
Arman Vaziri, who leads product and sales led growth engineering at Ramp, details go-to-market orchestration: an intent — offering Pro V1 golf balls to golfers at East Coast construction companies — becomes targeted outbound, paid creative, landing page and in-app nudges. He argues the bottleneck was never ideas but messy data, rep busy work, and coordination cost; Ramp's answer was an internal CDP on Kafka and Postgres, plus embedded unstructured data for search. Pre-meeting briefs for account managers map attendee emails to accounts and run as durable Temporal threads; a skill library for custom brief formats drove adoption. For smaller teams, build narrow automations first — three years ago two people used GPT-3.5 for outbound — because nobody gets a year to design a perfect architecture.

AI in GTM at Notion — Flora Liu
Aug 26, 2026 · 21:15
Flora Liu, an engineer on Notion's GTM team, argues GTM is now a distributed systems problem, not marketing ops, and details how Notion built a unified decisioning system across self-serve and sales-assist. Workflows reduce to four questions — what do we know, what should happen next, how to execute safely, did it work — built as Know, Decide, Act, Learn layers. Snowflake computes the profile, DynamoDB serves it in milliseconds, and Notion is the substrate where humans and agents operate on the same context; signals become Temporal workflows. Agents research and draft, humans approve customer actions, and contact-sales forms are untrusted input. After 13 weeks, enterprise reps log more qualified opportunities, and context-aware recommendations made users 63% more likely to take next step.

Reverse-Engineering the AI Buyer — Aliisa Rosenthal, Acrew Capital
Aug 26, 2026 · 19:10
Aliisa Rosenthal, who helped take OpenAI's enterprise business from a couple million to several billion in revenue, argues founders should build the automated sales machine before hiring humans. ChatGPT launched with no enterprise features; nine months later they shipped an expensive product, but self-serve cannibalized it four months on. Advice: capture phone numbers at signup, reply to every inbound (10,000 a day), never give buyers homework, avoid pilots via reference calls, data evals, or 90-day opt-outs. $60 per user per month was too high; a low base fee plus usage spread. Automate security, don't pull up-market, hire AI-native sellers when buyers need contact, and use forward deployed engineers sparingly—scarce but sticky. First 10 customers should be handpicked design partners, and POCs aren't self-serve.

Why Your Enterprise Tech Stack Isn’t Ready for AI Agents — Christopher Lovejoy & Saul Howard
Aug 19, 2026 · 19:15
Christopher Lovejoy of Anthropic and Saul Howard of Anteria argue that enterprise stacks aren't ready for AI agents; regulated industries need primitives, not bolted-on compliance. An audit trail isn't a developer log: under HIPAA it must record every action, data access, and authorization, so they use an immutable append-only event log. Patient data lives in schema-driven object storage referenced by events, letting devs debug without PHI, enabling zero trust against prompt injection. Escalation treats humans and models as equivalent agents, enabling the same actions by either. These primitives make privacy-preserving evals a byproduct, enabling replay of production data and evaluation in customer environments without exposing data; take constraints first and rebuild toward POC accuracy.

Healthcare’s Agent Bytecode: X12 as the Harness for AI Agents — Vasant Kearney, Onlay
Aug 19, 2026 · 20:25
Vasant Kearney of Onlay argues that X12, the healthcare claims data standard, is the harness for AI agents: every step of the claim lifecycle—from eligibility (270) through acknowledgment (999) to payment (835)—has an X12 correspondence, and an agent calling a payer or driving a portal is emitting the same transaction by another route. He says phone, portal, and X12 surfaces are built by different teams and can all agree on the wrong answer, so no surface is ground truth; his system keeps a semi-correct internal X12 representation, correct only until downstream evidence says otherwise. Enterprise memory must live in a database rather than local disk, and swapping in a stronger model is not automatically better inside a system built around the old one—evals and validation must be redone. The goal is cutting insurance costs and improving patient experience, and that means being AI pilled but also AI skeptical.

Local Models: Trust, Control, Optimization — Carter Abdallah, NVIDIA
Aug 7, 2026 · 43:21
NVIDIA's Carter Abdallah, Prime Intellect's Vincent Weisser, Arcee's Lucas Atkins, and NVIDIA's Chris Alexiuk argue open-weight models are the trustworthy foundation for enterprise and local AI. Atkins separates trust from safety: when Anthropic pulled Fable, enterprises chose Chinese open models for guaranteed availability, and open models are inspectable unlike closed APIs. Arcee pretrained a 400B model in six months; Weisser cites a customer that specialized an open model for finance in a week or two, beating Opus at a fraction of Haiku's cost. Alexiuk calls open weights the fix for 'mismanaged genius' and expects capable local models on MacBooks within a year; the panel predicts Fable-level open models within a year and hopes local-model use rises from a rounding error to 10–15%.

AI tools for Forward Deployed Engineering — Vasuman Moza, Varick Agents
Jul 28, 2026 · 20:23
Varick Agents CEO Vasuman Moza and head of engineering JD Pruitt explain how forward deployed engineers (FDEs) design AI agents that sit on top of existing enterprise systems like SAP or NetSuite rather than requiring migrations—one customer spent $5M and five years on NetSuite, so Varick drops agents onto those systems. The bottleneck is understanding each business’s unique, undocumented workflows (e.g., when AP fails, Sarah sends to Chris, adding four days of cycle time). FDEs map these processes, re-engineer them around AI (automating four of eight steps, keeping human-in-the-loop for three, fully human for one), then deploy using Varick OS. To scale FDEs without hiring exponentially, Varick built an AI FDE agent that ingests granola notes and Slack threads, uses a Postgres-based dependency graph as a single source of truth, and post-trains open-source models (Kimi K26) to extract the right context and strip redundancy. JD outlines three stages: an engagement agent for querying documentation, a workflow agent that shadows FDEs inside the platform, and a future autonomous agent that handles client change requests (e.g., rerouting a QC report) without FDE involvement.

How Forward Deployed Engineering is done at Cognition — Jia Wu
Jul 28, 2026 · 17:38
Jia Wu, deployed engineering lead at Cognition, argues that their forward deployed engineers measure outcomes—not token usage—by embedding Devin in customer environments and delivering measurable productivity gains like an 82% reduction in delivery timelines. The job involves deeply understanding customer problem spaces, mapping Devin’s capabilities to highest-leverage work, and feeding product feedback from deployments. Wu distinguishes Cognition from single-point tools by claiming organizations achieve 10x faster output, not just individual engineers. He cites anonymized case studies: 150% headcount increase in three months, double the PRs compared to single-point tools, and one-third the timeline for an ETL migration at a bank. Wu insists that as coding becomes commoditized, deployed engineers must prioritize business and people skills, relentlessly tying into customer success and communicating back to the product roadmap.

How Forward Deployed Engineering is done at Decagon — Sunny Rekhi
Jul 28, 2026 · 18:09
Sunny Rekhi, CTO of Forward Deployed Engineering at Decagon, explains that forward deployed engineering is identical to product engineering, with two kinds of work: configuring the AI agent's brain (instructions and handoff rules) and solving customer asks in a way that scales to all customers. Decagon, which builds 24/7 AI customer service agents, grew from 50 to 500 people in a year, breaking the original role into specialized lanes: agent builders who configure within the UI, and agent software engineers who productize customer requests. Rekhi stresses restraint—avoiding one-off patches—and proving value fast in the first weeks of a partnership. A key ethos is that custom work never stays custom: every integration built for one customer gets upstreamed into the platform, becoming self-serve for the next. Forward deployed engineers also act as advisors, using their cross-customer knowledge to guide enterprises on where they will see the highest ROI based on historical data.

Can Oncology Workflows Run Without Human Touch? - Anant Shankhdhar, Risa Labs
Jul 20, 2026 · 16:41
Anant Shankhdhar, an AI engineer at Risa, explains how his team automates oncology workflows end-to-end using four AI agents—EV, Auth, Necessity, and Submission—to eliminate human touch in prior authorization processes. The EV Agent handles eligibility and benefits verification via a unified service that connects to payer APIs and RPA portals, using LLM-driven config generation and self-healing loops to scale. The Auth Agent determines drug authorization status by reconciling evidence from patient notes, authorization letters, and a payer rule knowledge base, enabling no-touch handling for drugs that are already authorized or don't require authorization. The Medical Necessity Agent answers clinical questions per patient, attaching confidence scores and escalating only cases needing human review. The Submission Agent submits orders to payers using customized integrations. Risa's agents are deployed across 20+ hospitals, supporting care for over 100,000 patients, and the no-touch share…

Forward Deployed Engineering at Cursor — Pauline Brunet
Jul 14, 2026 · 20:47
Pauline Brunet, VP of Forward Deployed Engineering at Cursor, explains how to build an effective FDE team by mapping customer digital maturity and product customization to the right engagement model—self-service, traditional SaaS deployment, advisory, or embedded transformation. She defines the ideal FDE as a highly technical engineer with strong customer-facing skills, capable of scoping impactful projects, co-developing with customer teams, and feeding insights back to product. Cursor’s FDE team uses a project-based approach, hires experienced unicorns (5+ year software engineers), and structures around geographies, industries, and product SMEs. Best practices include solving the right problem, defining success metrics upfront, keeping scope directional, involving customers at every step, measuring ROI (revenue, cost, risk), and leaving documentation behind. Brunet advises learning fast, pivoting, partnering with system integrators, and not being afraid to say no to mismatched use cases.

AI-Driven Multi-Document Correlation for Financial Compliance - Varsha Shah, Independent
Jun 28, 2026 · 19:00
Varsha Shah presents an AI-driven framework combining graph-based entity correlation, adaptive probabilistic risk modeling, and cross-jurisdictional normalization to detect hidden compliance risks across payroll, tax, procurement, and financial records. Evaluated on approximately 3 million anonymized records across four jurisdictions, the framework achieved 91% precision, 87% recall, and an F1 score of 0.89, while reducing false positives by 76% and manual audit effort by 40%. Shah argues that many sophisticated fraud patterns exist between documents rather than within them, and that traditional rule-based systems analyzing documents in isolation fail to capture these cross-document risks. The framework enables continuous learning from audit outcomes, shifting compliance from reactive validation to predictive, intelligence-driven governance. Key to deployment are seamless integration with existing enterprise systems, jurisdiction-specific configuration, alignment with audit frameworks, and scalability to process millions of records.

AI System Design: From Idea to Production - Apoorva Joshi, MongoDB
Jun 28, 2026 · 28:53
Apoorva Joshi presents a structured framework for designing AI systems from idea to production, arguing that thorough product specification and evaluation are now more critical than coding itself. Using a health insurance claims review system for MDB Health as an example, she walks through four phases: defining business problems with measurable success metrics (e.g., reducing urgent claim processing from 2 days to 1 hour within 90 days), designing data strategy and architecture using patterns like RAG, control flow, and human-in-the-loop, establishing guardrails and evaluation metrics such as faithfulness and cost per recommendation, and optimizing for accuracy, cost, latency, and reliability before shipping. She emphasizes building evaluation in from the start and iterating from the simplest system.

Why Can't Anyone Answer Questions About the Business? — Garrett Galow, WorkOS
Jun 11, 2026 · 19:06
Garrett Galow from WorkOS built Studio, an internal workspace where anyone can ask natural language questions against Snowflake, Linear, and Notion, and get reusable widgets instead of filing a request. The LLM generates declarative JavaScript widgets that call data sources directly, making subsequent runs deterministic and cheap. Three techniques made it reliable: preflight sequencing injects schema context only when a tool is invoked, a layering rule tells the model to distrust its own knowledge about WorkOS and use primary sources, and query validation catches valid SQL that returns zero rows before hardcoding it into a widget.

How to Leverage Domain Expertise — Chris Lovejoy, Notius Labs
May 16, 2026 · 24:45
Chris Lovejoy argues that winning in vertical AI is an organizational problem solved by domain experts acting as Oracle (directly improving AI), Evaluator (defining metrics for engineers), or Architect (building self-improving systems). Granola's first employee, a writer, reviews meeting notes and tweaks prompts as an Oracle because there is no objectively perfect note. Tandem used decentralized Oracles—doctors per specialty and country—to handle variation in medical scribe outputs. Anteria progressed from Oracle to Evaluator to Architect as prior authorization required measurable correctness and automated learning from usage variation. Lovejoy advises hiring a principal domain expert early, giving them ownership, and hiring for breadth (domain expertise plus adjacent skills like data science or engineering) to avoid slow progress and turnover.

Contact Center Voice AI: Low-Latency Intelligence Extraction from Messy Audio Streams — Dippu Singh
Apr 8, 2026 · 22:56
Dippu Singh of Fujitsu North America presents an architecture for real-time voice intelligence in contact centers that reduces post-call work by 50% through structured intent extraction from messy audio streams. The pipeline comprises four stages: voice capture with stereo channel splitting and PII masking, speech-to-text with domain-specific dictionaries, a generative AI core that uses system prompts to output separate JSON bullet points for customer intent and operator actions, and a customer data sync layer that maps LLM output to CRM fields via REST APIs. Key results show after-call work (ACW) dropped from 6.3 to 3.1 minutes, while data entry quality became standardized. Current constraints include STT accuracy for heavy accents, API token costs for long transcripts, and security compliance overhead. The roadmap targets explainable AI for agent coaching, predictive staffing from categorized intent data, and real-time abusive-call detection to protect operators.

AI That Pays: Lessons from Revenue Cycle — Nathan Wan, Ensemble Health
Jul 24, 2025 · 18:19
Nathan Wan, head of AI at Ensemble Health Partners, argues that Revenue Cycle Management (RCM) is a critical yet overlooked domain for AI disruption in healthcare, where 40% of hospitals operate at negative margins due to broken manual financial processes. He explains that RCM's vast structured and unstructured data, rule-based workflows, and direct financial impact make it ripe for AI, contrasting it with more publicized clinical AI applications. Wan details how Ensemble uses AI to predict and prevent claim denials—often caused by technical errors like missing data—and to accelerate clinical denial appeals via GenAI, achieving a 40% reduction in processing time and improved overturn rates. He emphasizes that Ensemble's end-to-end view of the revenue cycle, supported by their EIQ data platform, enables longitudinal error correction and agent-driven automation, turning upstream fixes into measurable ROI for providers facing increasing payer denial rates.

How agents will unlock the $500B promise of AI - Donald Hruska, Retool
Jul 23, 2025 · 16:22
Donald Hruska, engineering lead for Retool's Agents product, argues that AI agents will unlock the $500B promise of AI by moving enterprises beyond toy chatbots into production-grade systems. He explains that building a basic agent is easy (e.g., 100 lines of code using React) but getting it into production requires solving security, cost overruns, observability, and compliance. Hruska outlines four options—build from scratch, use a framework like LangGraph, a managed platform like Retool Agents, or verticalized agents—and advises building for core differentiators and buying for commodity workflows. He cites Retool customers like ClickUp saving over $200,000 in vendor costs and Descript saving hundreds of hours weekly, while noting inference costs dropped 99.7% from 2022-2024. Retool charges $3 per hour for its cheapest agent and supports on-prem deployment.

POC to PROD: Hard Lessons from 200+ Enterprise GenAI Deployments - Randall Hunt, Caylent
Jul 23, 2025 · 19:16
Randall Hunt from Caylent shares hard lessons from over 200 enterprise GenAI deployments, arguing that evals, embeddings, and prompt engineering matter far more than fine-tuning. He emphasizes that speed and UX are critical; a slow inference kills adoption, while techniques like generative UI and caching can mitigate latency. Hunt details real-world examples: using audio amplitude spectrographs for sports highlight reels, pooling multimodal embeddings for nature footage search, and noting that nurses prefer chat over voice bots in noisy hospitals. He reports zero regressions moving from Claude 3.7 to 4, and advises optimizing context and economics, such as leveraging Amazon Bedrock batch for 50% cost reduction. The talk underscores that knowing your end customer and minimizing irrelevant context are key to production success.

Building agent fleet architectures your CISO doesn't hate — Lou Bichard, Gitpod
Jun 27, 2025 · 13:52
Lou Bichard explains how Gitpod evolved from a managed SaaS to a 'bring your own cloud' architecture that satisfies CISOs in regulated industries by running secure dev environments—and now agent fleets—on customer infrastructure via a simple runner (a single ECS task) instead of complex Kubernetes. The platform, used by banks and healthcare firms for 37 hours per week per developer, reduces operational overhead through a cloud-formation-based setup that takes three minutes. For agents, the same infrastructure provides source code access and audit logging, ensuring privacy and compliance. Bichard argues that vendors should simplify architectures to lower customers' day-two costs, and advises buyers to prioritize security and ownership models when selecting AI tools.

Anthropic in the Enterprise — Alexander Bricken & Joe Bayley
Apr 13, 2025 · 20:55
Alexander Bricken and Joe Bayley from Anthropic's Applied AI team argue that enterprise AI implementation often fails due to overengineering, poor data infrastructure, or lack of testing—but industry leaders achieve transformative results with Claude. They detail Anthropic's deployment models (API, cloud partnerships, enterprise solutions) and real-world case studies like Intercom's Fin agent, which solved 86% of support volume using Claude. Best practices include building evals early as intellectual property, identifying intelligence/cost/latency trade-offs based on use-case stakes, and avoiding premature fine-tuning by trying prompt caching, contextual retrieval, and agentic architectures first. They also highlight interpretability research and the Model Context Protocol for reliable AI deployments.

Accelerate your AI journey with Azure AI model catalog: Sharmila Chokalingam
Feb 6, 2025 · 23:14
Shubhi and Sharmila present the Azure AI model catalog as a platform offering over 1,600 models including GPT-4, Mistral, Llama, Cohere, and Phi3, with a standardized inference API enabling easy model swapping. They demonstrate deployment via serverless API (pay-per-token) and managed compute, emphasizing that customer prompts and completions are not shared with model providers or used for training. The platform includes model benchmarks, playground for RAG, and Prompt Flow for building generative AI apps with evaluation and variant comparison. Customer success stories include EY's EYQ chat adopted by 275,000 employees, CMA CGM's reduced response latency using Mistral, and Bridgestone's 30% reduction in forecasting errors with Nixtla's TimeGen.

Real ROI: Lessons from Enterprises that have already succeeded with LLMs at Scale: Raza Habib
Dec 31, 2024 · 20:01
Raza Habib, CEO of Humanloop, shares lessons from enterprises like Duolingo, Filevine, and Ironclad that have achieved real ROI with LLMs at scale. He argues that success hinges on centering domain experts (e.g., Duolingo's linguists do all prompt engineering), needing less ML expertise than expected, and breaking down evaluation into small, testable components rather than chasing a single metric. Habib emphasizes collecting end-user feedback (acceptance rates, edits, thumbs up/down) and building tooling for logging, regression testing, and team collaboration. Filevine doubled revenue by launching six LLM-powered products; Ironclad's open-source Rivet logging tool enabled agents to auto-negotiate 50% of contracts. The talk covers how to optimize the four-component chain (base model, prompt, data selection, function calling) and why simpler systems will win as models improve.

Build enterprise generative AI apps using Llama 3 at 1,000 tokens/s on the SambaNova AI platform
Sep 11, 2024 · 54:34
SambaNova’s Michelle Matern and Petro Milan present their full-stack AI platform, demonstrating how the SN40L RDU chip enables Llama 3 inference at 1,000 tokens per second. They introduce Samba-1, a composition of 92 expert models behind a single endpoint, and benchmark it against GPT-3.5 and GPT-4 on enterprise tasks like information extraction and text-to-SQL. The workshop then builds a RAG-based Q&A system using LangChain, Unstructured, E5-large-v2 embeddings, ChromaDB, and Llama-3-8B-Instruct at 1,000 tokens per second. Attendees set up the environment, load documents, and run inference with real-time metrics showing time to first token of 0.09 seconds and total inference time of 0.65 seconds.
Powered by PodHood