Episodes from AI Engineer about AI Strategy.

The Signal Layer: What to Build When Anything Can Be Built — Lena Hall, Akamai
Aug 29, 2026 · 19:44
Lena Hall of Akamai says when AI lets everyone build anything, the scarce skill is choosing what to point it at and keeping it undistorted — a 'signal layer' — since competitors get the same leverage the same morning. She calls AI a 'convergence machine' that makes average work worthless, and rejects taste as a trainable differentiator; what survives is judgment about events not yet happened and relationships models can't observe. Citing Hamming, she says AI gave everyone an attack on every problem, making the rare skill choosing which problem deserves one. Signal distorts via founders compressing context, org layers rerounding to average, and AI remixing claims into promises — a 94% eval became a promise. Trust has no grader; define and protect your signal and use AI for the rest.

From Tokenmaxxing to Trusted Throughput — Mingsheng Hong, Ironclad
Aug 29, 2026 · 23:04
Mingsheng Hong, VP of Engineering focused on AI at Ironclad, argues token dashboards are smoke detectors, not leaderboards, and the goal is trusted throughput—merged, customer-validated PRs weighted by complexity—not token minimization. Don't cut cost before measuring value; Ironclad's metric evolved from lines of code to open PRs to merged PRs to merges weighted by an LLM-assigned complexity score, because a ten-line concurrency fix beats boilerplate. He identifies review and CI as new bottlenecks, where slow pipelines encourage giant batched PRs; he recommends AI as first-pass reviewer, killing flaky tests, capping agent retry loops, and measuring ready-to-merged time. He advises prompt caching, context pruning, and buying infrastructure while building context-specific playbooks in-house.

The Half Life of Agent Infrastructure — Ben Kus, Box
Aug 29, 2026 · 19:26
Ben Kus, CTO of Box, explains why AI agent infrastructure has a half-life of months, not the three-to-five years typical of enterprise software. He retracts his own graph-based agent approach from last year's talk, and recounts asking an engineer to rebuild working agentic search twice in quick succession. At Box's scale — over an exabyte of data and around a trillion tokens — models, agent design, and retrieval have each shifted repeatedly. His advice: prepare teams to expect change, build abstractions that let underlying layers be swapped, and switch only when eval sets show measurable gains in cost, speed, quality, or capability. Box now reviews every AI technology on a six-month clock, and he suggests judging vendors by how well they handled past changes.

Which AI startups actually land enterprise contracts? — Brian Lewis, Millennium
Aug 29, 2026 · 18:45
Brian Lewis of Millennium, a hedge fund, explains why only ~5% of AI vendor demo calls end in signed contracts: on the buying side, 10-15 startups per pain point become two or three demos, zero or one pilot, and one contract per four pilots. He catalogs failures—a vendor asking customers to self-report gateway telemetry so it could charge margin on traffic it never carried, a 'zero retention' vendor revealing data it shouldn't have, betas engineered around permissive retention, and a core API down for hours on a trading day. His larger claim: 60% of becoming AI-native is unsexy work—entitlements, integration, change management—and agents inherit whatever foundations you have, so fix the boring 60% first. He advises startups to treat Millennium-grade security and ops as the bar.

How do you diffuse AI into the real world? — Varun Shenoy, Long Lake
Aug 28, 2026 · 17:46
Varun Shenoy, co-founder of Long Lake, argues that AI diffusion into real-world services requires operator-owners, not vendors, and explains how his firm buys and runs 35 services businesses — including a $6.3 billion take-private of American Express Global Business Travel — to make agents complete economically relevant tasks. He frames adoption as a generational shift like electricity, and details a ladder from copilots to async agents to AI coworkers, earned by proving value. Long Lake represents knowledge work as code, captures traces and ground truth from work like roof repairs and book closing to build evals, and treats continual learning and enablement as one snowball loop. He concludes that co-designing software with a 100-year-old firm cannot happen over Zoom; it requires showing up in person, running conference stands, and asking questions on a mountain bike.

Knowledge Systems: The New GTM Stack — Jeffrey Wang, Exa
Aug 26, 2026 · 18:49
Jeffrey Wang, cofounder of Exa, argues go-to-market is an AI engineering problem and details Exa's stack: an ICP dashboard classifying nearly every company in its addressable market with anticipated spend, and Request Lens alerting on meaningful customer events. The team runs about a dozen Slack agents plus Jeffbot, an AI clone of Wang trained on 760 emails—he averages 18 words and signs 'best' not 'sincerely'—and limited to drafts when others use it. He closes on three principles: agent-first means API-first, not everything should be a chatbot, and buy-versus-build is false—Salesforce exposed as MCP is arbitrarily customizable. He also notes an eight- or nine-person FDE org runs deals and builds the sales systems.

The Building Blocks of GTM Orchestration — Arman Vaziri, Ramp
Aug 26, 2026 · 19:55
Arman Vaziri, who leads product and sales led growth engineering at Ramp, details go-to-market orchestration: an intent — offering Pro V1 golf balls to golfers at East Coast construction companies — becomes targeted outbound, paid creative, landing page and in-app nudges. He argues the bottleneck was never ideas but messy data, rep busy work, and coordination cost; Ramp's answer was an internal CDP on Kafka and Postgres, plus embedded unstructured data for search. Pre-meeting briefs for account managers map attendee emails to accounts and run as durable Temporal threads; a skill library for custom brief formats drove adoption. For smaller teams, build narrow automations first — three years ago two people used GPT-3.5 for outbound — because nobody gets a year to design a perfect architecture.

AI in GTM at Notion — Flora Liu
Aug 26, 2026 · 21:15
Flora Liu, an engineer on Notion's GTM team, argues GTM is now a distributed systems problem, not marketing ops, and details how Notion built a unified decisioning system across self-serve and sales-assist. Workflows reduce to four questions — what do we know, what should happen next, how to execute safely, did it work — built as Know, Decide, Act, Learn layers. Snowflake computes the profile, DynamoDB serves it in milliseconds, and Notion is the substrate where humans and agents operate on the same context; signals become Temporal workflows. Agents research and draft, humans approve customer actions, and contact-sales forms are untrusted input. After 13 weeks, enterprise reps log more qualified opportunities, and context-aware recommendations made users 63% more likely to take next step.

Reverse-Engineering the AI Buyer — Aliisa Rosenthal, Acrew Capital
Aug 26, 2026 · 19:10
Aliisa Rosenthal, who helped take OpenAI's enterprise business from a couple million to several billion in revenue, argues founders should build the automated sales machine before hiring humans. ChatGPT launched with no enterprise features; nine months later they shipped an expensive product, but self-serve cannibalized it four months on. Advice: capture phone numbers at signup, reply to every inbound (10,000 a day), never give buyers homework, avoid pilots via reference calls, data evals, or 90-day opt-outs. $60 per user per month was too high; a low base fee plus usage spread. Automate security, don't pull up-market, hire AI-native sellers when buyers need contact, and use forward deployed engineers sparingly—scarce but sticky. First 10 customers should be handpicked design partners, and POCs aren't self-serve.

Coding Agents Don't Scale Themselves. Neither Do Your Teams. — Patrick Debois, Tessl
Aug 22, 2026 · 22:06
Patrick Debois of Tessl argues that the 'dark factory' of autonomous coding agents won't work in most organizations because teams and platforms aren't set up for it, not because the technology fails. Developers who rebelled at 'writing better prompts' re-engaged once tooling opened a technical path; skeptics are ideal for context authoring, and retros target repeated agent failures, not code. Planning splits into well-scoped agent tasks and conversational human work; Debois tracks two metrics: human touches per result (down) and shared fixes that help everyone, not one 10x person. Platform teams must own paved roads, registries, eval systems, and spend visibility; that's the shift from solo developer to multiplayer system, and hiring tests AI leverage, engineering taste, and collaboration.

Preferences Over Benchmarks: Model Routing — Archana Kamath & Tyler Gillam, DigitalOcean
Aug 22, 2026 · 15:54
DigitalOcean's Archana Kamath and Tyler Gillam argue that no single best AI model exists, so choosing models by benchmark leaderboards is the wrong instinct; the right model depends on the request's task, cost, latency, system prompts, and end-user preferences. They demo their open-source inference router, built into DigitalOcean's AI-native cloud, which uses a purpose-built mixture-of-experts model to decide in under 200 milliseconds, costs nothing extra, and needs no code changes. In a live coding-agent comparison, the router matched premium quality while spending 14 cents versus 44 cents over a session and scored 90% correctness versus Opus's 95% with fewer tokens and faster speed. They frame routing as a foundation layer for evals, caching, and personalization, with no vendor lock-in.

Your Fine-Tuned Model Is Tech Debt: A 50x ROI House of Cards — Dan Bjornn, Lease End
Aug 20, 2026 · 16:39
Dan Bjornn, senior data scientist at Lease End, explains why his team's fine-tuned LLM for customer intent—despite bringing in $12 million at a 50x ROI—became tech debt. The fine-tuning pipeline took a week per retrain, with training itself the shortest step, and each fix caused regressions, so bugs were triaged by tolerable customer pain. He calls this the calcification tax: the model locked them into one provider and an outdated architecture, preventing upgrades. The rebuild swapped the tuned model for skills, prompts, and context on a model-agnostic framework, letting fixes ship in under an hour as uploaded files. Accuracy went up, cost per message rose, but total cost fell. Bjornn concludes fine-tune only when you cannot call a frontier model, and even then the decision must beat the tax.

Prototyping as Leadership: How a CTO Ships with AI Agents — Hursh Agrawal, The Browser Company
Aug 20, 2026 · 18:18
Hursh Agrawal, CTO and co-founder of The Browser Company, argues AI coding agents let leaders keep building: despite 15+ meetings, 7 direct reports, and a toddler, he ships 2–10 PRs/week. He says frontier models turn over every three months, so hands-on building is the only way to judge them and show engineers working prototypes. His method is an overnight loop: a coworker agent gathers Slack/Jira/Notion context into a prompt at 5pm, coding agent runs for hours, and a morning hour reviews tests, CI, and AI code review. He details three overnight uses: building features, hill-climbing evals from feedback JSONs, and training custom models like a PII classifier on AWS. He warns leaders to avoid critical path work and to rely on trustworthy CI, feature flags, a prototype branch, and readable PRs.

How to build an AI-Native Health Company — Dan Feng, Maven Clinic
Aug 19, 2026 · 17:19
Dan Feng of Maven Clinic says becoming AI-native means rewiring hiring, planning, code review, and reliability, not just adding tools. Maven hires engineers who solve ambiguous problems independently, keeps a one-year vision only directional, commits to two-to-four-week sprints, and treats three-to-six-month plans as unplannable because models change too fast. Engineers now write thousands of lines daily, so PRs are capped at 500 lines, engineers can self-certify simple changes but stay accountable, and rubber-stamp approvals are false confidence. Reliability is tiered: a scheduling failure one out of 1,000 times is tolerable because users retry, but reimbursement claims require multiple models to agree on the same receipt and integration tests to pass repeatedly.

Open Source Is Dead. Long Live Open Source. — Saoud Rizwan, Cline
Aug 7, 2026 · 17:30
Saoud Rizwan, founder of Cline, argues that AI has killed the community side of open source while open weights models win on economics. He cites Zig banning AI from PRs, curl weighing shutdown of its bug bounty over AI-generated reports, and tldraw auto-closing pull requests, plus a LiteLLM compromise that stole credentials for three hours. Rizwan makes the case that closed labs' subsidized subscriptions lead to lock-in and price gouging, while open models like GLM match or beat Opus on cost and code quality — GLM used twice the tokens at half the cost and fixed a real Cline bug that Opus's faster fix left broken. He compares open weights to Facebook's Open Compute project and urges American labs to release open weights before foreign models become the standard.

Why Off-the-Shelf AI Doesn't Understand Money — Udi Menkes, Intuit
Jul 29, 2026 · 19:50
Udi Menkes, principal PM at Intuit, argues that off-the-shelf frontier models deliver a 'fluent bluff' when advising on money: advice that sounds right but is dangerous because models have read about money but lack experience. He shows a rental property example where a frontier model told a landlord in negative cash flow to acquire a second property, while a model grounded in real outcomes recommended raising rent 5-10%. Intuit's head-to-head test across 100,000 businesses found frontier models gave advice that would harm businesses 40% of the time, while a mid-sized grounded model outperformed them by training on millions of state-action-outcome records from QuickBooks data. A Princeton study confirmed frontier models given $1M went bankrupt within 500 days, while a simple rule-based system beat them. Menkes says the moat belongs to whoever owns the best system of context, and advises leaders to find verified outcomes in their own data to ground AI.

The Dirty Secret of Forward Deployed Engineering — Natalie Meurer, Sierra
Jul 28, 2026 · 16:49
Natalie Meurer argues that forward deployed engineering (FDE) has become a meaningless label because it has stretched from DevOps at Palantir in 2008 to data integration, ontology work in Slate and Foundry, solution architecture, and enablement—but its durable core is customer accountability and outcome-based pricing. Tracing FDE's history at Palantir, she shows how the role evolved from keeping the platform stable (2008) to data integration (2012), custom dashboarding in Slate (2016), and finally customer enablement in Foundry (2020). As coding agents make software cheap, she contends the lasting value lies in integrating data, understanding customers, and owning outcomes. Pricing tells the story: seat-based assumes a tool, while usage or outcome pricing puts the provider on the hook—exactly what FDEs have always done. She concludes agent engineering is FDE reborn, and that product, infra, and AI engineering are all trending toward the same customer-accountable model.

How Forward Deployed Engineering is done at Kepler — Vinoo Ganesh
Jul 28, 2026 · 22:20
Vinoo Ganesh, a former Palantir engineer who built the Frontline rotation program, argues that Forward Deployed Engineering is a product strategy, not a go-to-market motion, and shows how Kepler applies this philosophy. He illustrates with stories from Palantir: solving a shipping customer's 47-page requirements with a four-hour Slack alert, building a Parquet viewer after watching a data quality engineer manually spot-check CSVs, and the Groovy script that became a product supporting 100,000 people. Ganesh emphasizes detecting real problems by observing users' actions—any repeated task hints at a missing feature, and pulling out a phone mid-workflow is a bug report you'll never find in documentation. He also explains defining ontology: when different teams call the same entity 'clients,' 'billing,' or 'accounts,' the FDE must canonicalize terms to become the linguistic foundation. The hardest skill is discarding—ship everything as if it will run 18 months, because every hack goes into production. Ganesh concludes that FDEs drive product leverage by solving small problems on-site, then generalizing solutions into the core product.

Notion's Token Town — Sarah Sachs, Notion
Jul 23, 2026 · 23:55
Sarah Sachs, Notion's AI engineering lead and contract negotiator, argues that AI companies must stop competing on token economics and instead build model-agnostic products that win on data flywheels, orchestration, and security. She advises treating every model supplier as a competitor, because frontier labs charge a markup on a markup for tokens they sell for first-party use. Notion's auto model routes 75% of traffic through a Switzerland-like system that swaps providers underneath, avoiding vendor lock-in. Sachs advocates routing by cost per capability per second, using open weight models for the moderate middle, and reaching for CPUs over GPUs (e.g., no LLM needed to turn a CSV into a PDF). She highlights the 'lethal trifecta' of private data, untrusted content, and external communication as the next security challenge, and demos Notion agents scoping a task, tagging teammates, and opening a PR. Her core message: optionality is leverage, and the product must transcend tokens.

Stop Renting Your Cognitive Infrastructure - Thiyagarajan Maruthavanan, Kalmantic Labs
Jul 18, 2026 · 7:52
Thiyagarajan Maruthavanan of Kalmantic Labs argues that AI teams should stop renting inference from providers like Anthropic and build their own infrastructure, coining "rent to learn, own to earn." He recounts his app UltraSuno costing hundreds of thousands of dollars in inference, a stolen key hitting $10,000, and moving to his own DGX Sparks hardware. Three enterprises—a fund, hospital, and tax practice—each hit walls with renting: control, audit redlining, and reproducibility. He open-sourced JustTokenMax (benchmarked better than Netflix's Headroom) and wrote the book "Peak Inference" on building your own inference infra. He notes the market's conflicting pitches (Jensen's token factory, Nadella's unmetered, NeoClouds' endpoints) but concludes owning is essential post-PMF.

Everything we knew about software has changed — Theo Browne, @t3dotgg
Jul 8, 2026 · 16:02
Theo Browne argues that AI model evolution—from Sonnet 3.5’s tool-calling to Opus 4.5's long-running tasks and Mythos's orchestration—requires engineers to think bigger and wider. He compares current developer habits to skeuomorphism in iOS 7, urging rejection of legacy constraints like Git's inability to commit environment files and terminal-centric workflows. Browne introduces a shifted tier system: what was a startup is now a side project; a Markdown file running on a cron job can replace a company's product. His own PR triage service became a Markdown file updated daily via cron. He advocates building breadth over depth—architecting products so users can extend features, enabling small teams to compete with AWS or Salesforce. 'If your idea doesn't feel stupid, it's because your idea is not big enough,' he concludes.

Sovereign Escape Velocity: Ownership w Open Models — Gus Martins, & Ian Ballantyne, Google DeepMind
Jun 10, 2026 · 20:52
Gus Martins and Ian Ballantyne of Google DeepMind introduce Gemma 4, a family of open-weight models that deliver high quality per parameter, enabling deployment on a single GPU or even a phone. They argue that the models' efficiency — a 31B model rivals those twenty times larger — and the shift to Apache 2.0 licensing remove barriers for sovereign institutions like those in Ukraine, Bulgaria, and Brazil. Ian demonstrates multi-agent translation running locally on an M4 Mac, showcasing ownership and control over agentic workloads.

20 days of compute vs 7 hours: rethinking what state-of-the-art means — Bertrand Charpentier, Pruna
Jun 1, 2026 · 19:36
Bertrand Charpentier, cofounder and chief scientist at Pruna AI, argues that state-of-the-art is not a single model but multiple on a Pareto front that balances quality and efficiency. He highlights that public leaderboards disagree—Hunyuan ranks 10th on Artificial Analysis but 5th on Arena—and most models lose 40% of head-to-head battles, meaning the top-ranked model is wrong for nearly half of use cases. Evaluating a large model like ChatGPT image on Design Arena (26k battles, 62 seconds each) costs $5,000 and 20 days of compute, consuming energy equivalent to 400 marathons, while a fast compressed model completes the same evaluation in 7 hours for $265. Charpentier advocates plotting quality against latency or cost to find the frontier, which often surfaces small specialized models instead of large foundation models, with 20x efficiency differences at similar quality scores.

Most Enterprise Agentic Projects Are Doomed, Here's Why — Jess Grogan-Avignon & Jack Wang, Accenture
May 28, 2026 · 20:35
Jess Grogan-Avignon and Jack Wang, Accenture, argue that most enterprise agentic projects fail not because of bad code but because organizational scaffolding designed for human pace chokes machine-speed AI delivery. They built an agentic application in two weeks, then spent twelve months aligning infrastructure, security, AI gateway, data governance, and application teams before it shipped—a pattern they say will worsen as GitHub's 275 million weekly commits flood approval infrastructure built for humans. They name five predictive tensions: human approval chains must become executable code, not longer meetings; finance should back a portfolio of AI bets like a VC, not demand committed returns from each project; delivery should run on hypothesis-driven loops building statistical confidence, not fixed milestones; trust is built through progressive autonomy—shadow, advisory, controlled autonomy—gated by outcome evidence, not plan completion; and the real moat is living memory from real customer signals, not static CRM or ERP data. Their prescription: bet like a VC, upgrade governance for machine speed, and engineer for trust with feedback loops from day one.

Does GenAI "belong" to data scientists? — Phil Hetzel, Braintrust
May 25, 2026 · 18:54
Phil Hetzel of Braintrust argues that generative AI development should not be isolated to data scientists or ML engineers, but instead requires a diverse team including product engineers, systems engineers, and non-technical domain experts. Because models are already built by OpenAI and Anthropic, the remaining work is prompt and context engineering, distributed systems, human annotation, and functional evaluation—not traditional training pipelines. Data scientists add value through rigorous testing and LLM-as-judge evaluation, but must move beyond precision/recall toward broader functional metrics. Traditional enterprises often mistakenly hand GenAI to ML platform teams, while AI natives use small cross-functional groups with closer problem proximity. The ideal mix combines technical roles for implementation and system design with non-technical experts for prompt engineering and human annotation, keeping the agent relevant through continuous feedback.

Bounded Autonomy: Between Free Will and Determinism — Angus J. McLean, Oliver
May 25, 2026 · 16:52
Angus J. McLean, AI Director at Oliver, argues that developers should intentionally constrain large language models rather than maximizing their capabilities, because models naturally tend toward complexity and verbosity. He shares how Oliver generates 4,000 creative assets daily for 200+ brands using agents, but warns that context windows will never be sufficient as world knowledge doubles every 12 hours. McLean advises replacing internet access with curated documentation, asking how little context can be used to complete a task, and never automating a job you cannot do yourself. He illustrates this with his own experience building a complex agent application for his CV that was outperformed 100x by a simple HTML page. The talk frames AI as fundamentally a translation process between representations—text to image, image to audio, etc.—and recommends using multiple representation structures like Markdown, graphs, and folders. McLean's core message is that abundance stops scrappiness, so self-imposed constraints drive creativity and better results.

Rewiring the State — Eoin Mulgrew, No. 10 (Downing Street)
May 18, 2026 · 28:18
Eoin Mulgrew, from the Number 10 Data Science team, details how a small insurgent unit at the center of UK government bypasses bureaucracy to rapidly deploy AI. The team recruits exclusively outsiders (0.7% acceptance rate), pays market rates, and ships tools in weeks. Examples include an engineer who saved £1.5M by building a statute-book analysis tool in two weeks, a policy simulation platform for Universal Credit, a delivery red-teaming PMO, and a public service that went from idea to live in two months. The team also placed fellows in the AI Safety Institute, the Incubator for AI (co-creating Xtract with DeepMind to digitize planning applications), and Justice AI, which embeds engineers in prisons. Mulgrew closes with Will, a Y Combinator founder and Harvard dropout, standing outside HMP Wormwood Scrubs with the keys two weeks into the job, urging talented technologists to 'join us, and we'll give you the keys to the state.'

One Registry to Rule them All - Sonny Merla, Mauro Luchetti, & Mattia Redaelli, Quantyca
Apr 10, 2026 · 22:47
Amplifon's AI transformation led to the Amplify program, for which Quantyca built an enterprise-grade registry system for MCP servers and A2A agents. Sonny Merla, Mauro Luchetti, and Mattia Redaelli explain how three registries—MCP, A2A, and use case—are linked via a catalog to provide full lineage, including ownership, environment, authentication, cost attribution, and use case linkage. The solution uses an AI gateway for unified LLM access with Entra ID authentication and budgeting, and provides template repositories on GitHub with CI/CD pipelines that automatically publish agent cards and server metadata to the registries. This enables discovery, governance, and impact analysis across 26 countries and multiple teams, letting developers focus on business logic while avoiding reinventing security and deployment infrastructure.

Small Bets, Big Impact Building GenBI at a Fortune 100 – Asaf Bord, Northwestern Mutual
Dec 23, 2025 · 22:50
Asaf Bord, AI Product Lead at Northwestern Mutual, shares how his team built GenBI, an LLM-powered analytics copilot, by flipping the logic from a single big bet to an incremental roadmap of small, fundable projects. Using real, messy data from the 160-year-old company, they deployed a modular architecture with metadata, RAG, SQL, and BI agents, each productizable independently. The RAG agent alone automated 80% of the 20% of BI team capacity spent on finding and sharing reports, saving roughly two full-time employees. Bord explains how a crawl-walk-run release strategy built trust with both users and leadership, starting with BI experts before expanding to business managers, and how each six-week sprint delivered tangible business value—like proving the ROI of enriched metadata through A/B tests against a semantic layer initiative. He also explores the future of SaaS pricing in the GenAI era, questioning whether per-seat models still make sense when individuals become 10x more effective.

Leadership in AI Assisted Engineering – Justin Reock, DX (acq. Atlassian)
Dec 19, 2025 · 18:11
Justin Reock, Deputy CTO at DX (acquired by Atlassian), argues that AI's impact on engineering productivity varies wildly and that leaders must move beyond top-down mandates to focus on psychological safety, measurement of actual outcomes, and targeted integration across the SDLC. He presents data showing a 2.6% average increase in change confidence but extreme variability across companies, with some seeing 20% drops. Emphasizes that writing code is rarely the bottleneck; instead, leaders should identify and fix bottlenecks like context switching, citing Morgan Stanley's DevGenAI saving 300,000 hours annually by converting legacy code specs and Zapier reducing engineer onboarding to two weeks via AI agents. Introduces DX's AI Measurement Framework covering utilization, impact, and cost, and stresses trust-building through system prompt feedback loops and temperature settings. The episode delivers actionable guidance on measuring AI's true impact and enabling engineers through education, time to learn, and creative unblocking of usage.

AI Consulting in Practice – NLW, Superintelligent, @AIDailyBrief
Dec 18, 2025 · 18:18
NLW, host of the AI Daily Brief and CEO of Superintelligent, presents findings from an ROI survey of enterprise AI adoption, revealing that 44.3% of organizations report modest ROI and 37.6% report high ROI, with 67% expecting high growth next year. Agent adoption jumped from 11% to 42% in 2024, yet only 7% of organizations are fully at scale, and most remain in pilot phases. Time savings accounts for 35% of use cases, but automation and agentic use cases significantly outperform others in self-reported ROI. Risk reduction use cases, though only 3.4% of submissions, are most likely to yield transformational impact at 25%. Larger organizations and those using AI across multiple functions see greater benefits, while coding and software-related use cases show above-average ROI.

Moving away from Agile: What's Next – Martin Harrysson & Natasha Maniar, McKinsey & Company
Dec 12, 2025 · 21:55
McKinsey's Martin Harrysson and Natasha Maniar argue that enterprises must overhaul their people and operating models—moving beyond Agile to AI-native workflows—to capture the full potential of AI in software development. They identify bottlenecks like work allocation, manual code review, and tech debt that limit gains to 5-15% despite promising individual productivity stories. To scale, they advocate AI-native workflows (spec-driven development, continuous planning) and AI-native roles (smaller pods of 3-5 with consolidated roles). In a client study with a bank, interventions led to 60x increase in agent consumption, 51% increase in code mergers, and faster delivery tied to business priorities. They emphasize change management: 70% of companies haven't changed roles, and top performers invest in hands-on upskilling, measurement systems that track outcomes beyond adoption, and incentives to drive adoption.

Can you prove AI ROI in Software Eng? (Stanford 120k Devs Study) – Yegor Denisov-Blanch, Stanford
Dec 11, 2025 · 16:40
Stanford researcher Yegor Denisov-Blanch presents research based on 120,000+ developers across 600+ companies, arguing that AI ROI in software engineering is often negative despite apparent productivity gains. The study finds a median 10% productivity lift from AI, but a widening gap between top and bottom performers. Key drivers include codebase hygiene (environment cleanliness index with R² 0.40) and AI usage quality over volume. A company case study shows pull requests increased 14% but code quality dropped 9% and rework rose 2.5x, yielding no net effective output gain. Denisov-Blanch proposes a measurement framework using a primary metric (engineering output via ML model) paired with guardrail metrics, and emphasizes that companies can retroactively measure impact via git history without waiting for experiments.

State of Startups and AI 2025 - Sarah Guo, Conviction
Aug 2, 2025 · 23:52
Sarah Guo of Conviction argues that AI's value creation is massive and early, with companies like Cursor reaching $100M ARR in 12 months and Harvey exceeding $70M ARR. She predicts that by end of 2026, AI agents will ship code directly to production, voice AI will replace text for most business communication, and inference costs will drop below a cent per million tokens. Reasoning is a new scaling vector unlocking higher-stakes use cases, and agent startups have increased 50% in the last year. Multimodal models from HeyGen and Eleven are already rocketing past $50M ARR. The model market is more competitive than ever, with GPT-4 costs falling from $30 to $2 per million tokens in 18 months and open-source like DeepSeek competing. Guo advises builders to focus on thick wrappers around LLMs, leveraging domain and workflow knowledge, and warns against generic text boxes: 'The prompt is a bug, not a feature.' Execution, not first-mover advantage, is the moat.

Everything is ugly, so go build something that isn't — Raiza Martin, Huxe (ex NotebookLM)
Jul 28, 2025 · 25:15
Raiza Martin, former lead of Google's NotebookLM and founder of Huxe, argues that the current chaotic phase of AI product design is a once-in-a-career opportunity to rebuild from first principles, calling everything we use 'the ugliest that it will ever be.' Drawing from her experience forcing NotebookLM into existence against skepticism, she emphasizes that personal clarity of vision fuels product building, and purpose must be relentlessly focused on a single outcome—for NotebookLM, enabling users to upload 50 files and interact with them. She stresses earning trust by nailing deterministic behavior first, noting that 90% of first queries were summarization and failures drove users away forever, then layering on delightful probabilistic features like podcast generation. Finally, she warns against the 'kitchen sink' approach, citing her own Huxe app that did everything but users only used one feature, advocating restraint as an innovation multiplier and focus on one excellent outcome to avoid building ugly products.

Rise of the AI Architect — Clay Bavor, Cofounder, Sierra w/ Alessio Fanelli
Jul 24, 2025 · 18:55
Clay Bavor, cofounder of Sierra, and Alessio Fanelli discuss the rise of the AI Architect—a new role combining technology, brand, and business outcomes to build customer-facing AI agents. Sierra serves hundreds of millions of consumers this year. Bavor defines the AI Architect as wearing three hats: understanding AI capabilities, defining the agent's voice (e.g., Chubbies' irreverent Duncan Smothers), and driving business outcomes. Successful AI Architects embrace risk, start with narrow problems like processing a single return, and re-architect teams to coach the AI. On build vs. buy, Bavor warns of the "agent iceberg"—hundreds of hidden complexities like regression testing and model migration. He advises tracking model improvement in a Google Doc and anticipating future capabilities, predicting glasses as the ultimate interface for trusted personal AI.

Structuring a modern AI team — Denys Linkov, Wisedocs
Jul 24, 2025 · 17:40
Denys Linkov, who leads ML at Wisedocs, argues that building a modern AI team hinges on identifying your company's bottleneck—shipping features, acquiring users, or scalability—rather than reflexively hiring AI researchers. He introduces Ampere's Wager: trading your entire domain-savvy team for five top-lab researchers is usually a losing bet. For early-stage AI strategy, generalists who blend model training, serving, and business acumen outperform specialists; Linkov lived this in 2021 building a custom MLOps platform for a conversational AI startup and again in 2024 using advanced open-source tools for medical record processing. He stresses reskilling existing teams through weekly learning cadences and moving domain experts from giving feedback to writing evaluations. Hiring should hold context and act on it, verifying trends like 'don't hire juniors' against YC's AI school drawing 2,000 young people.

The Rise of Open Models in the Enterprise — Amir Haghighat, Baseten
Jul 24, 2025 · 16:50
Amir Haghighat, CTO of Baseten, argues that enterprises are increasingly moving from closed frontier models like OpenAI and Anthropic toward open source models, driven by four specific cracks in the assumption that closed models will work indefinitely: quality for specialized tasks (e.g., medical document extraction), latency requirements (especially for voice), unit economics ballooning from agentic use cases where a single user action triggers 50 inference calls, and the desire for competitive differentiation. Drawing on conversations with over 100 enterprises, he explains that while most started with dedicated deployments on Azure/AWS for toying around in 2023, by 2024 about 40-50 had production use cases, and in 2025 the shift accelerated. However, adopting open models forces enterprises to build inference infrastructure, facing challenges like speculative decoding, prefix caching, guaranteeing four-nines reliability with hardware failures and VLM crashes, and scaling replicas—with one Fortune 50 soft drink company reporting an eight-minute spin-up time. Haghighat concludes by contrasting the simple API-call world with the complexities of mission-critical inference, where…

From Hype to Habit: How We’re Building an AI-First SaaS Company—While Still Shipping the Roadmap
Jul 23, 2025 · 19:29
Rossella Blatt Vital and Deepsha, AI leaders at Sprout Social, share a candid, real-time look at transforming a SaaS company into an AI-first organization, arguing that the journey is messy, non-linear, and requires evolving across strategy, ways of working, and people. They explain that being AI-first means shifting from adding AI to existing features to reimagining entirely new experiences, while managing the innovator's dilemma of balancing current customer needs with future investments. Key shifts include adopting ritualized discovery—making experimentation a repeatable process—and embracing 'smart velocity' (speed with purpose) over chaotic fast shipping. On the people side, they advocate for T-shaped talent (deep expertise plus broad vision) and building org-wide AI fluency so every team feels empowered to use AI. Despite the disruption, they emphasize that fundamentals like solving real customer problems, user experience, and trust remain non-negotiable. The talk provides a practical framework but no simple playbook, encouraging leaders to start with the right questions and the conviction to evolve.

Build Dynamic Products, and Stop the AI Sideshow — Eliza Cabrera (Workday) + Jeremy Silva (Freeplay)
Jul 23, 2025 · 18:10
Eliza Cabrera (Principal AI Product Manager at Workday) and Jeremy Silva (Product Lead at Freeplay) argue that companies must stop treating AI as a separate 'sideshow' and instead deeply integrate it into product strategy using a crawl-walk-run approach to build dynamic, differentiated experiences. They explain that bolt-on AI products result from centralized AI strategies quarantined from core product, leading to features like chatbots that demonstrate capability but don't solve real customer problems. Workday's example shows starting with Gen AI content generation and translations for knowledge bases (crawl), then a contextually aware assistant that processes PII (walk), and finally agentic capabilities that autonomously act on policy changes (run). The north star is AI products that feel like natural, cohesive parts of the experience, not bolt-on enhancements.

The Billable Hour is Dead; Long Live the Billable Hour — Kevin Madura + Mo Bhasin, Alix Partners
Jul 23, 2025 · 17:04
Kevin Madura and Mo Bhasin from AlixPartners argue that AI is reshaping knowledge work by compressing upfront data ingestion from 50% to 10-20% human effort, enabling analysis of 100% of data rather than the top 20%. They detail three use cases: categorization via structured outputs achieving 95% accuracy on 10,000 vendors in minutes; enterprise-scale RAG that democratizes access to siloed data by teaching LLMs to call APIs; and structured extraction from documents using schemas and log-prob-based confidence scoring. They caution that AI investments face a paradox—89% of CEOs plan agentic AI but NBER finds no earnings impact—and stress that success requires people skills, demos, and a focus on NPS and ROI over chasing shiny new tools. The episode concludes that once Excel-powered LLMs work reliably, AGI will be here.

Critical AI Inference your CIO can Trust — Sahil Yadav, Hariharan Ganesan, Telemetrak
Jul 22, 2025 · 19:04
Sahil Yadav and Hariharan Ganesan of Telemetrak present a three-pillar framework—explainability, adaptive guardrails, and human-in-the-loop on a foundation of traceability—to operationalize trust in enterprise AI inferences, arguing that without these, silent failures can cost millions. They introduce XTOPS as an integrated upgrade to MLOps, with trust-specific dashboards and dynamic guardrails. A case study of GuardHat, a worker safety platform, shows how GPS drift caused 70% false positives; applying XTOPS reduced resolution time from eight months to seven days. They propose metrics MTTRE (mean time to resolve explainable errors) and trust-adjusted risk in dollars, and demonstrate convincing CIOs by quantifying savings, e.g., $500K per site per year in fines avoided.

The Bitter Layout or: How I Learned to Love the Model Picker — Maximillian Piras, Yutori
Jul 21, 2025 · 14:24
In this talk, Maximillian Piras argues that conversational interfaces remain dominant in AI apps not because they are natural, but because they are 'conformable' — able to absorb the next model's capabilities without redesign. He calls this pattern 'The Bitter Layout': an input field, turn-by-turn flow, and a model picker, which prioritizes flexibility over usability. Applying Clayton Christensen's theory of commoditization, Piras claims that as long as scaling laws keep models from commoditizing, the interface itself is the commodity, and designers must conform to the next model. He traces the debate over chat UX from Linus Lee (2022) to Julian Lear (2024), and criticizes the model picker as a 'mode selector' that creates usability issues. Looking ahead, Piras suggests designers shift from procedural thinking to setting goals and constraints, and speculates that future AI UX will be 'grown' like a garden rather than built.

Building a 10 person unicorn - Max Brodeur-Urbas, Gumloop
Jul 15, 2025 · 12:03
Max Brodeur-Urbas, founder of Gumloop, explains how his company scaled to millions in ARR as a team of two that raised a Series A and grew to only nine people, by being super picky in hiring, using product-led hiring where customers like those from Instacart and Webflow join the team, and requiring work trials such as hacking together in Airbnbs. They eliminate almost all meetings to give engineers deep focus time and automate every internal process with Gumloop itself, from customer research reports to chatbot monitoring. Culture-wise, they balance intense 45-minute shipping challenges with fun retreats and a public company handbook. Brodeur-Urbas argues that every hire must be a no-brainer, and small teams can outpace larger ones by avoiding meetings and leveraging AI tools.

Bolt.new: How we scaled $0-20m ARR in 60 days, with 15 people — Eric Simons, Bolt
Jul 15, 2025 · 17:33
Eric Simons, CEO of Bolt.new, explains how his team of less than 20 people scaled the company from $0 to $20M ARR in just 60 days, achieving the second fastest product growth in history by staying lean and focusing on high-impact decisions. He describes building a remote team with high context and agency, resisting pressure to hire more during the 2021 boom. They relied on 'things that don't scale' like weekly office hours to build user love, and leveraged AI support tools like PeraHelp to handle 90% of tickets. Community initiatives such as a Guinness World Record hackathon (over 80,000 participants) further amplified growth without adding headcount. Simons emphasizes taking consistent shots on goal and making independent bets rather than following VC trends.

Survive the AI Knife Fight: Building Products That Win — Brian Balfour, Reforge
Jul 14, 2025 · 14:10
Brian Balfour, CEO of Reforge, argues that winning in today's AI knife-fight requires answering 'What do I build and why will it win?' by focusing on proprietary data, unique functionality, and unmet customer needs rather than building custom AI. He illustrates with Granola, which entered a crowded AI note-taker market by understanding users wanted help taking better notes, not full automation, and assembled off-the-shelf AI (DeepGram, Anthropic, OpenAI) with unique data (user notes plus transcription) and functionality (Mac app, calendar integration) to create a competitive edge. Balfour warns competitive advantages now last only 2-3 weeks, so teams must sequence smaller moats continuously, each buying time to execute faster. The talk emphasizes treating AI as Lego blocks—assembling pre-trained models, data, and product superpowers into a system that spins a data flywheel.

The Agent Native Company — Rick Blalock, Agentuity
Jun 3, 2025 · 20:58
Rick Blalock of Agentuity argues that an agent-native company, built from the ground up with AI agents at the core of product, operations, and culture, is fundamentally different from an AI-enhanced one that merely uses AI as a tool. He contrasts the two: removing agents from an agent-native company would halt productivity, while an AI-enhanced business would just become less efficient. Blalock describes the agent-native workday, where humans oversee agents that handle routine tasks, and notes the rise of roles like 'Agent Manager' and the importance of AI fluency in hiring. He shares how his 7-person team built an entire agentic cloud infrastructure in weeks using agents like Devin, arguing that this paradigm shift requires founders to rethink org charts, roles, and skills. The episode concludes that businesses must decide whether they are just using AI or ready to be built around it.

Agentic Enterprise - What your CEO must know about AI - Hubert Misztela
Jun 3, 2025 · 28:04
Hubert Misztela, an AI research lead at Novartis, argues that organizations may be run by AI agents within three years and must pivot from traditional roles to persona-based workflows to harness agentic automation. He explains that AI agents glue multiple cognitive steps, automating workflows where only humans could operate. Understanding deep context around each workflow is crucial, as this knowledge is often undocumented. Misztela introduces five employee personas (silent achiever, individual contributor, connector, multiplier, knowledge hub) to project how agents merge tasks, magnify impact, and replace roles. He warns that intelligence and domain knowledge become cheap commodities, so companies need multidisciplinary or deeper specialization, and new roles like workflow miner will emerge. Employees must build their own agents using no-code tools, and ethical questions around value alignment remain.

From PM at Stripe to Building an AI startup, a recent founder's journey - Mounir Mouawad
Jun 3, 2025 · 11:59
Mounir Mouawad, CEO and co-founder of Porsche AI, explains how building an AI startup differs from product roles at Stripe, Google, and Amazon, using video game analogies. He argues user problems are an 'emergent property' requiring hypothesis-driven iteration rather than conventional roadmaps. Product development is gratifying with releases in hours or days, but velocity is a 'stable stick' as opportunities like MCP come and go quickly. The hardest part is outreach without big brand support—like playing Crash Bandicoot without boosters—so he finds people followers, advocates, and partnerships (e.g., with Browserbase) essential. He asks listeners to star Porsche AI's GitHub repo.

Stop Ordering AI Takeout A Cookbook for Winning When You Build In House - Jan Siml
Jun 3, 2025 · 10:45
Jan Siml argues that small in-house teams can generate millions in revenue by focusing on one job-to-be-done, tracking dollar outcomes, and pushing proactive insights instead of chasing multi-agent systems and expensive models. Over 10 sprint weeks with two developers, his team built a sales alert system driving several million dollars ARR. He shares five lessons: go deep on one value event, trace everything to revenue (offline evals never sign contracts), push insights proactively (daily digests had 20-point higher NPS than chat UI), convert time saved into guided action, and invest in data and UX over bigger models (changing models only affected costs and evals, not user outcomes). Owning data and tight feedback loops create a revenue flywheel.

Unlocking Africa's Potential with AI — Thabang Ledwaba
Jun 3, 2025 · 26:17
Thabang Ledwaba argues that Africa, often seen as a latecomer to AI, actually has immense potential to lead in AI-driven innovation by leveraging its unique challenges and creativity. He points to Kenya's third-highest daily ChatGPT usage and fintech successes like M-Pesa as evidence of immersion in technology. Criticizing over-engineered solutions, he contrasts ticket systems for home office queues with his idea to auto-initiate ID applications at age of eligibility. He highlights Africa's 30% of earth minerals, noting the irony of exporting raw materials only to import finished goods, and calls for a mindset shift akin to China's 'serve yourself first' strategy. Ledwaba showcases African innovations like Nigerian pharma wings and Moroccan Project Cumulus, urging Africans to see themselves as producers, not just consumers, and to harness AI for sustainable development.

The missing pieces of workflow automation — Shirsha Chaudhuri, Thomson Reuters Labs
Apr 23, 2025 · 14:37
Shirsha Chaudhuri, head of co-innovation at Thomson Reuters Labs, identifies eight missing pieces preventing true AI workflow automation in the enterprise. She argues that while 71% of Fortune 500 companies still run mainframes and 68% of IT production workloads remain on mainframe, current agentic efforts lack connectors to bridge legacy systems, standardized agent architectures, and reliable ROI metrics. She highlights the need for domain experts to reimagine processes alongside AI practitioners, collaborative UX design, and AI governance that translates ethics into agent architecture. Control balance between deterministic and agent-driven steps remains unresolved, and the fast-evolving agent lifecycle lacks a clear update strategy. Drawing from Thomson Reuters' own journey—from a 2023 open AI arena through RAG and prompt engineering to 2024's agentic experiments—she calls for a reimagined workflow design rather than just task-level automation.

Insights on Building AI Teams — Heath Black, SignalFire
Apr 15, 2025 · 20:30
Heath Black, Managing Director of Product at SignalFire, uses Beacon platform data to guide AI team building, arguing that credentialism is declining—only 7% of AI engineers had PhDs in 2023 versus 16% in 2015—and that work experience now outweighs education. He shows AI talent concentrates in San Francisco (35% of AI engineers), Seattle (22%), and New York (10%), and that tracking retention rates (e.g., Anthropic at 66% four-year retention vs. Perplexity at 43%) helps time outreach. Black advises hiring based on a candidate's body of work, removing academic requirements from postings, and understanding generational job-hopping (27% of Gen Z left jobs in 2023). He warns against relying solely on salary and equity, as AI engineers command 5% salary and 10–20% equity premiums, and recommends narratives centered on mission, speed, and collaborative teams. The talk emphasizes using data to filter, locate, time, and close hires effectively.

How to Fail at AI Strategy: Hamel Husain & Greg Ceccarelli
Apr 13, 2025 · 17:03
Greg Ceccarelli and Hamel Husain argue that the most reliable path to AI failure is to follow a set of inverted worst practices, including cultivating disconnect between executives and builders, promising unrealistic AI capabilities, drowning communication in jargon, and avoiding data analysis. They describe how to divide your company by incentivizing secrecy and using jargon like 'agents' to exclude domain experts, ensuring that AI projects are disconnected from real needs. The speakers advocate faking strategy by highlighting random paragraphs from last year's report, announcing vague goals like 'become the global AI leader in everything,' and creating a massive backlog with no timeline. They recommend throwing tools at problems—buying expensive vector databases or switching frameworks—without understanding root causes, and blindly trusting off-the-shelf evaluation metrics like BLEU and ROUGE. Crucially, they insist on never looking at data, using complex systems inaccessible to domain experts, and trusting gut feelings over evidence. This inverted guide guarantees wasted resources, alienated teams, and spectacular failure.

OpenAI for VP's of AI + Advice for Building Agents
Mar 5, 2025 · 16:52
OpenAI's Toki Sherbakov and Prashant Mital explain how enterprises adopt AI through a three-phase journey: building an AI-enabled workforce with ChatGPT, automating operations with APIs, and infusing AI into end products. They detail a Morgan Stanley case study where retrieval methods improved an internal knowledge assistant's accuracy from 45% to 98%. The pair define agents as models with instructions, tools, and self-terminating execution loops, then share four field lessons: build with primitives before frameworks, start with a single purpose-built agent, graduate to a network of specialized agents with handoffs for complex tasks, and keep prompt instructions simple while running guardrails in parallel using fast models like GPT-4o mini for safety and reliability.

WTF do people use Open Models for??
Feb 22, 2025 · 28:01
Eugene Cheah of Featherless.ai breaks down how individuals and enterprises actually use open-source AI models, based on platform data. DeepSeek R1 dominates individual usage, but Mistral Nemo 8B remains the top enterprise model due to production stickiness and Apache 2.0 licensing. Creative writing and roleplay account for 30–40% of all traffic, with over 60% of users in that segment being women; coding copilots and agents make up 20–30%, driven by 'vibe coding' and token-hungry workflows like Kline. RAG and ChatGPT clones represent 20%, while agentic workflows (10–20%) succeed with human-in-the-loop designs. Cheah advises enterprises to aim for 80% automation with escape hatches, and warns against chasing 100% reliability. He concludes by introducing Quirky, a post-transformer hybrid built for $100k.

Privacy First Enterprise AI: Building AI Agents that Never Leave Your Security Boundary
Feb 22, 2025 · 7:10
Steven Moon, founder of Aech AI, argues that enterprise AI agents should be deployed within existing security boundaries by treating them like human employees—using existing identity management, compliance frameworks, and audit tools rather than building parallel systems. He explains that IT departments will evolve into HR departments for AI agents, provisioning them through active directory and applying standard security policies. Moon highlights email as a powerful medium for agent-to-agent communication, where every interaction is logged and auditable through existing systems. He advocates for enhancing current enterprise platforms like Microsoft 365 and Azure with AI agents instead of creating new interfaces, noting that the era of mandatory translation layers between humans and machines is ending.

Reverse Conway's law and GenAI: How agents will take over the organisation - Patrick Debois
Feb 22, 2025 · 28:12
Patrick Debois argues that generative AI will reverse Conway's law, reshaping organizations around agents instead of human teams. He traces a progression from AI as a copilot to a team member, then a peer, and eventually a manager, with each stage unbundling human tasks and shrinking team sizes. Debois cites Amazon's AI pricing glitch as a reminder that humans remain needed for failure cleanup, and notes that LLMs mimic human collaboration behaviors, as shown in a multi-agent simulation. He warns of agent toxic behavior and the need for guardrails akin to human codes of conduct, while speculating that companies may replace SaaS with internally built AI services, making performance reviews and ROI calculations for agents inevitable. Debois advises engineers to focus on building the AI that builds their current work, not the work itself.

Where AI is superhuman: The right jobs to automate with LLMs
Feb 22, 2025 · 11:53
Andy Treadman, partner at Theory Ventures, argues that LLMs are already superhuman in high-volume, low-complexity jobs, making them the prime targets for full automation. He breaks down LLM capabilities into transformation, synthesis, and reasoning, and maps jobs on a spectrum of volume and complexity. At the high-volume end, he says LLMs can be 10x better than humans because scale is the challenge, not reliability. He highlights two portfolio companies: Dropzone AI automates security alert investigations, performing better than rules-based systems and covering 24/7, while Amp personalizes customer engagement at scale, discovering new cohorts like late-night snackers. Treadman predicts that organizations will shrink and invert from pyramids to diamonds, with fewer entry-level roles. He advises founders to consider technology-problem fit and other factors when building AI workflow automation.

Cohere for VPs of AI: Vivek Muppalla
Feb 5, 2025 · 16:11
Vivek Muppalla, Director of Engineering at Cohere, details the company's enterprise AI strategy centered on security, customization, and deployment flexibility. He presents Cohere's product line: Command R and R+ for generation, plus advanced retrieval models like embeddings and the ReRanker, which reduces RAG costs by narrowing context. Key claims include a focus on enterprise-specific eval suites (health, HR, finance), out-of-the-box citations, and multilingual performance. Partnerships with Accenture and McKinsey bridge the last-mile gap, while wins often stem from private cloud deployment and data control. In Q&A, he recommends purpose-built classifiers for high-throughput production and notes current 128k context windows.

Cooking with fire without burning down the kitchen: Dominik Kundel
Dec 31, 2024 · 20:35
Dominik Kundel, who leads product and design for Twilio's Emerging Tech & Innovation team, explains how the company balances disruptive AI innovation with its existing communications and customer data platform businesses. He distinguishes sustaining innovation (e.g., Apple's AI features) from disruptive innovation (e.g., agents not yet enterprise-ready due to quality and cost), arguing that ignoring disruptive AI is increasingly dangerous because quality improves daily. Kundel shares three key lessons from Twilio's AI journey: first, ship early and often—even rough prototypes—to gather real feedback, which led to the Twilio Alpha sub-brand for setting expectations; second, build a curious, problem-owning team rather than requiring existing AI expertise; third, share learnings internally and externally to avoid operating in a silo and to help customers be thought leaders. He recounts how an initial AI personalization engine was too disruptive for mainstream R&D, so the team iterated on an AI Assistance agent builder, using internal hackathons and dogfooding with low-risk use cases like IT helpdesk to improve quality before broader release.

Hiring & Building an AI Engineering Team: Dr. Bryan Bischof
Dec 31, 2024 · 29:07
Dr. Bryan Bischof, Head of AI at Hex, argues that building an AI engineering team requires hiring based on product stage—starting with full-stack engineers and data profiles, then adding designers and later MLEs—while avoiding the 'mythical man month' trap in early AI products. He advocates for data intuition over LeetCode in interviews, using a take-home data exercise to assess candidates' ability to extract meaning from data and give feedback. Key attributes he looks for are curiosity, urgency, and product-mindedness, noting that enthusiasm alone is insufficient. Bischof also recommends working directly with domain experts to model AI behavior and suggests a centralized AI platform team to support multiple product teams, rather than having each team build AI infrastructure independently.

AI Platform Engineering: Patrick Debois
Dec 31, 2024 · 28:18
Patrick Debois, who coined DevOps in 2009, argues scaling GenAI requires an AI platform team, mirroring the cloud and DevOps pattern. He lists shared infrastructure: model access, vector databases, RAG connectors, version control, proxy, observability, monitoring, caching, feedback services, and stresses enablement through prototyping and frameworks. He warns of pitfalls like chasing GenAI without use case, overfocus on fine-tuning, and cost obsession. Citing the Ironies of GenAI Automation, he notes that as engineers shift from producing to reviewing code—copying co-pilot suggestions more on weekends—they risk losing situational awareness. For governance, he advises awareness programs, opt-out training, license checks, and layered guardrails: central rules plus team-specific overlays. He recommends combining cloudops, secops, devx, data platform, and AI platform teams for collaboration.

Decoding Mistral AI's Large Language Models: Devendra Chaplot
Nov 21, 2024 · 18:16
Devendra Singh Chaplot of Mistral AI details the company's open-source large language models, including Mistral 7B, Mixtral 8x7B, Mixtral 8x22B, and CodeStral 22B, arguing that open models complement rather than compete with profit by serving as branding tools and driving customer acquisition for proprietary upgrades. He explains the three-stage LLM training process—pre-training on trillions of tokens, instruction tuning with prompt-response pairs, and learning from human feedback via preference optimization—emphasizing that more data does not guarantee better performance due to noise. The episode highlights Mistral's focus on optimizing the performance-to-cost ratio, with CodeStral 22B outperforming larger models like Code LLaMA 70B while being smaller and multilingual across 80+ programming languages. Practical guidance is given: prototype with high-end commercial models, then fine-tune open models for specific tasks to balance performance and cost.

Lessons From A Year Building With LLMs
Jul 19, 2024 · 35:21
The six authors of the O'Reilly article "Lessons From A Year Building With LLMs" — Bryan Bischof, Jason Liu, Hamel Husain, Eugene Yan, Shreya Shankar, and Swix — argue that the model itself is not a moat and that success comes from continuous improvement centered on evals and data. They stress that AI engineers should treat models as SaaS, quickly swapping for better ones, and focus on product and user interactions. The talk warns against toxic practices like prematurely hiring ML engineers without data or blindly adopting tools, advocating instead for deliberate eval practice and data literacy. On tactics, they compare LLM-as-judge (quick to prototype) vs fine-tuned evaluators (more precise and faster), and emphasize looking at real user data regularly with automated guardrails. Ultimately, they conclude that going from demo to production requires sustained investment in infrastructure and evaluation, echoing MLOps lessons from a decade ago.

The AI Evolution: Mario Rodriguez, GitHub
Nov 7, 2023 · 19:32
Mario Rodriguez, VP of Product at GitHub, recounts the history and future of GitHub Copilot, the first at-scale AI programmer. Copilot now serves over 20,000 organizations and 1 million-plus developers, generating over $100M in ARR, with 46% of code written via completions. He shares insider lessons: ghost text, low latency under 100 milliseconds, and prompt engineering were key to success. He warns that 'syntax is not software' and that global presence and offline/online scorecards are essential at scale. Looking ahead, Rodriguez envisions moving from procedures to goals and constraints, enabling AI to reason on code, and designing immersive UIs for human-AI collaboration. He concludes that GitHub has evolved from a version control system into an end-to-end platform infused with AI.

The 1,000x AI Engineer: Swyx
Oct 23, 2023 · 9:27
Swyx argues that AI engineers are just in time for a 1,000x opportunity, drawing on historical tech cycles and compute scaling laws. He cites Carlota Perez's tech revolution cycles, placing the AI revolution's start at AlexNet in 2012. With compute growing 600x by decade's end, GPT-3 took one person-year of compute, GPT-4 took 100, and GPT-10 will exceed all human compute ever. He defines three AI engineer types: AI Enhanced (Copilot-like), AI Product (Midjourney), and AI Agent (Auto-GPT). To achieve 1,000x, he advises scaling knowledge from O(n) (attending talks) to O(n²) (teaching others) to O(2^n) (building networks).
Powered by PodHood