A product discussed on AI Engineer.

Why Off-the-Shelf AI Doesn't Understand Money — Udi Menkes, Intuit
Jul 29, 2026 · 19:50
Udi Menkes, principal PM at Intuit, argues that off-the-shelf frontier models deliver a 'fluent bluff' when advising on money: advice that sounds right but is dangerous because models have read about money but lack experience. He shows a rental property example where a frontier model told a landlord in negative cash flow to acquire a second property, while a model grounded in real outcomes recommended raising rent 5-10%. Intuit's head-to-head test across 100,000 businesses found frontier models gave advice that would harm businesses 40% of the time, while a mid-sized grounded model outperformed them by training on millions of state-action-outcome records from QuickBooks data. A Princeton study confirmed frontier models given $1M went bankrupt within 500 days, while a simple rule-based system beat them. Menkes says the moat belongs to whoever owns the best system of context, and advises leaders to find verified outcomes in their own data to ground AI.

How Intuit uses LLMs to explain taxes to millions of taxpayers - Jaspreet Singh, Intuit
Jul 23, 2025 · 18:59
Jaspreet Singh, Senior Staff Engineer at Intuit, explains how TurboTax uses Anthropic's Claude and OpenAI models to generate personalized tax explanations for 44 million customers. The system, built on Intuit's GenOS platform, combines prompt-engineered static explanations with dynamic question-answering using RAG and graphRAG for tax-specific queries. Singh details their fine-tuning of Claude Haiku on AWS Bedrock, which improved quality but proved too specialized. A key focus is evaluation: manual reviews by tax analysts followed by automated LLM-as-a-judge scoring for accuracy, relevancy, and coherence, with a golden dataset and safety guardrails to prevent hallucinated numbers. He highlights challenges like latency spikes on tax day (3–10 seconds), vendor lock-in through expensive contracts, and the difficulty of upgrading models even within the same vendor. Singh emphasizes that evaluations are essential for launching any GenAI feature in a regulated domain like tax.
Powered by PodHood