AIAI EngineerJul 29, 2026· 19:50

Why Off-the-Shelf AI Doesn't Understand Money — Udi Menkes, Intuit

Udi Menkes, principal PM at Intuit, argues that off-the-shelf frontier models deliver a 'fluent bluff' when advising on money: advice that sounds right but is dangerous because models have read about money but lack experience. He shows a rental property example where a frontier model told a landlord in negative cash flow to acquire a second property, while a model grounded in real outcomes recommended raising rent 5-10%. Intuit's head-to-head test across 100,000 businesses found frontier models gave advice that would harm businesses 40% of the time, while a mid-sized grounded model outperformed them by training on millions of state-action-outcome records from QuickBooks data. A Princeton study confirmed frontier models given $1M went bankrupt within 500 days, while a simple rule-based system beat them. Menkes says the moat belongs to whoever owns the best system of context, and advises leaders to find verified outcomes in their own data to ground AI.

  1. 0:00The Problem
  2. 2:32Bad Advice
  3. 5:15Fluent Bluff
  4. 6:36Princeton Study
  5. 9:44Context ≠ Experience
  6. 11:05Causality
  7. 13:42State-Action-Outcome
  8. 15:17Head-to-Head
  9. 16:32Outcome Era
  10. 17:56Takeaways

Powered by PodHood

Transcript

The Problem0:00

Udi Menkes0:13

So, I have a 3-year-old daughter and she's absolutely adorable, and the parents here in the room know how insightful that age can be. And she has a complete theory about money by now. And I'll give an example. So, a couple weeks ago, I was driving the car, I was coming into park, and there was something in my dead end, and I scratched the car.

I went out, I'm like, oh man, I can't believe I scratched the car. And then I hear my daughter from the back, and she was like, "Daddy, what's happened?" And, you know, I'm explaining it to her, and then she says, "What's the problem?

Just buy another one." So, anyway, I want to ask you today, with a raise of hand, who uses LLMs as used LLMs for getting financial advice, a recommendation on something in the financial world? Great. Almost everyone. Wait, keep your hand up if you trusted the answer and you actually followed the advice.

Okay. A lot of hands are going down. And that's the core problem. So, I had the same thing a couple, a couple months ago. I had a big decision I was looking to take. Should I invest in, you know, real estate niche or in the stock market niche on a specific area?

And, you know, I do AI for finance for a living, so I went the full-blown way. Context, brain, the books, the knowledge, all my finances combined, all the latest models, and it gave me a recommendation, "You should do A," with great reasoning.

And then I changed just a little bit one of the assumptions, and it completely flipped. "You should do B. Never do A." And then I tweaked one more small thing, and it went all the way back to A.

And at that moment I understood that the advice sounds good, it sounds sound, but I can't really trust it. And by the end of the talk today, you will understand why off-the-shelf LLMs don't understand money and what you need to do about it.

Bad Advice2:32

Udi Menkes2:32

So, I'm going to show you a couple of real examples from a study we're doing at Intuit on thousands and thousands of businesses, around 100,000 situations and timeframes. This is an example of a small business, a new landlord that is building a rental property business.

His first property, and he's down, he's in a negative cash flow, there's an open loan, and the profit is basically trending into the red. And a question comes up, "How do I improve my profit?" And a frontier model gives the following response: go and acquire a second rental property because that'll bring more income and compensate for the deficit.

And that model had all of the business's data. Now, that's very risky for someone in the negative, in the red, to be doing. On the other hand, a model that is grounded in real outcomes, and what I mean by that is a model that has seen similar situations of such businesses, what they did and what was the outcome, actually recommended to raise prices on the existing tenant by 5 to 10 percent.

And to do it before the renewal. Now, some of you are thinking that it's just a matter of context. Just give it more context. And the thing is that this advice is coming based off on real situations of similar businesses and what actually moved them into profitability in this case.

And this is not a one-off. I can go on and on showing you a lot of examples. This is the second one. This is an egg supplier where one customer is 70 percent of the revenue and one vendor is almost all of its cost.

Same question, "How do I improve my profit?" A frontier model says, "Raise your prices 15 to 20 percent on that customer." Now, you understand that is very risky because you might lose almost all your revenue. On the other hand, the same grounded model in real outcomes went actually to the cost side and recommended to negotiate, negotiate the vendor cost pricing for a 5 to 10 percent reduction.

So, it went to the cost side, it took into account the constraints. So, what we saw are two examples of frontiers. I'm talking about leading models in the world today that had all the business context actually give advice that could be very harmful for the business.

And that's what I call the fluent bluff. The fluent bluff is a generic, fluent, and confident answer that frontier LLMs can give you around money because of what they learn on the internet, blogs, books, advice columns, what people wrote about money, but not based on what actually happened.

Fluent Bluff5:15

Udi Menkes5:43

And I would argue that almost every answer that you see related to money and finances is such. And we're soon going to release the research I talked about, but I'll give you a highlight from there. Across these 100,000 businesses and timeframes, 40 percent of the time, the essence of the advice that the frontier models were giving was, "Acquire a new customer," which is, you know, everyone would wish they could do that.

And 14 additional percent were increase basically the revenue from your product. So, combined, more than half of the essence of the advice that was given by the frontier models was acquire new customers and try to increase revenue from the profit.

Princeton Study6:36

Udi Menkes6:36

And this is not just me saying, this is very interesting research coming that just came out a couple weeks ago from researchers at Princeton. What they tried to do is simulate and answer the question, "Can the leading models drive long-horizon business decisions?"

And what they did is they gave the models a harness with tools and data and everything they would need to take decisions across a simulation of 500 days. Can they turn a profit? Each model got a million dollars to start with, and guess what happened?

Most of the models drove the company bankrupt, and it didn't even take 500 days. And the interesting part is that they also ran a simple rules-based system, and that rules-based system outbeat almost all of the models. Even the very, very few models that were able to generate some profit, it was in a specific instance, and when you rerun that, they also actually drove bankruptcy.

So, how can it be that a simple rule-based system outbeats the frontier models today on real business decisions?

And here's what I want you to think about. A frontier model has read about money, but a grounded model in real outcome has actually watched what happens. And let me be precise with my argue here. Even if you take all of a company's data, and for example, we at Intuit have all the financial data from QuickBooks, for example, the general ledger, the P&L, the cash flows, everything, the invoices of the business, and you give it to a frontier LLM, it's still just one group of data points on a company.

And that's a difference between soundingright and actually beingright. So, I'm Udi Menkes, and I've been in the AI and finance world for the past 15 years. I started in the AI science world, leading AI and data teams, and shifted into product management, becoming in it about four years ago, an AI product manager at Intuit, long before, by the way, it was cool to become an AI PM.

And today I'm a principal product manager at Intuit. I lead financial intelligence and advisory systems that help Intuit's customers to take better decisions and grow their business. And the question that I fixate on, on a daily basis, is not which model is the best now that I can use.

It's what do we fundamentally have that no model access can replicate, and how can I transform that into AI-native, transformative, and delightful experiences for our customers? Now, I want to develop some intuition from three different angles on why these models bluff.

So, the first angle is around context is not experience. And I'll give you another real example. This is an apparel company where 80 percent of the cost is coming from this one vendor. And the textbook answer is, "Go cut your biggest cost,"right?

Context ≠ Experience9:44

Udi Menkes10:03

Textbook, book answer. But the issue here is, you cut this cost, that same vendor was actually enabling the generation of 97 percent of that company's revenue. So, cut that biggest cost and you lose almost all the revenue. And that can be great margins on zero dollars.

And think about it, if I give you an option to work with two different advisors, one advisor is a very experienced one, years of experience working with businesses, guiding them, and another advisor, which is very, very smart. They know all the textbook, they're fresh, they read everything, they know AI in and out, but they don't have experience.

I would bet you would always go with the experienced one. And that's the same thing with the models and what I just showed you. And it turns out experience is very hard to measure. And I'll give you an example.

So, let's look at a restaurant. A restaurant, let's say, raises price. And after six months, becomes much more profitable. Is it because they raise prices, or is it because they're just naturally successful? And the challenge is, obviously, you can't run the business twice,right?

Causality11:05

Udi Menkes11:25

So, what you do is you take two groups of similar companies, similar businesses, that have the same propensity to raise prices, the same likelihood to raise prices. One group raised prices while the other didn't. And then we look after some time at the results.

So, the group that raised prices actually gained 4,200 a day profit. And the group that didn't raise prices actually gained $2,800 a day. So, the question is, what is the impact of raising prices? So, a naive answer would be the difference,right?

1,400 is the impact of raising prices. But actually, you need to account for the fact that the companies that raised prices are actually naturally more successful businesses, which is also why they could raise the prices. So, the real difference is more like $1,150 for this illustration.

And we measure the impact of actions on the outcome through a measure called CATE, Conditional Average Treatment Error, which looks at that connection.

And here's what I want all the AI and finance leaders here in the room to pay attention to. So, find where you can see a lot of different situations across entities that you have in your data, in our case, it's businesses, what they did, and verify the outcomes and how things turned out, if you can see that in the data.

Because that's the one thing that frontier off-the-shelf models do not have. And don't get me wrong, frontier models are amazing, and we actually use them, and I'll show you how we use them. So, we use them to generate hypotheses, candidates for actions we would suggest a business to do.

But then we would use a model that we trained using reinforcement learning in order to figure out which one of those is actually aright move to do versus a mistake that could drive the business down. Now, how we do it, a little bit into our approach is we look at what we call, we actually create from the data, millions of business trajectories.

So, we have data at Intuit across our products, QuickBooks, TurboTax, Credit Karma, MailChimp. So, think about a business in the financial data. There's the general ledger, the P&L, the cash flow statements, the invoices, like I mentioned before. So, we take all of that data and we create what we call business states.

State-Action-Outcome13:42

Udi Menkes14:00

A state of a business is, think about a very detailed summary at a given point of time. And then we derive actions. So, for example, in your general ledger, I can look and see that you have invested in a campaign, in a marketing campaign, or you paid someone.

So, I know you're paying payroll, I know how much

your hiring costs are, and so on. So, we derive all of these actions. And we look at what are the outcomes in different timeframes. And outcomes can be increase in profit, in revenue, in cash flow, in time, combination of those.

So, we create millions of vectors of state, action, and outcome. And then we train in our model to be able to understand, in given situations of similar businesses, which actions lead to the best outcomes. And then the third step is we actually train an LLM to be able to generate that better advice.

That's where opinions are going in and evidence is going out. So, the model you saw in the examples at the beginning were actually a model that we developed with researchers at Intuit across millions of small and medium businesses.

Head-to-Head15:17

Udi Menkes15:17

And we actually tested it head to head with all of the leading models in the world. And we were able, with a mid-sized, cheaper model, to outperform the frontier models because of the grounding that I just showed you.

And the interesting part, as a product person, you would think that it's all about the model size and the bigger and better model, obviously I would have a lot better chance. But it doesn't turn out to be true.

And the moat here is that it's not about the model access, it's about the data itself that you have.

And then we went ahead and built an experience out of it. And this is an AI business advisor that is currently in beta with research, in a research preview with our customers, where we use the LLM that I just described that we created to proactively raise opportunities for businesses at every given point of time.

Here's what you should do, here is why, grounded in who is like you, who's similar to you, what they did, and why we're actually recommending you to do it. And you can drill down into it, understand, and create action plans that will lead to your business actually growing in theright direction.

Now, I want to take a step back and zoom out, because this isn't just about money. We are entering the era of outcome-driven AI. And the question stops being which model is better and starts becoming, how can we steer AI to actually make it achieve the outcomes we want in our domains?

Outcome Era16:32

Udi Menkes16:54

And it doesn't matter if you're building an anti-fraud system or a healthcare system, logistics, developer tools. The winners, in my opinion, are going to be those with the best system of records, creating unique data sets out of them, and then training the models to achieve the outcomes.

And as the product person here, it's not just about the science. The science is very important. But a great advisor, think about the great advisors and mentors that you had in your life. They understand you,right? They understand your preferences, what you like, what you don't like.

So, a great advisory experience needs to have two things. It needs to have great grounded science, the best science, but also it needs to understand you, what you prefer. And it needs to even make you feel as if you were part of the decision to create a trusted experience.

So, three things I want you to remember today. Every model has read about your domain, but none has actually watched the plays and the moves and their outcomes. And that gap is the whole game. Now, in coding agents and coding models, we're seeing it very advanced, a lot of verified outcomes, and creating models that actually lead to better outcomes in coding.

Takeaways17:56

Udi Menkes18:23

But it's very much unexplored in the financial domain and in other domains as well. And you don't close the gap with bigger models. You close the gap with experience, embedding experience into the model by looking at verified outcomes in your data.

What actually worked at scale? So, off-the-shelf models don't understand money, but grounded in real outcomes, it does. And that's what we were able to figure out. So, here's the one thing I want you to do tomorrow. Well, actually, you know what, go ahead and enjoy the 4th of July weekend.

Butright after that, look in your data where you can see situations across entities and outcomes you can verify.

Get deep into that data and start grounding your AI in that. Think about those angles. And that's for you to build. Thank you very much. Thank you for listening to me. Happy to connect, LinkedIn, Twitter, in the hallway.

Thank you very much.