# Persona Engineering: A Field Guide to AI Synthetic Personas — Ishan Anand, InsightSciences.ai

AI Engineer · 2026-07-29

<https://aiengineer.podhood.com/e3418a8b-f274-47ef-bdd9-26059edb66cf>

Ishan Anand of InsightSciences argues that synthetic personas, powered by LLMs, can predict human survey responses with 83% alignment when normalized against human noise (humans are only 80% consistent with themselves over time). However, they fail in three critical ways: models invent confounders (e.g., price as a proxy for quality, creating an inverted U-shaped purchase curve), exhibit extreme order bias in answer choices, and predict stated attitudes far better than actual behaviors. Techniques like fine-tuning on human distributions (the subpop paper) or mapping model-generated text to human-scaled responses via semantic similarity can recover accurate distributions, not just averages. To validate, Anand recommends using correlation plus shape metrics and establishing a noise floor by splitting human data against itself. The takeaway: synthetic personas are forecasts, not ground truth, and work best as a complement to human research—extending data to unasked questions and simulating human-plus-agent ecosystems.

## Questions this episode answers

### How accurately can AI agents replicate human responses in market research surveys?

Ishan Anand describes a study where about 1,000 humans were extensively interviewed, and AI agents given those transcripts then took the same surveys. The agents matched the humans on 83% of responses. However, this figure is normalized: when the same humans retook the surveys two weeks later, they were only 80% consistent with themselves, setting a noise floor that limits how accurate any method can be.

[4:03](https://aiengineer.podhood.com/e3418a8b-f274-47ef-bdd9-26059edb66cf?t=243000)

### Why would a synthetic persona's purchase probability increase as the price goes up, contradicting basic economics?

In a market research experiment, Ishan Anand notes that while real humans showed decreasing purchase probability as price increased, LLM agents produced an inverted U-shaped curve. Through additional tests, researchers found the AI was using price as a proxy for latent product qualities like quality or competing prices, effectively inventing confounders because the prompt lacked sufficient context to fix those variables.

[5:39](https://aiengineer.podhood.com/e3418a8b-f274-47ef-bdd9-26059edb66cf?t=339000)

### How can you capture the full distribution of responses from synthetic personas, rather than just average likelihood scores?

Ishan Anand describes a technique where instead of asking a model for a 1-to-5 rating, you have it output free-text reactions. Human annotators write example texts for each rating, and you measure the semantic similarity between the model's text and those examples to build a probability vector over the possible scores. This preserves response variation and more accurately mirrors the distribution shape found in real human data.

[13:42](https://aiengineer.podhood.com/e3418a8b-f274-47ef-bdd9-26059edb66cf?t=822000)

## Key moments

- **[0:00] Introduction**
  - [0:13] Synthetic personas are forecasts bounded by their data and compute regime, like weather forecasts limited to a certain forecast horizon.
- **[1:43] Motivation**
- **[2:46] Weather analogy**
  - [2:48] Simulmatics in the 1950s promised to predict the electorate with raw statistics but failed, cautioning against overhyping synthetic personas today.
  - [4:03] AI agents replicated human survey responses with 83% accuracy after extensive interviews, normalized against humans' own 80% self-consistency rate.
- **[4:04] Agent replicas**
  - [5:15] LLM purchase probability formed an inverted U-curve as price rose, inferring latent confounders like product quality from price changes.
- **[5:22] Confounders**
  - [7:10] "It's a little like the LLM is playing improv with you."
- **[7:44] Order bias**
  - [7:44] Swapping the order of yes/no options caused extreme LLM response bias, averaging to 50/50 noise, far stronger than human first-order bias.
- **[8:36] Attitudes vs actions**
  - [8:36] LLMs predict stated attitudes more accurately than behaviors because they are trained on what people say, not what people do.
- **[9:52] Prompting**
  - [9:52] Early persona construction used text completion prompts like 'In 2016, I voted for' to sample political voting preferences.
- **[11:47] Fine-tuning**
  - [11:47] Fine-tuning on survey data improved alignment for both seen and unseen demographic groups, suggesting LLMs have latent group knowledge.
- **[13:08] Semantic mapping**
  - [13:08] Using free-text responses and semantic similarity to human examples captured choice variation as a probability distribution, not just point estimates.
- **[15:28] Measuring alignment**
  - [15:53] Re-running a synthetic persona without changing inputs does not increase statistical significance, just as re-running a weather forecast doesn't improve its accuracy.
  - [17:28] Human self-consistency measured at 80% sets a noise floor for synthetic persona accuracy, achievable via split-half testing of ground truth data.
- **[18:38] Conclusion**
  - [18:38] Synthetic personas are economic forecasts, not people, and should be validated against real outcomes, not treated as ground truth.
  - [20:12] Generative agent-based modeling will enable simulating interactions between personas, turning human data into a living queryable asset.

## Speakers

- **Ishan Anand** (guest)

## Topics

Synthetic Persona Engineering

## Mentioned

InsightSciences (company)

## Transcript

### Introduction

**Ishan Anand** [0:13]
Hello, I'm Ishan, and welcome to "Can AI Predict People Like We Predict the Weather," a field guide to the nascent field of synthetic personas. Now, I'm sure most of you in this room, at some point or another, have prompted a large language model with a role prompt: "You are a..."

fill in the blank, and then the task. Believe it or not, that core principle of steering a model's outputs as if it were a particular person or persona has turned into an entire category that companies are using to test product concepts and messaging against synthetic respondents.

And it has moved from a novelty to market momentum, as you can see from both these headlines as well as the increase in funding for the last few years. And the running analogy I want to leave you with is that synthetic personas are like weather forecasting.

Like weather forecasting, they were unlocked thanks to an increase in compute and data. And like weather forecasting, they operate within a particular regime, and going past that sometimes can go outside of where they're accurate. So, for example, you can only predict the weather a certain number of days in advance.

Similarly, with synthetic personas, there's only so far you can go before you'll run into issues. And understanding those issues is as important as understanding their promise and their potential.

### Motivation

**Ishan Anand** [1:44]
So, I'm Ishan Anand. I'm the Chief AI Officer at InsightSciences. We construct LLM synthetic personas for market research and market insights teams. And the reason for this talk is that most of the coverage in this space is very shallow, doesn't go into the technical details.

It's either outright hype or outright dismissal. And, uh, it's really hard to separate the noise from what's real. And so what I want to cover is that messy middle of the technical details, and you don't have to take my word for it, even though I'm a vendor in the space, because everything I'm going to talk about today is going to be based on published research.

So we're going to cover why now for synthetic personas, how they fail, some techniques to inspire you, and then some metrics to judge whether your synthetic persona is accurate or not. Speaking of weather forecasting, another parallel is just like in the 1950s and '60s, when you got computers that promised us, correctly, a future of accurate weather forecasts.

### Weather analogy

**Ishan Anand** [2:48]
We were also promised, believe it or not, people forecasts. This company, Simulmatics, were, uh, extensively covered by Jill Laporte, promised that they could simulate and predict the electorate using raw statistics and the computational power at the time. Fortunately, that turned out not to be the case.

So we should approach claims like this with some humility. But we have something they did not have then. And that unlock is, again, more computational power, but also better modeling thanks to LLMs. And LLMs unlock a new kind of simulation.

For the longest time, to simulate something meant to mathematize it in formulas or equations. But certain things—how we feel, how we act, what choices we make—aren't always succumbing to the equations. And what LLMs offer us is a new medium, a new atomic unit of language itself that we can model against.

Now, granted, they are based on math under the hood, but it gives us this intermediary layer that we can construct and simulate against that we couldn't before. And the process can work. I want to share with you one of the most, uh, well-known demonstrations of this in the field.

Um, what they did is they took about 1,000 humans. They put them through about 2 and a half hours of extensive interviews about their background and their views and their attitudes. And they put those people through a battery of personality tests and surveys.

### Agent replicas

**Ishan Anand** [4:20]
Then they took those transcripts and they passed it to an AI agent. And they had the AI agent take the same set of surveys and personality tests. And what they found was that the agents were basically about 83% aligned and predictive to the corresponding humans they were modeled against.

Now, one caveat is that number is normalized against the uncertainty and noise of the humans themselves. It's a theme we're going to come back to at the end of this talk. But don't get too excited, because synthetic personas are different from regular experiments, and they're liable to confuse and fool you if you don't know how they fail.

So I'm going to cover three important failure modes that you need to know about when dealing with synthetic personas. To understand the first one, I want to consider this prompt these researchers gave. It's a very im it's a very ambiguous and very unsophisticated prompt.

It basically says, "You are a customer. I'm going to show you a product. I'm going to tell you the category. I'm going to give you its price." Those are going to be the variables in the template. And then I'm going to ask you to say whether you're going to purchase or not purchase.

### Confounders

**Ishan Anand** [5:27]
Willingness to pay, willingness to purchase is basically the test. And what the researchers did is they recruited a panel of humans and put them through the same test, and then they put the synthetic personas through the same test.

And what they found is very interesting. So the humans are here in red. They do exactly what you would expect from basic economic theory. As the purchase price increases, we see that the purchase probability goes down, slopes downward.

But the LLMs did something different. They had this inverted U-shaped curve. And particularly problematic is this arearight here, where as the price is increasing, the purchase probability is going up. That seems really bizarre. Through a series of additional experiments, what they discovered was that the LLM was using the price as a proxy for other properties about the product that the humans were considering were fixed.

Things like the expiration date based on the price, what the price of competing products were also as the price changed. And those correlations, those latent confounders that weren't clear and immediate were actually confusing the result. And the way to think about this is when an LLM is missing context, it has to potentially infer or invent confounders,right?

When we do a human experiment, if I put, like, a gold watch on a table, I ask a human to walk in and estimate the price of it, everything about the environment is fairly fixed. The human and their decisions are the random variable.

In a synthetic experiment, if you don't set it up properly, other parts of it actually become part of the random variable itself. I like to say if it's a poorly grounded persona, it's a little like the LLM is playing improv with you.

It's like, "Gold watch on a table? Oh, well, we must be in a jewelry store,"right? It has to infer what's likely. And maybe this is a rich person, so they're more likely to purchase. And so the lesson is we need to richly ground our personas in the personality, the context, and bizarrely, even the study's own construction.

In a human subject experiment, you want to hide the study construction from the participant. But in the case of an LLM, they have no universe other than what's in the prompt. And you have to use the prompt to paint the world to prevent any type of confounders.

Another failure mode is prompt sensitivity. So here's a researcher that took a question. They gave the same question, same choices, they just swapped the order of the choices. Yes was the first on the first question. Yes was the second option in the second question.

### Order bias

**Ishan Anand** [7:59]
And what they found was that the model had an extremely strong order bias. Basically, when they took the two results and they averaged them together, it washed out into noise, into 50/50. Now, humans do have a first-order bias, but not to this extent.

And so the lesson here is that we need to durability test our personas to understand how they will change under reorderings, under rewordings, and even adversarial challenges to their opinions. The third and final area that I want to highlight is that LLMs are trained on what people say, and they're not trained on what people do.

So as a consequence, predicting stated attitudes tends to be easier than predicting actions or behaviors. Both because they're clear and likely to be in the text, but also because they are natively text themselves. So this chart is from a bunch of researchers that used an LLM to try and predict known social science experiments.

### Attitudes vs actions

**Ishan Anand** [8:57]
The original point of this chart is to show that the LLMs are about as good as the experts. LLM is in black in a circle. The experts are in blue. And you can see they're both doing about equally well in making the prediction.

But the point I want to draw you to is that there are two categories of experiments here. The top are surveys. Those are natively language and text-based. And those reflect attitudes. And on the whole, the models tend to do better there.

The bottom half is field experiments. Those are behaviors, and those are things that need to be transcribed into actions. They're less likely to be in the training data. And correspondingly, the LLM doesn't do as well. So the lesson we often tell our clients is consider questions that triangulate to behavior from attitudes.

As a hypothetical example, if you want to know about gym attendance, you might be better off well, you can ask about both, but asking about attitudes towards working out rather than asking about attendance, and see if that's a suitable proxy.

Okay. Now let's talk about three example techniques to kind of inspire your own synthetic personas. So the first one is just prompting the model. Uh, thisright here is from the Argyle paper, which is really one of the seminal papers in this field.

### Prompting

**Ishan Anand** [10:09]
In fact, it's so early that the model they used was a text completion model. That's why this prompt isn't in the form of a chat. It's a statement of, "I am." So they gave it a prompt that said, for example, the middle column is basically where the context is.

"I am a strong liberal. I support progressive values," et cetera, et cetera. And at the end, it says, "In 2016, I voted for." And they basically gen let the model sample its completions, and it says, uh, Hillary Clinton, Bernie Sanders, Hillary Clinton, and so forth.

And you can see what happens for the conservative case on the top. Since this time, obviously, there have been a lot more prompting techniques and a lot more models. And I can't tell you which prompting technique and which model is going to work best for your use case.

What you are going to have to do is figure it out empirically by validating against some known human ground truth data. You'll have to do what these guys did. So, for example, here in this research, they're trying to figure out how well they can construct personas to represent voting patterns.

What they found was they compared here on the left is reality, and on theright is there four different types of persona constructions. And they didn't realize it at the time, but their persona construction was actually amplifying bias within the model as they got more and more detailed.

And they found it was actually throwing it further and further astray from reality. So you probably have a bunch of different ideas. The answer is you're going to have to test it and validate it against ground truth. The other natural thing you might expect is, well, hey, we can fine-tune it, especially if it's missing data that isn't there, especially, for example, if it's behaviors, uh, or something that wouldn't be in the training text.

And this is, uh, a great paper to be inspired by for this. This is the subpop paper. Basically, they construct a prompt template, which is the demographic information, then the survey question they want to ask, and then they compare the known human data distribution to the distribution that comes out of the model.

### Fine-tuning

**Ishan Anand** [12:04]
And they do fine-tuning until the model and the human data align. Now, here's the interesting thing. When they did this, as you'd expect, the results that were from the populations they gave to the model, that's the ones in blue, improved.

But very interestingly, the ones in white also improved by almost the same degree. Alignment improved even for the unseen groups. That seems almost magical. And some subsequent research has hinted that what might be really happening here is that the model itself has a latent understanding of these groups.

It just didn't know how to express it in the format of surveys. And if you think about it, LLMs aren't used to doing surveys as a task. And so they aren't going to be as good as fitting it, especially to a prompt format they may not have seen on the first go-round.

But fine-tuning actually is helping it learn the task or how to express itself. So a lesson you can kind of take away is that your persona that you're looking for is in there. We just need to figure out the way to summon it or elicit it.

### Semantic mapping

**Ishan Anand** [13:08]
And that lesson actually takes us to the third technique, which I want to highlight to show how sophisticated your techniques can get if you're just using so-called prompting alone, but using careful calibration and thinking. So, uh, in this one, this team did something very clever.

They set up a system prompt that was demographics. They showed a product concept. And then they asked, "How likely would you be to purchase this product?" And they gave it the same scale from 1 to 5, 5 being the most likely, 1 being the least likely to purchase, like you'd expect, kind of your basic naive prompting pattern.

And then they said, "Well, you know," hearkening back to that paper, although I don't know if they were inspired by it, they said, "Well, large language models aren't used to doing surveys, but they are more used to expressing themselves in text."

So what they said is, "Instead of giving us a 1 to 5 rating, give us a set of text." So the example here is, "I'm somewhat interested. If it works well and isn't too expensive, I might give it a try."

And then to map that text to the 1 through 5 willingness to pay, they had humans write out corresponding text for what they would expect. So if it's a 1, "Hell no, I'll never buy that. 5, absolutely, I'll buy 20,"right?

They had them write out examples of each one of the different options. And then they measured the semantic similarity between the text that came out of the model and those human examples. And that gave them a vector over which they can basically measure a probability distribution of where this text that came out of the model lands.

So what I like about this is it's actually a distribution. It kind of feels like, you know, humans. Some days I might say 4. Some some days I might say 5 in this graph, but rarely would I say 1, 2, or 3 in this example.

And what they were able to show is that they were not able to only reconstruct accurate values for willingness to pay. They were able to capture the distribution. Because one of the important failure modes we haven't talked about is that LLMs, even when they get the persona averagesright, they very often lose the details.

The variations get muddled together in the middle. This chart at the bottom, basically that horizontal axis, is a measure of the entire shape similarity. And 1 means perfectly identical and 0 means not. And what you can see is the naive way in the purple, or I guess pink, uh, doesn't do as well as the yellow, which is up near the top of the range.

### Measuring alignment

**Ishan Anand** [15:28]
So that means it really did a good job not only understanding what the ultimate choice was, but how well that choice varied. Okay. Let's talk about how to measure alignment from a synthetic persona. Um, one of the things that our traditional market research, uh, clients are sometimes surprised by and disappointed is that you cannot use statistical synthetic personas to boost statistical significance.

You can take an underrepresented population and get more values out of it, but you can't say it's statistically significant. And to understand this, it helps to go back to that weather analogy. If I want to know how much it rains today in San Francisco, and I used to live here, so I know it rains a lot, I'd stick a weather gauge.

And if I wanted to know with more certainty, I'd stick 1,000 weather gauges. And those would increase the accuracy of my estimate. But if I want to know if it's going to rain tomorrow, if I take a forecast and I rerun it 1,000 times without changing the input, that doesn't change my certainty of that forecast.

It improves my estimate of what the model is telling me, but it doesn't make the forecast itself more accurate. And that's what happens when you basically are rerunning a synthetic persona with no changes to input. So the lesson is more synthetic samples aren't actually going to improve your statistical significance for the most part.

So what you need to do is you need to do what you do with weather forecasts. You'd basically check against what actually happened, or in our case, what humans actually said. And that's where we're going to basically be measuring distributions of data.

Unlike classic evals, where there's clearly aright and wrong, and you can score how many wereright and how many were wrong, now we need to measure the data as a comparison of distributions. And there are many ways for distributions to get wrong.

They could be completely wildly off. They can, as we mentioned, get the averageright, but the shape of the distribution wrong. And so you're going to need multiple metrics to capture how well your model is reflecting different personas. Um, I recommend using a correlation-type metric along with one of these shape-type metrics which capture what the underlying shape of the distribution is.

The other thing you need to do is estimate the fundamental noise in your ground truth data. That experiment I talked about in the beginning, where they got 83% accuracy, the key smart thing they did is they took those humans and they brought them back two weeks later.

And they redid the battery of surveys and personality tests. And they found that the humans, on average, were only 80% consistent to themselves. So that sets a noise floor as to how accurate our models could ever get because the humans themselves are fundamentally noisy.

And so the 83% is actually normalized against that. If you can do this, bring your humans back, that's great. Very often you can't. So the way you can kind of artificially do this is take your ground truth human data, break it into two chunks, and then pretend one is synthetic and one is human, and then measure the correlation and repeat that hundreds and thousands of times and average it.

And that'll set kind of a noise floor that your ground truth data, where half of it's synthetic, half of it real, could be the level of accuracy you could hope to get. So hopefully by now you have an appreciation for why I think weather forecasts are the best lens to understand synthetic personas.

### Conclusion

**Ishan Anand** [18:38]
They are not people. They are forecasts. And we should treat them accordingly. Both systems are bounded. Both systems will be improving over time. And they're most trustworthy when they are validated against reality. Um, now synthetic personas are very often cast in the market against human research.

And I think that's unfortunate because they're actually complementary to each other. And I'll give you two reasons why. One is that we're entering an era where humans are no longer the sole economic actor. Every action your human customer is taking in terms of awareness, consideration, or a purchase decision to buy is being increasingly mediated by AI agents.

So a human-only study is actually not the gold truth. What we really need to understand is what does the human-plus-agent ecosystem look like?

And then finally, the alternative to a synthetic persona is not human research. In most cases, it's no research, or it's somebody's opinion. What really happens is you've done a survey of humans, and you get a question, and if it's in the survey, you can just answer it.

That's very simple to do. But what typically happens is it's two months later, and you're like, "We need to answer this question," which we didn't ask. Well, then somebody needs to be like, "Uh, I think it would be this by extrapolation."

A expert plus a synthetic persona is going to give you a better result to that. So what we like to tell customers is synthetic extends your human data to more phases of your development process. It can go more places your existing research can't.

One of the most exciting directions is to actually run simulations. We didn't get time for this, but it's called generative agent-based modeling, where we can take each of these personas and simulate what the dynamics and how they'll interface and interact with each other.

And ultimately, what this will let you do is turn your human data into a living queryable asset. If you're interested in doing that with your data, feel free to reach out to us. We help market research and insights teams generate and use synthetic personas in AI.

You can find us on the web at insightsciences.ai. And my contact information is on the slide. I hope you have a good conference. Thank you.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
