Synthetic consumers can produce answers. But can you bet a product launch on them?

People contradict themselves. They forget why they bought something. They claim to care about sustainability and then choose the product that arrives tomorrow (hello world!). They say they want fewer notifications and continue opening apps built around…

Idea 03 of 09

Idea 03 of 09

Plausibility is not prediction

All ideas

All ideas

  1. 01Introduction
  2. 02Synthetic answers are not human evidence
  3. 03Plausibility is not prediction
  4. 04The melody incident is a useful warning
  5. 05Consider a protein brand choosing its next flavour
  6. 06Where else can synthetic research mislead?
  7. 07The accuracy debate is a distraction
  8. 08Synthetic data is useful—but only within limits
  9. 09The question is not whether synthetic consumers are impressive

Introduction

Synthetic consumer research is becoming one of the most heavily funded ideas in AI.

Companies such as Simile are building digital representations of people that claim to simulate how consumers will respond to products, prices, features and marketing messages. The appeal is obvious.

Instead of spending weeks recruiting participants, conducting interviews and analysing responses, a company can generate hundreds or thousands of synthetic reactions within hours.

That makes synthetic research faster, cheaper and significantly easier to scale (without talking to a single customer!).

But speed is not the same as evidence.

The central problem with synthetic consumers is not that their answers are (always) wrong. It is that businesses often cannot determine when those answers are wrong. And that distinction becomes critical when the research is being used to make an expensive, irreversible decision.

Synthetic answers are not human evidence

Most synthetic research systems combine large language models with demographic information, behavioural datasets, social-media content, transaction data or previous research. The model then generates responses that resemble what a person from a particular segment might say.

But the answer still comes from a model.

There is no actual customer behind the quotation. There is no respondent whose circumstances can be examined. There is no interview recording to revisit, no behaviour to observe and no person whom the researcher can question further.

A synthetic respondent might say:

“I would pay ₹2,499 for a premium protein supplement because I value clean ingredients.”

That sounds useful. But what does it prove?

It does not tell you whether a real buyer will pay ₹2,499 when a competing product is available for ₹1,799. It does not reveal whether the buyer will abandon the cart after seeing the delivery fee. It does not account for advice from a gym trainer, distrust of an unfamiliar brand, concerns about taste or a spouse questioning the monthly expense.

The model has generated a plausible explanation. It has not observed a purchase.

That makes synthetic data difficult to use as evidence for high-stakes decisions. A company cannot confidently tell its board, product team or investors that customers demanded a feature when no customers were actually involved in producing the finding.

Plausibility is not prediction

Large language models are very good at generating the answer a reasonable person might give.

Unfortunately, humans are not consistently reasonable.

People contradict themselves. They forget why they bought something. They claim to care about sustainability and then choose the product that arrives tomorrow (hello world!). They say they want fewer notifications and continue opening apps built around notifications. They demand privacy but exchange personal information for a small discount.

Research has found that synthetic respondents can produce less variation than real consumers and may exaggerate the relationship between demographics and attitudes. In other words, they can make consumer segments appear more internally consistent—and more different from one another—than real people actually are.

This is a structural problem, not merely an accuracy problem.

A synthetic consumer is constructed from patterns. A real consumer lives inside constraints, relationships, habits, contradictions and moments of irrationality. Those forces are often what determine the purchase.

The melody incident is a useful warning

In May 2026, a video involving Indian Prime Minister Narendra Modi, Italian Prime Minister Giorgia Meloni and Melody toffees became popular on social media.

Retail investors subsequently bought shares of Parle Industries, apparently associating the listed company with Melody. But Melody is manufactured by Parle Products, a completely different and privately held company. Parle Industries had no connection to the chocolate.

The sequence was irrational but recognisably human:

PM Modi gifts Melody to Italian PM, Melodi
Parle makes Melody.
A listed company has “Parle” in its name.
Buy the damn stock.

Similar mistakes have happened elsewhere. After Elon Musk tweeted “Use Signal,” investors drove up the shares of Signal Advance, an unrelated medical-device company. Investors also repeatedly confused Zoom Video Communications with the unrelated Zoom Technologies.

These are extreme examples, but they expose the underlying problem.

Human decisions are shaped by availability bias, mistaken associations, social proof, fear of missing out and whatever happens to be salient at that moment.

A model trained to generate coherent behaviour may systematically underestimate incoherent behaviour.

Consider a protein brand choosing its next flavour

Suppose a protein company must decide whether to launch mango, chocolate hazelnut or unflavoured whey.

A synthetic panel can analyse category trends, reviews, demographic preferences and social-media conversations. It might conclude that mango will perform strongly among young Indian consumers because it is familiar, culturally relevant and differentiated from existing chocolate products.

That is a reasonable hypothesis.

But the actual purchase may depend on factors the model cannot experience:

Does mango whey taste artificial when mixed with water?

Does its smell become unpleasant after sitting in a shaker for an hour?

Do consumers associate mango with a refreshing drink rather than a heavy protein product?

Will gym trainers recommend it?

Does the bright packaging make it look less serious than competing products?

Will customers enjoy the first serving but become tired of the flavour after ten days?

These are not merely data points. They are experiences.

A synthetic consumer can describe what consuming mango protein might feel like. It cannot taste it repeatedly, become bored with it, regret buying a one-kilogram pack or leave the half-used container at the back of a kitchen shelf.

The distinction matters because the company is not deciding which idea sounds most appealing. It is deciding which product to manufacture, inventory, distribute and promote.

Where else can synthetic research mislead?

Consider pricing.

A synthetic respondent may make a rational trade-off between price, ingredients and brand reputation. A real buyer may choose the most expensive product because a fitness influencer recommends it—or the cheapest one because payday is still a week away.

Consider packaging.

An AI persona can evaluate the visual design presented on a screen. It cannot discover that the lid is difficult to open, the scoop gets buried in the powder or the container does not fit inside a kitchen cabinet.

Consider customer churn.

A model may infer that customers cancel because the product is expensive. Interviews might reveal that customers actually felt embarrassed asking the support team the same question repeatedly, or that a spouse objected to another subscription appearing on the credit-card statement.

Consider a new feature.

Synthetic users may consistently prefer more control and customisation. Real users may never discover the settings, may find them overwhelming or may continue using the default because changing established behaviour requires effort.

Consider advertising.

A synthetic audience may correctly understand the intended message. Real consumers may misinterpret one phrase, turn a screenshot into a meme or associate the campaign with a controversy that did not exist when the research data was collected.

Synthetic research is weakest precisely where businesses most need research: where context, behaviour and consequences matter.

The accuracy debate is a distraction

Synthetic-research companies frequently discuss whether their predictions are 80%, 90% or 95% accurate.

But an aggregate accuracy number tells a decision-maker very little.

Accurate at predicting what?

Under which conditions?

For which population?

Compared with which human sample?

Is the data citable?

Does the system predict average survey responses, individual choices, market share, repeat purchases or actual behaviour under financial constraints?

A system might reproduce broad consumer sentiment with 90% accuracy and still fail badly on the minority behaviour that determines whether a specific product succeeds. It could correctly identify chocolate as the safest flavour while completely missing the small but valuable group willing to pay substantially more for an unflavoured, clean-label product.

Even a highly accurate model can be dangerous when users do not know where the remaining errors are concentrated.

The issue is not whether synthetic data can approximate human responses. It clearly can. The issue is whether the approximation remains reliable when the market changes, the product is unfamiliar or the decision depends on an unusual human reaction.

Those are often the exact situations in which companies commission research.

Synthetic data is useful—but only within limits

Synthetic research should not be dismissed entirely.

It can help teams generate hypotheses, explore possible segments, stress-test questionnaires, identify obvious objections and narrow a large set of concepts before speaking with customers. It can also help researchers determine which questions deserve deeper investigation.

In these situations, synthetic data is being used as a thinking tool rather than as proof.

The mistake begins when generated responses are presented as customer evidence.

There is a major difference between saying:

“The simulation suggests that price may be a concern.”

and saying:

“Our customers told us that price is the primary barrier.”

Only the second statement requires actual customers.

A sensible research process can use synthetic consumers at the beginning to explore possibilities. But before making decisions about product launches, pricing, positioning, inventory or major investments, those possibilities must be tested with real people and, wherever possible, real behaviour.

The question is not whether synthetic consumers are impressive

The question is whether a company should manufacture ten thousand units, change its pricing or enter a new market based on what those consumers say.

For low-risk exploration, synthetic data may be sufficient.

For decisions where being wrong is expensive, plausible answers are not enough. Businesses need evidence that can be traced to real people, real circumstances and real decisions.

Synthetic consumers can tell you what might happen.

Human research tells you what people are actually experiencing.

And behavioural data tells you what they ultimately did.

What’s your take?

Showing Plausibility is not prediction, idea 3 of 9.