/
/
Science
Science
Synthetic Consumer
Synthetic Consumer
Synthetic Consumer

How Opensurvey Built Its Synthetic Panel

Talk with consumers across time and space, and explore more opportunities and possibilities. That is why Opensurvey builds synthetic consumers.

How Opensurvey Built Its Synthetic Panel

You have probably heard by now that AI can answer surveys in a consumer's place. Abroad, several services already go by names like synthetic respondent and synthetic panel, and in academia there is active research on simulating opinion polls with LLMs.

But even when two things share the same word, "synthetic panel," the character of what you get depends entirely on how it was built. In this piece we place the market's common approach and Opensurvey's approach side by side, and explain why the Opensurvey team chose the approach it did.

The market's approaches: starting from a plausible fictional person

The most widely used approach in the market today looks like this.

You give an LLM a prompt with demographics and a few disposition traits, such as "You are a woman in your early thirties, an office worker living in Seoul, and you care a lot about your health." Then you show that fictional person a concept and ask for a reaction. In academia this is sometimes called silicon sampling. It is fast, it is cheap, and it lets you ask about any topic instantly. Those advantages are real.

The market is aware of the limits of this approach, of course. So recent efforts have tried to give the fictional person more grounding by combining real survey data or a client's own data. That is clear progress. But if you look closely at the shape of the data being combined, most of it amounts to laying group-level summaries onto the prompt, along the lines of "consumers in their thirties who lean health-conscious tend to behave this way."

So whether you use the prompt alone or add group summary data, a question arises.

Who, exactly, is this fictional person?

What an LLM knows about the "woman in her early thirties, office worker living in Seoul" in the prompt is the average conventional wisdom about that group as learned from the internet. Adding group summary data only corrects that conventional wisdom toward the average of real data. It is still an average. The complexity of a real consumer, the person who talks about valuing health yet actually enjoys spicy malatang and donuts, gets erased inside the average.

And while you can combine individuals' data and take an average, you cannot break a person molded from averages back down into individuals. Even if you create a hundred fictional people who all start from the average, they are not a hundred individuals. They are a hundred variations on the same average.

Opensurvey's approach: we don't invent, we describe

Opensurvey defined the problem differently. Not as the problem of inventing a plausible fictional person, but as the problem of accurately describing each real respondent, one by one.

Opensurvey's synthetic panel is an AI virtual respondent built to correspond to the data of each real respondent, one person at a time. This difference in definition is decisive because the thing it has to resemble actually exists. Of a person molded from averages, you can only ask whether they seem plausible. A synthetic panel that starts from a real person is built and refined by asking how closely it resembles that person. In practice, Opensurvey sets aside real survey items that were not used in building the panel, then checks how well the panel reproduces those answers.

Behind the fact that this approach was possible lies a data asset unique to Opensurvey.

① A difference in materials: words (surveys) together with behavior (Food Diary)

For years, Opensurvey has used its Food Diary to record what consumers actually ate, when, and how. To this we add the same respondents' profiles and their surveys on food and purchase attitudes.

This combination matters because, often without meaning to, consumers say one thing and do another.

Let me show you a real example. There is a male respondent in his late thirties who answered in a survey that he values health. Going by the attitude survey alone, he would be classified as a textbook "health-oriented" consumer. But open his Food Diary and, right next to the brown rice and protein bars, you find spicy instant noodles, donuts, and burgers.

Opensurvey's synthetic panel captures even this gap between words and behavior. The persona generated from this respondent keeps the contradiction of "conscious of health but not one to hold back" intact, which makes possible a reading with real texture: for this person, a message that eases guilt is likely to work better than one that demands restraint. That is a dimensionality hard to get from a fictional person invented from demographic prompts alone.

Throughout this entire process, personally identifying information such as names and contact details is not used, and the persona contains nothing that could identify a specific individual. A synthetic panel's persona is not a copy of a particular person. It is a description of the eating patterns that one real respondent showed.

② A difference in how it is built: core design principles

First, a statistical model and an LLM divide the roles. This structure was not there from the start, to be honest. We tried leaving it all to the LLM, and we also tried predicting directly with a statistical model alone. Leaving it all to the LLM sent us back to the average conventional wisdom described earlier. A statistical model alone could give accurate numbers but never a "person" you could talk with. Each of those failures became the basis for the current structure. The statistical model quantitatively picks out only the signals that actually move this person's responses, and the LLM integrates those signals into a person with contradictions and preferences intact. Rather than training a giant new model on respondent data, we build each persona, one by one, using verified signals as the material.

Second, we build at the level of the individual respondent. Instead of an average representative of a segment, we build a persona that corresponds to each respondent, one by one. That is what makes it possible to analyze not only the overall average but also how reactions differ across subgroups such as gender and age.

Suppose there is a concept called an "animal-welfare egg sandwich." Two people who both answered positively to the same concept may have reacted for different reasons: one to the value of "animal welfare," the other to the fact that it is "easy to eat on the go." Approach it by segment average and these different reasons get lumped into a single number. Approach it at the respondent level and you can look into why each individual judged as they did. And that "why" becomes the hint for how to develop the concept.

The market and Opensurvey's synthetic panel, at a glance


The market's common approach

Opensurvey's synthetic panel

How it is built

A fictional person generated from demographic prompts, or from group-level summary data

A virtual respondent built to correspond to each real respondent, one by one

Source data

General knowledge from the public web that the LLM learned

General knowledge from the public web that the LLM learned, plus the data assets Opensurvey holds

Individual dimensionality

Tends to converge toward the average

By combining surveys (words) with the Food Diary (behavior), it reflects each respondent's tastes and even the gaps and contradictions between words and behavior

What we have built so far, and what to watch for

Frankly, the synthetic panel cannot answer every survey right away. In principle it is structured to answer any survey, but doing everything well at once is difficult. So we started with a problem we can do well and whose value is clear: preference testing, or screening, for new product ideas in the food and beverage domain.

We chose food and beverage as the first domain for two reasons. Generating enough ideas and picking the promising ones is a clear bottleneck for food companies, and our behavioral data asset, the Food Diary, supports exactly this. After validating the methodology in this domain, we plan to expand into other categories.

There is also something we want to be clear about. The synthetic panel solves a different problem from the human panel. Questions that must be asked of real consumers still have to be asked of real consumers. The synthetic panel is a tool for the stage before that: the stage of narrowing and filtering what to ask, so those decisions come faster.

How you can use it

You can meet Opensurvey's synthetic consumers in two forms. Right now, you can broaden your understanding of consumers by talking with a synthetic persona inside Dataspace AI, and there is Concept Studio, where you generate new product ideas and screen their appeal with the synthetic panel. In Concept Studio, the synthetic panel plays the role of simulating which of hundreds of candidate ideas to turn into concepts and validate with real consumers.

Synthetic Consume

AI

Opensurvey

Opensurvey is an AI research tech company. We connect the entire research process—from research planning to data collection and analysis—with AI, and we offer a platform, expert research services, and a consumer panel all together. We work alongside industry-leading companies such as Samsung Electronics, P&G, CJ CheilJedang, and Woowa Brothers, and over the past 14 years we have served some 3,000 corporate clients across 25,000 projects. With ISMS-P, ISO/IEC 27001·27701, and ISO 20252 certifications, along with full membership in ESOMAR, we meet international standards in both security and research quality.

Dataspace

Dataspace is an AI-powered consumer intelligence platform provided by Opensurvey. An orchestrator that understands research context, together with specialized agents for each stage, accompanies the entire research process—from planning to data collection, analysis, insight reporting, and sharing. Its Dual Layer architecture, which separates statistical computation from AI inference, ensures analytical accuracy, and every insight is presented with evidence grounded in real consumer responses. You can connect with consumer panels in 20 countries including Korea, collect data directly from your own customers, or use APIs to integrate with external platforms such as CRM systems. The consumer data and research context accumulated in Dataspace remain as a company's intelligence asset. Building on this, you can create synthetic consumers tailored to your own brand to hold conversations with them and predict market responses.

ⓒ Opensurvey

All copyrights to Opensurvey's content belong to Opensurvey Inc. When citing, you must always indicate the source; even with attribution, copying content in full, unauthorized reproduction, and redistribution are prohibited. When citing the source, please include the name of the Opensurvey blog and the URL or link of the article cited. ex) Opensurvey Blog, article URL or link

CraftofResearch

Opensurvey Inc. | CEO Hwang Hee-young | 12-15th floor, 13, Gangnam-daero 84-gil, Gangnam-gu, Seoul | 02-2070-2110 | Business Registration Number: 106-86-77081 | E-commerce Sales Business Report Number: 2022-Seoul Gangnam-04044

Certification Number: ISMS-P-KISA-2023-027
Certification Scope: Research Platform Service
Validity Period: 2023.7.5 ~ 2026.7.4

Copyright © Opensurvey Inc.