
Transcending time and space, you can talk with consumers and explore more opportunities and possibilities. That's why Opensurvey is building synthetic consumers.
In an earlier piece in this synthetic consumer series, we defined the synthetic persona as a tool for exploring why a consumer might feel a certain way. That naturally raises the next question: How is this persona created, and how far can its answers be used in real business decisions?
This piece walks through how Opensurvey builds its synthetic personas. We're transparent about how they're made for a reason — so users can judge for themselves how much to trust the results, and use them wisely.
The difference between a "plausible answer" and a "grounded answer"
Ask a general-purpose LLM to "answer like a woman in her 30s with an office job," and a plausible response appears within seconds. But there's a critical problem: there's no way to verify where the answer comes from. The model is only imitating general patterns from the web data it was trained on — it doesn't directly reflect target data from today's Korean market.
At the other extreme, a company can model its entire internal dataset to build virtual customers. The grounding is solid, but the upfront cost is steep — too heavy a solution for the light, sharp questions that come up constantly in day-to-day work.
Opensurvey took a third path between the two: grounded in real consumer data, yet generated lightly the moment a user enters a target.
The starting point is real consumer data
Opensurvey's synthetic personas aren't pre-built fictional characters. When a user enters the target they're curious about, the process begins by searching for the real data best suited to represent it. Opensurvey's trend report data — built from years of analysis across diverse industries and channels — serves as the base source, and customers who have accumulated their own data can draw on that internal data as well.
The reason for anchoring the starting point in real data is simple: only when users know which data a persona and its answers come from can they judge how far to trust and reference them.
A four-step generation process
1. Data Sourcing. When you enter a question like "I want to understand the target for a Korean skincare brand entering the U.S. market," the system searches for highly relevant data — such as U.S. K-beauty trend data — and sets it as the source.
2. Segmentation. Two things happen at once within the retrieved data. The AI proposes market-appropriate segments the way a researcher would, while the analytics engine cross-validates them to confirm whether they are genuinely distinct groups. Only the segments that are clearly separable in the data remain as final candidates.
3. Persona Generation. Drawing on the confirmed segment's gender, age, lifestyle, and purchasing behavior, the system extracts the group's core characteristics to generate a persona. The attributes it applies are the traits that stand out in that segment compared to the general consumer — producing a persona that feels true to the target, rather than a generic, average figure.
4. Grounding. Two roles operate together in conversation: the persona that speaks like a consumer, and the analytics engine that keeps its answers within the range of the actual data. Behind every answer sits a data analysis result, and even when there's no direct answer, the persona states which traits its inference is based on — for example, by showing the response tendencies of the underlying data beneath the reply.
How far to rely on it, and what to watch for
Even generated this way, a synthetic persona's answers are not the direct words of a real consumer. Rather than treating them as quotes that represent the market, the best use is to broaden your understanding of a market or consumer you don't know well, or to form hypotheses to confirm in your next study.
The richer the source data on a topic, the firmer the grounding; the further you move from the data, the greater the weight of inference. Opensurvey clearly separates the two, so users can see which answers are data-based and which are inferred as they talk. That said, metrics tied directly to real decisions — price acceptance, concrete purchase intent — still need final validation through actual consumer research.
Start now on Dataspace AI
Synthetic personas run on Dataspace AI. When an interview ends, you can ask DS AI — which understands the conversation's context — for a summary report, then move in one stop to designing a survey that validates the resulting hypotheses with real consumers. The conversation with a persona doesn't end as a simple Q&A; it's built to connect to your next business action.
In the next piece, we'll take a closer look — with actual interview screens — at what a conversation with a synthetic persona can give you.
Opensurvey
Opensurvey is an AI research tech company. We connect the entire research process—from research planning to data collection and analysis—with AI, and we offer a platform, expert research services, and a consumer panel all together. We work alongside industry-leading companies such as Samsung Electronics, P&G, CJ CheilJedang, and Woowa Brothers, and over the past 14 years we have served some 3,000 corporate clients across 25,000 projects. With ISMS-P, ISO/IEC 27001·27701, and ISO 20252 certifications, along with full membership in ESOMAR, we meet international standards in both security and research quality.
Dataspace
Dataspace is an AI-powered consumer intelligence platform provided by Opensurvey. An orchestrator that understands research context, together with specialized agents for each stage, accompanies the entire research process—from planning to data collection, analysis, insight reporting, and sharing. Its Dual Layer architecture, which separates statistical computation from AI inference, ensures analytical accuracy, and every insight is presented with evidence grounded in real consumer responses. You can connect with consumer panels in 20 countries including Korea, collect data directly from your own customers, or use APIs to integrate with external platforms such as CRM systems. The consumer data and research context accumulated in Dataspace remain as a company's intelligence asset. Building on this, you can create synthetic consumers tailored to your own brand to hold conversations with them and predict market responses.
ⓒ Opensurvey
All copyrights to Opensurvey's content belong to Opensurvey Inc. When citing, you must always indicate the source; even with attribution, copying content in full, unauthorized reproduction, and redistribution are prohibited. When citing the source, please include the name of the Opensurvey blog and the URL or link of the article cited. ex) Opensurvey Blog, article URL or link



