AI in Market Research
Synthetic Data, Digital Twins & Digital Personas
TRC Insights uses AI-powered synthetic data to help teams reduce uncertainty earlier, faster, and at lower cost – without replacing the human research that anchors a decision. Our two synthetic-data solutions, Digital Twins and Digital Personas, are built from your own validated human survey data, tested against real respondents, and developed in collaboration with academic researchers.
At TRC, AI Is Not New
We have integrated machine learning, natural language processing, and simulation modeling into our research process for years, and our synthetic-respondent work extends that foundation.
What Are Synthetic Data in Market Research?
Synthetic data are AI models trained on real human survey data that simulate how those people – or segments of them – would answer new research questions. They are built to mirror the relative differences and conclusions found in human data, not to replace it.
In market research, the credible use of synthetic data is to reduce uncertainty – to refine concepts, screen ideas, and explore more pathways before committing scarce human-research budget – rather than to cut cost by replacing real people.
There is an important distinction the industry often blurs: synthetic data is only as good as what it is grounded in. As TRC President Rajan Sambandam and Columbia Business School professor Oded Netzer write in their Quirk’s article, “The Case for Synthetic Data” (November 2025), synthetic data “done well” can reorient the research process and deliver better insights – but it has a low barrier for entry and a high barrier for excellence. Off-the-shelf synthetic panels trained on generic data, or AI personas conjured from a one-line prompt, produce plausible-sounding answers that lack real grounding. TRC’s approach is built on the opposite principle: every synthetic respondent is modeled from real, targeted human data and validated where validation is possible.
Three common uses of synthetic data and where TRC stands
- Boosting sample size (adding synthetic data to “fill” a quota) – TRC does not do this. As Sambandam and Netzer explain, augmenting human data with synthetic data masks uncertainty rather than reducing it; the apparent gain in sampling precision is an illusion.
- Off-the-shelf synthetic panels (pre-generated respondents on generic data) – TRC does not rely on these, because they risk misaligned audiences and answers that sound right but lack relevance.
- Simulating respondent answers from real, targeted data (Digital Twins and Digital Personas) – this is TRC’s approach. It is grounded in actual survey responses, preserves relative differences and tradeoffs, and extends the value of data you’ve already collected.
Digital Twins - Simulate Individual Respondents, Grounded, and Validated
A TRC Digital Twin is a virtual model of an individual real respondent, built from that person’s own survey responses. A panel of twins lets you ask new quantitative questions and get simulated individual-level answers in days rather than weeks – useful for screening concepts, narrowing lists, testing messages, and getting an early read on market changes before you field a full human study.
How TRC Builds Digital Twins
- Identify the target audience – consumers, B2B buyers, or specific segments.
- Conduct human research – comprehensive individual-level data collection (attitudes, behaviors, needs, pain points, preferences, demographics, and more), typically around 200 respondents per subgroup.
- Train the model – each twin learns from its own respondent’s data; models stay independent from one another.
- Validate against holdout data – twin answers are compared to real human answers on questions the model never saw, producing a per-project accuracy scorecard. But our team of analysts also measures how much of a twin’s answer actually derives from the client’s survey data versus the LLM’s own general knowledge. This is why it matters. If the model leans on its own priors rather than the real respondent data, you get stereotyping and generic answers. Hold-out testing alone won’t catch that – you have to measure the data’s contribution directly.
- Establish scope and guardrails – our AI R&D reviews prompts for relevance and appropriateness for the model, to flag out-of-scope questions.
- Deploy – the twin panel is “always on,” with access managed by TRC.
Built from your research
Our digital twins are built using synthetic data generated from research conducted specifically for your organization. Because the underlying data comes from your own customers, employees, or stakeholders, the resulting models remain grounded in your business context rather than generic consumer datasets.
This is an important distinction. Some AI research solutions rely on proprietary synthetic panels built from data their providers own. While those approaches can produce rapid results, organizations have limited visibility into how representative that underlying data is for their specific business. Our approach starts with research you own, creating a direct connection between your original data and the AI twins built from it.
Why validation is the differentiator
TRC validates twins on a per-project basis using holdout testing: a variable is withheld during training, the twins are asked about it, and their answers are compared to the real human answers on mean scores, top-box, top-two-box, and ranking. This matters because, as the industry and the academic literature both note, there is no universal “expected accuracy” benchmark for synthetic data – accuracy figures quoted in magazines and on LinkedIn vary widely and are derived differently. The only honest benchmark is comparison against your own human data, every time.
What Digital Twins are good for and what to watch
Twins are directional. They replicate the patterns, rankings, and tradeoffs in human data, not exact point estimates. TRC’s testing has consistently found that twins tend to skew more favorably than humans, can mute or rationalize price sensitivity, overweight tech-forward attributes, and are weaker on deep emotional nuance and lived experience. The implication, which we build into every engagement: treat twin outputs as directional ranks and deltas, use them to screen and refine, and reserve final pricing, deeply emotional, and experiential questions for human research.
Digital Twins replicate the direction of human data – the rankings, deltas, and tradeoffs – not the exact scores. They are built to tell you which option wins and why, faster and cheaper, so you know where deeper human validation is most worth it.
Digital Personas - converse with a segment, for inspiration and exploration
Digital Persona is a single AI chatbot that represents the aggregate voice of a real audience segment, built from your quantitative and qualitative data. Rather than producing individual data points, a persona is something your team talks to – a conversational way to stress-test messaging, explore segment-specific reactions, and spark new perspectives in real time, one persona at a time (much like a series of in-depth interviews).
How TRC Builds Digital Personas
- Identify definable segments (not general population).
- Collect and synthesize data – surveys, qualitative interviews, existing personas and segmentation models, aggregated at the segment level.
- Train one persona per segment to represent that segment’s aggregate voice.
- Test persona responses against known segment behaviors and attitudes.
- Establish scope – our AI R&D reviews prompts for relevance and appropriateness for the model, to flag out-of-scope questions.
- Deploy as a secure, login-based chatbot, available on a monthly subscription.
Digital Twins vs. Digital Personas - which should inform a decision?
The core difference is the level of grounding and what each can be validated against. Digital Twins are modeled on individual respondents and can be statistically validated, so they’re suited to informing decisions. Digital Personas represent a group, so they cannot be validated with the same statistical rigor – they’re best for inspiration, exploration, and bringing a segment to life, not for decision-grade measurement.
This is the crucial, often-misunderstood point: a Digital Persona cannot be statistically validated the way a Digital Twin can, because it is built from summarized data about a group of people rather than from any one individual’s responses. Its validation is qualitative and more subjective. There is no individual “ground truth” to hold out and test a persona’s answer against – by design, a persona is a single voice standing in for many. Its validation is qualitative (does it behave consistently with what we know about the segment?), not statistical. That makes personas excellent for teaching teams about a segment, exploring hypotheses, and generating ideas – and unsuitable as a substitute for measured, decision-grade data.
As Rajan Sambandam and Oded Netzer note, modeling and predicting individuals is hard, but the business question is usually about the segment, not the individual – twins let you start at the individual level and aggregate up to the segment that actually drives the decision.
Both simulate audiences – but they’re built differently and answer different questions.
Twins model real individuals for directional decisions; personas model segments to spark thinking.
Here’s how they compare.
| Dimension | Digital Twins Modeled from real individuals |
Digital Personas Modeled from segments |
|---|---|---|
| Built from | Individual-level survey data | Summarized, aggregated segment data |
| Represents | One real individual | One segment – a defined group |
| Output | One response per individual – primarily quantitative | One response per persona – primarily qualitative |
| Validation | Statistically validated, holdout-tested per project | Qualitative only – more subjective, not statistically validated |
| Best for | Screening, message testing, narrowing lists, understanding the “why” behind an answer, and faster decisions | Teaching teams about segments, exploring hypotheses, and sparking new ideas |
| Decision role | Inform strategy – directional | Inspire and explore |
How TRC ensures quality and responsible use
TRC’s synthetic data sit at the “high barrier for excellence” end of the spectrum, distinct from two weaker approaches common in the market:
- Basic prompting (“act like a telehealth patient”) – no real data foundation, high drift risk; useful for brainstorming, not decision support.
- Contextual persona prompting (“here’s a profile, respond as them”) – better grounded, but still prone to drifting beyond the data; good for inspiration, inconsistent for research decisions.
- TRC’s validated digital agents – survey-grounded, holdout-tested against humans, bias-aware with guardrails, and developed in academic collaboration. Designed to augment, not replace, human research.
Security, Privacy, and Compliance
TRC Insights is HITRUST CSF® certified and SOC 2 attested – third-party validations that place TRC in an elite group of organizations worldwide and demonstrate our commitment to protecting sensitive and regulated data, including in healthcare and financial services. TRC’s HITRUST certification incorporates HIPAA, NIST, ISO, and COBIT requirements.
Our synthetic-respondent work is built inside that same security regime:
- All modeling is done in a secure development environment.
- Client data and information are secured within TRC’s network.
- Models are trained independently, with nothing shared between them.
- Our personas chatbots are also HITRUST CSF® and SOC 2 complaint.
For full details on TRC’s certifications and controls, see our Data Security page.
Frequently Asked Questions
How can I trust the results generated by Digital Twins?
Trust starts with transparency. Our digital twins are built using data that our clients own and understand; they are not anonymous, third-party datasets. Every twin is grounded in research collected specifically for your organization, allowing results to be evaluated against known dataset and business context.
We also prioritize rigorous validation throughout the model development process and apply enterprise-grade security standards to protect client data. The result is an AI research approach that can withstand scrutiny – from research teams, analytics leaders, and executive stakeholders who need confidence in how insights were generated.
What is the difference between synthetic respondents and real survey respondents?
Real respondents are people answering a survey. Synthetic respondents are AI models trained on real human data that simulate how those people, or a segment, would answer. At TRC, synthetic respondents reduce uncertainty and refine research before fielding – they do not replace human data.
Can synthetic respondents replace human survey research?
No. TRC’s position, shared by its President and academic collaborators, is that human data remains the primary source of truth. Synthetic data helps explore and refine concepts so that the human study has the best chance of succeeding.
Are TRC’s Digital Twins accurate?
Accuracy is measured per project through holdout testing against your own human data, because there is no universal benchmark for synthetic-data accuracy. Twins are directionally accurate – strong on rankings and tradeoffs, weaker on exact scores, deep emotion, and price sensitivity.
Why can’t a Digital Persona be validated like a Digital Twin?
Because a persona represents a group rather than an individual, there is no single person’s “ground truth” to test its answers against. Personas are validated qualitatively (consistency with known segment behavior), not statistically, which is why they suit exploration and inspiration rather than decision-grade measurement.
Is synthetic data secure?
Yes. TRC is HITRUST CSF® certified and SOC 2 attested. All synthetic-respondent modeling happens in a secure environment, data stays within TRC’s network, and models are trained independently.
When should I use Digital Twins vs. Digital Personas?
Use Digital Twins when you need directional, decision-supporting answers to new quantitative questions (concept screening, message testing, narrowing lists). Use Digital Personas when you want to converse.
Explore AI-Powered Research with TRC
Talk to our team about whether Digital Twins or Digital Personas fit your next decision and where human research should lead.
"*" indicates required fields