in brief

  • Synthetic respondent validation should test the intended task on genuinely held-out evidence and compare it with a useful baseline.
  • Matching a person’s answer is different from predicting a purchase; leakage, calibration, and performance across categories matter.
  • relvo’s internal persona backtest improved one reason-prediction result but did not improve action prediction overall, so purchase forecasts are not shipped.
how should you validate synthetic respondents? tests and limits | relvo article cover

Validate synthetic respondents by defining the task, withholding the evidence needed to test it, comparing with a useful baseline, and examining where the result breaks. A fluent response or plausible persona description does not establish predictive value.

The first question is what the synthetic respondent is meant to do. Reproduce an interview answer? Describe a group’s concerns? Predict a new purchase? These are different tasks and need separate evaluations.

what should a held-out test contain?

Do not give the model a paraphrase of the answer and call the original question held out. Other posts from the same author, later summaries, and category notes can also reveal the target indirectly. The split should match the deployment claim.

A model that reconstructs an answer from related material has passed a different test from a model that predicts a future decision.

research note from relvo

what did relvo’s persona backtest find?

relvo’s internal product-definition record reports a backtest across 236 buying decisions in three categories. It built one persona per customer group, hid a real decision, rebuilt the persona without that decision, and tested the prediction.

reported prediction taskpersonamodel with market data onlybase
Why they acted38%19%107 decisions
What they did36%38%236 decisions

These are relvo-reported internal evaluation results, not an independently audited benchmark. The product-definition record does not provide the complete scoring rubric or uncertainty intervals. The results should therefore be read at the scope reported, rather than treated as a general accuracy claim.

why do the limitations change the interpretation?

The same record says the credit category went the other way on reason prediction: 23% for personas versus 33% for the market-data model, with a base of 30. It also notes that part of the gain elsewhere came from authors’ other posts remaining in their group’s evidence. That is a leakage concern for stronger predictive claims.

Persona probabilities were reported as worse calibrated than simple counting. Calibration concerns whether stated probabilities match outcomes over repeated cases. A convincing answer does not establish that match. Without a full evaluation artifact, readers cannot reconstruct that comparison from this article.

what the test supports

The reported evaluation gives relvo a reason to restrict use: group personas can help examine message and claim hypotheses, but relvo does not ship simulated purchase behavior or customer purchase forecasts. It does not prove that every synthetic respondent system will fail.

what should you ask a supplier to show?

supplier claimevidence to request
Predicts customer purchasesRelevant unseen outcomes, time order, and a baseline.
Matches human answersQuestion-level holdouts and scoring method.
Represents your audienceGrounding source, coverage, and excluded groups.
Provides reliable confidenceCalibration results for the intended task.
Works across categoriesResults outside the training venue and brand.

An FMCG concept test illustrates the distinction. A model may reproduce reasons people gave for liking a snack while still failing to predict which snack they buy at its normal price. That is a hypothetical example of two tasks, not a measured consumer-brand result.

how should simulated material appear in a report?

Label it as simulated. Keep generated answers separate from human interviews, surveys, and purchase records. State what it was used for, what validated that use, and which decision remains untested. Several personas are not automatically several independent respondents.

For related distinctions, read what customer digital twins can and cannot establish and why hypothetical purchase intent needs care.

Ask for the evaluation behind the intended decision before treating simulated agreement as customer evidence. Bring that question to relvo.