How to evaluate synthetic research
A practical worksheet for comparing methods, evidence and the study you actually need.
By InstaSights ·
Start with the decision, then compare the evidence
A provider may generate convincing answers without having demonstrated that those answers fit your audience or question. Before comparing prices, write down the decision, the population and the result that would change your next step.
This worksheet is for evaluating research methods, including our own. InstaSights sells consumer studies using simulated responses. It is not an independent ranking of vendors, and no provider earns a recommendation simply by completing it.
Download the evaluation worksheet. There is no form to fill out.
Ask what the system actually does
Ask whether you are buying a model-generated survey, a model grounded in your own customer data, augmentation of existing survey responses, or AI analysis of answers collected from people. Those are different research workflows.
For every proposed study, request the supported countries, languages, audiences, question formats and stimulus types. If a tool can read a product description, do not assume it can evaluate a photograph, video, physical product or live website.
Our consumer research guide explains these distinctions with a naming example. The InstaSights product page shows the supported study workflow.
Put the evidence beside the claim
Use this table in a provider discussion. Keep an evidence link or file alongside each answer so another person can inspect it.
| Ask the provider | Record | What remains unresolved without it |
|---|---|---|
| What audience and questions were validated? | Population, language, topic, wording and response formats | Whether the comparison applies to your study |
| Where did the human reference come from? | Source, fieldwork date, recruitment and sample sizes | What the model was compared against |
| Were evaluation answers held out? | Training and tuning exclusions; anything unknown | Whether the test was independent of model development |
| What does the accuracy measure mean? | Metric definition, denominator, per-question and subgroup errors | Whether one attractive percentage hides important misses |
| What happens when the study is repeated? | Model version, run settings and variation between runs | How stable a close ranking is |
| Which studies failed? | Failure cases and unsupported applications | Where you should choose another method |
These are evaluation questions, not a universal scoring formula. Set your acceptance criteria before seeing the result, based on the consequences of the decision.
AAPOR's survey-research guidance emphasizes distinguishing generated answers from human observations and evaluating relationships and subgroup behavior, as well as aggregate results. That supports inspecting several kinds of evidence rather than treating a headline score as sufficient. AAPOR report
Read a benchmark without turning it into a guarantee
Our methodology page presents one provider-reported retail-returns comparison: 76% in the human survey and 88% in the simulation agreed on the importance of free returns. The difference is 12 percentage points.
The two results put a majority on the same side, but they do not give the same estimate. Whether that error changes a decision depends on the decision. A comparison about returns does not validate a new product name, a different population or a forecast of sales.
For each benchmark, write two sentences: what it demonstrates, and what it leaves untested. If the evidence does not answer an important question, record “not established” rather than filling the gap with an assumption.
For B2B, check the professional audience
A consumer demographic profile does not establish professional buying experience. For synthetic B2B research, ask how the system represents job function, seniority, industry, company size, purchase authority and the people involved in the buying decision.
Then ask for evidence from comparable professional audiences and buying situations. Knowing the language of enterprise software is different from having measured the preferences of its buyers.
InstaSights currently offers consumer studies. This checklist helps evaluate a B2B provider; it is not an offer of a B2B study through InstaSights.
Check costs, data handling and outputs
Compare the full study scope: question limits, generated-response count, preparation allowance, reruns, exports, subscriptions and any additional analysis. Do not compare a short simulation with a recruited interview project as if they delivered the same thing.
If you plan to supply customer data or a confidential brief, request the provider's actual terms for retention, access, training use and deletion. A generated output does not, by itself, establish privacy protection for the information you upload.
Ask to inspect the questionnaire, response distributions and export format before buying. Our example study shows those screens and response previews; pricing explains the InstaSights offer.
Make the result of the evaluation explicit
Your evaluation should end with one of three decisions: suitable for this defined exploratory task; suitable only after a relevant validation exercise; or a different method is needed.
Record the unanswered questions and the next human-research or market-testing step. You can use the research brief template to keep the decision, audience and follow-up plan together.