The Fidelity Bar for Synthetic Users Is Now Published
Simulation vendors now sell against a measurable accuracy standard, which turns every prompted persona sitting in your discovery folder into an unlabeled bet on a number nobody checked.
Why a better prompt cannot close the gap
A researcher asks a prompted persona what it would pay, and gets an answer that is articulate and internally consistent. That is the tell. The shortfall is structural, not a prompt-quality problem. Web-scale training text is mostly self-exposed attitudinal data, meaning what people say about themselves, with observed behavior only sprinkled in. Frontier labs then post-train that base toward expertise, sourcing professional programmers and scientists through vendors like Mercor and Scale. The output is a super-rational reasoner. Real users are not super-rational. Simile's stated objective is the inverse: reproduce human biases and mistakes, which it argues requires changing weights rather than writing better instructions. So "Improve our persona prompts" is not a viable initiative. A capability the base model was optimized against does not come back through the instruction field.
What the early buyers are doing with it
What teams tell themselves they are buying is a survey panel that never sleeps. Cheaper focus groups. That is not the interesting usage. Wealthfront had agents reason over multimodal input and traverse Figma mockups and live URLs. Shopify's SimGym searches a shopping trajectory for the intervention that lifts conversion. Both are pre-build intervention search, not post-ship measurement. The experiment moves upstream of the build gate instead of replacing the A/B test after it. Latent Space reports tens of millions of simulations run for Fortune 100 clients, with contracts quoted in the millions per customer, so enterprise competitors are already operating at volume. Person-models are domain-agnostic and reusable, so marginal cost per study falls with each study run. Procured as a standing panel, that compounds. Procured as one-off statements of work, it is a treadmill.
| Approach | Behavioral fidelity | Best PM use | Failure mode |
|---|---|---|---|
| Prompted frontier personas | 50-60% general, 20-30% niche | Hypothesis generation, survey wording, straw-man objections | Optimized for rationality; misses irrational real behavior |
| Vertical sim environments (SimGym) | Domain-tuned on platform trajectories | Intervention search on one flow, e.g. checkout conversion | Domain-locked; does not transfer off commerce |
| Behavioral foundation models (Simile) | 85% of human test-retest | Concept tests, surveys, agents walking mockups | Inherits noise from source studies; consent exposure |
The two caveats that set the price
Validity inheritance comes first. One Simile model was post-trained on tens of thousands of pre-registered randomized trials from the Open Science Framework, a literature that just spent five years in a replication crisis, where a p<0.05 threshold still implies a 5% false-positive floor by construction. The question for any vendor is which corpora they used and whether replication status was filtered. Attribute decay is quieter and more expensive. Stable traits like risk tolerance persist. Frequency behaviors such as how often someone visits a CVS drift. A vendor naming which attributes it treats as stable is also naming the refresh budget and the shelf life of every conclusion drawn from the panel. The founder's own maturity read puts simulation at roughly the GPT-3.5/GPT-4 stage: good enough to do real work, not good enough to be the decision.
Install a gate, not a replacement
The only honest scorecard is a blind back-test. Take three experiments already completed with known outcomes, have a vendor predict them cold alongside a control arm of raw prompted personas, and score both against your own test-retest reliability. That reliability is the ceiling, not vendor marketing accuracy. Then place the survivor before the build gate, where agents walk mockups and candidate flows before an engineering slot is allocated. The metric that proves it worked is concepts screened per build slot, plus at least one concept killed pre-build that would otherwise have burned a sprint.
Synthetic users are good enough to kill a bad idea and not good enough to pick a winner — put them before your build gate, never after it.
What to do
Tag every prompted persona, AI-generated survey respondent and LLM-authored persona doc in your discovery folder with the segment it claims to represent by end of next week, and mark every niche one hypothesis-only until real users confirm it.
Run a blind back-test this sprint on three completed experiments — vendor simulation plus a raw prompted-persona control arm — scored against your own test-retest reliability rather than vendor-claimed accuracy.
Inventory proprietary behavioral data (transactions, session logs, support transcripts) and get Legal's consent-scope read this quarter, before any simulation pilot touches first-party data.