UX

AI Personas as Hypotheses: Making Contradictions Visible

How synthetic personas can produce usage hypotheses and test questions without replacing interviews, behavioral data, or real users.

Meet “Perfect Pete.” Twenty-eight years old, obsessed with efficiency, drinks only fair-trade coffee, and uses your app every day exactly the way you pictured it. The problem: Pete does not exist.

A persona like Pete can be completely coherent and still miss how users actually behave. Language models amplify this risk: they quickly produce a smooth narrative even when the underlying data has gaps.

Paper profile with an overlaid statement and an action pointing in the opposite direction
What people say and what they do can diverge. The explanation has to be tested.

The useful application therefore does not begin with “Simulate a realistic user for me.” It begins with a narrower question: What plausible explanations and open questions might account for an already observed tension?

A synthetic persona is not user research

LLMs can organize existing observations, vary scenarios and generate new questions for a workshop. Without real user data, they cannot measure attitudes, segment sizes, or behavioral distributions.

Research supports precisely this narrow role. In a DIS study of human–AI workflows for personas, collaboration worked better when people first grouped user data by relevant characteristics and the model then drafted personas from those groups.1 SimUser positions simulated feedback as a heuristic aid in prototyping, not as a replacement for usability testing.2

More recent work also calls for calibration. A methodological paper on LLM simulations distinguishes exploratory heuristics from confirmatory evidence.3 In a preprint on simulated preference tests, synthetic distributions diverged from human results in a substantial share of the tasks examined.4 This does not establish a general error rate for every persona. It is enough, however, to serve as a warning: more details, names, and contradictions do not automatically make an output more realistic.

A workable process

1. Prepare the observations

You can draw on anonymized interview statements, support reasons, search queries, analytics, or published research. Each observation receives a source ID. Personal or sensitive raw data belongs only in an approved environment.

Selection remains human work. The team decides which differences are relevant and which data gaps remain open. The GOV.UK Service Manual likewise bases personas on what has been learned about users and requires ongoing testing with real or likely users.

2. Frame tensions as hypotheses

Instead of prescribing two “hard contradictions” to the model, I weigh an observation against several possible explanations. The result is not a supposed biography but a hypothesis map. Here is a prompt I use for that:

You receive anonymized observations with source IDs.

Task:
1. Describe the observed tension without adding new facts.
2. Formulate two plausible alternative explanations.
3. Name the supporting source IDs for each explanation.
4. Explicitly mark uncertainty and missing data.
5. Derive an open research question and a suitable test task.

Rules:
- Do not invent quotes, figures, motives, or experiences.
- Do not estimate segment sizes.
- Do not assign design priority.

Observations:
- [SOURCE A]: [observed statement or action]
- [SOURCE B]: [observed context or counterevidence]

3. Read the output against the sources

The same question applies to every line: Is this in the data, is it a clearly labeled interpretation, or has the model filled a gap? Fabricated material is removed. Even apparently empathetic direct speech does not belong in the persona if it is not grounded in an interview.

Here is what an output might look like, a synthetic example, not a participant quote: Lisa, 50, a teacher. Observed tension: she worries about data theft, yet posts photos of her grandson every Sunday. Explanation 1: the pull to stay connected within the family outweighs the abstract worry about privacy. Explanation 2: the worry is really about strangers, not her close family circle. Open question: which explanation holds, and what feature would increase her sense of security without costing her that family connection?

4. Decide with real users

The synthetic output can prepare a workshop, a scenario, an edge case, or the next test. Priorities should be set only on the basis of observed behavior, user feedback or reliable existing data. Content Testing shows how to turn a hypothesis into a testable task and decision criterion.

Synthetic personas are hypothesis generators. They do not replace interviews, observations, or usability tests and must not be quoted as participant statements.

Stop looking for the perfect user

The perfect user does not exist. Look for the cracks and fault lines, the small contradictions real people trip over. But be careful: if an AI persona irritates you, that is not a seal of realism. It is a prompt to investigate, nothing more.

An irritating output may reveal an overlooked context of use. It may just as easily be a stereotype or a prompt artifact. Only tracing it back to sources and real usage can show whether the hypothesis holds.

This boundary does not make AI personas worthless. It defines their value more precisely: the model expands the search space. You remain responsible for the data, validation, and decision.

Sources & References

🌐