What synthetic users can and can't tell you
AI-generated research participants are fast and cheap. What they can't do is tell you how real people actually behave, and that gap is exactly where the risk lives.
The recruiting round for a usability study takes two weeks if nothing goes wrong. A screener, a recruiting platform, a scheduler, a consent form, and then the calendar fills in slowly as participants confirm or no-show or reschedule once. By the time you're sitting with your first participant, you've spent most of the sprint you had, and the design has moved twice since you wrote the discussion guide.
Synthetic users look like the answer to that problem. You describe your user type to an LLM, give it your prototype or discussion questions, and within an hour you have ten "participants" who have answered your questions, navigated your flows, and identified friction points. No recruiting. No scheduling. No no-shows.
What you've done is ask the model to roleplay a user type based on its training data and your description. That is a useful thing to do. It is also a different thing than user research, and the gap between them is where programs get into trouble.
What the training data carries
The best synthetic users are built on models trained on writing from people who describe their experiences publicly: tech workers, product people, the users who file detailed feedback and write long posts about their workflows. That population isn't your users. It's probably not close to your users. The question isn't whether the synthetic response is well-reasoned and internally consistent, it's whether the reasoning maps to how your actual users think, which requires your actual users.
This is a version of an old problem. Surveys and focus groups have always had this character: people describe what they think they do rather than what they do. Synthetic users have the same shape but with a new wrinkle. The gap isn't between what the participant says and what they do; it's between what the model predicts about a population and what that population is. The model has never met your users and cannot ask what it doesn't know to ask.
The failure mode isn't bad feedback. It's confident feedback that maps cleanly to your mental model of your users, because you described your mental model of your users when you wrote the persona prompt. You get back a mirror, slightly polished by the model's ability to be coherent.
Where synthetic users help
The places synthetic users are genuinely useful share one quality: the bottleneck is speed and the question isn't whether the design works for humans.
Testing a discussion guide is one of them. The mistakes in a discussion guide (a question that leads, a task that's ambiguous, a flow that skips a step the researcher forgot to model) usually only surface when a real participant is confused by them. Running the guide against a synthetic participant first is a fast way to find the structural problems before you've spent your recruitment budget.
Pressure-testing a concept before it's worth recruiting for is another. If you have five directions and you want to narrow to two or three before you spend two weeks recruiting, synthetic users can tell you which concepts are underspecified, which questions they can't answer because the concept doesn't have an answer yet, and where the probe questions need to go deeper. You're not validating the concept; you're validating whether the concept is developed enough to validate.
Early work on hard-to-recruit edge cases is a third. If you're building for a population that's difficult to find (a rare professional role, a specific disability, an unusual workflow), synthetic users can help you develop hypotheses that a small number of real sessions can then confirm or reject. They're a way to make the most of limited access to a hard-to-reach group, not a substitute for reaching them.
What only real participants can do
The nonverbal signal is the first thing synthetic users can't produce. The pause before the answer, the sigh, the moment a participant reaches for the wrong button and then self-corrects, the workaround they invent without realizing it, the thing they don't mention because it didn't occur to them to mention it. These are not in the transcript. They live in the session, and they're often where the most important findings are.
The empathy function is the second. Design teams that have sat with real users in real difficulty build a relationship to those users that changes how they make decisions. It's harder to argue against fixing the accessibility issue when you've watched someone struggle with it. It's harder to deprioritize the onboarding confusion when you can picture the specific person who was confused. Synthetic users produce findings; real users produce faces, and the faces are what move design teams to act.
How to use both
The teams that use synthetic users well treat them as the front of the research pipeline. Synthetic users in the first 20% of a cycle, for concept-vetting, guide pressure-testing, and sharpening the questions before the real sessions. Real participants for the validation work that drives the decision.
The less obvious use is ongoing. Most product teams ship multiple times a week, often without research coverage on every change. Experimentation, A/B tests, incremental releases: the shipping cadence has outrun what a research team can realistically cover with recruited participants. An automated synthetic user test running on the same schedule as the engineering team's deploys gives the team a standing pulse on the product between formal research cycles. It won't tell you whether a new feature is right; it will tell you whether something that used to work has quietly broken, or whether a change has introduced friction that nobody in the release process noticed.
For this to be useful, the synthetic users need to be narrow. One synthetic user per persona the team has already identified, scoped to the product area that persona uses. A synthetic user representing a power user of the analytics dashboard shouldn't be roaming the onboarding flow; the signal gets diluted and the output gets hard to interpret. When the scope is tight, a failed weekly test points somewhere specific, which is the only kind of result worth acting on.
The risk isn't that a design team runs one synthetic session and ships a broken product. It's that the speed and confidence of synthetic sessions make it easy to skip the recruiting step over and over, until the team's understanding of their users is a layer of AI inference over a layer of their own assumptions, with no actual users anywhere in the picture.
Synthetic users are a research accelerant. The work is making sure they're accelerating toward something real.
How insightful was this article?
Let's work together.
Open to select projects and collaborations — design systems, accessibility, and AI-native product work.