ai-agentssynthetic-usersproduct-researchux-researchsimulationevaluation

Synthetic Users Are Simulators, Not Customers: A Guide to Agentic Product Research

By John Babich9/22/20267 min read
Beginner to Intermediate
Synthetic Users Are Simulators, Not Customers: A Guide to Agentic Product Research

Synthetic Users Are Simulators, Not Customers: A Guide to Agentic Product Research

The pitch is irresistible.

Why spend weeks recruiting research participants when one hundred AI personas can test your product before lunch? Give them demographics, goals, and frustrations. Let them click through the prototype. Ask what they loved. Export a chart.

The chart will look authoritative. That is the dangerous part.

Synthetic-user agents can be extremely useful. They can also create industrial quantities of plausible fiction. The difference depends on what question you ask them to answer.

TL;DR

Use synthetic-user agents as simulators for coverage, hypothesis generation, journey rehearsal, and interface testing. Do not use them as a substitute population for measuring demand, preference, willingness to pay, or emotional response. Calibrate synthetic behavior against real users, preserve uncertainty, and treat every simulated insight as a candidate for validation.

What is a synthetic-user agent?

A synthetic-user agent is a model configured to interact from a particular context: a goal, level of expertise, device, constraints, history, or behavioral profile. Unlike a static persona document, the agent can take actions, encounter state, revise its plan, and explain where it became confused.

For example, a team might simulate:

  • a first-time administrator configuring SSO
  • a rushed shopper comparing three complex products
  • a screen-reader user navigating a checkout flow
  • a finance lead reviewing an unfamiliar dashboard
  • a customer attempting to cancel without contacting support

The agent can traverse the product, record decisions, and produce a structured trace. That makes synthetic users more useful than invented quotes pasted onto a persona slide.

The sweet spot: breadth before evidence

Synthetic users are good at exploring possibility space.

They can run the same flow with different goals, starting states, vocabulary, and interruptions. They can expose missing branches, inconsistent labels, brittle instructions, and assumptions that only work for an expert.

Useful applications include:

  • generating edge cases for a research plan
  • rehearsing task flows before human testing
  • checking whether interface copy supports multiple mental models
  • exploring accessibility and localization scenarios
  • producing candidate objections for sales or onboarding
  • stress-testing support documentation
  • identifying where additional instrumentation is needed

This is fast, inexpensive pre-work. It improves the questions you bring to real people.

The category error: simulation is not sampling

A language model is not a randomly selected customer. It is a learned generator conditioned on its training data and your prompt.

Giving it a persona does not create a statistically representative member of that group. Running the prompt one thousand times does not produce a survey with a sample size of one thousand. The outputs may be diverse, but they share model priors, blind spots, and prompting artifacts.

Research on synthetic social agents warns against treating fluent predictions as calibrated population evidence. A useful framing is that these systems are powerful pattern matchers operating under explicit scope conditions, not drop-in replacements for probabilistic inference.

That means synthetic users should not be your primary evidence for:

  • market size
  • feature demand
  • willingness to pay
  • brand perception
  • emotional impact
  • prevalence of a behavior
  • differences between demographic groups

Those claims require real-world data.

Ask behavioral questions, not theatrical ones

"How do you feel about this landing page?" invites a convincing monologue.

"Can you find the annual price, and what did you inspect before deciding?" creates an observable task.

Design simulations around actions:

  1. Give the agent a goal and only the knowledge a user would have.
  2. Let it interact with the actual interface or a faithful prototype.
  3. Record clicks, tool calls, errors, backtracks, and completion state.
  4. Score the trace against explicit task criteria.
  5. Collect explanations as supporting material, not ground truth.

Behavioral traces are still simulated, but they are more diagnostic than free-form opinions.

Build profiles from evidence

Do not invent personas because they sound like people your product might have.

Ground profiles in existing research:

  • interview themes
  • support-ticket clusters
  • product analytics
  • search queries
  • accessibility reviews
  • CRM segments
  • known workflow differences

Strip personal information and encode the minimum characteristics needed for the task. A useful profile describes relevant constraints, not a fictional biography.

For a billing workflow, company size, approval authority, accounting cadence, and familiarity with the product may matter. Favorite coffee and a stock-photo name probably do not.

Calibrate against a human baseline

The safest way to use synthetic users is to learn where they agree and disagree with actual users.

Run a small human study, then replay the same tasks with synthetic agents. Compare:

  • task completion
  • time or step count
  • navigation path
  • error categories
  • points of confusion
  • severity ranking

Do not ask only whether the agent found the same problems. Check false positives and false negatives. A simulator that reliably finds form-label issues but misses trust concerns can still be valuable, provided you know its boundary.

Recent work such as PerceptUI explores how agent-based UI evaluation can align more closely with human judgments. The important product lesson is not that calibration is solved; it is that calibration must be measured.

Use diversity carefully

You can vary models, prompts, profiles, temperatures, and tool strategies to avoid one narrow synthetic voice. That is useful for generating a wider set of hypotheses.

But diversity of output is not representativeness.

A panel of five models may disagree in interesting ways while all failing to reflect your actual customer base. Treat model disagreement as a signal of uncertainty and an invitation to test, not as an election result.

A three-lane research model

A practical product-research program separates evidence into three lanes:

Lane Best use What it can claim
Synthetic simulation Coverage, rehearsal, hypothesis generation "This failure mode is plausible"
Behavioral data Funnels, usage, retention, experiments "This happened in the product"
Human research Motivation, context, interpretation, unmet needs "This is how participants experienced it"

The lanes should inform one another. Synthetic runs suggest what to instrument. Analytics reveal where to recruit. Interviews explain behavior. New evidence updates the simulation profiles.

Problems begin when one lane impersonates another.

Where synthetic users become genuinely powerful

The most interesting use is not replacing interviews. It is building a continuously updated product simulator.

Imagine every proposed onboarding change running through a regression suite of evidence-grounded user profiles. The system flags journeys that now require more steps, instructions that depend on expert language, or states that strand a user without recovery. Human researchers review the highest-risk changes and decide what needs real testing.

That turns research knowledge into reusable infrastructure while preserving human validation.

It also connects naturally to Simulation-First Testing for Agents: scenarios, fixtures, traces, and pass/fail gates are useful whether the actor is your agent or a simulated customer using your product.

A practical operating policy

Before a synthetic-user study, write down:

  • the decision the study supports
  • which profiles are evidence-grounded
  • what the simulator has been calibrated to detect
  • which claims are prohibited
  • what human or behavioral validation follows

In the final report, label synthetic findings clearly. Avoid fake respondent counts and precise percentages that imply population statistics. Report scenarios, observed traces, recurring failure modes, and confidence based on calibration.

The language should be: "Seven of ten simulated journeys encountered this issue; verify with novice users," not "70% of customers will struggle."

Summary

Synthetic users are neither magic focus groups nor useless toys.

They are simulators. A simulator helps you explore, rehearse, and find failure modes before the expensive or consequential test. It does not become the world merely because its dashboard has percentages.

Use agentic product research to arrive at better human research faster. Let real customers keep the final vote.

Related Tools

Useful tools for this topic

If you want to turn this article into a concrete next step, start with one of these.

Subscribe to AgentForge Hub

Get weekly insights, tutorials, and the latest AI agent developments delivered to your inbox.

No spam, ever. Unsubscribe at any time.

Loading conversations...