How designers run their own user research with AI

Chris Hlavaty
Chris Hlavaty
Co-founder, Sera
Updated 6 min read
A translucent wireframe surface marked by many small glowing points and ripples

A designer can now run their own user research. Not a hallway-test approximation of it — the real thing: paste a Figma prototype link, launch a study in about ten minutes, run 20–50 sessions in parallel, and read quantified findings you can trust by the end of the day.

That sentence would have been absurd three years ago. AI made it true, and it changes where evidence sits in the design process: inside the crit cycle, not weeks behind it.

This article is about what that unlocks — how fast the loop actually runs, why the results are more trustworthy than the old way, and which design decisions to point it at.

What does AI user research look like for a designer?

Paste your Figma prototype link. Or a live URL. Or describe the design question in a paragraph.

From that, the AI authors the full study — goals, screener, discussion guide, tasks — in about two minutes. It arrives as an editable document, and you treat it like one: cut a task that misses the point of your design, sharpen a probe, tighten the screener. Sera is built around exactly this flow, and the total launch effort is five to fifteen minutes.

Then it runs without you. AI-moderated voice and video sessions with screen capture, all in parallel, participants thinking aloud while they work through your flows. No recruiting spreadsheet, no scheduling across time zones, no moderating eight calls yourself.

Results typically land in under a day. Ship the prototype to a study in the morning; read the answer before you sign off.

Why is that a big deal?

Because the old process was the reason designers didn't test.

One round of usability testing used to mean writing a screener, recruiting, scheduling, moderating every session personally, and synthesizing a pile of notes — weeks of logistics for a design due this sprint. Most designers never had a researcher to hand the question to anyway: at well-staffed companies the ratio is one researcher to many designers, and at most startups it's zero.

So design decisions ran on the next best thing — crit opinion, a hallway test, five coworkers who already knew what the interface was for. Real usability data arrived after launch, as support tickets.

AI collapsed the cost on every front at once: how fast a study launches, how many people it can talk to, and how quickly sessions become findings. For the first time, the research cycle fits inside the design cycle instead of trailing it.

Can you trust the results?

This is the right question, and the answer is stronger than "yes, roughly." The output is better than what most self-run design testing produced — for three reasons.

Sound methodology is built in. The AI drafts the study the way a researcher would: screener logic that recruits the right segment, tasks stated as goals rather than routes, probes phrased to avoid leading the witness. You don't need research training to launch something a researcher would sign off on — and you're no longer the moderator nudging participants toward the path you designed.

Scale turns observations into numbers. When sessions run in parallel, forty take roughly the same elapsed time as four — so 20–50 in a day is practical, not heroic. At that scale, a usability finding stops being "a couple of people seemed confused" and becomes a proportion: how many of forty participants missed the primary action, stalled at the same step, or read the screen the same wrong way.

Every claim carries its evidence. Colleagues argue with your interpretation; almost nobody argues with the recording. The synthesis cites the timestamped moments behind each theme, so every finding in the readout is one click from a participant hitting the problem on video — which means the clips for your next crit are already cut.

A design crit with no user evidence is a taste contest, and the most senior taste in the room wins.

Two minutes of cited clips at the top of a crit changes what the room argues about: from whose instinct is right to what the design should do about what everyone just watched.

What design decisions can you point it at?

Once a study costs ten minutes to launch, every recurring decision moment in a designer's week becomes testable. The big five:

Decision momentThe questionWhat you paste inN (parallel, same elapsed time)Fits in
Pre-build checkCan people complete the core flows before engineering starts?Figma prototype link20–40 per roundOvernight, before handoff
Two directions on the tableWhich direction lands with users, and why in their words?Both prototypes20–40Before the next crit
Redesign validationDoes the new direction beat the pattern users already know?Prototype or staging URL20–40Inside one sprint
First-run comprehensionDo new users understand what this screen is and what to do first?Prototype or live URL20–40Under a day
Post-ship diagnosisWhy is this flow underperforming?Live URL, plus the question in a paragraph20–40Before the retro, not after

The N column would have looked absurd recently; those numbers used to require a research operations team. Now N is set by the confidence the decision needs, not by logistics.

And notice what the second row settles: direction debates. A preference question needs enough people for the split to mean something, which is exactly why crits used to deadlock on taste — nobody could afford thirty sessions to break the tie. Now the tie costs a day, and it comes back with reasons attached in users' own words.

How does the analysis work?

The hidden reason design testing stayed small was never the sessions. It was the synthesis — a designer who somehow ran thirty sessions then faced thirty hours of replay review, so nobody did.

That step is now automated, multi-pass, and exhaustive. Themes are extracted across every transcript. Each theme is aligned with the specific sessions that support it. Insights and recommendations are drawn from the pattern, and the whole thing rolls up into an executive summary.

It's ready about fifteen minutes after the last session ends.

So what lands in your hands is not raw footage. It's a readout where every theme carries its count and its cited moments — the artifact that survives a skeptical stakeholder review, generated while you were doing your actual job.

What changes when testing fits inside the design cycle?

Your habits, mostly.

You test the prototype before the build instead of finding out from support tickets. When two directions split the crit, you run both overnight instead of deadlocking for two weeks. When a shipped flow underperforms, you watch the users behind the metric before the retro, not after.

The cadence that becomes possible: test between crits. A direction survives crit, goes into a study that same day, and the next crit opens with evidence instead of opinion. Rounds compound — three fast rounds on an improving design find more than one big study on a frozen one, and now three rounds fit in a week.

The scarce resource is no longer recruiting, scheduling, or synthesis time. It's the quality of your design question — which is the part of the job you're good at.

Because that part stays yours. Deciding which flow deserves a test, editing the tasks so they aim at your actual design risk, and judging what the findings mean for the work: exactly as valuable as ever. Everything around that judgment — authoring, moderating, analyzing, citing — is now handled.

This is the loop Sera runs end to end. Paste the prototype link, spend ten minutes on the draft, and read quantified, cited findings tonight.

The fastest way to believe it is to launch one study. Pick the design currently stuck in debate — by this time tomorrow, it won't be.

Frequently asked questions

Can AI run usability tests on a Figma prototype?

Yes — the full loop. Paste the prototype link and the AI authors the study in about two minutes: goals, screener, and tasks aimed at the flows the prototype exercises. AI-moderated voice and video sessions then run in parallel with screen capture and think-aloud, and the synthesis cites the timestamped moments behind every finding.

How fast can a designer launch a user research study?

About ten minutes of your time. Paste a Figma prototype link or live URL — or describe the design question in a paragraph — and the AI drafts the full study in roughly two minutes. You edit it like a doc and launch. Total effort typically runs five to fifteen minutes, with results back in under a day.

How many usability sessions can AI-moderated testing run in a day?

20–50 is practical, because sessions run in parallel rather than one at a time on your calendar. Forty sessions take roughly the same elapsed time as four. At that scale, a usability finding arrives as a quantified proportion — how many participants stalled at the same step — instead of a handful of anecdotes.

Can you trust AI-moderated usability testing?

Yes, for three reasons. Sound methodology is built in — screener logic, goal-stated tasks, unbiased phrasing — so the study is well-constructed without research training. Samples of 20–50 turn themes into measured proportions rather than hallway impressions. And every claim in the synthesis cites timestamped session moments, so the evidence is one click away.

Do designers need a researcher to run user research?

For everyday design questions, no. Testing prototypes, comparing directions, and checking comprehension against real users is now fully designer-operable, at sample sizes that once required a research operations team. A dedicated researcher still earns their cost on generative work and the largest strategic bets — but design decisions no longer wait on headcount.

How does AI analyze usability test sessions?

With automated multi-pass analysis. Themes are extracted across every transcript, each theme is aligned with the specific sessions that support it, insights and recommendations are drawn from the pattern, and everything rolls into an executive summary — ready about fifteen minutes after the last session ends. You read a cited readout, not thirty recordings.

Hear an AI-moderated interview
on your own product.

Paste a URL. Sera drafts the study, recruits participants, and runs the interviews — usually within 24 hours.

Your first 7 interviews are on us — no credit card required.