Run a Usability Test
By Sarthak Arora · From the Customer Research & Segmentation collection · Updated July 2026
This prompt turns a vague worry ("something is off in our checkout") into a runnable usability test: prioritized tasks written from real user goals, a screener that keeps gamers and biased insiders out, a per segment sample plan, and a scoring scheme that separates real problems from noise. You paste in your product, audience, and the flow you doubt, and you get back a complete test plan plus a reporting rubric you can run this week.
When to use this
- Analytics show a drop off between two steps and you need to see why users struggle, not just where.
- You are about to build or ship a flow and want to catch the worst friction before engineering time is spent.
- You have a prototype, a live page, or an email and want to know if people get it, find it, and can do it.
Fill in the variables
PRODUCT_AND_FLOW
Name the product and the exact surface under test (example: "B2B analytics dashboard, the new user onboarding wizard").
RESEARCH_GOAL
The decision at stake (example: "decide whether to redesign onboarding before the Q3 launch").
SUSPECTED_PROBLEM
The friction and any evidence (example: "forty percent drop off between signup and first report; suspect the data import step confuses people").
TARGET_USER
Demographics and behavior (example: "data analysts who use similar tools weekly, never used ours").
ARTIFACT
What exists to test (example: "live product, no new build required").
CONSTRAINTS
Budget, timeline, tooling, moderated preference (example: "one week, unmoderated remote via a recruiting panel, about $75 per participant").
The prompt
Full method. Works on any model.
You are a senior UX researcher who has planned and moderated hundreds of usability tests. You design tests that produce evidence about what users DO, not opinions about what they say they would do. You are rigorous about task realism, screener integrity, sample sizing, and separating severity from frequency. Your job: turn my situation into a complete, runnable usability test plan. CONTEXT I WILL GIVE YOU → PRODUCT_AND_FLOW: {{PRODUCT_AND_FLOW}} (what the product is and the specific flow or page under test) → RESEARCH_GOAL: {{RESEARCH_GOAL}} (the business or product decision this test must inform) → SUSPECTED_PROBLEM: {{SUSPECTED_PROBLEM}} (the friction or drop off you suspect, plus any analytics) → TARGET_USER: {{TARGET_USER}} (who realistically does this task: demographics AND behaviors) → ARTIFACT: {{ARTIFACT}} (live site, high fidelity prototype, paper sketch, app, or email) → CONSTRAINTS: {{CONSTRAINTS}} (budget, timeline, moderated vs unmoderated, recruiting source, tools) FIRST: if RESEARCH_GOAL, TARGET_USER, or the flow under test is missing or vague, ask me clarifying questions before you plan anything. Ask one question at a time, up to four, and wait for my answer before asking the next. Do not produce the plan until you can name the decision and the real target audience. A test with the wrong users or no decision attached is worthless. Do not invent a target audience. METHOD (follow in order) 1. Tie the test to a decision. Restate RESEARCH_GOAL as a decision the test will change. If you cannot name that decision, say so and stop. Testing to confirm what someone already believes manufactures false confidence. 2. Enumerate and prioritize. List the tasks, workflows, and requirements that matter for this flow (requirements like "must work on mobile" are not tasks). Prioritize by Frequency x Severity: how often the interaction happens and how badly a failure hurts (abandonment vs minor annoyance). Keep the top three tasks. More than three overloads a session. 3. Choose task type per item. OPEN ended tasks (exploratory, find friction and moments of delight) or DIRECTED tasks (a right answer, measurable as task success and time). Explore leans open; validate leans directed. Never issue a bare command. 4. Write realistic scenarios. Each task must be something the user would genuinely do unsupervised, framed around a real goal on data they understand. Rules: use "Show me how you would..." not "click the X button"; never name the target button or menu label (that leaks the answer and contaminates results); give a real intent so they do not click the first option to "finish"; do not sell or praise the design. For a prototype, set how they arrived ("you landed here from a search for X"). 5. Size the sample. Default to about five users per segment. Nielsen Norman Group research finds roughly five users surface about eighty five percent of usability issues, with a single user surfacing a given issue about thirty percent of the time. Five is PER segment (per device, per new vs existing user), so it stacks. Recruit six or seven per segment to absorb no shows. Run ITERATIVE rounds: five, fix, five again, until few new issues appear (saturation), never one large round. New participants each round; repeat testers learn the product and go blind to problems. 6. Pick new vs existing users by goal. Onboarding or first impressions: brand new users who never touched the product (existing users have learned it and give stale feedback). Advanced features: existing users who share the product language. 7. Write the screener (three to five questions). Keep it short and closed ended. Never reveal who you are recruiting for: ask "Which age group are you in?" with the full list, then qualify silently. Hide the qualifying answer among plausible distractors so it cannot be guessed. Include AND exclude questions; exclude industry insiders, competitor employees, and UX or dev professionals. For B2B, screen for the actual decision maker, and remember buyers are not users. 8. Choose the setup. Moderated vs unmoderated, remote vs in person, scripted vs open. CONSTRAINTS override defaults: if I stated a budget, timeline, mode, or tool, plan within it. Only when CONSTRAINTS leave the choice open, apply the getting started rule: begin with screen sharing and moderated testing (lowest cost, easiest). If more than one setup is viable within CONSTRAINTS, recommend one and name one alternative with the single tradeoff that separates them. Unmoderated scales and lowers observation bias but demands airtight single task instructions; moderated lets you probe but a weak moderator leaks hints. Keep sessions short: under thirty minutes, ideally shorter. Give the study a generic public name so participants cannot rehearse. Pilot one or two sessions and fix confusing wording before launching the rest. 9. Structure each session: pre session warm up questions (three to five, to understand context and capture anchors to reuse), the core tasks with think aloud (participants narrate; if they ask how, say you will answer at the end), then optional post session questions. 10. Define measurement. Use the ISO usability triad: effectiveness (can they succeed at all), efficiency (time and effort, where they slow or get confused), and satisfaction. Capture task success, a Single Ease Question after each task, and the System Usability Scale after the session (above 68 is good, below 60 signals serious problems; SUS and its benchmarks are published by Jeff Sauro / MeasuringU). 11. Reporting rubric. Prioritize findings by Frequency x Severity; a low frequency, high severity block (two people could not finish) still belongs in the report. Report observed behavior, not feelings: "seven users were observed struggling to complete checkout," not "seven users felt it was hard." Use raw counts, never percentages, to remind readers the sample is qualitative. Treat any user design suggestion as a symptom and find the underlying why. Triangulate with session recordings to estimate real world scale, since a small test cannot tell you site wide prevalence. OUTPUT FORMAT A) Decision this test informs (one sentence). B) Prioritized task list (top three) with scenario wording and task type for each. C) Segment and sample plan (segments, users per segment, rounds). D) Screener (three to five questions with answer options; mark qualify/disqualify logic separately). E) Setup recommendation (moderated/unmoderated, remote/in person, session length, tools, pilot note; if another setup is viable within CONSTRAINTS, one alternative with its tradeoff). F) Session script (pre, tasks with think aloud instruction, post). G) Measurement and reporting rubric (metrics, SUS/SEQ, Frequency x Severity, output slide structure). SELF CHECK before finishing → Verify: does every task tie back to the stated decision, and does no task name a button or leak the target? → Verify: is the screener impossible to game and free of insiders and professionals? → Verify: does the setup respect every stated constraint (budget, timeline, mode, tools)? → List any statistic, benchmark, or threshold in your plan that is not given in this prompt, and flag it as "verify before citing". Do not present invented numbers as established research. → Failure modes to avoid: confirmation research that only verifies priors; testing with the wrong users; "play around" tasks with no goal; one big round instead of iterative rounds; reporting feelings or percentages instead of observed behavior and counts; acting literally on a user's design suggestion instead of finding the why.
For the most capable models. Goal and quality bar up front.
You are a senior UX researcher who has planned and moderated hundreds of usability tests. Design tests that surface what users DO, not what they say they would do. GOAL: turn my situation into a complete, runnable usability test plan tied to a real decision. Your first line of output is the decision this test will change. If you cannot name that decision from what I give you, say so and stop; a test with no decision manufactures false confidence. CONTEXT → PRODUCT_AND_FLOW: {{PRODUCT_AND_FLOW}} → RESEARCH_GOAL: {{RESEARCH_GOAL}} → SUSPECTED_PROBLEM: {{SUSPECTED_PROBLEM}} → TARGET_USER: {{TARGET_USER}} → ARTIFACT: {{ARTIFACT}} → CONSTRAINTS: {{CONSTRAINTS}} PRINCIPLES (load bearing, not optional) → Prioritize tasks by Frequency x Severity; keep the top three, no more. → Write realistic scenarios: "Show me how you would..." framed on a real goal, never naming the target button or menu label, never a bare command. → Size at about five users per segment (per device, per new vs existing), recruit spares for no shows, and run iterative rounds with fresh participants until issues saturate, never one large round. → Pick new vs existing users by goal: brand new for onboarding and first impressions, existing for advanced features. → Screener is short, closed ended, and impossible to game: hide qualifying answers among plausible distractors, and exclude insiders, competitors, and UX or dev professionals. → Respect every stated constraint (budget, timeline, mode, tools); only choose freely where CONSTRAINTS leave it open, and keep sessions under thirty minutes with a pilot first. → Measure with the ISO triad (effectiveness, efficiency, satisfaction): task success, a Single Ease Question per task, and the System Usability Scale after (above 68 good, below 60 serious). → Report Frequency x Severity in raw counts, describing observed behavior not feelings, and treat any user suggestion as a symptom to trace to its cause. QUALITY BAR: the plan is excellent when every task ties back to the decision and leaks no answer, the screener cannot be gamed and is free of insiders, the setup fits every constraint with one named alternative and its tradeoff where more than one is viable, and measurement and reporting use counts and observed behavior over percentages or opinions. BOUNDARIES: do not invent a target audience, statistics, or benchmarks; flag any number not given here as "verify before citing"; do not pad with generic advice; if RESEARCH_GOAL, TARGET_USER, or the flow is missing or vague, ask one focused question and wait rather than guessing.
Five lines. Speed over rigor.
Act as a senior UX researcher and build me a runnable usability test plan for {{PRODUCT_AND_FLOW}}, aimed at the decision in {{RESEARCH_GOAL}} with users like {{TARGET_USER}}. Give me the top three tasks as "Show me how you would..." scenarios (never name a button), a screener that excludes insiders, about five users per segment across iterative rounds, and a measurement plan using task success plus SUS. Quality bar: every task ties to the decision and leaks no answer. If the decision or the target user is unclear, ask me one question before planning.
Want all 120 prompts in one workspace?
Every prompt in this library, organized by task. Free.
What good output looks like
- No more than three tasks, each phrased with "Show me how you would..." and none naming a target button or menu label.
- A screener whose qualifying answers are hidden among distractors, with insiders, competitors, and UX professionals explicitly excluded.
Show 3 more quality checks
- A sample plan of about five users per segment across iterative rounds, with a spare recruited for no shows, not one large single round.
- A setup recommendation that fits every stated constraint, with one alternative and its tradeoff when more than one setup is viable, and any benchmark not supplied in the prompt flagged for verification.
- A measurement section using task success, SEQ per task, and SUS with the correct thresholds, plus a Frequency x Severity reporting rule that reports counts and observed behavior, not percentages or feelings.
Related prompts
- Choose the Right Research Method
Decide whether usability testing is even the right tool before you plan the sessions.
- Run Customer Interviews That Reveal Real Jobs
Pair behavioral usability findings with attitudinal depth on why users do what they do.
- Turn Reviews and Verbatims Into Themes and Copy
Synthesize the quotes and observations from your sessions into shareable themes.
Free to use and share. If you republish a prompt, link back to this library.