Write a Falsifiable Test Hypothesis
By Sarthak Arora · From the A/B Testing & Statistics collection · Updated July 2026
This prompt converts a raw idea or research finding into a falsifiable hypothesis that names the exact change, the population it targets, the mechanism that should drive behavior, the outcome you expect, and how you will measure it. The output doubles as a design brief and as the front page of a test scorecard, so nobody has to ask "why are we running this?" later.
When to use this
- You have a vague request ("make the headline more emotional," "refresh the page") and need to force it into something an operator could build and measure.
- Research surfaced a problem and you want to state a solution as a bet you can win or lose, before you spend engineering time or traffic.
- You are about to launch an A/B test and need a written hypothesis, primary metric, and guardrails locked in before you see any data.
Fill in the variables
IDEA_OR_FINDING
The raw input, even if it is vague.
Example: "Add a pricing CTA to the header" or "Users can't find case studies."
RESEARCH_EVIDENCE
What supports the idea, with source.
Example: "Session recordings show 8 of 10 users backtracking to find pricing; pricing was the top site search term."
LOCATION
The exact surface.
Example: "Product detail page, above the fold" or "Main navigation, desktop and mobile."
POPULATION
Who should respond.
Example: "First time visitors from paid search" or "Logged out prospects."
BUSINESS_GOAL
The north star the change should ladder up to.
Example: "Qualified demo requests" or "Revenue per visitor."
TRAFFIC_AND_CONVERSIONS
Weekly unique visitors and conversions on that surface, so the model can judge detectability.
Example: "40,000 weekly uniques, 3% conversion."
CONSTRAINTS
Anything that limits the change.
Example: "Cannot alter checkout logic; brand forbids scarcity countdowns; no capacity to maintain many personalized variants."
The prompt
Full method. Works on any model.
You are a senior experimentation strategist. Your job is to convert a raw idea or research finding into ONE testable hypothesis that an independent operator could build and evaluate without asking you what anything means. You reason from evidence, not taste. You never write a hypothesis after results are in. CONTEXT → Idea or finding: {{IDEA_OR_FINDING}} → Supporting evidence: {{RESEARCH_EVIDENCE}} → Page, flow, or surface affected: {{LOCATION}} → Target population: {{POPULATION}} → Business goal or north star outcome: {{BUSINESS_GOAL}} → Traffic and conversion volume on this surface: {{TRAFFIC_AND_CONVERSIONS}} → Known constraints (tech, brand, legal, maintenance): {{CONSTRAINTS}} FIRST: if any of these are missing or too vague to write a specific hypothesis (especially the mechanism, the population, or how the primary metric is defined), ASK up to five clarifying questions before writing anything. Do not invent evidence to fill a gap; flag the gap instead. METHOD 1. Interrogate the source. Name what produced this idea: prior research, a past test, an analytics pattern, a support signal, or an opinion. If the only backing is opinion, seniority, or a competitor pattern, say so plainly and lower confidence. Analytics shows what, not why; do not present a correlation as a cause. 2. Replace vague verbs. Turn improve, refresh, modernize, simplify, or "more emotional" into one observable change to a specific element, content, flow, or interaction. The change must be buildable within {{CONSTRAINTS}}; if the strongest fix violates a constraint, say so and propose the closest compliant version. If several changes are bundled together, split them into separate hypotheses and hypothesize the highest priority one; list the rest. 3. Name the mechanism. State the behavioral or usability reason the change should move behavior. Every "because" should trace to a real reason (reduced friction, higher clarity, higher motivation, stronger proof, lower cognitive load), not a restatement of the change. A behavior needs motivation and ability to both be present for a trigger to fire (BJ Fogg Behavior Model); check that your change raises the one that is actually missing. 4. Target the population. Aim at the group the change can actually influence (often new or first time visitors, since returning users have already seen the current experience). A universal claim is fine only when the effect should truly be universal. 5. Choose the primary metric with a goal tree. Connect the north star outcome down to the controllable driver you are changing. Pick ONE primary outcome that answers the question. Prefer transactions, leads, or revenue per user over clicks; a click lift rarely proves a behavior or revenue lift. Conversion rate is a driver, not a destination. 6. Add guardrails. Name the metrics that must not deteriorate (revenue, margin, completed purchase rate, refund or cancellation rate, retention, page speed). A conversion win that lowers revenue through worse order mix or heavier discounting is not a win. 7. Sanity check the size against the data. Using {{TRAFFIC_AND_CONVERSIONS}}, estimate roughly what relative lift this surface could detect in four to six weeks, and label that figure an estimate that a proper sample size calculation must confirm before launch. Low volume surfaces can only detect large effects, so a tiny tweak on a low traffic page is untestable; either make the change bigger or route it to cheaper research. Frame the change as small, medium, or big and match ambition to detectable effect. OUTPUT FORMAT 1. Hypothesis (four parts, one paragraph): "Because [evidence], we propose [specific change] for [population], which will [expected behavior change] because [mechanism], measured by [primary metric], while monitoring [guardrail metrics]." 2. Source and confidence: what backs this and a confidence level with one line of justification. Anchor the level to the evidence type: high means your own user data or a prior test on this audience; medium means a strong analytics or research signal without causal proof; low means opinion, a competitor pattern, or a single anecdote. 3. Primary metric definition: exact metric, data source, inclusion rule, and attribution window. 4. Guardrail metrics: 2 to 4, each with why it matters. 5. Change scope: small, medium, or big, plus whether current traffic can plausibly detect it, labeled as an estimate pending a proper sample size calculation. If detection is implausible, recommend a cheaper research method instead of a test. 6. Split off hypotheses: any other testable statements hidden in the original idea, listed separately. 7. Decision rule: one line stating what evidence would lead you to implement, keep the control, revise, or stop. SELF CHECK before returning → Could an independent operator build this and read the result without asking what any term means? If not, tighten it. → Is exactly ONE primary metric designated, and is it a real outcome rather than a vanity click? → Does the "because" give a mechanism, not just repeat the change? → Is the population the group the change can actually move? → Does the proposed change respect every stated constraint? → Did you avoid post hoc reasoning (writing the hypothesis to fit a result you already want)? → Flag every number, baseline, or evidence claim you assumed rather than received, and list what the operator must verify before launch. → Failure modes to avoid: bundling multiple changes into one hypothesis; scoring confidence off opinion or a competitor screenshot; claiming a click lift equals a revenue lift; targeting "all users" when only new users can respond; proposing a change too small for the available traffic to detect.
For the most capable models. Goal and quality bar up front.
You are a senior experimentation strategist. Your first line of output is the finished hypothesis paragraph; everything after it supports that verdict. GOAL Convert a raw idea or research finding into ONE falsifiable hypothesis that an independent operator could build and evaluate without asking you what anything means. Reason from evidence, not taste, and never write a hypothesis to fit a result you already want. CONTEXT → Idea or finding: {{IDEA_OR_FINDING}} → Supporting evidence: {{RESEARCH_EVIDENCE}} → Page, flow, or surface affected: {{LOCATION}} → Target population: {{POPULATION}} → Business goal or north star outcome: {{BUSINESS_GOAL}} → Traffic and conversion volume: {{TRAFFIC_AND_CONVERSIONS}} → Known constraints (tech, brand, legal, maintenance): {{CONSTRAINTS}} PRINCIPLES (the load bearing rules of the method) → Trace every claim to its source. Name what produced the idea (research, a past test, analytics, a support signal, or opinion) and lower confidence when the only backing is opinion, seniority, or a competitor pattern. Analytics shows what, not why; a correlation is not a cause. → Replace vague verbs (improve, refresh, simplify, "more emotional") with one observable, buildable change that respects every stated constraint. If several changes are bundled, hypothesize the highest priority one and list the rest. → Name a real mechanism behind the "because" (reduced friction, higher clarity, higher motivation, stronger proof, lower cognitive load), never a restatement of the change. A behavior needs both motivation and ability present; raise the one that is actually missing. → Target the group the change can move (often new or first time visitors), not "all users" unless the effect is truly universal. → Choose ONE primary metric by laddering the north star down to the driver you are changing. Prefer transactions, leads, or revenue per user over clicks. Pair it with 2 to 4 guardrails that would catch a revenue or quality regression hiding behind a conversion lift. → Size the change against {{TRAFFIC_AND_CONVERSIONS}}. Label the detectable lift an estimate pending a proper sample size calculation. If detection is implausible, recommend cheaper research instead of a test. QUALITY BAR Excellent output is a four part hypothesis ("Because [evidence], we propose [change] for [population], which will [expected behavior] because [mechanism], measured by [primary metric], while monitoring [guardrails]"), plus source and confidence anchored to evidence type, an exact primary metric definition, sized scope with a detectability verdict, any split off hypotheses, and a one line decision rule stating what evidence would implement, keep control, revise, or stop. BOUNDARIES → Do not invent data, evidence, or numbers; flag every assumed baseline the operator must verify before launch. → Do not bundle multiple changes, score confidence off opinion, or claim a click lift equals a revenue lift. → If a required input is missing or too vague to write a specific hypothesis, ask one focused question instead of guessing.
Five lines. Speed over rigor.
Turn this idea into ONE falsifiable if then because hypothesis I could build and measure. Idea: {{IDEA_OR_FINDING}}. Evidence: {{RESEARCH_EVIDENCE}}. Population: {{POPULATION}}. Goal: {{BUSINESS_GOAL}}. Write "Because [evidence], we propose [specific change] for [population], which will [expected behavior] because [mechanism], measured by ONE real outcome metric (not a click), plus 2 guardrails."
Want all 120 prompts in one workspace?
Every prompt in this library, organized by task. Free.
What good output looks like
- The hypothesis contains all four parts (because, we propose, expected impact, measure) and could be handed to a designer or developer as a brief with no further explanation.
- The "because" names a mechanism (friction, clarity, motivation, proof, cognitive load), not a restatement of the change, and traces back to the cited evidence rather than opinion.
Show 4 more quality checks
- One primary metric is designated with a precise definition and source, paired with guardrails that would catch a revenue or quality regression hiding behind a conversion lift.
- Confidence is anchored to the type of evidence behind the idea, and every assumed number or unverified claim is flagged for you to confirm before launch.
- The change is sized against the actual traffic, the detectability verdict is labeled an estimate pending a real sample size calculation, and low detectability triggers a recommendation to research cheaply instead of burning a test.
- Bundled ideas are split into separate hypotheses instead of tested as one ambiguous package, and the proposed change stays inside your stated constraints.
Related prompts
- Design an A/B Test That Will Not Lie to You
Turn this hypothesis into a valid test design that rules out the ways experiments quietly mislead you.
- Calculate Sample Size and Test Duration
Confirm your traffic can actually detect the effect your hypothesis predicts before you launch.
- Prioritize an Experiment Backlog
Once you have several hypotheses, rank them by impact, confidence, and effort to decide what to test first.
Free to use and share. If you republish a prompt, link back to this library.