AIkratesAIkrates
Book a call
Prompt Library/A/B Testing & Statistics

Write a Falsifiable Test Hypothesis

By Sarthak Arora · From the A/B Testing & Statistics collection · Updated July 2026

This prompt converts a raw idea or research finding into a falsifiable hypothesis that names the exact change, the population it targets, the mechanism that should drive behavior, the outcome you expect, and how you will measure it. The output doubles as a design brief and as the front page of a test scorecard, so nobody has to ask "why are we running this?" later.

When to use this

  • You have a vague request ("make the headline more emotional," "refresh the page") and need to force it into something an operator could build and measure.
  • Research surfaced a problem and you want to state a solution as a bet you can win or lose, before you spend engineering time or traffic.
  • You are about to launch an A/B test and need a written hypothesis, primary metric, and guardrails locked in before you see any data.

Fill in the variables

IDEA_OR_FINDING

The raw input, even if it is vague.

Example: "Add a pricing CTA to the header" or "Users can't find case studies."

RESEARCH_EVIDENCE

What supports the idea, with source.

Example: "Session recordings show 8 of 10 users backtracking to find pricing; pricing was the top site search term."

LOCATION

The exact surface.

Example: "Product detail page, above the fold" or "Main navigation, desktop and mobile."

POPULATION

Who should respond.

Example: "First time visitors from paid search" or "Logged out prospects."

BUSINESS_GOAL

The north star the change should ladder up to.

Example: "Qualified demo requests" or "Revenue per visitor."

TRAFFIC_AND_CONVERSIONS

Weekly unique visitors and conversions on that surface, so the model can judge detectability.

Example: "40,000 weekly uniques, 3% conversion."

CONSTRAINTS

Anything that limits the change.

Example: "Cannot alter checkout logic; brand forbids scarcity countdowns; no capacity to maintain many personalized variants."

The prompt

Full method. Works on any model.

You are a senior experimentation strategist. Your job is to convert a raw idea or
research finding into ONE testable hypothesis that an independent operator could
build and evaluate without asking you what anything means. You reason from
evidence, not taste. You never write a hypothesis after results are in.

CONTEXT
→ Idea or finding: {{IDEA_OR_FINDING}}
→ Supporting evidence: {{RESEARCH_EVIDENCE}}
→ Page, flow, or surface affected: {{LOCATION}}
→ Target population: {{POPULATION}}
→ Business goal or north star outcome: {{BUSINESS_GOAL}}
→ Traffic and conversion volume on this surface: {{TRAFFIC_AND_CONVERSIONS}}
→ Known constraints (tech, brand, legal, maintenance): {{CONSTRAINTS}}

FIRST: if any of these are missing or too vague to write a specific hypothesis
(especially the mechanism, the population, or how the primary metric is defined),
ASK up to five clarifying questions before writing anything. Do not invent
evidence to fill a gap; flag the gap instead.

METHOD
1. Interrogate the source. Name what produced this idea: prior research, a past
   test, an analytics pattern, a support signal, or an opinion. If the only
   backing is opinion, seniority, or a competitor pattern, say so plainly and
   lower confidence. Analytics shows what, not why; do not present a correlation
   as a cause.
2. Replace vague verbs. Turn improve, refresh, modernize, simplify, or "more
   emotional" into one observable change to a specific element, content, flow, or
   interaction. The change must be buildable within {{CONSTRAINTS}}; if the
   strongest fix violates a constraint, say so and propose the closest compliant
   version. If several changes are bundled together, split them into separate
   hypotheses and hypothesize the highest priority one; list the rest.
3. Name the mechanism. State the behavioral or usability reason the change should
   move behavior. Every "because" should trace to a real reason (reduced
   friction, higher clarity, higher motivation, stronger proof, lower cognitive
   load), not a restatement of the change. A behavior needs motivation and
   ability to both be present for a trigger to fire (BJ Fogg Behavior Model);
   check that your change raises the one that is actually missing.
4. Target the population. Aim at the group the change can actually influence
   (often new or first time visitors, since returning users have already seen the
   current experience). A universal claim is fine only when the effect should
   truly be universal.
5. Choose the primary metric with a goal tree. Connect the north star outcome down
   to the controllable driver you are changing. Pick ONE primary outcome that
   answers the question. Prefer transactions, leads, or revenue per user over
   clicks; a click lift rarely proves a behavior or revenue lift. Conversion rate
   is a driver, not a destination.
6. Add guardrails. Name the metrics that must not deteriorate (revenue, margin,
   completed purchase rate, refund or cancellation rate, retention, page speed).
   A conversion win that lowers revenue through worse order mix or heavier
   discounting is not a win.
7. Sanity check the size against the data. Using {{TRAFFIC_AND_CONVERSIONS}},
   estimate roughly what relative lift this surface could detect in four to six
   weeks, and label that figure an estimate that a proper sample size calculation
   must confirm before launch. Low volume surfaces can only detect large effects,
   so a tiny tweak on a low traffic page is untestable; either make the change
   bigger or route it to cheaper research. Frame the change as small, medium, or
   big and match ambition to detectable effect.

OUTPUT FORMAT
1. Hypothesis (four parts, one paragraph):
   "Because [evidence], we propose [specific change] for [population], which will
   [expected behavior change] because [mechanism], measured by [primary metric],
   while monitoring [guardrail metrics]."
2. Source and confidence: what backs this and a confidence level with one line of
   justification. Anchor the level to the evidence type: high means your own user
   data or a prior test on this audience; medium means a strong analytics or
   research signal without causal proof; low means opinion, a competitor pattern,
   or a single anecdote.
3. Primary metric definition: exact metric, data source, inclusion rule, and
   attribution window.
4. Guardrail metrics: 2 to 4, each with why it matters.
5. Change scope: small, medium, or big, plus whether current traffic can plausibly
   detect it, labeled as an estimate pending a proper sample size calculation. If
   detection is implausible, recommend a cheaper research method instead of a test.
6. Split off hypotheses: any other testable statements hidden in the original
   idea, listed separately.
7. Decision rule: one line stating what evidence would lead you to implement,
   keep the control, revise, or stop.

SELF CHECK before returning
→ Could an independent operator build this and read the result without asking what
  any term means? If not, tighten it.
→ Is exactly ONE primary metric designated, and is it a real outcome rather than a
  vanity click?
→ Does the "because" give a mechanism, not just repeat the change?
→ Is the population the group the change can actually move?
→ Does the proposed change respect every stated constraint?
→ Did you avoid post hoc reasoning (writing the hypothesis to fit a result you
  already want)?
→ Flag every number, baseline, or evidence claim you assumed rather than received,
  and list what the operator must verify before launch.
→ Failure modes to avoid: bundling multiple changes into one hypothesis; scoring
  confidence off opinion or a competitor screenshot; claiming a click lift equals
  a revenue lift; targeting "all users" when only new users can respond; proposing
  a change too small for the available traffic to detect.

For the most capable models. Goal and quality bar up front.

You are a senior experimentation strategist. Your first line of output is the
finished hypothesis paragraph; everything after it supports that verdict.

GOAL
Convert a raw idea or research finding into ONE falsifiable hypothesis that an
independent operator could build and evaluate without asking you what anything
means. Reason from evidence, not taste, and never write a hypothesis to fit a
result you already want.

CONTEXT
→ Idea or finding: {{IDEA_OR_FINDING}}
→ Supporting evidence: {{RESEARCH_EVIDENCE}}
→ Page, flow, or surface affected: {{LOCATION}}
→ Target population: {{POPULATION}}
→ Business goal or north star outcome: {{BUSINESS_GOAL}}
→ Traffic and conversion volume: {{TRAFFIC_AND_CONVERSIONS}}
→ Known constraints (tech, brand, legal, maintenance): {{CONSTRAINTS}}

PRINCIPLES (the load bearing rules of the method)
→ Trace every claim to its source. Name what produced the idea (research, a past
  test, analytics, a support signal, or opinion) and lower confidence when the
  only backing is opinion, seniority, or a competitor pattern. Analytics shows
  what, not why; a correlation is not a cause.
→ Replace vague verbs (improve, refresh, simplify, "more emotional") with one
  observable, buildable change that respects every stated constraint. If several
  changes are bundled, hypothesize the highest priority one and list the rest.
→ Name a real mechanism behind the "because" (reduced friction, higher clarity,
  higher motivation, stronger proof, lower cognitive load), never a restatement
  of the change. A behavior needs both motivation and ability present; raise the
  one that is actually missing.
→ Target the group the change can move (often new or first time visitors), not
  "all users" unless the effect is truly universal.
→ Choose ONE primary metric by laddering the north star down to the driver you
  are changing. Prefer transactions, leads, or revenue per user over clicks.
  Pair it with 2 to 4 guardrails that would catch a revenue or quality
  regression hiding behind a conversion lift.
→ Size the change against {{TRAFFIC_AND_CONVERSIONS}}. Label the detectable lift
  an estimate pending a proper sample size calculation. If detection is
  implausible, recommend cheaper research instead of a test.

QUALITY BAR
Excellent output is a four part hypothesis ("Because [evidence], we propose
[change] for [population], which will [expected behavior] because [mechanism],
measured by [primary metric], while monitoring [guardrails]"), plus source and
confidence anchored to evidence type, an exact primary metric definition, sized
scope with a detectability verdict, any split off hypotheses, and a one line
decision rule stating what evidence would implement, keep control, revise, or
stop.

BOUNDARIES
→ Do not invent data, evidence, or numbers; flag every assumed baseline the
  operator must verify before launch.
→ Do not bundle multiple changes, score confidence off opinion, or claim a click
  lift equals a revenue lift.
→ If a required input is missing or too vague to write a specific hypothesis, ask
  one focused question instead of guessing.

Five lines. Speed over rigor.

Turn this idea into ONE falsifiable if then because hypothesis I could build and
measure. Idea: {{IDEA_OR_FINDING}}. Evidence: {{RESEARCH_EVIDENCE}}. Population:
{{POPULATION}}. Goal: {{BUSINESS_GOAL}}. Write "Because [evidence], we propose
[specific change] for [population], which will [expected behavior] because
[mechanism], measured by ONE real outcome metric (not a click), plus 2 guardrails."

Want all 120 prompts in one workspace?

Every prompt in this library, organized by task. Free.

What good output looks like

  • The hypothesis contains all four parts (because, we propose, expected impact, measure) and could be handed to a designer or developer as a brief with no further explanation.
  • The "because" names a mechanism (friction, clarity, motivation, proof, cognitive load), not a restatement of the change, and traces back to the cited evidence rather than opinion.
Show 4 more quality checks
  • One primary metric is designated with a precise definition and source, paired with guardrails that would catch a revenue or quality regression hiding behind a conversion lift.
  • Confidence is anchored to the type of evidence behind the idea, and every assumed number or unverified claim is flagged for you to confirm before launch.
  • The change is sized against the actual traffic, the detectability verdict is labeled an estimate pending a real sample size calculation, and low detectability triggers a recommendation to research cheaply instead of burning a test.
  • Bundled ideas are split into separate hypotheses instead of tested as one ambiguous package, and the proposed change stays inside your stated constraints.

Related prompts

  • Design an A/B Test That Will Not Lie to You

    Turn this hypothesis into a valid test design that rules out the ways experiments quietly mislead you.

  • Calculate Sample Size and Test Duration

    Confirm your traffic can actually detect the effect your hypothesis predicts before you launch.

  • Prioritize an Experiment Backlog

    Once you have several hypotheses, rank them by impact, confidence, and effort to decide what to test first.

Free to use and share. If you republish a prompt, link back to this library.

Contact

IGNIPC Private Limited

1st Floor, Flat No. 111, Hemkunt Chambers

Nehru Place, New Delhi 110019

India

+91 8700187916

sarthak@aikrates.com

Pipeline

  • P3 Sprint
  • Pipeline Clarity Audit

AI Strategy

  • Voice Agents
  • Business Audit
  • Enterprise

Solutions

  • Real Estate

Resources

  • Blog
  • Prompt Library
  • Book a call

Company

  • Privacy Policy
  • Terms of Service
  • LinkedIn

© 2026 IGNIPC Private Limited.