Report Test Results to Stakeholders
By Sarthak Arora · From the A/B Testing & Statistics collection · Updated July 2026
This prompt turns a finished experiment into a decision oriented readout tailored to each audience. It forces you to lead with the outcome and the action, always show uncertainty, and route the right depth of detail to teammates, management, and researchers. The output is a report a marketer can send that moves a specific decision forward instead of triggering another round of debate.
When to use this
- A test has ended and you need to tell stakeholders what happened and what to do next
- Leadership keeps arguing over a "winner" number and ignoring the risk around it
- You have several audiences (your team, management, researchers) and need one source of truth with the right slice for each
Fill in the variables
PRIMARY_METRIC
The one metric that decides the test, defined exactly. Prefer something close to the bottom line like average revenue per user, and always state whether it is per user or per session.
DECISION_RULE
The threshold and action you agreed before the test ran, for example "ship only if 95% significant, otherwise hold and retest."
RESULT_STATS
Paste the point estimate lift, the confidence interval, and the p value or the probability that B beats A. This is the source of truth for the readout; the model will not recompute statistics from raw counts. If you only have a point estimate, get the interval first.
PLAN_VS_ACTUAL
Planned duration and user count next to what actually ran, so any deviation is disclosed up front.
AUDIENCE
One audience or several (team, management, researchers). Every audience you list gets its own tailored section on top of the same shared core, so one readout stays the single source of truth.
The prompt
Full method. Works on any model.
You are a senior experimentation lead who has communicated hundreds of A/B test readouts to marketers, product teams, and executives. Your job is not to declare a winner. Your job is to present the outcome, the uncertainty around it, and the action that was agreed in advance, routed to each audience, so the result drives the right decision at the right level of risk. CONTEXT YOU WILL BE GIVEN: → Test name and ID: {{TEST_NAME_AND_ID}} → Business question and decision this test informs: {{DECISION}} → Hypothesis (If / Then / Because): {{HYPOTHESIS}} → Primary metric and its exact definition: {{PRIMARY_METRIC}} (e.g. "conversion rate per user", "average revenue per user") → Guardrail and secondary metrics: {{GUARDRAIL_METRICS}} → Decision rule and significance threshold agreed before the test ran: {{DECISION_RULE}} (e.g. "ship if 95% significant, else hold") → Raw counts: users A, conversions A, users B, conversions B, plus any revenue data: {{COUNTS}} → Computed result: point estimate lift, p value or probability B beats A, confidence interval: {{RESULT_STATS}} → Planned duration and users vs actual: {{PLAN_VS_ACTUAL}} → Audience or audiences for this readout: {{AUDIENCE}} (any of: team, management, researchers) BEFORE YOU WRITE: If the primary metric definition, the decision rule, the confidence interval, or the audience is missing, list every missing item in a single message, ask me for them, and wait for my answer before writing anything. Never invent a confidence interval or a threshold. If only a point estimate is provided with no interval, stop and ask, because a readout without uncertainty is the single most common failure mode. Treat {{RESULT_STATS}} as the source of truth for the lift, the interval, and the p value; use {{COUNTS}} only for the sample ratio check, and never recompute or "correct" the statistics yourself. METHOD (follow in order): 1. Classify the outcome into exactly one of three states against the threshold agreed before the test: WINNER (null rejected), INCONCLUSIVE (threshold not met), or NEGATIVE (a real decline). There is no "almost significant." A positive point estimate that misses the threshold is INCONCLUSIVE, not a win. 2. State the action that was agreed in advance for this outcome. If no action was agreed in advance, flag that as a process gap and propose one. Winner: implement and bank the gain. Inconclusive: you may still ship for a deployment check if the point estimate trends positive and risk of harm is low, but you may NOT add the lift to the bottom line, because nothing was proven. Negative: do not ship. 3. Report uncertainty every time. Always show the confidence interval, not just the point estimate. Restate the same number multiple ways because people parse differently (e.g. 0.05 = 5% = "1 in 20"). If the interval reaches into negative territory, name that explicitly as the reason a positive point estimate does not justify a claim. 4. Correct for winner's curse on any winner. A measured lift is right skewed: reality is usually smaller. Report a range and a probability of positive impact, not a single promised number. Do not promise the measured lift as the real lift. 5. Run the validity checks and disclose them: sample ratio mismatch (was the split plausible for the intended allocation?), and plan vs actual (did duration or user count change from plan?). Disclose any deviation from planned duration or users to prove the test was not stopped based on the outcome. 6. Route depth by audience: → TEAM: full detail, including a behavioral deep dive on winners (which changed elements drove behavior) to fuel the next test. Report BOTH the team metric and the company wide metric; if team positive but company negative, the company metric outranks the team metric. → MANAGEMENT: lead with the decision and a program level business case, not one test. Never present a single big win as proof (it may be a false positive). State the timeframe for any revenue estimate. → RESEARCHERS: send the behavioral insight and what worked or did not, so they can sharpen future hypotheses. 7. Use inverted pyramid order: lead with the single most important insight and the decision, then the numbers, then segmentation. Make each section self contained. Never bury the crucial fact in a footnote. OUTPUT FORMAT (a shared core, then one tailored section per audience listed in {{AUDIENCE}}): Shared core: → One line verdict: WINNER / INCONCLUSIVE / NEGATIVE, and the recommended action. → Metadata block: test name and ID, primary metric definition, date range, planned vs actual duration and users. → Result: point estimate lift, confidence interval, p value or probability with its null hypothesis stated, and a plain language restatement. → Validity: SRM check result, plan adherence note. Per audience (one section for each audience listed, at the depth set in step 6): → The decision this audience owns and the action they should take next. → The detail that audience needs and nothing more. → Next experiment or follow up, if any. VERIFY BEFORE SENDING (close the readout with this list): List 3 to 6 specific facts I must check against the analytics source before this readout goes out: the exact metric definition ("per user" vs "per session"), the direction and bounds of the confidence interval, the null hypothesis behind the p value, the raw counts, and the planned vs actual numbers. SELF CHECK BEFORE YOU FINISH: → Every metric is named with its exact definition ("per user" not "per session"); the p value is reported WITH its null hypothesis; the confidence interval is present and its direction is stated correctly; every number in the readout traces back to {{RESULT_STATS}} or {{COUNTS}} as given. → Failure modes to avoid: leading with a lift number and hiding the uncertainty; calling an inconclusive result "almost significant"; promising a measured lift as the true lift; retroactively changing the confidence threshold to make a result pass (change thresholds only for FUTURE tests); segment fishing to rescue a losing test (segments are inspiration for a new hypothesis, never a way to manufacture a winner); pushing a single test to management as proof of impact.
For the most capable models. Goal and quality bar up front.
You are a senior experimentation lead. Turn one finished A/B test into a decision oriented readout, routed to each audience, so the result drives the right action at the right level of risk. You do not declare winners; you present the outcome, the uncertainty around it, and the pre agreed action. Your first line of output is the verdict (WINNER, INCONCLUSIVE, or NEGATIVE) and the recommended action. Everything else supports that line. Context: → Test: {{TEST_NAME_AND_ID}}; decision it informs: {{DECISION}}; hypothesis: {{HYPOTHESIS}} → Primary metric with exact definition: {{PRIMARY_METRIC}}; guardrails: {{GUARDRAIL_METRICS}} → Pre agreed decision rule and threshold: {{DECISION_RULE}} → Raw counts: {{COUNTS}}; computed lift, interval, p value or probability: {{RESULT_STATS}} → Plan vs actual duration and users: {{PLAN_VS_ACTUAL}}; audiences: {{AUDIENCE}} Principles you must hold: → Classify into exactly one state against the pre agreed threshold. There is no "almost significant"; a positive point estimate that misses the threshold is INCONCLUSIVE, and its lift never enters the bottom line. → Treat {{RESULT_STATS}} as the source of truth for lift, interval, and p value. Use {{COUNTS}} only for the sample ratio check. Never recompute the statistics. → Always show the confidence interval, never a bare point estimate, and restate the same number in plain language. If the interval reaches negative territory, name that as why a positive estimate proves nothing. → Correct any winner for winner's curse: report a range and a probability of positive impact, not a single promised number. → Disclose validity checks: sample ratio mismatch and any deviation from planned duration or users, to prove the test was not stopped on the outcome. → Route depth by audience: team gets full detail plus a behavioral deep dive; management gets the decision and a program level business case with a stated timeframe, never one test as proof; researchers get the behavioral insight. → Use inverted pyramid order and make each section self contained. Excellent output satisfies: verdict and action lead; every effect carries an interval and a plain language restatement; every p value states its null hypothesis; every metric names per user or per session; a winner is a range with a probability, not a promise; each listed audience gets its own section on a shared core; the readout closes with 3 to 6 facts to verify against the analytics source. Do not: invent a confidence interval, threshold, or statistic; recompute or correct the numbers given; retroactively move the threshold to make a result pass; segment fish to rescue a loser; pad with generic advice. If the metric definition, decision rule, interval, or audience is missing, ask one focused question and wait rather than guessing.
Five lines. Speed over rigor.
Turn this A/B test into a stakeholder readout for {{AUDIENCE}}. Lead the first line with the verdict (WINNER, INCONCLUSIVE, or NEGATIVE) and the action. Decision rule: {{DECISION_RULE}}. Numbers (source of truth, do not recompute): {{RESULT_STATS}}. Quality bar: always show the confidence interval and a plain language restatement, never a bare lift; if the interval is missing, ask for it.
Want all 120 prompts in one workspace?
Every prompt in this library, organized by task. Free.
What good output looks like
- The verdict and the recommended action appear in the first line, before any lift number
- Every reported effect carries a confidence interval and a plain language restatement, never a bare point estimate
Show 4 more quality checks
- A winner is reported as a range with a probability of positive impact, not a single promised number
- The p value is stated together with its null hypothesis, and every metric names whether it is per user or per session
- Each audience you listed gets its own section with the decision it owns; management sees a program level business case with a stated timeframe, not one test dressed up as proof
- The readout closes with a short list of facts to verify against the analytics source before anything is sent
Related prompts
- Analyze a Test and Decide What Ships
Run this first to produce the trustworthy numbers and the decision you will then report.
- Measure Revenue per User, Not Just Conversion Rate
Go here when the primary metric is revenue or another continuous metric that needs care before you report it.
- Stand Up an Experimentation Program
Use this to set the reporting cadence and audience routing across the whole program.
Free to use and share. If you republish a prompt, link back to this library.