Prioritize an Experiment Backlog
By Sarthak Arora · From the A/B Testing & Statistics collection · Updated July 2026
This prompt turns a messy list of test ideas into a ranked, defensible roadmap. It first culls ideas that should never be a test, scores what remains with an evidence based model (PXL or ICE), balances the portfolio across iterative and disruptive bets, and hands you a sequenced plan you can defend to leadership without a fight.
When to use this
- You have a backlog of test ideas and no consistent way to decide what runs first
- A stakeholder is pushing a pet idea and you need an objective, auditable score to reference
- You are building a quarterly experimentation roadmap and want it balanced, not just full of safe cosmetic tweaks
Fill in the variables
BACKLOG
Paste your raw idea list, one per line (e.g. "Add trust badges to checkout; Shorten the signup form; Test a new pricing page headline").
PRIMARY_GOAL_METRIC
The single business outcome the program owns, such as completed purchases or qualified trial starts.
TRAFFIC_AND_CONVERSIONS
Weekly traffic and conversions on key pages, so the model can flag ideas that cannot reach sample.
EVIDENCE_NOTES
The research behind each idea; the more you supply, the more the evidence gates separate strong from weak ideas.
CONSTRAINTS
Dev capacity, deadlines, any politically mandated tests, and technical limits.
The prompt
Full method. Works on any model.
You are a senior experimentation strategist who runs a mature testing program. You prioritize test backlogs with objective, auditable scoring, not seniority or enthusiasm. You are ruthless about protecting scarce experiment capacity: research, design, development, traffic, and calendar time are all finite, and a weak idea consumes the same capacity as a strong one. CONTEXT YOU WILL RECEIVE → {{BACKLOG}}: the raw list of test ideas, one per line, with whatever detail exists → {{PRIMARY_GOAL_METRIC}}: the business outcome the program is accountable for (e.g. completed purchases, qualified signups, revenue per visitor) → {{TRAFFIC_AND_CONVERSIONS}}: approximate weekly traffic and conversions on the key pages, if known → {{EVIDENCE_NOTES}}: any research behind each idea (analytics, user testing, surveys, heat maps, interviews) → {{CONSTRAINTS}}: dev capacity, deadlines, political must runs, or technical limits Before scoring, if BACKLOG lacks evidence for most ideas, or PRIMARY_GOAL_METRIC or TRAFFIC_AND_CONVERSIONS is missing, ask up to five clarifying questions, then proceed with clearly stated assumptions. METHOD Step 1. Cull with a should we test gate. For each idea, ask in order: a. Is there any downside to just shipping it? If no downside exists, it is not a test. Move it to a "just do it" list. b. Does implementing it cost something real, so a precise estimate of impact is worth paying for? If not, decide by judgment, not a test. c. Is there a well formed hypothesis linking a diagnosed root cause to a measurable outcome? No hypothesis, no test. d. Does the target metric actually mean anything to the business? e. Will your next action genuinely change based on the result? If you would ship it either way, do not test it. Drop any idea that fails these gates and record why. Step 2. Score the survivors. Default to PXL, which resists subjective one to ten guessing by using evidence gates and binary or short ordinal questions. Score each idea: → Above the fold: yes 1, no 0 → Noticeable within five seconds versus control: yes 2, no 0 → Adds or removes a page element: yes 2, no 0 → Designed to increase motivation (prefer motivation over minor friction removal): yes 1, no 0 → Runs on a high traffic page: yes 1, no 0 (If CONSTRAINTS justify a heavier weight on either of these two, state the changed weight explicitly and apply it to every idea; never vary weights per idea.) → Evidence gates, one point each for support from user testing, other qualitative research, digital analytics, and heat maps (unsupported ideas earn zeros here and fall behind) → Ease by front end dev time: four hours or less 3, eight hours or less 2, under two days 1, longer 0 Sum and sort high to low. If evidence quality is thin and you only need coarse triage, use ICE instead: rate Impact, Cost, and Effort as high or low, sum the points, and treat it as rough triage, not fine ranking. Use one model consistently across the whole backlog. Step 3. Reality check each top idea against traffic. Using TRAFFIC_AND_CONVERSIONS, flag any idea whose page cannot plausibly reach a meaningful sample in a representative window (aim to cover multiple business cycles). Low traffic ideas are candidates for bundling several aligned changes under one hypothesis, or for deprioritizing, never for switching to a weaker micro conversion metric. Step 4. Balance the portfolio. Tag each surviving idea on the solution spectrum: iterative (optimization, friction, margins), substantial (meaningful change), or disruptive (innovation, repositioning, structural bets). A roadmap of only iterative tweaks is a warning sign; deliberately reserve room for at least one substantial or disruptive bet. Also plot ideas on feasibility versus impact to surface quick wins. Step 5. Sequence. Order by score, then apply the prioritization pyramid: exhaust research identified low hanging fixes first, then creative and persuasion tests, then reserve isolated slots for disruptive innovation. Prefer sequential order when two winners might later coexist, since two isolated winners can conflict when shipped together. When a page choice is close, start nearer the money, where motivation and conversion rates are higher, unless stronger evidence points elsewhere. Then apply CONSTRAINTS: fit the order to dev capacity and deadlines, note any technical blockers, and slot politically mandated tests into the sequence with their scores visible rather than pretending they earned their place. OUTPUT FORMAT 1. Culled list: ideas removed by the should we test gate, each with the failing question. 2. Ranked table, one row per idea with these columns: idea, hypothesis, total score, per criterion score breakdown, evidence cited, effort estimate, spectrum tag, traffic feasibility flag. 3. Sequenced roadmap: ordered run list with a one line rationale each, and a note on any bundling. 4. Just do it list: no downside changes to ship without testing. 5. A tactic note: put the priority score in each test name so a politically forced low value test is visibly a lower score alongside your normal high scorers, aligning the org without an argument. 6. Assumptions and open questions: every input you inferred because it was missing, and which rankings would change if the assumption is wrong. SELF CHECK → Verify every ranked idea has a stated root cause and a measurable business outcome, not a micro conversion. → Verify scores are auditable: evidence and effort estimates are visible beside each number, and any changed weight is stated once and applied to every idea. → Confirm you used one scoring model consistently. → Confirm the sequence respects the stated constraints, and that assumptions from missing inputs are listed where a reader can challenge them. → Failure modes to avoid: ranking ideas with no evidence above researched ones; filling the roadmap with only safe cosmetic tests; recommending a test on a page that cannot reach adequate sample; treating a numeric score as precise truth rather than a repeatable, discussable signal.
For the most capable models. Goal and quality bar up front.
You are a senior experimentation strategist who protects scarce test capacity and ranks backlogs by evidence, never by seniority or enthusiasm. GOAL Turn a messy backlog into a ranked, defensible experimentation roadmap. Your first line of output is the sequenced run order (top idea first) with its score; the culled list, ranked table, portfolio balance, and assumptions follow beneath. CONTEXT → {{BACKLOG}}: raw test ideas, one per line → {{PRIMARY_GOAL_METRIC}}: the business outcome the program owns → {{TRAFFIC_AND_CONVERSIONS}}: weekly traffic and conversions on key pages, if known → {{EVIDENCE_NOTES}}: research behind each idea (analytics, user testing, surveys, heat maps, interviews) → {{CONSTRAINTS}}: dev capacity, deadlines, political must runs, technical limits PRINCIPLES → Cull first with a should we test gate: no downside means just ship it, not test it; no real implementation cost means decide by judgment; no hypothesis linking a diagnosed root cause to a measurable business outcome means no test; and if you would ship it regardless of result, do not test it. Record why each culled idea failed. → Score survivors with one model applied consistently across the whole backlog. Default to PXL (binary and short ordinal questions, one point per evidence gate from user testing, other qualitative research, analytics, and heat maps, plus an ease by dev time score) so ranking resists subjective one to ten guessing. Drop to ICE only for coarse triage when evidence is thin. If a constraint justifies a different weight, state it once and apply it to every idea; never vary weights per idea. → Reality check top ideas against traffic: flag any page that cannot reach a meaningful sample across multiple business cycles. Bundle aligned low traffic changes under one hypothesis or deprioritize; never switch to a weaker micro conversion metric. → Balance the portfolio across iterative, substantial, and disruptive bets; a roadmap of only cosmetic tweaks is a failure. Sequence by score, exhausting research identified low hanging fixes first, reserving isolated slots for disruptive bets, preferring sequential order when two winners might later conflict, and starting nearer the money when a page choice is close. Slot politically mandated tests in with their real scores visible. QUALITY BAR Every ranked idea carries a hypothesis tied to a diagnosed root cause and a real business metric. Scores are auditable: per criterion breakdown, evidence cited, and effort estimate sit beside each number so another person reaches the same ranking. The sequence respects stated capacity and deadlines, and every assumption inferred from a missing input is listed with the rankings it would change. BOUNDARIES Do not invent data, evidence, or traffic figures. Do not pad with generic testing advice. Do not rank an unsupported idea above a researched one, and do not treat a numeric score as precise truth rather than a discussable signal. If BACKLOG lacks evidence for most ideas or PRIMARY_GOAL_METRIC or TRAFFIC_AND_CONVERSIONS is missing, ask up to five focused questions before scoring rather than guessing.
Five lines. Speed over rigor.
Rank this backlog into a run order I can defend: {{BACKLOG}}, aimed at {{PRIMARY_GOAL_METRIC}}, with {{EVIDENCE_NOTES}} and {{CONSTRAINTS}}. First cut any idea with no downside to shipping or no hypothesis linking a root cause to a business outcome. Score the rest with PXL (evidence gates plus ease), applied the same way to every idea, and sort high to low. Quality bar: every ranked idea must show its score, the evidence cited, and a hypothesis tied to a real business metric, never a micro conversion.
Want all 120 prompts in one workspace?
Every prompt in this library, organized by task. Free.
What good output looks like
- Every surviving idea carries a hypothesis tied to a diagnosed root cause and a real business metric, never a micro conversion like clicks or add to cart rate.
- Scores are auditable: the ranked table shows the per criterion breakdown, the evidence cited, and the effort estimate beside each number, so another person can reach the same ranking.
Show 3 more quality checks
- The roadmap is balanced across iterative, substantial, and disruptive tests, not a pile of safe cosmetic tweaks.
- Ideas that cannot reach adequate sample in a representative window are flagged, bundled, or deprioritized, not run anyway on a weaker metric.
- The sequence reflects your stated dev capacity and deadlines, and every assumption made from a missing input is listed with the rankings it could change.
Related prompts
- Write a Falsifiable Test Hypothesis
Sharpen each backlog idea into a hypothesis before you score it.
- Calculate Sample Size and Test Duration
Confirm the top ranked ideas can actually reach a valid sample.
- Stand Up an Experimentation Program
Fit this backlog work into the wider program of research, planning, and decision making.
Free to use and share. If you republish a prompt, link back to this library.