A/B Test Winner Analyzer
Turn raw A/B test metrics into a clear winner, confidence level, rationale, and next experiment.
Customise & copy
Add the context you have. The prompt updates as you work.
Select the single outcome that determines the winner.
Include duration and any relevant sample-size context.
Provide the raw metrics available for Variant A.
Provide the raw metrics available for Variant B.
Live prompt
4 fieldsInterpret the A/B test data below and return only a valid JSON object with exactly these four fields: "winner", "why", "confidence", and "next_test". Requirements: - "winner": Explicitly name Variant A or Variant B. Choose the variant that best supports the stated business goal. If the evidence is weak or nearly even, still choose the leading variant and set confidence to "Low". - "why": Concisely compare the relevant rates or outcomes, including the calculations when possible. Consider sample size, test duration, effect size, and data quality. Do not claim statistical significance unless it is supported by the provided data. - "confidence": Use exactly one of: "High", "Medium", or "Low". Base it on sample size, result difference, duration, and consistency with the business goal. - "next_test": Propose one specific follow-up variant or experiment that builds on the result. Do not request more information, use external data, or add greetings, caveats, or filler. Primary business goal: Conversion rate Test duration or sample size details: 14 days; 25,000 impressions per variant Variant A performance metrics: Impressions: 50,000; clicks: 2,400; conversions: 720 Variant B performance metrics: Impressions: 50,000; clicks: 2,650; conversions: 795
Proof before you use it
A real example
Tested on gpt-5.6-luna on 2026-09-10
{
"test_details": "21 days; 8,200 sessions per variant; paid and organic SaaS signup traffic",
"business_goal": "Lead generation",
"variant_a_metrics": "Sessions: 8,200; pricing-page visits: 1,640; demo requests: 164",
"variant_b_metrics": "Sessions: 8,050; pricing-page visits: 1,530; demo requests: 191"
}A small ritual that works
How to use it
- 01Select the primary business goal.
- 02Enter test duration and sample-size details.
- 03Paste Variant A performance metrics.
- 04Paste Variant B performance metrics.
Why it works
The prompt gives the model a clear decision task, a defined business objective, and separate inputs for timing, sample context, and performance metrics. Its four fixed JSON fields create a compact reporting structure: a selected variant, a concise rationale, a confidence label, and one follow-up action. The instructions also encourage rate calculations, comparison of effect size, consideration of data quality, and restraint around significance claims. Requiring a winner even when evidence is close prevents an undecided response while the Low confidence option preserves uncertainty.
Where it fails
The prompt is less suitable when the supplied metrics do not clearly map to the selected business goal or when the free-text details omit important context such as exposure definitions, duplicate events, eligibility rules, or attribution windows. Its fixed confidence vocabulary also compresses nuanced uncertainty into only three categories. Because the follow-up must be one specific experiment and cannot request more information, the structure favors decisive guidance over extended diagnosis.
Customisation tips
Add goal-specific calculation instructions if different metrics require different denominators, such as revenue per visitor or lead rate. Define the evidence thresholds that distinguish High, Medium, and Low confidence, and specify any minimum duration or volume requirements. If the workflow needs more diagnostic detail, add fields for assumptions, data-quality issues, or segmentation while preserving the required top-level structure. Make the follow-up direction more precise by naming the audience, experience element, or constraint that should guide the next comparison.
Watch-outs
- The prompt accepts free-text metric inputs, so inconsistent labels or denominators can make comparisons difficult to interpret.
- The instruction to choose a leading variant can encourage a selection even when the difference is practically negligible; the confidence field should communicate that limitation.
- The business goal is selected separately from the metric descriptions, so the supplied numbers should directly support that goal.
- The JSON requirement is strict: extra fields, explanatory text outside the object, or invalid formatting can make the response unusable to a parser.
Works in
Was this useful?
Submit your variation
If your version helps, it may be published with your name and a link back to you.
Want to make AI useful across your team? Explore Atul's training.