
Attribution
Part of Ecommerce experimentation
Reporting a test result without overstating significance
Write a store-test result with group rates, effect size, uncertainty, data checks, guardrails and a bounded decision.
Report a store test as a group comparison, an uncertainty estimate and a bounded decision. Give each group's primary outcome, the difference and any reason the comparison may be unreliable. A significance label alone does not show how large the effect could be or whether it is worth acting on.
Put the estimate first
Use the population and metric defined before launch. State the assigned eligible units per group, observation dates and the order-status rule. For a purchase rate, show both rates and their difference in percentage points. A relative percentage is useful only with the control rate and its denominator visible.
As an arithmetic illustration, rates of 4.0% and 4.4% differ by 0.4 percentage points, or 10% relative to control. These are hypothetical rates, not a test result. Without group counts, an uncertainty calculation and data-quality checks, they establish neither statistical significance nor commercial value.
Report an effect estimate with a confidence interval or another uncertainty measure from the stated analysis method. Explain what the range means for the decision: does it still allow an unacceptable loss, or does it exclude the improvement needed to justify the change? A confidence interval is not a guaranteed range for every future rollout.
Say whether the comparison held
Before interpreting the estimate, report whether observed assignment counts were consistent with the configured split and whether outcomes were collected comparably for both variants. Note missing orders or events, implementation deviations and any narrower subgroup analysis. An unexplained sample-ratio mismatch can make an apparently precise lift unreliable.
Show guardrails beside the primary result, with their observation cut-offs. An immediate checkout-error measure and a later value measure affected by refunds or cancellations have different maturity. Mark the later measure provisional when needed. “No detected difference” does not establish no harm if the interval still permits a material loss.
Use significance language carefully
A p-value is not the probability that the treatment works. Passing a threshold does not measure effect size or business importance. A non-significant result does not prove equivalence; a wide interval may cover both a worthwhile gain and an unacceptable loss. Call the result inconclusive for the planned decision when that is what the evidence shows.
Identify the primary analysis and label additional metrics, variants or segments as exploratory where appropriate. More comparisons create more chances for false positives unless the method accounts for them. If the team repeatedly examined fixed-horizon results and could stop early, do not report the selected p-value as though there had been only one planned analysis.
Interpreting Significance: Key Considerations
- Pros of Using p-valuesStandardised threshold for statistical significance; useful in initial screening.
- Cons of Overrelying on p-valuesDoes not reflect effect size or business impact; can lead to false positives without correction.
Record the decision
A closing statement should name the action, population and remaining uncertainty. A hypothetical example is: “Hold the product-page change for the eligible product group; the purchase-rate estimate is too imprecise to rule out the loss set as material.” That is a reporting example, not an observed finding.
Keep the brief version, dates, allocation, metric formula, result table, method, deviations and decision owner with the internal record. A later test of another version or product group adds new evidence; it does not change the original result.



