
Measurement Planning
Ecommerce experimentation
Plan an ecommerce experiment around a clear decision, eligible shoppers, one primary outcome, guardrails and a trustworthy result.
Ecommerce experimentation compares a store change with a randomly assigned control. It helps a trading team decide whether a change affected an outcome rather than coincided with a rise or fall in orders. The useful result is a decision supported by an effect estimate, data checks and clear limits.
Choose the decision
Name the proposed change, the shoppers eligible for it and the action the result will inform. Keep the treatment specific. If a product-page test changes delivery information, price and the purchase button together, the result cannot isolate the contribution of each change.
Choose an assignment unit that suits the experience, such as a shopper or account, and document what happens when someone returns. Device or cookie identifiers can change, so a person may encounter both variants. Define eligibility under a rule applied to both groups; selecting only people who interacted with the treatment would bias the comparison.
Set one primary outcome with a numerator, denominator, order-status rule and observation window. Purchase rate per assigned eligible shopper answers a different question from purchases per session. Include assigned eligible shoppers who did not click or buy. Analysing purchasers alone would select people based on an outcome the treatment may have changed.
Primary Outcome Metrics: Purchase Rate vs Purchases per Session
- Purchase Rate per Assigned Eligible Shopper
- Measures conversion from eligible users; includes non-clickers and non-buyers
- Purchases per Session
- Measures transaction frequency within a session; excludes non-purchasers
Make the hypothesis testable
State the expected effect of the treatment and choose metrics that could show whether that effect occurred. A claim that a change will improve brand perception is difficult to test unless a metric can measure that outcome. Microsoft Research recommends keeping a hypothesis simple; a complex change can be split into simpler changes, each with its own hypothesis.
Plan before launch
Record the control, treatment, allocation, expected eligible traffic and ordinary end condition. Check whether the planned sample could detect a change small enough to matter to the decision.
A minimum detectable effect is a planning measure under stated assumptions, not a predicted result. If traffic is insufficient, acknowledge that a small but important effect may remain unresolved.
Choose a few guardrails for plausible harms, and specify when their data will be available. Keep data-quality checks separate: missing outcome records or a sample-ratio mismatch can call the comparison itself into question.
For monetary outcomes, specify whether value is measured at checkout or after later adjustments. Shopify's sales-report documentation says pending, cancelled and unpaid orders are included in gross sales. A paid-order outcome needs a suitable order-status rule. A Google Analytics purchase event depends on the store's implementation and should not be assumed to represent that milestone.
Pros and Cons of Using Shopify Sales Report for Monetary Outcomes
- ProsIncludes pending, cancelled and unpaid orders in gross sales; useful for early trend analysis
- ConsDoes not reflect final paid orders; may overstate actual revenue if not adjusted
Size the test with relevant data
A power analysis can help estimate the traffic or duration needed to detect a chosen effect. Statsig's calculator uses a metric's historical mean and variance, along with traffic volume, to estimate the relationship between minimum detectable effect, exposures or days, and allocation.
Use a population that resembles the shoppers who will be eligible for the test. The calculator can use the whole user base, a targeting gate or data from a previous experiment; a previous experiment is most relevant when it affects a similar user base or part of the product. Statsig requires a targeting gate used for this calculation to have been active for at least seven days.
Key Experiment Planning Metrics
- Minimum Detectable Effect (MDE)
- Smallest meaningful change the test can detect
- Historical Variance
- Used in power analysis for metric stability
- Targeting Gate Duration
- Must be active for at least 7 days for reliable calculation
Run and read the comparison
Check that both variants render on the relevant routes and that assignment and outcomes can be observed for both groups. Record stock gaps, promotions and releases during the run.
Monitor faults and data quality; an urgent checkout problem can justify stopping the experience. For an ordinary success decision, follow the planned end rule and analysis method. Repeated early decisions based on a conventional fixed-horizon significance threshold raise false-positive risk.
Before interpreting lift, compare observed assignment counts with the configured allocation using an appropriate sample-ratio check. Investigate a persistent mismatch and any variant-specific logging or outcome gaps. Then report each group's primary outcome, assigned count, absolute difference, useful relative difference and uncertainty interval under the stated method. Read guardrails alongside the primary result.
A non-significant result does not establish no effect when its interval still permits a material gain or loss. A statistically significant effect may be too small or costly to justify rollout. State whether to launch, revise, repeat or leave the question open for the tested population and period. Do not extend a result from one product group or trading period to every future shopper without further evidence.
Monitor the whole experience
A result can improve one part of the shopping experience while making another worse. Microsoft Research recommends examining a broad set of product metrics so the decision weighs improvements against unintended regressions, such as increased checkout drop-off or longer page load times.
Metrics are more useful when they are sensitive, trustworthy, efficient, debuggable and interpretable. Check them during the run as well as at the end, so an unintended regression can be noticed while there is still an opportunity to intervene. Examine relevant shopper segments, too, rather than assuming the overall result describes every group.
Validate that the measured effect fits the change
Before relying on an outcome movement, check that the treatment is reflected in metrics in a way that fits the change made. For example, adding content to a page may affect page size and load time; adding or removing a product element may affect measures of its coverage. Unexpected or absent movements can point to a setup or measurement issue.
Check that the users included in the analysis match the intended audience, including any market or browser eligibility rules. If the measured population or the observed treatment effects do not fit the test design, investigate the setup before using the result to make a rollout decision.
In this guide
- Defining a conversion experiment before changing the storeWrite a store-test brief that fixes eligibility, assignment, the conversion formula, decision threshold and end rule before launch.
- Choosing a guardrail metric for a store testChoose store-test guardrails from plausible harms, with the right denominator, timing, data source and response rule.
- Checking for sample imbalance in an ecommerce experimentCompare observed experiment units with the planned split, trace where imbalance begins and decide whether the store-test result can be used.
- Reporting a test result without overstating significanceWrite a store-test result with group rates, effect size, uncertainty, data checks, guardrails and a bounded decision.



