Why do your site tests keep coming back inconclusive?

Wine bottles lined up on a cellar shelf

Written by

in

Most site tests come back inconclusive because the change tested is smaller than your traffic can measure, not because the team lacks discipline. A comparison can only report a difference larger than the ordinary variation in the metric, and a typical mid-tier DTC site’s order volume creates a wide band of that variation. A 0.1-second load-speed gain lifted retail conversion 8.4% (Google and Deloitte, 2020) — an effect a small site can see.

Open the last four quarters of test results and count the ones that ended in a decision. For most mid-tier DTC programs, the honest count is a small fraction of the attempts, and the rest were called off, ran past their window, or produced a difference nobody was willing to defend in a meeting. The usual conclusion drawn from that record is that the team lacks discipline.

It is almost never a discipline problem. It is an arithmetic one. A comparison between two versions of a page can only report a difference that is larger than the ordinary variation in the underlying numbers. A site at your order volume has a wide band of ordinary variation, so a small change disappears into it no matter how carefully the test is run, how long it runs, or how much the team wants an answer.

That gives you a useful reframe to carry into your next planning meeting. Your program does not have a testing capability problem. It has a resolution limit, just as a scale that reads to the nearest kilogram cannot weigh a letter, and the correct response is to change what you put on the scale.

The Detectable Change Rule

The rule states one thing: run a test only when the smallest outcome you would find interesting is larger than the noise in the metric you are measuring. Three components put it into practice, and none of them require a new platform.

Component 1: Size the change before anyone builds it

Before a page variant enters a sprint, the question on the table is how large a change this could plausibly produce if it works perfectly. Not how much you hope for; the ceiling.

Move a button, adjust a headline, tighten a paragraph, and the ceiling is small by construction. Remove a required account creation step, cut three fields from checkout, or change what a first-time visitor is asked to do, and the ceiling is high. Same engineering week, two entirely different odds of learning anything.

When a ceiling looks small, you can still ship the change; just make it a judgment call and label it as such, rather than spending six weeks pretending to measure it. That distinction is worth defending out loud, because a test log full of honest judgment calls is more credible in front of ownership than a test log full of results nobody believes.

Component 2: Run tests where the events are

Your surfaces do not generate events at anything like the same rate. A product page sees a modest number of sessions a week and a handful of orders. Your email service reaches your entire list on a set schedule, generates thousands of opens and clicks, and delivers results within days rather than quarters.

So the sequencing changes. Questions about copy, offer framing, subject construction, and send timing move to the email surface where they resolve quickly. Questions that can only be answered on the site are reserved for changes large enough to clear the resolution limit.

The timing lever is measurable, not theoretical. Triggered messages click near 5%, while batch sends run near 1.5% to 2%, and welcome or automated messages open in the 43% to 83% range, compared to roughly 31% for food and beverage campaigns generally (Klaviyo Email Benchmarks 2024; GetResponse Email Marketing Benchmarks 2024). Those are effect sizes you can see with the volume you already have.

Component 3: Choose documented causes over hunches

The strongest candidate list is not generated through brainstorming. It comes from the places where the loss is already known and quantified by somebody with a larger sample than yours.

Average cart abandonment sits near 70% across 50 studies, with a range of roughly 55% to 84%, and about 19% of abandonments cite having to create an account, making forced account creation one of the top avoidable causes rather than a matter of taste (Baymard Institute). Load speed is the other well-documented one: conversion peaks between one and two seconds and then degrades continuously, and a 0.1-second gain lifted retail conversion by 8.4% (Portent, 2019; Google and Deloitte, Milliseconds Make Millions, 2020). Note what is absent there: no threshold, no cliff, no magic second. Continuous degradation means every tenth of a second you remove is worth something, which is a far more useful brief for your developer than a target number.

The Objection You Will Hear Internally

Somebody will point out that the big changes are the risky ones, and that removing a required account step or restructuring checkout touches revenue on a live site during a quarter you are accountable for. That is a fair objection, and it deserves a real answer rather than reassurance.

The answer is staging and reversibility. A large change with a documented cause, shipped to a portion of traffic, with a one-click rollback and a named threshold at which to roll it back, is a smaller exposure than a year of small changes that were never measured and are all still live. Your site is currently carrying an accumulated stack of unmeasured decisions made by people who have since left. That is the actual risk position, and nobody in the building describes it that way.

What the Rule Produces

The first output is subtraction, and it will feel like less work rather than more. A long test backlog usually contains a handful of items that could never clear your resolution limit. Naming everything else as judgment calls and shipping it without ceremony recovers weeks of calendar in a quarter.

The second output is a defensible line in a review. “We ran a lot of tests” invites the question of what came of them. “We shipped the small changes as judgment calls and ran the tests that were sized to be measurable at our volume, and here is what each one settled” is a different conversation, and it is the one where your work survives contact with a skeptical owner.

The third output is quieter and takes a year to appear. Once a team’s sizing changes, honestly, the estimates themselves get better, because every rough ceiling written down in advance is eventually compared against a real outcome. After four quarters, you will know whether your group habitually overrates copy changes, underrates friction removal, or misjudges which surface a change belongs on. That calibration is not available on any dashboard, and it is the part of the practice that continues to pay after the individual findings have aged out.

The broader context supports spending the recovered time on conversion rather than on traffic. DTC shipments fell 15% in volume and 6% in value in 2025, the worst year in the report series, and the rise in average bottle price is explicitly a mix shift rather than buyers trading up (Sovos ShipCompliant and WineBusiness Analytics, DTC Wine Shipping Report 2026). Meanwhile, the spread between operators widened sharply: top-quartile wineries grew DTC revenue 22% while the median was flat and the bottom quartile fell 13% (Silicon Valley Bank, DTC Wine Report 2026). The gap between those groups is not explained by traffic volume, because everyone’s traffic is under the same pressure.

This Week’s Action

Take your current test backlog and add one column: the ceiling, meaning the largest result this change could produce if it worked perfectly. Fill it in from judgment, not from research; the exercise works even when the estimates are rough.

Then sort by that column and draw a line under the top three. Everything above the line is a test. Everything below it ships as a judgment call this month, with a one-line note in your log about why. You will have converted a stalled backlog into a short list of real questions and a long list of shipped decisions in an afternoon.

P.S. There is a second reason inconclusive tests are worth taking seriously rather than burying. Each one consumed a slot on a live surface that could have carried a change with a real ceiling, so the cost is not the wasted analysis; it is the six weeks of traffic spent answering a question that was never answerable. Wednesday’s email is about the opposite failure: the change that clearly worked, on a number that clearly moved, which you still cannot prove was responsible.