"Let us test a new headline" is not a hypothesis. It is a coin flip with extra steps and a fortnight of traffic attached. A real hypothesis names the evidence that prompted it, the change it proposes, the number that should move, and the result that would prove the whole idea wrong. Most test backlogs contain none of that, which is why most readouts turn into a discussion about what the test was really about.
This post owns one artefact: the hypothesis document itself. Not the wider process of running a programme, and not the post-mortem of why results disappoint, both of which live elsewhere. Just the paragraph you write before anything gets built, and the test that tells you whether what you wrote qualifies.
The template
Fill this in, in this order, in one paragraph: Because [evidence], we believe [change] will cause [effect] for [audience], measured by [primary metric]. We will know we are wrong if [failure condition].
Six slots, and each one is a trap for a different bad habit. Evidence first, because a hypothesis that opens with the change is a preference with grammar around it, and the evidence slot is where opinions quietly fail to qualify. Audience matters because a change that helps returning mobile buyers and hurts first-time desktop visitors nets to nothing, and you want to have said which group you meant before the segmentation argument starts. The primary metric is singular by design: two primaries is zero primaries, and the second one exists so somebody can declare victory later. And the failure condition, the slot almost every template omits, is what converts a story into an experiment.
The falsifiability test
Here is the check, and it takes ten seconds. Read your hypothesis, then write down the result that would make you abandon the belief behind it rather than the execution of it. If you cannot name one, you do not have a hypothesis. You have a plan to gather evidence for something you have already decided.
The tell is easy to spot once you know it. A team with a real hypothesis says "if trust badges under the buy button do not lift add-to-cart rate by at least 3%, we stop treating badge placement as a lever". A team without one says "if it does not win, we will try a different badge design", which is the same belief wearing a new outfit and a second fortnight of traffic. Both sentences sound like learning. Only the first one can lose.
There is a second-order benefit here that is easy to miss. A falsifiable hypothesis retires a whole class of ideas when it loses, not just the one variation that was built. Kill badge placement properly and the six badge variations queued behind it leave the backlog with it, which is worth more than most winners: a testing programme is limited by how many questions it can ask per quarter, and the fastest way to ask better ones is to stop asking questions you have already answered.
Three rewrites
These are the shapes we see most often in client backlogs, alongside what they become once the template has been applied to them:
- Before: "Test a sticky add-to-cart bar on mobile." After: "Because 34 session recordings show mobile visitors scrolling back up to find the buy button, we believe a sticky add-to-cart bar will lift add-to-cart rate by 4% for first-time mobile visitors, measured on add-to-cart rate. We are wrong if the lift is under 2% or checkout starts fall."
- Before: "Our product pages need better copy." After: "Because 41% of pre-purchase support tickets ask about fit, we believe adding a fit summary above the fold will lift revenue per visitor by 3% for first-time visitors on apparel pages. We are wrong if fit tickets do not fall alongside it, which would mean the copy was not read."
- Before: "Remove a checkout step." After: "Because analytics shows 22% of visitors abandoning at the account step, we believe offering guest checkout will lift checkout completion by 6% for new customers. We are wrong if completion rises while revenue per visitor does not, which would mean we moved abandonment rather than removed it."
Notice what each rewrite costs: about ninety seconds, and one specific commitment you can be held to. Notice what it buys: a developer who can build without a meeting, an analyst who knows which number grades the test, and a readout that cannot be relitigated by whoever is least happy with the answer. If the evidence slot is where your backlog stalls, our list of test ideas ranked by win rate is a reasonable place to borrow priors from.

The metric is part of the hypothesis
A hypothesis that names its metric after the results arrive is not a hypothesis, and this is the most common way a well-written one still fails. Choose the metric while you are writing, choose one, and choose the one closest to money that your traffic can actually resolve. Conversion rate is convenient and gameable; revenue per visitor is harder to move and much harder to argue with. Whichever you pick, write down the guardrails in the same breath, because a hypothesis that predicts a lift somewhere should also predict what must not fall.
From hypothesis to backlog
A hypothesis written this way prioritises itself, which is the quiet benefit. The evidence slot gives you confidence, the effect slot gives you expected impact, and the change slot gives you build effort, so scoring a backlog stops being a workshop and becomes a sort. Weak hypotheses cannot be scored at all, which is useful information about them: an idea that resists scoring is usually one nobody has found evidence for yet, and it belongs in research rather than in a build queue.
Keep the document as one paragraph and keep it with the test for its whole life, from backlog row to build brief to readout slide. The version that matters is the one written before launch, so store it somewhere it cannot be quietly improved afterwards. Hypotheses that get edited once the numbers are in stop being predictions and become captions, and a programme full of captions cannot tell you what it has learned.
The habit is cheap and the discipline is not. Teams that write hypotheses like this argue less about results, because the arguments happened up front while they were still cheap, and they retire beliefs rather than accumulating them. If you would rather have someone else hold the pen and the build capacity, experiment engineering is the service that covers both, and we are easy to reach.



