Skip to content
6 min read · A/B testing

The Experiment Archive: Documenting Tests So Learnings Compound

HWritten byHusbeyBusiness Analyst & QA
Updated on 24 June 2026
A filled experiment archive entry for a collection page filter test showing the hypothesis and its evidence, a 21-day window, revenue per visitor as the primary metric, a result of plus 2.8% at 91% confidence, a decision of not shipped, and a one-sentence learning field

Somewhere in your company, somebody is about to propose a test that already lost two years ago. The cost of that is not the test. It is proof that the programme has no memory, and a programme without memory pays full price for the same lesson twice, with the second bill usually arriving from a different agency.

The operating system around this, the owner, the cadence and the backlog that make a programme survive its first quarter, is covered in our guide to building a CRO programme, and the archive is one of its four parts. This post goes down a level into the artefact itself: what a single entry has to contain, and the retrieval habit without which the whole thing is a filing cabinet nobody opens.

The entry: twelve fields

One row per test, twelve fields, filled in during the readout rather than as homework afterwards. The list looks long and takes about eight minutes when the test is fresh, which is the only moment any of it is cheap to write.

  • Title and date range. A plain name a stranger could search for, plus exact start and end dates, because a result you cannot place in the calendar cannot be checked against whatever else was happening that fortnight.
  • The hypothesis exactly as it was launched, not a tidied version written with the benefit of the result.
  • The evidence behind it: the recording, the ticket volume, the funnel step, the earlier test. An entry without this cannot be judged later, because you cannot tell a reasoned bet from a guess.
  • Surface and segment: which template, which devices, which traffic sources, and who was deliberately excluded.
  • Primary metric and the sample plan agreed before launch, so a future reader can see whether the test ran to its plan or to somebody’s patience.
  • Guardrails with their thresholds, and whether any of them went red during the run.
  • The result: relative change, confidence, and the absolute baseline it moved from. All three, because a percentage with no baseline is unusable a year later.
  • The decision and who made it: ship, iterate, kill or hold. Decisions without names get relitigated.
  • Build notes: what the variation actually did in code, plus any defect found mid-run that might have affected the reading.
  • Screenshots of control and variant. Six months on, nobody can reconstruct a variation from a written description, and this is the field teams skip and regret.
  • Tags, drawn from a fixed vocabulary rather than typed freehand.
  • One sentence of learning, written for a stranger who was not in the room.

Who fills it in matters as much as what goes in. The person who built the variation writes the build notes, and the person who called the result writes the decision, because an archive assembled afterwards by whoever had time is an archive of summaries rather than facts. Those eight minutes belong inside the readout, before anybody leaves the call. Entries written a fortnight later are shorter, vaguer and quietly reshaped by whatever happened next, which is the failure mode nobody notices because the archive still looks full.

The learning sentence is the asset

Eleven of those fields are record-keeping. The twelfth is the only one that compounds, and it is the one most often written badly. The rule is that the sentence must be a claim about shoppers, not a report about the test. The variant lost is not a learning. Removing the delivery estimate from the buy box costs more on mobile than the tidier layout gains is a learning, because it is portable: it applies to the next template, the next redesign and the next person who suggests the same tidy-up. Test names decay within a year. Claims about how your customers behave keep paying.

Inconclusive results need the same treatment and rarely get it. A collection page filter test that ran 21 days and returned +2.8% on revenue per visitor at 91% confidence is not a winner and not nothing. It goes in as not shipped, with a learning that says the effect, if real, is smaller than this store can detect in three weeks, so the next attempt needs a bigger change rather than a longer run. That sentence saves somebody a month.

An experiment archive search filtered to the tags product page and urgency, returning six past entries with dates, results and decisions, two shipped, one flat and three killed, with the top row expanded to show its one-sentence learning

The retrieval ritual

An archive that is only ever written to is a cost. The habit that turns it into an asset takes ten minutes and attaches to a moment that already exists: before any hypothesis is written, search the archive by surface and mechanism. Searching product page and urgency on a mature archive might return six entries, two shipped, one flat and three killed. Three outcomes follow. The idea is already answered, so it comes off the backlog. It was answered in a different context, on desktop or before a redesign, so it gets rewritten to say what has changed since. Or nothing comes back, and the idea proceeds with the search itself recorded in the evidence field.

Making that visible is what keeps it alive. Add an archive checked line to the hypothesis template so a blank one is obvious to whoever reviews the backlog. A rule that adds a field to a document people already fill in survives contact with a busy quarter. A rule that adds a meeting does not.

Tagging so retrieval actually works

Free-text tags rot inside a quarter, because three people will type urgency, scarcity and countdown for the same idea and none of them will find the others. Use three fixed axes and nothing else. Surface: home, collection, product page, cart, checkout, account. Mechanism: information, urgency, trust, price, navigation, speed, personalisation. Outcome: shipped, flat, killed. Every entry carries exactly one value per axis, and anything that does not fit lives in the prose fields where search will still find it. Twenty controlled values beat two hundred organic ones, and the constraint is what makes the search above return six entries instead of nothing.

Where to keep it, and the second job it does

The tool matters less than teams expect. Notion, Airtable, a spreadsheet or a purpose-built experimentation platform will all work; a folder of slide decks will not, because decks are unsearchable and each one is somebody’s summary rather than a record. Four requirements decide it: full-text search across the learning sentences, filtering on the three tag axes, a stable link per entry so a readout can point at one, and a clean export. That last one is not a nicety. Your archive should outlive your testing tool, your analytics stack and your agency, and portability is the only thing that guarantees it.

The second job is one almost nobody plans for: onboarding. A new designer, developer or marketer who reads forty entries in an afternoon arrives with priors about your customers that would otherwise take them a year and several avoidable arguments to acquire. Our own archive holds 1640+ ecommerce experiments, and its real value is not the winners, it is that every engagement starts from a set of questions somebody has already answered. If you want that discipline running on your programme, along with the engineering that produces entries worth keeping, get in touch and bring whatever record you already have.

Sharein𝕏f

The questions people ask first

An experiment archive is the written record of every test a team has run, winners and losers alike, with one row per experiment and one sentence of learning on each. It is the asset a testing programme accumulates, and the first thing lost when staff or agencies change.

Twelve fields cover it: title, dates, hypothesis as launched, the evidence behind it, surface and segment, primary metric and sample plan, guardrails, result with confidence and baseline, decision and decision-maker, build notes, screenshots of both arms, tags, and one sentence of learning.

Especially those. A losing test is a fence, and it is the only thing that stops the same idea returning with fresh enthusiasm every eighteen months. An archive holding only winners is a sales deck, not a record.

Anywhere that gives you full-text search, tag filters, a stable link per entry and a clean export. A table in Notion, Airtable or a spreadsheet beats a folder of slide decks, and portability matters more than features because the archive should outlive the tool.

Attach retrieval to a moment that already exists: no hypothesis is written until the archive has been searched, and the search itself gets recorded on the hypothesis. A rule that adds a step to an existing ritual survives; a rule that adds a meeting does not.

HWritten byHusbeyBusiness Analyst & QA

Husbey handles analysis and QA at Optyv, and is the person who asks whether the number actually says what the deck claims it says.

Do you like what you see?

Optimize your store