The experiment archive template
Somewhere in your company, someone is about to re-run a test that already lost. The cost of that is not the test, it is proof the programme has no memory. This is the record we keep per experiment, the vocabulary that makes it searchable, and the habit that makes it pay.
The template
The entry: twelve fields
One row per test, filled in during the readout rather than as homework afterwards. It looks long and takes about eight minutes when the test is fresh, which is the only moment any of it is cheap to write.
| Field | What goes in it |
|---|---|
| Title and dates | Title and date range. A plain name a stranger could search for, plus exact start and end dates, because a result you cannot place in the calendar cannot be checked against whatever else was happening that fortnight. |
| Hypothesis as launched | The hypothesis exactly as it was launched, not a tidied version written with the benefit of the result. |
| Evidence | The evidence behind it: the recording, the ticket volume, the funnel step, the earlier test. An entry without this cannot be judged later, because you cannot tell a reasoned bet from a guess. |
| Surface and segment | Surface and segment: which template, which devices, which traffic sources, and who was deliberately excluded. |
| Metric and sample plan | Primary metric and the sample plan agreed before launch, so a future reader can see whether the test ran to its plan or to somebody’s patience. |
| Guardrails | Guardrails with their thresholds, and whether any of them went red during the run. |
| Result | The result: relative change, confidence, and the absolute baseline it moved from. All three, because a percentage with no baseline is unusable a year later. |
| Decision and owner | The decision and who made it: ship, iterate, kill or hold. Decisions without names get relitigated. |
| Build notes | Build notes: what the variation actually did in code, plus any defect found mid-run that might have affected the reading. |
| Screenshots | Screenshots of control and variant. Six months on, nobody can reconstruct a variation from a written description, and this is the field teams skip and regret. |
| Tags | Tags, drawn from a fixed vocabulary rather than typed freehand. |
| Learning | One sentence of learning, written for a stranger who was not in the room. |
Three filled entries
Illustrative entries, not client records. The first is the inconclusive collection-filter test published in our article on experiment documentation. The second and third are composites built to show the format, so their dates, sample sizes and result figures are made up rather than taken from anyone’s programme.
| Field | Collection filter panel | Delivery estimate in the buy box | Guest checkout |
|---|---|---|---|
| Title and dates | Collection filter panel, 3 to 24 March | Delivery estimate removed from buy box, 7 to 28 April | Guest checkout at the account step, 5 to 26 May |
| Hypothesis as launched | Because visitors on collection pages scroll past the third row without narrowing, we believe a persistent filter panel will lift revenue per visitor by 5% for all collection traffic. We are wrong if revenue per visitor moves less than 3%. | Because the buy box tests badly on mobile for density, we believe removing the delivery estimate will lift add-to-cart rate by 3% for mobile visitors. We are wrong if revenue per visitor falls. | Because 22% of visitors abandon at the account step, we believe offering guest checkout will lift checkout completion by 6% for new customers. We are wrong if completion rises while revenue per visitor does not. |
| Evidence | Scroll-depth analysis plus 18 recordings of visitors leaving collection pages without using a filter. | Five moderated mobile sessions where the buy box was described as cluttered. | Funnel analytics showing the account step as the largest single drop in checkout. |
| Surface and segment | Collection templates, all devices, all sources. Paid landing pages excluded. | Product templates, mobile only. Tablet excluded. | Checkout, new customers only. Returning logged-in customers excluded. |
| Metric and sample plan | Revenue per visitor. 21 days planned, 21 days run. | Add-to-cart rate. 21 days planned, 21 days run. | Checkout completion. 21 days planned, 21 days run. |
| Guardrails | Add-to-cart rate must not fall more than 2%. Not breached. | Revenue per visitor must not fall. Breached in week two and again in week three. | Revenue per visitor must not fall. Not breached. |
| Result | +2.8% on revenue per visitor at 91% confidence, from a baseline of £1.94. | +2.1% on add-to-cart rate at 96% confidence, with revenue per visitor down 3.4%. | +5.7% on checkout completion at 98% confidence, from a baseline of 61.2%. |
| Decision and owner | Not shipped. Called at the readout by the programme lead. | Killed. Called at the readout by the programme lead. | Shipped. Called at the readout by the programme lead. |
| Build notes | Client-side filter panel, no template change. No defects found during the run. | Single element hidden on mobile breakpoints. No defects found during the run. | New checkout step configuration. A flicker on slow connections was fixed on day two and the first day was excluded. |
| Screenshots | Control and variant, mobile and desktop, attached. | Control and variant, mobile only, attached. | Control and variant, mobile and desktop, attached. |
| Tags | Collection · navigation · flat | Product page · information · killed | Checkout · navigation · shipped |
| Learning | The effect, if real, is smaller than this store can detect in three weeks, so the next attempt needs a bigger change rather than a longer run. | Removing the delivery estimate from the buy box costs more on mobile than the tidier layout gains. | Requiring an account before payment costs this store more new customers than the account data is worth. |
Tagging so retrieval actually works
Free-text tags rot inside a quarter, because three people will type urgency, scarcity and countdown for the same idea and none of them will find the others. Three fixed axes and nothing else, exactly one value per axis per entry. Anything that does not fit lives in the prose fields, where search will still find it.
- Surface
- home
- collection
- product page
- cart
- checkout
- account
- Mechanism
- information
- urgency
- trust
- price
- navigation
- speed
- personalisation
- Outcome
- shipped
- flat
- killed
Take it with you
The Word file carries the twelve fields, the three filled entries and the tag vocabulary, ready to be pasted into whatever you keep records in. No address needed.
Download Word (.docx)Generated from this page, so the two cannot disagree.
The PDF is the printable version, useful as a one-page reference next to a readout. It is the only thing here that asks for an address.
How we use it
Who fills it in matters as much as what goes in. The person who built the variation writes the build notes, and the person who called the result writes the decision, because an archive assembled afterwards by whoever had time is an archive of summaries rather than facts. Those eight minutes belong inside the readout, before anybody leaves the call. Entries written a fortnight later are shorter, vaguer and quietly reshaped by whatever happened next, which is the failure mode nobody notices because the archive still looks full.
Eleven of the fields are record-keeping. The twelfth is the only one that compounds, and it is the one most often written badly. The sentence has to be a claim about shoppers, not a report about the test. "The variant lost" is not a learning. "Removing the delivery estimate from the buy box costs more on mobile than the tidier layout gains" is, because it is portable: it applies to the next template, the next redesign and the next person who suggests the same tidy-up.
An archive that is only ever written to is a cost. The habit that turns it into an asset takes ten minutes and attaches to a moment that already exists: before any hypothesis is written, search by surface and mechanism. Either the idea is already answered and comes off the backlog, or it was answered in a different context and gets rewritten to say what has changed, or nothing comes back and the search itself goes in the evidence field. Adding an "archive checked" line to the hypothesis template is what keeps the habit alive, because a rule that adds a field to a document people already fill in survives a busy quarter and a rule that adds a meeting does not.
The tool matters less than teams expect. Four requirements decide it: full-text search across the learning sentences, filtering on the three tag axes, a stable link per entry, and a clean export. That last one is not a nicety, because your archive should outlive your testing tool, your analytics stack and your agency. The reasoning, and the second job an archive does at onboarding, is in the full article on documenting tests so learnings compound.
Version
- v1.0 · 2026-08-07First published. Twelve fields, three fixed tag axes, three filled entries.
Questions about the archive
Yes, and they are the entries most often skipped. A test that returned a small positive at 91% confidence is not a winner and not nothing. It goes in as not shipped, with a learning that says the effect, if real, is smaller than the store can detect in three weeks, so the next attempt needs a bigger change rather than a longer run. That sentence saves somebody a month.
Because free-text tags rot inside a quarter. Three people will type urgency, scarcity and countdown for the same idea and none of them will find the others. Three axes with about twenty controlled values between them beat two hundred organic ones, and the constraint is the whole reason a search before writing a hypothesis returns anything at all.
Less important than teams expect. Notion, Airtable, a spreadsheet or a purpose-built platform all work; a folder of slide decks does not, because decks are unsearchable and each one is somebody’s summary rather than a record. Four things decide it: full-text search across the learning sentences, filtering on the three axes, a stable link per entry, and a clean export.
Only partly, and the page says which. The first is the inconclusive collection-filter test published in our article on experiment documentation, numbers included. The second and third are composites written to demonstrate the format: their subjects come from our published writing, but the dates, sample sizes and result figures are made up rather than taken from anyone’s programme.
Using this on your own site
The template is free to quote, screenshot and link to. A credit link keeps it findable for the next team trying to work out what to record per test.
<p>Experiment archive template by <a href="https://optyv.com/templates/experiment-archive">Optyv</a></p>No tracking in either snippet, and no nofollow on the link. It is a credit, not a beacon.
Or use the badge
For a team wiki, a process page or the front of the archive itself, to show which record shape it keeps.
<a href="https://optyv.com/templates/experiment-archive"><img src="https://optyv.com/badge/experiment-archive-used.svg" alt="Experiments archived to the Optyv template" width="250" height="40"></a>An archive is worth more than its winners
Ours holds 1640+ ecommerce experiments, and its real value is that every engagement starts from a set of questions somebody has already answered. If you want that discipline running on your programme, bring whatever record you already have.
Talk to us ↗