Skip to content

Store audits

A store audit your team can actually build from

We walk your store the way your customers meet it, on the hardware and in the states they actually arrive in, and read it against published usability heuristics and WCAG 2.2 at level AA.

What lands is a ranked roadmap rather than a list of observations: every finding carries the evidence it came from, the template and state it lives on, and an honest view of what fixing it costs.

  • Read against WCAG 2.2 at level AA
  • Ranked on your evidence, not our taste

The problem

Why most store audits change nothing

Almost every merchant has been sent an audit. Very few have acted on one. The document is rarely wrong, exactly. It is unusable, and usually for one of four reasons.

01

A checklist wearing a report

Best practices applied in order, with no evidence that any of them is costing this store anything. Every line is defensible and none of it is prioritised, so the list gets skimmed once.

02

Findings nobody owns

A problem stated without the template it lives on, the state it appears in, or who would fix it. It cannot become a ticket, so it never becomes a change.

03

Ranked by the auditor’s taste

Severity assigned by how much the reviewer disliked it. The loud cosmetic issue outranks the quiet one on the step that every order passes through.

04

Written from a desk, at one width

Reviewed on a large screen, signed in, on a fast connection, with an empty cart. Most of your customers meet the store in none of those states.

An audit is only worth what gets built from it. So the unit of work here is not the finding, it is the finding with its evidence, its location and its rank attached.

What gets read before anything is judged

Opinion is cheap and a store is specific, so the reading comes first. Four sources, and the interesting findings almost always sit where two of them disagree.

Your analytics say where people stop. Session recordings say what they were doing when they stopped, which is usually a different story than the funnel implies. The store itself, walked as a customer rather than as an owner, says what is actually on the screen at each of those moments. And your support inbox says what people gave up and asked a human instead, which is the only one of the four that arrives already written in the customer’s own words.

The walking is the part most audits skip, and it is deliberately awkward here. Signed out and signed in. On a mid-range Android handset, not a flagship. With one item in the cart and with nine. On a product that is out of stock in the size most people want. In the currency and to the country your second-biggest market actually orders from. A store behaves differently in each of those states, and a report written from only the first one is describing a shop nobody visits.

  • Analytics: where sessions end, by template, device and traffic source
  • Session replay: what happened in the last thirty seconds before they ended
  • The live store: walked in the states above, on real hardware
  • Support tickets and on-site search: the questions the interface failed to answer

Heuristics, and why a checklist is not an audit

The evaluation framework is the published one: Nielsen’s ten usability heuristics, which have been the standard vocabulary for this work since 1994 and are not ours to reinvent. What matters is not the list, it is that a heuristic is a lens, not a rule. It tells you where to look. It does not tell you whether what you found is costing you anything, and an audit that stops at the lens has done half the job.

So each one is read against a place your store actually has, rather than in the abstract. The table below is the translation, and it is also roughly the running order: visibility problems tend to be cheap and high-reach, while recovery problems tend to be the expensive ones nobody has looked at.

Four usability heuristics, what each one looks like on an ecommerce store, and the template it is usually found on.
The heuristicWhat it looks like on a storeUsually on
Visibility of system statusWhether an add to cart, an applied discount or a shipping estimate visibly confirms itself, or whether the customer has to guess it workedProduct page, cart drawer
Match with the real worldWhether sizes, materials, delivery windows and returns are described in the words a customer would use rather than the ones your PIM storesProduct page, delivery info
Recognition over recallWhether a customer four steps into checkout can still see what they are buying, what it costs delivered, and what changed since the cartCheckout, cart
Error prevention and recoveryWhat happens on a declined card, an undeliverable postcode, an out-of-stock variant or a mistyped email, and whether the entered data survives itCheckout, forms

Accessibility is read in the same pass rather than filed as a separate concern, against WCAG 2.2 at level AA, which is the standard our design work already commits to. The overlap with usability is not a coincidence: contrast that fails a ratio is also contrast nobody can read on a phone outdoors, and a focus order that strands a keyboard user is usually an order that confuses everyone. Findings that are purely conformance are labelled as such, so a legal obligation is never smuggled in as a conversion argument.

What we build

What the audit covers

The journey, end to end

  • Landing, collection, search and filtering, read as the paths traffic actually takes rather than the one the navigation implies
  • Product page: the questions a buyer has to answer before they can commit, and which of them your page leaves open
  • Cart, delivery expectation and checkout, walked to a real declined-card and out-of-stock state

Mobile as the primary case

  • Reviewed on mid-range hardware and a throttled connection, because that is the median customer rather than the median reviewer
  • Thumb reach, tap target size and what the on-screen keyboard covers at the moment it opens
  • What the page looks like part-loaded, which on a slow connection is what most people actually see

Accessibility, in the same pass

  • WCAG 2.2 AA read alongside usability rather than filed as a separate compliance exercise
  • Keyboard path, focus visibility and focus order through the flows that take money
  • Findings labelled where they are conformance obligations rather than conversion arguments

The trust layer

  • Whether delivery cost and date are knowable before checkout, which is the single most common cause of a cart nobody comes back to
  • Returns, sizing and stock honesty, read as the objections they are rather than as content
  • Consent and cookie interruption, judged on what it does to a first-time visitor mid-task

How a finding earns its rank

This is the part that decides whether the document is usable, and it is the part most audits leave to instinct. A finding is ranked on three things, stated separately so you can disagree with any one of them without discarding the whole rank.

  • Reach: how many sessions actually meet it, read off your own analytics rather than assumed. A flaw on a template three per cent of traffic sees is a real flaw and a low rank.
  • Severity: what it costs the customer who meets it. Blocked outranks confused, confused outranks ugly, and a thing that silently loses entered data outranks all three.
  • Cost to fix: a theme setting, a template edit, a design change, or a change that needs engineering. Stated because a roadmap nobody can afford is a roadmap nobody runs.

Two things follow from ranking this way, and both are worth saying out loud. The first is that cheap and boring wins a lot: the top of a roadmap is usually two theme settings and a piece of copy, not a redesign. The second is that some findings arrive with an honest verdict of leave it alone. An audit that recommends everything it noticed is not an audit, it is a list of everything the reviewer noticed.

What a rank is not is a revenue forecast. You will not find a predicted percentage attached to a finding here, because attaching one to a heuristic observation is arithmetic dressed as evidence. Where a finding is worth proving rather than assuming, it is written as a testable hypothesis and handed across, and where the traffic cannot support a test, the audit says so instead of prescribing one.

What actually lands

A roadmap your team can run without us, in the order it should be run. Each finding carries the same five things, because a finding missing any of them is one somebody has to reconstruct before it can become work.

  • The evidence: the recording, the screen or the number it came from, not a description of it
  • The location: the template, the state and the breakpoint where it appears
  • The rank, with reach, severity and fix cost shown separately rather than blended into one score
  • The recommendation, specific enough to brief: what changes, not that something should
  • Where it is uncertain, the test that would settle it, written as a hypothesis

It is yours, whoever runs it. Some clients take the roadmap to their in-house designers, some to another studio, and some ask us to design the fixes, which is the parent service rather than this one. A roadmap that only works if we build it is a sales document, and we would rather hand over something you can act on without us than something that makes us necessary.

How it runs

How the audit runs

Read-only access is enough to start, and nothing on your store changes at any point in it.

  1. 1

    Scope and access

    We agree which templates and which markets are in, and get read access to analytics and session replay. Scoping matters more than it sounds: an audit spread evenly across forty templates is thinner everywhere than one aimed at the four that carry the orders.

    • Templates and markets agreed up front
    • Read-only analytics and replay access
    • The states to walk, named before we start
  2. 2

    Evidence

    Where sessions end, by template and device. Enough replay to see patterns rather than anecdotes. Support tickets and on-site search read for the questions the interface failed to answer.

    • Exit points read per template, not site-wide
    • Replay watched to pattern, not to one clip
    • Search and ticket language collected
  3. 3

    Evaluation

    The store walked in the customer states agreed at scoping, on real hardware, against the heuristics and WCAG 2.2 AA. Each finding is captured with its evidence attached at the moment it is found, rather than reconstructed afterwards.

    • Signed out and signed in, mobile and desktop
    • Real declined-card and out-of-stock paths walked
    • Every finding captured with its evidence
  4. 4

    Ranking and readout

    Findings ranked on reach, severity and fix cost, then walked through live with the people who would build them. The call is where a rank gets argued with, which is the point of holding one.

    • Reach, severity and cost shown separately
    • Walked through with the team who would build it
    • The roadmap is yours, whoever runs it

This audit, or the CRO audit?

We publish two, they are genuinely different, and the wrong one is a waste of everybody’s time. The difference is what gets read and what comes out, not which is more thorough.

The UX audit and the CRO audit compared by what each one reads, what it produces and when to choose it.
This UX auditThe CRO audit
ReadsThe interface, walked in real customer states, against usability heuristics and WCAG 2.2Your analytics, sessions and order data, reconciled against each other
ProducesA ranked design roadmap: what to change and in what orderA prioritised test backlog: what to prove and in what order
Best whenYou suspect the experience is the problem, or a redesign is being argued about on tasteYou have enough traffic to test and want to know what to test first
Traffic neededNone. Heuristics do not need a sample sizeEnough to run experiments, which not every store has

If you are below roughly a thousand orders a month, the honest answer is usually this one: at that volume a testing programme spends most of its time waiting, while the interface problems are readable today without a sample size. Above it, the two work well in sequence, and the free CRO audit is the cheaper way to find out which of them you needed. If you already know the answer is testing, A/B test development is where that goes.

What happens after the roadmap

Nothing has to. The roadmap is a finished piece of work, and plenty of them get run entirely in-house, which is the outcome we design the document for.

When a client does want the fixes designed rather than described, that is our UX design service, the parent of this page, and it is scoped separately after the readout rather than assumed on the way in. Where the roadmap concentrates on the two templates that carry the money, PDP and checkout design is the narrower engagement; where the findings are mostly the same inconsistency repeated across forty pages, that is usually a design system problem wearing forty costumes, and no amount of page-by-page fixing will hold it.

The answers to your questions.

A UX audit reads the interface and produces a design roadmap; a CRO audit reads your data and produces a test backlog. This one walks your store in real customer states against usability heuristics and WCAG 2.2, and needs no traffic volume to be useful. The CRO audit reconciles analytics against orders and needs enough traffic to support experiments. Below roughly a thousand orders a month the UX audit is usually the one that pays, because the interface problems are readable today while a testing programme would spend most of its time waiting.

Read-only access is enough, and nothing on your store changes at any point. We ask for analytics and session replay access to size how many people actually meet each problem, because reach is what separates a real finding from a defensible opinion. Without that access an audit is still possible, but every rank in it becomes an estimate, and we would tell you that rather than present it as measured.

On three things, shown separately: reach, how many sessions actually meet it, read off your own analytics; severity, what it costs the customer who does; and cost to fix, from a theme setting to a change that needs engineering. They stay separate so you can disagree with one without discarding the rank. What you will not find is a predicted revenue percentage attached to a heuristic observation, because that is arithmetic dressed as evidence.

Yes, in the same pass, against WCAG 2.2 at level AA rather than as a separate compliance exercise. Usability and accessibility overlap more than they differ: contrast that fails a ratio also fails a customer on a phone outdoors, and a focus order that strands a keyboard user usually confuses everyone. Findings that are purely conformance obligations are labelled as such, so a legal requirement is never presented to you as a conversion argument.

No, and be wary of one that does. A heuristic observation cannot support a revenue forecast, and a number invented to make a roadmap persuasive is the fastest way to lose the roadmap’s credibility at the first fix that does not deliver it. What the audit gives you instead is reach and severity, both evidenced, so you can judge the size yourself. Where a finding genuinely needs proving rather than assuming, it is written as a testable hypothesis and handed across to testing.

No. The roadmap is written to be run by whoever has your design capacity, and plenty are run entirely in-house. Some clients take it to their own team, some to another studio, and some ask us to design the fixes, which is our UX design service and is scoped separately after the readout rather than assumed on the way in. A roadmap that only works if we build it would be a sales document, not an audit.

Been sent an audit nobody could act on?

Tell us your store, the template you most suspect and what your team can realistically build. We will say what we would read first, and whether an audit is the thing you need at all.

Tell us what your store is doing