Skip to content
5 min read · A/B testing

A/B Testing Outsourcing: When It Works and What It Costs

MWritten byM R Q MotinFounder & CEO
Updated on 6 August 2026
A sprint board titled the build queue where one coral-outlined task bar labelled test build, slipped twice stretches across three sprint columns, with a chip reading feature work wins every sprint

Most testing programmes do not stall for lack of ideas. They stall in the queue: every experiment needs developer time, developer time belongs to the product roadmap, and the product roadmap wins every sprint, because features have owners and experiments have hopes. That queue, not strategy, is the real reason A/B testing outsourcing exists as a market.

We are on one side of that market, so read this the way you would read any supplier on its own category. It is the decision framework we actually use on sales calls, including the branches where the right answer is not us, because a client who should not have outsourced is a renewal that never comes.

The scale of the queue problem is worth stating plainly: in most of the stalled programmes we audit, the gap between hypothesis written and test live is measured in months, and nothing about the hypothesis caused the delay. The research was done, the design was approved, and the build sat behind four sprint reviews in a row while smaller feature tickets sailed past it. By the time the test shipped, the person who wrote the hypothesis had stopped believing anyone would ever run it.

What outsourcing actually buys

Three things, concretely. Throughput that does not negotiate with your roadmap: reserved build capacity means the test ships whether or not the feature does. Specialist craft: flicker engineering, cross-browser QA and tracking validation are learned on hundreds of builds, and a generalist team learns them on yours, expensively. And elasticity: testing volume is naturally spiky, heavy before peak season and quiet inside it, and buying capacity beats staffing for your busiest month.

Notice what is missing from that list: judgement about your customers. The question of what to test is downstream of research into your specific visitors, and that knowledge lives with whoever holds the relationship. Vendors that promise to own strategy and build and analysis from a standing start are promising to know your customers better than you do by week two. The wider engineering half of CRO transfers well; the why does not.

The three models, and who fits each

The brief deserves a sentence before the models, because it is where every one of them succeeds or fails. A good brief names the evidence, the change, the audience, the primary metric and the minimum effect worth detecting; it takes twenty minutes to write and saves a week of clarification. A screenshot with make this section better underneath is not a brief, and no engagement model survives a steady diet of them.

  • Keep it all in-house. Right at sustained high velocity, 15 or more tests a month, with a dedicated developer and analyst who are not shared with the roadmap. Below that cadence the fixed cost is paying for idle months.
  • Hybrid: your strategy, their build. The most common healthy shape. You own research, hypotheses and verdicts; a partner owns build, QA and implementation. Works because the interface between the two is a brief, which is writable.
  • Full outsource. Right when nobody internal owns conversion at all and the honest choice is between an external programme and no programme. Accept the trade: the vendor is learning your customers from a cold start, and the first quarter is tuition.
Two side-by-side panels titled keep it in-house when and bring in a partner when, listing fifteen plus tests a month, a dedicated developer and analyst and velocity needs against tests queueing behind features, QA bouncing and spiky volume

What good looks like in the first ninety days

A healthy engagement has a recognisable shape early. In the first fortnight, the partner asks more questions than they answer: your traffic numbers, your QA expectations, your brief format, who owns which decision. By week four, the first build has arrived with a QA document you did not have to request. By the end of the quarter, you can name the throughput, the QA bounce rate and the number of missed dates, because you kept your own scorecard rather than relying on theirs.

The unhealthy shape is just as recognisable: a fast yes to every brief, no pushback on unreadable tests, results that are always about to be written up, and a relationship that lives entirely in calls because nothing is documented. Six months of that costs more than the invoices show, because the programme has been standing still while appearing to move.

What it costs, and the red flags

The market ranges are covered properly in our agency cost guide: per-test builds from £150 to £900, reserved capacity from £2,500 to £6,500 a month, full service above that. The number matters less than what arrives with it. A quote produced without a single question about your traffic is a template. A vendor who cannot show you a real QA document has QA in the brochure, not the process. A guaranteed win rate is a guarantee that the statistics will be read generously. And a contract where the variation code is not yours at exit converts a supplier into a hostage negotiation.

The failure modes on the buyer side deserve equal billing: outsourcing as a substitute for research, briefs that are screenshots with feelings, and treating losing tests as vendor failure. Most of the ways tests fail before launch are decided in the brief, wherever the build happens.

Whatever the model, keep three things in-house permanently: the research relationship with your customers, the decision log of what was tested and why, and the analytics account. The first is your judgement, the second is your memory, and the third is your evidence. A vendor can borrow all three; the day one of them owns any of them, switching costs stop being a line item and become the relationship.

The honest decision path

Start from the constraint. If ideas are weak, buy research before capacity, from anyone. If throughput is the constraint and your volume justifies a hire who will stay busy year-round, hire; the maths crosses over around 15 tests a month, sustained. Between those, the hybrid is the default for a reason, and the way to try it carries no cleverness at all: send one real brief, pay for it, and judge the scope response, the build, the QA document and the communication against what you have today. If you understand what a well-run test involves, one pilot brief tells you nearly everything. Ours can be found at A/B test development, and the framework above still applies when the vendor being evaluated is us.

Sharein𝕏f

The questions people ask first

When ideas are not your constraint, throughput is: tests queue behind the product roadmap, QA keeps bouncing builds, and volume is spiky. Buying capacity fixes those three. It fixes nothing about weak research or an unwillingness to act on losing results.

Ownership of the question: what to test, why, and what the business does with the answer. Customer knowledge does not outsource well. The build, QA, tracking validation and implementation are the parts that transfer cleanly, because they are craft rather than context.

Guaranteed win rates, QA described but never documented, no questions about your traffic before quoting, silence about flicker, and code you do not own at exit. Any one of these predicts the relationship; two or more predict the ending.

It is the most common healthy shape we see. The internal team keeps the work it wants and knows best; the partner absorbs spikes, specialist builds and QA-heavy experiments. The failure mode is unclear ownership of tracking and implementation, so write that down first.

MWritten byM R Q MotinFounder & CEO

Motin founded Optyv and still sits in on the readouts. He is happiest when a test proves him wrong in public.

Do you like what you see?

Optimize your store