Most testing programmes do not stall for lack of ideas. They stall in the queue: every experiment needs developer time, developer time belongs to the product roadmap, and the product roadmap wins every sprint, because features have owners and experiments have hopes. That queue, not strategy, is the real reason A/B testing outsourcing exists as a market.
We are on one side of that market, so read this the way you would read any supplier on its own category. It is the decision framework we actually use on sales calls, including the branches where the right answer is not us, because a client who should not have outsourced is a renewal that never comes.
The scale of the queue problem is worth stating plainly: in most of the stalled programmes we audit, the gap between hypothesis written and test live is measured in months, and nothing about the hypothesis caused the delay. The research was done, the design was approved, and the build sat behind four sprint reviews in a row while smaller feature tickets sailed past it. By the time the test shipped, the person who wrote the hypothesis had stopped believing anyone would ever run it.
What outsourcing actually buys
Three things, concretely. Throughput that does not negotiate with your roadmap: reserved build capacity means the test ships whether or not the feature does. Specialist craft: flicker engineering, cross-browser QA and tracking validation are learned on hundreds of builds, and a generalist team learns them on yours, expensively. And elasticity: testing volume is naturally spiky, heavy before peak season and quiet inside it, and buying capacity beats staffing for your busiest month.
Notice what is missing from that list: judgement about your customers. The question of what to test is downstream of research into your specific visitors, and that knowledge lives with whoever holds the relationship. Vendors that promise to own strategy and build and analysis from a standing start are promising to know your customers better than you do by week two. The wider engineering half of CRO transfers well; the why does not.
The three models, and who fits each
The brief deserves a sentence before the models, because it is where every one of them succeeds or fails. A good brief names the evidence, the change, the audience, the primary metric and the minimum effect worth detecting; it takes twenty minutes to write and saves a week of clarification. A screenshot with make this section better underneath is not a brief, and no engagement model survives a steady diet of them.
- Keep it all in-house. Right at sustained high velocity, 15 or more tests a month, with a dedicated developer and analyst who are not shared with the roadmap. Below that cadence the fixed cost is paying for idle months.
- Hybrid: your strategy, their build. The most common healthy shape. You own research, hypotheses and verdicts; a partner owns build, QA and implementation. Works because the interface between the two is a brief, which is writable.
- Full outsource. Right when nobody internal owns conversion at all and the honest choice is between an external programme and no programme. Accept the trade: the vendor is learning your customers from a cold start, and the first quarter is tuition.

What good looks like in the first ninety days
A healthy engagement has a recognisable shape early. In the first fortnight, the partner asks more questions than they answer: your traffic numbers, your QA expectations, your brief format, who owns which decision. By week four, the first build has arrived with a QA document you did not have to request. By the end of the quarter, you can name the throughput, the QA bounce rate and the number of missed dates, because you kept your own scorecard rather than relying on theirs.
The unhealthy shape is just as recognisable: a fast yes to every brief, no pushback on unreadable tests, results that are always about to be written up, and a relationship that lives entirely in calls because nothing is documented. Six months of that costs more than the invoices show, because the programme has been standing still while appearing to move.
What it costs, and the red flags
The market ranges are covered properly in our agency cost guide: per-test builds from £150 to £900, reserved capacity from £2,500 to £6,500 a month, full service above that. The number matters less than what arrives with it. A quote produced without a single question about your traffic is a template. A vendor who cannot show you a real QA document has QA in the brochure, not the process. A guaranteed win rate is a guarantee that the statistics will be read generously. And a contract where the variation code is not yours at exit converts a supplier into a hostage negotiation.
The failure modes on the buyer side deserve equal billing: outsourcing as a substitute for research, briefs that are screenshots with feelings, and treating losing tests as vendor failure. Most of the ways tests fail before launch are decided in the brief, wherever the build happens.
Whatever the model, keep three things in-house permanently: the research relationship with your customers, the decision log of what was tested and why, and the analytics account. The first is your judgement, the second is your memory, and the third is your evidence. A vendor can borrow all three; the day one of them owns any of them, switching costs stop being a line item and become the relationship.
The honest decision path
Start from the constraint. If ideas are weak, buy research before capacity, from anyone. If throughput is the constraint and your volume justifies a hire who will stay busy year-round, hire; the maths crosses over around 15 tests a month, sustained. Between those, the hybrid is the default for a reason, and the way to try it carries no cleverness at all: send one real brief, pay for it, and judge the scope response, the build, the QA document and the communication against what you have today. If you understand what a well-run test involves, one pilot brief tells you nearly everything. Ours can be found at A/B test development, and the framework above still applies when the vendor being evaluated is us.



