Every best A/B testing tools list you have read was written by a tool vendor or an affiliate. We are neither: we build tests in these platforms every week for client stores, we hold no reseller deals, and the links below earn us nothing. What follows is how the platforms actually compare when the invoice is your own.
One framing before the list. The tool decides how you inject variations, target audiences and read results. It does not decide whether your hypotheses are good, your builds are clean or your tracking is honest, and those three decide the programme. We have watched strong teams win on cheap tools and expensive tools host eighteen months of nothing.
How we judged them, since a comparison is only as honest as its criteria: the build experience for a developer writing real variation code, not the visual editor demo; how each handles flicker when configured by someone who knows to ask; whether the statistics engine is trustworthy at the sample sizes ecommerce actually has; how well it behaves on a Shopify theme full of app scripts; and whether you can discover the price without a sales call. Weightings are ours, and your backlog might weight differently.
The list is seven rather than seventeen on purpose. The long-tail platforms exist and some are good, but a shortlist you can actually evaluate beats a directory you can only scroll, and these seven cover every programme shape we meet in ecommerce, from a first template split on a standard Shopify plan to a server-side programme spanning pricing and search on a headless stack.
The seven worth knowing, honestly
- Optimizely. The enterprise reference: deep server-side SDKs, feature flagging, mature stats. What we like: it handles complex programmes without creaking. Friction we hit: the price of entry makes sense only when you will actually use that depth, and most ecommerce programmes never do.
- VWO. The broadest mid-market platform: visual editor, heatmaps, recordings and testing in one place. What we like: a team can run research and testing without stitching tools. Friction we hit: the all-in-one surface tempts teams into shallow tests because launching one feels so easy.
- AB Tasty. Strong visual editing and personalisation with an ecommerce lean. What we like: merchandising-flavoured targeting that maps to how retail teams think. Friction we hit: complex builds still end up in custom code, where every platform converges anyway.
- Convert. The quiet professional choice: solid engine, clean API, straightforward pricing published on the site, privacy-forward. What we like: it does the job without theatre, and agencies love it for exactly that. Friction we hit: fewer bundled extras, which is also rather the point.
- Kameleoon. Testing and personalisation with genuine technical depth on the server side. What we like: one of the better answers when a programme genuinely spans client and server. Friction we hit: less name recognition means more internal selling in larger organisations.
- Intelligems. The Shopify price-testing specialist: cohorted price experiments, profit-aware readouts. What we like: it solves a problem the general platforms handle badly or not at all. Friction we hit: it is a specialist, so it joins a stack rather than replacing one.
- Shoplift. Shopify-native theme and template testing. What we like: genuinely accessible for merchants without developers, and correct about where Shopify puts the boundaries. Friction we hit: hypothesis depth is capped by what a theme split can express.

The quick picks, if you want the shortcut
If you are a Shopify merchant without a developer, start with Shoplift and graduate when your questions get smaller than a template. If price is the question, Intelligems, full stop. If you are a mid-market DTC brand building a first serious programme, Convert or VWO cover everything you will do in year one at a price you can explain. If your roadmap genuinely spans client and server, Kameleoon and Optimizely are the conversations to have, in that order unless procurement enjoys the second one. And if you are an agency buying for clients, weight the API and the workspace model over everything else, because you will live in ten of these accounts at once.
How to actually choose
Start from the tests you intend to run, not the feature grid. A store whose backlog is visual product-page and cart work needs a dependable client-side editor and honest stats, which every platform above provides. A programme heading into pricing, search or checkout logic needs server-side capability, which narrows the field fast. And on Shopify specifically, the platform question is tangled up with theme architecture and app conflicts, which is why our Shopify testing guide treats the method choice before the tool choice.
Then weigh three unglamorous things. Pricing transparency, because at the time of writing only a minority of these platforms publish numbers, and a contact sales button predicts renewal negotiations you will not enjoy. Exit cost, because your test archive and event wiring accumulate in the tool, and moving platforms two years in is a genuine project. And support quality on a bad day, which no feature grid shows and every reference call reveals.
What does not matter
The AI features, mostly. Every platform now advertises AI-generated hypotheses, AI copy suggestions and AI insights. We have yet to watch one produce a hypothesis that survives contact with real session recordings and order data, because the AI has never met your customers. Buy the engine, the editor and the SDKs. Treat the AI garnish as garnish.
The heatmap-and-recording bundles matter less than the demo suggests, too, for a different reason: the teams that use them well would have bought them separately anyway, and the teams that would not are buying a feature they will open twice. The bundling question is about invoices, not capability.
And do not switch tools to fix a stalled programme. In our experience the bottleneck is almost never the platform; it is build quality and velocity, which travel with the team, not the licence. Pick a competent tool, then spend the energy where the results actually come from. If you want a recommendation against your specific stack and backlog, get in touch and we will give you one with no reseller margin behind it.



