Two or three landing variants split at ad level, sized before launch, with the stopping rule written into the brief.
About this service
At three hundred qualified conversions a month, a two-arm test sized to detect a twenty percent difference runs about five weeks. At four arms it eats the quarter. The first thing this engagement produces is that arithmetic against your real numbers, and roughly a third of the time it ends with us recommending you do not run the test at all.
How the split is served:
Variants are separate URLs split at the ad level, through Google Ads experiments or a Meta A/B test object, not swapped in the browser by a testing script. Client-side tools flicker, they delay the largest contentful paint, and they change the page after the ad platform's own systems have measured it. Each arm is a static build off one template, and before launch we verify that the LCP spread between arms sits inside a hundred milliseconds. Otherwise the test measures page speed wearing a headline's clothes and you act on it for a year.
What varies:
One axis: the promise in the headline and the proof sitting directly beneath it. Not layout, not colour, not button microcopy. Three arms at most and usually two. A variant set where fourteen things differ produces a winner nobody can explain and a follow-up test that fails to replicate. For a freight brand the arms might be transit-time certainty, total landed cost, and customs handling: three different reasons to pick up the phone, each carrying its own evidence. Whichever wins, the business learns something it can act on outside the page.
The stopping rule:
Written before launch, into the brief, with the date in it. A fixed horizon, three weeks or the computed sample, whichever arrives first, and no interim peek that counts as a result. Conversions are measured at the qualified event, meaning a confirmed booking, a returned quote, or a call past ninety seconds. Form submissions are not the unit; every inherited variant set we have unpicked was optimising to form fills and moving the wrong number. The ad platform is fed the same qualified event, so the auction and the test do not disagree about what happened.
Language:
Hebrew and English arms are never compared against each other. They are different audiences arriving from different auctions at different costs, and pooling them produces a result that is really a mix shift. Each language gets its own test or it does not get tested this quarter.
What we do not use:
Optimizely, VWO and their equivalents on this class of traffic. Multi-armed bandits when the goal is a decision rather than yield during the test, because a bandit leaves you without a clean answer to carry into the next quarter. Personalisation rules. Heatmap screenshots as evidence for anything at all.
Not included:
We do not run media, set budgets, or manage the campaign the split lives inside, though we write the exact experiment configuration for whoever does. We do not run ongoing monthly testing programmes. We do not build or license a testing platform for you.
Not for you if:
Fewer than roughly a hundred and fifty qualified conversions a month. At that volume the honest recommendation is to write one good page and change the offer, and we will make that recommendation rather than sell a test that cannot resolve. A board date that needs a winner regardless of significance. Or a fixed intention to test something we consider unmeasurable at your traffic, in which case we will say what we would test instead and you are entirely free to disagree.
What you receive:
The sizing model with your numbers in it, the variant set built and instrumented, the experiment configuration for your ad platform, and a readout stating the decision, the confidence behind it, and what it does not tell you.
Scope
- Target market
- Worldwide, United Kingdom, Israel
- Working language
- English, Hebrew
- Industry
- Ecommerce and DTC, Travel and hospitality, Local services, Logistics
- Engagement model
- One-off project
- Turnaround
- 1 month or more
- Seller type
- In-house-grade specialist