A creative testing system sized to what your conversion volume can actually resolve, with written decision rules and two cycles run alongside your team.
About this service
Comparing two ads inside one ad set is not a test. Delivery decides who sees what inside the first few hundred impressions, and the winner you read is the winner the system already chose. The arithmetic is the other half of the problem. Detecting a 20 percent difference in conversion rate at 80 percent power takes roughly 300 conversions per cell. At 30 conversions per cell, which is where most accounts declare winners, the test can only resolve a difference of around 70 percent, and no creative change produces that. We build the testing system around what your volume can actually resolve, which for many accounts means testing concepts rather than executions.
Layers, and what each is worth:
We separate creative into four layers and test them in order of expected effect. The claim, meaning what you are asserting is true. The proof, meaning why anyone should believe it. The format, meaning how it is delivered. The execution, meaning cut, colour and typeface. Claim and proof are where differences large enough to measure actually live. Execution is where most testing budget goes, because it is the easiest thing to produce, and it is where the differences are smaller than your sample size can see. An account that cannot resolve a 20 percent difference and spends its testing on execution is manufacturing confident noise.
How tests are run:
Cell-level campaigns with one variable moving, budget split, randomised at the person level where the platform supports it. Meta's A/B test tool for the comparisons it handles honestly, and campaign-level cell structures where it does not. Conversion lift studies when spend clears the platform's minimum and the question is incrementality rather than ranking. Hook rate, hold rate and thumbstop are read as diagnostics that explain why a cell lost. They are never the deciding metric, and an agency that reports them as results is reporting the easy number.
The operating system you keep:
A naming taxonomy enforced through the marketing API, so every ad carries its concept, claim, format and test identifier and reporting can be cut by any of them without manual tagging. A test register holding hypothesis, cell design, required sample, planned duration, and the decision written down before the test runs. Decision rules with real thresholds: what promotes to always-on, what earns a rematch, what is retired, at what frequency and on what fatigue signal. A brief template that forces a claim and a proof instead of a mood board. And a rotation cadence matched to your production capacity rather than ours.
Sector notes we bring:
In energy, a price or saving claim in an ad has to be substantiated to the standard the Swedish Marketing Act expects, and the substantiation belongs in the brief rather than in a scramble after a complaint. In sports and fitness the calendar moves more than the creative does; a test running across a race registration window is measuring the calendar and reporting it as a creative result. We time around both.
Not included:
We do not produce creative. We do not contract for a number of assets per month and will not quote that way. No brand strategy, no positioning workshop, no naming work.
Who this is not for:
Accounts under roughly 25,000 euro a month in paid social, where the honest advice is to stop testing and put the budget behind one strong concept. Teams that can produce one new concept a quarter, where a framework will only document how little there is to test. And anyone who wants rules nobody is obliged to follow: the decision rules are worth nothing unless the person holding veto power was in the room when they were written.
What you receive:
The test register, the taxonomy applied to the live account through the API, the brief template, the decision rules signed by whoever owns them on your side, and two full test cycles run with us in the room.