One asset per cell, an event floor before anything is called a result, and tagging that makes a quarter of rounds add up to knowledge.
About this service
Inside a single Meta ad set, ads are not randomised. The delivery system concentrates impressions on whichever asset wins the first few hours, so a five-creative ad set measures the algorithm's opening guess and then confirms it. Every creative round we build runs one asset per cell with the budget split at cell level, and no round is read below roughly 50 optimisation events per cell per week. Under that floor the round is not called; it is extended or cancelled.
What we build:
A protocol your media team runs without us, plus the tagging layer that makes a quarter of rounds add up to knowledge. Every asset is tagged on six attributes at upload: format, hook device, claim, proof type, offer position, and whether a person appears. The taxonomy is enforced in the naming convention, in the brief template and in the warehouse view, so after four rounds the account answers which claim wins rather than which video won in March. Tags survive a change of buyer, which is the actual test of whether the framework was built or merely described.
Measurement:
Cell-level split tests on Meta at campaign level. Drafts and experiments on Google Search. For Performance Max, we do not treat an asset-group swap as a controlled comparison, because it is not one; PMax gets a campaign-level experiment split or it gets left out of the readout. Where brand search absorbs the effect, we read incrementality with a geo design using GeoLift on your own regional data, and on accounts spending enough to support it, a platform conversion lift study as the cross-check. Fatigue is read by fitting cost per acquisition against cumulative reach, not by the folklore that frequency 3.0 means refresh.
What we deliberately do not report:
In-platform ROAS as the result of a test. It is attributed, not measured, and the attribution moves when you change the creative you are testing. Thumbstop rate as a win condition; it is a diagnostic for the first three seconds and nothing more. Statistical significance on a cell that has not passed its event floor, whatever the platform's own bar says on the screen.
What is not included:
Creative production. We design the round, write the brief the round requires and read the outcome; your studio or in-house team makes the assets, and we work with them directly. Media buying and budget management stay with your team. Landing page work belongs in a different engagement, and if the test's real variable is the page rather than the ad, we will say so and stop.
This is not for you if:
Paid social spend is under roughly 60,000 euro a month, where the event floor makes clean rounds arithmetically impossible and the honest advice is to spend the fee on production instead. You need a winner declared every week for a standing report. You want the same partner to buy the media and grade the results; that is a conflict we decline rather than manage. You are running the round across a promotional window, where the promotion moves more than the creative and the readout would be worthless.
How it runs:
Four weeks to a working framework: taxonomy, brief template, cell design, event floors per market, the warehouse view, and one round designed and read end to end with your team watching. Rounds after that either run in-house on the protocol, which is the intended outcome, or under retainer where the account carries several markets at once. English or Dutch, Netherlands or worldwide accounts.