Seed design, warehouse-side propensity models and an activation plan whose deliverable is incremental lift against a holdout, not an AUC score.
About this service
A lookalike built from every past converter usually beats broad targeting on reported cost per acquisition and adds nothing incremental, because the seed is dominated by people who would have arrived anyway. The seeds I ship are cut to a single event inside a stated window, an application started in the last fourteen days, a booking above a value threshold, a demo attended, and expansion stays between one and three percent until a holdout earns more.
The seed argument comes first:
Most of the model's quality is decided before a line of SQL is written. Which event stands in for value, how far back it stays useful, and which converters must be excluded because they were already in market. Under roughly two thousand qualifying conversions there is no model worth training and I will say so and build rule-based audiences instead, which is a cheaper engagement and the correct one.
How the model is built:
Training runs in BigQuery ML or on your warehouse, on first-party behavioural features only. No purchased attributes enter the feature set, partly on consent grounds and partly because they degrade quietly and nobody notices. Leakage is checked explicitly: features that only exist after conversion are the most common reason a model with an impressive area under the curve produces a useless audience. Scoring produces deciles, and the decile lift chart is the working artefact, since it tells your team where to cut the audience when the reach target changes.
Activation and its limits:
Audiences are pushed through Ads Data Hub, Customer Match, The Trade Desk first-party segments or Amazon Marketing Cloud. Where a platform's own similar-audience feature is available it runs as a baseline arm rather than as a competitor, because knowing that your DSP's built-in expansion matches a custom model is worth the fee it saves you. Retraining cadence is set by observed decay, weekly for travel demand, monthly for a degree admission cycle, and the decay is measured rather than assumed.
The holdout is part of the deliverable:
A model that has not been tested against a holdout is a hypothesis. Where user-level holdouts are possible the design uses them; where they are not, geographic holdouts or ghost-bid style measurement carry it. The reported number is incremental conversions against control, with a confidence interval and the sample size that produced it. If the interval crosses zero I will report that, and it happens.
Not included:
Media buying, pacing and bid strategy. Creative. Warehouse construction. Model deployment into your production systems as a live service, which is an engineering project with different obligations than an audience build. Any model trained on data you cannot show me a consent basis for.
Who this is not for:
Advertisers below the conversion floor, who are better served by rules and contextual work. Teams that want the model rebuilt monthly as a subscription without a test between versions, which is spending on activity. Medtech advertisers who want a seed derived from condition-page visitors or symptom searches, since expanding that seed distributes a health inference across a million people who never made one, and I refuse it in every form it is proposed.
What you keep:
The training query, the feature definitions, the scoring job, the decile cuts and the holdout analysis, in your warehouse under your account. Nothing here is held hostage in a tool of mine, and your team should be able to run the second version without me.