An intake, statistics and decision structure where tests are powered before launch and no one is allowed to stop one early.
About this service
Detecting a 2 percent relative improvement on a 5 percent conversion rate takes roughly 750,000 users per arm at 80 percent power and a two-sided 95 percent threshold. Most programs I am called in to repair are running four-week tests on 40,000 users and declaring winners in about half of them. That is the whole problem, and no amount of ideation fixes it.
How a test enters the program:
Backwards from the effect that would justify shipping. The proposer states the smallest change worth the engineering cost; I compute the users and the weeks required at your current traffic; if the answer is longer than the business will wait, the test does not launch. It becomes a research question, a sequential launch with a rollback, or a decision somebody makes on judgement and signs their name to. Around a third of proposed tests exit here, and that is the single largest saving in the engagement.
What gets fixed before launch:
One primary metric, chosen at the money end of the funnel. In gambling that is net revenue after bonus cost at day 30, never registrations. On a marketplace it is completed orders net of cancellations, with a supply-side guardrail so a buyer-side win that starves sellers is caught. Guardrails, the decision rule, the run length and the analysis method are written down and circulated before a single user is bucketed. After that, the numbers cannot change anybody's mind about what the numbers mean.
The statistics, and why:
Sequential analysis with always-valid intervals, available in Statsig and in GrowthBook, so that people can look at a running test without corrupting it. Nobody stops looking, so the method has to survive being looked at. CUPED using a pre-period covariate, which on established funnels cuts the required sample materially and is the only free power you will get. Sample ratio mismatch checked on every test with a chi-square threshold at p below 0.001, treated as a stop-and-investigate rather than a footnote. Benjamini-Hochberg across secondary metrics, because reading fifteen of them and reporting the interesting one is how programs manufacture wins.
The rules that do not bend:
No early stop on a positive reading, ever. No launch without a pre-registration document. No promotion of a secondary metric to primary after the fact. A shipped winner needs a named code owner within 30 days or it gets reverted, because winners rotting in a feature flag is how a program quietly becomes a fork. And the quarterly reconciliation: projected uplift from shipped winners against what actually appeared in revenue. That number is usually smaller than the sum of the wins, and putting it on the page in front of the finance lead is what makes the program credible for the following year.
What I will not do:
Run the tests for you. This is program design and governance; your team executes and I sit in the readouts. I will not set a target for tests per quarter or report velocity, and if that metric already exists on somebody's scorecard I will ask for it to be removed as a condition of the work. I do not sign a readout on a test I was not allowed to power properly.
Who should not buy this:
Companies whose largest funnel step sees under 25,000 users a month. A/B testing is the wrong instrument at that volume and I will recommend interviews, session review and staged rollouts instead. Also not for teams wanting one supplier for both program and execution.
On your side afterwards:
A one-page pre-registration template, a decision rule set your leadership has signed, a readout format of one page per test carrying hypothesis, MDE, achieved n, the primary result with its interval, guardrails and a decision expiry date, and an archive where losers stay findable.