Experiment Readouts That Survive a Finance Review

Nuria IglesiasProNew0 orders on this service
CRO and Experimentation · Statistical analysis and readout

Re-analysis and written decision memos: SRM first, pre-registered metrics, CUPED where it helps, and four verdicts including no answer.

About this service

What a readout has to do: End in a decision. Every readout I write closes with one of four verdicts: ship, kill, iterate on the same hypothesis, or the answer did not arrive. The fourth appears in roughly a third of the tests I am asked to re-analyse. Teams find that share uncomfortable for about a week, until they compare it with the alternative, which is a shipped change with no real effect entered into the record as evidence and used to justify the next three tests. Order of operations: Sample ratio mismatch first, before anything else is looked at. A chi-square p below 0.001 means the test is not read, it is fixed. Nothing downstream of a broken randomisation deserves interpretation, and the most expensive mistake in this discipline is the careful analysis of a corrupted split. Then the pre-registered primary metric, at the unit the randomisation actually used. Ratio metrics computed per user with delta-method standard errors rather than naive session ratios, which understate variance and manufacture significance that evaporates on replication. Revenue winsorised at the 99th percentile, stated openly, with the unwinsorised figure shown next to it so you can see whether the conclusion depends on that choice. Variance reduction: CUPED against a 14 to 28 day pre-period covariate where the traffic supports it. On returning-user-heavy products this typically cuts variance on revenue metrics by 10 to 25 percent, which is the difference between a four-week test and a three-week one. On new-user acquisition traffic it does close to nothing, and I will not bill you for applying it there. Stopping rules: Fixed horizon by default, with the horizon fixed before launch. Sequential methods, always-valid intervals or mSPRT, where the business genuinely needs the option to stop early and only where that was chosen in advance. Peeking at a fixed-horizon test and stopping on a p-value is the single most common way a serious programme generates confident nonsense. Secondary and guardrail metrics: Declared as a family before launch and corrected with Benjamini-Hochberg. Guardrails read as one-sided harm checks, never reported as wins. Segment analysis only where the segment was pre-registered, and reported as a hypothesis for the next test rather than as a finding. What you receive: A decision memo of no more than three pages: the verdict, the number it rests on, the interval around that number, what would change the verdict, and what to run next. Behind it, the SQL against your warehouse and a Python notebook that reproduces every figure in the memo. Your analysts should be able to rerun it without me, and should be able to disagree with me using it. Not included: No test design or hypothesis work in this engagement. No dashboard build. No instrumentation repair: if the analysis shows the data is wrong I tell you exactly where and stop, because fixing it is a different piece of work with different people in the room. No case study rights and no use of your results in my own material. What I refuse: Re-cutting a test until it produces a winner. Reading a test that never had the power to answer its question, which I identify from the design and decline before doing the work rather than after invoicing for it. Reporting a posterior probability of being best without the loss function that would make it actionable. And presenting a result to your executives that I would not defend, in the same words, to a sceptical statistician sitting in the same room. Who this is not for: Teams who need the answer to be yes. The memo has no adjective in it that a number cannot support, and that is the point of buying it from outside.

Scope

Target market
Worldwide, Spain
Working language
English, Spanish
Industry
Developer tools, Ecommerce and DTC, Marketplaces, Gaming
Engagement model
Audit only
Turnaround
1 week
Seller type
Fractional executive

What the seller needs from you

  1. 1Which test or tests should I read, and what decision is waiting on the answer?
  2. 2Was the primary metric and stopping horizon fixed before launch? Where is that written?
  3. 3Warehouse access to exposure and metric tables.
  4. 4Baseline rate and observed variance on the primary metric.

Asked at checkout. Delivery time starts once you answer, not when you pay.

Reviews

No reviews on this service yet.

Reviews appear only after an order completes, and both sides review each other. Nothing here is seeded or bought.

Other sellers offering statistical analysis and readout

See all →

Starting at €5,500