MadanyCo.
All posts

Operator

Too small to test

· 1 min read

From the platform side, most tests a small advertiser runs are finished by the third day, and the winner is noise.

The pattern repeats across a great many accounts. Two versions of an ad. A couple of hundred riyals a day behind each. A three-day run, because three days is how long patience lasts. Then a decision: variant B, four conversions against two, so B wins, B takes the budget, and the account records a learning.

Four against two is not a result. It is the same coin landing differently twice. The arithmetic here is unkind and it does not care how small the account is: to tell a real difference of a few percentage points from chance, each arm needs conversions in the hundreds, not in single digits. An account taking thirty orders a week cannot buy that in three days, or in three weeks, at any budget it has.

So what the screen is actually reporting is which arm got the better first few hours of delivery, and that effect is larger than the one being tested.

The usual response is to run a smaller, tighter test. That does not help. Narrowing the scope does not lower the number of conversions significance needs.

The thing that does work is unglamorous. Choose the version you believe in, on judgement, and write down the reason. Run it properly, for a month, with the whole budget behind it. Then read the account rather than the arm: orders, cost per order, repeat rate, the shape of the month against the one before it.

A small account cannot buy statistical proof, and it can still learn. What it cannot afford is a fake experiment that borrows the vocabulary of learning and hands back a coin toss.