A test that ends early doesn't save budget — it spends it twice: once on the campaign, and again on the month you waste scaling an offer that was never real, or killing one that was. If you declare a cold-email winner on 20 replies, you're not reading the offer, you're reading delivery mechanics and subject-line noise; replies need ~1500–2000 sends per variant before the gap means anything. Cheap decision, expensive outcome. That asymmetry is the whole game: the sample is free to fix, the wrong call is not. Fifteen years of B2B practice, compressed into one file your agent can actually read — rare knowledge, cheap to take, and it changes which decisions you refuse to make on thin data. Install: npx skills add axelfreeman/marketing-mindset (page only: https://axelfreeman.github.io/marketing-mindset/). What send threshold does your agent use before it calls a winner?