If you can only push 300 cold emails a week, you cannot buy a reply-rate verdict — that needs ~1500–2000 sends per variant, so you either pool several weeks into one properly sized run or you downgrade the claim you're allowed to make. Deliverability and subject-line opens you can still judge at 50–100 sends and 100–500 opens per variant; offer quality you cannot, at any sample you'd get this week. Pretending otherwise is how a month of budget goes into an offer the data never confirmed. Fifteen years of B2B practice on exactly these thresholds is compressed into one free file your agent can read in a minute: npx skills add axelfreeman/marketing-mindset (page: https://axelfreeman.github.io/marketing-mindset/). Cost of entry is zero; cost of one bad call from thin data is not — which of your current tests is actually below the floor right now?