How Long to Run Ad A/B Tests: Sample Size Math Made Practical
A/B test duration is arithmetic, not folklore: required sample divided by daily volume. Here is the sample-size math made usable, plus the floors, ceilings, and kill rules that keep tests affordable.

Run an ad A/B test until each variant has collected enough conversions to make the difference you care about visible — not for a fixed number of days. As a working rule, that means at least one full week to cover day-of-week cycles, and somewhere between 50 and a few thousand conversions per variant depending on how small a difference you want to detect. Duration is an output of arithmetic, not a calendar preference: days needed = required sample per variant ÷ your daily volume per variant. This article gives you that arithmetic, the floors and ceilings that override it, and the cheaper testing sequence to use when the math says you can't afford significance.
Duration is an output, not a calendar choice#
Most advice hands you a fixed answer — seven days, fourteen days, a month. Days are the wrong unit. A campaign generating 40 conversions a day resolves a lander test in about a week; the same test on a campaign generating 4 a day takes over two months, by which point the result is contaminated by everything that changed in between. Three inputs decide your real duration:
- The baseline rate of the metric you are testing — CTR for creatives, conversion rate for landers and offers.
- The smallest difference worth detecting. Detecting a 10% lift costs far more traffic than detecting a 50% lift, and most small lifts aren't worth the spend it takes to see them.
- Daily volume per variant — set by your budget and bids.
Fix the first two and the required sample falls out of a formula. Divide by the third and you have your duration. If the answer comes out longer than about four weeks, the correct move is usually to change the test, not to grind it out — more on that below.
The sample-size math, minus the statistics degree#
The standard shortcut for a two-variant test at 95% confidence and 80% power:
n per variant ≈ 16 × p × (1 − p) ÷ d²
where p is your baseline rate as a decimal and d is the absolute difference you want to detect. Two worked examples:
- Creative test at CTR level. Baseline CTR 0.5%, detecting a 20% relative lift (0.1 points absolute): n ≈ 16 × 0.005 × 0.995 ÷ 0.001² ≈ 80,000 impressions per variant. Native placements serve impressions fast and cheap, so this can resolve in days.
- Lander test at conversion level. Baseline conversion rate 2%, detecting a 25% relative lift (0.5 points absolute): n ≈ 16 × 0.02 × 0.98 ÷ 0.005² ≈ 12,500 clicks per variant. At a $0.40 CPC that is a five-figure test for a single pair of pages.
| What you're testing | Baseline | Lift to detect | Sample per variant |
|---|---|---|---|
| CTR (creative) | 0.5% | 20% relative | ~80,000 impressions |
| CTR (creative) | 0.5% | 50% relative | ~13,000 impressions |
| CVR (lander) | 2% | 25% relative | ~12,500 clicks |
| CVR (lander) | 2% | 50% relative | ~3,100 clicks |
| CVR (offer page) | 5% | 25% relative | ~4,900 clicks |
Two lessons fall out of that table. First, small lifts are brutally expensive: halving the detectable difference quadruples the sample. Second, testing at impression level costs a small fraction of testing at click level. That asymmetry should shape your whole testing program: test creative angles at CTR level, test funnels at conversion level, and don't test button-color-sized changes at all — the sample needed to detect them rarely pays for itself.
The floor: one full week, whatever the math says#
Even when volume delivers your sample in two days, run at least seven. Weekday and weekend audiences behave differently — different people, different devices, different intent. Publisher mix on native networks rotates through the week as news cycles and content schedules shift, so Tuesday's traffic is not Saturday's traffic. Conversion lag adds a second bias: a variant's Thursday clicks may convert on Saturday, so an early cut systematically favors whichever variant happened to get its impressions first. A test that ends mid-week has sampled a skewed slice of your real traffic, and the "winner" may simply be the weekday specialist.
The ceiling: long tests rot from the inside#
Past three to four weeks, an A/B test stops being a controlled experiment. Creative fatigue decays both variants — rarely at the same speed, which quietly reverses rankings mid-test. Competitors enter and leave the auction and shift your traffic quality. Seasonality drifts underneath the whole comparison. If the formula says you need six weeks of traffic, read that as a verdict: the difference you're hunting is too small to detect at your volume. Test a bigger swing instead — a different hook or angle, a different funnel structure — where the detectable difference is large and the sample is affordable.
Don't peek — or at least don't act on peeks#
The classic failure mode: check the dashboard daily and declare victory the first day the numbers look significant. Repeated looks at accumulating data inflate false positives dramatically — with daily peeking, an "it hit significance at some point" stopping rule fires far more often than the advertised 5% error rate. Regression to the mean then explains most "my winner stopped winning" complaints: the variant that leapt ahead on day two was riding noise, and it drifts back. The discipline is simple: pre-commit your sample size, monitor daily only for stop-loss conditions — broken tracking, runaway spend, a variant performing catastrophically — and evaluate the winner once, at the end.
Kill rules for when you can't afford significance#
Most native and affiliate buyers cannot fund 12,500 clicks per variant, and pretending otherwise produces fake rigor. The honest alternative is staged culling — cruder than significance testing, and standard practice among buyers who test dozens of creatives a month:
- CTR cull (days 1–3). Give each creative a few thousand impressions, then kill everything clearly below the pack. You aren't proving anything — you're triaging so budget concentrates on plausible winners.
- Spend guardrail (rolling). A common practitioner rule: pause any variant that has spent 2–3× your target CPA with zero conversions. That isn't statistics; it's survival.
- Significance on the finalists. Once two or three survivors carry most of the spend, run a proper conversion-level test between them using the math above. You fund one real test instead of ten fake ones.
Shortcut the sample: start from ads that already survived#
Every variant you don't need to test is money saved, and the cheapest information about what works is what competitors keep paying to run. An ad that has run for 30+ days has passed its owner's profitability check every day of that run. OpenAdLibrary's index of 725,000+ live native creatives across 49 networks (June 2026) records first-seen and last-seen dates for every creative, which turns ad longevity into a filterable signal rather than a hunch — the longest-running native ads in your vertical are a pre-tested starting grid. Build your test around two or three angles already proven in the wild and you're testing variations on a winner instead of cold guesses; you can browse live competitor creatives free with the native ad research tool. Fewer, better variants mean smaller samples, shorter tests, and less budget burned on discovery.
A worked example, end to end#
Say you run a $60-payout lead-gen offer with one campaign doing roughly 1,500 clicks a day across your test cells.
- Creative stage: four creatives compete at CTR level. A few thousand impressions each resolves in 2–3 days; cull to the top two.
- Lander stage: two landers at conversion level. Baseline CVR around 3%, detecting a 30% relative lift (0.9 points absolute): n ≈ 16 × 0.03 × 0.97 ÷ 0.009² ≈ 5,700 clicks per variant, about 11,400 total. At 1,500 clicks a day that is eight days — call it a full week plus weekend coverage.
- Verdict: one test cycle runs 10–12 days from launch to a defensible winner, with losing spend capped by the guardrail rule rather than left running on hope.
That is the realistic shape of "how long": a few days of cheap triage, a week-plus of focused measurement, and a calendar length dictated by your click volume — not by anyone's magic number of days.







