An ad is a hypothesis until the evidence says otherwise.

Responsive search ads can compare combinations of headlines and descriptions, but the comparison needs multiple RSAs and enough unpinned assets to create meaningful variation. We define the reason for each variant, launch a controlled comparison, and wait for enough evidence before interpreting the result. You end up with an ad-testing rhythm where each live variant answers a specific question and the answer remains available for the next test. For ad groups with enough traffic to reach a real read, and reviewers available to check copy before launch.

A specialist comparing two responsive search ad variants on a testing board with a hypothesis written above each

Some of the 500+ brands we've worked with

See all references
  • MediaMarkt
  • LC Waikiki
  • Bayer
  • PWC Türkiye
  • Sigortam.net
  • Apsiyon
  • Yatsan
  • Amazon
  • BMW
  • Shell
  • Hyundai
  • PepsiCo
  • Red Bull
  • Decathlon
  • Sanofi
  • EY
  • KPMG
  • GE
  • 3M
  • Domino’s
  • Lexus
  • Trendyol
  • Hepsiburada
  • Yandex

AI can draft variant copy and track the sample size. A person reviews brand-sensitive claims and signs off before a live ad launches.

How we hold ourselves to it

  • Every new asset starts as a stated hypothesis, not a guess
  • Pinning is a tradeoff you choose on purpose, every time
  • A test needs a real sample before anyone calls it
  • Every result gets written down, win or not
  1. Form the hypothesis

    Name the specific signal driving the test, whether that is low Ad Strength, weak click-through rate, or an unproven claim, and state exactly what the variant changes. The paid-search lead confirms the hypothesis is worth testing.

    Written hypothesis naming the driving signal and exactly what the variant changes.

  2. Draft the variant assets

    Write headlines and descriptions against the hypothesis, deciding deliberately what, if anything, gets pinned and accepting the combination trade-off that comes with it. The copy lead selects and edits the assets that actually launch.

    Variant headline and description set with the pinning decision and its combination trade-off recorded.

  3. Run it through policy and brand review

    Check claims, restricted terms, and brand voice before anything goes live, so a complaint or a disapproval never gets the first look. Brand or compliance reviewer signs off before launch.

    Cleared policy and brand review covering claims, restricted terms and voice.

  4. Launch a controlled comparison

    Run at least two RSAs per ad group with matched targeting and unique final URLs, so the comparison actually isolates the variable being tested. The paid-search lead approves the live comparison before it starts.

    Controlled comparison running at least two responsive search ads per ad group with matched targeting and unique final URLs.

  5. Hold for a real sample

    Let the test run until it reaches a defensible sample size and duration for the ad group's actual traffic. The paid-search lead approves any early stop.

    Run record showing the test reached a defensible sample size and duration for that ad group's traffic.

  6. Read the result and record it

    Call it a win, a loss, or inconclusive against the original hypothesis, then write the learning into the record for the next ad group. The paid-search lead signs the learning before it gets applied elsewhere.

    Learning record calling the test a win, a loss or inconclusive against the original hypothesis.

Each test keeps its hypothesis, comparison, and result together, giving the next ad group evidence it can consult before repeating the same question.

  • Test matrix

    Every hypothesis, its variant, the ad group it ran in, and its current status, in one running view.

  • Asset library with pinning rules

    What's pinned, why, and the combination trade-off accepted for each ad group's asset set.

  • Policy and brand pre-flight checklist

    The specific claims, terms, and brand checks cleared before each variant went live.

  • Learning record

    Win, loss, or inconclusive, with the reasoning, kept searchable so a past result can inform the next ad group.

An unchanged ad stops producing useful learning. Google recommends running at least two responsive search ads with a Good or Excellent Ad Strength rating per ad group, each with its own final URL, precisely so there's something to compare. Pin every headline out of caution and you shrink the very combination pool the format is built to test. Call a result a win after two days and a handful of clicks, and all you've actually measured is noise.

A good fit when

  • Your ad group draws enough traffic for two responsive search ads, so the comparison can move well beyond a handful of clicks.
  • A brand or policy reviewer can check claims and restricted terms before launch, so every cleared variant enters the learning record with its approval attached.
  • The planned sample and hold period can govern the read, even when an early click-through swing makes one headline look convincing.
  • An ad group has run one RSA for months with no second variant to compare it against.
  • Every headline and description is pinned, so the combinations the format is built to test barely happen.
  • A new variant gets called a "win" after a couple of days and a small number of clicks.
  • Nobody flagged a policy-sensitive claim before the ad went live.

Better handled as other work when

  • Traffic is too thin for two responsive search ads to reach a meaningful comparison, so the test matrix would remain inconclusive.
  • A specific headline has to launch today, before a hypothesis, controlled comparison, and planned sample can be completed.
  • No brand or policy reviewer can clear sensitive claims before launch, so the variant needs copy approval outside this testing method first.

Paid Search, Paid Social, CRO, and Programmatic each run under a named owner at Zeo. The consultants below are matched to the channel this page is about, so you can see who you'd actually work with.

  • Adalysis

    runs the RSA variant comparison itself, from cloning assets to reading which one actually won

Bring us the current ad groups. We'll set the hypotheses and comparisons, then keep each result in a learning record the next test can use.
Start testing your ad copy

How many ad variants do you run in one ad group at a time?

At least two, matching Google's own recommendation, each with a unique final URL so the comparison is real, though more can run if the ad group's traffic supports reading them separately.

Does pinning every headline ruin Ad Strength?

It can. Pinning reduces the number of combinations the system can test, so we pin deliberately, only for a legal, brand, or clarity reason.

How long before you call a test result?

Until it reaches a sample size and duration set for that ad group's actual traffic. A low-traffic ad group needs longer for the same confidence a high-traffic one reaches quickly. Cutting it short on day two just measures noise.

Can AI write and launch ad copy on its own?

AI can draft variant copy against a stated hypothesis. A named copy lead edits and approves what launches, and a policy or brand reviewer signs off before anything goes live.