An experiment cannot be more trustworthy than its QA record.

A visual editor shows one browser and one happy path. It doesn't show the mobile keyboard covering the CTA, the tracking script that quietly stopped firing, or the accessibility regression a screen-reader user hits first. We run every variant through the same checklist, verify assignment and tracking before exposure, and exercise the rollback path in advance. You get a signed QA record for every experiment before launch, plus a rollback path that has already passed a live drill. A standing release gate makes sense once experiments run often enough that checks by feel create avoidable risk.

A signed release checklist covering functional, tracking, performance, and accessibility QA for an experiment variant

Some of the 500+ brands we've worked with

See all references
  • GE
  • Yves Rocher
  • Peak Games
  • eOfis
  • Hotiç
  • Turna.com
  • Amazon
  • BMW
  • Shell
  • Hyundai
  • PepsiCo
  • Red Bull
  • Decathlon
  • MediaMarkt
  • Bayer
  • Sanofi
  • EY
  • KPMG
  • 3M
  • Domino’s
  • Lexus
  • Trendyol
  • Hepsiburada
  • Yandex

We apply the approved checklist to each variant and keep the evidence with the release record. AI can draft checks and flag mismatches. The experiment owner is the person who approves release.

How we hold ourselves to it

  • Every variant clears one release checklist
  • Assignment logs must match the configured split before exposure begins
  • QA reviews performance and accessibility against the base page
  • The rollback mechanism is exercised in advance
  1. Build the release checklist for this program

    Before the first experiment enters QA, define the checks required for this site or app, including functional behavior, cross-device rendering, accessibility, performance, tracking, and assignment. The CRO strategist and engineering owner approve the checklist before it's used.

    Program release checklist covering functional, cross-device, accessibility, performance, tracking and assignment checks.

  2. Verify functional and cross-device behavior

    Check the paused variant across the browsers, viewports, and devices that matter for this audience. Confirm that the intended change renders and behaves as designed in the full matrix. An independent reviewer who did not build the variant signs off on functional QA.

    Signed device and browser matrix result with every discrepancy noted.

  3. Verify tracking and assignment integrity

    Confirm every relevant event fires correctly in each variant and that the assignment mechanism splits traffic the way it was configured, before any traffic is exposed. The analytics owner confirms tracking is correct before the experiment owner authorizes exposure.

    Event-firing and assignment-split verification recorded per variant before exposure.

  4. Check performance and accessibility regressions

    Measure the variant's effect on load performance and layout stability. Re-run an accessibility check against the same criteria used for the base page before release. The engineering owner reviews flagged regressions before sign-off.

    Performance and accessibility comparison of the variant against the control page.

  5. Prove the rollback path works

    Exercise the rollback mechanism before the experiment goes live. The drill must leave a working exit if a guardrail breach occurs. The experiment owner confirms the rollback test passed before authorizing launch.

    Rollback drill record showing the exit worked before launch.

  6. Sign, launch, and log the record

    Collect sign-off from every required reviewer, launch only the approved scope, and file the complete QA record so it's available if a guardrail trips or a post-mortem is needed later. The experiment owner authorizes launch only once every required sign-off is on file.

    Complete, signed QA record filed with the launch scope it approved.

These four records show which checks ran, what they found and who approved the variant before launch.

  • Program-level release checklist

    The standing checklist every experiment on this program gets checked against, covering functional, cross-device, tracking, performance, and accessibility QA.

  • Functional and cross-device QA record

    The signed result of running each variant through the configured device and browser matrix, with any discrepancies noted.

  • Tracking, assignment, performance, and accessibility findings

    The QA record confirms that events fired, assignment split as configured, and performance and accessibility were reviewed against the control before release.

  • Tested rollback record

    The signed result of exercising the rollback mechanism and confirming it worked before launch.

A variant that looks correct in the builder can still fail in production. The builder preview is only the starting point. Once variant code meets real browsers and devices, it can shift the layout, interrupt tracking, slow the page, or create an accessibility regression. QA has to inspect that production behavior before visitors are exposed.

A good fit when

  • Experiments launch often enough that checks by feel create avoidable release risk, yet no repeatable checklist shows what every variant must clear.
  • Variants come from an internal team or agency, but nobody independent has checked their tracking, device behavior, performance, and accessibility before exposure.
  • A failed experiment release would disrupt real traffic, so the team needs the rollback path tested before a guardrail ever trips.
  • One desktop browser is the entire review.
  • Assignment splits go unchecked until the results look odd.
  • A new script or style changes page load or layout stability before anyone measures it.
  • The team has no tested rollback path once exposure starts.

Better handled as other work when

  • You run only a handful of experiments each year. The preflight inside test design already covers them.
  • The variants still exist only as ideas, so there is no staged build, assignment log, or event firing to inspect yet.
  • QA keeps finding issues, but no one owns the release or rollback decision, so every flag waits without a route to action.

Paid Search, Paid Social, CRO, and Programmatic each run under a named owner at Zeo. The consultants below are matched to the channel this page is about, so you can see who you'd actually work with.

  • BrowserStack

    catches the mobile keyboard covering the CTA on a real device, not an emulator's guess at one

  • Google Tag Manager

    shows which tags actually fired, and in what order, before the variant goes live

We can build or run the release gate, catch tracking, performance, accessibility and functional problems before exposure, then exercise the rollback path.
Plan experiment QA

Isn't QA already part of the test design method?

Test design includes a preflight step for that one experiment. This method is the standing, program-level discipline for teams running enough experiments that a repeatable, independently reviewed checklist matters more than a one-off check.

Who builds the variants, Zeo or our team?

Either team can build the variant, and an independent reviewer who was not the builder signs off before launch.

What happens if QA finds a real problem close to launch?

The launch waits. The issue gets fixed. The affected checks run again. Exposure begins only after the required owners sign off.

Do you test the rollback for every single experiment?

Yes. A rollback that has never been exercised is only a guess about whether it works. We run the drill before launch and confirm the mechanism executes.