The rules that decide the test are written down before the first visitor is assigned.

A test loses credibility when people check it every day and stop as soon as the result looks favorable. We set the decision rules before launch and apply them consistently, whatever the result shows. You end up with a defensible test design and readout, with guardrails that reveal when one metric improves at another's expense.

Two Zeo testers at an A/B door pair, measuring foot traffic

Some of the 500+ brands we've worked with

See all references
  • Domino’s
  • Silverline
  • Odeabank
  • Duru
  • Gusto
  • Elele
  • Amazon
  • BMW
  • Shell
  • Hyundai
  • PepsiCo
  • Red Bull
  • Decathlon
  • MediaMarkt
  • Bayer
  • Sanofi
  • EY
  • KPMG
  • GE
  • 3M
  • Lexus
  • Trendyol
  • Hepsiburada
  • Yandex

We agree on the rules before anyone can adjust them to fit the emerging data. Four steps take the experiment from a defined hypothesis to a defensible decision.

How we hold ourselves to it

  • What counts as success, decided before anyone looks — We define the hypothesis and primary success metric before anyone sees the results.
  • Enough traffic that a real effect won't hide in the noise — We estimate the traffic or duration needed to detect a meaningful effect, reducing the temptation to stop early because of noise.
  • A metric that must not get worse — We choose guardrail metrics upfront so an improvement in the primary metric cannot conceal harm elsewhere.
  • When and how the test gets read, fixed in advance — We agree when and how to analyze the test, removing the option to stop at a convenient-looking result.
  1. Design the experiment

    We define the hypothesis, unit of assignment, sample size, guardrails, and stopping rule before launch. Your product owner confirms the primary metric.

    Experiment design document

    Illustrated figure sketching plans at a drafting table
  2. Verify the setup

    Before the full launch, we verify that assignment works as intended and that the test has not already been contaminated. The analyst confirms the test isn't already contaminated.

    Setup verification

    Illustrated figure reading an oversized measurement dial
  3. Monitor without peeking-driven decisions

    During the run, we monitor technical health without using interim performance to make stop or continue decisions. No one calls the test before the rule date.

    Monitoring log

    Illustrated figure watching a monitor full of tracked rows
  4. Call the result

    We analyze the outcome using the agreed rule and report it plainly, including a null result when that is what the test produced. The analyst confirms every guardrail before recommending a rollout.

    Experiment readout

    Illustrated figure presenting a bar chart on an easel

Sample size, guardrails, and stopping rules are fixed before launch.

Automation estimates the sample size from your traffic, checks the assignment split against the intended ratio, flags technical anomalies during the run, and drafts the readout from the pre-agreed rule. The discipline is human: nobody calls the test before the rule date, and an analyst clears every guardrail before a rollout is recommended.

The readout remains useful whether the result is positive, negative, or inconclusive.

  • An experiment readout sheet and a decision log card

    Working document

    Experiment design document

    The hypothesis, sample size calculation, guardrails, and stopping rule, written before results existed.

  • An experiment readout sheet and a decision log card

    Decision memo

    Experiment readout

    The result, analyzed according to the pre-agreed rule, with guardrail performance included.

  • An experiment readout sheet and a decision log card

    Reference document

    Decision record

    What was decided based on the result and why, so it's traceable later.

We call it done when: the readout follows the pre-registered rule, every guardrail is reported alongside the primary metric, and a decision to ship nothing is written up as carefully as a win.

This work fits teams that need a controlled product experiment and an honest readout.

A good fit when

  • You're testing a specific in-product feature or change and need a properly designed experiment rather than an informal split.
  • Past tests have been stopped early or re-analyzed after the fact, and results no longer feel trustworthy.
  • You need guardrail metrics to catch a change that helps one number while quietly hurting another.

Better handled as other work when

  • You're testing marketing spend or channel effectiveness rather than an in-product change. That is Incrementality & Lift Measurement.
  • You need us to build and ship the treatment itself. We design the test and read the result. Your product or engineering team implements the change.

If one of these is closer to your situation, start here instead: All Product & Customer Analytics tasks

We call it done when: the hypothesis, primary metric, sample size, guardrails, and stopping rule are recorded and the product owner has signed off on the primary metric.

  • Evan Miller's A/B Tools

    the neutral calculator and reference essay behind the fixed sample size and stopping rule

  • Amplitude

    watches the guardrail metrics during the run, so nobody has to peek at the primary result to stay informed

  • Optimizely

    runs the test itself against the sample size and stopping rule fixed before launch

Tell us what you want to test and which decision the result needs to support. We will set the rules before launch and apply them consistently, whatever the outcome.
Plan an experiment

Can we stop the test early if it looks like a clear win?

Only if we planned for that upfront with a proper sequential testing method. Stopping an unplanned test the moment it looks good is one of the most common ways experiments produce false positives.

What if the test comes back inconclusive?

We report that honestly rather than searching for a subgroup where it "worked." An inconclusive result is still useful information. It often means you need more traffic or a larger effect to detect.

Do you build the feature we're testing?

No. We design the test and analyze the result. Your product or engineering team builds and ships the actual treatment.

What do you need from our testing setup?

Access to your experimentation platform or feature flag system, the hypothesis you want to test, and enough traffic to reach a meaningful sample size.