Web Analytics · Product & Customer
Product Experimentation Measurement
The rules that decide the test are written down before the first visitor is assigned.
A test loses credibility when people check it every day and stop as soon as the result looks favorable. We set the decision rules before launch and apply them consistently, whatever the result shows. You end up with a defensible test design and readout, with guardrails that reveal when one metric improves at another's expense.


Some of the 500+ brands we've worked with
See all referencesHow we run it
No result is called before the agreed rule allows it.
We agree on the rules before anyone can adjust them to fit the emerging data. Four steps take the experiment from a defined hypothesis to a defensible decision.
How we hold ourselves to it
- What counts as success, decided before anyone looks — We define the hypothesis and primary success metric before anyone sees the results.
- Enough traffic that a real effect won't hide in the noise — We estimate the traffic or duration needed to detect a meaningful effect, reducing the temptation to stop early because of noise.
- A metric that must not get worse — We choose guardrail metrics upfront so an improvement in the primary metric cannot conceal harm elsewhere.
- When and how the test gets read, fixed in advance — We agree when and how to analyze the test, removing the option to stop at a convenient-looking result.
Design the experiment
We define the hypothesis, unit of assignment, sample size, guardrails, and stopping rule before launch. Your product owner confirms the primary metric.
Experiment design document


Verify the setup
Before the full launch, we verify that assignment works as intended and that the test has not already been contaminated. The analyst confirms the test isn't already contaminated.
Setup verification


Monitor without peeking-driven decisions
During the run, we monitor technical health without using interim performance to make stop or continue decisions. No one calls the test before the rule date.
Monitoring log


Call the result
We analyze the outcome using the agreed rule and report it plainly, including a null result when that is what the test produced. The analyst confirms every guardrail before recommending a rollout.
Experiment readout


Sample size, guardrails, and stopping rules are fixed before launch.
Automation estimates the sample size from your traffic, checks the assignment split against the intended ratio, flags technical anomalies during the run, and drafts the readout from the pre-agreed rule. The discipline is human: nobody calls the test before the rule date, and an analyst clears every guardrail before a rollout is recommended.
What you get
The primary result always comes with its guardrail performance.
The readout remains useful whether the result is positive, negative, or inconclusive.


Working document
Experiment design document
The hypothesis, sample size calculation, guardrails, and stopping rule, written before results existed.


Decision memo
Experiment readout
The result, analyzed according to the pre-agreed rule, with guardrail performance included.


Reference document
Decision record
What was decided based on the result and why, so it's traceable later.
We call it done when: the readout follows the pre-registered rule, every guardrail is reported alongside the primary metric, and a decision to ship nothing is written up as carefully as a win.
Fit and readiness
A favorable result on the day you stop is not reliable evidence.
This work fits teams that need a controlled product experiment and an honest readout.
A good fit when
- You're testing a specific in-product feature or change and need a properly designed experiment rather than an informal split.
- Past tests have been stopped early or re-analyzed after the fact, and results no longer feel trustworthy.
- You need guardrail metrics to catch a change that helps one number while quietly hurting another.


Better handled as other work when
- You're testing marketing spend or channel effectiveness rather than an in-product change. That is Incrementality & Lift Measurement.
- You need us to build and ship the treatment itself. We design the test and read the result. Your product or engineering team implements the change.
If one of these is closer to your situation, start here instead: All Product & Customer Analytics tasks
We call it done when: the hypothesis, primary metric, sample size, guardrails, and stopping rule are recorded and the product owner has signed off on the primary metric.
People who build your measurement system
Zeo designs measurement systems that connect a business decision to governed collection and reporting you can check. The people shown here work on the part of that system this page covers.

Zafer Yıldız
Web Analytics Manager

İlker Emir
Senior Performance Marketing Executive

Abdullah Tanıdır
Performance Marketing Team Lead

Sevda Yurtvermez
Performance Marketing Team Lead

Serap Yurtvermez
Performance Marketing Team Lead

İpek Ezer
Performance Marketing Executive

Onur Durdağı
Performance Marketing Executive

Deniz Çağın Demirci
Frontend Developer

Mirzamin Aghazada
UI/UX Designer

Gülşah Şahin Özkan
Senior SEO Analyst

Metehan Urhan
New Business & Partnership Manager
Tools we use
Tools behind this work
Evan Miller's A/B Toolsthe neutral calculator and reference essay behind the fixed sample size and stopping rule
Amplitudewatches the guardrail metrics during the run, so nobody has to peek at the primary result to stay informed
Optimizelyruns the test itself against the sample size and stopping rule fixed before launch
Next step
Run an experiment your team can trust


Before we start
Questions teams ask before booking
Can we stop the test early if it looks like a clear win?
Only if we planned for that upfront with a proper sequential testing method. Stopping an unplanned test the moment it looks good is one of the most common ways experiments produce false positives.
What if the test comes back inconclusive?
We report that honestly rather than searching for a subgroup where it "worked." An inconclusive result is still useful information. It often means you need more traffic or a larger effect to detect.
Do you build the feature we're testing?
No. We design the test and analyze the result. Your product or engineering team builds and ships the actual treatment.
What do you need from our testing setup?
Access to your experimentation platform or feature flag system, the hypothesis you want to test, and enough traffic to reach a meaningful sample size.





















































