An experiment has little lasting value if nobody can find its record.

We rebuild every reconstructable past experiment as a standard record covering its hypothesis, mechanism, evidence, result, and decision. Tags show what the test examined and who saw it. When two records point in opposite directions, a reviewer sees the conflict before the next hypothesis reaches approval. Your team can retrieve the evidence and decision behind a past experiment before the same dead hypothesis consumes another testing slot. This fits a testing program with enough history that memory, old decks, and former employees' folders can no longer carry the record.

A tagged, searchable archive of past experiments organized by mechanism and audience

Some of the 500+ brands we've worked with

See all references
  • Shiftdelete
  • ETS Tur
  • Sompo Sigorta
  • Hisar
  • Amazon
  • BMW
  • Shell
  • Hyundai
  • PepsiCo
  • Red Bull
  • Decathlon
  • MediaMarkt
  • Bayer
  • Sanofi
  • EY
  • KPMG
  • GE
  • 3M
  • Domino’s
  • Lexus
  • Trendyol
  • Hepsiburada
  • Yandex
  • Pegasus Airlines

Before a hypothesis moves forward, the proposer searches by mechanism and audience. Reviewers then compare the evidence behind any conflicting records. AI can suggest tags and surface possible matches, but a person decides whether two past results really contradict each other.

How we hold ourselves to it

  • Wins, nulls, negatives, and invalid results share one record format
  • Every new hypothesis goes through an archive search before approval
  • Reviewers add a resolution note when past results conflict
  • Search follows mechanism and audience
  1. Inventory what's already run

    Collect every past experiment that can still be reconstructed, including its hypothesis, design, result, and decision, even when it was not documented properly at the time. The research owner confirms the inventory before archiving begins.

    Inventory of every past experiment that can still be reconstructed, with its hypothesis, design, result and decision.

  2. Standardize the record format

    Rebuild each past experiment into the same structure: hypothesis, mechanism, evidence, result, confidence, decision, so a test from three years ago reads the same way as one from last week. The CRO strategist confirms a standardized record accurately represents what actually happened.

    Rebuilt archive where every record uses the same hypothesis, mechanism, evidence, result, confidence and decision fields.

  3. Tag by mechanism and audience

    Tag each record by the mechanism it tested, such as urgency, trust, clarity, or friction, and by the audience it ran on. A new hypothesis about urgency can then surface every past urgency test across the archive. The research owner approves the tags before the record is searchable.

    Mechanism and audience tag index across the whole archive.

  4. Surface contradictions across the archive

    Cross-check records with similar tags and opposite directional results. Each possible contradiction goes to a reviewer for a resolution note. The CRO strategist reviews flagged contradictions and adds a resolution note.

    Contradiction log pairing similar-tag records with opposite directional results, each routed to a reviewer.

  5. Build the reuse-check step into new hypothesis approval

    Require every new candidate hypothesis to be checked against the archive before roadmap approval. The search catches dead or contradicted ideas before they consume another testing slot. The CRO strategist reviews the surfaced history and decides whether the new hypothesis still stands.

    Approval gate requiring a completed archive search before a hypothesis reaches the roadmap.

  6. Keep the archive current

    Add every newly completed experiment to the repository in the same standard, tagged format. The archive stays current as the program keeps testing. The research owner approves each new entry before it's added.

    Standing intake that adds each completed experiment to the archive in the same tagged format.

The archive has to help during approval. These linked records let proposers find old evidence and give reviewers enough context to understand conflicts before another test is approved.

  • Standardized experiment archive

    Every reconstructable past experiment in one consistent format: hypothesis, mechanism, evidence, result, decision.

  • Mechanism and audience tag index

    The index lets someone search by the mechanism tested and the audience that saw it.

  • Contradiction log

    Cases where two past tests reached opposite conclusions on a similar mechanism, with a resolution note where one exists.

  • Reuse-check record

    The log of new hypotheses checked against the archive before approval, including any revised or rejected after the search.

When the first result cannot be found, the same hypothesis gets tested again. A year after a test runs, its record is often a slide in a deck nobody can find, inside a folder named after someone who left the company. The idea returns as a new proposal. Sometimes it produces the same null result. Sometimes it conflicts with the old test because the first record's data-quality problem was forgotten.

A good fit when

  • Your experiment history now spans more records than anyone can remember, so past evidence disappears into old decks when a new idea arrives.
  • New hypotheses keep repeating tests from a year or two ago, but the proposer cannot retrieve the earlier mechanism, audience, and result.
  • Several teams contribute experiments, yet no searchable archive lets reviewers compare nulls, negatives, wins, and conflicting results before approval.
  • A "new" hypothesis turns out to be a test that already ran and came back null.
  • The evidence behind last year's "win" is missing.
  • Two tests reached opposite conclusions about the same page element, and nobody connected them.
  • Wins get shared. Null and negative results fade from memory.

Better handled as other work when

  • Only one or two experiments have finished, so a standardized archive would contain too little history to change a new hypothesis decision.
  • The team skips the repository, so archived evidence cannot stop the same hypothesis from consuming another test slot.
  • Past decks and tickets lack enough evidence to reconstruct the hypothesis, implementation, or result, so the archive would preserve guesses.

Paid Search, Paid Social, CRO, and Programmatic each run under a named owner at Zeo. The consultants below are matched to the channel this page is about, so you can see who you'd actually work with.

  • VWO

    the one place old test insights go instead of a doc nobody can find eight months later

We can inventory the tests you can still reconstruct, put them into one consistent format and add the tags that expose a repeat or contradiction before another testing slot is spent.
Build a learning repository

What happens to old tests that were never documented properly?

We start with the decks, tickets, and analytics history that remain. Each record includes only what those sources support. Missing pieces stay marked. The uncertainty remains visible when someone retrieves the test later.

Does the archive include tests that didn't win?

Null, negative and invalid results often do the most useful work because they stop an unsupported hypothesis from returning as a new idea. The archive keeps those outcomes alongside the wins.

Who has to check the archive before proposing a new test?

Everyone who submits a hypothesis to the roadmap completes an archive search during approval.

How do you handle two past tests that contradict each other?

We flag the contradiction, then a reviewer compares the audience, timing, implementation, and sample size before adding a resolution note.