An inventory of what should be indexed, evidence for what is actually happening, and cohort decisions that each carry a name and a reason.

Robots rules, sitemaps, canonical tags, and redirects are usually written at different times by different people, so after a site has run a few years they routinely disagree. A rule blocks a page the sitemap still lists as a priority, or a canonical points one way while a redirect sends bots another. We pull server logs and Search Console coverage to see what's actually happening, compare it against every current rule, and decide cohort by cohort which URLs to allow, consolidate, redirect, or retire. The pages worth finding get found. The rest stop competing for attention.

A Zeo crawler figure sorts glowing URL tiles into keep, redirect, and retire piles beside a sitemap console.

Some of the 500+ brands we've worked with

See all references
  • Hepsiburada
  • Yves Rocher
  • ETS Tur
  • Tosla
  • Desa
  • Koleksiyon Mobilya
  • Elele
  • Amazon
  • BMW
  • Shell
  • Hyundai
  • PepsiCo
  • Red Bull
  • Decathlon
  • MediaMarkt
  • Bayer
  • Sanofi
  • EY
  • KPMG
  • GE
  • 3M
  • Domino’s
  • Lexus
  • Trendyol

Four stages, each one grounded in log data and Search Console evidence.

  1. Get the real data

    We pull your URL inventory, server logs, and Search Console coverage. This shows what bots and search engines are actually doing instead of a crawler's best guess at what's live. Our technical lead confirms the log window and URL scope are genuinely representative before the dataset becomes the agreed baseline.

    A dated, scoped dataset everyone agrees is the starting point.

  2. Compare crawled against indexed

    We line up bot requests against index coverage, sitemap membership, and canonical signals for the same URLs, so we can see exactly where discovery and indexing disagree. Each disagreement gets read in its own context, and our technical lead decides whether it is a real problem or a deliberate exclusion.

    A clear map of which URLs are eligible, which are excluded, and which are stuck in between.

  3. Decide, cohort by cohort

    For each group of URLs we choose allow, consolidate, redirect, noindex, or retire, and write down why, so the next person doesn't have to guess. The allow, redirect, consolidate, or noindex call for each cohort is made by our technical lead, who signs their name to the reason.

    A change list where every item traces back to evidence instead of a hunch.

  4. Ship it small, then watch

    We roll out the smallest version of the change, recrawl, and check whether the index actually responded the way we expected. Whether the pilot is strong enough to expand, or needs another round first, is our technical lead's decision.

    Either the pattern holds and we expand it, or it doesn't and we find out why before it spreads.

AI reconciles the datasets; a named technical lead decides what each cohort deserves.

AI reconciles the CMS URL export, the routing layer, and the raw server-log sample into one matched dataset, joins bot-request logs, sitemap membership, canonical signals, and Search Console coverage per URL, groups URLs by shared template, parameter pattern, or mismatch signature, and diffs the recrawled pilot cohort against the prediction. The cohort call stays human. We will not force every URL into the index to inflate a coverage number, we do not ship a bulk robots, canonical, or redirect change without sampling who it actually touches first, and we never build directives that show search engines something different from what your visitors see.

Four things you can hold in your hand.

  • Brief

    Crawl & index baseline

  • Decision matrix

    URL-state evidence

  • Redirect map

    Robots, sitemap & redirect changes

  • Audit report

    Post-launch validation

We call it done when: The baseline, URL-state evidence, directive changes, and post-launch validation are done when the baseline names the cohorts, the owner, and what is out of scope, every URL's crawl and index status is backed by a log line or a Search Console record, every rule change carries a reason, an owner, and a way to reverse it, and the validation says plainly whether the index moved as expected and what happens next.

Getting crawled isn't the same as getting indexed, and getting indexed isn't the same as getting found for the right thing.

A good fit when

  • You're not sure how much of your site search engines actually keep in the index, versus just visit once and move on.
  • Your crawl budget is going to parameters, duplicates, or pages nobody should be looking at in the first place.
  • You need robots, sitemap, and canonical rules that someone actually owns and can explain.

Better handled as other work when

  • You want every URL forced into the index regardless of whether it deserves to be there.
  • You want to ship a bulk redirect or noindex change today without checking who and what it touches first.

If one of these is closer to your situation, start here instead: Technical SEO

We call it done when: You finish with an inventory of what should be indexed, evidence of what is actually happening, and a monitoring habit that catches drift before it costs you traffic.

  • Screaming Frog

    directive crawl for canonicals, robots, sitemaps, and redirects

  • JetOctopus

    bot-request evidence by directory, pattern, and response code

  • Oncrawl

    joined crawl, log, analytics, and Search Console states

  • Sitebulb

    visual crawl maps for traps, depth, and isolated cohorts

  • Google Search Console

    index coverage, sitemap status, and sampled URL decisions

  • Bing Webmaster Tools

    second-engine index checks and controlled URL submission evidence

  • Yoast SEO

    index directives and sitemap membership traced to their WordPress source

Bring your Search Console access and whatever server logs you've got. We'll help you see the gap between what's crawled and what's indexed.
Look inside your index

How much of this does AI actually do?

Agents compare thousands of URLs against their expected crawl and index state and flag the mismatches. This covers the comparison work that takes hours by hand. Deciding what a mismatch means, and whether to change a directive, stays with a Zeo specialist.

What do you need from us to start?

A URL export from your CMS or routing layer, a server log sample that covers a normal crawl cycle, and read access to Search Console. That's usually enough for a first pass.

How long before we see index movement?

Recrawling and reindexing run on the search engine's own schedule. It's often one to a few weeks before changes show up in coverage reports. We check field data weekly rather than promise a date.

What if we can't give you full server log access?

We can still work from Search Console coverage data and whatever log sample you do have, but the URL-state evidence will be thinner and some cohort decisions will carry more uncertainty. Full log access is what makes the URL-state evidence airtight. A partial sample just means we flag more items as 'needs more evidence' instead of deciding outright.