Indexation policy, sharded sitemaps and a robots rewrite
SEO and Organic Search · XML sitemap and robots.txt architecture
An indexation policy written as thresholds, sitemaps sharded so lastmod means something, and a robots file that stops contradicting your canonicals.
About this service
Large sites are rarely under-indexed because Google cannot find the pages. They are under-indexed because the site says three different things about the same URL: the sitemap submits it, the canonical points somewhere else, and robots.txt blocks the path the canonical depends on. On a marketplace we audited last year, 41 percent of sitemap entries carried a canonical to a different URL. Crawl rate was never the constraint, and nothing about crawl budget mattered until that contradiction was removed.
The decision this forces:
Before a single file is written, somebody has to say which pages deserve to be in the index. That is a commercial decision rather than a technical one, and it is usually why this work has stalled before. We write the indexation policy as rules with thresholds, so the answer stops depending on who happens to be in the room. A facet combination enters the sitemap when it has at least eight in-stock items and at least twenty organic sessions in the trailing ninety days, and leaves when either condition fails for thirty consecutive days. Expired listings leave on a stated schedule with a stated status code. Every rule carries an owner's name.
The recommendation we make most often is to stop submitting between a third and half of the estate. If page count is a number reported upward in your company, that conversation belongs at the start, and we will have it with the person who owns the number rather than with the engineer who would otherwise have to defend the deletion.
What we build:
Sitemaps sharded by lifecycle instead of by arbitrary chunk: recently added, recently changed, stable. lastmod then means something, which is the only condition under which it gets read. An index sitemap that agrees with the shards. A generator specification precise enough for your platform team to implement without a follow-up call, covering inputs, ordering, shard boundaries and what happens when the source query times out, plus the validation job that fails the build when a sitemap contains a non-canonical or non-200 URL.
robots.txt is rewritten line by line, and every line we remove is accompanied by what it was blocking and what changes when it stops. Where a disallow has been hiding a real problem, such as a faceted crawl trap, a session parameter or a staging host that leaked, we say what actually fixes it, because a disallow on a URL that also carries a canonical means the canonical is never read at all.
Evidence we work from:
Search Console through the API into BigQuery, so coverage states can be joined to revenue rather than looked at as a chart. Server or CDN logs sampled across a full week, to compare what is being crawled with what is being submitted. A crawl checking canonical, hreflang and status agreement against each other. We do not let a third-party index estimate reach a recommendation.
Where we say no:
On the two project tiers we do not write the generator into your codebase; we specify it and review the implementation. We do not do content, internal link building or anything downstream of what belongs in the index. We do not run the retainer alongside a project tier, so pick one. And we will not produce a sitemap that submits pages we have told you should not be submitted, even on instruction. If that is the requirement, the engagement ends with the audit delivered and paid.
The wrong service if:
Your site is under roughly five thousand URLs. A single sitemap and your CMS default are fine, and you would be paying us to agree with a plugin. Your platform team cannot change robots.txt outside a quarterly release train, in which case the work is still real but the sequencing changes, and you should tell us that before we start rather than after. Or you want a sitemap submitted for its own sake, with nothing else touched. That is paperwork, and it does not need us.
Scope
- Target market
- Worldwide, United States, United Kingdom, DACH, Nordics
- Working language
- English
- Industry
- B2B SaaS, Ecommerce and DTC, Marketplaces, Travel and hospitality, Media and publishing
- Engagement model
- One-off project, Monthly retainer
- Turnaround
- 2 weeks
- Seller type
- Boutique agency
What the seller needs from you
- 1How are URLs generated, and who can change what the sitemap outputs?
- 2Links to your current robots.txt and sitemap index, plus any regional variants.
- 3Read-only Search Console access, and analytics or BigQuery if the data already lands there.
- 4Is indexed page count reported as a KPI internally, and to whom?
- 5Can you export a week of server or CDN logs including bot traffic?
Asked at checkout. Delivery time starts once you answer, not when you pay.
Reviews
No reviews on this service yet.
Reviews appear only after an order completes, and both sides review each other. Nothing here is seeded or bought.
Starting at €6,900