Sitemap Partitioning and Robots.txt Architecture

Nathan SinclairTop ratedNew0 orders on this service
SEO and Organic Search · XML sitemap and robots.txt architecture

Sitemaps split by template so coverage data becomes diagnosable, lastmod tied to real content revisions, robots.txt rewritten line by line.

About this service

A sitemap is a measuring instrument before it is anything else. One file listing every URL tells you 68 percent are indexed and nothing whatever about which ones or why. Split the same URLs into files by template — product detail, category, editorial, location, help — and the Search Console coverage report becomes a per-template indexation rate your team can act on the same afternoon. That split is usually the highest-value hour of the engagement and it is the first thing I do. How the files get built: Inside your build or publishing pipeline, from the same source of truth the router uses, never from a plugin that crawls your own site on a cron. A crawl-generated sitemap inherits whatever is broken in your internal linking, which defeats the reason for having one. Index file at the root, children held under 50,000 URLs and 50MB uncompressed, gzipped, each regenerated when its own section changes rather than nightly for everything. lastmod, honestly: Google uses lastmod where a site's values prove consistently accurate and ignores it otherwise, across the whole property, for a long time. A CMS stamping every page with today's date on a template deploy has already spent that credibility. I wire lastmod to the content revision timestamp — the moment body, title or price changed — and explicitly not to a build or cache event. Where that distinction cannot be made in your stack, we omit lastmod rather than lie in it. robots.txt, and what it does not do: It governs fetching, not indexing. A disallowed URL with links pointing at it can still surface in results, described from anchor text, and it will never be recrawled to discover the noindex you added, because the two directives cancel each other. I go through your file line by line, separate what belongs there from what belongs in an X-Robots-Tag header, and remove the rules pasted in four years ago that now block a rendering asset. Along the way: the staging host disallowed instead of password-protected, the faceted navigation rules blocking the crawl path to pages you want found, and the AI crawler question — I will implement whichever policy you choose and tell you plainly what each option costs you in retrieval visibility. What arrives: The generator, in your repository, with tests. A robots.txt where every line carries a comment stating why it exists and who asked for it. A temporary sitemap for URLs being removed, kept live exactly as long as the deindexing needs and then deleted, which is the one legitimate use of the pattern and nobody does it. IndexNow and ping submission wired into deploy where your platform supports it. Then a second reading of coverage data at 30 days, per template, against the numbers we started from. Not included: Fixing the underlying reason a template is not indexed. If a category page type is excluded as duplicate or thin, the sitemap will document that cleanly and it stays a content and internal linking problem outside this scope. Also excluded: HTML sitemap pages, image and video sitemaps for assets hosted on a third-party platform I have no access to, and news sitemaps unless you are genuinely in Google News. Who this is not for: Sites of a few thousand URLs on a managed platform whose built-in sitemap is already correct. Those platforms exist, and if you are on one I will tell you rather than bill you to replace something that works. Teams who want the robots.txt hardened without touching anything else — blocking is the easy half and rarely the half that is wrong.

Scope

Target market
Worldwide, Canada
Working language
English, French
Industry
Ecommerce and DTC, Marketplaces, Energy, Media and publishing
Engagement model
One-off project
Turnaround
2 weeks
Seller type
In-house-grade specialist

What the seller needs from you

  1. 1Where the URL source of truth lives and how pages are published.
  2. 2Your current robots.txt, plus any rule nobody can explain.
  3. 3Search Console coverage export for the property.
  4. 4Which URL groups you want out of the index, if any.

Asked at checkout. Delivery time starts once you answer, not when you pay.

Reviews

No reviews on this service yet.

Reviews appear only after an order completes, and both sides review each other. Nothing here is seeded or bought.

Other sellers offering xml sitemap and robots.txt architecture

See all →

Starting at $5,000