XML Sitemap Sharding and Robots Rules for Large Catalogs

Marlow CommerceVerified agencyNew0 orders on this service
SEO and Organic Search · XML sitemap and robots.txt architecture

Sitemaps sharded per template so Search Console reports an indexing rate you can act on, plus robots rules that do what people assume they do.

About this service

A sitemap is not a request. It is the one instrument you own that reports back, and most catalogs waste it by submitting a single flat file containing every URL on the site. Split the same URLs into shards by template and lifecycle and the Sitemaps report changes from a vanity number into a per-template indexing rate. When eight thousand of nine thousand variant URLs sit unindexed while the core catalog shard reads 96 per cent, you stop guessing which template Google is refusing. How we shard: Well below the format ceiling. The limits are fifty thousand URLs and 50MB uncompressed per file, but we build files of five to ten thousand, because the report resolves per file and a file is the smallest unit of evidence Google will give you. Shards follow templates and lifecycle rather than the alphabet: new arrivals, core catalog, long-tail variants, editorial. lastmod is emitted only when content materially changed. A timestamp that moves on every deploy is noise, and Google learns to ignore the field for your domain entirely, which costs you the one recrawl hint that still works. Robots.txt, and the two things it cannot do: It cannot remove a page from the index. A disallowed URL still collects links and can be indexed URL-only, appearing in results with no title. And it cannot enforce a noindex, because a crawler forbidden from fetching the page never reads the tag. That combination is the one teams ship most confidently and it guarantees the page stays. We separate the tools properly: robots.txt for crawl waste, noindex on a crawlable page for removal, canonical for consolidation, 410 for things that should be gone. Facets, which is where the budget actually goes: On a marketplace the parameter space is effectively infinite and Googlebot will spend real budget inside it. We read logs first, at least fourteen days, and count the share of Googlebot requests landing on parameterised, sorted or paginated URLs rather than on the canonical set. Then we decide per parameter: crawlable and indexable where the combination has demand behind it, crawlable and canonicalised where it does not, blocked only where the combination is a machine artefact. Google stopped using rel=next and rel=prev in 2019; if your pagination still leans on them it needs a different answer. What we hand over: A generator wired into your build, not a static file someone regenerates by hand and then forgets during a busy quarter. Shard definitions in code, an IndexNow endpoint for Bing and Yandex with the honest caveat that Google does not consume it, and a short note per shard stating what belongs in it and what its indexing rate was on the day we left, so the next drop is visible to whoever is looking. Not included: We do not submit URLs in order to make them index. When a template is not being indexed, the sitemap is reporting a content or duplication problem and the answer is upstream: fewer pages, better ones, or a consolidation we will scope separately. We also will not use robots.txt to hide pages you are embarrassed by. Blocking thin pages leaves them on the site, in the internal links and in the way. Deleting or merging them is the work. Hiding them is how a site ends up with four hundred thousand URLs and eleven thousand sessions. Who should not buy this: Sites under roughly five thousand URLs. Google will find your pages. An afternoon from your own developer and the default sitemap output of your framework covers it, and we would rather say that now than take the fee and dress it up.

Scope

Target market
Worldwide, India
Working language
English, Hindi
Industry
Ecommerce and DTC, Marketplaces, Home and furniture, Media and publishing
Engagement model
One-off project
Turnaround
2 weeks
Seller type
Boutique agency

What the seller needs from you

  1. 1Roughly how many URLs exist per template, including parameterised variants?
  2. 2Can you provide 14 to 30 days of raw access or CDN logs?
  3. 3How are your sitemaps generated today?
  4. 4List the URL parameters in use and what each one does
  5. 5Search Console access, including any domain-level property

Asked at checkout. Delivery time starts once you answer, not when you pay.

Reviews

No reviews on this service yet.

Reviews appear only after an order completes, and both sides review each other. Nothing here is seeded or bought.

Other sellers offering xml sitemap and robots.txt architecture

See all →

Starting at $5,500