Log-file forensics: which crawl requests your site wastes
SEO and Organic Search · Log-file analysis
Verified Googlebot requests only, split by template and status, so you can see which crawl waste is worth an engineering quarter and which is not.
About this service
Crawl waste is measurable, and it is almost always larger than the team expects. Across the log sets Renderlab parsed last year, a median of 38 per cent of verified Googlebot requests landed on a redirect, a 404, or a parameterised duplicate of a page Google had already fetched that same week. That figure is not a statistic to admire. It is the reason a newly published template waits nineteen days for indexation instead of three: the requests were spent elsewhere, and Google does not owe you more of them.
Only logs can tell you that. The crawl stats report in Search Console is sampled, aggregated to a date, and silent about which template. A raw access log gives the exact URL, the second, the source IP, the response, and whether the requester was Google at all.
Verification comes before analysis:
Between a tenth and a third of self-declared Googlebot traffic in a typical log file is not Google. Every claimed bot request is verified by reverse DNS on the source IP and a forward lookup back to it, and whatever fails is discarded before a single chart exists. Analysis of unverified logs produces confident conclusions about a scraper's habits. We have watched an agency present a crawl strategy built entirely on traffic from a rank tracker.
Formats accepted:
nginx and Apache combined, Cloudflare Logpush, Fastly, Akamai DataStream, CloudFront, AWS ALB, Vercel log drains, Kinsta and WP Engine exports. Compressed or raw. If a line carries a timestamp, a path, a status and a user agent, it is usable. Where your CDN terminates before origin we need both sides, and we will explain on the first call why one alone will mislead you.
What the engagement decides:
Where Google spends its requests, what it gets back, which of those requests you would rather it made somewhere else, and whether fixing that is worth an engineering quarter. The last clause is the one that matters, and it is where most log reports stop.
The output is a decision document rather than a dashboard. Crawl distribution by template and directory. Status mix per template, with redirect chains named and hops counted. Discovery latency, the gap between a URL first existing and Googlebot first requesting it, per template, which is the number that shows whether your sitemaps and internal linking do anything. Bot separation: Googlebot Smartphone against Desktop, Google-InspectionTool, GoogleOther, Bingbot, and the AI crawlers, counted apart because they behave differently and one of them may be the largest single consumer of your origin CPU.
Who reads them:
The same two people every time. We have read logs for a marketplace with eleven million URLs and for a publisher with four thousand, and the hard part has never been the size of the file. It is deciding which of the things you could block, canonicalise or delete actually returns requests to pages that earn money. We have told one client the honest answer was to remove most of their catalogue pages from the index, and told another that their crawl waste was real, small, and not worth a sprint.
Not included:
No implementation. We do not commit robots.txt, deploy redirect maps or write CDN rules.
No tool. The retainer below is two people reading your data every month, not a login and a chart.
No ranking or content analysis. If the question is why a page sits fourth, a log file cannot answer it.
Not for you if:
Your logs are retained for under fourteen days and retention cannot be extended. There is nothing to read.
You are pre-launch. Crawl behaviour is observed, never predicted.
You need the number that proves the last migration went well. Twice we have reported the opposite in writing to the person who commissioned the work.
What we need, and when:
Thirty consecutive days of logs at minimum, ninety preferred; Search Console at owner level; your current sitemap index; and one call with whoever owns the CDN configuration. Two weeks to the document from the day the last file lands. On the retainer, a fixed day each month and a standing thirty minutes with your platform team.
Scope
- Target market
- Worldwide, United States, DACH, Netherlands, Poland
- Working language
- English, Polish
- Industry
- B2B SaaS, Ecommerce and DTC, Marketplaces, Travel and hospitality, Media and publishing
- Engagement model
- Monthly retainer
- Turnaround
- 2 weeks, 1 month or more
- Seller type
- Boutique agency
What the seller needs from you
- 1Where do your logs live, in what format, and how long are they retained?
- 2Can you export at least 30 consecutive days, ideally 90?
- 3Search Console access at owner level, plus the URL of your current sitemap index.
- 4Which sections of the site earn money, and which exist for other reasons?
- 5Is a decision waiting on this?
Asked at checkout. Delivery time starts once you answer, not when you pay.
Reviews
No reviews on this service yet.
Reviews appear only after an order completes, and both sides review each other. Nothing here is seeded or bought.
Starting at €7,200