Server Log Analysis for Crawl Budget and Indexation

Nathan SinclairTop ratedNew0 orders on this service
SEO and Organic Search · Log-file analysis

Ninety days of raw origin logs read against verified Googlebot: where crawl actually goes, which templates go unrequested, how long indexing takes.

About this service

Why pages published a quarter ago are still not indexed is almost always answerable from the access log, and the answer is usually uncomfortable: in the last eleven engagements, between 40 and 70 percent of verified Googlebot requests hit URLs nobody at the company would defend — parameter permutations, pagination past page five, redirect chains left behind by a platform migration. Meanwhile the template that carries revenue was requested every 30 to 60 days. That ratio, not a crawl tool's score, is what the work is about. What I need from you: Raw origin access logs, 90 days minimum, in whatever format your server already writes. nginx or Apache combined, CloudFront or Cloudflare Logpush, Akamai DataStream, ALB logs — all fine. What I will not accept as a substitute is a crawler tool's export, Search Console Crawl Stats screenshots, or a 5 percent sample. Sampling destroys the tail and the tail is where the problem lives. If your CDN terminates before the origin and you keep only edge logs, say so at the start; the analysis changes and I will tell you what it can and cannot conclude. How the analysis runs: Verification first. Every user agent claiming to be Googlebot is checked by reverse then forward DNS. On a payments property sitting behind a WAF, 15 to 30 percent of self-declared Googlebot traffic turns out not to be Google, and counting it silently inflates every conclusion that follows. Each surviving request is then joined to a template class taken from your route definitions rather than guessed from URL patterns, and to the status code, response time and bytes actually served. From there: crawl allocation by template, request-to-first-index latency per template, the response-time percentile at which Googlebot demonstrably backs off on your infrastructure, orphan URLs receiving crawl, and indexed URLs receiving none. What you get: A findings document written for your engineering lead, ordered by the crawl volume each fix recovers, with the log lines attached to every claim. A DuckDB query set that runs against your own exports — the same queries I ran, so your team reproduces every number without me in the room. On the larger engagements the fixes themselves: redirect map collapse, parameter and pagination handling, sitemap partitioning, delivered as pull requests against your repository and re-measured in the following log window. Not included: Content, editorial internal linking, link acquisition, keyword work. No Search Console-only diagnosis; if log access is impossible for vendor or legal reasons I will say the engagement cannot be done properly rather than substitute weaker data and charge for it anyway. No list of 200 prioritised issues. Most of what a crawler flags does not move crawl allocation, and padding a document with it wastes the one week your engineering lead will actually give this. Who this is not for: Sites under roughly 5,000 URLs. At that size crawl allocation is rarely the binding constraint and you would be paying me to confirm it. Teams who cannot ship a change to robots.txt, the sitemap generator or the redirect layer inside a quarter — the analysis will be correct and inert. Anyone who needs the finding to be that the current setup is fine. Working languages are English and French. Logs are processed on Canadian infrastructure by default; if your data residency policy forbids that, tell me before kickoff and I will run the pipeline inside your environment instead, which is why the queries are portable SQL in the first place.

Scope

Target market
Worldwide, Canada
Working language
English, French
Industry
Ecommerce and DTC, Marketplaces, Fintech, Energy
Engagement model
One-off project
Turnaround
2 weeks
Seller type
In-house-grade specialist

What the seller needs from you

  1. 1Which log sources exist, and who can authorise an export?
  2. 2Search Console access for the property, at user or API level.
  3. 3The route or template map for the site.
  4. 4Any migration, replatform or CDN change in the last 12 months, with dates.
  5. 5Data residency or PII constraints on where logs may be processed.

Asked at checkout. Delivery time starts once you answer, not when you pay.

Reviews

No reviews on this service yet.

Reviews appear only after an order completes, and both sides review each other. Nothing here is seeded or bought.

Other sellers offering log-file analysis

See all →

Starting at $6,500