How to Get Log File Analysis Right
Reading server logs to see exactly which URLs crawlers requested, when, and what they received.
The key figures
- What logs record
- every request, including crawler requests, with status and timestamp
- Verification
- crawler identity should be verified by reverse DNS, since user agents are spoofed
- Typical findings
- crawl waste on parameters, redirects and error pages
- Value threshold
- most useful on large sites
- Retention
- logs are often rotated quickly and need collecting deliberately
Why this is worth getting right
It is the only direct evidence of crawler behavior, as opposed to inference from reports, and on large sites it reveals waste nothing else shows.
Do this, not that
Do
- Verify crawler identity rather than trusting the user agent string
- Compare crawled URLs against the sitemap and the known URL set
- Look for crawl budget spent on parameters, redirects and errors
- Arrange retention before you need the data
- Combine with Search Console rather than reading either alone
Don’t
- Trusting user agent strings without verification
- Analyzing a period too short to be representative
- Drawing conclusions from logs alone
- Assuming logs will still be there when you need them
When to bring in help
Our advice Bring in help when a large site has crawl problems, when important pages are rarely crawled, or when crawl budget is being consumed by URLs nobody intended to publish.
Where this comes from
- Google Search Central — Verifying Googlebot
- Google Search Central — Crawl budget management
- Bing Webmaster Tools — Crawl reporting
The figures and practices above come from the sources listed.
Working on something like this?
We take on SEO Services work for teams who want it done once, properly. Tell us what you are building and we will tell you honestly whether we are the right studio for it. Start a project.
Where to go next
Spotted something wrong? Report an error on this page. We correct on the page and say what changed.