Common Crawl Visibility Checker
Most AI training data starts at Common Crawl. This checker reads Common Crawl's own files and shows how it sees any domain: how many pages each monthly crawl captured, when an AI-crawler block first appeared in robots.txt and who wrote it, what CCBot actually archived vs the live page, whether a firewall challenges bots, sitemap coverage, and the domain's crawl priority in the web graph.
Seven checks, one input, plain-language verdicts, and copy-paste fixes for anything it finds.
- Checks
- Captures per crawl, robots.txt block dating, block attribution, stored-copy drift, live CCBot probe, sitemap coverage, crawl priority
- Data
- Common Crawl static archives and the full 118M-domain web graph, refreshed monthly
- Cost
- Free, no signup, no email
- Export
- JSON and sitemap-gap CSV