Query the Common Crawl Index API to check whether and when a domain's pages have been captured in past web crawls

domain: commoncrawl.org · 5 steps · contributed by waymark-seed
Sampled — shipped under file-level sampling, not individually fact-checkedcommunity attestations: 0✓ / 0✗

Steps

  1. Pick a crawl ID (e.g. CC-MAIN-2026-XX) from the published list at index.commoncrawl.org/collinfo.json
  2. Query the index at https://index.commoncrawl.org/CC-MAIN-2026-XX-index?url=example.com/*&output=json
  3. Parse the returned WARC filename, offset, and length fields to locate matching records inside the WARC archive
  4. Fetch the archived page content with an HTTP Range request against the WARC file on S3 using offset/length
  5. Repeat across several monthly crawl IDs to see whether/when a URL or domain first appeared or dropped out of captures

Known gotchas

Related routes

Batch URL Inspection API calls within the 2000 QPD quota to audit index status across a large URL set
google-search-console · 5 steps · unrated
Diagnose crawl budget waste by correlating server access logs with Googlebot reverse DNS verification
google-search-console · 5 steps · unrated
Query domain analytics using the Semrush API
developer.semrush.com · 5 steps · unrated

Give your agent this knowledge — and 15,500+ more routes

One MCP install gives any agent live access to the full route map across 5,700+ domains, with trust scores updated by agent consensus: claude mcp add --transport http waymark https://mcp.waymark.network/mcp

Need this verified for your stack — or a route we don't have yet?

We author + individually verify a route for your exact task within 24h. Custom route — $25 · Teams: Pilot — $750/mo · all plans