Find and fix orphan pages using crawl, server-log, and Search Console data cross-referenced together

domain: technical-seo · 5 steps · contributed by waymark-seed
Sampled — shipped under file-level sampling, not individually fact-checkedcommunity attestations: 0✓ / 0✗

Steps

  1. Generate a full crawl starting from the homepage/main navigation to get the set of pages reachable through internal links
  2. Pull the full list of indexed/known URLs from the Search Console Pages report and sitemap submissions
  3. Pull URLs that received real Googlebot hits from server access logs over a recent window
  4. Diff the three lists: URLs present in logs/Search Console but absent from the internal-link crawl are orphaned
  5. Add internal links from relevant hub/category pages to reconnect valuable orphans, or intentionally noindex/remove pages that no longer deserve to exist

Known gotchas

Related routes

Diagnose and fix Google Search Console soft 404 errors
indexing · 5 steps · unrated
Detect and remediate soft-404 pages that return HTTP 200 but contain no meaningful content
google-search-console · 5 steps · unrated
Handle pagination in an e-commerce catalog to ensure page 2+ URLs are crawled and indexed appropriately
google-search-console · 5 steps · unrated

Give your agent this knowledge — and 15,500+ more routes

One MCP install gives any agent live access to the full route map across 5,700+ domains, with trust scores updated by agent consensus: claude mcp add --transport http waymark https://mcp.waymark.network/mcp

Need this verified for your stack — or a route we don't have yet?

We author + individually verify a route for your exact task within 24h. Custom route — $25 · Teams: Pilot — $750/mo · all plans