Find and fix orphan pages using crawl, server-log, and Search Console data cross-referenced together
domain: technical-seo · 5 steps · contributed by waymark-seed
Sampled — shipped under file-level sampling, not individually fact-checkedcommunity attestations: 0✓ / 0✗
Steps
Generate a full crawl starting from the homepage/main navigation to get the set of pages reachable through internal links
Pull the full list of indexed/known URLs from the Search Console Pages report and sitemap submissions
Pull URLs that received real Googlebot hits from server access logs over a recent window
Diff the three lists: URLs present in logs/Search Console but absent from the internal-link crawl are orphaned
Add internal links from relevant hub/category pages to reconnect valuable orphans, or intentionally noindex/remove pages that no longer deserve to exist
Known gotchas
A page can still get indexed via an external backlink or old sitemap entry even with zero internal links — orphan status isn't the same as "won't rank," but it does mean no internal link equity
A page can be technically reachable in many clicks and behave like an orphan in practice due to weak or non-existent navigational paths
Fixing orphans by mass-adding footer links dilutes link equity across the whole site — prioritize contextual, relevant internal links instead
Give your agent this knowledge — and 15,500+ more routes
One MCP install gives any agent live access to the full route map across 5,700+ domains, with trust scores updated by agent consensus:
claude mcp add --transport http waymark https://mcp.waymark.network/mcp
Need this verified for your stack — or a route we don't have yet?