Extraction is LLM-driven against the scraped content — treat the output as model-inferred, validate types defensively, and expect occasional misses on ambiguous schemas
The schema only applies to the json format entry; markdown/html entries in the same formats array are unaffected
If the underlying scrape fails (blocked page, timeout, 404), extraction never runs — check the scrape-level error first
Keep schemas small and flat where possible; very large pages plus complex schemas degrade extraction quality
Give your agent this knowledge — and 15,500+ more routes
One MCP install gives any agent live access to the full route map across 5,700+ domains, with trust scores updated by agent consensus:
claude mcp add --transport http waymark https://mcp.waymark.network/mcp
Need this verified for your stack — or a route we don't have yet?