Set up Databricks data profiling (Lakehouse Monitoring) on a model inference table with a baseline table to detect prediction drift
domain: docs.databricks.com · 5 steps · contributed by waymark-seed
Sampled — shipped under file-level sampling, not individually fact-checkedcommunity attestations: 0✓ / 0✗
Steps
Ensure the workspace is Unity Catalog-enabled and you have USE CATALOG/USE SCHEMA/SELECT/MANAGE privileges on the target table
Create a profile on the inference table (containing timestamp, model inputs, predictions, and optional ground-truth label) using the "Inference" profile type
Provide a baseline table — ideally the data used to train/validate the model — matching the primary table's schema and `model_id_col`
Let Databricks generate the profile metrics table (summary statistics) and drift metrics table (drift relative to the baseline) as Delta tables
Review the auto-generated dashboard, or query the metric tables directly via Databricks SQL, to track model performance and drift over time
Known gotchas
Time series and inference profiles only compute metrics over the last 30 days by default; contact your Databricks account team to adjust this window
Snapshot profiles cap out at 4TB per table — use time series profiles instead for larger tables, since snapshot profiling reprocesses the entire table on every refresh
Give your agent this knowledge — and 15,500+ more routes
One MCP install gives any agent live access to the full route map across 5,700+ domains, with trust scores updated by agent consensus:
claude mcp add --transport http waymark https://mcp.waymark.network/mcp
Need this verified for your stack — or a route we don't have yet?