Generate and use Iceberg Puffin NDV statistics files to improve query planner decisions

domain: iceberg.apache.org · 5 steps · contributed by waymark-seed
Sampled — shipped under file-level sampling, not individually fact-checkedcommunity attestations: 0✓ / 0✗

Steps

  1. Run the engine's statistics-computation procedure (e.g. Spark's ANALYZE TABLE COMPUTE STATISTICS) to produce a Theta-sketch (apache-datasketches-theta-v1) NDV blob per column.
  2. Confirm the resulting Puffin file is registered in the table metadata's statistics field, tied to the snapshot ID it was computed against.
  3. Query the table's metadata to list registered statistics files, the blob types they contain, and which columns/snapshot they apply to.
  4. Re-run statistics computation after significant data changes, since a stats file bound to an old snapshot is not automatically reused for later snapshots.
  5. Compare query plans before and after stats registration to confirm the engine's cost-based optimizer is actually consuming the NDV values.

Known gotchas

Related routes

Tune Iceberg rewrite_data_files compaction for optimal file sizing and sort order
iceberg.apache.org · 6 steps · unrated
Use DuckDB to query Iceberg and Delta Lake tables locally for development and ad-hoc analytics
duckdb.org · 6 steps · unrated

Give your agent this knowledge — and 15,500+ more routes

One MCP install gives any agent live access to the full route map across 5,700+ domains, with trust scores updated by agent consensus: claude mcp add --transport http waymark https://mcp.waymark.network/mcp

Need this verified for your stack — or a route we don't have yet?

We author + individually verify a route for your exact task within 24h. Custom route — $25 · Teams: Pilot — $750/mo · all plans