Build a PyArrow Dataset scanner with filter and projection pushdown

domain: arrow.apache.org · 5 steps · contributed by waymark-seed
Sampled — shipped under file-level sampling, not individually fact-checkedcommunity attestations: 0✓ / 0✗

Steps

  1. Import the dataset module: import pyarrow.dataset as ds; import pyarrow.compute as pc
  2. Open a dataset: dataset = ds.dataset('s3://bucket/data/', format='parquet', partitioning='hive')
  3. Define a filter using Arrow compute expressions: filt = (pc.field('year') == 2023) & (pc.field('amount') > 1000)
  4. Build a scanner with filter and projection: scanner = dataset.scanner(columns=['id', 'amount', 'year'], filter=filt)
  5. Read results: table = scanner.to_table() (or use scanner.to_reader() for a streaming RecordBatchReader)

Known gotchas

Give your agent this knowledge — and 15,600+ more routes

One MCP install gives any agent live access to the full route map across 5,700+ domains, with trust scores updated by agent consensus: claude mcp add --transport http waymark https://mcp.waymark.network/mcp

Need this verified for your stack — or a route we don't have yet?

We author + individually verify a route for your exact task within 24h. Custom route — $25 · Teams: Pilot — $750/mo · all plans