You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
read_dta now sets TableFunction::filter_pushdown = true. At scan init the pushed TableFilterSet is converted into one boolean expression over the projected columns via TableFilter::ToExpression, so all filter shapes DuckDB pushes are handled uniformly — constant comparisons, AND/OR conjunctions, IS [NOT] NULL, IN lists, and dynamic/optional filters (dynamic filters degrade to true when uninitialized, which is safe since they're advisory). Each scan thread evaluates the predicate with its own ExpressionExecutor and slices the output chunk by the selection vector; row ranges that filter to zero rows loop on to the next range instead of returning an empty chunk (which would end the scan).
Plan verification — the filter lands inside the scan, no separate FILTER operator:
Tests in test/sql/filter_pushdown.test cover range/equality/conjunction filters, IS [NOT] NULL against Stata missing values, string and ENUM (value-label) comparisons, DATE filters on converted %td columns, filters that skip whole chunks, and filters matching nothing. Full suite: 599 assertions passing.
Note this composes with the projection pushdown and parallel scan that landed on the same branch (66eabf1). Since the row format is row-major, filtered-out rows are still read from disk within each claimed range — the win is that they're dropped at the scan and never flow through the pipeline. Skipping decode of non-filter columns for rejected rows (filter_prune + two-phase decode) is a possible follow-up refinement. Leaving open until merged.
duckdb/duckdb#9644