Skip to content

typing: give init_sd's display config keys a type, type-check the styling files - #980

Merged
paddymul merged 7 commits into
mainfrom
typing/init-sd-display-config
Sep 28, 2026
Merged

paddymul merged 7 commits into
mainfrom
typing/init-sd-display-config

Conversation

@paddymul

@paddymul paddymul commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

Closes #978.

Display config reached styling through init_sd, typed as the stats bag ColMeta. Its SDVals value type accepts a str for ag_grid_specs, and a str for delete_keys because a str is an Iterable[str]. None of the styling files were type-checked in CI.

Types

  • InitColMeta (styling_core.py) types the keys styling reads from one column's init_sd entry: displayer_args, ag_grid_specs, delete_keys, highlight_phrase / highlight_regex / highlight_color, merge_rule and column_config_override, all NotRequired. Any other key is a summary stat, typed like ColMeta's values through PEP 728 extra_items=SDVals, so ColMeta itself is unchanged. InitSD = Dict[ColIdentifier, InitColMeta].
  • delete_keys is List[str]. Sequence or Iterable would accept a bare str.
  • init_sd's displayer_args is shallow-merged over the computed one, so it's typed DisplayerArgsOverride, which has every displayer key and all of them optional.
  • PartialColConfig is now a closed TypedDict (any subset of a column config) and is the value type of OverrideColumnConfig. column_config_overrides used to require a full BaseColumnConfig, so the {'merge_rule': 'hidden'} overrides in compare.py, extension_utils.py and the pandera integration didn't type-check.
  • init_sd is annotated on BuckarooWidgetBase, BuckarooInfiniteWidget, DFViewerInfinite, PolarsDFViewerInfinite, CustomizableDataflow, create_dataflow and create_polars_dataflow.
  • The Python displayer types gain what DFWhole.ts already had: DurationDisplayerA, InheritDisplayerA, and highlight_* on StringDisplayerA.

ag_grid_specs stays Dict[str, Any], now named AGGridColDef. It's handed to AG-Grid as a ColDef. Mirroring ColDef in Python would mean copying a large third-party interface, most of which is callbacks Python can't send. The str-for-a-dict mistake from the issue is already caught by Dict.

PEP 728 at runtime. typing_extensions before 4.13 raises on closed and extra_items (4.12.2: TypeError: TypedDict takes either a dict or keyword arguments, but not both), and marimo's WASM export runs Pyodide 0.27.5, which ships 4.12.2. So PartialColConfig and InitColMeta are defined under TYPE_CHECKING and are Dict[str, Any] at runtime. basedpyright 1.39.8 handles both.

Type-checking the styling files

styling_core.py, customizations/styling.py and styling_helpers.py are added to pyrightconfig.typecheck.json. They had 39 errors on main and have 0 now. reportUnnecessaryTypeIgnoreComment is on for the whole scope and flagged nothing in the files that were already there.

tests/unit/dataflow/styling_typecheck_cases.py holds the type-level cases. basedpyright checks it, pytest doesn't collect it. Each line with a # pyright: ignore is a mistake the types have to reject, and an ignore that no longer suppresses anything is an error.

There are three casts. Two are in style_columns, where orig_col_name and column_config_override come out of the untyped merged sd. The third is in fix_column_config, which turns a BaseColumnConfig into a ColumnConfig by swapping identity keys.

Behaviour changes

  • get_left_col_configs now str()s a single-level index or columns name, as the MultiIndex branches already did. An int-named index used to reach col_path as an int ([3, 7]), while JS types col_path as string[]. The function also builds the last index column's col_path before appending it instead of patching ccs[-1] afterwards. The output is the same, and the existing index-styling tests cover it.
  • An unnamed column level next to a named one used to get the header text "None". pivot_table(index='region', columns='year', values=['revenue', 'units']) does this: its values level is unnamed and its year level is named. get_index_level_names str()'d every level name once any was set, so the index column's top header read "None". It's now blank, as it already was when no level is named. I checked this in a rendered static embed, before and after.
  • The same for rows. pd.concat({'actual': df1, 'budget': df2}) gives an unnamed outer index level above a named one. When the columns axis is named, every index column gets a col_path, and the unnamed level's header read "None". It's blank now. I checked this in a rendered static embed as well.

Int-labelled MultiIndex frames

The ddd gets three frames built the way real code gets int column levels:

  • get_multiindex_int_cols_df: pivot_table with a values list, giving ('revenue', 2023), ...
  • get_multiindex_int_levels_df: unstack onto year/quarter, giving (2023, 1), ...
  • get_multiindex_int_names_df: a pivot of a headerless CSV, which has int level names as well as int values.
  • get_multiindex_partly_named_index_df: pd.concat of a dict of year pivots, giving an unnamed outer row level above region and a columns axis named year.

The tests cover their left col_paths, the data columns' col_path, and column_config_overrides keyed by int-containing tuples.

A data column's col_path is still the frame's own label, ints included. That's what merge_column_config looks overrides up by, and in the rendered page an int header such as 2023 displays fine. So the str() above only applies to level names on the index side.

Not covered here

  • For any row MultiIndex, the pinned summary rows show None in the index cells instead of the stat names (dtype, mean, ...), whether or not the levels are named. That's a separate, pre-existing bug and isn't touched here.

  • Widget and dataflow constructor calls still aren't argument-checked. They're traitlets HasTraits classes, and HasDescriptors.__new__ returns Any, so pyright never evaluates __init__ for a BuckarooWidget(...) call. Editors show the new annotations, and the constructor body is typed. A caller's bad init_sd is only flagged when it goes through create_dataflow / create_polars_dataflow. A TYPE_CHECKING-only __new__ -> Self on the base classes would fix that. It would also start checking pinned_rows, which is annotated PinnedRowConfig although callers pass a list.

  • style_column still receives the merged sd as ColMeta. InitColMeta types init_sd where it enters, not where styling reads it.

  • The typecheck job is still non-blocking. The scoped files are at 0 errors, so it could be made blocking.

  • Runtime validation of these keys is the fast follow from the issue.

Test

  • 1d90531 adds the type cases, adds the styling files to the typecheck config, and adds test_index_styling_non_str_level_names. On CI, every Python test job failed on that test and nothing else (3.13: 1 failed, 1163 passed). The typecheck job is non-blocking and stays green, but its log showed 51 errors.
  • b280eac is the fix. Locally, basedpyright reports 0 errors, pytest tests/unit gives 1275 passed and 6 skipped, ruff and paddy_format are clean, and buckaroo.dataflow.styling_core imports under typing_extensions 4.12.2.
  • f6bbf8f adds the int-labelled ddd frames and test_index_styling_int_cols_unnamed_level, which fails on the "None" header. On CI, every Python test job failed on that test and nothing else (3.13: 1 failed, 1168 passed).
  • 5dd7349 adds int-label coverage tests that already pass.
  • 7556af5 is the "None" header fix. Locally, pytest tests/unit gives 1280 passed and 6 skipped, and every CI check passes.
  • 326ad64 adds the partly named row index frame and two failing tests, one under a named columns axis and one under MultiIndex columns. On CI, every Python test job failed on those two tests and nothing else (3.13: 2 failed, 1169 passed).
  • fbe7f89 is the row fix. Locally, pytest tests/unit gives 1282 passed and 6 skipped, and every CI check passes.

🤖 Generated with Claude Code

…ck the styling files

styling_typecheck_cases.py is checked by basedpyright, not pytest. Each line
with a `# pyright: ignore` is a mistake the types have to reject; with
reportUnnecessaryTypeIgnoreComment on, an ignore that suppresses nothing is an
error. Today init_sd has no type, so every one of them is accepted, and the
good cases fail because merge_rule-only overrides, the inherit and duration
displayers and the string displayer's highlight_* keys aren't in the Python
types.

styling_core.py, customizations/styling.py and styling_helpers.py join
pyrightconfig.typecheck.json.

test_index_styling_non_str_level_names: an int-named single-level index or
columns level reaches col_path as an int, but col_path is string[] on the JS
side and the MultiIndex branches already str() their level names.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

📦 TestPyPI package published

pip install --index-strategy unsafe-best-match --index-url https://test.pypi.org/simple/ --extra-index-url https://pypi.org/simple/ buckaroo==0.15.6.dev36153886336

or with uv:

uv pip install --index-strategy unsafe-best-match --index-url https://test.pypi.org/simple/ --extra-index-url https://pypi.org/simple/ buckaroo==0.15.6.dev36153886336

MCP server for Claude Code

claude mcp add buckaroo-table -- uvx --from "buckaroo[mcp]==0.15.6.dev36153886336" --index-strategy unsafe-best-match --index-url https://test.pypi.org/simple/ --extra-index-url https://pypi.org/simple/ buckaroo-table

📖 Docs preview

🎨 Storybook preview

…ling files

InitColMeta types the keys styling reads out of an init_sd entry
(displayer_args, ag_grid_specs, delete_keys, highlight_*, merge_rule,
column_config_override); every other key is a summary stat typed like
ColMeta's values, via PEP 728 extra_items. delete_keys is a List because
a bare str is an Iterable[str]. init_sd is annotated InitSD on the
widgets, CustomizableDataflow and the two server create_* functions.

PartialColConfig becomes a closed TypedDict and is what
column_config_overrides now takes, so a merge_rule-only override (compare,
extension_utils, pandera) type-checks. Both PEP 728 types are defined
under TYPE_CHECKING: typing_extensions raises on closed / extra_items
before 4.13 and Pyodide 0.27 ships 4.12, so at runtime they're plain dicts.

ag_grid_specs stays Dict[str, Any] under an AGGridColDef alias: it's
handed to AG-Grid as a ColDef, a large third-party interface that's
mostly callbacks Python can't send.

The Python displayer types pick up what the JS side already has: the
duration and inherit displayers and the string displayer's highlight_*
keys.

styling_core.py, customizations/styling.py and styling_helpers.py now
type-check clean. get_left_col_configs str()s a single-level index or
columns name like the MultiIndex branches already did, and builds the
last index column's col_path before appending it instead of patching it
afterwards.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
paddymul and others added 2 commits September 25, 2026 10:33
…t for an unnamed column level's header

Three ddd frames built the way real code gets int column levels:
pivot_table with a values list (('revenue', 2023), ...), unstack onto
year/quarter, and a pivot of a headerless CSV (int level names too).

The pivot_table frame leaves the values level unnamed next to a named
'year' level. get_index_level_names str()s every name once any is set,
so the index column's top header renders the text "None".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
These pass already. Left col_path for int level values and int level
names, data columns keeping the frame's own int-containing labels as
col_path, and column_config_overrides keyed by those labels (a hidden
merge_rule and a color_map_config).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
get_index_level_names str()'d every level name once any level was named,
so pivot_table(values=[...], columns='year') put the text "None" in the
index column's top header. An unnamed level now gets '', as it already
did when no level is named.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…med ones

get_multiindex_partly_named_index_df is pd.concat of a dict of year
pivots: the dict keys make an unnamed outer index level above 'region',
and the columns axis is named 'year'. With a named columns axis every
index column takes the col_path branch, which str()s the None level name,
so the first index column's header reads "None". The same happens under
MultiIndex columns.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d ones

With a named columns axis, get_left_col_configs gives every index column a
col_path and appended str(idx_name), so the unnamed outer level pd.concat
makes from a dict's keys got the header "None". It's '' now, matching the
column-level fix in get_index_level_names.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@paddymul
paddymul added this pull request to the merge queue Sep 28, 2026
Merged via the queue into main with commit fd58dad Sep 28, 2026
28 checks passed

This branch was successfully deployed

1 active deployment
testpypi — fbe7f891 Deployed Sep 25, 2026 by paddymul via Publish to TestPyPI #1582
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

typing: display config rides in the untyped stats dict — init_sd keys have no type, styling files aren't type-checked

1 participant