Skip to content

Migrate MatchVariables to narwhals, add polars support - #1063

Merged
solegalli merged 3 commits into
narwhals-migrationfrom
narwhals-match-variables
Sep 19, 2026
Merged

solegalli merged 3 commits into
narwhals-migrationfrom
narwhals-match-variables

Conversation

@solegalli

Copy link
Copy Markdown
Collaborator

Summary

Migrates MatchVariables to narwhals. fit and transform take pandas, polars or any other narwhals-supported dataframe and return the same backend.

  • pandas: adds, drops and reorders the columns with a single X.reindex(columns=feature_names_in_, fill_value=fill_value), replacing the previous drop, setitem and select. match_dtypes keeps pandas dtypes and astype, because narwhals dtypes can't hold the categories of a pandas CategoricalDtype. dtype_dict_ still holds pandas dtypes.
  • other backends: with_columns(lit) for the added columns, then select(feature_names_in_). match_dtypes stores dict(nw_X.schema) and casts with narwhals. Where polars would otherwise raise or differ from pandas, the code:
    • sets values outside the categories of an Enum to null. pandas sets them to NaN; polars' strict cast raises.
    • parses strings into Datetime or Date with str.to_datetime() or str.to_date(). polars' plain cast from string to datetime is deprecated and fails on "2020-02-24 00:00:00".
  • fill_value=np.nan adds null Float64 columns on polars instead of NaN. polars treats NaN as a value, so imputers and is_null would miss it. An integer fill_value adds Int64 columns; nw.lit(1) would give Int32 on polars, and pandas gives int64.
  • Init validation uses _check_param_missing_values, and every message now ends with Got {param} instead.. The old messages had a stray ' and a missing space.
  • Tests are rewritten to the conventions: make_df, frame_to_dict, null_count, full-message match=, init tests first. The data dicts stay in the test file, so the shared conftest is untouched.
  • The docstring gets a polars example, and a broken np.nan output table is fixed. The user guide gets a "Working with polars" section, and its existing outputs are refreshed from a real run (see below).

Benchmarks

Median of repeats, candidates alternated in rounds, run on a machine that was also running other jobs. Whole transform() with missing_values="ignore" isolates the column work. Times in ms.

Synthetic data: half float, a quarter int and a quarter string columns. Scenarios:

  • identical: same columns.
  • reordered: reversed column order.
  • add2_drop2: 2 train columns missing, 2 extra columns, interleaved order.

pandas (before = previous drop/setitem/select code; narwhals = with_columns + select(nw.col(*names)), which integer column names need):

rows cols scenario before reindex (chosen) narwhals
10k 10 identical 0.17 0.07 0.22
10k 10 add2_drop2 0.31 0.12 0.58
10k 200 identical 0.36 0.13 2.52
10k 200 add2_drop2 0.69 0.39 5.50
100k 50 identical 0.21 0.08 0.72
100k 50 reordered 0.23 0.11 0.71
100k 50 add2_drop2 1.01 0.79 1.65
100k 200 add2_drop2 2.80 2.38 5.65
500k 10 add2_drop2 0.97 0.78 0.69
500k 50 identical 0.22 0.08 0.71
500k 50 add2_drop2 3.89 3.67 1.78
500k 200 identical 0.38 0.13 2.68
500k 200 add2_drop2 12.07 11.48 5.91

reindex is fastest in 24 of 27 cells, and 2-3x faster than before on the common case where the columns already match. narwhals only wins when columns are added at 500k rows. When pandas has to fill new columns, reindex copies the frame; narwhals keeps the new columns as separate blocks.

polars, whole transform() (narwhals vs the same with_columns + select written in native polars):

rows cols scenario narwhals (chosen) polars-native
500k 10 identical 0.08 0.06
500k 10 add2_drop2 0.18 0.11
500k 50 identical 0.16 0.15
500k 200 add2_drop2 0.47 0.39
2M 50 reordered 0.17 0.15
2M 200 add2_drop2 0.42 0.36

Native polars is 0.01-0.08 ms faster. That is a fixed per-call overhead that doesn't grow with the number of rows. With the default missing_values="raise", the NaN check alone takes 5-20 ms at these sizes, so the gap is under 1% of transform. A third code path isn't worth it, and it would still need the narwhals path for the other backends. Narwhals and native cast for match_dtypes take the same time (2M x 200, 50 casts: 49.6 vs 50.9 ms).

With the default missing_values="raise", _check_contains_na takes over 95% of the pandas transform time: 380 ms at 500k x 200, against 0.1 ms for the column work. That is the shared helper, which this PR doesn't touch; see "Pre-existing issues".

Behaviour

pandas: identical. Before migrating, I recorded outputs of the code from before #1019, since origin/narwhals-migration currently fails on every MatchVariables test. I checked 40 scenarios:

  • every fill_value (NaN, int, float, str, bool, negative), with and without match_dtypes
  • reorder only, identical columns, no column in common
  • NaN in fit, in transform, in dropped columns and in columns missing from transform
  • integer column names, and a non-default index
  • match_dtypes for string/number/datetime/object in both directions, float↔int, a bad string to int, and categories that are missing, extra, from string and unseen

Frames, dtypes, dtype_dict_, feature_names_in_, get_feature_names_out, raised errors and "input not modified" are all identical. The only difference is the order of names in the verbose messages. It used to come from set differences, so it could change between runs. It is now fixed: added variables in train order, dropped variables in input order. The user guide outputs are updated to match.

polars vs pandas: the same values in all scenarios, except the items under "Needs decision". The dtypes map as follows:

pandas polars
float64 Float64
int64 Int64
str/string String
datetime64[us] Datetime(us)
category Enum

Tests

tests/test_preprocessing, one pytest call, base origin/narwhals-migration vs this branch:

  • before: 27 failed, 8 passed. All MatchVariables tests failed with AttributeError: 'list' object has no attribute 'tolist', because check_X now returns a narwhals frame.
  • after: 7 failed, 90 passed. None of the remaining failures are new:
    • 5 are MatchCategories (migrated in a separate PR): 3 in test_match_categories.py, plus its check_estimator_from_feature_engine and set_output tests.
    • 2 are test_check_estimator_from_sklearn, which fails for every migrated transformer because sklearn passes numpy arrays to check_X.

No other module imports match_columns. flake8 feature_engine tests is clean. mypy feature_engine reports the same 2 errors as the base, in datetime_subtraction.py and log.py.

Needs decision

  1. fill_value=np.nan on polars adds nulls, not NaN. I chose null because it is polars' missing value; with NaN, match_dtypes would also turn the column into the string "NaN" when the train dtype is String. The alternative is to add the literal NaN (Float64).
  2. dtype_dict_ holds narwhals dtypes for non-pandas input ({"a": Int64, ...}). They print like the polars dtypes, but dtype_dict_["a"] == pl.Int64 is False. Storing native polars dtypes would need a polars-only branch for both fit and cast.
  3. An int train column that is missing from transform, with match_dtypes=True and the default NaN fill_value:
    • pandas raises IntCastingNaNError, as before.
    • polars returns a null Int64 column, because polars integers can hold nulls.
  4. Datetime cast to string: polars writes 2020-02-24 00:00:00.000000 and pandas writes 2020-02-24 00:00:00. This is each library's default format.
  5. Failed casts (for example "tom" to int) raise ValueError on pandas and polars' InvalidOperationError on polars.

Pre-existing issues, not fixed

  • The pandas IntCastingNaNError above: with match_dtypes=True, a missing integer variable filled with NaN makes the whole transform fail. This hits the typical use case, where a variable is absent at prediction time. A fix, such as casting to a nullable integer or skipping added columns, would change behaviour.
  • _check_contains_na (shared helper) dominates the pandas transform time: 380 ms at 500k x 200. A pandas-native isna().any() fast path there would likely speed up every transformer that uses it.
  • fill_value accepts bool (it is an int), and it rejects numpy scalars such as np.int64(1). I left both unchanged.

fit() and transform() take any narwhals-supported dataframe and return
the same backend. pandas adds, drops and reorders the columns with a
single reindex (faster than the previous drop/setitem/select and than
narwhals at 10k-500k rows); other backends use with_columns + select.

With match_dtypes, pandas keeps its dtypes and astype, since narwhals
dtypes don't hold the categories of pandas categoricals. Other
backends store narwhals dtypes and cast with narwhals, turning values
outside Enum categories into nulls and parsing strings into dates, as
pandas does.

With polars, np.nan fill values add null Float64 columns and integer
fill values add Int64 columns. The verbose messages list the variables
in a fixed order: training order for added variables, input order for
dropped ones. Init error messages now end with "Got {param} instead.".

Rewrite the tests to the make_df conventions and add a polars example
to the docstring and the user guide, refreshing its outputs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@solegalli
solegalli force-pushed the narwhals-match-variables branch from f9841b8 to 084354a Compare September 19, 2026 09:19
contain a dictionary of variables and their corresponding dtypes.
contain a dictionary of variables and their corresponding dtypes. With pandas,
these are pandas dtypes. With other dataframe libraries, like polars, these
are narwhals dtypes, which carry the same names as the polars dtypes.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we can't have narwhals in user facing documentation.

if nwd.is_pandas_dataframe(X) is True:
self.dtype_dict_: Dict = X.dtypes.to_dict()
else:
self.dtype_dict_ = dict(nw_X.schema)

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

do we need this parameter to be visible to the user? It feels like it's of no use for them, just for the transformer to work, so we could make this a hidden parameter and then we don't need to document this in the rst, where we mention narwhals at the moment.

…ata in integer and boolean variables

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@solegalli

Copy link
Copy Markdown
Collaborator Author

Changes after review (commit ebb8715):

  • The learned dtypes are now private. dtype_dict_ became _dtype_dict, and its entry in the docstring is removed. It only holds what transform needs, so users don't need to see it, and the docs no longer mention narwhals.
  • Integer and boolean variables with missing data no longer fail on pandas. This covers variables added at transform time with a NaN fill_value, and any other integer or boolean variable with missing values, when match_dtypes=True. pandas raised IntCastingNaNError for integers, and silently turned NaN into True for booleans. They now use pandas' nullable dtypes (Int64, UInt8, ..., boolean), which is the pandas equivalent of polars' null integers and booleans. Variables without missing data keep their numpy dtypes. The new test test_match_dtypes_of_added_integer_and_boolean_variables runs on both backends and fails on pandas without the fix.
  • Accepted as they are: points 1, 4 and 5 of "Needs decision".

Note for the release notes: dtype_dict_ was a public attribute in 1.9.4, so removing it is an API change.

@solegalli

Copy link
Copy Markdown
Collaborator Author

waiting for #1068 I imagine that the changes in column order in the rst would not occur once we have that functionality in

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@solegalli
solegalli merged commit c8ae828 into narwhals-migration Sep 19, 2026
4 of 10 checks passed
@solegalli
solegalli deleted the narwhals-match-variables branch September 19, 2026 10:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant