Migrate MatchVariables to narwhals, add polars support - #1063
Merged
Merged
Conversation
fit() and transform() take any narwhals-supported dataframe and return
the same backend. pandas adds, drops and reorders the columns with a
single reindex (faster than the previous drop/setitem/select and than
narwhals at 10k-500k rows); other backends use with_columns + select.
With match_dtypes, pandas keeps its dtypes and astype, since narwhals
dtypes don't hold the categories of pandas categoricals. Other
backends store narwhals dtypes and cast with narwhals, turning values
outside Enum categories into nulls and parsing strings into dates, as
pandas does.
With polars, np.nan fill values add null Float64 columns and integer
fill values add Int64 columns. The verbose messages list the variables
in a fixed order: training order for added variables, input order for
dropped ones. Init error messages now end with "Got {param} instead.".
Rewrite the tests to the make_df conventions and add a polars example
to the docstring and the user guide, refreshing its outputs.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
solegalli
force-pushed
the
narwhals-match-variables
branch
from
September 19, 2026 09:19
f9841b8 to
084354a
Compare
solegalli
commented
Sep 19, 2026
| contain a dictionary of variables and their corresponding dtypes. | ||
| contain a dictionary of variables and their corresponding dtypes. With pandas, | ||
| these are pandas dtypes. With other dataframe libraries, like polars, these | ||
| are narwhals dtypes, which carry the same names as the polars dtypes. |
Collaborator
Author
There was a problem hiding this comment.
we can't have narwhals in user facing documentation.
solegalli
commented
Sep 19, 2026
| if nwd.is_pandas_dataframe(X) is True: | ||
| self.dtype_dict_: Dict = X.dtypes.to_dict() | ||
| else: | ||
| self.dtype_dict_ = dict(nw_X.schema) |
Collaborator
Author
There was a problem hiding this comment.
do we need this parameter to be visible to the user? It feels like it's of no use for them, just for the transformer to work, so we could make this a hidden parameter and then we don't need to document this in the rst, where we mention narwhals at the moment.
…ata in integer and boolean variables Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Collaborator
Author
|
Changes after review (commit ebb8715):
Note for the release notes: |
Collaborator
Author
|
waiting for #1068 I imagine that the changes in column order in the rst would not occur once we have that functionality in |
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Migrates
MatchVariablesto narwhals.fitandtransformtake pandas, polars or any other narwhals-supported dataframe and return the same backend.X.reindex(columns=feature_names_in_, fill_value=fill_value), replacing the previous drop, setitem and select.match_dtypeskeeps pandas dtypes andastype, because narwhals dtypes can't hold the categories of a pandasCategoricalDtype.dtype_dict_still holds pandas dtypes.with_columns(lit)for the added columns, thenselect(feature_names_in_).match_dtypesstoresdict(nw_X.schema)and casts with narwhals. Where polars would otherwise raise or differ from pandas, the code:Enumto null. pandas sets them to NaN; polars' strict cast raises.DatetimeorDatewithstr.to_datetime()orstr.to_date(). polars' plain cast from string to datetime is deprecated and fails on"2020-02-24 00:00:00".fill_value=np.nanadds nullFloat64columns on polars instead of NaN. polars treats NaN as a value, so imputers andis_nullwould miss it. An integerfill_valueaddsInt64columns;nw.lit(1)would giveInt32on polars, and pandas gives int64._check_param_missing_values, and every message now ends withGot {param} instead.. The old messages had a stray'and a missing space.make_df,frame_to_dict,null_count, full-messagematch=, init tests first. The data dicts stay in the test file, so the shared conftest is untouched.np.nanoutput table is fixed. The user guide gets a "Working with polars" section, and its existing outputs are refreshed from a real run (see below).Benchmarks
Median of repeats, candidates alternated in rounds, run on a machine that was also running other jobs. Whole
transform()withmissing_values="ignore"isolates the column work. Times in ms.Synthetic data: half float, a quarter int and a quarter string columns. Scenarios:
identical: same columns.reordered: reversed column order.add2_drop2: 2 train columns missing, 2 extra columns, interleaved order.pandas (before = previous drop/setitem/select code; narwhals =
with_columns+select(nw.col(*names)), which integer column names need):reindexis fastest in 24 of 27 cells, and 2-3x faster than before on the common case where the columns already match. narwhals only wins when columns are added at 500k rows. When pandas has to fill new columns,reindexcopies the frame; narwhals keeps the new columns as separate blocks.polars, whole
transform()(narwhals vs the samewith_columns+selectwritten in native polars):Native polars is 0.01-0.08 ms faster. That is a fixed per-call overhead that doesn't grow with the number of rows. With the default
missing_values="raise", the NaN check alone takes 5-20 ms at these sizes, so the gap is under 1% oftransform. A third code path isn't worth it, and it would still need the narwhals path for the other backends. Narwhals and nativecastformatch_dtypestake the same time (2M x 200, 50 casts: 49.6 vs 50.9 ms).With the default
missing_values="raise",_check_contains_natakes over 95% of the pandastransformtime: 380 ms at 500k x 200, against 0.1 ms for the column work. That is the shared helper, which this PR doesn't touch; see "Pre-existing issues".Behaviour
pandas: identical. Before migrating, I recorded outputs of the code from before #1019, since
origin/narwhals-migrationcurrently fails on everyMatchVariablestest. I checked 40 scenarios:fill_value(NaN, int, float, str, bool, negative), with and withoutmatch_dtypestransformmatch_dtypesfor string/number/datetime/object in both directions, float↔int, a bad string to int, and categories that are missing, extra, from string and unseenFrames, dtypes,
dtype_dict_,feature_names_in_,get_feature_names_out, raised errors and "input not modified" are all identical. The only difference is the order of names in the verbose messages. It used to come from set differences, so it could change between runs. It is now fixed: added variables in train order, dropped variables in input order. The user guide outputs are updated to match.polars vs pandas: the same values in all scenarios, except the items under "Needs decision". The dtypes map as follows:
Tests
tests/test_preprocessing, one pytest call, baseorigin/narwhals-migrationvs this branch:MatchVariablestests failed withAttributeError: 'list' object has no attribute 'tolist', becausecheck_Xnow returns a narwhals frame.MatchCategories(migrated in a separate PR): 3 intest_match_categories.py, plus itscheck_estimator_from_feature_engineandset_outputtests.test_check_estimator_from_sklearn, which fails for every migrated transformer because sklearn passes numpy arrays tocheck_X.No other module imports
match_columns.flake8 feature_engine testsis clean.mypy feature_enginereports the same 2 errors as the base, indatetime_subtraction.pyandlog.py.Needs decision
fill_value=np.nanon polars adds nulls, not NaN. I chose null because it is polars' missing value; with NaN,match_dtypeswould also turn the column into the string"NaN"when the train dtype is String. The alternative is to add the literal NaN (Float64).dtype_dict_holds narwhals dtypes for non-pandas input ({"a": Int64, ...}). They print like the polars dtypes, butdtype_dict_["a"] == pl.Int64isFalse. Storing native polars dtypes would need a polars-only branch for both fit and cast.inttrain column that is missing fromtransform, withmatch_dtypes=Trueand the default NaNfill_value:IntCastingNaNError, as before.Int64column, because polars integers can hold nulls.2020-02-24 00:00:00.000000and pandas writes2020-02-24 00:00:00. This is each library's default format."tom"to int) raiseValueErroron pandas and polars'InvalidOperationErroron polars.Pre-existing issues, not fixed
IntCastingNaNErrorabove: withmatch_dtypes=True, a missing integer variable filled with NaN makes the wholetransformfail. This hits the typical use case, where a variable is absent at prediction time. A fix, such as casting to a nullable integer or skipping added columns, would change behaviour._check_contains_na(shared helper) dominates the pandastransformtime: 380 ms at 500k x 200. A pandas-nativeisna().any()fast path there would likely speed up every transformer that uses it.fill_valueacceptsbool(it is anint), and it rejects numpy scalars such asnp.int64(1). I left both unchanged.