Skip to content

Migrate Winsoriser/Winsorizer to narwhals, add polars support - #1036

Merged
solegalli merged 7 commits into
narwhals-migrationfrom
narwhals-winsorizer
Sep 19, 2026
Merged

solegalli merged 7 commits into
narwhals-migrationfrom
narwhals-winsorizer

Conversation

@solegalli

Copy link
Copy Markdown
Collaborator

Migrates Winsoriser / Winsorizer to narwhals with polars support.

WinsorizerBase.fit/transform (shared base) are migrated on narwhals-outliers-base. This covers the Winsoriser-specific piece: transform()'s add_indicators path, which compares the capped output against the original input to build per-tail boolean flag columns and previously only worked on pandas. Module-level import pandas/numpy removed; X type hints use IntoDataFrame.

Merge vs split: benchmarked the add_indicators comparison+concat step at 10k/50k/100k rows × 1/2/10 cols. pandas-native (boolean comparison + pd.concat) is up to ~3x faster than the narwhals with_columns equivalent on pandas input, and the loss grows with column count (10 cols: ~2–3x slower). That crosses the "keep pandas fast path" threshold, so transform() splits on nwd.is_pandas_dataframe, matching MissingIndicator's precedent: pandas keeps its comparison+concat logic (obtaining pd via nw.from_native(...).__native_namespace__() instead of importing it); a new narwhals with_columns path (per-column Series comparison, cast to Float64) covers polars and other backends.

Deprecation preserved exactly: Winsoriser is the current public name (British spelling, #967); Winsorizer is a deprecated subclass raising FutureWarning, removal in 2.1.0 — the reverse of what the class names suggest.

Tests: test_winsorizer.py converted from pandas-only fixtures to local dicts parametrized over make_df in [pd.DataFrame, pl.DataFrame], asserting identical capping values, indicator columns and get_feature_names_out() on both backends. Missing-value dicts use None not np.nan in string columns (polars rejects float NaN in a string column).

Verified: test_winsorizer.py 93 passed; full tests/test_outliers 123 passed / 3 pre-existing check_estimator failures (identical against a narwhals-outliers-base baseline run). flake8 / mypy clean, sphinx -W clean. Winsoriser.rst examples verified against real output (house_prices dataset available), "With polars" sections added. Full polars fit_transform incl. add_indicators runs with pandas import blocked.


Stacked on narwhals-outliers-base (its own PR). Until that merges this PR's diff also contains the shared BaseOutlier / WinsorizerBase commit; review that one first.

@solegalli

Copy link
Copy Markdown
Collaborator Author

Updated this branch:

Locally: test_winsorizer.py 93 passed; no new failures in tests/test_outliers. flake8 and mypy clean.

@solegalli
solegalli force-pushed the narwhals-winsorizer branch 2 times, most recently from 0195196 to 4e39d57 Compare September 19, 2026 06:42
solegalli and others added 6 commits September 19, 2026 08:53
Removed the module-level `import pandas as pd` and `import numpy as np`;
X type hints now use narwhals' IntoDataFrame. WinsorizerBase.fit/transform
(shared base) were already migrated on origin/narwhals-outliers-base; this
change covers the Winsoriser-specific piece: transform()'s add_indicators
path, which compares the capped output against the original input to build
per-tail boolean flag columns and previously only worked on pandas.

Benchmarked the add_indicators comparison+concat step at 10k/50k/100k rows
x 1/2/10 columns: pandas-native (boolean comparison + pd.concat) is up to
~3x faster than the narwhals with_columns equivalent on pandas input, and
the loss grows with column count (1 col: narwhals-on-pandas was actually
faster; 10 cols: ~2-3x slower). That crosses the "keep pandas fast path"
threshold, so transform() splits on `nwd.is_pandas_dataframe`, matching
MissingIndicator's precedent for its own indicator-building step: pandas
keeps its existing comparison+concat logic (now obtaining the `pd` module
via `nw.from_native(...).__native_namespace__()` instead of importing it),
and a new narwhals with_columns path (per-column Series comparison, cast to
Float64) covers polars and other backends.

Preserved the Winsoriser/Winsorizer deprecation exactly as-is: Winsoriser
is the current public name (renamed to the British spelling in #967);
Winsorizer is a deprecated subclass that raises the same FutureWarning on
__init__ and will be removed in 2.1.0. Note this is the reverse of what
one might guess from the class names alone.

Tests: converted tests/test_outliers/test_winsorizer.py from pandas-only
fixtures (df_normal_dist, df_vartypes, df_na) to local dicts parametrized
over `make_df` in [pd.DataFrame, pl.DataFrame], asserting identical capping
values, indicator columns, and get_feature_names_out() on both backends for
the same input. Missing-value dicts use None instead of np.nan in string
columns, since polars' DataFrame constructor rejects a float NaN mixed into
a string column. A helper filters both pandas' NaN and polars' None
representations of a missing value when comparing outputs cross-backend.

Docs: verified every doc example in docs/user_guide/outliers/Winsoriser.rst
against actual output (network access to fetch_openml's house_prices
dataset was available; outputs matched exactly, no changes needed) and
added a "With polars" section covering add_indicators, matching the
pattern used in other migrated user guides. Added a verified "With polars"
example to the class docstring.

Verified: tests/test_outliers/test_winsorizer.py 93 passed. Full
tests/test_outliers suite: 123 passed / 3 pre-existing failures in
test_check_estimator_outliers.py (confirmed identical against a baseline
run of origin/narwhals-outliers-base: 83 passed / same 3 failures -
sklearn's check_estimator feeds raw numpy arrays, which check_X() has
always rejected per the narwhals migration's dataframe-only contract;
predates this change). flake8 and mypy clean. sphinx -W build clean (only
the pre-existing unrelated linkcode_resolve warning, confirmed present on
the base branch too). Confirmed winsorizer.py and base_outlier.py import
successfully and a full polars fit_transform (including add_indicators)
runs correctly with pandas' own import blocked at the builtins level.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
With add_indicators=True, transform() compared the capped output against
check_X(X), which is now a narwhals frame, so the pandas path mixed pandas
and narwhals objects (broadcast errors, wrong indicators). Compare against
the user's native X instead; _transform() already validates it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Replace the file-local data dicts and _col/_cols/_shape/_drop_missing
helpers with the shared test structure: make_df and data_normal_dist /
data_na fixtures, isinstance(X, make_df) plus to_dict() checks, and
pytest.raises(match=...).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…nsorizer

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@solegalli
solegalli merged commit 16262df into narwhals-migration Sep 19, 2026
4 of 10 checks passed
@solegalli
solegalli deleted the narwhals-winsorizer branch September 19, 2026 07:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant