Skip to content

Migrate DropDuplicateFeatures to narwhals, add polars support - #1085

Open
solegalli wants to merge 3 commits into
narwhals-migrationfrom
narwhals-drop-duplicate-features
Open

solegalli wants to merge 3 commits into
narwhals-migrationfrom
narwhals-drop-duplicate-features

Conversation

@solegalli

Copy link
Copy Markdown
Collaborator

Stacked on #1070 (selection base). Its commit shows in the diff until #1070 merges; the change of this PR is the last commit.

Summary

Migrates DropDuplicateFeatures to narwhals: fit() accepts pandas and polars dataframes and transform() (inherited from BaseSelector) returns the same library it receives.

The old fit() transposed the selected columns and hashed each row of the transposed frame with pd.util.hash_pandas_object(X[variables].T). That needs pandas, and it is very slow: the transposed frame has one column per row of the data (object dtype as soon as the data mixes dtypes), so pandas hashes 10k-500k tiny columns one by one.

Now each variable gets a fingerprint of its values, and variables with the same fingerprint are duplicates. As before, the first variable of each group is kept:

  • numbers and booleans: compared as float64 values, so 1, 1.0 and True are equal, as they were before. -0.0 becomes 0.0 and all NaNs share one bit pattern. Integer variables with values from 2**53 on are hashed exactly, because float64 can't tell consecutive integers apart at that size.
  • datetimes: compared in nanoseconds (UTC), so datetime64[us] and datetime64[ns] variables with the same timestamps are duplicates. Time-zone-aware and naive datetimes are never duplicates of each other.
  • durations: compared in nanoseconds.
  • other values (strings, categoricals, objects): pandas uses hash_pandas_object, which hashes categoricals by value. Polars casts categoricals and enums to strings.
  • variables with only missing values: duplicates of each other, whatever their dtype (as before).

pandas: the fingerprint is hash(values.tobytes()) of the float64 / int64 values, or of the hash_pandas_object hashes, computed one column at a time.
polars: narwhals has no hash function, so a single polars select computes (pl.col(v).hash() * odd_random_weights).sum() and null_count() for all variables in parallel. The weights make the sum depend on row order.

missing_values is now validated with the shared _check_param_missing_values, so its error message ends with Got ... instead. and non-string values raise the same error.

Benchmarks

Data: 60% float (5% NaN), 15% int, 15% string (50 categories, 5% missing), 5% datetime, 5% bool, plus 10% duplicated columns (half of the int duplicates cast to float). Timings are medians of the repeats. Another job was running on the machine at the same time, so read the numbers as orders of magnitude.

Whole fit(), pandas, old (transpose + hash) vs new. Old: 3 repeats. New: 5 repeats. Each cell is the best of 2 runs.

rows x cols old new speed-up
10k x 11 1.91 s 3.7 ms ~500x
10k x 55 1.87 s 12.5 ms ~150x
10k x 220 6.56 s 42.5 ms ~150x
100k x 11 14.1 s 11.4 ms ~1200x
100k x 55 11.4 s 90.1 ms ~125x
100k x 220 not run 477 ms
500k x 11 41.3 s 238 ms ~170x
500k x 55 not run 548 ms
500k x 220 not run 2.22 s

Whole fit(), polars (new). Before this PR, polars input wasn't supported.

rows x cols new
500k x 11 19 ms
500k x 55 46 ms
500k x 220 133 ms
1M x 11 38 ms
1M x 55 101 ms
1M x 220 305 ms
2M x 55 185 ms

Candidate fingerprint implementations (the fingerprinting only; 5 repeats in alternating order, median). The chosen implementation is in bold.

pandas:

  • pd_hash: hash_pandas_object on every normalised column.
  • pd_np: numbers as canonical float64 bytes, datetimes/durations as int64 ns, hash_pandas_object only for other values.
  • pd_np2d: numbers as one 2D float64 block with a weighted uint64 sum.
  • nw: narwhals casts, then a pandas hash per column.
rows x cols pd_hash pd_np pd_np2d nw
10k x 55 16.6 ms 10.5 ms 12.0 ms 30.2 ms
10k x 220 76 ms 64 ms 54 ms 138 ms
100k x 55 130 ms 114 ms 115 ms 146 ms
100k x 220 545 ms 506 ms 453 ms 624 ms
500k x 55 645 ms 545 ms 539 ms 643 ms
500k x 220 2.96 s 2.87 s 4.39 s 4.63 s

pd_np and pd_np2d tie up to 100k rows. The 2D block needs an n x p float64 copy plus an equally large uint64 temporary, which makes it much slower at 500k x 220. pd_np is also the simpler of the two. String columns dominate the pandas time: factorising them inside hash_pandas_object accounts for about 40% of the total. A Python hash(tuple(values)) was no faster.

polars:

  • hash_sum: one select of (col.hash() * weights).sum() for all columns.
  • hash_bytes: polars hashes, then hash(bytes) in Python.
  • numpy: float64 bytes for numbers, polars hash for the rest.
rows x cols hash_sum hash_bytes numpy
500k x 11 21 ms 35 ms 34 ms
500k x 55 47 ms 135 ms 179 ms
500k x 220 111 ms 414 ms 790 ms
1M x 55 82 ms 186 ms 377 ms
2M x 55 172 ms 475 ms 778 ms

I also tried pl.col(v).implode().hash(), which needs no weights. It was 2.5x slower than hash_sum at 500k x 55 (64 ms vs 26 ms). A pure-narwhals implementation isn't possible because narwhals has no hash function.

Behaviour

Identical to the old pandas output (duplicated_feature_sets_, including group order, features_to_drop_, variables_ and the transformed dataframe). I compared the old code on main with the new code on 20 cases: random mixed-dtype frames (200-20k rows, 15-88 columns, with and without NaN, 4 seeds), Titanic with duplicated columns (with and without missing data), a frame of edge cases (int/float, bool/0-1, None/NaN strings, all-empty columns, NaN floats), variables, confirm_variables and missing_values="raise". Polars gives the same groups and columns as pandas on all 20 cases.

Other cases that match the old output: integer column names, nullable Int64 vs float with NaN, datetime64[ns] vs [us], NaT, tz-aware vs naive (not duplicates), two time zones with the same instants (duplicates), timedeltas, periods, category vs string, all-empty columns of different dtypes, uint64 vs int64, inf.

Differences (all in rare pandas edge cases where the old result came from pandas hashing the object-dtype transposed frame):

  1. Pre-existing bug, fixed: integers from 2**53 on were rounded to float when any float variable was selected, so [2**60, 1] and [2**60 + 1, 1] counted as duplicates. They are now compared exactly. test_large_integers_are_compared_exactly fails with the old code.
  2. 0.0 vs -0.0: the old code treated them as duplicates only when a non-numeric variable was also selected. They are now always duplicates. See "Needs decision".
  3. The cases under "Needs decision" below.

Tests

I rewrote tests/test_selection/test_drop_duplicate_features.py to the conventions:

  • init errors come first, then test_init_param_assignment;
  • every behaviour runs on both backends through make_df, with frame_to_dict and full-message match=;
  • a polars-only test checks NaN vs null, and a pandas-only test checks integer column names.

41 tests pass.

tests/test_selection, base branch (origin/narwhals-selection-base) vs this branch:

  • before: 137 failed, 408 passed;
  • after: 129 failed, 453 passed;
  • no new failures. 8 tests are fixed: the 4 old tests of this file and 4 test_check_estimator_selectors.py tests (test_check_estimator_from_sklearn, test_check_multivariate_estimator_from_feature_engine, test_confirm_variables, test_transformers_in_pipeline_with_set_output_pandas) that failed on this selector.

tests/parametrize_with_checks_selection_v16.py has 289 failures before and after, the same list. They are the numpy-input checks, and they are pre-existing.

flake8 feature_engine tests is clean. mypy feature_engine shows the same 2 pre-existing errors as the base (in datetime_subtraction.py and log.py).

Docs:

  • The docstring and user guide explain how values of different dtypes are compared.
  • New "With polars" examples.
  • The user guide's train_t.columns output now shows dtype='str', which is what pandas 3 returns.
  • I ran every example and pasted the real output.

Needs decision

  1. Numbers stored as text vs numeric variables. The old code treated [1, 2, 3] (int) and ["1", "2", "3"] (string) as duplicates. It did this by accident: pandas turns mixed object values into strings when it hashes them. The result was also inconsistent: 1.0 vs "1" were not duplicates, and in mixed groups the result depended on the order of the columns. They are no longer duplicates. Keeping the old result would need the transposed-frame hashing (100-1000x slower) or a string cast of every numeric column.
  2. pandas object columns holding numbers or datetimes, and categoricals with numeric categories. Examples are object [1, 2] vs int64 [1, 2], object Python datetimes vs datetime64, and pd.Categorical([1, 2]) vs int [1, 2]. These were duplicates before and are not now, because they are compared as "other" values. Supporting them would mean inspecting object and category contents in pandas only.
  3. -0.0 vs 0.0 are now always duplicates. Before, they were duplicates only when non-numeric variables were also selected.
  4. Backends other than pandas and polars. The non-pandas branch uses polars' hash() through nw.get_native_namespace, because narwhals has no hash function. PyArrow, modin and cuDF input would fail in fit(). If those backends are in scope, we'd need a generic (slower) path.

Pre-existing issues, not fixed

  • The shared _missing_values_docstring in feature_engine/_docstrings/selection/_docstring.py says missing values are raised or ignored "when determining correlation", which is wrong for this selector. The file is shared with other selectors, so I left it.

solegalli and others added 3 commits September 19, 2026 11:44
…s support

BaseSelector.transform() returns the retained features in the train set
order, in the same library as the input (pandas X[features], narwhals
select otherwise). BaseRecursiveSelector.fit() trains the estimators on
native frames and returns (nw_X, y). The helpers in
base_selection_functions no longer import pandas: correlations are
computed with numpy (np.corrcoef, or matrix products for pairwise
complete observations when there are missing values), and feature
importances are pandas Series for pandas input and dicts otherwise.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
fit() no longer transposes the dataframe. Each variable gets a fingerprint
of its values instead: numbers (booleans included) as float64, datetimes and
durations in nanoseconds, and other values through pandas'
hash_pandas_object, or, for polars, a weighted sum of the polars row hashes
computed in a single select. Integers from 2**53 on are hashed exactly.
Variables with the same fingerprint are duplicates and the first one is
kept, as before. Missing values are equal whatever the data type.

missing_values is validated with the shared check, so its error message now
ends with "Got ... instead.".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant