Conversation
…s support BaseSelector.transform() returns the retained features in the train set order, in the same library as the input (pandas X[features], narwhals select otherwise). BaseRecursiveSelector.fit() trains the estimators on native frames and returns (nw_X, y). The helpers in base_selection_functions no longer import pandas: correlations are computed with numpy (np.corrcoef, or matrix products for pairwise complete observations when there are missing values), and feature importances are pandas Series for pandas input and dicts otherwise. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
fit() no longer transposes the dataframe. Each variable gets a fingerprint of its values instead: numbers (booleans included) as float64, datetimes and durations in nanoseconds, and other values through pandas' hash_pandas_object, or, for polars, a weighted sum of the polars row hashes computed in a single select. Integers from 2**53 on are hashed exactly. Variables with the same fingerprint are duplicates and the first one is kept, as before. Missing values are equal whatever the data type. missing_values is validated with the shared check, so its error message now ends with "Got ... instead.". Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #1070 (selection base). Its commit shows in the diff until #1070 merges; the change of this PR is the last commit.
Summary
Migrates
DropDuplicateFeaturesto narwhals:fit()accepts pandas and polars dataframes andtransform()(inherited fromBaseSelector) returns the same library it receives.The old
fit()transposed the selected columns and hashed each row of the transposed frame withpd.util.hash_pandas_object(X[variables].T). That needs pandas, and it is very slow: the transposed frame has one column per row of the data (object dtype as soon as the data mixes dtypes), so pandas hashes 10k-500k tiny columns one by one.Now each variable gets a fingerprint of its values, and variables with the same fingerprint are duplicates. As before, the first variable of each group is kept:
1,1.0andTrueare equal, as they were before. -0.0 becomes 0.0 and all NaNs share one bit pattern. Integer variables with values from 2**53 on are hashed exactly, because float64 can't tell consecutive integers apart at that size.datetime64[us]anddatetime64[ns]variables with the same timestamps are duplicates. Time-zone-aware and naive datetimes are never duplicates of each other.hash_pandas_object, which hashes categoricals by value. Polars casts categoricals and enums to strings.pandas: the fingerprint is
hash(values.tobytes())of the float64 / int64 values, or of thehash_pandas_objecthashes, computed one column at a time.polars: narwhals has no hash function, so a single polars
selectcomputes(pl.col(v).hash() * odd_random_weights).sum()andnull_count()for all variables in parallel. The weights make the sum depend on row order.missing_valuesis now validated with the shared_check_param_missing_values, so its error message ends withGot ... instead.and non-string values raise the same error.Benchmarks
Data: 60% float (5% NaN), 15% int, 15% string (50 categories, 5% missing), 5% datetime, 5% bool, plus 10% duplicated columns (half of the int duplicates cast to float). Timings are medians of the repeats. Another job was running on the machine at the same time, so read the numbers as orders of magnitude.
Whole
fit(), pandas, old (transpose + hash) vs new. Old: 3 repeats. New: 5 repeats. Each cell is the best of 2 runs.Whole
fit(), polars (new). Before this PR, polars input wasn't supported.Candidate fingerprint implementations (the fingerprinting only; 5 repeats in alternating order, median). The chosen implementation is in bold.
pandas:
pd_hash:hash_pandas_objecton every normalised column.pd_np: numbers as canonical float64 bytes, datetimes/durations as int64 ns,hash_pandas_objectonly for other values.pd_np2d: numbers as one 2D float64 block with a weighted uint64 sum.nw: narwhals casts, then a pandas hash per column.pd_npandpd_np2dtie up to 100k rows. The 2D block needs an n x p float64 copy plus an equally large uint64 temporary, which makes it much slower at 500k x 220.pd_npis also the simpler of the two. String columns dominate the pandas time: factorising them insidehash_pandas_objectaccounts for about 40% of the total. A Pythonhash(tuple(values))was no faster.polars:
hash_sum: oneselectof(col.hash() * weights).sum()for all columns.hash_bytes: polars hashes, thenhash(bytes)in Python.numpy: float64 bytes for numbers, polars hash for the rest.I also tried
pl.col(v).implode().hash(), which needs no weights. It was 2.5x slower thanhash_sumat 500k x 55 (64 ms vs 26 ms). A pure-narwhals implementation isn't possible because narwhals has no hash function.Behaviour
Identical to the old pandas output (
duplicated_feature_sets_, including group order,features_to_drop_,variables_and the transformed dataframe). I compared the old code onmainwith the new code on 20 cases: random mixed-dtype frames (200-20k rows, 15-88 columns, with and without NaN, 4 seeds), Titanic with duplicated columns (with and without missing data), a frame of edge cases (int/float, bool/0-1, None/NaN strings, all-empty columns, NaN floats),variables,confirm_variablesandmissing_values="raise". Polars gives the same groups and columns as pandas on all 20 cases.Other cases that match the old output: integer column names, nullable
Int64vs float with NaN,datetime64[ns]vs[us], NaT, tz-aware vs naive (not duplicates), two time zones with the same instants (duplicates), timedeltas, periods, category vs string, all-empty columns of different dtypes,uint64vsint64, inf.Differences (all in rare pandas edge cases where the old result came from pandas hashing the object-dtype transposed frame):
[2**60, 1]and[2**60 + 1, 1]counted as duplicates. They are now compared exactly.test_large_integers_are_compared_exactlyfails with the old code.0.0vs-0.0: the old code treated them as duplicates only when a non-numeric variable was also selected. They are now always duplicates. See "Needs decision".Tests
I rewrote
tests/test_selection/test_drop_duplicate_features.pyto the conventions:test_init_param_assignment;make_df, withframe_to_dictand full-messagematch=;41 tests pass.
tests/test_selection, base branch (origin/narwhals-selection-base) vs this branch:test_check_estimator_selectors.pytests (test_check_estimator_from_sklearn,test_check_multivariate_estimator_from_feature_engine,test_confirm_variables,test_transformers_in_pipeline_with_set_output_pandas) that failed on this selector.tests/parametrize_with_checks_selection_v16.pyhas 289 failures before and after, the same list. They are the numpy-input checks, and they are pre-existing.flake8 feature_engine testsis clean.mypy feature_engineshows the same 2 pre-existing errors as the base (indatetime_subtraction.pyandlog.py).Docs:
train_t.columnsoutput now showsdtype='str', which is what pandas 3 returns.Needs decision
[1, 2, 3](int) and["1", "2", "3"](string) as duplicates. It did this by accident: pandas turns mixed object values into strings when it hashes them. The result was also inconsistent:1.0vs"1"were not duplicates, and in mixed groups the result depended on the order of the columns. They are no longer duplicates. Keeping the old result would need the transposed-frame hashing (100-1000x slower) or a string cast of every numeric column.[1, 2]vs int64[1, 2], object Pythondatetimes vsdatetime64, andpd.Categorical([1, 2])vs int[1, 2]. These were duplicates before and are not now, because they are compared as "other" values. Supporting them would mean inspecting object and category contents in pandas only.hash()throughnw.get_native_namespace, because narwhals has no hash function. PyArrow, modin and cuDF input would fail infit(). If those backends are in scope, we'd need a generic (slower) path.Pre-existing issues, not fixed
_missing_values_docstringinfeature_engine/_docstrings/selection/_docstring.pysays missing values are raised or ignored "when determining correlation", which is wrong for this selector. The file is shared with other selectors, so I left it.