Skip to content

Migrate SelectByInformationValue to narwhals, add polars support - #1088

Open
solegalli wants to merge 2 commits into
narwhals-migrationfrom
narwhals-select-by-information-value
Open

solegalli wants to merge 2 commits into
narwhals-migrationfrom
narwhals-select-by-information-value

Conversation

@solegalli

Copy link
Copy Markdown
Collaborator

Stacked on #1070 (selection base classes). Its commit shows in this diff until #1070 is merged.

Summary

SelectByInformationValue now works with pandas and polars, and returns the same library it receives (transform is inherited from BaseSelector).

  • fit: the information values are computed with numpy on pandas (Series.factorize() + np.bincount), and with one lazy narwhals query over all the variables on other backends (polars runs the per-variable group-bys in parallel).
  • Numerical variables are discretised into integer codes (return_boundaries=False) instead of interval labels: the IV only needs to know which interval each value falls in, and grouping integers is faster than building and grouping strings.
  • The numerical variables are found with find_categorical_and_numerical_variables(..., exclude_datetime=False): variables_ has already been filtered for datetime by _select_all_variables, and the datetime check (which parses the values) doesn't change the numerical list. This check was the largest single cost of fit.
  • Zero counts are replaced by 0.5, the same rule WoEEncoder now uses on narwhals-migration. The selector no longer calls WoE._calculate_woe, so the rule is written in the selector too (see "Needs decision").
  • confirm_variables is now validated at init via super().__init__(confirm_variables), like the other selectors.
  • strategy gets the isinstance(..., str) check before the membership test.
  • No pandas import left in the module.
  • Docstring and user guide: added a "With polars" example and a section on categories with no positive or no negative cases. Fixed user guide outputs that didn't match the code (see below).

Benchmarks

Machine under load from other jobs during some runs, so I repeated the runs when load was low (load average about 3–5). Numbers below are from those runs. Median of 3 subprocess runs, each the median of 3 fits, alternating old and new. Data: 12 categories per categorical variable, normal numerical variables, binary target, bins=5.

Whole fit(): base (#1070) vs this PR

backend rows cat+num vars base this PR speedup
pandas 10k 5+5 0.047s 0.013s 3.7x
pandas 10k 10+10 0.094s 0.024s 3.9x
pandas 10k 25+25 0.224s 0.051s 4.4x
pandas 100k 5+5 0.157s 0.071s 2.2x
pandas 100k 10+10 0.302s 0.134s 2.2x
pandas 100k 25+25 0.748s 0.325s 2.3x
pandas 500k 5+5 0.615s 0.315s 2.0x
pandas 500k 10+10 1.243s 0.628s 2.0x
pandas 500k 25+25 3.108s 1.553s 2.0x
pandas, equal_frequency 500k 25+25 3.234s 1.737s 1.9x
polars 500k 5+5 0.303s 0.138s 2.2x
polars 500k 10+10 0.606s 0.267s 2.3x
polars 500k 25+25 1.523s 0.669s 2.3x
polars 2M 5+5 1.169s 0.533s 2.2x
polars 2M 10+10 2.349s 1.063s 2.2x
polars 2M 25+25 5.940s 2.601s 2.3x
polars, equal_frequency 2M 25+25 6.588s 3.377s 2.0x

Most of the remaining time in fit is spent in shared helpers (_select_all_variables' datetime check, the NaN/inf checks, the discretiser fit).

IV aggregation only (numerical variables already discretised), median of 7, rotating order

  • current: WoE._calculate_woe per variable, then numpy sum
  • narwhals lazy: one lazy query, a group-by per variable, concatenated (chosen for polars)
  • pandas groupby: y.groupby(X[var], observed=True, sort=False).agg(["sum", "size"])
  • numpy: X[var].factorize() + np.bincount (chosen for pandas)
  • polars native: pl.collect_all of one lazy query per variable
backend rows vars current narwhals eager loop narwhals lazy pandas groupby numpy polars native
pandas 10k 5+5 0.0305 – 0.0215 0.0038 0.0022 –
pandas 10k 25+25 0.1495 – 0.1055 0.0188 0.0106 –
pandas 100k 5+5 0.0631 – 0.0324 0.0199 0.0184 –
pandas 100k 25+25 0.3090 – 0.1593 0.1001 0.0913 –
pandas 500k 5+5 0.2054 – 0.0800 0.0935 0.0901 –
pandas 500k 25+25 1.0073 – 0.4002 0.4614 0.4463 –
polars 500k 5+5 0.0255 0.0184 0.0107 – 0.1032 0.0077
polars 500k 25+25 0.1298 0.0882 0.0487 – 0.4849 0.0327
polars 2M 5+5 0.0868 0.0514 0.0359 – 0.4340 0.0321
polars 2M 25+25 0.4859 0.2687 0.1517 – 2.2560 0.1302
  • pandas: numpy is 1.7x–10x faster than narwhals from 10k to 100k rows. At 500k, narwhals is about 10% faster: factorizing the string columns takes most of the time on both paths. I chose numpy because it wins across most of the pandas range.
  • polars: the lazy narwhals query is within about 15% of native polars collect_all and 1.7x–2.7x faster than a per-variable eager loop, so polars uses the narwhals path.

Discretising into integer codes instead of interval labels: pandas 500k 25 numerical variables 0.377s → 0.210s, polars 2M 25 variables 2.216s → 0.840s.

Behaviour

I compared outputs on 37 cases (credit-approval data from the user guide, synthetic data with every parameter branch, NaN, inf, datetime, bool, constant, integer column names, pandas category dtype, non-default index, list/array/str/bool/{1,2} targets, zero-count categories, empty intervals, reordered columns). Old versions compared: v1.9.4, base (#1070), and this PR on pandas and polars.

  • This PR vs base (pandas): same variables, features to drop, support and transform output in every case. The IVs are the same, or differ by at most 4e-16 relative, because the per-category terms are summed in a different order (first appearance or interval order instead of sorted category labels). For example, in the user guide, A6 goes from 0.6006252129425703 to 0.6006252129425705.
  • pandas vs polars (this PR): same results, with IVs within 1e-15 relative. The exception is when every feature is dropped (see "Pre-existing issues").
  • information_values_ values are now Python float on both backends. Before, pandas gave np.float64, which is a float subclass.

Differences vs v1.9.4 (all already present on base, through the WoE changes on narwhals-migration)

  1. Zero counts: v1.9.4 raised ValueError: The proportion of one of the classes for a category in variable X is zero... when a category or interval had no positive or no negative cases. Now the zero count is replaced by 0.5 and the IV is computed. This affects many realistic inputs: the credit-approval data with default parameters, bins=10, numerical variables whose outer equal-width intervals hold only a few observations, and any rare category. In v1.9.4 the selector never accepted a fill_value-style option: it always passed fill_value=None to _calculate_woe, so it always raised. There is nothing to deprecate on the selector's API.
  2. pandas category dtype with unused categories: v1.9.4 grouped with observed=False, so an unused category had 0/0 counts and raised the error above. Now unused categories are ignored, and the IV equals the IV of the same column without them.
  3. Bug fix (inherited from WoE._check_fit_input): in v1.9.4, a target that wasn't 0/1 (for example 1/2 or strings), used with a dataframe with a non-default index, lost its index when remapped. All IVs came out 0.0, and every feature was dropped. Now the IVs are correct. This case is covered by test_target_not_0_1_with_pandas_index.

User guide fixes

  • variables_ was shown as 7 variables including A7. It returns the 6 variables passed.
  • The last digit of 3 IVs changed (summation order, see above).
  • The note that said the transformer raises on zero counts is replaced by a section with a worked example of the 0.5 replacement.

Tests

tests/test_selection/test_information_value.py is rewritten to the conventions: init tests first (one per error message, including the new confirm_variables validation), then make_df tests on both backends with explicit expected values (including hand-written IV formulas for the regular and zero-count cases and for empty intervals), plus pandas-only tests for integer column names, category dtype with unused categories, and a non-default index. I removed the test of the private _calculate_iv method along with the method; the IV formula is now tested through fit.

When I run the new test file against the base code, only the 5 confirm_variables validation tests fail.

pytest tests/test_selection:

flake8 feature_engine tests: clean. mypy feature_engine: 2 errors, the same as base.

Needs decision

  1. Zero counts (behaviour change vs 1.9.4, inherited from the WoE change). I kept what base does: replace by 0.5, don't raise. The docstring and user guide describe this. Options:

    • (a) Keep it as it is.
    • (b) Restore the error for the selector only.
    • (c) Keep the 0.5 replacement and add a variables_with_zero_counts_ attribute, as WoEEncoder has, so users can see which variables were affected.

    Also: should the user guide get a "New in version 2.0" note about this change, as the WoEEncoder guide has?

  2. Zero-count rule written twice. The 0.5 rule is now in WoE._calculate_woe and in the selector's IV computation, because reusing _calculate_woe per variable was 2x–14x slower on pandas and about 3x slower on polars than computing the IV directly ("current" column above). If you prefer one place, the IV computation could move to the WoE mixin. That would touch woe.py, which is outside this PR.

  3. _check_variable_number(): the addendum says to call it after choosing the variables, but this selector never did, and v1.9.4 accepts a single variable. I didn't add it, because doing so would be a behaviour change.

Pre-existing issues, not fixed

  • polars, all features dropped: BaseSelector.transform calls nw_X.select(nw.col(*features)) with an empty list when every feature is dropped. polars then raises TypeError: Col.__call__() missing 1 required positional argument: 'name', while pandas returns a dataframe with no columns. This happens easily with this selector (for example threshold=1). The fix belongs in base_selector.py (Migrate the selection base classes and helpers to narwhals, add polars support #1070): guard the empty list, or raise a clear error as DropConstantFeatures does.
  • bins=True and threshold=True pass the init checks, because bool is a subclass of int.
  • The _more_tags comment explaining _skip_test says the transformer raises on zero counts, which is no longer true. I left the tag alone. Whether the sklearn checks now pass without it is worth a separate look.

solegalli and others added 2 commits September 19, 2026 11:44
…s support

BaseSelector.transform() returns the retained features in the train set
order, in the same library as the input (pandas X[features], narwhals
select otherwise). BaseRecursiveSelector.fit() trains the estimators on
native frames and returns (nw_X, y). The helpers in
base_selection_functions no longer import pandas: correlations are
computed with numpy (np.corrcoef, or matrix products for pairwise
complete observations when there are missing values), and feature
importances are pandas Series for pandas input and dicts otherwise.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- Compute the information values with numpy on pandas and with one lazy
  narwhals query on other backends.
- Discretise numerical variables into integer codes instead of interval
  labels, and skip the datetime parsing when finding the numerical variables.
- Validate confirm_variables at init, like the other selectors.
- Rewrite the tests to run on pandas and polars, and update the docstring and
  user guide.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant