Skip to content

Migrate BaseTargetMeanEstimator to narwhals, add polars support - #1066

Open
solegalli wants to merge 1 commit into
narwhals-migrationfrom
narwhals-prediction-base
Open

solegalli wants to merge 1 commit into
narwhals-migrationfrom
narwhals-prediction-base

Conversation

@solegalli

Copy link
Copy Markdown
Collaborator

Summary

Migrates BaseTargetMeanEstimator (feature_engine/_prediction/base_predictor.py) to narwhals. fit and _predict now accept pandas and polars dataframes. TargetMeanClassifier and TargetMeanRegressor are not changed here; they get their own PRs stacked on this one.

Before this PR, the base built a Pipeline of a discretiser (return_boundaries=True) and one or two MeanEncoders. Every numerical value became an interval string, and the encoder grouped the target by those strings. On polars, fit ran but _predict failed at DataFrame.mean(axis=1). On pandas, integer column names failed in fit.

The new implementation computes the means directly:

  • Numerical variables: the migrated EqualWidthDiscretiser / EqualFrequencyDiscretiser is fitted to get binner_dict_ (reused as is). Its _digitize turns values into integer bin codes. The target mean per code is computed, and the keys of encoder_dict_ are the same interval labels the discretiser writes (_format_bin_labels). A per-bin array of means (_bin_means, NaN for bins that had no training rows) makes _predict a numpy lookup: means[codes].
  • Categorical variables: a group-by mean of the target per category.
  • _predict: sums the encoded columns in numpy and divides by the number of variables. It raises the same "NaN values were introduced" error for unseen categories and for values that fall in bins that were empty in the training set.
  • nw_X = check_X_y(X, y); the native X goes to the variable, NaN and inf helpers; the target is paired with add_target_to_X.
  • The private _pipeline, _make_*_pipeline and _transform are removed. _transform was only used by _predict.
  • Init: strategy checks the type before the membership test, and bins must be a positive integer (see "Pre-existing issues").

Benchmarks

Median of 7 runs, versions run in alternation, times in ms. Data: num standard-normal columns, cat string columns with 20 categories, continuous target, bins=5, equal_width. "old" is the pipeline on the base branch. Its polars _predict fails, so that column shows "-".

backend rows num cat fit old fit new predict old predict new
pandas 10k 2 2 16.6 5.8 7.9 1.9
pandas 10k 10 0 27.4 7.5 15.0 3.3
pandas 10k 0 10 30.6 13.0 9.6 3.9
pandas 10k 10 10 60.0 19.7 25.1 6.8
pandas 100k 2 2 72.5 27.5 32.2 9.3
pandas 100k 10 0 135.0 25.8 74.5 13.9
pandas 100k 0 10 139.5 97.0 57.5 29.5
pandas 100k 10 10 336.9 121.8 131.1 43.1
pandas 500k 2 2 312.1 114.3 132.4 40.1
pandas 500k 10 0 592.0 97.9 323.1 54.5
pandas 500k 0 10 631.7 480.6 278.3 143.0
pandas 500k 10 10 1630.7 602.3 614.1 218.4
polars 500k 2 2 143.4 75.6 - 22.9
polars 500k 10 0 352.6 95.7 - 58.8
polars 500k 0 10 354.9 255.1 - 47.8
polars 500k 10 10 740.3 317.9 - 90.1
polars 2M 2 2 470.2 267.3 - 74.9
polars 2M 10 0 1435.5 370.3 - 237.2
polars 2M 0 10 1277.8 1056.2 - 131.8
polars 2M 10 10 3539.8 1585.1 - 439.5

On categorical variables, most of the remaining fit time is spent in find_categorical_and_numerical_variables, which is shared and not changed here (see "Pre-existing issues").

How each step was chosen, timed per variable, isolated, median of repeats:

step options (500k rows unless noted) chosen
numerical: target mean per bin, pandas pandas y.groupby(codes).mean() 2.53 ms; narwhals group_by on pandas 3.05; np.bincount 2.31 (10k: 0.10 / 0.75 / 0.05) pandas groupby. bincount is about 8% faster but differs from the old means in the last bits (see "Needs decision")
numerical: target mean per bin, polars whole fit, 10 variables: narwhals group_by 121.5 ms vs bincount 124.7 (2M: 320 vs 336; 5M: 810 vs 904) narwhals group_by
categorical: target mean, pandas pandas groupby(observed=True, dropna=False) vs narwhals group_by: 10k 0.22 vs 0.48; 100k 1.9 vs 2.0; 500k 9.5 vs 9.2 (20 categories) pandas. It is 2x faster at 10k and about 5% slower at 500k. dropna=False is safe after the NaN check and is about 40% faster than the default
categorical: target mean, polars narwhals group_by 1.2 ms vs polars-native 1.1 (2M: 3.0 vs 3.2) narwhals
numerical: bin codes _digitize 5.8 ms; np.searchsorted 5.5; polars search_sorted 5.1 (2M: 23.5 / 22.1 / 20.6) _digitize on both backends, to reuse the discretiser. A polars-only branch would save about 10% of this step
categorical: encode at predict, pandas factorize(use_na_sentinel=False) + lookup 8.7 ms; .map(dict) 12.5; narwhals replace_strict 14.1 factorize
categorical: encode at predict, polars narwhals replace_strict 2.6 ms vs polars-native 2.3 narwhals
row mean numpy running sum / k (same bits as the old DataFrame.mean(axis=1)) numpy

Behaviour

pandas: I compared outputs of BaseTargetMeanEstimator, TargetMeanClassifier and TargetMeanRegressor (fit attributes, _predict, predict, predict_proba and error messages) with the base branch on 44 cases. The cases covered both strategies, 3/5/10 bins, binary and continuous targets, list/array/bool/int targets, reordered columns, constant columns, bins left empty in training, unseen categories, out-of-range values, NaN and inf at fit and at predict, a wrong number of columns, datetime columns, category dtype with unused categories, object columns holding integers, more bins than rows, a custom index and nullable pandas dtypes. binner_dict_, encoder_dict_ values, predictions and error messages are bit-identical. The differences:

  • pandas integer column names now work; fit raised InvalidIntoExprError before.
  • Key order inside encoder_dict_ changed (dict equality is unchanged). Numerical variables now list their intervals in bin order. Before, the order was alphabetical by label string, for example '(-inf, 9.8]', '(19.6, 29.4]', ..., '(9.8, 19.6]'. Categorical variables are sorted on pandas; before, they were sometimes sorted by count, depending on which pipeline branch ran. On polars the order follows group_by.

polars: gives the same values as pandas on all cases (rtol 1e-12). Group-by means on polars can differ from pandas in the last bit, so the tests use pytest.approx.

Tests

  • New tests/test_prediction/test_base_predictor.py, following the conventions. It starts with init errors (bins, strategy) and test_init_param_assignment. The fit and predict tests run on both backends through make_df: attributes per strategy, predictions, list and array targets, numerical-only and categorical-only data, variable selection with a datetime column, reordered columns, out-of-range values, a constant variable, unseen categories, bins that were empty in training, NaN/None and inf at fit and predict, the column count check and NotFittedError. Three tests are pandas-only: integer column names, a custom index and category dtype.
  • test_check_estimator_prediction.py: test_attributes_upon_fitting checked _pipeline.named_steps. It now checks _discretiser.
  • tests/test_prediction: 58 passed / 0 failed before; 122 passed / 0 failed after.
  • tests/test_selection (imports the estimators): 177 failed / 251 passed before and after, with the same list of failures. test_target_mean_selection.py already fails on the base branch, in SelectByTargetEncoding itself ('list' object has no attribute 'to_list').
  • flake8 feature_engine tests is clean. mypy feature_engine reports the same 2 errors as the base branch.

Docs: the prediction module has no user guide page, and the base class is private, so only the docstring changed. The user-facing examples belong in the classifier and regressor PRs.

Needs decision

  1. Numerical means on pandas with np.bincount: it is about 8% faster on numerical-only fits. It changes the means in the last bits (relative difference up to about 1e-13) because pandas uses a compensated sum. I kept pandas groupby so the output stays identical.
  2. Unseen values in both numerical and categorical variables: the error still names only the numerical variables, as before. Should it name all of them in a single error?

Pre-existing issues

  • Fixed: bins=0 gave a misleading NaN error at fit, and negative bins raised IndexError. They now raise at init with bins must be a positive integer. Got {bins} instead. Before, the message said "bins must be an integer".
  • Not fixed: find_categorical_and_numerical_variables scans the full values of every string column to check whether they parse as datetimes or numbers. It takes about 60% of fit on categorical data (for example, 315 ms of 583 ms at 500k rows with 10 categorical variables). This is shared code outside this PR.
  • Not fixed: polars Object columns, such as the output of a discretiser with return_object=True, can be fitted but _predict fails in replace_strict. MeanEncoder.transform fails the same way. On the base branch, fit already rejected them.

Compute the target mean per bin and per category directly instead of
through a Pipeline of discretiser and MeanEncoders, reusing the fitted
discretiser's bin edges and labels. Fit and _predict accept pandas and
polars dataframes, and pandas integer column names now work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant