Skip to content

Migrate TargetMeanClassifier to narwhals, add polars support - #1076

Open
solegalli wants to merge 2 commits into
narwhals-migrationfrom
narwhals-target-mean-classifier
Open

solegalli wants to merge 2 commits into
narwhals-migrationfrom
narwhals-target-mean-classifier

Conversation

@solegalli

Copy link
Copy Markdown
Collaborator

Stacked on #1066 (BaseTargetMeanEstimator). Please review and merge #1066 first. Until then this PR's diff also shows #1066's commit. The only commit that belongs to this PR is the last one.

Summary

TargetMeanClassifier now works with pandas and polars dataframes. The target can be a pandas or polars Series, a numpy array or a list.

The base from #1066 does the dataframe work: fit, binning, encoding, _predict and the input checks. This PR changes only the class-handling code, and each part was checked with pandas, polars, numpy and list targets:

  • Binary-target check and classes_: sklearn's check_classification_targets and unique_labels already accept polars Series, so they are kept. unique_labels now runs once instead of twice (the second call was only used to remap the labels).
  • Remapping labels to 0/1: this now runs on np.asarray(y). That fixes a bug: a list of string labels raised TypeError, because list == "label" is a single False and not an elementwise comparison.
  • predict_proba, predict_log_proba, predict: unchanged. They are numpy operations on the 1-D array returned by _predict, so they don't depend on the backend.
  • The type hints and docstrings no longer mention pandas only. The docstring gains a pandas and a polars example (I ran both and pasted the real output). There is no user guide page for this class, since it lives in the private _prediction module.

Benchmarks

I timed the whole fit(), median of 9 runs (5 at 2M rows), with the three versions alternating. Columns are half numerical and half categorical, and bins=5. predict/predict_proba were not changed and were not benchmarked against alternatives, because the thresholding and stacking are trivial numpy operations on the base's output.

  • base ref: the current code, which calls check_classification_targets once and unique_labels twice when the labels are not 0/1.
  • this PR: unique_labels called once.
  • attach_unique option: not implemented, see "Needs decision".

With 0/1 labels, this PR and the base ref run the same code, so the gap between them in those rows is noise. The noise is about ±10–30%, because the machine was shared with other jobs.

The target handling only matters for string labels. For those, sklearn sorts an object array to find the unique labels, and this dominates fit. Removing one of the three sorts saves roughly 35%. For example, pandas with 500k rows and 1 column goes from 877 ms to 531 ms.

Full table
backend rows columns target labels base ref (ms) this PR (ms) attach_unique option (ms)
pandas 10,000 1 0/1 3.0 2.8 2.5
pandas 10,000 1 1/2 2.9 2.7 2.4
pandas 10,000 1 strings 15.1 9.8 4.9
pandas 10,000 5 0/1 6.2 6.3 5.9
pandas 10,000 5 1/2 6.6 6.3 5.9
pandas 10,000 5 strings 19.1 14.0 8.9
pandas 10,000 20 0/1 31.0 32.4 29.9
pandas 10,000 20 1/2 27.0 24.9 22.1
pandas 10,000 20 strings 30.7 26.0 21.0
pandas 100,000 1 0/1 5.2 5.2 4.2
pandas 100,000 1 1/2 6.2 5.3 4.2
pandas 100,000 1 strings 153.1 93.3 34.4
pandas 100,000 5 0/1 27.8 27.8 26.9
pandas 100,000 5 1/2 28.5 27.6 26.6
pandas 100,000 5 strings 172.4 113.5 56.7
pandas 100,000 20 0/1 113.3 113.0 112.1
pandas 100,000 20 1/2 115.2 113.9 112.8
pandas 100,000 20 strings 261.7 202.1 142.2
pandas 500,000 1 0/1 17.8 17.9 13.7
pandas 500,000 1 1/2 22.0 17.9 14.1
pandas 500,000 1 strings 876.5 530.8 187.5
pandas 500,000 5 0/1 121.8 122.4 117.3
pandas 500,000 5 1/2 126.0 122.5 118.3
pandas 500,000 5 strings 1569.4 1072.2 500.8
pandas 500,000 20 0/1 907.7 891.3 891.2
pandas 500,000 20 1/2 754.2 875.6 717.9
pandas 500,000 20 strings 1411.4 1057.3 718.4
polars 500,000 1 0/1 25.0 21.0 14.7
polars 500,000 1 1/2 23.1 18.7 13.7
polars 500,000 1 strings 349.4 326.5 59.3
polars 500,000 5 0/1 101.4 133.2 108.6
polars 500,000 5 1/2 75.3 65.9 62.4
polars 500,000 5 strings 484.3 536.9 141.0
polars 500,000 20 0/1 302.8 340.3 367.3
polars 500,000 20 1/2 383.7 482.4 368.0
polars 500,000 20 strings 890.1 783.8 377.3
polars 2,000,000 1 0/1 71.2 62.6 52.9
polars 2,000,000 1 1/2 84.9 61.1 55.4
polars 2,000,000 1 strings 1710.3 1224.0 246.7
polars 2,000,000 5 0/1 195.7 192.9 173.7
polars 2,000,000 5 1/2 225.0 196.0 186.5
polars 2,000,000 5 strings 1602.9 1229.6 373.8
polars 2,000,000 20 0/1 685.5 669.8 702.6
polars 2,000,000 20 1/2 639.1 633.9 613.7
polars 2,000,000 20 strings 2085.7 1847.9 842.0

Behaviour

Before the change, I recorded the outputs of the base ref on 296 cases. These cover fit, classes_ and its dtype, encoder_dict_, binner_dict_, predict, predict_proba and predict_log_proba (values, dtype, memory layout), score and every error message. The inputs were:

  • label sets 0/1, 1/2, -1/1, strings, booleans, floats, a single class, 3 classes, a continuous target, and ""/"z"
  • each label set as a pandas Series, list or numpy array (pandas X), and as a polars Series or numpy array (polars X)
  • 4 parameter combinations
  • special targets: None, a scalar, a string, category dtype, Int64 with and without NA, string with NA, object with None, mixed types, NaN, a dataframe, a 2-D list, a mismatched index, polars categorical, polars with nulls, polars boolean
  • integer column names, and NaN, unseen, wrong-column, not-a-dataframe and not-fitted errors at predict time

Results:

  • pandas: identical in every case, except a list of string labels, which now fits (bug fix above).
  • polars: every result is the same as pandas: every label set, target type and error.

Tests

tests/test_prediction/test_target_mean_classifier.py is rewritten to the conventions:

  • The make_df fixture runs every test on both backends. The data is a plain dict in the file, and the expected values are explicit.
  • Targets are built with make_series, plus one test with list and numpy targets (with integer and string labels).
  • Every error uses match=re.escape(full message), including NotFittedError for predict, predict_proba and predict_log_proba.
  • df_classification is removed from tests/test_prediction/conftest.py. Only this file used it.

The tests cover classes_ and the predictions for 6 label types, sorted classes (the probability is that of the larger label), a probability of 0.5 returning the first class, predict_log_proba, score, the not-binary and continuous-target errors, and pandas-only tests for integer column names and a non-default index. The init-parameter errors are tested in test_base_predictor.py, which is part of #1066.

Before/after (base ref = #1066 head, 6d284b4):

folder base ref this PR
tests/test_prediction 122 passed 172 passed
tests/test_selection (uses TargetMeanClassifier) 177 failed, 251 passed 177 failed, 251 passed (same failing tests)
tests/parametrize_with_checks_prediction_v16.py (not collected by default) 75 failed, 50 passed 75 failed, 50 passed

On the base ref, the new test file fails only on the list-of-string-labels cases.

flake8 feature_engine tests is clean. mypy feature_engine reports the same 2 errors as on the base ref, both in other modules.

Needs decision

Find the unique labels once with sklearn.utils._unique.attach_unique. This is the "attach_unique option" column in the benchmark. The code would be:

y_arr = attach_unique(np.asarray(y))
check_classification_targets(y_arr)
self.classes_ = unique_labels(y_arr)

sklearn's checks reuse the unique values attached to the array, so the object array is sorted once instead of twice. With string labels, fit becomes 2–5x faster than this PR:

  • pandas, 500k rows, 1 column: 531 → 188 ms
  • pandas, 500k rows, 5 columns: 1072 → 501 ms
  • polars, 2M rows, 1 column: 1224 → 247 ms

I did not implement it for two reasons:

  1. It imports from a private sklearn module. attach_unique is available in every sklearn version we support (1.7 or later), but it is private.
  2. It changes some error messages for invalid targets:
    • mixed-type labels whose first value is not a string (for example pd.Series([1, "a", ...])) raise TypeError: '<' not supported between instances of 'str' and 'int' instead of ValueError: Unknown label type: unknown...
    • a 2-D list target ([[0], [1], ...]) fits instead of raising TypeError: cannot use 'list' as a set element
    • None or a scalar target raises "Unknown label type: unknown" instead of "Expected array-like...", unless we add an if np.asarray(y).ndim == 0 guard.

Do you want it?

Merge conflict with the TargetMeanRegressor PR. That PR is being migrated in parallel and probably edits tests/test_prediction/conftest.py too, since df_regression lives there. The conflict is trivial.

Pre-existing issues, not fixed

  • Single-class target: fit succeeds, but predict and score raise IndexError: index 1 is out of bounds for axis 0 with size 1, because predict reads classes_[1]. predict_proba returns probabilities for 2 classes. Should fit raise a clear error instead? That would be a behaviour change.
  • pandas target with an index that differs from X's: it raises "The indexes of X and y do not match." only when the labels are 0/1. With other labels, the target is remapped to a numpy array before the base's check, so the index is ignored and rows pair by position.
  • User guide link: the docstring's :ref:`User Guide <targetmeanclassifier>` points to a label that doesn't exist. No user guide page exists for this class.
  • Truthiness check: the remap condition any(x for x in self.classes_ if x not in [0, 1]) tests whether each label is truthy rather than just membership. It gives a different result only for a single-class target with label "", which is already covered by the single-class issue above.

solegalli and others added 2 commits September 19, 2026 11:33
Compute the target mean per bin and per category directly instead of
through a Pipeline of discretiser and MeanEncoders, reusing the fitted
discretiser's bin edges and labels. Fit and _predict accept pandas and
polars dataframes, and pandas integer column names now work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- fit, predict, predict_proba and predict_log_proba accept pandas and polars
  dataframes, and the target as a Series of any backend, numpy array or list.
- Find the target labels once instead of twice when remapping them to 0 and 1.
- Fix fit with a list of string labels, which raised a TypeError.
- Rewrite the tests to the conventions, with both backends.
- Add pandas and polars examples to the docstring.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant