Conversation
Compute the target mean per bin and per category directly instead of through a Pipeline of discretiser and MeanEncoders, reusing the fitted discretiser's bin edges and labels. Fit and _predict accept pandas and polars dataframes, and pandas integer column names now work. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- fit, predict, predict_proba and predict_log_proba accept pandas and polars dataframes, and the target as a Series of any backend, numpy array or list. - Find the target labels once instead of twice when remapping them to 0 and 1. - Fix fit with a list of string labels, which raised a TypeError. - Rewrite the tests to the conventions, with both backends. - Add pandas and polars examples to the docstring. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #1066 (
BaseTargetMeanEstimator). Please review and merge #1066 first. Until then this PR's diff also shows #1066's commit. The only commit that belongs to this PR is the last one.Summary
TargetMeanClassifiernow works with pandas and polars dataframes. The target can be a pandas or polars Series, a numpy array or a list.The base from #1066 does the dataframe work: fit, binning, encoding,
_predictand the input checks. This PR changes only the class-handling code, and each part was checked with pandas, polars, numpy and list targets:classes_: sklearn'scheck_classification_targetsandunique_labelsalready accept polars Series, so they are kept.unique_labelsnow runs once instead of twice (the second call was only used to remap the labels).np.asarray(y). That fixes a bug: a list of string labels raisedTypeError, becauselist == "label"is a singleFalseand not an elementwise comparison.predict_proba,predict_log_proba,predict: unchanged. They are numpy operations on the 1-D array returned by_predict, so they don't depend on the backend._predictionmodule.Benchmarks
I timed the whole
fit(), median of 9 runs (5 at 2M rows), with the three versions alternating. Columns are half numerical and half categorical, andbins=5.predict/predict_probawere not changed and were not benchmarked against alternatives, because the thresholding and stacking are trivial numpy operations on the base's output.check_classification_targetsonce andunique_labelstwice when the labels are not 0/1.unique_labelscalled once.With 0/1 labels, this PR and the base ref run the same code, so the gap between them in those rows is noise. The noise is about ±10–30%, because the machine was shared with other jobs.
The target handling only matters for string labels. For those, sklearn sorts an object array to find the unique labels, and this dominates
fit. Removing one of the three sorts saves roughly 35%. For example, pandas with 500k rows and 1 column goes from 877 ms to 531 ms.Full table
Behaviour
Before the change, I recorded the outputs of the base ref on 296 cases. These cover fit,
classes_and its dtype,encoder_dict_,binner_dict_,predict,predict_probaandpredict_log_proba(values, dtype, memory layout),scoreand every error message. The inputs were:""/"z"None, a scalar, a string, category dtype,Int64with and without NA,stringwith NA, object with None, mixed types, NaN, a dataframe, a 2-D list, a mismatched index, polars categorical, polars with nulls, polars booleanResults:
Tests
tests/test_prediction/test_target_mean_classifier.pyis rewritten to the conventions:make_dffixture runs every test on both backends. The data is a plain dict in the file, and the expected values are explicit.make_series, plus one test with list and numpy targets (with integer and string labels).match=re.escape(full message), including NotFittedError forpredict,predict_probaandpredict_log_proba.df_classificationis removed fromtests/test_prediction/conftest.py. Only this file used it.The tests cover
classes_and the predictions for 6 label types, sorted classes (the probability is that of the larger label), a probability of 0.5 returning the first class,predict_log_proba,score, the not-binary and continuous-target errors, and pandas-only tests for integer column names and a non-default index. The init-parameter errors are tested intest_base_predictor.py, which is part of #1066.Before/after (base ref = #1066 head,
6d284b4):TargetMeanClassifier)On the base ref, the new test file fails only on the list-of-string-labels cases.
flake8 feature_engine testsis clean.mypy feature_enginereports the same 2 errors as on the base ref, both in other modules.Needs decision
Find the unique labels once with
sklearn.utils._unique.attach_unique. This is the "attach_unique option" column in the benchmark. The code would be:sklearn's checks reuse the unique values attached to the array, so the object array is sorted once instead of twice. With string labels,
fitbecomes 2–5x faster than this PR:I did not implement it for two reasons:
attach_uniqueis available in every sklearn version we support (1.7 or later), but it is private.pd.Series([1, "a", ...])) raiseTypeError: '<' not supported between instances of 'str' and 'int'instead ofValueError: Unknown label type: unknown...[[0], [1], ...]) fits instead of raisingTypeError: cannot use 'list' as a set elementNoneor a scalar target raises "Unknown label type: unknown" instead of "Expected array-like...", unless we add anif np.asarray(y).ndim == 0guard.Do you want it?
Merge conflict with the TargetMeanRegressor PR. That PR is being migrated in parallel and probably edits
tests/test_prediction/conftest.pytoo, sincedf_regressionlives there. The conflict is trivial.Pre-existing issues, not fixed
fitsucceeds, butpredictandscoreraiseIndexError: index 1 is out of bounds for axis 0 with size 1, becausepredictreadsclasses_[1].predict_probareturns probabilities for 2 classes. Shouldfitraise a clear error instead? That would be a behaviour change.:ref:`User Guide <targetmeanclassifier>`points to a label that doesn't exist. No user guide page exists for this class.any(x for x in self.classes_ if x not in [0, 1])tests whether each label is truthy rather than just membership. It gives a different result only for a single-class target with label"", which is already covered by the single-class issue above.