Imputation Comparison

Imputation comparison utilities.

This module provides two complementary benchmarking functions:

compare_imputations()

Quick exploration tool. Requires a complete DataFrame, masks a fraction of values, imputes, and measures reconstruction RMSE / accuracy. Fast and deterministic, but answers the wrong question for production use: it measures how well an imputer reconstructs masked values rather than how well the downstream model generalises.

cv_compare_imputations()

Production-quality benchmark. Accepts a DataFrame with real missing values and a supervised target vector. For each CV fold it fits the imputer on the training split (no leakage), transforms both splits, and scores the downstream estimator. Returns per-imputer mean and std of the cross-validated score so you can select the imputer that yields the best model — not the best RMSE on masked cells.

Compatibility

Requires Python 3.9+ and pandas 2.0+.

missingly.compare.compare_imputations(df, methods=None, mask_frac=0.2, random_state=42, missing_values=None)[source]

Compare imputation methods by masking a complete DataFrame and measuring reconstruction.

Artificially masks a fraction of values in each column, applies each imputation method, and evaluates reconstruction accuracy.

Note

This function answers “which imputer best reconstructs masked values?”, not “which imputer gives the best downstream model?”. For the latter use cv_compare_imputations().

Scoring:

  • Numeric columns → RMSE (lower is better).

  • Categorical columns → accuracy (higher is better).

  • Mixed → normalised composite Score (lower is better).

Parameters:
  • df (pd.DataFrame) – A complete DataFrame (no missing values after sentinel replacement) with at least one column. Mixed numeric/categorical dtypes are supported.

  • methods (list of callables, optional) – Imputation functions to compare. Each must accept a DataFrame and return a fully-imputed DataFrame. Defaults to all seven built-in methods: mean, median, mode, knn, mice, rf, gb.

  • mask_frac (float, optional) – Fraction of values to mask per column (default 0.20). Must be in (0, 1).

  • random_state (int, optional) – Random seed for reproducible masking. Default 42.

  • missing_values (list, optional) – Sentinel values (e.g. [-99, "N/A"]) to treat as missing before checking completeness. They are replaced with np.nan in a copy of df.

Returns:

DataFrame indexed by method name, sorted ascending by Score (or RMSE / Accuracy when only one column type exists). Methods that raise an exception during imputation are assigned NaN scores and flagged with an Error column.

Return type:

pd.DataFrame

Raises:
  • ValueError – If df still has missing values after sentinel replacement.

  • ValueError – If mask_frac is not in (0, 1).

Examples

>>> import pandas as pd, numpy as np
>>> df = pd.DataFrame({'age': [25, 30, 35, 40], 'city': ['A','B','A','B']})
>>> compare_imputations(df)
missingly.compare.cv_compare_imputations(X, y, estimator, strategies=None, n_splits=5, scoring=None, random_state=42, imputer_kwargs=None)[source]

Compare imputation strategies via cross-validated downstream model score.

This is the correct way to choose an imputer for production. For each fold of a K-Fold CV:

  1. Fit the imputer on X_train only (no leakage).

  2. Transform X_train and X_test using those fitted parameters.

  3. Fit the downstream estimator on imputed X_train.

  4. Score on imputed X_test.

The final result is the mean and standard deviation of the CV scores across all folds, giving a direct comparison of how each imputation strategy affects model generalisation.

Parameters:
  • X (pd.DataFrame) – Feature matrix, may contain real missing values. The imputer is fit on each training fold and applied to the test fold.

  • y (array-like) – Target vector. Must have the same length as X.

  • estimator (sklearn estimator) – Any fitted sklearn estimator with a predict (and optionally predict_proba) method. A fresh clone is used for each fold.

  • strategies (list of str, optional) – List of imputation strategy names to compare. Each must be one of the strategies supported by MissinglyImputer: "mean", "median", "mode", "knn", "mice", "rf", "gb", "pmm", "logreg", "polyreg", "polr". Defaults to ["mean", "median", "mode", "knn", "mice"].

  • n_splits (int, optional) – Number of CV folds. Default 5.

  • scoring (callable, optional) – A function scoring(estimator, X, y) -> float. Defaults to estimator.score (accuracy for classifiers, R² for regressors).

  • random_state (int, optional) – Random seed for KFold shuffling. Default 42.

  • imputer_kwargs (dict, optional) – Dict mapping strategy name to keyword arguments forwarded to MissinglyImputer. E.g. {"knn": {"n_neighbors": 3}}.

Returns:

Indexed by strategy name, with columns:

  • mean_score — mean CV score across folds.

  • std_score — standard deviation of CV scores.

Sorted descending by mean_score.

Return type:

pd.DataFrame

Raises:
  • ValueError – If strategies contains an unrecognised strategy name.

  • ValueError – If X and y have different lengths.

Examples

>>> import pandas as pd, numpy as np
>>> from sklearn.linear_model import LogisticRegression
>>> from missingly.compare import cv_compare_imputations
>>>
>>> rng = np.random.default_rng(0)
>>> X = pd.DataFrame({'a': rng.normal(size=100), 'b': rng.normal(size=100)})
>>> X.loc[rng.choice(100, 20, replace=False), 'a'] = np.nan
>>> y = rng.integers(0, 2, size=100)
>>>
>>> results = cv_compare_imputations(
...     X, y,
...     estimator=LogisticRegression(),
...     strategies=['mean', 'knn', 'pmm'],
...     n_splits=3,
... )
>>> print(results)