Imputation Functions

Imputation utilities for missing data.

This module provides two layers of imputation API:

  1. Stateless functions (impute_mean, impute_median, etc.) Accept a DataFrame, fit on that same DataFrame, and return a filled copy. Convenient for exploration, but they re-fit on every call so they must not be used across train/test splits (data leakage).

  2. FittedImputer (make_imputer factory) A lightweight fit / transform / fit_transform wrapper that stores per-column fill values (mean, median, or mode) computed on training data and applies them to unseen data. Supports only {mean, median, mode} — strategies whose fit-state is a simple scalar per column. For model-based strategies (knn, mice, rf, gb) use MissinglyImputer instead.

Key design decisions

  • Python None in object-dtype columns is normalised to np.nan before any sklearn estimator sees the data.

  • Numeric columns in ML-based imputers use the provided regressor.

  • Categorical columns use a classifier, avoiding the error of treating category codes as continuous values.

  • GradientBoostingRegressor / GradientBoostingClassifier do not accept NaN in feature matrices. Any remaining NaN in the feature side is filled with column means computed from the training rows.

Error handling

All public functions validate their inputs eagerly and raise specific missingly.exceptions exceptions on failure. No exception is swallowed silently. When strict_mode=True (see missingly.config or the global missingly.config.strict_mode setting) the column-by-column imputer raises ImputationError rather than falling back to a column mean.

Large-data warnings

Functions that are O(n²) or slow on large DataFrames emit a UserWarning when the input exceeds the configured large_df_threshold (default 50 000 rows; configure via missingly.config.large_df_threshold).

missingly.impute.impute_mean(df, numeric_only=False)[source]

Impute missing values with column means.

Parameters:
  • df (pd.DataFrame) – Input DataFrame.

  • numeric_only (bool, default False) – If True, only numeric columns are imputed; non-numeric columns are left unchanged.

Returns:

Copy of df with missing values replaced by column means.

Return type:

pd.DataFrame

Raises:

Examples

>>> import pandas as pd, numpy as np
>>> df = pd.DataFrame({"a": [1.0, np.nan, 3.0]})
>>> impute_mean(df)
     a
0  1.0
1  2.0
2  3.0
missingly.impute.impute_median(df, numeric_only=False)[source]

Impute missing values with column medians.

Parameters:
  • df (pd.DataFrame) – Input DataFrame.

  • numeric_only (bool, default False) – If True, only numeric columns are imputed.

Returns:

Copy of df with missing values replaced by column medians.

Return type:

pd.DataFrame

Raises:

Examples

>>> import pandas as pd, numpy as np
>>> df = pd.DataFrame({"a": [1.0, np.nan, 3.0]})
>>> impute_median(df)
     a
0  1.0
1  2.0
2  3.0
missingly.impute.impute_mode(df)[source]

Impute missing values with column modes (most frequent values).

Parameters:

df (pd.DataFrame) – Input DataFrame.

Returns:

Copy of df with missing values replaced by column modes.

Return type:

pd.DataFrame

Raises:

Examples

>>> import pandas as pd, numpy as np
>>> df = pd.DataFrame({"city": ["Berlin", np.nan, "Berlin"]})
>>> impute_mode(df)
    city
0  Berlin
1  Berlin
2  Berlin
missingly.impute.impute_knn(df, n_neighbors=5, metric='euclidean')[source]

Impute missing values using k-Nearest Neighbours.

Parameters:
  • df (pd.DataFrame) – Input DataFrame.

  • n_neighbors (int, default 5) – Number of nearest neighbours to use.

  • metric ({"euclidean", "mixed"}, default "euclidean") –

    Distance metric to use.

    • "euclidean" — Ordinal-encode categorical columns and apply sklearn’s KNNImputer with Euclidean distance. Fast and works well for purely numeric data.

    • "mixed" — Use Gower distance which handles numeric and categorical columns natively. Statistically sounder for heavy-categorical datasets, but O(n²). Not recommended for n > 10 000.

Returns:

Copy of df with missing values imputed via KNN.

Return type:

pd.DataFrame

Raises:

Examples

>>> import pandas as pd, numpy as np
>>> df = pd.DataFrame({"age": [25.0, np.nan, 35.0, 40.0]})
>>> impute_knn(df, n_neighbors=2)
    age
0  25.0
1  30.0
2  35.0
3  40.0
missingly.impute.impute_mice(df, max_iter=10, random_state=0, estimator=None, n_imputations=1, return_history=False)[source]

Impute missing values using MICE (Multiple Imputation by Chained Equations).

When n_imputations > 1 runs independent chains each with a distinct random seed so the returned DataFrames differ from one another.

Parameters:
  • df (pd.DataFrame) – Input DataFrame with missing values.

  • max_iter (int, default 10) – Maximum number of MICE imputation iterations per chain.

  • random_state (int, default 0) – Base random seed.

  • estimator (sklearn estimator or None, default None) – Regression estimator used per column. Defaults to BayesianRidge() with posterior sampling enabled.

  • n_imputations (int, default 1) – Number of independently imputed DataFrames to generate. Returns a single DataFrame when 1, a list when > 1.

  • return_history (bool, default False) –

    If True, also return per-iteration imputed means for each variable that had missing values. Used by mice_convergence() for trace plots and Gelman-Rubin R-hat computation.

    When n_imputations == 1 and return_history=True the return value is (imputed_df, history) where history is Dict[str, List[float]] mapping column name to a list of per-iteration mean imputed values.

    When n_imputations > 1 and return_history=True the return value is (list_of_dfs, list_of_histories).

    With the default Bayesian ridge estimator, each chain uses posterior predictive draws. Histories therefore preserve between-chain stochasticity and are suitable for diagnostic use; they are still not proof that the imputation model or missing-data assumptions are valid.

Returns:

Imputed copy (or list of m copies) of df. When return_history=True a tuple is returned instead — see above.

Return type:

pd.DataFrame or list of pd.DataFrame

Raises:

Examples

>>> import pandas as pd, numpy as np
>>> df = pd.DataFrame({"a": [1.0, np.nan, 3.0, 4.0]})
>>> result = impute_mice(df, max_iter=2, random_state=0)
>>> result["a"].isna().any()
False
>>> imputed, history = impute_mice(df, max_iter=5, return_history=True)
>>> len(history["a"])  # one entry per iteration
5
missingly.impute.impute_pmm(df, max_iter=10, random_state=0, n_nearest_donors=5)[source]

Impute missing values using Predictive Mean Matching (PMM).

PMM is the most commonly used method in R’s mice package. It works by:

  1. Fitting a predictive model for each variable with missing values.

  2. Using the model to predict values for missing observations.

  3. Finding the closest observed values (donors) to each predicted value.

  4. Randomly selecting one donor as the imputed value.

This preserves the distribution of the observed data better than direct regression imputation.

Parameters:
  • df (pd.DataFrame) – Input DataFrame with missing values.

  • max_iter (int, default 10) – Maximum number of MICE iterations per chain.

  • random_state (int, default 0) – Random seed for reproducibility.

  • n_nearest_donors (int, default 5) – Number of nearest observed values to consider as donors for each missing value. Larger values increase randomness.

Returns:

Imputed copy of df.

Return type:

pd.DataFrame

Raises:

Examples

>>> import pandas as pd, numpy as np
>>> df = pd.DataFrame({"a": [1.0, np.nan, 3.0, 4.0, 5.0]})
>>> result = impute_pmm(df, random_state=0)
>>> result["a"].isna().any()
False

Notes

PMM is preferred over direct regression imputation because it:

  • Preserves the distribution of observed values.

  • Does not create impossible values (e.g., negative ages).

  • Handles non-linear relationships through the matching step.

  • Is the gold standard method in R’s mice package.

missingly.impute.impute_logreg(df, max_iter=10, random_state=0)[source]

Impute missing values using Logistic Regression (for binary variables).

This method is equivalent to R’s mice logreg method. It fits a logistic regression model for each binary variable with missing values, then samples from the predicted Bernoulli probability.

Parameters:
  • df (pd.DataFrame) – Input DataFrame with missing values.

  • max_iter (int, default 10) – Maximum number of MICE iterations.

  • random_state (int, default 0) – Random seed for reproducibility.

Returns:

Imputed copy of df.

Return type:

pd.DataFrame

Raises:

Examples

>>> import pandas as pd, numpy as np
>>> df = pd.DataFrame({"a": [1.0, np.nan, 0.0, 1.0, 0.0]})
>>> result = impute_logreg(df, random_state=0)
>>> result["a"].isna().any()
False

Notes

This method is best suited for binary variables (0/1 or two-level categorical). For categorical variables with more than 2 levels, use impute_polyreg instead.

missingly.impute.impute_polyreg(df, max_iter=10, random_state=0)[source]

Impute missing values using Multinomial Logistic Regression.

This method is equivalent to R’s mice polyreg method. It fits a multinomial logistic regression model for each categorical variable with more than 2 levels.

Parameters:
  • df (pd.DataFrame) – Input DataFrame with missing values.

  • max_iter (int, default 10) – Maximum number of MICE iterations.

  • random_state (int, default 0) – Random seed for reproducibility.

Returns:

Imputed copy of df.

Return type:

pd.DataFrame

Raises:

Examples

>>> import pandas as pd, numpy as np
>>> df = pd.DataFrame({"color": ["red", np.nan, "blue", "green", "red"]})
>>> result = impute_polyreg(df, random_state=0)
>>> result["color"].isna().any()
False

Notes

Best suited for nominal categorical variables with 3+ levels. For binary variables use impute_logreg.

missingly.impute.impute_polr(df, max_iter=10, random_state=0)[source]

Impute missing values using Ordinal Logistic Regression.

This method is equivalent to R’s mice polr method. It fits an ordinal logistic regression model for each ordered categorical variable. Falls back to multinomial logistic regression when statsmodels is unavailable or the ordinal model fails to converge.

Parameters:
  • df (pd.DataFrame) – Input DataFrame with missing values.

  • max_iter (int, default 10) – Maximum number of MICE iterations.

  • random_state (int, default 0) – Random seed for reproducibility.

Returns:

Imputed copy of df.

Return type:

pd.DataFrame

Raises:

Examples

>>> import pandas as pd, numpy as np
>>> df = pd.DataFrame({"grade": ["A", np.nan, "B", "C", "A"]})
>>> result = impute_polr(df, random_state=0)
>>> result["grade"].isna().any()
False

Notes

Best suited for ordinal categorical variables (natural ordering). For nominal categoricals use impute_polyreg instead.

Uses statsmodels.miscmodels.ordinal_model.OrderedModel when available. If statsmodels is not installed or the model fails to converge, falls back to sklearn multinomial logistic regression with an audible UserWarning.

missingly.impute.impute_rf(df, max_iter=1, random_state=0, strict_mode=None, **rf_kwargs)[source]

Impute missing values using Random Forest.

Parameters:
  • df (pd.DataFrame) – Input DataFrame.

  • max_iter (int, default 1) – Number of imputation passes over all columns.

  • random_state (int, default 0) – Random seed passed to the Random Forest estimators.

  • strict_mode (bool or None, default None) – If True, raise ImputationError on estimator failure instead of falling back to the column mean/mode. If None, uses the global missingly.config.strict_mode setting.

  • **rf_kwargs – Additional keyword arguments forwarded to RandomForestRegressor and RandomForestClassifier.

Returns:

Copy of df with missing values imputed via Random Forest.

Return type:

pd.DataFrame

Raises:
  • TypeError – If df is not a pandas.DataFrame.

  • ValueError – If df is empty.

  • ImputationError – If an estimator fails and strict_mode=True.

  • InsufficientDataError – If a column has no observed values.

Examples

>>> import pandas as pd, numpy as np
>>> df = pd.DataFrame({"a": [1.0, np.nan, 3.0, 4.0, 5.0]})
>>> result = impute_rf(df, random_state=0)
>>> result["a"].isna().any()
False
missingly.impute.impute_gb(df, max_iter=1, random_state=0, strict_mode=None, **gb_kwargs)[source]

Impute missing values using Gradient Boosting.

Parameters:
  • df (pd.DataFrame) – Input DataFrame.

  • max_iter (int, default 1) – Number of imputation passes over all columns.

  • random_state (int, default 0) – Random seed passed to the Gradient Boosting estimators.

  • strict_mode (bool or None, default None) – If True, raise ImputationError on estimator failure instead of falling back to the column mean/mode. If None, uses the global missingly.config.strict_mode setting.

  • **gb_kwargs – Additional keyword arguments forwarded to GradientBoostingRegressor and GradientBoostingClassifier.

Returns:

Copy of df with missing values imputed via Gradient Boosting.

Return type:

pd.DataFrame

Raises:
  • TypeError – If df is not a pandas.DataFrame.

  • ValueError – If df is empty.

  • ImputationError – If an estimator fails and strict_mode=True.

  • InsufficientDataError – If a column has no observed values.

Examples

>>> import pandas as pd, numpy as np
>>> df = pd.DataFrame({"a": [1.0, np.nan, 3.0, 4.0, 5.0]})
>>> result = impute_gb(df, random_state=0)
>>> result["a"].isna().any()
False
class missingly.impute.FittedImputer(strategy='mean')[source]

Bases: object

Lightweight fit/transform imputer for simple strategies.

Learns per-column fill values (mean, median, or mode) from a training DataFrame and applies them to any subsequent DataFrame. This prevents data leakage when imputing test data inside a cross-validation loop.

Note

Only {mean, median, mode} are supported — strategies whose fit-state is a scalar per column. For model-based strategies (knn, mice, rf, gb) use MissinglyImputer instead, which implements the full sklearn BaseEstimator / TransformerMixin contract including get_params / set_params / clone.

Parameters:

strategy ({"mean", "median", "mode"}) – Fill strategy to use.

Raises:

InvalidStrategyError – If strategy is not one of {"mean", "median", "mode"}.

Examples

>>> import pandas as pd, numpy as np
>>> from missingly.impute import make_imputer
>>> train = pd.DataFrame({"a": [1.0, 2.0, 3.0, np.nan]})
>>> test  = pd.DataFrame({"a": [np.nan, 5.0]})
>>> imp = make_imputer("median").fit(train)
>>> imp.transform(test)
     a
0  2.0
1  5.0
fit(df)[source]

Learn fill values from df.

Parameters:

df (pd.DataFrame) – Training data. Only non-missing values are used.

Returns:

self (for method chaining).

Return type:

FittedImputer

transform(df)[source]

Apply learned fill values to df.

Parameters:

df (pd.DataFrame) – Data to impute. Must have the same columns as the training DataFrame passed to fit.

Returns:

Imputed copy of df.

Return type:

pd.DataFrame

Raises:

RuntimeError – If transform is called before fit.

fit_transform(df)[source]

Fit on df then transform df.

Parameters:

df (pd.DataFrame) – Data to fit and impute.

Returns:

Imputed copy of df using fill values learned from df itself.

Return type:

pd.DataFrame

missingly.impute.make_imputer(strategy='mean')[source]

Factory function that returns a FittedImputer.

Only simple strategies are supported: {mean, median, mode}. For model-based strategies use MissinglyImputer.

Parameters:

strategy ({"mean", "median", "mode"}, default "mean") – Fill strategy.

Return type:

FittedImputer

Raises:

InvalidStrategyError – If strategy is not one of the supported values.

Examples

>>> imp = make_imputer("median")
>>> type(imp)
<class 'missingly.impute.FittedImputer'>