Imputation Functions¶
Imputation utilities for missing data.
This module provides two layers of imputation API:
Stateless functions (
impute_mean,impute_median, etc.) Accept a DataFrame, fit on that same DataFrame, and return a filled copy. Convenient for exploration, but they re-fit on every call so they must not be used across train/test splits (data leakage).FittedImputer (
make_imputerfactory) A lightweightfit/transform/fit_transformwrapper that stores per-column fill values (mean, median, or mode) computed on training data and applies them to unseen data. Supports only{mean, median, mode}— strategies whose fit-state is a simple scalar per column. For model-based strategies (knn, mice, rf, gb) useMissinglyImputerinstead.
Key design decisions¶
Python
Nonein object-dtype columns is normalised tonp.nanbefore any sklearn estimator sees the data.Numeric columns in ML-based imputers use the provided regressor.
Categorical columns use a classifier, avoiding the error of treating category codes as continuous values.
GradientBoostingRegressor/GradientBoostingClassifierdo not accept NaN in feature matrices. Any remaining NaN in the feature side is filled with column means computed from the training rows.
Error handling¶
All public functions validate their inputs eagerly and raise specific
missingly.exceptions exceptions on failure. No exception is
swallowed silently. When strict_mode=True (see missingly.config
or the global missingly.config.strict_mode setting)
the column-by-column imputer raises ImputationError
rather than falling back to a column mean.
Large-data warnings¶
Functions that are O(n²) or slow on large DataFrames emit a
UserWarning when the input exceeds the configured
large_df_threshold (default 50 000 rows; configure via
missingly.config.large_df_threshold).
- missingly.impute.impute_mean(df, numeric_only=False)[source]¶
Impute missing values with column means.
- Parameters:
df (pd.DataFrame) – Input DataFrame.
numeric_only (bool, default False) – If True, only numeric columns are imputed; non-numeric columns are left unchanged.
- Returns:
Copy of df with missing values replaced by column means.
- Return type:
pd.DataFrame
- Raises:
TypeError – If df is not a
pandas.DataFrame.ValueError – If df is empty.
Examples
>>> import pandas as pd, numpy as np >>> df = pd.DataFrame({"a": [1.0, np.nan, 3.0]}) >>> impute_mean(df) a 0 1.0 1 2.0 2 3.0
- missingly.impute.impute_median(df, numeric_only=False)[source]¶
Impute missing values with column medians.
- Parameters:
df (pd.DataFrame) – Input DataFrame.
numeric_only (bool, default False) – If True, only numeric columns are imputed.
- Returns:
Copy of df with missing values replaced by column medians.
- Return type:
pd.DataFrame
- Raises:
TypeError – If df is not a
pandas.DataFrame.ValueError – If df is empty.
Examples
>>> import pandas as pd, numpy as np >>> df = pd.DataFrame({"a": [1.0, np.nan, 3.0]}) >>> impute_median(df) a 0 1.0 1 2.0 2 3.0
- missingly.impute.impute_mode(df)[source]¶
Impute missing values with column modes (most frequent values).
- Parameters:
df (pd.DataFrame) – Input DataFrame.
- Returns:
Copy of df with missing values replaced by column modes.
- Return type:
pd.DataFrame
- Raises:
TypeError – If df is not a
pandas.DataFrame.ValueError – If df is empty.
Examples
>>> import pandas as pd, numpy as np >>> df = pd.DataFrame({"city": ["Berlin", np.nan, "Berlin"]}) >>> impute_mode(df) city 0 Berlin 1 Berlin 2 Berlin
- missingly.impute.impute_knn(df, n_neighbors=5, metric='euclidean')[source]¶
Impute missing values using k-Nearest Neighbours.
- Parameters:
df (pd.DataFrame) – Input DataFrame.
n_neighbors (int, default 5) – Number of nearest neighbours to use.
metric ({"euclidean", "mixed"}, default "euclidean") –
Distance metric to use.
"euclidean"— Ordinal-encode categorical columns and apply sklearn’sKNNImputerwith Euclidean distance. Fast and works well for purely numeric data."mixed"— Use Gower distance which handles numeric and categorical columns natively. Statistically sounder for heavy-categorical datasets, but O(n²). Not recommended forn > 10 000.
- Returns:
Copy of df with missing values imputed via KNN.
- Return type:
pd.DataFrame
- Raises:
TypeError – If df is not a
pandas.DataFrame.ValueError – If df is empty.
InvalidStrategyError – If metric is not one of the supported values.
TypeError – If n_neighbors is not a positive integer.
Examples
>>> import pandas as pd, numpy as np >>> df = pd.DataFrame({"age": [25.0, np.nan, 35.0, 40.0]}) >>> impute_knn(df, n_neighbors=2) age 0 25.0 1 30.0 2 35.0 3 40.0
- missingly.impute.impute_mice(df, max_iter=10, random_state=0, estimator=None, n_imputations=1, return_history=False)[source]¶
Impute missing values using MICE (Multiple Imputation by Chained Equations).
When
n_imputations > 1runs independent chains each with a distinct random seed so the returned DataFrames differ from one another.- Parameters:
df (pd.DataFrame) – Input DataFrame with missing values.
max_iter (int, default 10) – Maximum number of MICE imputation iterations per chain.
random_state (int, default 0) – Base random seed.
estimator (sklearn estimator or None, default None) – Regression estimator used per column. Defaults to
BayesianRidge()with posterior sampling enabled.n_imputations (int, default 1) – Number of independently imputed DataFrames to generate. Returns a single DataFrame when 1, a list when > 1.
return_history (bool, default False) –
If True, also return per-iteration imputed means for each variable that had missing values. Used by
mice_convergence()for trace plots and Gelman-Rubin R-hat computation.When
n_imputations == 1andreturn_history=Truethe return value is(imputed_df, history)where history isDict[str, List[float]]mapping column name to a list of per-iteration mean imputed values.When
n_imputations > 1andreturn_history=Truethe return value is(list_of_dfs, list_of_histories).With the default Bayesian ridge estimator, each chain uses posterior predictive draws. Histories therefore preserve between-chain stochasticity and are suitable for diagnostic use; they are still not proof that the imputation model or missing-data assumptions are valid.
- Returns:
Imputed copy (or list of m copies) of df. When
return_history=Truea tuple is returned instead — see above.- Return type:
pd.DataFrame or list of pd.DataFrame
- Raises:
TypeError – If df is not a
pandas.DataFrame.ValueError – If df is empty or n_imputations < 1.
Examples
>>> import pandas as pd, numpy as np >>> df = pd.DataFrame({"a": [1.0, np.nan, 3.0, 4.0]}) >>> result = impute_mice(df, max_iter=2, random_state=0) >>> result["a"].isna().any() False
>>> imputed, history = impute_mice(df, max_iter=5, return_history=True) >>> len(history["a"]) # one entry per iteration 5
- missingly.impute.impute_pmm(df, max_iter=10, random_state=0, n_nearest_donors=5)[source]¶
Impute missing values using Predictive Mean Matching (PMM).
PMM is the most commonly used method in R’s mice package. It works by:
Fitting a predictive model for each variable with missing values.
Using the model to predict values for missing observations.
Finding the closest observed values (donors) to each predicted value.
Randomly selecting one donor as the imputed value.
This preserves the distribution of the observed data better than direct regression imputation.
- Parameters:
df (pd.DataFrame) – Input DataFrame with missing values.
max_iter (int, default 10) – Maximum number of MICE iterations per chain.
random_state (int, default 0) – Random seed for reproducibility.
n_nearest_donors (int, default 5) – Number of nearest observed values to consider as donors for each missing value. Larger values increase randomness.
- Returns:
Imputed copy of df.
- Return type:
pd.DataFrame
- Raises:
TypeError – If df is not a
pandas.DataFrame.ValueError – If df is empty.
Examples
>>> import pandas as pd, numpy as np >>> df = pd.DataFrame({"a": [1.0, np.nan, 3.0, 4.0, 5.0]}) >>> result = impute_pmm(df, random_state=0) >>> result["a"].isna().any() False
Notes
PMM is preferred over direct regression imputation because it:
Preserves the distribution of observed values.
Does not create impossible values (e.g., negative ages).
Handles non-linear relationships through the matching step.
Is the gold standard method in R’s mice package.
- missingly.impute.impute_logreg(df, max_iter=10, random_state=0)[source]¶
Impute missing values using Logistic Regression (for binary variables).
This method is equivalent to R’s mice
logregmethod. It fits a logistic regression model for each binary variable with missing values, then samples from the predicted Bernoulli probability.- Parameters:
- Returns:
Imputed copy of df.
- Return type:
pd.DataFrame
- Raises:
TypeError – If df is not a
pandas.DataFrame.ValueError – If df is empty.
Examples
>>> import pandas as pd, numpy as np >>> df = pd.DataFrame({"a": [1.0, np.nan, 0.0, 1.0, 0.0]}) >>> result = impute_logreg(df, random_state=0) >>> result["a"].isna().any() False
Notes
This method is best suited for binary variables (0/1 or two-level categorical). For categorical variables with more than 2 levels, use
impute_polyreginstead.
- missingly.impute.impute_polyreg(df, max_iter=10, random_state=0)[source]¶
Impute missing values using Multinomial Logistic Regression.
This method is equivalent to R’s mice
polyregmethod. It fits a multinomial logistic regression model for each categorical variable with more than 2 levels.- Parameters:
- Returns:
Imputed copy of df.
- Return type:
pd.DataFrame
- Raises:
TypeError – If df is not a
pandas.DataFrame.ValueError – If df is empty.
Examples
>>> import pandas as pd, numpy as np >>> df = pd.DataFrame({"color": ["red", np.nan, "blue", "green", "red"]}) >>> result = impute_polyreg(df, random_state=0) >>> result["color"].isna().any() False
Notes
Best suited for nominal categorical variables with 3+ levels. For binary variables use
impute_logreg.
- missingly.impute.impute_polr(df, max_iter=10, random_state=0)[source]¶
Impute missing values using Ordinal Logistic Regression.
This method is equivalent to R’s mice
polrmethod. It fits an ordinal logistic regression model for each ordered categorical variable. Falls back to multinomial logistic regression when statsmodels is unavailable or the ordinal model fails to converge.- Parameters:
- Returns:
Imputed copy of df.
- Return type:
pd.DataFrame
- Raises:
TypeError – If df is not a
pandas.DataFrame.ValueError – If df is empty.
Examples
>>> import pandas as pd, numpy as np >>> df = pd.DataFrame({"grade": ["A", np.nan, "B", "C", "A"]}) >>> result = impute_polr(df, random_state=0) >>> result["grade"].isna().any() False
Notes
Best suited for ordinal categorical variables (natural ordering). For nominal categoricals use
impute_polyreginstead.Uses
statsmodels.miscmodels.ordinal_model.OrderedModelwhen available. If statsmodels is not installed or the model fails to converge, falls back to sklearn multinomial logistic regression with an audibleUserWarning.
- missingly.impute.impute_rf(df, max_iter=1, random_state=0, strict_mode=None, **rf_kwargs)[source]¶
Impute missing values using Random Forest.
- Parameters:
df (pd.DataFrame) – Input DataFrame.
max_iter (int, default 1) – Number of imputation passes over all columns.
random_state (int, default 0) – Random seed passed to the Random Forest estimators.
strict_mode (bool or None, default None) – If True, raise
ImputationErroron estimator failure instead of falling back to the column mean/mode. If None, uses the globalmissingly.config.strict_modesetting.**rf_kwargs – Additional keyword arguments forwarded to
RandomForestRegressorandRandomForestClassifier.
- Returns:
Copy of df with missing values imputed via Random Forest.
- Return type:
pd.DataFrame
- Raises:
TypeError – If df is not a
pandas.DataFrame.ValueError – If df is empty.
ImputationError – If an estimator fails and
strict_mode=True.InsufficientDataError – If a column has no observed values.
Examples
>>> import pandas as pd, numpy as np >>> df = pd.DataFrame({"a": [1.0, np.nan, 3.0, 4.0, 5.0]}) >>> result = impute_rf(df, random_state=0) >>> result["a"].isna().any() False
- missingly.impute.impute_gb(df, max_iter=1, random_state=0, strict_mode=None, **gb_kwargs)[source]¶
Impute missing values using Gradient Boosting.
- Parameters:
df (pd.DataFrame) – Input DataFrame.
max_iter (int, default 1) – Number of imputation passes over all columns.
random_state (int, default 0) – Random seed passed to the Gradient Boosting estimators.
strict_mode (bool or None, default None) – If True, raise
ImputationErroron estimator failure instead of falling back to the column mean/mode. If None, uses the globalmissingly.config.strict_modesetting.**gb_kwargs – Additional keyword arguments forwarded to
GradientBoostingRegressorandGradientBoostingClassifier.
- Returns:
Copy of df with missing values imputed via Gradient Boosting.
- Return type:
pd.DataFrame
- Raises:
TypeError – If df is not a
pandas.DataFrame.ValueError – If df is empty.
ImputationError – If an estimator fails and
strict_mode=True.InsufficientDataError – If a column has no observed values.
Examples
>>> import pandas as pd, numpy as np >>> df = pd.DataFrame({"a": [1.0, np.nan, 3.0, 4.0, 5.0]}) >>> result = impute_gb(df, random_state=0) >>> result["a"].isna().any() False
- class missingly.impute.FittedImputer(strategy='mean')[source]¶
Bases:
objectLightweight fit/transform imputer for simple strategies.
Learns per-column fill values (mean, median, or mode) from a training DataFrame and applies them to any subsequent DataFrame. This prevents data leakage when imputing test data inside a cross-validation loop.
Note
Only
{mean, median, mode}are supported — strategies whose fit-state is a scalar per column. For model-based strategies (knn,mice,rf,gb) useMissinglyImputerinstead, which implements the full sklearnBaseEstimator/TransformerMixincontract includingget_params/set_params/clone.- Parameters:
strategy ({"mean", "median", "mode"}) – Fill strategy to use.
- Raises:
InvalidStrategyError – If strategy is not one of
{"mean", "median", "mode"}.
Examples
>>> import pandas as pd, numpy as np >>> from missingly.impute import make_imputer >>> train = pd.DataFrame({"a": [1.0, 2.0, 3.0, np.nan]}) >>> test = pd.DataFrame({"a": [np.nan, 5.0]}) >>> imp = make_imputer("median").fit(train) >>> imp.transform(test) a 0 2.0 1 5.0
- fit(df)[source]¶
Learn fill values from df.
- Parameters:
df (pd.DataFrame) – Training data. Only non-missing values are used.
- Returns:
self (for method chaining).
- Return type:
- transform(df)[source]¶
Apply learned fill values to df.
- Parameters:
df (pd.DataFrame) – Data to impute. Must have the same columns as the training DataFrame passed to
fit.- Returns:
Imputed copy of df.
- Return type:
pd.DataFrame
- Raises:
RuntimeError – If
transformis called beforefit.
- missingly.impute.make_imputer(strategy='mean')[source]¶
Factory function that returns a
FittedImputer.Only simple strategies are supported:
{mean, median, mode}. For model-based strategies useMissinglyImputer.- Parameters:
strategy ({"mean", "median", "mode"}, default "mean") – Fill strategy.
- Return type:
- Raises:
InvalidStrategyError – If strategy is not one of the supported values.
Examples
>>> imp = make_imputer("median") >>> type(imp) <class 'missingly.impute.FittedImputer'>