Summary Functions

Statistical diagnostics and summary statistics for missing data.

This module is the canonical home for all missingness analysis that does not involve imputation or visualisation. It merges what was previously split across stats.py and summary.py.

Public API — Summary

bind_shadow

Attach a boolean shadow matrix to a DataFrame.

n_miss / n_complete

Dataset-level missing / present counts.

pct_miss / pct_complete

Dataset-level missing / present percentages.

miss_var_summary

Per-column missingness table.

miss_case_summary

Per-row missingness table.

summarise

One-row dataset-level summary DataFrame.

miss_scan_count

Count specific sentinel values per column.

miss_summary

Combined dict of all three summary tables.

Public API — Statistical Tests

mcar_test

Little’s MCAR test (chi-square + p-value).

missingness_association_test

Exploratory observed-data association screen; does not identify MAR or MNAR.

mar_mnar_test

Deprecated compatibility wrapper for missingness_association_test.

diagnose_missing

MCAR evidence summary and assumption-aware planning guidance.

Public API — MICE Convergence Diagnostics

mice_convergence

Assess convergence of multiple MICE chains using iteration history.

plot_mice_convergence

Plot per-chain trace plots with Gelman-Rubin R-hat annotation.

missingly.diagnostics.bind_shadow(df, missing_values=None)[source]

Bind a boolean shadow matrix to a DataFrame.

The shadow matrix has one column per original column named <col>_NA; True means the value is missing.

Parameters:
  • df (pd.DataFrame)

  • missing_values (list, optional) – Sentinel values treated as missing in addition to NaN.

Returns:

Shape (n_rows, 2 * n_cols).

Return type:

pd.DataFrame

Examples

>>> import pandas as pd, numpy as np
>>> df = pd.DataFrame({'A': [1.0, np.nan], 'B': [np.nan, 2.0]})
>>> bind_shadow(df)
     A    B   A_NA   B_NA
0  1.0  NaN  False   True
1  NaN  2.0   True  False
missingly.diagnostics.n_miss(df, missing_values=None)[source]

Count total missing values in a DataFrame.

Parameters:
  • df (pd.DataFrame)

  • missing_values (list, optional) – Sentinel values treated as missing in addition to NaN.

Return type:

int

missingly.diagnostics.n_complete(df, missing_values=None)[source]

Count total non-missing values in a DataFrame.

Parameters:
  • df (pd.DataFrame)

  • missing_values (list, optional)

Return type:

int

missingly.diagnostics.pct_miss(df, missing_values=None)[source]

Percentage of missing values (0–100).

Parameters:
  • df (pd.DataFrame)

  • missing_values (list, optional)

Return type:

float

missingly.diagnostics.pct_complete(df, missing_values=None)[source]

Percentage of non-missing values (0–100).

Parameters:
  • df (pd.DataFrame)

  • missing_values (list, optional)

Return type:

float

missingly.diagnostics.miss_var_summary(df, missing_values=None)[source]

Summarise missingness by variable (column).

Parameters:
  • df (pd.DataFrame)

  • missing_values (list, optional)

Returns:

Columns: variable, n_miss, pct_miss.

Return type:

pd.DataFrame

missingly.diagnostics.miss_case_summary(df, missing_values=None)[source]

Summarise missingness by case (row).

Parameters:
  • df (pd.DataFrame)

  • missing_values (list, optional)

Returns:

Columns: case, n_miss, pct_miss.

Return type:

pd.DataFrame

missingly.diagnostics.summarise(df, missing_values=None)[source]

High-level one-row summary of missing data for the whole DataFrame.

Parameters:
  • df (pd.DataFrame)

  • missing_values (list, optional)

Returns:

One row with columns: n_rows, n_cols, n_cells, n_miss, n_complete, pct_miss, pct_complete.

Return type:

pd.DataFrame

Examples

>>> import pandas as pd, numpy as np
>>> df = pd.DataFrame({'a': [1.0, np.nan, 3.0], 'b': [np.nan, 2.0, 3.0]})
>>> summarise(df)
   n_rows  n_cols  n_cells  n_miss  n_complete   pct_miss  pct_complete
0       3       2        6       2           4  33.333333     66.666667
missingly.diagnostics.miss_scan_count(df, search, missing_values=None)[source]

Count occurrences of specific sentinel values per variable.

Parameters:
  • df (pd.DataFrame)

  • search (scalar or list) – Value(s) to search for (e.g. [-99, -98, "N/A"]).

  • missing_values (list, optional)

Returns:

Long-format with columns variable, value, n. Only rows where n > 0 are returned.

Return type:

pd.DataFrame

Examples

>>> import pandas as pd
>>> df = pd.DataFrame({'a': [-99, 1, -99], 'b': [2, -99, 3]})
>>> miss_scan_count(df, search=[-99])
  variable  value  n
0        a    -99  2
1        b    -99  1
missingly.diagnostics.miss_summary(df, missing_values=None, digits=2)[source]

Combined missing-data summary: dataset, variable, and case views.

Parameters:
  • df (pd.DataFrame)

  • missing_values (list, optional)

  • digits (int, default 2) – Decimal places for pct_miss / pct_complete columns.

Returns:

Keys: "summary", "by_variable", "by_case".

Return type:

dict

Examples

>>> import pandas as pd, numpy as np
>>> df = pd.DataFrame({'a': [1.0, np.nan], 'b': [np.nan, 2.0]})
>>> list(miss_summary(df).keys())
['summary', 'by_variable', 'by_case']
missingly.diagnostics.mcar_test(X, max_iter=500, tol=1e-07, ridge=1e-08, missing_values=None)[source]

Perform Little’s MCAR test.

Tests the null hypothesis that data are Missing Completely At Random (MCAR) under a multivariate-normal working model. A small p-value gives evidence against MCAR; a large p-value does not prove MCAR.

Parameters:
  • X (pd.DataFrame) – DataFrame with missing values. All columns must be numeric, finite when observed, and contain at least one observed value.

  • max_iter (int, default 500) – Maximum EM iterations.

  • tol (float, default 1e-7) – Maximum absolute EM parameter change required for convergence.

  • ridge (float, default 1e-8) – Non-negative diagonal stabilizer used only for conditional covariance solves. A positive value changes the numerical approximation.

  • missing_values (list, optional) – Extra sentinel values treated as missing.

Returns:

Keys: chi_square, df, p_value, missing_patterns, amount_missing, and em_iterations.

Return type:

dict

Raises:
  • TypeError – If X is not a pandas.DataFrame.

  • ValueError – If data are invalid, non-numeric, non-finite when observed, contain an all-missing column, or do not have enough informative rows.

  • RuntimeError – If observed-data EM does not converge within max_iter.

Examples

>>> import numpy as np
>>> import pandas as pd
>>> frame = pd.DataFrame({"x": [1.0, 2.0, np.nan, 4.0], "y": [1.0, np.nan, 3.0, 4.0]})
>>> "p_value" in mcar_test(frame)
True

Notes

This is not a general test for MAR versus MNAR, and non-rejection is not a basis for claiming MCAR or selecting a single imputation method.

References

Little, R. J. A. (1988). A test of missing completely at random for multivariate data with missing values. Journal of the American Statistical Association, 83(404), 1198-1202. https://doi.org/10.1080/01621459.1988.10478722

missingly.diagnostics.missingness_association_test(X, outcome, missing_values=None)[source]

Screen observedness associations with an observed auxiliary outcome.

For each feature with both observed and missing values, this diagnostic compares two logistic models for its observedness indicator: one using the other columns in X and one additionally using outcome. It tests whether that observed outcome is associated with observedness conditional on the supplied columns.

This is an exploratory association screen. It does not test or identify MAR versus MNAR: those mechanisms are not identifiable from observed data alone.

Parameters:
  • X (pandas.DataFrame) – DataFrame whose column-observedness indicators are screened.

  • outcome (array-like) – Fully observed auxiliary outcome of length len(X). Non-numeric values are deterministically factorized for the one-degree-of-freedom screening model.

  • missing_values (list, optional) – Extra sentinel values treated as missing in X.

Returns:

Columns are feature, lrt_statistic, p_value, and n_obs.

Return type:

pandas.DataFrame

Raises:
  • TypeError – If X is not a DataFrame.

  • ValueError – If outcome has a different length from X or contains missing values.

Examples

>>> import numpy as np
>>> import pandas as pd
>>> frame = pd.DataFrame({"age": [20.0, np.nan, 40.0], "group": [0, 1, 1]})
>>> list(missingness_association_test(frame, [0, 1, 1]).columns)
['feature', 'lrt_statistic', 'p_value', 'n_obs']

Notes

The likelihood-ratio reference distribution is an approximation and this screen is not a replacement for a pre-specified missing-data model. Treat p-values as exploratory, especially when screening many features.

missingly.diagnostics.mar_mnar_test(X, Y, missing_values=None)[source]

Deprecated wrapper for missingness_association_test().

Parameters:
  • X (pandas.DataFrame) – DataFrame whose observedness indicators are screened.

  • Y (array-like) – Fully observed auxiliary outcome retained for backward compatibility.

  • missing_values (list, optional) – Extra sentinel values treated as missing in X.

Returns:

Legacy (feature, lrt_statistic, p_value) tuples.

Return type:

list of tuple

Examples

>>> import numpy as np
>>> import pandas as pd
>>> mar_mnar_test(pd.DataFrame({"x": [1.0, np.nan, 3.0]}), [0, 1, 1])
[]

Warning

FutureWarning

The name is deprecated because observed data cannot identify MAR versus MNAR.

missingly.diagnostics.diagnose_missing(df, significance=0.05, missing_values=None, max_iter=100, tol=1e-05, ridge=1e-06)[source]

Summarise observed-data evidence relevant to a missing-data plan.

Runs Little’s MCAR test when its numeric-data preconditions are met and reports missingness prevalence plus descriptive nullity correlation. The result never identifies MAR versus MNAR: those mechanisms cannot be distinguished from observed data alone.

Parameters:
  • df (pd.DataFrame)

  • significance (float, default 0.05) – p-value threshold for the MCAR test.

  • missing_values (list, optional)

  • max_iter (int, default 100)

  • tol (float, default 1e-5)

  • ridge (float, default 1e-6)

Returns:

All keys from mcar_test() plus mcar_evidence, legacy-compatible mechanism, recommendation, strategy_hint, high_missingness_cols, and descriptive max_nullity_corr. The only valid values of mcar_evidence are "MCAR_not_rejected", "MCAR_rejected", and "insufficient_data".

Return type:

dict

Raises:

Examples

>>> import pandas as pd, numpy as np
>>> df = pd.DataFrame({
...     'age':    [25, np.nan, 35, 40, np.nan],
...     'income': [50000, 60000, np.nan, 80000, 55000],
... })
>>> result = diagnose_missing(df)
>>> result['mcar_evidence'] in ('MCAR_not_rejected', 'MCAR_rejected', 'insufficient_data')
True

Notes

A non-rejection of MCAR is not proof of MCAR. Nullity correlation describes joint observedness patterns; it does not identify MAR or MNAR. Choose the imputation and analysis model from substantive knowledge, then assess sensitivity to defensible missing-data assumptions.

missingly.diagnostics.mice_convergence(histories, variables=None, rhat_threshold=1.1)[source]

Assess convergence of multiple MICE chains using iteration history.

This replaces the old snapshot-based approach. The histories argument must come from impute_mice(..., return_history=True) which tracks the per-iteration mean of imputed values for each column.

Parameters:
  • histories (list of dict) – One history dict per chain as returned by impute_mice(..., n_imputations=m, return_history=True). Each dict maps column_name -> [mean_iter_1, ..., mean_iter_k].

  • variables (list of str, optional) – Specific variables to diagnose. If None, all variables found in the histories with at least 2 iterations are used.

  • rhat_threshold (float, default 1.1) – Threshold below which a variable is considered converged. R-hat < 1.1 is the standard criterion (Gelman & Rubin, 1992); R-hat < 1.05 is preferred for publication-quality analysis.

Returns:

Dictionary containing:

  • "rhat": Gelman-Rubin R-hat per variable (dict).

  • "trace_data": per-chain iteration means per variable (dict of dict).

  • "converged_by_variable": bool per variable (dict).

  • "converged": overall convergence flag (bool).

  • "n_chains": number of chains used (int).

  • "n_iter_used": minimum iterations used per variable (dict).

Return type:

dict

Raises:

ValueError – If fewer than 2 histories are provided.

Examples

>>> import pandas as pd, numpy as np
>>> from missingly.impute import impute_mice
>>> df = pd.DataFrame({"a": [1.0, np.nan, 3.0, 4.0, np.nan]})
>>> _, histories = impute_mice(
...     df, n_imputations=3, max_iter=5, return_history=True
... )
>>> result = mice_convergence(histories)
>>> "rhat" in result
True
>>> result["n_chains"]
3

Notes

R-hat is computed using per-iteration means of imputed values across chains, analogous to the trace-plot diagnostic in R’s mice package (van Buuren & Groothuis-Oudshoorn, 2011, J. Stat. Software 45(3)). The Gelman-Rubin statistic compares within-chain to between-chain variance (Gelman & Rubin, 1992, Statistical Science 7(4), 457-472). It is a diagnostic, not proof that the imputation model or missing-data assumptions are valid. Identical chain histories are reported as non-converged because they contain no observed between-chain stochasticity.

missingly.diagnostics.plot_mice_convergence(convergence_result, variables=None, figsize=(10, 4), rhat_threshold=1.1)[source]

Plot MICE convergence trace plots with Gelman-Rubin R-hat annotation.

Each panel shows per-chain iteration means for one imputed variable, with the R-hat value annotated in green (converged) or red (not converged). This is analogous to the plot() method on a mids object in R’s mice package.

Parameters:
  • convergence_result (dict) – Output from mice_convergence().

  • variables (list of str, optional) – Variables to plot. If None, all variables in the result are plotted.

  • figsize (tuple, default (10, 4)) – Figure size per variable panel (width, height).

  • rhat_threshold (float, default 1.1) – R-hat threshold for green/red annotation colouring.

Returns:

Displays matplotlib figure.

Return type:

None

Examples

>>> import pandas as pd, numpy as np
>>> from missingly.impute import impute_mice
>>> df = pd.DataFrame({"a": [1.0, np.nan, 3.0, 4.0, np.nan]})
>>> _, histories = impute_mice(
...     df, n_imputations=3, max_iter=5, return_history=True
... )
>>> result = mice_convergence(histories)
>>> plot_mice_convergence(result)