Summary Functions¶
Statistical diagnostics and summary statistics for missing data.
This module is the canonical home for all missingness analysis that
does not involve imputation or visualisation. It merges what was
previously split across stats.py and summary.py.
Public API — Summary¶
- bind_shadow
Attach a boolean shadow matrix to a DataFrame.
- n_miss / n_complete
Dataset-level missing / present counts.
- pct_miss / pct_complete
Dataset-level missing / present percentages.
- miss_var_summary
Per-column missingness table.
- miss_case_summary
Per-row missingness table.
- summarise
One-row dataset-level summary DataFrame.
- miss_scan_count
Count specific sentinel values per column.
- miss_summary
Combined dict of all three summary tables.
Public API — Statistical Tests¶
- mcar_test
Little’s MCAR test (chi-square + p-value).
- missingness_association_test
Exploratory observed-data association screen; does not identify MAR or MNAR.
- mar_mnar_test
Deprecated compatibility wrapper for
missingness_association_test.- diagnose_missing
MCAR evidence summary and assumption-aware planning guidance.
Public API — MICE Convergence Diagnostics¶
- mice_convergence
Assess convergence of multiple MICE chains using iteration history.
- plot_mice_convergence
Plot per-chain trace plots with Gelman-Rubin R-hat annotation.
- missingly.diagnostics.bind_shadow(df, missing_values=None)[source]
Bind a boolean shadow matrix to a DataFrame.
The shadow matrix has one column per original column named
<col>_NA;Truemeans the value is missing.- Parameters:
df (pd.DataFrame)
missing_values (list, optional) – Sentinel values treated as missing in addition to NaN.
- Returns:
Shape
(n_rows, 2 * n_cols).- Return type:
pd.DataFrame
Examples
>>> import pandas as pd, numpy as np >>> df = pd.DataFrame({'A': [1.0, np.nan], 'B': [np.nan, 2.0]}) >>> bind_shadow(df) A B A_NA B_NA 0 1.0 NaN False True 1 NaN 2.0 True False
- missingly.diagnostics.n_miss(df, missing_values=None)[source]
Count total missing values in a DataFrame.
- missingly.diagnostics.n_complete(df, missing_values=None)[source]
Count total non-missing values in a DataFrame.
- missingly.diagnostics.pct_miss(df, missing_values=None)[source]
Percentage of missing values (0–100).
- missingly.diagnostics.pct_complete(df, missing_values=None)[source]
Percentage of non-missing values (0–100).
- missingly.diagnostics.miss_var_summary(df, missing_values=None)[source]
Summarise missingness by variable (column).
- Parameters:
df (pd.DataFrame)
missing_values (list, optional)
- Returns:
Columns:
variable,n_miss,pct_miss.- Return type:
pd.DataFrame
- missingly.diagnostics.miss_case_summary(df, missing_values=None)[source]
Summarise missingness by case (row).
- Parameters:
df (pd.DataFrame)
missing_values (list, optional)
- Returns:
Columns:
case,n_miss,pct_miss.- Return type:
pd.DataFrame
- missingly.diagnostics.summarise(df, missing_values=None)[source]
High-level one-row summary of missing data for the whole DataFrame.
- Parameters:
df (pd.DataFrame)
missing_values (list, optional)
- Returns:
One row with columns:
n_rows,n_cols,n_cells,n_miss,n_complete,pct_miss,pct_complete.- Return type:
pd.DataFrame
Examples
>>> import pandas as pd, numpy as np >>> df = pd.DataFrame({'a': [1.0, np.nan, 3.0], 'b': [np.nan, 2.0, 3.0]}) >>> summarise(df) n_rows n_cols n_cells n_miss n_complete pct_miss pct_complete 0 3 2 6 2 4 33.333333 66.666667
- missingly.diagnostics.miss_scan_count(df, search, missing_values=None)[source]
Count occurrences of specific sentinel values per variable.
- Parameters:
- Returns:
Long-format with columns
variable,value,n. Only rows wheren > 0are returned.- Return type:
pd.DataFrame
Examples
>>> import pandas as pd >>> df = pd.DataFrame({'a': [-99, 1, -99], 'b': [2, -99, 3]}) >>> miss_scan_count(df, search=[-99]) variable value n 0 a -99 2 1 b -99 1
- missingly.diagnostics.miss_summary(df, missing_values=None, digits=2)[source]
Combined missing-data summary: dataset, variable, and case views.
- Parameters:
- Returns:
Keys:
"summary","by_variable","by_case".- Return type:
Examples
>>> import pandas as pd, numpy as np >>> df = pd.DataFrame({'a': [1.0, np.nan], 'b': [np.nan, 2.0]}) >>> list(miss_summary(df).keys()) ['summary', 'by_variable', 'by_case']
- missingly.diagnostics.mcar_test(X, max_iter=500, tol=1e-07, ridge=1e-08, missing_values=None)[source]
Perform Little’s MCAR test.
Tests the null hypothesis that data are Missing Completely At Random (MCAR) under a multivariate-normal working model. A small p-value gives evidence against MCAR; a large p-value does not prove MCAR.
- Parameters:
X (pd.DataFrame) – DataFrame with missing values. All columns must be numeric, finite when observed, and contain at least one observed value.
max_iter (int, default 500) – Maximum EM iterations.
tol (float, default 1e-7) – Maximum absolute EM parameter change required for convergence.
ridge (float, default 1e-8) – Non-negative diagonal stabilizer used only for conditional covariance solves. A positive value changes the numerical approximation.
missing_values (list, optional) – Extra sentinel values treated as missing.
- Returns:
Keys:
chi_square,df,p_value,missing_patterns,amount_missing, andem_iterations.- Return type:
- Raises:
TypeError – If X is not a
pandas.DataFrame.ValueError – If data are invalid, non-numeric, non-finite when observed, contain an all-missing column, or do not have enough informative rows.
RuntimeError – If observed-data EM does not converge within
max_iter.
Examples
>>> import numpy as np >>> import pandas as pd >>> frame = pd.DataFrame({"x": [1.0, 2.0, np.nan, 4.0], "y": [1.0, np.nan, 3.0, 4.0]}) >>> "p_value" in mcar_test(frame) True
Notes
This is not a general test for MAR versus MNAR, and non-rejection is not a basis for claiming MCAR or selecting a single imputation method.
References
Little, R. J. A. (1988). A test of missing completely at random for multivariate data with missing values. Journal of the American Statistical Association, 83(404), 1198-1202. https://doi.org/10.1080/01621459.1988.10478722
- missingly.diagnostics.missingness_association_test(X, outcome, missing_values=None)[source]
Screen observedness associations with an observed auxiliary outcome.
For each feature with both observed and missing values, this diagnostic compares two logistic models for its observedness indicator: one using the other columns in
Xand one additionally usingoutcome. It tests whether that observed outcome is associated with observedness conditional on the supplied columns.This is an exploratory association screen. It does not test or identify MAR versus MNAR: those mechanisms are not identifiable from observed data alone.
- Parameters:
X (pandas.DataFrame) – DataFrame whose column-observedness indicators are screened.
outcome (array-like) – Fully observed auxiliary outcome of length
len(X). Non-numeric values are deterministically factorized for the one-degree-of-freedom screening model.missing_values (list, optional) – Extra sentinel values treated as missing in
X.
- Returns:
Columns are
feature,lrt_statistic,p_value, andn_obs.- Return type:
- Raises:
TypeError – If
Xis not a DataFrame.ValueError – If
outcomehas a different length fromXor contains missing values.
Examples
>>> import numpy as np >>> import pandas as pd >>> frame = pd.DataFrame({"age": [20.0, np.nan, 40.0], "group": [0, 1, 1]}) >>> list(missingness_association_test(frame, [0, 1, 1]).columns) ['feature', 'lrt_statistic', 'p_value', 'n_obs']
Notes
The likelihood-ratio reference distribution is an approximation and this screen is not a replacement for a pre-specified missing-data model. Treat p-values as exploratory, especially when screening many features.
- missingly.diagnostics.mar_mnar_test(X, Y, missing_values=None)[source]
Deprecated wrapper for
missingness_association_test().- Parameters:
X (pandas.DataFrame) – DataFrame whose observedness indicators are screened.
Y (array-like) – Fully observed auxiliary outcome retained for backward compatibility.
missing_values (list, optional) – Extra sentinel values treated as missing in
X.
- Returns:
Legacy
(feature, lrt_statistic, p_value)tuples.- Return type:
Examples
>>> import numpy as np >>> import pandas as pd >>> mar_mnar_test(pd.DataFrame({"x": [1.0, np.nan, 3.0]}), [0, 1, 1]) []
Warning
- FutureWarning
The name is deprecated because observed data cannot identify MAR versus MNAR.
- missingly.diagnostics.diagnose_missing(df, significance=0.05, missing_values=None, max_iter=100, tol=1e-05, ridge=1e-06)[source]
Summarise observed-data evidence relevant to a missing-data plan.
Runs Little’s MCAR test when its numeric-data preconditions are met and reports missingness prevalence plus descriptive nullity correlation. The result never identifies MAR versus MNAR: those mechanisms cannot be distinguished from observed data alone.
- Parameters:
- Returns:
All keys from
mcar_test()plusmcar_evidence, legacy-compatiblemechanism,recommendation,strategy_hint,high_missingness_cols, and descriptivemax_nullity_corr. The only valid values ofmcar_evidenceare"MCAR_not_rejected","MCAR_rejected", and"insufficient_data".- Return type:
- Raises:
TypeError – If df is not a
pandas.DataFrame.ValueError – If df is empty.
Examples
>>> import pandas as pd, numpy as np >>> df = pd.DataFrame({ ... 'age': [25, np.nan, 35, 40, np.nan], ... 'income': [50000, 60000, np.nan, 80000, 55000], ... }) >>> result = diagnose_missing(df) >>> result['mcar_evidence'] in ('MCAR_not_rejected', 'MCAR_rejected', 'insufficient_data') True
Notes
A non-rejection of MCAR is not proof of MCAR. Nullity correlation describes joint observedness patterns; it does not identify MAR or MNAR. Choose the imputation and analysis model from substantive knowledge, then assess sensitivity to defensible missing-data assumptions.
- missingly.diagnostics.mice_convergence(histories, variables=None, rhat_threshold=1.1)[source]
Assess convergence of multiple MICE chains using iteration history.
This replaces the old snapshot-based approach. The histories argument must come from
impute_mice(..., return_history=True)which tracks the per-iteration mean of imputed values for each column.- Parameters:
histories (list of dict) – One history dict per chain as returned by
impute_mice(..., n_imputations=m, return_history=True). Each dict mapscolumn_name -> [mean_iter_1, ..., mean_iter_k].variables (list of str, optional) – Specific variables to diagnose. If None, all variables found in the histories with at least 2 iterations are used.
rhat_threshold (float, default 1.1) – Threshold below which a variable is considered converged. R-hat < 1.1 is the standard criterion (Gelman & Rubin, 1992); R-hat < 1.05 is preferred for publication-quality analysis.
- Returns:
Dictionary containing:
"rhat": Gelman-Rubin R-hat per variable (dict)."trace_data": per-chain iteration means per variable (dict of dict)."converged_by_variable": bool per variable (dict)."converged": overall convergence flag (bool)."n_chains": number of chains used (int)."n_iter_used": minimum iterations used per variable (dict).
- Return type:
- Raises:
ValueError – If fewer than 2 histories are provided.
Examples
>>> import pandas as pd, numpy as np >>> from missingly.impute import impute_mice >>> df = pd.DataFrame({"a": [1.0, np.nan, 3.0, 4.0, np.nan]}) >>> _, histories = impute_mice( ... df, n_imputations=3, max_iter=5, return_history=True ... ) >>> result = mice_convergence(histories) >>> "rhat" in result True >>> result["n_chains"] 3
Notes
R-hat is computed using per-iteration means of imputed values across chains, analogous to the trace-plot diagnostic in R’s
micepackage (van Buuren & Groothuis-Oudshoorn, 2011, J. Stat. Software 45(3)). The Gelman-Rubin statistic compares within-chain to between-chain variance (Gelman & Rubin, 1992, Statistical Science 7(4), 457-472). It is a diagnostic, not proof that the imputation model or missing-data assumptions are valid. Identical chain histories are reported as non-converged because they contain no observed between-chain stochasticity.
- missingly.diagnostics.plot_mice_convergence(convergence_result, variables=None, figsize=(10, 4), rhat_threshold=1.1)[source]
Plot MICE convergence trace plots with Gelman-Rubin R-hat annotation.
Each panel shows per-chain iteration means for one imputed variable, with the R-hat value annotated in green (converged) or red (not converged). This is analogous to the
plot()method on amidsobject in R’smicepackage.- Parameters:
convergence_result (dict) – Output from
mice_convergence().variables (list of str, optional) – Variables to plot. If None, all variables in the result are plotted.
figsize (tuple, default (10, 4)) – Figure size per variable panel
(width, height).rhat_threshold (float, default 1.1) – R-hat threshold for green/red annotation colouring.
- Returns:
Displays matplotlib figure.
- Return type:
None
Examples
>>> import pandas as pd, numpy as np >>> from missingly.impute import impute_mice >>> df = pd.DataFrame({"a": [1.0, np.nan, 3.0, 4.0, np.nan]}) >>> _, histories = impute_mice( ... df, n_imputations=3, max_iter=5, return_history=True ... ) >>> result = mice_convergence(histories) >>> plot_mice_convergence(result)