Case Study: Analyzing Missing Data

In this notebook, we’ll walk through a case study of how to use missingly to analyze a dataset with missing values.

1. The Dataset

We’ll create a simulated dataset representing sensor readings from a fictional industrial process. The dataset has some missing values, which could be due to sensor malfunction or other issues.

[1]:
import pandas as pd
import numpy as np
import missingly as mi

np.random.seed(42)

data = {
    'temperature': np.random.normal(100, 5, 100),
    'pressure': np.random.normal(50, 2, 100),
    'humidity': np.random.normal(40, 5, 100),
    'vibration': np.random.normal(10, 1, 100)
}
df = pd.DataFrame(data)

# Introduce some missing values
df.loc[df.sample(frac=0.1).index, 'temperature'] = np.nan
high_temp_idx = df[df['temperature'] > 105].index
df.loc[high_temp_idx, 'pressure'] = np.nan
both_missing_idx = df.sample(frac=0.05).index
df.loc[both_missing_idx, ['humidity', 'vibration']] = np.nan

df.head()
[1]:
temperature pressure humidity vibration
0 102.483571 47.169259 41.788937 9.171005
1 99.308678 49.158709 42.803923 9.439819
2 103.238443 49.314571 45.415256 10.747294
3 107.615149 NaN 45.269010 10.610370
4 98.829233 49.677429 33.111653 9.979098

2. Initial Assessment

[2]:
mi.miss_var_summary(df)
[2]:
variable n_miss pct_miss
0 temperature 10 10.0
1 pressure 10 10.0
2 humidity 5 5.0
3 vibration 5 5.0
[3]:
mi.matrix(df)
[3]:
<Axes: >

3. Deeper Dive with Visualizations

[4]:
mi.upset(df)
[4]:
{'intersections': <Axes: title={'center': 'Missingness pattern intersections'}, ylabel='Rows'>,
 'matrix': <Axes: xlabel='Pattern'>,
 'totals': <Axes: xlabel='Missing'>}
[5]:
mi.dendrogram(df)
[5]:
<Axes: title={'center': 'Dendrogram of Variables by Missing Data Patterns'}>

4. Testing for MCAR

[6]:
mi.mcar_test(df)
[6]:
{'chi_square': np.float64(40.1573415928144),
 'df': np.int64(10),
 'p_value': 1.58970524641866e-05,
 'missing_patterns': 6,
 'amount_missing':                  temperature  pressure  humidity  vibration
 Number Missing          10.0      10.0      5.00       5.00
 Percent Missing          0.1       0.1      0.05       0.05,
 'em_iterations': 21}