Data Manipulation

Data manipulation utilities for missing data workflows.

Compatibility

Compatible with Python 3.9+.

missingly.manipulation.replace_with_na(df, replace)[source]

Replace specified values in a DataFrame with NaN.

Parameters:
  • df (pd.DataFrame) – The dataframe to modify.

  • replace (dict) –

    A dictionary whose keys are column names and whose values describe which entries to replace. Each value may be:

    • a single scalar — replace that exact value;

    • a list of scalars — replace any value in the list;

    • a callable — replace where callable(cell) returns True.

Returns:

A new dataframe with the specified values replaced with NaN.

Return type:

pd.DataFrame

Examples

>>> import pandas as pd, numpy as np
>>> df = pd.DataFrame({'a': [1, -99, 3], 'b': ['x', 'N/A', 'z']})
>>> replace_with_na(df, replace={'a': -99, 'b': 'N/A'})
     a     b
0  1.0     x
1  NaN  None
2  3.0     z
missingly.manipulation.replace_with_na_all(df, condition)[source]

Replace all values in a DataFrame with NaN if they meet a condition.

Parameters:
  • df (pd.DataFrame) – The DataFrame to modify.

  • condition (callable) – A function that accepts a single cell value and returns True if that cell should be replaced with NaN.

Returns:

A new DataFrame with matching values replaced by NaN.

Return type:

pd.DataFrame

Examples

>>> import pandas as pd
>>> df = pd.DataFrame({'a': [1, -99, 3], 'b': [-99, 2, -99]})
>>> replace_with_na_all(df, condition=lambda x: x == -99)
     a    b
0  1.0  NaN
1  NaN  2.0
2  3.0  NaN
missingly.manipulation.add_any_miss_var(df, missing_values=None, col_name='any_miss')[source]

Add a boolean column indicating whether each row has any missing value.

Appends a single boolean column (any_miss by default) that is True for every row that contains at least one NaN (or any additional sentinel value supplied via missing_values).

Inspired by naniar::add_any_miss() in R.

Parameters:
  • df (pd.DataFrame) – Input DataFrame. Not modified in place.

  • missing_values (list, optional) – Additional scalar values to treat as missing alongside NaN (e.g. [-99, "N/A"]).

  • col_name (str, default "any_miss") – Name of the boolean indicator column to append.

Returns:

Copy of df with one extra boolean column appended on the right.

Return type:

pd.DataFrame

Raises:

ValueError – If col_name already exists in df to prevent silent overwrites.

Examples

>>> import pandas as pd, numpy as np
>>> df = pd.DataFrame({'a': [1.0, np.nan, 3.0], 'b': [4.0, 5.0, np.nan]})
>>> add_any_miss_var(df)
     a    b  any_miss
0  1.0  4.0     False
1  NaN  5.0      True
2  3.0  NaN      True

With a sentinel value:

>>> df2 = pd.DataFrame({'a': [1, -99, 3], 'b': [4, 5, 6]})
>>> add_any_miss_var(df2, missing_values=[-99])
   a  b  any_miss
0  1  4     False
1 -99  5      True
2  3  6     False
missingly.manipulation.bind_shadow_matrix(df, missing_values=None)[source]

Return the shadow matrix of a DataFrame as a standalone DataFrame.

Each column of the returned DataFrame corresponds to one column of the input, renamed <col>_NA, and contains True where the original value is missing and False where it is present.

Unlike bind_shadow() (which concatenates the shadow alongside the original data), this function returns only the shadow matrix — useful when you want to analyse or visualise the missingness pattern independently.

Parameters:
  • df (pd.DataFrame) – Input DataFrame. Not modified in place.

  • missing_values (list, optional) – Additional scalar sentinels treated as missing alongside NaN.

Returns:

Shape (n_rows, n_cols) with boolean dtype, column names ["<original_col>_NA", ...], and the same index as df.

Return type:

pd.DataFrame

Examples

>>> import pandas as pd, numpy as np
>>> df = pd.DataFrame({'a': [1.0, np.nan, 3.0], 'b': [np.nan, 2.0, 3.0]})
>>> bind_shadow_matrix(df)
    a_NA   b_NA
0  False   True
1   True  False
2  False  False

With a sentinel value:

>>> df2 = pd.DataFrame({'x': [0, -99, 2], 'y': [1, 2, -99]})
>>> bind_shadow_matrix(df2, missing_values=[-99])
    x_NA   y_NA
0  False  False
1   True  False
2  False   True
missingly.manipulation.clean_names(*args, **kwargs)[source]

Legacy shim — emits FutureWarning.

Deprecated since version 0.2.0: Moved to data_quality_toolkit.cleaning.

missingly.manipulation.remove_empty(*args, **kwargs)[source]

Legacy shim — emits FutureWarning.

Deprecated since version 0.2.0: Moved to data_quality_toolkit.cleaning.

missingly.manipulation.coalesce_columns(*args, **kwargs)[source]

Legacy shim — emits FutureWarning.

Deprecated since version 0.2.0: Moved to data_quality_toolkit.cleaning.

missingly.manipulation.miss_as_feature(df, columns=None, *, missing_values=None, suffix='_NA', keep_original=True)[source]

Encode missingness as binary indicator columns (experimental).

Deprecated since version 0.2.0: Experimental — may be moved to a separate package.

Parameters:
Return type:

DataFrame