Multiple-Imputation Contracts

The contracts in this module define validated boundaries for future FCS engines. They do not alter existing impute_mice return values yet; use them to build and validate explicit plans and result artifacts while migration proceeds.

Typed, validated contracts for multiple-imputation workflows.

These objects are intentionally independent from a particular imputation algorithm. They establish the data and provenance boundary that a future fully conditional specification (FCS) engine will consume without changing the legacy missingly.impute.impute_mice() return contract prematurely.

class missingly.mi_contracts.FCSKernel(*args, **kwargs)[source]

Bases: Protocol

Protocol for one conditional model step in a future FCS engine.

Implementations receive the current chain state and return values only for the target column’s missing rows. Kernels must be explicit about their stochastic generator and must not mutate state in place.

impute(state, target, mask, rng)[source]

Generate one target-column conditional draw.

Parameters:
Returns:

Draws aligned to the True entries of mask.

Return type:

pandas.Series

class missingly.mi_contracts.ImputationPlan(methods, n_imputations, max_iter, seed_sequence, constraints=<factory>, provenance=<factory>)[source]

Bases: object

Describe a reproducible, explicit plan for multiple imputation.

Parameters:
  • methods (mapping of str to str) – Per-column method identifiers. Keys must be schema column names and values must be nonempty public method identifiers.

  • n_imputations (int) – Number of completed data sets to generate.

  • max_iter (int) – Maximum conditional-model iterations per chain.

  • seed_sequence (tuple of int) – One deterministic seed per imputation chain.

  • constraints (mapping of str to str, optional) – Human-readable bounds, passive terms, or other future kernel rules.

  • provenance (mapping of str to Any, optional) – Raw-data-free reproducibility metadata such as an analysis digest.

Examples

>>> plan = ImputationPlan({"x": "pmm"}, 2, 5, (101, 202))
>>> plan.n_imputations
2
methods: Mapping[str, str]
n_imputations: int
max_iter: int
seed_sequence: Tuple[int, ...]
constraints: Mapping[str, str]
provenance: Mapping[str, Any]
validate_schema(schema)[source]

Check that all planned methods target columns in a schema.

Parameters:

schema (MissingnessSchema) – Input schema against which method keys are validated.

Raises:
Return type:

None

Examples

>>> schema = MissingnessSchema.from_dataframe(pd.DataFrame({"x": [1.0]}))
>>> ImputationPlan({"x": "pmm"}, 1, 2, (7,)).validate_schema(schema)
class missingly.mi_contracts.ImputationResult(data, histories=(), warnings=(), validation=<factory>)[source]

Bases: object

Attach diagnostic traces and validation metadata to completed MI data.

Parameters:
  • data (MIData) – Validated original and completed data sets.

  • histories (tuple of mappings, optional) – Per-chain, per-variable numerical diagnostic traces.

  • warnings (tuple of str, optional) – Explicit non-fatal algorithm warnings.

  • validation (mapping of str to bool, optional) – Named validation outcomes such as convergence or bounds checks.

Examples

>>> original = pd.DataFrame({"x": [1.0, np.nan]})
>>> schema = MissingnessSchema.from_dataframe(original)
>>> plan = ImputationPlan({"x": "pmm"}, 1, 2, (9,))
>>> data = MIData(original, schema, plan, (pd.DataFrame({"x": [1.0, 2.0]}),))
>>> ImputationResult(data, histories=({"x": (1.5, 1.6)},)).validation["result_valid"]
True
data: MIData
histories: Tuple[Mapping[str, Tuple[float, ...]], ...] = ()
warnings: Tuple[str, ...] = ()
validation: Mapping[str, bool]
class missingly.mi_contracts.MIData(original, schema, plan, imputations)[source]

Bases: object

Bind an incomplete input, schema, plan, and validated completed data sets.

Parameters:
  • original (pandas.DataFrame) – Original incomplete data, structurally matching schema.

  • schema (MissingnessSchema) – Original structure and eligible-cell mask.

  • plan (ImputationPlan) – Reproducible per-column method and seed configuration.

  • imputations (tuple of pandas.DataFrame) – Completed data sets. Observed original cells must be unchanged.

Examples

>>> original = pd.DataFrame({"x": [1.0, np.nan]})
>>> schema = MissingnessSchema.from_dataframe(original)
>>> plan = ImputationPlan({"x": "pmm"}, 1, 2, (9,))
>>> MIData(original, schema, plan, (pd.DataFrame({"x": [1.0, 2.0]}),)).n_imputations
1
original: DataFrame
schema: MissingnessSchema
plan: ImputationPlan
imputations: Tuple[DataFrame, ...]
property n_imputations: int

Return the number of completed data sets in the contract.

class missingly.mi_contracts.MissingnessSchema(columns, dtypes, index, original_mask)[source]

Bases: object

Capture the original DataFrame structure and cells eligible for imputation.

Parameters:
  • columns (tuple of str) – Ordered unique column names from the original data.

  • dtypes (tuple of str) – Ordered pandas dtype strings corresponding to columns.

  • index (pandas.Index) – Exact original index. Duplicate labels are not allowed because result alignment must be unambiguous.

  • original_mask (pandas.DataFrame) – Boolean mask with True exactly where original values were missing.

Examples

>>> frame = pd.DataFrame({"x": [1.0, np.nan]}, index=["a", "b"])
>>> schema = MissingnessSchema.from_dataframe(frame)
>>> schema.original_mask["x"].tolist()
[False, True]
columns: Tuple[str, ...]
dtypes: Tuple[str, ...]
index: Index
original_mask: DataFrame
classmethod from_dataframe(frame)[source]

Create an immutable structural contract from a DataFrame.

Parameters:

frame (pandas.DataFrame) – Original incomplete input. It is never modified.

Returns:

Schema with a deep-copied boolean missing-value mask.

Return type:

MissingnessSchema

Examples

>>> frame = pd.DataFrame({"score": [1.0, np.nan]})
>>> MissingnessSchema.from_dataframe(frame).columns
('score',)
validate_frame(frame, *, name='frame')[source]

Ensure a frame has the exact structure required by this schema.

Parameters:
  • frame (pandas.DataFrame) – Frame to validate against this schema.

  • name (str, default="frame") – Name included in actionable validation errors.

Raises:
  • TypeError – If frame is not a DataFrame.

  • ValueError – If columns, index, or dtype strings differ from the original.

Return type:

None

Examples

>>> original = pd.DataFrame({"x": [1.0, np.nan]})
>>> schema = MissingnessSchema.from_dataframe(original)
>>> schema.validate_frame(original.copy())