Multiple-Imputation Contracts¶
The contracts in this module define validated boundaries for future FCS engines.
They do not alter existing impute_mice return values yet; use them to build and
validate explicit plans and result artifacts while migration proceeds.
Typed, validated contracts for multiple-imputation workflows.
These objects are intentionally independent from a particular imputation
algorithm. They establish the data and provenance boundary that a future
fully conditional specification (FCS) engine will consume without changing
the legacy missingly.impute.impute_mice() return contract prematurely.
- class missingly.mi_contracts.FCSKernel(*args, **kwargs)[source]¶
Bases:
ProtocolProtocol for one conditional model step in a future FCS engine.
Implementations receive the current chain state and return values only for the target column’s missing rows. Kernels must be explicit about their stochastic generator and must not mutate
statein place.- impute(state, target, mask, rng)[source]¶
Generate one target-column conditional draw.
- Parameters:
state (pandas.DataFrame) – Current completed chain state.
target (str) – Column being imputed.
mask (pandas.Series) – Boolean target-column missingness mask.
rng (numpy.random.Generator) – Chain-specific random generator.
- Returns:
Draws aligned to the
Trueentries ofmask.- Return type:
- class missingly.mi_contracts.ImputationPlan(methods, n_imputations, max_iter, seed_sequence, constraints=<factory>, provenance=<factory>)[source]¶
Bases:
objectDescribe a reproducible, explicit plan for multiple imputation.
- Parameters:
methods (mapping of str to str) – Per-column method identifiers. Keys must be schema column names and values must be nonempty public method identifiers.
n_imputations (int) – Number of completed data sets to generate.
max_iter (int) – Maximum conditional-model iterations per chain.
seed_sequence (tuple of int) – One deterministic seed per imputation chain.
constraints (mapping of str to str, optional) – Human-readable bounds, passive terms, or other future kernel rules.
provenance (mapping of str to Any, optional) – Raw-data-free reproducibility metadata such as an analysis digest.
Examples
>>> plan = ImputationPlan({"x": "pmm"}, 2, 5, (101, 202)) >>> plan.n_imputations 2
- validate_schema(schema)[source]¶
Check that all planned methods target columns in a schema.
- Parameters:
schema (MissingnessSchema) – Input schema against which method keys are validated.
- Raises:
TypeError – If
schemais not aMissingnessSchema.ValueError – If a method or constraint references an unknown column.
- Return type:
None
Examples
>>> schema = MissingnessSchema.from_dataframe(pd.DataFrame({"x": [1.0]})) >>> ImputationPlan({"x": "pmm"}, 1, 2, (7,)).validate_schema(schema)
- class missingly.mi_contracts.ImputationResult(data, histories=(), warnings=(), validation=<factory>)[source]¶
Bases:
objectAttach diagnostic traces and validation metadata to completed MI data.
- Parameters:
data (MIData) – Validated original and completed data sets.
histories (tuple of mappings, optional) – Per-chain, per-variable numerical diagnostic traces.
warnings (tuple of str, optional) – Explicit non-fatal algorithm warnings.
validation (mapping of str to bool, optional) – Named validation outcomes such as convergence or bounds checks.
Examples
>>> original = pd.DataFrame({"x": [1.0, np.nan]}) >>> schema = MissingnessSchema.from_dataframe(original) >>> plan = ImputationPlan({"x": "pmm"}, 1, 2, (9,)) >>> data = MIData(original, schema, plan, (pd.DataFrame({"x": [1.0, 2.0]}),)) >>> ImputationResult(data, histories=({"x": (1.5, 1.6)},)).validation["result_valid"] True
- class missingly.mi_contracts.MIData(original, schema, plan, imputations)[source]¶
Bases:
objectBind an incomplete input, schema, plan, and validated completed data sets.
- Parameters:
original (pandas.DataFrame) – Original incomplete data, structurally matching
schema.schema (MissingnessSchema) – Original structure and eligible-cell mask.
plan (ImputationPlan) – Reproducible per-column method and seed configuration.
imputations (tuple of pandas.DataFrame) – Completed data sets. Observed original cells must be unchanged.
Examples
>>> original = pd.DataFrame({"x": [1.0, np.nan]}) >>> schema = MissingnessSchema.from_dataframe(original) >>> plan = ImputationPlan({"x": "pmm"}, 1, 2, (9,)) >>> MIData(original, schema, plan, (pd.DataFrame({"x": [1.0, 2.0]}),)).n_imputations 1
- schema: MissingnessSchema¶
- plan: ImputationPlan¶
- class missingly.mi_contracts.MissingnessSchema(columns, dtypes, index, original_mask)[source]¶
Bases:
objectCapture the original DataFrame structure and cells eligible for imputation.
- Parameters:
columns (tuple of str) – Ordered unique column names from the original data.
dtypes (tuple of str) – Ordered pandas dtype strings corresponding to
columns.index (pandas.Index) – Exact original index. Duplicate labels are not allowed because result alignment must be unambiguous.
original_mask (pandas.DataFrame) – Boolean mask with
Trueexactly where original values were missing.
Examples
>>> frame = pd.DataFrame({"x": [1.0, np.nan]}, index=["a", "b"]) >>> schema = MissingnessSchema.from_dataframe(frame) >>> schema.original_mask["x"].tolist() [False, True]
- classmethod from_dataframe(frame)[source]¶
Create an immutable structural contract from a DataFrame.
- Parameters:
frame (pandas.DataFrame) – Original incomplete input. It is never modified.
- Returns:
Schema with a deep-copied boolean missing-value mask.
- Return type:
Examples
>>> frame = pd.DataFrame({"score": [1.0, np.nan]}) >>> MissingnessSchema.from_dataframe(frame).columns ('score',)
- validate_frame(frame, *, name='frame')[source]¶
Ensure a frame has the exact structure required by this schema.
- Parameters:
frame (pandas.DataFrame) – Frame to validate against this schema.
name (str, default="frame") – Name included in actionable validation errors.
- Raises:
TypeError – If
frameis not a DataFrame.ValueError – If columns, index, or dtype strings differ from the original.
- Return type:
None
Examples
>>> original = pd.DataFrame({"x": [1.0, np.nan]}) >>> schema = MissingnessSchema.from_dataframe(original) >>> schema.validate_frame(original.copy())