Cross-tool Benchmark Evidence

Use a BenchmarkManifest to commit the precise metadata and scalar oracle results for a cross-tool comparison. The manifest deliberately stores dataset identity and expected metrics, never raw dataset records.

Versioned, privacy-preserving contracts for cross-tool benchmarks.

Benchmark manifests capture the evidence needed to reproduce a numerical comparison without embedding the benchmark’s raw records. They are immutable, strictly validated, and have a canonical JSON representation suitable for reviewed fixtures and artifact indexes.

class missingly.benchmark.BenchmarkManifest(dataset_name, dataset_license, dataset_sha256, dataset_source_url, reference_tool, reference_tool_version, reference_platform, reference_script_sha256, method, missing_value_policy, estimand, expected_metrics, tolerances, known_differences=(), seed=None, schema_version=2, missingly_version='1.0.0')[source]

Bases: object

Describe one reproducible cross-tool benchmark without raw records.

Parameters:
  • dataset_name (str) – Stable dataset identifier used in the repository or public catalogue.

  • dataset_license (str) – Licence or terms that permit the benchmark fixture to be used.

  • dataset_sha256 (str) – Lowercase SHA-256 of the exact source dataset bytes.

  • dataset_source_url (str) – Canonical source URL for the dataset, not a path to a local copy.

  • reference_tool (str) – Tool that generated the frozen reference results, such as "R mice".

  • reference_tool_version (str) – Exact version of the reference tool or package.

  • reference_platform (str) – Operating system and runtime/platform used for the reference output.

  • reference_script_sha256 (str) – Lowercase SHA-256 of the script that produced the reference output.

  • method (str) – Fully qualified reference method and material configuration summary.

  • missing_value_policy (str) – How each tool interprets sentinels, nulls, and dropped rows.

  • estimand (str) – Quantity being compared, for example "Little MCAR chi-square".

  • expected_metrics (mapping of str to float) – Frozen scalar reference results. These are outputs, not raw records.

  • tolerances (mapping of str to float) – Non-negative absolute or relative tolerances keyed by metric name.

  • known_differences (tuple of str, optional) – Intentional algorithmic differences that reviewers must consider.

  • seed (int, optional) – Random seed used by the reference tool. None means deterministic.

  • schema_version (int, default=1) – Manifest schema version. Only the current version is accepted.

  • missingly_version (str, optional) – Missingly version used to reproduce the comparison. Defaults to the installed package version.

Raises:
  • TypeError – If a field has an incompatible type.

  • ValueError – If an identity field is empty, a digest is invalid, a metric is not finite, or a tolerance is negative.

Examples

>>> manifest = BenchmarkManifest(
...     dataset_name="airquality",
...     dataset_license="CC0-1.0",
...     dataset_sha256="a" * 64,
...     dataset_source_url="https://example.test/airquality.csv",
...     reference_tool="R naniar",
...     reference_tool_version="1.1.0",
...     reference_platform="R 4.4 on Linux",
...     reference_script_sha256="b" * 64,
...     method="naniar::mcar_test",
...     missing_value_policy="NA values retained",
...     estimand="Little MCAR chi-square",
...     expected_metrics={"chi_square": 25.1},
...     tolerances={"chi_square": 1e-6},
... )
>>> len(manifest.sha256())
64
>>> "raw records" in manifest.to_json()
False
dataset_name: str
dataset_license: str
dataset_sha256: str
dataset_source_url: str
reference_tool: str
reference_tool_version: str
reference_platform: str
reference_script_sha256: str
method: str
missing_value_policy: str
estimand: str
expected_metrics: Mapping[str, float]
tolerances: Mapping[str, float]
known_differences: Tuple[str, ...] = ()
seed: int | None = None
schema_version: int = 2
missingly_version: str = '1.0.0'
to_dict()[source]

Return a JSON-compatible manifest dictionary without raw records.

Returns:

Strict schema fields only. Mappings are copied so callers cannot mutate the manifest through the returned representation.

Return type:

dict

Examples

>>> manifest = BenchmarkManifest(
...     "toy", "CC0", "b" * 64, "https://example.test/toy",
...     "reference", "1.0", "Linux", "b" * 64, "method", "NA", "mean",
...     {"estimate": 1.0}, {"estimate": 0.01},
... )
>>> sorted(manifest.to_dict())[0]
'dataset_license'
to_json()[source]

Serialize the manifest canonically for committed evidence artifacts.

Returns:

UTF-8-safe JSON with sorted keys and compact, deterministic spacing.

Return type:

str

Examples

>>> manifest = BenchmarkManifest(
...     "toy", "CC0", "c" * 64, "https://example.test/toy",
...     "reference", "1.0", "Linux", "c" * 64, "method", "NA", "mean",
...     {"estimate": 1.0}, {"estimate": 0.01},
... )
>>> manifest.to_json() == manifest.to_json()
True
sha256()[source]

Return the SHA-256 identity of the canonical manifest JSON.

Returns:

Lowercase, 64-character hexadecimal digest of to_json().

Return type:

str

Examples

>>> manifest = BenchmarkManifest(
...     "toy", "CC0", "d" * 64, "https://example.test/toy",
...     "reference", "1.0", "Linux", "d" * 64, "method", "NA", "mean",
...     {"estimate": 1.0}, {"estimate": 0.01},
... )
>>> len(manifest.sha256())
64
classmethod from_dict(payload)[source]

Create a manifest from an exact-schema JSON-compatible mapping.

Parameters:

payload (mapping of str to Any) – Decoded manifest data. Unknown fields, including fields that could hold raw rows or records, are rejected rather than silently kept.

Returns:

Validated immutable benchmark manifest.

Return type:

BenchmarkManifest

Raises:
  • TypeError – If payload is not a mapping.

  • ValueError – If the mapping keys differ from the published schema.

Examples

>>> manifest = BenchmarkManifest(
...     "toy", "CC0", "e" * 64, "https://example.test/toy",
...     "reference", "1.0", "Linux", "e" * 64, "method", "NA", "mean",
...     {"estimate": 1.0}, {"estimate": 0.01},
... )
>>> BenchmarkManifest.from_dict(manifest.to_dict()) == manifest
True