Insurance case companion

Sources, boundaries and recomputation

This companion checks how fixed file fields become training targets. It does not establish the reasons for missing amount rows, real accident counts, model performance, or whether the processing is suitable for a pricing or underwriting decision.

Inputs

Obtain these original files from OpenML, keeping the filenames:

Their expected byte counts, SHA-256 hashes and literal headers are in INPUTS.json. A changed download is refused, not silently accepted as another version. No raw data files are redistributed in this companion; consult their original source and terms.

The transformation follows scikit-learn 1.7.2 at commit 25dee604bae18205b01548348388baf7a1cdfe0e. This is a separately authored reconstruction, not execution of that example or its trained models. The case page links the book discussion and the alternative R count rule; those sources are context, not alternative runs reproduced here.

What is compared

Keys are matched as exact positive integers, including integer-valued scientific notation. This does not establish the identity of real policies, persons or accidents. File versions and row ordinals are the time coordinates; no event-time alignment is claimed.

All used counts must be nonnegative integers, exposures positive and amounts finite and nonnegative. Missing values in present amount rows, missing keys, duplicate frequency keys and invalid numeric domains stop the computation. This prevents an all-missing group or cancellation of positive and negative amounts from being called a recorded zero under this method. Other legitimate data conventions need another contract, not forced acceptance here.

Amount rows are summed by key, then left-joined to the unique frequency keys. A missing match remains None and carries its origin. Six states are compared:

  1. Join while keeping missing amounts.
  2. Fill missing amounts with zero.
  3. Cap counts at 4.
  4. Cap exposure at 1.
  5. Cap aggregate amount at 200,000.
  6. Where amount is zero and count is positive, reset count to zero.

Frequency is count/exposure; pure premium is amount/exposure; average claim amount is amount/max(count, 1). Undefined-to-defined changes are counted separately from numeric changes. Sums are unweighted field sums, not population estimates or monetary valuations. Arithmetic uses exact rational values from decimal tokens, not the upstream binary floating-point runtime.

Run locally

Extract the companion, open a terminal in that folder, and use Python 3.10 or later. No third-party Python packages, model downloads or network calls are required by the scripts.

python -B test_companion.py
python -B check_before_fill.py
python -B check_before_fill.py --sources PATH_TO_ARFF_FILES
python -B recompute.py --sources PATH_TO_ARFF_FILES

The first check-before-fill command is synthetic; the second uses only the pinned real files. Both preserve coverage separately from values. They do not change labels or input files.

Full replay compares the entire reconstructed numerical result with EXPECTED_RESULT.json, including six stage fingerprints, changes and cohort counts. The engine also generates affected-row memberships in memory; these are not exported by the command. Inspect engine.analyse to recover them when needed. Synthetic controls test row reordering, disabling the reset, distinguishing absence from zero, and reassigning a key while preserving the amount total. They are arithmetic controls, not causal evidence.

A passing replay means that the declared computation matches these inputs and expected results. It does not certify source semantics or decision suitability.

Credit, reuse and corrections

The case text and companion adaptation are presented by the Claim Qualification project. Underlying data and example code retain their original attribution; no affiliation with or endorsement by OpenML, scikit-learn, the book authors or the insurer is implied. The known count/amount inconsistency is not claimed as a new discovery.

No general open-reuse licence has been selected for the project-authored materials in this edition. Website availability is not a claim to authorship of the source data. No DOI has been assigned to this edition.

For a correction, identify the file versions, source lines or computation that change the result. Reports of a different data version or use can reopen the question without erasing this edition. Read the case or use the short check card.

Return to Cases →