Insurance case companion

Before you fill with zero, keep where the zero came from

A missing row and a recorded zero are different inputs. Once both become 0, a later rule cannot distinguish them from that value alone.

Before calculating, ask:

  1. Was there a matching row? Keep row coverage separately from the amount.
  2. Was its value recorded? A present row with a missing amount is not a recorded zero. Check this before aggregation.
  3. What will the zero trigger next? A filled value may later change a label, rate or decision rule.

Keep the original value and an origin column through later processing. A row-presence flag is not a reason for absence.

Try the distinction

Download the code and reproducibility companion. With Python 3.10 or later, run:

python -B check_before_fill.py

This runs a constructed example, not insurance observations. It keeps four rows distinguishable: no amount row, an amount recorded as zero, a positive amount, and a positive sum that includes a zero-valued amount row. It does not fill values or reset counts.

With the two original files listed in the methods, run:

python -B check_before_fill.py --sources PATH_TO_ARFF_FILES

The fixed insurance files contain 653,069 frequency rows without a matching amount row, including 9,116 with a positive claim count. The amount file has no recorded zero amounts. These counts do not explain the missing rows.

Read the case · Methods and full result replay

This is not a rule against filling zero. It is a way to retain the distinction needed to ask whether that fill is justified for the next use. The supplied checker is specific to the pinned files and nonnegative-amount rules; it is not a validated checker for arbitrary datasets.

Return to Cases →