Insurance records · worked example

How missing amounts become zero claim counts in a training target

freMTPL2 (OpenML 41214 / 41215, version 1) · scikit-learn 1.7.2 example Tweedie regression on insurance claims · Claim Qualification project

We traced a known data inconsistency along the example code into the training target, and quantified what it changes and what it does not.

Following the example's processing order, 9,116 records that have a positive claim count but no matching row in the amount table have their claim count set to 0: 26.8% of all records with a positive count. This step changes their frequency target; their pure-premium target was already 0 at the earlier fill step.

This is a reconstruction of the example's rules on the two fixed files: we did not run the original example, and we trained no model.

How the chain connects

The steps below show only the main path by which an absence becomes zero; between the zero-fill and the count reset there are three caps (lines 226–228).

  1. The amount table, freMTPL2sev, is summed by IDpol and left-joined onto the count table, freMTPL2freq (lines 77, 79). Of the count table's 678,013 rows, 653,069 have no amount row.
  2. Missing amounts are filled with 0 (line 80). All 26,639 rows of the amount table are positive and none is recorded as 0, so every zero amount after the join comes from this fill.
  3. Where the amount is 0 and the count is at least 1, the count is set to 0 (line 232). The example's comment gives the reason: the severity model needs strictly positive amounts, and this keeps frequency and severity more consistent (lines 229–231).
  4. The processed counts, amounts and exposures then form the frequency, average-claim-amount and pure-premium targets.

The key lies between steps 2 and 3. Step 2 does not pass on where a zero came from, so step 3 can only act on the value. The original files still exist, so the path can be traced back; but in the processed amount column alone, "no amount row" and "an amount of zero" can no longer be told apart. What is worth seeing is an operation with an explicitly stated rationale that depends on an input whose meaning was changed one step earlier.

What changed and what did not

Before vs after all processing stepsBeforeAfter
Rows with a positive count34,06024,944
Sum of the ClaimNb field (not a population estimate)36,10226,406

Of the field-sum reduction of 9,696, 9,650 comes from step 3 and 46 from the count cap of 4. The frequency target changes in 9,175 rows: 9,124 decrease, and 51 increase because of the exposure cap of 1. Step 3 does not change the pure premium or average claim amount of these 9,116 rows, because step 2 had already filled their amounts with 0. In other words, their pure premium of 0 comes from the fill, not from the reset.

Concentrated in one ID range

  • Of the 9,116 reset rows, 8,851 (97.1%) have IDpol ≤ 24,500.
  • Within that ID range, ClaimNb in the file is positive for every one of its 9,139 rows, and 96.8% of them have no matching row in the amount table. Among positive-count rows outside the range, the share is 1.1%.
  • After processing, the share of rows with a positive count is about 3.2% inside the range and 3.7% outside. That ratio alone would not reveal the earlier large difference. Similar ratios do not show that the data are sound, nor that the processing is correct.
  • An ID range does not stand for time, a real population or a cause.

What is already known

  • The example's authors state their reason for the reset. This analysis describes the transformation; it does not establish that the reset is erroneous.
  • Wüthrich and Merz, Statistical Foundations of Actuarial Learning and its Applications (2023), Appendix B, flag a correspondence problem between counts and amounts for IDpol ≤ 24,500 and recount claims from the amount table instead. The book cites private communication as part of its basis, which we have not obtained. That the book discusses this does not mean every user of the example knows it.
  • Around freMTPL2, three public materials use different rules to construct the claim count: the R example linked to the original study uses only ClaimNb from the count table; the book above recounts from the amount table; the scikit-learn example keeps the original count and then resets it by amount. Their historical inputs have not been matched one to one with these OpenML files, so what is compared here is rules, not three runs on the same input.

What is still unknown

  • Why these 9,116 records have no amount row.
  • How ClaimNb and ClaimAmount are each formed in this version of the files: their definitions, inclusion conditions and valuation dates.
  • The CASdatasets repository has a June 2022 commit labelled as removing inconsistencies, which changed these two files. Which records it changed, and how it relates to the OpenML files, has not been checked.
  • A further 6 amount keys (195 rows) have no match in the count table. The left join does not bring them in, and the reason is likewise unknown.
  • Whether this example, or similar processing, has been used in any real pricing or underwriting decision.

The question we would like to ask

In this version of freMTPL2, by what definitions, inclusion conditions and valuation dates were ClaimNb and ClaimAmount formed? Is there a data dictionary, export note or maintainer record that would support setting these positive-count records without a matching amount row to zero?

Industry experience can suggest candidate explanations. But explaining these 9,116 specific records needs source rules that apply to these rows.

How to check

The core counts can be checked with the two public files alone, by exact matching on IDpol across OpenML 41214 and 41215 (version 1).

For checking the calculation, download the reproducibility companion and follow the methods. It contains the full expected numerical result, an offline replay path and synthetic controls. Obtain the two original data files separately from their public source; they are not redistributed here. The replay reconstructs exact rational values from decimal tokens; it is not a bit-for-bit reproduction of the example's original floating-point run. Earlier working notes and review correspondence are not included.

For a short check before filling missing amounts, use Before you fill with zero.

Return to Cases →