LFRResearch note · 0.1.4

From BMS to Dataset: What Remains Visible?

Publish the meaning of the timestamps alongside the readings. Say which values may have been carried forward. For questions about control, keep requests and limits alongside recorded power, where available.

Pingxin Wang / LFR · Research note · Edition 0.1.4 · 1 October 2026

Start with a change

Imagine the current falling during a battery discharge. Did the power request fall too? Did a limit change? Or did the recorder leave us an older reading? These are different questions, and each needs a different connection between records.

Begin with one practical step: choose the comparison you want to make, then check what each record means before putting the records together. If a needed connection is unknown, write down that particular gap. Other comparisons may still be useful.

This matters when comparing operating periods, training a model or preparing a dataset for someone else. Otherwise, a difference in how readings were taken can enter the analysis as a difference in the battery.

The record has a history too

Follow a reading through five places. This is a question map, not an installation drawing:

  1. 01AcquisitionWhat was sensed, where, and when?
  2. 02Estimation & controlWhich values were estimated? Which actions were issued?
  3. 03CommunicationWhat was updated, delayed or retained?
  4. 04StorageWhat was archived, combined or omitted?
  5. 05ReleaseWhat was selected, transformed and documented?

Control also feeds back into operation: a reported state can affect what the system permits next. The map follows records towards the analyst, while that feedback loop continues at the installation.

When information is missing from a release, ask where it might still exist: in the recorder, the operator's archive or another document. A missing published field tells us about the release, not automatically about what the original BMS knew or used.

Three examples from published documentation

Documented identifies what the named source says. Interpretation explains the comparison it helps us make. The small illustrations use invented values; they are not operating data from these installations.

1. Weihai: where does each rack's count begin?

Documented. The release describes nine racks, operation start/end fields, and separate time/value arrays for current, voltage, power and SOC. The listed fields do not supply a request–limit–execution event chain. Dataset description, §§5–7.

Interpretation. Before comparing two racks, check what their clocks and operation boundaries mean. Separate arrays alone do not prove that acquisition was asynchronous; a shared row would not prove that it was simultaneous either.

If you calculate “charge passed so far” starting at each rack's first available current reading, each count has its own starting point. Record that starting point and the stretches with readings used in the calculation. Then two labels saying “100 Ah” can be read with their different origins still visible.

Illustration — invented times and counts, not Weihai data. These origins are calculation choices, not observed physical beginnings:

Count Starts at Label later reached
Rack A First available reading, 10:00 100 Ah since 10:00
Rack B First available reading, 10:05 100 Ah since 10:05

Keep with the result: timestamp definitions, counting origins, gaps and the rule used to select or pair readings. Those records let another reader reconstruct the comparison.

2. OpenCEM: a new row can contain an old reading

Documented. The acquisition README describes polling a fixed set of register blocks plus a rotating group. Unpolled registers keep their previous values; an exception can also leave a previous value unchanged. The two inverter snapshots are written with a shared timestamp. Acquisition README, in the deposited support archive.

Illustration — invented snapshots and values, not OpenCEM data. The last column is supplied for this explanation; it is not a claim that the release includes a per-field update log. The row spacing does not represent the recorder's polling interval.

Snapshot Value written What happened in this invented example
t0 50.0 A new reading was received.
t1 50.0 No new reading; the previous value was kept.
t2 50.0 A new reading was received, with the same value.

Interpretation. The numbers look identical. Their histories differ. A row timestamp alone cannot tell us whether a repeated number was read again or carried forward. Here, reading age means time since that field was last acquired, not time since the row was saved.

OpenCEM also distinguishes when a context record applies from when it was entered. Keeping both times helps a later reader ask what information was available at a particular moment. Release README.

Keep with the result: field-update information, if recorded, plus polling and error rules. Otherwise leave the reading's age unknown.

3. M5BAT: put the setpoint beside recorded power

Documented. The report defines unit-level active-power setpoint and recorded active-power fields, plus inverter modes. At station level, SPA_ask_P means “Request for active power for setpoint adjustment”; SPA_exec_P means “Active power for setpoint adjustment”. Both are in kW. The table's time coordinate is UTC. Field tables, pp. 6–7; dataset record.

Page 14 describes SPA's SOC-management purpose and 15-minute intraday continuous-market context. That context does not turn the two power fields into market-order or trade-execution records. Report, p. 14.

Invented example: a setpoint changes at one labelled point; recorded power changes at a later point. Two separate traces let a reader notice this difference.

Illustration — not M5BAT data. Two invented traces in arbitrary units show a possible difference in timing. The shape is not a measured delay or a controller model.

Interpretation. With both fields, we can ask whether the recorded setpoint and power changed together, then inspect the reported mode. Explaining a difference still needs timing definitions and relevant limits or events; the two lines alone do not identify its cause.

Keep with the result: requests, applicable limits, modes, command acknowledgements and execution feedback, with their times and definitions. For a controller comparison, also identify which request is upstream of the controller being replaced: a unit setpoint may already be another controller's output.

Keep three distinctions visible

Distinction Possible reading mistake What to retain
Zero, missing, or carried forward Reading a gap as zero current, or an old value as a new measurement. Value codes, error rules and update information.
An estimate and its inputs Treating agreement as independent confirmation when both use the same measurements. Which system produced the estimate, its reference or model, and recorded corrections.
Battery response and how it was measured Reading a calibration or logging change as a battery change. Sensor, rack identity, configuration and logging histories, with effective dates.

A data-chain card, starting from one question

Use the card for one dataset version and one intended comparison. Cite what answers the question, say what remains unknown, and name a useful next record.

Your question First records to check
Did two racks change together? Rack identities, reading times, reading ages and pairing rule.
Did current fall after a request or a restriction? Request, limits and execution records on defined clocks.
Why do similar reported SOC values accompany different paths? How SOC was produced and the preceding current/voltage records.
Has a familiar response changed? Operating conditions and changes in sensors, recording or control settings.

A filled fragment: does this row mean every field is fresh?

Here, fresh means newly acquired, rather than copied from an earlier reading.

Question. Does an OpenCEM row timestamp mean every field was just updated?

Answer — documented behavior. No such guarantee follows from the documented polling rule: some register blocks rotate between polls, and unpolled fields keep their previous values.

Basis. OpenCEM v1.0.0, Electrical snapshot acquisition, polling-cycle and persistent-output-dictionary paragraphs.

What this supports. Check for a field update before treating a repeated value as a new measurement.

Still not confirmed. Which fields were updated in a particular row, and the age of each retained reading. This fragment reads the documentation; it does not reconstruct a row.

Next step. Choose one field and one snapshot. Look for its update record or polling log. If unavailable, record its age as unknown.

Use the same question inside a monitoring system

Imagine comparing today's rack voltage with last month's, after the sensor was recalibrated. Keep the recalibration date beside the readings and flag the voltage comparison for another look. Other comparisons may still be usable. This is a hypothetical example of a record to preserve, not an event found in the three datasets.

A useful system could remember both the earlier response and how it was measured. To test whether that helps, choose one judgment—such as whether an apparent voltage shift remains after accounting for a documented recalibration—and compare the judgment with and without that information. That experiment remains to be done here.

You can start small: a publisher can explain one timestamp; an analyst can record the origin of one calculation; a monitoring developer can preserve one configuration change alongside the history it affects.

Keep the readings, and enough of their history for the next person to understand the comparison.

Scope, sources and what remains open

What this note provides. Three selected documentation examples and a proposed record card. The examples illustrate questions to ask; they are not a survey of how common a practice is.

What was checked. The cited descriptions and retained documents. No new telemetry analysis, causal identification or control validation was performed. The card is a proposal, not a standard or an operating instruction.

What remains open. Actual acquisition behaviour, the cause of any operating change, which information the original BMS used, and whether preserving more information improves a decision. This note makes no health, safety or operational-benefit assessment.

The source page records inspection limits and exclusions, including unverified standards comparisons and near-zero-current findings. A useful contribution can be very specific: identify one timestamp definition, update log or configuration record that changes an answer on the card.

See sources.html for document locators and inspection limits, and selected-sources.json for the reference register. An editable copy is provided as research-note.md.

Back to top ↑

Scope: Documentary examples and a proposed documentation template. No new telemetry analysis, causal identification or control validation.

Research note; not independent scientific validation.

Original LFR research-note and data-chain-card content: CC BY 4.0. Third-party material, fonts, stylesheets and code are excluded; their own terms apply. Reuse scope and attribution.

Sources retain their original authorship and licence terms. Attribution and inspection limits. LFR · edition 0.1.4 · 1 October 2026.