Data & Provenance

Say you don't know

Most data specifications treat a missing field as a failure and an optional field as a kindness. Both produce the same result: records full of invisible holes. We have changed ours to do the opposite.

By Common MeasurePublished September 2026Updated September 20266 min read

A person filling out a form on a tablet while sitting on a wooden lake dock

Two records that look the same

Here are two assurance records for the same water sample.

In the first, the field for reaction volume is not there.

In the second, it reads not recorded.

To a validator these are nearly identical — one field short of a complete document. To anyone trying to use the result they are entirely different situations.

The first tells you nothing. Was the field forgotten? Judged irrelevant? Withheld? Does the submitter not know, or did they decide you did not need to? You cannot tell, and each possibility carries different consequences for how much weight the result deserves.

The second tells you exactly where you stand. The information does not exist. That is a fact about the world and you can reason about it.

A declared gap is data. An empty field is an absence of data about the absence of data.

Optional fields do not get filled in

The conventional handling is to mark the field recommended rather than required. It sounds reasonable and it performs badly.

An optional field is skipped by default, and skipped without leaving a trace. Nobody records that they chose not to record something.

Over time that produces a corpus in which the most difficult, most informative and most expensive-to-capture fields are systematically the emptiest — and nobody knows by how much, because the skipping is invisible.

The fields that go missing are never random. They are the ones somebody would have had to make an effort to capture.

Mandatory fields fail worse

So make them required. That fails differently, and more expensively.

Consider a lake association holding a sequence file from 2023. They commissioned a survey, a laboratory ran it, they received a report. Three years later they want it to be part of a comparable record.

They cannot know what reaction volume the laboratory used. Nobody told them, they would not have known to ask, and the person who ran it may have left.

Under a conventional mandatory field that record can never conform. Not because the data is bad — because the format demands something its holder cannot supply.

Multiply that across the field. Most environmental DNA ever produced sits in exactly this condition: real data, commissioned in good faith, held by somebody who is not the laboratory.

A specification that excludes most existing data is not rigorous. It is useless.

And the real consequence is worse than exclusion. Faced with a mandatory field and no knowledge, a reasonable person enters a plausible number. Now the record is not merely incomplete. It is wrong, and it looks complete.

The third option

Require the field. Permit the answer to be not recorded.

That is the whole change, and it is now how our specification works. Sixteen fields moved from recommended to required, each accepting an explicit declaration of ignorance as a conformant value.

Conformance requires that you answer. It does not require that you know.

The lake association conforms by stating what it cannot know. The laboratory that does know is now obliged to say. And a reader holding any record can tell, immediately, the difference between a fact and a gap.

What it makes possible

Three things follow. The third we did not anticipate.

It lowers the barrier rather than raising it. Historical data becomes conformant instead of excluded. A 2023 survey with six declared gaps sits in the same corpus as a 2027 survey with none, and the difference is visible rather than hidden.

It removes the incentive to guess. When "I don't know" is an acceptable answer, nobody has to invent a plausible one.

And it makes completeness measurable. Because declared gaps are explicit, they can be counted. Across every record processed, per field, we will know what proportion carry real values and what proportion carry not recorded.

That is a dataset nobody currently has — an empirical picture of what this field actually captures, as against what everyone assumes it captures. It tells us which requirements are realistic and which are aspirational. When the time comes to decide whether reaction volume should demand a real number, that will be a decision made on evidence rather than on anyone's view of what good practice ought to look like.

What we have not settled

Two things, both published as open questions rather than resolved quietly.

Whether the pattern should apply everywhere. We drew the line at fields where the record still describes something useful without them. Coordinates, collection date, marker, reference database version and the record hash still demand real values, because a record without those describes nothing at all. We are not confident the line is in the right place.

Whether not recorded is too blunt. It currently collapses three situations: never captured, not applicable to this substrate or assay, and deliberately withheld for commercial or sovereignty reasons. A reader might reasonably want to tell those apart. Splitting them adds precision and adds burden.

Both are structural decisions about how the specification behaves rather than technical details, which makes them the kind of thing a governance body should own rather than something we should settle alone.

The general point

This is a small change to a young specification and it will not be the last thing we get wrong.

But it points at something broader about how records of measurement should work. The purpose of documentation is not to make every result look complete. It is to let a reader establish what is known, what is not, and how much confidence they are entitled to.

A record that hides its gaps serves the person who produced it. A record that declares them serves the person who has to use it.

Only one of those is evidence.

Published by Common Measure. We do not run samples.

Publishing our specification in the open.

Common Measure launches in 2027. If you run a laboratory, buy environmental data, or set policy that depends on it, we'd like to hear from you before then.

Join the waitlist