Evidence & Decisions

Reliability was never the problem

In 2020, four scientists concluded that environmental DNA already meets the legal standard for evidence in most American courts. Six years later, managers still hesitate to act on a detection. The gap between those two facts is the reason we exist.

By Common MeasurePublished September 2026Updated September 20268 min read

Institutional records storage

An unusual finding, buried in an unglamorous sentence

In 2020, Adam Sepulveda and colleagues at the US Geological Survey published a paper in Trends in Ecology & Evolution asking whether environmental DNA methods were ready for aquatic invasive species management. It has since been cited more than 160 times.

Their answer was not what most people expected.

They found that eDNA methods meet legal standards for being admissible as evidence in most courts, and concluded that method reliability is not the problem. Under the Daubert standard — the test American federal courts apply to scientific evidence — eDNA was, in their assessment, already sufficiently mature and reliable.

That is a remarkable thing to establish about a technique many practitioners still describe as emerging.

The science was not the obstacle. It had not been for some time.

So what was stopping anyone?

The same paper answers that too. Invasive species managers struggle to use eDNA results because a detection might not indicate that a species is present, and because the cost of acting can be very high. The authors were direct about where the work needed to happen: the interface between results and management needs attention, because there are few tools for integrating uncertainty into decision-making.

Read that again as a description of a market. The instrument works. The chemistry works. The statistics work. What does not exist is the layer between a laboratory's output and a decision that costs money.

That layer is not a scientific problem. It is an infrastructure problem, and nobody was building it.

The problem has a specific shape

A year later, a group led by Bettina Thalinger published a validation scale for eDNA assays in Environmental DNA. Their case study was signal crayfish, an invasive species with assays developed independently in several countries.

Their finding is worth quoting carefully. The assays had been developed using a variety of strategies for sampling, capture, extraction and qPCR across different genetic markers, applied to genetically diverse populations — and because of that methodological variability, direct comparisons between results obtained from these assays are impossible.

Not difficult. Not imprecise. Impossible.

They described the situation facing an end user as a minefield, and argued that guidelines determining assay suitability were needed so that management decisions could rest on analyses with quantifiable uncertainties and clear inference limits.

This is the same finding as Sepulveda's, expressed at the level of a single species. Each assay was defensible. Together they were incomparable. A manager holding two positive results from two providers had no principled way to treat them as the same kind of evidence.

Even the word "positive" was not agreed

In 2021, John Darling, Christopher Jerde and Adam Sepulveda published a short paper in Environmental DNA with an unusually blunt title: What do you mean by false positive?

Their argument was that misunderstandings about the term itself were a significant hurdle to adoption. A false-positive test — a laboratory result produced in error — is not the same thing as a false-positive inference, which is a conclusion about a site that direct observation later contradicts. The two get conflated constantly, by scientists and managers alike, and the confusion is expensive.

Papers published since have adopted their vocabulary as a matter of course. It became the field's shared language because the field did not have one.

That is worth pausing on. A community of competent scientists, working on the same problem for a decade, needed a paper to agree what a word meant.

And the design decisions are unresolved too

Two studies published this year make the same point from opposite directions.

A federal team sampling a single river site every fifteen minutes for nearly three weeks found that concentration at one location, from one laboratory, using one method, could plausibly land anywhere between roughly a quarter and four times the expected value. They also found that treating samples collected fifteen minutes apart as simultaneous replicates biased their variance estimates badly. Their conclusion was that there is unlikely to be a universal definition of near-simultaneous sampling.

A commercial laboratory, working on a marine transect, compared ten laboratory replicates from one field sample against one replicate from each of ten field samples. The richness was statistically indistinguishable. Their conclusion was that field replication remains essential in spatially heterogeneous environments, but that the spatial scale at which it becomes important is unknown.

Different questions, different systems, different organizations. Both arrived at the same place: the right guidance depends on the site, and nobody can specify it universally.

Which is not an argument against a specification. It is an argument for a particular kind of one. You cannot standardize a threshold that varies by system. You can standardize the disclosure — require that whoever produced the result records what they actually did, in a form somebody else can read — and let the person interpreting it decide whether two results belong side by side.

Fifteen years, one conclusion

The applied literature on eDNA is unusually consistent. In 2011, Darling and Mahon were writing about the move from molecules to management. In 2020 Sepulveda's group concluded that reliability was not the constraint. In 2021 Thalinger's group showed that assays for a single species could not be compared. In 2022, Abigail Keller and colleagues built a joint model in Ecological Applications precisely to quantify what adding eDNA to an existing monitoring program actually buys, because nobody could otherwise say. A 2026 paper in BioScience is still explaining false negatives to decision-makers using a pizza analogy.

Across fifteen years, nobody in this literature argues for better sequencing.

They argue, in different words each time, for a way to decide what a result means — consistently, across providers, in a form someone can act on.

What we took from it

Three things, and they became the design of our company.

First, do not try to improve the science. It works, and the people doing it are better at it than we would be. Recovering usable genetic material from a soil core, designing a primer set that amplifies what you intend, diagnosing inhibition, keeping a high-throughput laboratory clean — that is difficult, skilled work and it is not ours.

Second, build the layer the literature keeps describing. A persistent identity for a sequence that survives taxonomic revision. A record of method, versions and controls that travels with a result. A way to establish, with evidence rather than assertion, how far apart two laboratories actually are.

Third, and least obviously — it cannot be built by a laboratory. Thalinger's crayfish assays were not incomparable because anyone did poor work. They were incomparable because each was developed independently, for good local reasons, with no shared frame. Building that frame requires laboratories to send raw data to whoever operates it, and no laboratory can reasonably be asked to send that to a competitor bidding for the same contracts. That is not a failure of goodwill. It is a structural fact, and it is why the operator has to be someone who does not run samples.

The uncomfortable part

If reliability was settled in 2020, then every year since has produced data that could have supported decisions and largely did not — because the result could not be compared to anything, and the manager could not defend acting on it.

That is not a story about a young technology finding its feet. It is a story about missing infrastructure, and missing infrastructure is a solvable problem.

We would rather be six years late than not build it.

References

  1. Darling, J.A. & Mahon, A.R. (2011). From molecules to management: adopting DNA-based methods for monitoring biological invasions in aquatic environments. Environmental Research 111(7), 978–988. doi:10.1016/j.envres.2011.02.001 (opens in a new tab)
  2. Sepulveda, A.J., Nelson, N.M., Jerde, C.L. & Luikart, G. (2020). Are environmental DNA methods ready for aquatic invasive species management? Trends in Ecology & Evolution 35(8), 668–678. doi:10.1016/j.tree.2020.03.011 (opens in a new tab)
  3. Darling, J.A., Jerde, C.L. & Sepulveda, A.J. (2021). What do you mean by false positive? Environmental DNA 3, 879–883. doi:10.1002/edn3.194 (opens in a new tab)
  4. Thalinger, B., Deiner, K., Harper, L.R., Rees, H.C., Blackman, R.C., Sint, D., Traugott, M., Goldberg, C.S. & Bruce, K. (2021). A validation scale to determine the readiness of environmental DNA assays for routine species monitoring. Environmental DNA 3. doi:10.1002/edn3.189 (opens in a new tab)
  5. Keller, A.G. et al. (2022). Tracking an invasion front with environmental DNA. Ecological Applications 32, e2561. doi:10.1002/eap.2561 (opens in a new tab)
  6. Augustine, B.C., Hutchins, P.R., George, S.D., Huddleston Adrianza, C.C., Darling, M.J., Sadekoski, T.R. & Sepulveda, A.J. (2026). Temporal replication and sampling efficiency in high-frequency autonomous and manual eDNA sampling. Environmental DNA. doi:10.1002/edn3.70347 (opens in a new tab)

Quotations are drawn from the published abstracts and article text of the works cited. Where we have characterized an argument rather than quoted it, the citation points to the source so you can check us.

Published by Common Measure. We do not run samples.

Publishing our specification in the open.

Common Measure launches in 2027. If you run a laboratory, buy environmental data, or set policy that depends on it, we'd like to hear from you before then.

Join the waitlist