Bench lab calibration of the Lume 1.2 sensor against Colilert (IDEXX defined-substrate, MPN) under controlled dilutions, with comparison to membrane filtration (MF, CFU), for E. coli and total coliform quantification.
Lab Validation
Sensor-to-reference agreement is bounded by reference-method reproducibility, not by Lume hardware. The two EPA-approved methods agree with each other at only R² = 0.52 (membrane filtration reads ~2.2× higher than Colilert). Against that backdrop the Lume tracks whichever method it is trained on (R² = 0.86 vs Colilert, 0.70 vs MF) and flags ≥126 CFU/100 mL exceedances against Colilert at 0.90 balanced accuracy (94% sensitivity, 86% specificity, in-sample), while adding the continuous temporal coverage that grab-sample laboratory methods cannot provide.
How does the Lume compare against the two EPA-approved laboratory methods, Colilert (IDEXX) and membrane filtration (MF), and against itself when trained on each reference? Each comparison is examined three ways: regression (predicted vs. observed), Bland–Altman agreement, and a binary exceedance classifier at ≥126 CFU/100 mL. The three comparisons (1 MF vs. Colilert, 2 Lume vs. Colilert, 3 Lume vs. MF) run left-to-right within each panel below.
Each chart recomputes its own log-log fit in your browser. Gray error bars = 95% CI. Points are green under one consistent rule across all three panels: every value carries a Poisson-type 95% CI (Colilert = published QuantiTray MPN CI; MF = exact Poisson CI on the count, keyed to the reported value as a count from 100 mL), and green means the two CIs overlap. In comparison 1 both sides are real method measurements. In comparisons 2 and 3 the Lume prediction is given the same CI as the reference it is compared against, i.e. we assume Lume’s measurement CI is no worse than the reference method’s. This is deliberately conservative: Lume’s own regression prediction interval is ~7× wide (±0.43 log10) vs the reference CI’s ~1.5×, so it denies Lume the benefit of its true, larger uncertainty rather than letting it inflate agreement. Overlap: 33% / 84% / 63% (comparisons 1–3). Axes are log-log, square 1:1.
Per-comparison agreement on the log10 scale: each point is log10(compared) − log10(reference) against the pair mean. Solid line = mean bias, shaded band = 95% limits of agreement (bias ± 1.96·SD). A band centered on 0 with tight width = good agreement.
For each comparison a binary logistic regression is fit, P(reference ≥ 126) ~ log10(compared value), and the confusion matrix is read at the Youden-optimal probability cut. Metrics are in-sample. Single-event caveat: every ≥126 exceedance is a Boulder Creek field-stream sample from one high-flow episode (32, 17, and 32 positives for comparisons 1–3), so sensitivity is optimistic for an unseen new event; the bench dilution series contributes only negatives.
The same logistic classifier, but the cut is each reference’s own median value (shown in each panel) instead of the 126 threshold. This splits every session’s samples into balanced halves, so it is not dominated by the single high-flow event and gives a more robust read of discrimination across the ordinary concentration range. Metrics are in-sample.
The two EPA-approved methods are paired across the dedicated Colilert–MF replicate study (Colilert ×2–3 replicates, MF ×3; 8 zero-valued pairs dropped on the log scale). They agree at R² = 0.52 (slope 0.79) with a +0.34 log10 bias: MF reads ~2.2× higher than Colilert. Bland–Altman limits of agreement span [−0.66, +1.34] (paired lab samples differ by up to ~22×), and by ISO 17994:2014 the two methods are not equivalent (mean relative difference +79% vs a ±20% limit). The ≥126 classifier reaches balanced accuracy 0.86 (sens 0.94, spec 0.78, κ 0.56). This inter-method disagreement is the ceiling any sensor can reach against either reference.
The Colilert-trained Lume regression is evaluated against Colilert across bench (n = 176) and field-stream (n = 33) samples. It achieves R² = 0.861, zero bias, and tight limits of agreement [−0.42, +0.42] (within ~2.6× of the reference); CIs overlap on 84% of points. The ≥126 classifier reaches balanced accuracy 0.90 (sens 0.94, spec 0.86, κ 0.47). Against its training reference the Lume performs at or above the agreement the two EPA methods reach with each other.
Trained instead on membrane filtration (same per-sensor form, bench + field-stream), the Lume reaches R² = 0.702, zero bias, and limits of agreement [−0.75, +0.75]; CIs overlap on 63% of points. The ≥126 classifier reaches balanced accuracy 0.98 (sens 0.97, spec 0.99, κ 0.94), though every exceedance in this set is a single field-stream episode, so classifier sensitivity is optimistic for an unseen event. The lower R² against MF than against Colilert tracks MF’s own noisier agreement (comparison 1), not a sensor limitation.
Literature context (comparison 1). MF–Colilert agreement for E. coli is matrix-dependent: many freshwater studies report near-parity, while marine samples can diverge by 1–3 orders of magnitude. Reported freshwater correlations are high (e.g. Pearson r ≈ 0.96, slope ≈ 1 in a long-term river study), but are typically computed on raw concentrations, which are dominated by the few highest samples; on the log scale used here, method agreement is genuinely looser. Our R² = 0.52 and ~2× MF-high bias therefore sit at the modest-agreement end of published values, not outside the documented range, and better-replicated references could agree more tightly, so this ceiling is conservative.
Bench calibration sessions, January 2026, three sensors (50030 / 50031 / 50032); fluorescence signal (mon2_val) captured at the operating point led_power = 1024, sipm_bias = 3040. Each CSV below is exactly the points plotted above: the observed reference value and its 95% CI, the compared value (Lume prediction or MF) and its CI, and the overlap flag.