Virridy Home | Lume — Water Quality Sensing Water for Carbon

Daily-Calibration Sufficiency Test

The question was whether a daily wet two-point calibration, a blank cap and an FDOM puck both read in DI, can hold fixed TLF standards constant day over day. It cannot, and the reason turned out not to be the calibration aging. Seven ladders were run on 8 TLF units between 22 and 25 September, plus a dilution ladder from 6 August that predates any handling of these sensors. What came out is a different procedure, given first below, and a decomposition of the drift into a handling term that moves the zero and a slower term that moves the gain. Every reading is taken wet, because air readings are specious through total internal reflection. All times are Mountain.

Drift and rhodamine challenge

5 TLF units and 9 Chl-a units went into 25 L of metered DI water together at 11:10 on 1 October and were handled and dosed as one bath. The test closed at 16:15. Every logged event is a red dashed line on each chart, named in the legend. The record below stops at the close.

Findings: closed 16:15 MT, 1 October

Sequence (MT): 11:10 all 14 units into DI · 13:52 all units shaken in the bath · 15:03 all units twisted under water, caps on · 15:15 rhodamine WT to 5 ppb · 15:27 to 2 mg/L · 15:46 to 20 mg/L · 15:54 twisted again · 16:15 closed. The 20 and 100 ppb rungs of the draft ladder were skipped at 15:30: the question for the Chl-a batch was whether it responds at all, so the run went straight to the two doses that reference instruments had already confirmed.

Chl-a: no response at any dose. All nine reporting units (500184 never reported) sweep LED 2048 alone over 34–48.5 V. Absolute mon2 at the top bias step sat at 173–179 counts on every unit in every window: untouched DI, after the shake, after the twist, at 5 ppb, at 2 mg/L (18–19 sweeps per unit) and at 20 mg/L (28–29 sweeps). No unit moved more than one count from its own pre-dose level, and the span across the bias arc stayed at 3–9 counts with no trend. 2 mg/L is the dose at which the AquaTroll Chl-a channel read 8.5–11 RFU on 11 September; 20 mg/L is the top of the 2 September screen. The handling steps also produced no step on these units, which is what a signal chain registering no light from the water would show.

TLF: both handling steps in DI raised S-TLF on all five units, and the raised level held. SiPM temperature was flat at about 16.5 °C through every event, so the quench correction is not what moved. Median of the 6 min before each event against 2–8 min after, S-TLF at 20 °C:

UnitDI before shakeShake 13:52Twist 15:03Twist 15:54 at 20 mg/L
5002305.6+0.4 (+7%)+1.5 (+24%)0 after the dose settled
50010340.1+0.2 (+0.4%)+1.9 (+5%)confounded, still falling from the dose
5001234.8+0.5 (+10%)+1.3 (+24%)0.0
50012833.2+2.5 (+7%)+1.1 (+3%)confounded
5002274.6+1.0 (+21%)+0.7 (+14%)confounded

Four of five units stepped within one minute of the shake (500128 took 4 min; 500103, the brightest, barely moved) and then held flat for the 71 min to the twist. The twist stepped all five over 3–7 min and held through the 5 ppb rung. Every step on both events clears that unit's minute-to-minute scatter in the 6 min before it by 3× or more (most by 30–185×); the smallest, 500103's shake at +0.17, is 4× its scatter. Whether the twist did more than the shake depends on the unit: three stepped more on the twist (500230, 500103, 500123, by 3–11×), two stepped more on the shake (500128, 500227). What does not depend on the unit is the sum. On the three low-background units the two steps add to +1.85, +1.81 and +1.62, the same total within 15% although the shake took anywhere from a fifth to three fifths of it, and that total matches the September air-round-trip step of +1.88, the largest handling release in the record. At the second twist, the two units that had already settled after the 20 mg/L dose (500230, 500123) show no step at all. Against this record's tryptophan gain of roughly 0.10–0.16 slope per ppb, the shake is worth about 2–25 and the twist about 5–19 ppb-equivalent depending on the unit. The DI level read after both handling steps sat 10–40% above the untouched morning level on the three low-background units and 5–11% above on the two bright ones.

TLF and rhodamine. 5 ppb moved nothing. 2 mg/L split the batch: the three low-background units (500230, 500123, 500227) rose 15–30% over 15 min while the two bright units fell 14% (500103) and 27% (500128), onset 2–3 min after the dose. 20 mg/L drove all five down by 42–79% within 8 min. The sign reversal on the bright units and the collapse at 20 mg/L are the pattern of absorption by the dye, not a fluorescence response to it.

What today adds to the handling-step question

The open question in this record is what releases the extra light when a unit is handled. The September candidates still standing were elastomer stress relaxation at the cap joint, the window mount shifting, and a pinned capillary film at the contact line; SiPM saturation, temperature, photobleaching, trapped air and cap orientation had been eliminated. Today changes the picture in two places and confirms it in two more.

  1. The step is optical, not electrical. Nine Chl-a units with the same board, SiPM and bias sweep sat in the same bath and were shaken and twisted at the same instants. Their signal chain registers no light from the water, and they showed no step at all: every sweep within one count of the pedestal through both handling events. Whatever the handling releases, it reaches the detector as light through the optical path. Connector flex, bias wobble or any other board-level disturbance common to both builds is ruled out independently of the LED-proportionality argument made in September.
  2. The reservoir looks fixed, and the shake and the twist split it. September found that repeating a removal an hour later gave nothing and read that as a reservoir that needs 1–2 h to rebuild. Today a twist 71 min after a shake stepped every unit, so the shake had not emptied it. But on the three low-background units the shake and the twist together released +1.6 to +1.9 whatever share the shake took, and that total equals the largest single handling step in the September record (air round trip, +1.88). The simplest reading is one reservoir of fixed size per unit: a shake releases a variable part of it, a twist releases most of what is left, and after both there is nothing until it rebuilds. That is the September ordering by completeness of exchange at the window (shake < quarter turn < cap off/on < air round trip) with the sizes made additive. Three units is thin evidence for a constant sum; the two bright units gave +2.1 and +3.5, which cannot be compared on the same scale because their backgrounds are 7–8× higher. Of the three surviving candidates this does not pick one out: a cap-joint stress model, a shifting mount and a capillary film can each hold a finite amount to release.
  3. The raised level holds. No relaxation over the 71 min between the shake and the twist, nor over the 12 min to the first dose, on any unit. Same as September.
  4. The second twist does not discriminate. At 15:54, 51 min after an identical twist and 8 min after the 20 mg/L dose, the two units that had already settled showed no step. That is what the used-up rule predicts for a repeated action, but the dye had also cut their signal by 42–45%, so an attenuated step and no step cannot be told apart here.

One observation is noted without a conclusion. After the 2 mg/L dose the three low-background units ramped for about 15 min before levelling, while the two bright units completed their fall in about 4 min. If the slow ramp is water reaching the window through the cap rather than mixing in the tank, it is a direct measure of how fast the cavity exchanges; the unit positions in the bath were not logged, so tank mixing cannot be excluded.

For the procedure, today's numbers say the same thing September's did: a DI zero read after handling sits 5–24% above the untouched level and does not come back within the hour, so the zero and the reading have to share one handling state. The rule “read DI, then the sample, nothing touched in between” is what that requires. Nothing from today changes the conclusion that the step is additive light released by handling; what it adds is that the release is optical and comes from a reservoir of roughly fixed size per unit. The discriminating tests remain the ones already set out on this page under the relaxation record: press the cap without turning, relaxation rate at 10 versus 25 °C, and a dial indicator on the cap and window over 10 h.

Loading…

TLF: S-TLF at 20 °C

One point per sweep. S-TLF is the deployed fitTlfSlope, normalized to 20 °C by the deployed tlfSlopeAt20C using each unit's own sipm_temperature (nearest row within 90 s). Hover shows the raw slope and the temperature. ToF is plotted underneath it and is not used as a correction.

TLF: SiPM temperature

TLF: ToF signal per SPAD

Chl-a: sweep span at the highest LED

The production slope gate returns nothing on this channel in clean water, so each sweep is shown as its span: the highest mon2_val below 3250 at the sweep's top LED, minus that LED's pedestal (median of its cells below bias 2850, the same pedestal rule fitTlfSlope uses). Filled markers are sweeps run at LED 2048 alone, which is the configuration set for this test. Open markers are sweeps that ran any other LED set. Hover shows which LEDs ran.

Per unit: last sweep and LED set
Event log

The procedure

This is what the test arrived at. It is scored against every ladder in the record further down, alongside the six alternatives that were tried and rejected.

The two drift terms call for different treatment. The offset is a protocol problem and is fixed by never opening the cap between the zero and the reading. The gain is not, so it is measured once per sensor and stored, and refreshed on a cycle count.

Commissioning, once per sensor

  1. Put the unit in a vessel of DI and leave it there. It does not come out and the cap does not come off for the rest of the procedure.
  2. Settle, read DI. Spike the same vessel to 5 ppb, settle, read. Spike again to 50 ppb, settle, read. Record the vessel volume, since the rung now comes from the spike arithmetic.
  3. Gain is (S-TLF at 50 ppb − S-TLF at DI) / 50, on that unit's own quench coefficient against sipm_temperature. Repeat the ladder two or three times and store the median.
  4. Store the unit's cumulative sample-cycle count alongside the gain. This is the odometer reading the gain is anchored to.

Every run

  1. Read DI, then read the sample, with nothing touched in between: no cap opened and the unit not lifted, since both produce a step and lifting it produces the larger one.
  2. Age the stored gain to the unit's current odometer reading: gain = gain0 × ek(n − n0), where n is the cycle count now and n0 the count at commissioning.
  3. ppb = (S-TLF sample − S-TLF DI) / gain.

Every term is temperature corrected

Each reading is a full-sweep S-TLF normalized to 20 °C as exp(ln(slope) − ρu(T − 20)), with ρ fitted per sensor from that unit's own ramp and T taken from sipm_temperature, the thermometer next to the water, matched to the sweep within 15 minutes. No fleet or batch coefficient is used anywhere. The eight coefficients run from −0.89 to −2.61 %/°C. The correction is applied to the DI zero, to the sample, and to both ends of the commissioning ladder, so it enters the difference on both sides.

It is doing work. Window temperatures across the seven runs span 15.9 to 28.3 °C, and switching the correction off moves 50 ppb from 98% to 92% and the median error from 4.6 to 7.6 ppb.

ρ scaled by50 ppb reads Median error
0, correction off92%7.56 ppb
0.595%6.28 ppb
1, as used98%4.63 ppb
1.5102%4.16 ppb
2103%5.50 ppb

A small residual survives, and the fitted coefficients look slightly shallow. Regressing the procedure's 50 ppb reading on the window temperature within each unit leaves a median +0.43 %/°C, 6 of 8 units positive, and two units are clearly off in opposite directions: 500212 at −1.37 %/°C is over-corrected and 500227 at +2.42 under-corrected. The error also minimises at ρ × 1.5 rather than at 1. That is not a reason to rescale ρ: refitting it across runs absorbs between-run level steps as if they were temperature and returns coefficients as steep as −8.6 %/°C. The deliberate ramp under Still open is what settles it.

The odometer term

The gain ages with use rather than with the calendar, so the stored value is aged to the unit's current cycle count before it is divided by. The rate is fitted per sensor wherever a unit has two gain measurements, and only a unit with none is carried on the fleet rate, recorded as such so it is never mistaken for a characterized one.

UnitRate, % per 10,000 cycles SourceCycles at run 1at run 7
500230+1.1%own39,75744,173
500212+2.7%own36,74041,174
500103+3.4%own39,54343,969
500123+4.6%own40,62045,037
500181+9.1%own40,22744,652
500128−14.3%own18,65622,910
500140+3.3%fleet1,0315,450
500227+3.3%fleet22,59527,033

The spread in those rates is why the term is applied per sensor. 500128, powered down for four of the seven weeks, moves the other way at −14.3% per 10,000 cycles, and correcting it with the fleet rate would take its error from 4.4 ppb to 8.0.

It earns its place out of sample. Each unit's own rate comes from the same two gain measurements this page scores, so applying it here is in-sample by construction. Refitting leave-one-unit-out, and refitting at every choice of reference run, the term beats both alternatives on 7 of 8 choices. Median across those choices: 6.02 ppb with the odometer, 7.67 with a single fleet scalar, 9.67 with no correction at all. It cannot touch the day-to-day spread, which stays at 17%, because a constant divisor cannot change a coefficient of variation.

It is a correction that works, and its name is wrong. The fitted rate swings from −0.08 to +6.65% per 10,000 cycles depending on which run is taken as the reference, and at run 5 it changes sign. A rate that does that is not an identified physical quantity. Cycles vary only 1.11-fold among the five continuously running units, so for those five the term is near enough a constant, and its entire advantage over a plain fleet scalar is that 500128 receives a different correction. So what the term reliably does is remove a fleet-wide offset and treat one unit differently. It is kept because it does that better than the alternatives, on the record in hand, and not because aging with use has been demonstrated.

One rule it breaks. 500140 and 500227 have no earlier ladder at all and run on the fleet rate, which is a batch coefficient doing real work, and recording it as k_src does not make it per sensor. It becomes their own as soon as a second commissioning ladder is run. A third on each unit, spaced in cycles rather than in days, is what would let the shape of the curve be tested instead of assumed exponential, and it is the measurement that would decide whether this term deserves its name.

How it scores

Each run's own DI window supplies the zero. The three procedure rows differ only in where the stored gain came from, which separates the procedure itself from the staleness of its gain. Every row is scored on the identical 47 unit-runs that all arms can reach, run 1 excluded throughout because no arm can be scored against its own gain, and 500227's run 4 ladder excluded for the reason given under The outliers.

Stored gain fromn 50 ppb readsMedian errorDay-to-day CV5 ppb reads
Median of that unit's earlier ladders47102% 4.2 ppb18%76% (1.7 ppb)
Run 1 alone4797%4.8 ppb17%71% (1.6 ppb)
The 6 August ladder, 47 days and ~48,000 cycles earlier47 108%6.5 ppb17%81% (1.1 ppb)
  the same, aged on the odometer (leave-one-out) 4797%5.6 ppb17%73% (1.6 ppb)
the daily two-point calibration, as actually run47115% 10.9 ppb23%220% (6.8 ppb)

With a current gain the procedure reads 50 ppb at 101% with a median error of 4.2 ppb, against 115% and 10.9 ppb for the daily two-point calibration, and day-to-day reproducibility improves from 23% to 17%. At 5 ppb the absolute error falls from 6.8 ppb to 1.7 ppb, though the relative spread there stays wide.

Reproducibility is set by the reading, not by the gain. All three gain sources land at 17 to 18%, because dividing by a constant cannot change a coefficient of variation. The daily calibration sits at 23% only because its divisor is itself re-measured every run and carries its own movement. So no choice of gain source will get below about 17% here; that floor belongs to the ladder reading. The gain's age shows up in accuracy instead, 4.2 ppb on a gain from this week against 6.5 on one seven weeks old, and the odometer term recovers most of that difference without a new ladder, at 5.6 ppb.

Per unit. Six units carry a gain from 6 August; 500140 and 500227 have no usable earlier ladder, so they are commissioned from run 1 and scored on runs 2 to 7 only, never against their own gain. Ordered by error, seven of the eight land between 91% and 111% with 2.4 to 5.6 ppb. 500128 sits apart at 125% and 13.4 ppb.

UnitGain from Cycles at run 1Runs scoredReads at 50 ppbMedian error Day-to-day CV
5001236 Aug40,6207100%2.4 ppb15%
5002306 Aug39,7577100%2.7 ppb16%
500140run 11,031695%3.5 ppb18%
5001816 Aug40,227791%4.7 ppb6%
500227run 122,595592%5.0 ppb16%
5002126 Aug36,7407111%5.5 ppb9%
5001036 Aug39,543795%5.6 ppb20%
5001286 Aug18,6567125%13.4 ppb22%

Across the fleet that is 98% at 50 ppb with a median error of 4.6 ppb on 53 readings, and 72% at 5 ppb with 1.5 ppb. Six of the eight gains are seven weeks old and the procedure still beats the daily two-point calibration on every column, which is the measure of what the cap opening was costing.

The outliers

500227 in run 4 never saw the 50 ppb solution. Its 50 ppb rise came in at 1.81 S-TLF against 8.80 in every other run, and run 4 was ordinary for the other seven units, which sit within −5% to +6% of their own medians. Reading the sweeps one at a time says what happened. The unit climbed steadily through its 5 ppb window and past it, 7.03 to 8.09, and was still climbing when the window closed. At the 50 ppb window it stepped once, 8.09 to 9.16, then decayed to 8.83, where every other unit in that run rose through its 50 ppb window by +0.50 to +4.68.

A single step followed by decay, with no sustained rise, is what handling produces when the concentration does not change. The step is 1.0 S-TLF and this unit's measured cap-opening step is 1.31. The ratio that settles it is the 50 ppb rise over the 5 ppb rise, which for this unit runs 11.7 to 13.0 in the other runs and 3.71 in run 4. So the run 4 reading is its 5 ppb level plus one handling step, which makes it an operator error on one unit in one run.

That ladder is now dropped, and with it the earlier reading that 500227 was the unit to qualify out. With run 4 in it sat at 37% day to day; without it, 16%, in line with the fleet. The whole ladder goes rather than the 50 ppb rung alone, because the 5 ppb reading was taken while the unit was still equilibrating and is 45% low. The exclusion is recorded in driftcal/exclusions.json with its reason, carried into data/driftcal_ladder_cal.json, and applied in every builder, so nothing on this page is scored on a different set from anything else. It is the only point dropped: 53 of the 54 procedure readings remain, and the common basis for the comparison tables falls from 48 unit-runs to 47.

500128 is the unit with a real problem. It reads 125% at 50 ppb with a median error of 13.4 ppb, against 91% to 111% and 2.4 to 5.6 ppb for the other seven. Its blank is three to ten times noisier than the rest at σ = 0.29 S-TLF against 0.02 to 0.07, it is the unit whose gain fell over the seven weeks, and its 5 ppb rise comes out negative in two runs. Nothing about the record explains it away, and per-unit qualification is what would catch it.

Two drift terms

Every cap opening these eight sensors have had is logged in this test, the test start included. That makes 6 August to 22 September an interval of 47 days with no handling at all, and it is what lets the two terms be measured apart. There is a full tryptophan dilution ladder on 6 August (0, 0.1, 0.5, 5, 10, 50 ppb, dosed in place in one vessel, so the cap is never opened between rungs), which gives each unit a gain from before the handling began.

The two terms enter the measurement differently, which is what makes them identifiable. A gain is the difference of two levels taken with no cap opening between them, (50 ppb − DI) / 50, so an additive handling step cancels out of it. An offset is a level, so it carries the step and not the analyte response.

UnitSample cycles
6 Aug to 22 Sep
Gain 6 AugGain run 1Gain changeOffset change
50010348,5320.10920.1244+14%−9%
50012349,6810.13940.1674+20%−45%
50012824,6870.14350.1077−25%−16%
50018149,2990.09360.1326+42%−26%
50021243,0420.13770.1519+10%+36%
50023048,8080.17510.1828+4%−57%
cumulative cycles against the change, run 1 as reference r = 0.84r = −0.27

The gain moved, and the direction is not in doubt. All six units sat through the same 47 days. The five that ran continuously rose by +4% to +42%, a fleet-wide shift of about +15%. 500128 was powered down for four of those weeks, giving it 24,687 cycles against about 48,000 for the others, and it is the only one whose gain fell.

That it tracks cycles is not established, and an earlier version of this page overstated it. The r = 0.84 above takes run 1 as the current gain. Taking each of the other runs instead gives r from 0.20 to 0.94, and among the five units that actually ran continuously the correlation runs −0.62 to +0.55 with a median of −0.16. The within-unit run-to-run CV of the ladder gain is 6 to 21%, so a single reference run carries that noise. Cycles span 2.18-fold across the six units and only 1.11-fold among the five continuous ones, which leaves 500128 as the sole low-cycle point. With n = 1 at that end, "tracks cycles" cannot be separated from "500128 differs", and 500128 is also the noisy-blank outlier. What survives is the fleet-wide shift and one unit that went the other way after sitting idle. Its cause is not identified: a different tryptophan stock, dosing in place in August against bath transfers in September, and temperature are all uncontrolled between the two dates.

Handling moves the offset, and leaves the gain alone. Each cap opening adds a median 1.07 S-TLF on 8 of 8 units, flat across a 15-fold spread in pedestal, so it is additive rather than a fraction of the reading. It is used up, since a repeat opening an hour later moves the fleet −0.7%, and it relaxes at about −1 %/hour. Inside this week the gain has no systematic trend, +1% per run with mixed signs and r² near zero, and over the 47 days the offset moved between −57% and +36% with no relation to cycles at all.

What follows for each. Handling is a protocol problem and the protocol can reduce it: take the zero from a reading with no cap opening and no lift between it and the sample. The gain shift is not a protocol problem. It cannot be removed by any ordering of the windows, and a gain fixed at commissioning goes stale, though on what schedule this record does not say.

Further limits. The idle period makes cycles and powered hours the same variable for 500128, so those two cannot be told apart either. The interval has two endpoints and nothing in between, so the shape of the curve is unknown and the percentages should not be read as a rate and extrapolated. 500140 and 500227 have no 6 August ladder at all, and 500227's only earlier ladder is the 1 September spike, whose doses cannot be converted to concentration without the barrel volume, which is not in the recorded data.

What was run

One run, repeated seven times. The cap is opened during the puck block, which is what the next part turns out to be about.

  1. Sampling check: every unit produced samples in the last hour (v0.1.21 stall bug: power-cycle any stalled unit, log it in a note).
  2. Blank + DI (parallel, storage bath): de-bubble, settle 10 min, start Blank cap + DI (ALL); the 5-min window ends itself.
  3. Puck + DI (serial): rotate the two pucks over their fixed unit groups; per unit: swap cap, de-bubble, settle 10 min, start FDOM puck 13.59 or 12.94 RFU + DI (that barcode); 5-min window ends itself.
  4. Validation set (all units together, ascending): DI → 5 ppb → 50 ppb; per bath: de-bubble, settle 10 min, start (ALL); 5-min window ends itself. Fresh dilutions daily, same recipe/containers; record water temps.
  5. Return units to the storage bath, blank caps on, powered. No power cycles between sessions.

Run 7, 25 September — what makes the step

Every handling step measured before this had taken the optical window through the air–water boundary, so the test was whether the boundary is required. Run 7 was done with no unit leaving the bath, and the prediction was set in advance: if crossing is the mechanism, a submerged cap swap comes in at the shake level. It did not. A cap removed and put back entirely under water produces most of the step, so crossing is not required. It is not the whole story either, as the last column shows.

The step is additive, flat across a 15-fold spread in pedestal, so S-TLF is the metric that compares events and a percentage is not: events on different days sit on different baselines. Ranked on the additive scale, median over the 8 units, 2 to 8 min after against the 6 min before:

EventAirCap Step, S-TLFStep, %
Shake in the bath (three of them)nountouched+0.32 to +0.452.0 to 3.6%
Quarter turn under water, 12:11noloosened+0.472.7%
Cap off and on under water, 11:27noremoved+1.198.6%
Out and back by hand, 24 Sep 13:43yesuntouched+1.8813.6%

So neither air nor the cap is necessary and each is sufficient, and the air round trip with the cap never opened is the larger of the two, by about half again. What orders the four events is how completely the water at the window is exchanged: a shake moves the bath but not the enclosed volume, a quarter turn breaks the seal, a submerged removal exchanges the cavity, and lifting the unit out drains and refills it. An earlier version of this page said the crossing was not a mechanism at all. That was a comparison of percentages across days with different baselines, and on the additive scale it does not hold.

The step is used up. The same submerged removal repeated four times an hour later moved the fleet by −0.7% with 2 of 8 units up, and the blank → DI-test transfer that ran at +8.1% (1.50 S-TLF) across runs 1 to 5 and +11.8% (2.29) in run 6 came in at +0.4% (0.03 S-TLF) in run 7. Whatever is released has to rebuild before it can be released again.

Temperature does not account for it. The 11:27 step is +8.4% with the quench correction switched off entirely and +7.4% with it doubled, and it would take 5.1 °C of cooling at the median fitted coefficient to fake, against 0.30 °C of measured sipm_temperature movement. Nor is it a ramp read as a step: the trace is flat between events, and the undisturbed drift through the same session is −1.07 %/hour, downward.

What it does not settle. What accumulates, and how long it takes to rebuild. Full effect appeared 2.1 h after the previous cap disturbance and nothing appeared at 0.7 and 1.0 h. The overnight traces in the next section show this is not a recovery time: the raised level does not come down in that hour, it holds, and a second disturbance simply cannot raise it further. The level relaxes over about ten hours. The direct test is still one cap removal, under water, repeated at 30 min, 1, 2, 4 and 8 h with nothing else touched.

The practical consequence is already visible in the run 7 transfer. The step that has been setting the day-scale detection limit is discharged by opening the cap once and waiting, so a purge before the calibration, rather than a correction applied after it, is the thing to test next.

After the step — does the reading come back, and what is left to explain

Every unit was traced through the whole record at 30 to 60 min resolution, S-TLF at 20 °C on each unit's own coefficient against sipm_temperature, to ask whether a disturbed unit returns to where it was. It does, on a timescale of about ten hours, and to nearly the same place each morning. Read with the run 7 result, the step behaves as a two-state system: any disturbance of the cap joint or the housing drives the reading into a raised state that adds a fixed amount of returned excitation; that state saturates, so a second disturbance inside it adds nothing; left alone it relaxes back to a unit-specific baseline overnight.

Within the first one to two hours the raised level holds. After the 13:43 out-and-back on 24 September the low-background units sat at their new level from 13:30 to 15:00 with no decline: 500212 at 9.1, 9.0, 9.2, 9.2; 500227 at 6.4, 6.3, 6.7, 6.6. The “used up” result above is therefore not the effect fading; it is the ceiling.

Overnight it relaxes, the same way both nights the units were in DI, and the two morning baselines agree to within 2 to 10% on every unit:

Unit23 Sep 18:3024 Sep 08:30Overnight 24 Sep 18:0025 Sep 08:00Overnight
5002126.96.1−12%7.06.3−10%
5002274.74.4−6%4.84.5−6%
5001236.95.8−16%7.76.4−17%
50023012.19.5−21%10.18.4−17%
50018128.027.90%28.828.4−1%
50012836.835.1−5%38.335.2−8%
50010342.241.4−2%42.141.6−1%
50014075.475.10%75.975.60%

The decline is smooth and monotonic on every unit, about 1%/hour on the low-background units at the start and slowing. The same decay runs across the daytime gap between runs 3 and 4 on the 23rd: 500230 went 17.5 to 14.4 over five undisturbed hours. So the daily protocol read its calibration on a relaxed unit and its ladder, an hour and one cap opening later, on a raised one, and that is the offset error the record carries.

Two footings the runs do not share. From the end of run 2 until run 3 at 07:30 on 23 September the units were out of the water: every trace sat at its in-air level (42 to 116 on the units that show it most) with sipm_temperature at 21.7 °C, room temperature. Run 3's blank was therefore read within minutes of a re-immersion, on units in the raised state, and the readings decayed through that whole day. Runs 5 and 6 started from relaxed overnight baselines; runs 3 and 4 did not. And on the morning of 22 September, before run 1, the low-background units read about half their later baselines (500212 2.7 against 6.1 two mornings later, 500123 2.8 against 5.8, 500230 5.4 against 9.5) while the bright units were unchanged. The units had been powered up and immersed that morning; whether that was a true lower baseline or pre-equilibration cannot be told from this record, but runs 1 and 2 were scored against a level the units never returned to.

Also scored: the cap removal in the 50 ppb bath, 13:20 on 25 September (registry 100), the last handling step of run 7. Median 5 to 15 min after against 5 min before: fleet median about +0.9 S-TLF (500230 +3.0, 500212 +1.8, 500103 +1.9, 500140 +1.2, 500123 +1.0, 500227 +0.7, 500181 +0.2, 500128 0.0), LED-proportional and pedestal unchanged. The same order as the DI removal at 11:27, with 50 ppb of fluorophore at the window. Only the first ten minutes are scoreable: the units came out of the water at 13:30 (registry 101) and the test ended.

What the record has ruled out

CandidateKilled by
SiPM saturation and slow recovery after light exposureThe added signal scales 1 : 4 : 16 with LED drive on every unit and the no-gain pedestal is unchanged to the count. It is LED light returned from in front of the window, not detector behaviour; and the level holds flat for 40 min where recovery would decay.
TemperatureSteps survive with the quench correction off or doubled; sipm_temperature moves 0.3 °C against the 5 °C needed.
Photobleaching of the trapped cavity waterNothing to bleach in recirculated DI at the 7 to 14 ppb-equivalent the step would need; the removal in 50 ppb gave no larger step.
Anything in the cavity water: particulates settling, gas film, thermal gradientShaking exchanges the cavity water and gives +0.3 to +0.5 against +1.2 to +1.9 for opening or draining the joint.
Air carried in on reinstallWould be an occasional outlier; this is 8 of 8 units, every time, at the same size, with de-bubbling done every time.
Cap orientation or resting positionCaps have been keyed since run 5; the underwater removal and the run 6 transfer stepped by the usual amount.

What is left standing

Whatever it is lives in the mechanical assembly in front of the detector, has a memory of hours, saturates, and is untouched by water moving through the cell. Three candidates fit, and each is testable in a session.

  1. Stress relaxation in the compressed elastomer of the joint. Breaking and remaking the joint, or loading the housing by hand, puts the seal at a fresh state; it creeps back over hours, on exactly this timescale at 16 to 25 °C. A stray path from LED to detector through or around that seal tracks its state. Same seal on every unit, so the same size everywhere; a joint remade twice in an hour is at the same state either way, which is the saturation. The record hints at the thermal signature: 500230 fell 18% in five daytime hours at 23 °C and 21% in fourteen night hours at 16.5 °C.
  2. The window or filter shifting on its mount. The same seal, acting on the element it holds: a fraction of a degree of window tilt changes how much of the window's own front-surface reflection of the excitation reaches the detector. That reflection depends only on the LED and the geometry, so it is the same on every unit and independent of what is in the water, which the 50 ppb result requires.
  3. A pinned capillary film at the cap-to-bezel contact line, which bulk flow never exchanges, reformed when the joint is broken. Not the water's contents and not a bubble, but water geometry. Weakest of the three on the ten-hour timescale.

Tests that separate them. Press on the cap firmly under water without turning it: 1 and 2 predict a step, 3 does not. Run the same disturbance on two units in a 25 °C bath and two in a 10 °C bath: elastomer creep goes several times faster over that range, a film does not. Put a dial indicator on the cap face, and on the window if it can be reached, before a disturbance, immediately after, and at 1, 3 and 10 h: movement of tens of microns following the S-TLF decay is 1 or 2 in one number, and which surface moves says which.

What follows whichever it is. The relaxed overnight state is the reproducible one, so the reference and the reading must be taken in the same state, both relaxed or both freshly disturbed; the procedure above does this by touching nothing between the DI read and the sample. And a design that removes the elastomer from the optical path, or preloads it so handling cannot change its state, takes the step out at the source.

Is there drift that does not come from handling?

Not much, once a unit has settled, and this record cannot show any on its own: every trend in it follows a handling event within hours. Two other records can. The three days since the test ended, with the units untouched (in air, since 13:30 on 25 September), and the barrel's DI tank, where five of these eight units sat for 8.4 days from 21 to 29 August with nothing done to them after the first afternoon.

Since the test ended. Rate of change of S-TLF at 20 °C, per unit, by interval after the units came out of the water. The rate falls toward zero as the disturbance recedes; a single line through the three days would read −0.5 to −1.6 %/day and call a decaying transient a slope.

Interval after removal500212500227500123500230500181500128500103500140
first 15 h−5.8+0.5−6.4−4.3−2.2−1.5+11.3−3.2
day 2−1.0−4.2−0.6−1.9−1.0−0.4−1.2−1.1
day 3−0.25−0.2−0.4−0.9−0.3−2.1−0.3−0.5
day 4 (12 h)0.0−0.2−0.8−1.10.00.00.0−0.5

%/day. By day 3 six units are inside 0.5 %/day; 500230 is still moving at about −1 and 500128, the unit with 94% internal humidity, is noisy rather than trending. This is the same result the drift-rate factorial reached in mid-September on different units: the dominant drift is a settling transient in elapsed time after a disturbance, and it flattens in calendar days. The LED-current tilt that factorial found (the 512:32 response ratio falling 0.35 to 0.87 %/day) does not show here: over the three days it ran −0.4 to +0.5 %/day with mixed signs.

The barrel: 8.4 undisturbed days in water. Daily medians of the canonical amplitude, each unit against its own reference, from the frozen ditank record. The shake, heater ramp and UV dose all fell on 20 August; the registry shows nothing after it. 500128 went idle on the 21st and has no record here.

Date500212500123500230500181500103Bath °C
21 Aug0.800.840.891.181.2111
22 Aug0.860.900.921.031.0619
23 Aug0.960.980.991.001.0122
24 Aug1.011.011.000.991.0022
25 Aug1.071.051.031.001.0023
26 Aug1.111.081.051.001.0023
27 Aug1.121.091.061.001.0023
28 Aug1.121.091.061.001.0023
29 Aug1.121.091.051.001.0022

Read the first 2.5 days as temperature, not settling. The amplitude in this file is not temperature corrected, and the bath went from 34 °C under the heater on the afternoon of the 20th to 8 °C outdoors that night and back to 22 °C by the 23rd. The bright units' fall from 1.18 to 1.00 and most of the low-background units' rise over 21 to 23 August track that swing, in opposite directions, and cannot be separated from settling. The clean stretch is 23 to 29 August at a constant 22 to 23 °C.

Two behaviours at constant temperature. The bright units, 500181 and 500103, are flat from the 23rd on: over the last 3.4 days their rates are +0.1 and −0.1 %/day with temperature fitted jointly. The low-background units, 500212, 500123 and 500230, keep climbing for about four more days, by 17, 11 and 6% from the 23rd to the 27th, and then plateau; over the last 3.4 days they are at +0.65, +0.52 and +0.12 %/day and still slowing. From immersion to plateau is about seven days. A straight line through all eight days reports +4 to +5 %/day on 500212 and calls a slow approach to a plateau a drift.

The spike record confirms the plateau. The same units, still in the same barrel on 29 to 31 August, nine to eleven days after immersion, at a constant 22.4 °C and untouched: S-TLF at 20 °C moved −0.2 to −0.9 %/day over two days (500181 30.27 to 30.19, 500212 6.38 to 6.26, 500103 43.93 to 43.75). That is the settled state.

This record shows the same climb, larger. On the morning of 22 September, when the units were powered and immersed for the test, the low-background units read about half of what they read two mornings later (500212 2.7 against 6.1, 500123 2.8 against 5.8, 500230 5.4 against 9.5) while the bright units were unchanged. That is a bigger and faster rise than the barrel's, at a warmer 27 to 29 °C and from a first power-up, so the two are not the same measurement; both agree on the direction and on the unit type. So there are two transients, and they go opposite ways: the post-handling raised state, which decays over about ten hours, and the post-immersion climb on low-background units, which takes days.

How long each kind of relaxation takes

Three records, three different disturbances, and the spike test adds two that are not the instrument at all. The spike record (same barrel, 29 August to 2 September) was read for this: before the doses the units were untouched for two days; each dose was read 10 min after addition; and the test closed overnight with the pump off and 7 NTU of turbidity standard plus three fluorophores in the water.

What was disturbedHow long it takesWhere measured
Cap joint opened, or unit handledRaised level holds 1 to 2 h, then relaxes over about 10 to 14 h at constant temperature, to a baseline that repeats morning to morningThis record: both DI overnights and the 23 September daytime
Unit taken out of water into airRate of change falls from −6 %/day to under 0.5 %/day by day 3 and to near zero by day 4This record: the three days after the test
Unit immersed, low-background typeClimbs for about a week, then flatBarrel, constant-temperature stretch 23 to 27 August; confirmed settled by the spike pre-dose stretch on 29 to 31 August (−0.2 to −0.9 %/day); same direction here on 22 to 24 September
Unit immersed, bright typeUnder 2.5 days; not resolvable under the barrel's temperature swingBarrel
Water changed around a settled unitComplete within the 10-minute mixing wait; the hourly trace steps cleanly rung to rungSpike dosing, 1 September
Turbid, dosed water left still overnightReadings fell 13 to 28% by morning. This is particles settling and the analytes changing, not the sensors; the spike page opens a fresh baseline after the pump restarts for that reasonSpike, 1 to 2 September

Two floors the spike record sets on itself: its S-TLF is normalised on the fleet quench constant rather than per-unit coefficients, so the chiller setpoint change on 31 August (22.5 to 15.7 °C) moves the corrected trace by about −6% on 500181, and the chiller's 1 °C cycling afterwards leaves a residual oscillation of about 0.5%. That is correction residue, not relaxation, and it is the resolution limit of that record.

What this means for the runs. Runs 1 and 2 on 22 September were made in the first hours of that week-long climb, and runs 3 and 4 the next morning after a night in air, so the low-background units' baselines were moving under them through the first half of the study, independent of any handling. Runs 5 to 7 are the first on settled units and are the ones the procedure should be judged on. It also gives the procedure a number: a new or re-immersed low-background unit needs about a week in water before its zero is stable enough to store. Once settled, this fleet's drift is below roughly 0.5 %/day on the low-background units and indistinguishable from zero on the bright ones at 3-day resolution, and nothing in either record shows a persistent handling-independent slope larger than that.

The holding question, re-asked with handling instead of the clock

The three aged views index a calibration by how many hours old it is. Run 7 gave a reason to doubt that the clock is the variable, because an older calibration is also one with more cap openings between it and the reading, and the two cannot be separated by looking at hours. Every pairing was re-indexed by cap openings, counting each puck window as two (the cap comes off to fit the puck and off again to put the blank back) and marking whether each opening happened in air or under water.

The discriminator is the fresh calibration, where time is held constant. In every run the calibration and the ladder are about half an hour apart with exactly one cap opening between them, the one that takes the puck off and puts the blank back. Runs 1 to 6 did that in air and run 7 under water.

Fresh calibrationcal to read OpeningDI readsError at 50 ppb
Runs 1 to 6 pooled0.5 hin air10.4 ppb11.0 ppb
Run 70.5 hunder water0.9 ppb26.4 ppb

The zero error is handling. DI should read zero by definition and in runs 1 to 6 it reads 10.4 ppb; with the cap opened under water it reads 0.9 ppb. Nothing about the clock changed between those two rows. That is the single largest error on this page, and it is an artifact of taking the cap off, not of a calibration decaying.

The gain goes the other way, and the reason is the same step. Run 7 reads 50 ppb as 76 ppb, 153% recovery, against 103 to 132% in the air runs. Its span is 33% smaller (3.27 S-TLF against a 4.48 to 5.29 median), and 6 of the 8 units fall below their entire run 1-to-6 range, so this is not one unit moving. The span is puck − blank, and the blank is read before the cap is opened while the puck is read after it, so in the air runs the span carries one cap-opening step and is inflated by it. Subtracting each unit's own measured step, in S-TLF because the step is additive, reproduces run 7's span to a median factor of 1.15 (range 0.91 to 1.81). Most of the collapse is the step; the remainder is not accounted for.

So the daily calibration appeared to recover the gain in runs 1 to 6 because the step inflated the divisor and the readings together and the two errors partly cancelled. Removing the handling exposes both: the zero becomes right and the gain becomes wrong.

One confound this cannot separate. A puck cap fitted under water holds water between the puck face and the window where one fitted in air may not, which would lower the puck reading on its own and shrink the span without any step. One submerged run cannot tell that apart from the step explanation, and the experiment is closed, so both remain on the table. What does not depend on the choice is the zero, which the submerged run fixes outright.

The ladders

All seven ladders, one color per unit, run as line style (solid run 1, dashed 2, dotted 3, long dash 4, dash-dot 5, long dash-dot 6, heavy 7). Click a legend entry to hide a unit, double-click to isolate it. Three views: the readings as they come off the instrument, the procedure applied to them, and the daily two-point calibration for contrast.

Runs 1 to 5 read against a matte cap interior, runs 6 and 7 against bare aluminum, after the sticker came off on 24 September at 17:40, so blanks either side of that are not the same reference. Run 7 was done with no unit leaving the water.

What else was tried

Six alternatives, all scored on the same ladders. None of them beat the procedure above, and the reasons are specific rather than general.

Every scenario on one basis

The best scenario is the procedure with the gain taken as the median of that unit's earlier ladders: 101% recovery at 50 ppb, a median error of 4.3 ppb, and 18% day to day. All three versions of the procedure beat every other arm on reproducibility; where the stored gain came from changes the accuracy while leaving the spread alone.

Every arm below is scored on the identical 47 unit-runs, the set all of them can reach, so the ranking is not an artefact of which points each one happens to cover. Run 1 is excluded throughout, because an arm cannot be scored against its own gain, and so is 500227's run 4 ladder. DI reads is what the zero rung comes out as when it should be zero. Rec is the median reading over the nominal. Day-to-day CV is within a unit across runs, the criterion the protocol registered at 10%.

Scenarion DI reads5 ppb
rec / err
50 ppb
rec / err
Day-to-day CV
5 / 50 ppb
Between-unit
at 50 ppb
Procedure, gain = median of earlier ladders 470.076% / 1.7102% / 4.244% / 18%6%
Procedure, gain from run 1 only470.0 71% / 1.697% / 4.845% / 17%10%
Procedure, 6 Aug gain aged on the odometer (leave-one-out) 470.073% / 1.697% / 5.645% / 17%12%
Procedure, gain from the 6 Aug ladder, no odometer term 470.081% / 1.1108% / 6.545% / 17%14%
Puck span, ladder DI zero470.075% / 1.997% / 6.8 52% / 26%8%
Daily two-point, as run (blank zero, puck span)478.0 ppb220% / 6.8 115% / 10.947% / 23%8%
Procedure, but rescaled by the puck each run470.092% / 1.5 120% / 14.452% / 26%24%
Puck then DI then test (span = puck − DI)470.0103% / 2.8 136% / 18.292% / 31%73%
arms that produce no concentration, so CV only
Raw S-TLF level, nothing applied56n/an/a n/a9% / 7%67%
Zero only, from the ladder DI560.0n/a n/a41% / 17%15%

Rows are ordered by median error at 50 ppb. The four procedure rows differ only in where the stored gain came from, and they sit at 17 to 18% day to day whichever it is, against 23% for the daily calibration and 31% for the worst arm. The gain's age shows up in accuracy instead: 4.2 ppb on a gain from earlier this week, 6.5 ppb on one seven weeks old, and 5.6 ppb on that same old gain once the odometer term ages it.

Moving the zero to the ladder's own DI is what fixes the low rung. Every arm that does it reads 5 ppb between 71% and 103%, where the daily calibration reads 220%. Reading the puck immediately before the DI is the worst arrangement tested, at 31% day to day and 73% between units, because the cap opening lands between the puck and its zero and shrinks the divisor. Rescaling a stored gain by the puck each run undoes most of what the stored gain buys, 14.4 ppb against 5.6.

Nothing clears the 10% bar. The best day-to-day figure among arms that produce a concentration is 17%. The raw S-TLF level sits at 7%, but that is a level carrying a large per-unit pedestal, which flatters a coefficient of variation; the same units are 67% apart from each other on that basis, against 8% once a gain and a zero are applied.

Do the FDOM pucks earn their place?

The puck exists to track gain, and gain drift is now known to be real: it moved a median 20% over the 47 days before this test, ordered by sample cycles. So the question is whether the puck follows it. Three tests, scored on all seven runs.

It does not track the gain. Within a unit, across runs, the correlation between the puck response and that run's ladder gain is a median −0.17 against the blank and −0.08 against the DI-test, with 3 of 8 units positive. A gain standard would sit near +1. Expressed in ppb through each unit's own ladder, a fixed puck should read the same number every run; its within-unit CV is 28%.

Used end to end it makes the procedure worse. Take the stored gain from run 1 and rescale it each run by that run's puck response, which is what a gain standard is for.

Gain appliedn50 ppb reads Median errorDay-to-day CV5 ppb error
Stored gain, no puck4848.5 ppb (97%) 4.8 ppb18%1.7 ppb
Rescaled by the puck, referenced to the blank4852.5 ppb (105%)12.4 ppb30%1.5 ppb
Rescaled by the puck, referenced to the DI-test4842.7 ppb (85%)17.3 ppb41%2.6 ppb

Why it cannot help at any interval, which is the part that does not depend on the length of this test. For a gain standard to be worth applying, its own run-to-run noise has to be smaller than the drift it corrects. The puck's noise is 28%. The gain drift it would be correcting is about 0% within this week and 20% over 47 days. The noise exceeds the signal even at seven weeks, so rescaling by the puck adds more error than it removes, which is what the table shows: the median error goes from 4.8 ppb to 12.4 or 17.3 ppb.

The submerged swap fixes the level but not this. Through each unit's own ladder the puck is worth 33.0 ppb across runs 1 to 6, where the cap was opened in air, and 20.6 ppb in run 7, where it was opened under water. The 38% difference is the cap-opening step sitting in every air-run puck reading. Removing it corrects the number the puck is worth, and one run cannot show whether it also makes the puck track gain, so that stays open. What the submerged reading does not change is the 28% run-to-run noise that the end-to-end test is measuring.

The puck is still worth reading as a diagnostic, since it is the one fixed optical reference in the protocol and it is what exposed the cap-opening step. It should not be in the calibration path.

The recalibration interval is still not measurable. The record holds two intervals, about 25,000 and 48,000 cycles, and the gain had moved 10% to 42% at both, so nothing says where it is still good. A ladder repeated at 2,000, 5,000, 10,000 and 20,000 cycles would settle it, and it needs no handling, because the ladder is dosed in place.

What survives the calibration

Calibrating removes each unit's baseline and gain, so the eight units do move closer together. What it does to day-to-day reproducibility, which is the pre-registered criterion, is the opposite. Bars are the mean within-unit CV across the seven runs; the dashed line is the 10% bar the protocol set for success.

Is it the temperature correction? Partly, and run 4 is what showed it. Runs 1 to 3 sat inside 2.0 °C of each other, which gave this test almost no temperature leverage, and on those three runs no test could find any dependence: the calibrated error against the ladder-minus-calibration gap came back at 10.2, 6.4 and −2.5 %/°C, every one inside its own standard error. Run 4 ran at 16.4 °C against 25 to 27 °C for the others, a 10.7 °C spread and five times the leverage, and the dependence is now measurable.

DI is nominally the same water every run, so after a complete quench correction its S-TLF at 20 °C should not depend on how warm the run was. Regressed across the runs it still does, by a median 0.76 %/°C, which over the 10.7 °C from run 1 to run 4 is 8% of reading. 500140 is the clean case, with no level step and tight windows:

500140, DITraw S-TLF at 20 °C, ρ = −1.43% (in use)at 20 °C, ρ = −2.09% (refit)
Run 128.4 °C63.0771.1475.16
Run 227.8 °C64.8272.4776.27
Run 325.1 °C67.0772.1774.63
Run 415.9 °C83.2678.5576.49
spread across runs 1 to 4 SD 3.36, CV 4.6%SD 0.88, CV 1.2%

The coefficient in use is too shallow, so a cold run stays too high after correction, which is exactly what run 4 does. Its residual slope is −0.76 %/°C at SE 0.12, t = −6.08. Two limits on this. It is unit-specific: 500103 shows no residual dependence (+0.07 %/°C, t = 0.29) and needs no change. And refitting ρ across runs absorbs between-run level steps as if they were temperature, which on the low-signal units returns coefficients as steep as −8.6 %/°C, far outside anything this program has measured. A coefficient fitted on a deliberate ramp at constant analyte is the only clean way to settle it, which is the run described above.

The puck in ppb — the measurement behind that verdict

Each puck reading is converted into a TLF ppb-equivalent using that unit's own tryptophan ladder from the same run: ppb = (puck − DI) ÷ ((50 ppb − DI) / 50), all in S-TLF at 20 °C. Doing it per unit per run cancels both the unit's gain and any day-to-day drift, so if a puck is a fixed standard every marker should land on the same height regardless of which sensor read it or when. Temperature is normalized on each unit's own quench coefficient (fitted from that unit's own ramp over the undisturbed periods between sessions) against sipm_temperature — no fleet or batch constant is used anywhere.

The DI reference is bracketed, not trailing. Each run measures DI, then the eight pucks, then DI again — identical water either side. A unit's reading on identical water moves by about ±1.7 S-TLF over half an hour, and that drift is additive and much the same size on every unit (correlation with signal level 0.12), so it is a fixed offset wandering rather than a gain changing. It is not temperature — after the per-sensor correction the residual change still tracks ΔT at only r = 0.04. Referencing each puck to the DI measured after the puck block therefore injected that drift instead of cancelling it. Interpolating the DI baseline in time to each puck's own midpoint removes it: pooled CV falls from 35.5% to 18.1%, and the center moves up, because the trailing reference was biased low.

One marker per reading; dashed lines are each puck's median, the solid line and shaded band the adopted truth. Hover for the run and the raw step. The adopted value is the gain anchor every calibrated view below uses.

Adopted truth — … for both pucks
One value is used for both pucks on purpose. Nothing in the data separates them — the difference is +2.0 ppb with a standard error of 1.96 (Welch t = 1.03, 21 df), and their confidence intervals overlap heavily. Seven of the eight units saw a different puck on different days, so encoding a difference the data does not support would inject a step into exactly the comparison this test is measuring.

To update: the value lives in data/driftcal_puck.json under truth, regenerated by truth.mjs → build_puck_chart.mjs in lume-ecoli-data/driftcal/. It is a working number from — each new ladder should tighten it, and it should be re-derived rather than edited by hand. If a future run ever does separate the two pucks, split them then and say which run showed it.

Live signal — S-TLF slope at 20 °C

S-TLF = the production bias-response slope, fitTlfSlope imported live from the deployed model at thelume.ai/shared/ecoli-model.js: one value per sweep from all LED × bias cells, per-sweep pedestal, quadratic fit, derivative at bias 3000. Then normalized to 20 °C with the model's own tlfSlopeAt20C, using the diagnostics temperature nearest each sweep (≤ 90 s) and the model's quench coefficient for that unit (batch value −2.15 %/°C for these units; none has a measured one). Production joins the board temperature, which on the bench is driven by charging (25–45 °C today) rather than the water; the SiPM temperature (24–30 °C) is offered as an alternative so the two can be compared. Every chart and table on this page follows the selector. Hover shows the raw slope and the temperature used.

Idle.

Test points — one marker per unit per window

y = median S-TLF at 20 °C over the window after dropping its first 60 s (handling). Same colors as the time series below. Hover for window id, start time, and cycle count. As days accumulate, each condition column fills with one marker per unit per day, which is the day-to-day spread the test scores.

Step size vs signal level — run transfers against the handling events

Every run: the Blank cap + DI window against the DI (validation) window that came next (window medians, current temperature setting, first 60 s dropped) is one circle: y = how much the reading changed, x = where the unit was reading before. Events are the same thing for the 5 min before and after: crosses for shakes in water, diamonds for out-of-water and back. Color is the unit. An additive offset at the window plots as a flat band in slope units regardless of level; a gain change plots as a line rising with level. Switch y to percent to see the same steps as a fraction of reading. Fetched per window for every run in the registry, independent of the time-series range above. The 17:40 cap change on 24 Sep is marked on the traces below but is not scored here: the units were still in run 5's 50 ppb solution right up to it and returned to DI at the same moment, so its before-and-after is the medium changing rather than the cap.

Waiting for registry…

Time series

All 8 units on one axis, one color per unit; click a legend entry to hide it, double-click to isolate it. Shaded bands = marked ALL condition windows (green blank · dark teal puck 13.59 · light teal puck 12.94 · grey DI · amber 5 ppb · red 50 ppb); per-unit puck windows are not shaded here. A sweep that never rises above its own pedestal returns no slope and is skipped.

Temperature

Same units, same colors and the same shaded windows as the trace above, so a step in S-TLF can be read against what the temperature was doing at that moment. The channel plotted follows the selector, and sipm_temperature is the one next to the water; the board channel self-heats about 1 °C above it.

Limit of detection — what this test buys us, in CFU/100 mL

Computed for the procedure as it is actually run: the DI blank is read and nothing else, so the limit is 3 σ(zero) ÷ gain with the stored commissioning gain. The zero's noise is measured on the undisturbed stretches of 25 September, units in DI with blank caps on and nothing touched, either side of the 11:27 cap swap. It is derived twice here, once from this bench record and once from the deployed field record, and the two do not share an assumption. Field set is — paired Colilert grabs on — deployed sensors, model —.

The clock stops mattering once the handling is out. Reading the sample 20 minutes after the DI costs almost nothing, 1.21 ppb against 1.12 back to back, because an undisturbed zero moves about 0.08 S-TLF an hour, which is a small fraction of a rung. Under the old daily calibration the same interval was worth tens of ppb, and that was the cap-opening step rather than the passage of time.

The binding constraint is now the individual unit. The fleet median clears 10 CFU/100 mL, but only 4 of the 8 units do on their own. The spread runs from 4 to 43 CFU/100 mL, with 500128 at the top on a blank three to ten times noisier than the rest. Every unit clears 200. Per-unit qualification therefore decides which sensors can be used against a drinking water threshold, and no change to the procedure moves that.

Two limits on the number. The zero's noise is measured over about 70 minutes on one day, so it carries the conditions of that day. And the stored gain's age scales the result without adding any blank noise: a gain 47 days old is low by a median 14%, which inflates the computed limit by about the same fraction and leaves the ranking of units unchanged.

The field record does not follow the bench down. On deployed sensors against paired Colilert grabs the signal first clears chance at 100 CFU/100 mL, not 9. The instrument resolving a part per billion of tryptophan in DI is a different claim from a deployed sensor separating river water at a threshold. The gap between the two tiles measures that difference.

The two ways we use the instrument

The same units meet one target and sit on the edge of the other, because the two modes differ in how long the blank has to hold. Calibrate and read straight away and the limit is the sensor's precision. Calibrate once and come back a day later and the limit is its drift.

Each unit's own limit against both targets. The bar is the unit; the two dashed lines are the thresholds the two modes have to reach. A unit below a line clears that target.

How the limit is derived

Two routes that share no assumption, then a check on what the served number is actually made of.

Bench: blank noise divided by gain

Both terms measured on the same 8 units in the test above. Blank is the DI window, gain is (50 ppb − DI) / 50 in S-TLF at 20 °C on each unit's own quench coefficient against sipm_temperature. Detection limit is 3 × SD(blank) / gain, the ordinary 3σ construction. The bar is the limit; the three series are the three timescales over which the blank is allowed to move.

Field: where does the signal separate from chance?

For each candidate threshold, label every grab above or below it and ask how well the sensor's own TLF reading ranks them (AUC). The signal is centred within sensor first, so what is left is the unit's own movement at its own site, with the chronic level of that site removed. 0.5 is chance. A point whose 95% interval clears 0.5 is a concentration we can detect; one below 0.5 means the sensor reads the wrong way round. Intervals are 2,000-sample bootstraps.

How much of the dashboard number is the site and not the sensor

The same test run twice on the deployed prediction: once as served, and once after removing each sensor's own mean. What survives the second version is what the sensor contributes beyond knowing which site is chronically dirty.

Notes on the numbers

Converting ppb to CFU

Why the CBT set could not be used

Still open

Tryptophan cannot be held at one concentration long enough to ramp temperature, since it degrades. DI water and the FDOM pucks are stable, and between them they carry the quantity the calibration actually divides by.

  1. Split the fleet: four units blank-capped in DI, four with pucks on, all in one bath.
  2. Ramp and soak across the widest range the bath reaches, holding each setpoint until sipm_temperature is flat, so the water and the SiPM are not chasing each other.
  3. Swap caps and repeat the same ramp, so every unit is measured in both states and no result rests on one group of units or one ramp.
  4. Fit ln(S-TLF) ~ sipm_temperature per unit per state. The difference between the two states gives the temperature response of span = puck − blank, which is the divisor.

What this settles and what it does not. It measures the pedestal coefficient properly, over a deliberate range instead of the 2 to 6 °C of incidental drift the current fits rest on, where three of the eight units sit at r² 0.22, 0.27 and 0.42. It measures the span's own temperature response for the first time. It does not give tryptophan's coefficient, because the puck's fluorophore is a different material, and the puck-on reading is the sensor and the puck responding together rather than the sensor alone. For the calibration that is the right quantity, since the span is what gets divided. For the ladder readings the analyte coefficient stays open.

Every window is pre-set to 5 minutes and ends itself, and its first 60 s are dropped as handling. Dilutions are fresh daily, from one day-0 refrigerated stock. Scoring (pre-registered): per unit/day offset = blank median, span = puck − blank; the TLF standards are scored calibrated-daily vs calibrated-once-on-day-1 (ablation) vs raw. Success = daily-cal day-to-day CV of the standards ≤10% and materially below the frozen-cal CV. Span is a small difference of large numbers (puck adds 20–35% over blank), so the dilutions double as an independent gain check.

Condition windows — the bench tool

The recording interface the runs were made with. The experiment is closed; this is kept so the record can be read and extended.

Start a window when the units are settled in the condition. It runs a fixed 5 minutes and ends on its own (the end time is stamped server-side, so the page can be closed). Stop early only if something went wrong. ALL = every unit shares the bath (blank storage, DI test, TLF baths). Individual barcodes are for the serial puck rotations. One running window per barcode at a time; mis-clicks can be deleted.

Loading registry…
Backfill a window that already happened
Record an event (handling, shake, bath change)

Events are drawn as vertical lines on the S-TLF and temperature traces and listed in the registry. They are not scored.

Running windows

Registry (newest first) — …