What the drift experiments have established
A synthesis across three campaigns — the burn-in fleet (25 units, 15 days, now closed), the FDOM/Chl-a variant batches, and the vacuum degassing rig — answering the three questions the programme actually turns on: can we control temperature, do bubbles matter, and is there a real burn-in drift. All three serve one goal — correcting temperature and drift well enough to leave a flat baseline — so that outcome is stated first and the diagnosis follows.
Every number and chart below is computed from the archived
sweeps by scripts/build_findings.py. Nothing here is typed in by hand.
The short version
The three findings are not equally solid, and the page says which is which. Temperature is settled. Bubbles are proven as a mechanism but not as the dominant one. Drift is real, smaller than we had been quoting, and not predictable.
The goalDoes the correction give us a flat baseline?
Everything else on this page is diagnosis. This is the outcome. The barrel is the first window where all four batches sit in the same water at the same time, and the 08-20/21 heater-then-chiller ramp is the first that makes temperature and drift separable at all. So it is the first honest test of the only thing that matters: after correcting per sensor for quench, thermal lag and drift, is what remains flat?
Where the model is accepted, the answer is yes. Raw scatter of — falls to — once temperature is removed and to — once drift is removed as well. Temperature does most of the work and drift closes the rest; neither alone is sufficient. (These three ranges are read from the table below rather than quoted: an earlier version of this sentence said “roughly 0.5 %”, which was the barrel page’s narrower water-phase subset, not this page’s own figure.)
And it gets worse outside a controlled ramp. Run over the frozen burn-in campaign (Stage 10) the same model passes 24/24 units in the cold hold and 23/24 in air — windows spanning 0.3 and 3.4 °C, where there is almost nothing to correct — but only 7/25, 5/25 and 12/24 on the wide-swing outdoor windows, which is the condition a deployed sensor actually lives in. Flatness on a bench ramp is necessary, not sufficient.
1Temperature is a physical constant, and it transfers
The sensor’s temperature response is the same coefficient everywhere we have measured it — in water and in air, on two fleets, across two weeks. That makes it correctable in the strong sense: a coefficient measured once can be applied later, elsewhere, in a different medium.
bT per unit, by condition
Each dot is one sensor’s own coefficient, from repeated short windows. The four conditions sit on top of each other.
Do units actually differ from each other?
Spread between units against each unit’s own scatter. Below the line, per-unit calibration is fitting noise.
What follows practically.
Quench is a property of the individual sensor, and is fitted per sensor — from
that unit’s own sipm_temp_c, over the window being analysed. Everything on this page is
now built that way; there is deliberately no fleet constant anywhere in the pipeline, because a pooled
coefficient silently attributes one sensor’s response to another.
An earlier version of this section argued the opposite — that a single fleet number was
“the right choice” because between-unit spread came out smaller than within-unit scatter. That
was a statement about measurement quality in the windows available at the time, not about the
instrument. Those windows spanned 1.5–3 °C, where a per-unit coefficient is mostly noise,
and the conclusion did not survive contact with a proper temperature sweep.
The condition that matters is temperature range, not fleet size. Fit per sensor wherever
the window supports it; where it does not, the answer is not to borrow a neighbour’s coefficient but
to recognise that nothing is identifiable and nothing needs correcting — a window spanning
under 0.5 °C carries less than 1 % of quench at a typical −1.7 %/°C, which is
why the constant-temperature holds are simultaneously the least correctable and the most trustworthy
drift measurements we have. In air the measurement is roughly seven times cleaner than in water, so
an air window is the best place to characterise a unit’s own quench, and the
08-20/21 ramp (36.7 → 9.9 °C, ~27 °C on 44 units) is the first sweep wide
enough to determine it properly for every sensor at once.
2Bubbles are real, and rare
Pulling a vacuum on a sealed container of DI water grows something on the optical window; shaking the sensor sheds it. Both channels move together in both directions, which is what distinguishes a window effect from an electronics one. 500226 reached 13.4× its baseline and returned to 0.93× the moment it was shaken.
Why the reversal is the evidence, not the rise. A rise on its own is ambiguous: drift, ageing, fouling and contamination all raise a signal. None of them undo themselves when the unit is shaken. Temperature is excluded by three orders of magnitude — the SiPM moved ≤0.17 °C across the vacuum, worth ≤0.4 % at −1.7 %/°C, against +1238 %.
2bThe degassing rig — bubbles applied on purpose
500146 is the exception and its record ends early. It stopped producing sweeps at 2026-08-20 22:42Z and has 376 rows against 1,603–1,678 for the other four. It is not a sync backlog: a re-pull recovered exactly one further reading, at 08-21 19:35Z, and that reading is in air — ToF 141 against its own in-tank median of 55 (range 52–66), while all four peers still read 44–48. The unit left the water; its degassing record legitimately ends at 22:42Z and the late reading is not part of the experiment. The
degas cohort is marked frozen in
scripts/export_burnin_sweeps.js, so it is no longer refreshed and every number
below is stable; re-pulling needs an explicit --thaw. Sealed-tank readings that
buffered on-device are included up to that pull.Merged in from its own page 2026-08-21. Five units — 500220, 500225, 500226, 500108, 500146 — in a sealed pressure container of DI water, in at 2026-08-20 09:29 local with the vacuum applied at 10:50, 81 min later. This is the interventional half of section 2: the suspected cause is applied deliberately, at a known instant, to units that are otherwise untouched.
The testBubbles, applied on purpose
The record
This page's data begins at the test (09:29), so there is no dry-air record in it and no air→water entry step to quote — that transition lives on the burn-in page. Each unit is normalised to its own median over the 30 min before the vacuum, so 1.0 is the settled-in-water baseline and the vacuum step reads directly as a ratio. Rows the out-of-water detector flags are kept: at immersion that flag is the transition itself.
sTLF — normalised to the pre-vacuum baseline
Full-sweep amplitude (tlf_amp). Watch for steps, not slope.
ToF — the corroborating channel
Signal per SPAD. A ToF move that lines up in time with an sTLF move points at the window; an sTLF move alone does not.
SiPM temperature
A vacuum drives evaporative cooling and sTLF is temperature-quenched, so this is what separates a real optical change from quench. It did not move enough to matter here — at most 0.17 °C across the vacuum, worth ≤0.4 % against excursions of tens to hundreds of percent.
Battery
Pack voltage. These units are on battery, so this sets how long the test can run.
Per-unit vacuum step and current state
Event registry & reading notes
What was done to these five units, what each intervention tests, and how the interpretation changed as the cycles accumulated. The charts and table above are the record; this is the reasoning behind it.
Handling a submerged sensor moves sTLF by roughly +14 %; the same handling in air moves it by a median of −0.06 % (n=22, measured on the burn-in fleet). That asymmetry points at bubbles on the optical window, but only circumstantially, because nothing in it controlled the gas. Pulling a vacuum does: dropping the pressure drops the solubility of dissolved air, which leaves solution and nucleates preferentially on surfaces, the window among them. This test moves the bubble hypothesis from observational to interventional — the suspected cause is applied deliberately, at a known instant, to units that are otherwise untouched.
The 81 min between immersion and vacuum is the design, not a delay: it gives each unit a quiet, undisturbed water baseline, so a move at the vacuum line is attributable to the gas rather than to handling or to settling after immersion.
The reversal is what makes this conclusive. A rise on its own is ambiguous — drift, ageing, fouling and contamination can all raise a signal. None of them undo themselves when you shake the unit. Something physically sat on the optical window, was grown by lowering the pressure, and was knocked off by agitation, moving the fluorescence and the IR return together. That is a bubble.
Temperature cannot account for it: the SiPM moved at most 0.17 °C across the vacuum, worth ≤0.4 % at the project’s −2.4 %/°C quench slope, against a +1238 % excursion.
What it does not yet establish is prevalence. Only two of five units responded; 500108 and 500220 stayed flat throughout, and 500146 moved on sTLF alone with no ToF response and then rose further after the shake (1.42× → 1.93×), which is not a window effect and still needs its own explanation. So bubbles are demonstrated as a mechanism, not established as the mechanism for every handling step.
But 500226’s ToF never moved. Flat at 1.00–1.02 through a sevenfold change in sTLF. Our stated rule was that both channels must move for a step to count as a window effect — and by that rule this, the clearest bubble event in the programme, would have been rejected. The rule is wrong as a necessary condition. The ToF and the fluorescence optics do not share an aperture, so a bubble can sit in one path and not the other. Coupling remains strong evidence for a bubble; its absence is not evidence against one. Reversibility on pressure release is the better discriminator, because it keys on the gas physics rather than on where the bubble happens to sit.
Note on the window: the effect grows under sustained vacuum rather than stepping instantly, so the original ±15 min figure (median +0.9 %) badly understates it and is kept below only as the prompt response. The peak-under-vacuum column is the one to read.
3Drift is real, batch-dependent, and mostly water-side
Most published drift numbers from this programme do not survive a basic reliability check. Splitting each window in half and re-fitting, five of six windows disagree with themselves by more than the drift they report.
Cut instead so that every window is bounded by an event and opened 2 h after it, — of — windows are separable (time and temperature not collinear) and drift becomes measurable per sensor. The two that are not separable are the post-UV window and the whole barrel tank phase — which is exactly why the barrel page declines to quote a rate there. Tables below.
“Not predictable” still stands and is a different claim. What follows shows drift is consistently estimable across units at a point in time. The extrapolation test further down measures whether a fitted slope persists forward in time, and it does not. Consistent across units is not the same as persistent over time; neither is evidence about the other.
Per sensor, air against water — windows with no intervention in them
Temperature is fitted jointly per unit, so quench cannot leak into the time slope. w is how many separable windows the unit contributed; rel how many of those pass the split-half check. Immersion is taken as 16:45Z, verified from the data — ToF steps 140 → 46 kcps in all three batches at once. Reading the prose local time as UTC puts six hours of air data in the water column; local here is UTC−6.
Per-sensor drift — all units, air and water
The controlled comparison, and the batch spread
The burn-in fleet is the only batch measured in both media, so its two rows are the same sensors under the two conditions — the cleanest statement available about where drift comes from. Air is the only medium in which all four batches have data.
Which windows carry a number, and why
A joint fit only means something where time and temperature are separable inside the window. |corr| above 0.9 means the two regressors are collinear and the split between them is arbitrary, so the window is reported as not separable rather than quoted.
What drift actually looks like
The same 24 burn-in units in both panels, 20 h apart, each temperature-corrected and normalised to its own first settled hour, then binned on hours-since-start so the two media share an axis. The line is the fleet median, the band the interquartile range. Because it is one fleet measured twice, the difference between the panels is the medium, not the hardware.
Water
Air
Read the bands, not just the lines. In stable water the fleet loses about a percent and a half over ~44 h and the band stays narrow — units move together, which is what a real shared process looks like. In air over a comparable span the median barely moves and the band straddles zero: some units rise, some fall. That is the difference between a measurable drift and no measurable drift, on the same instruments.
And what an uncontrolled window looks like
Same units, in water, but riding a 15 °C diurnal swing. The fleet median ends 27 % down while the fitted rate comes out at −0.01 %/day and the per-unit IQR spans −8.3 to +5.5 %/day — the daily cycle dominates and no drift number can be recovered from it. This is the shape most of our historical drift figures were fitted to.
Bars are the drift each window reports; the marker is how far the two halves of that same window disagree. Where the marker exceeds the bar, the number is not measuring anything.
Even a good window decays
Successive 6 h blocks inside the one reliable hold. The early blocks are settling, not drift — a short hold overstates the rate.
Extrapolating drift makes prediction worse
Fit a slope on past data, predict the future. Compared with simply assuming the last baseline holds.
What to do instead: re-baseline
If drift cannot be predicted, the only control is a recent baseline. This is how fast one goes stale — error against a 6 h baseline, as a function of the baseline’s age.
A baseline is good for roughly 6–12 h at the few-percent level. Beyond a day it is worth more than ten percent of the reading, which is the entire signal range for many deployments.
4The four batches — where they agree and where they don’t
Three optics builds from the Aug-2026 variant run (FDOM, Chl-a, TLF) plus the older burn-in fleet. Comparing them separately matters because a fleet-wide number can hide a build that behaves differently — and on temperature, one of them does.
Temperature response is not common across builds
Each batch’s coefficient is tight internally — interquartile ranges of about 0.1–0.2 %/°C — but the batches sit apart from one another. FDOM quenches roughly 2.5× harder than Chl-a.
Mann-Whitney on the per-unit coefficients: FDOM differs from Chl-a (p = 0.0012), from new TLF (p = 0.0082) and from the burn-in fleet (p = 0.0006); Chl-a differs from burn-in (p = 0.0130). Chl-a vs new TLF is the one pair that is not separable (p = 0.064).
Drift in air, by batch
Same window for all four (08-19 → 08-20, in air, temperature-corrected), so the comparison is like-for-like. The reliability column is the split-half check — how far the two halves of the same window disagree, relative to the drift being reported.
Read this one cautiously. The medians differ — Chl-a steepest at −2.99 %/day, the burn-in fleet nearly flat at −0.28 — but only FDOM comes close to passing the split-half check, and the batches are confounded with age (the variant builds were 1–2 days old, the burn-in fleet 13–15). On the evidence so far this is a difference worth chasing, not a difference established.
Response to the barrel shake, by batch
All 44 units shaken at one instant in shared water, temperature-corrected. This is where the batches agree: nothing moved.
Every batch median sits within 2 % of zero and no unit in any batch moved more than 5 % on both channels. Whatever differences exist between these builds, bubble susceptibility in freshly-poured DI is not one of them.
6Offset and gain are still required — the corrections do not replace them
Four corrections now sit between a raw sweep and a number, and it is worth being precise about which job each one does, because they are not interchangeable and none of them substitutes for the others.
| Step | What it removes | Scope |
|---|---|---|
| Amplitude estimator | dependence on which cells survived screening — railing, pedestal pinning, cell dropout | one sensor, one sweep |
| Quench | temperature, on that unit’s own measured coefficient | one sensor, over temperature |
| Drift | that unit’s change over time | one sensor, over time |
| Offset + gain | the unit’s own sensitivity — counts per ppb | across sensors, absolute scale |
Only the last one makes two sensors comparable in absolute terms, and the amplitude does not
do it. The amplitude is exp(c − c0), a ratio to each unit’s
own reference sweep. That normalises a unit against itself — which is exactly what makes
it railing-proof — and leaves every unit reading ≈1 at its own reference state regardless of
how many counts per ppb it actually delivers.
conc = a + b·amp per sensor on
the 2026-07-16 dilution ladder, 13 sensors: gain spans 10.6× (92.8–983.1) and the
blank 2.2×. Across the wider burn-in fleet the gain spans 27×. Replacing
per-sensor offset and gain with one shared pair takes R² from 0.815 to 0.120 and RMSE
from 7.14 to 15.58 ppb — barely better than predicting the mean. Reproduce with
scripts/dilution_feature_compare.mjs.
The blanks fall into two groups — roughly 0.52–0.64 and 0.96–1.12 — which tracks the two-tier gain structure (~93–118 against ~404–983). That is a hardware population difference, not something any correction removes.
Where this bites operationally. The FDOM, Chl-a and new-TLF batches carry no dilution calibration at all — they were never run through a ladder — so they can be compared in shape but never in concentration, and no amount of correction closes that gap. One unit, 50091, has a negative fitted gain (−0.00054 at R² 0.30): more tryptophan reading as less signal, which is not a calibration.
offsetC / gainC are fitted on the ladder’s
absolute corrected amplitude (blank ≈ 1.38), while /barrel plots a
baseline-relative amplitude that sits at ≈ 1 by construction. Subtracting one
from the other put 69 % of the plotted points below zero (median
−9.3 ppb, floor −106 ppb). A calibration is only valid against a signal
carried on the same basis it was fitted on.
mon2 it is level (pooled R² 0.9087 vs 0.9097 on the 12 sensors
both fit). It is robustness: it fits 13 of 14 sensors where mon2 fits 12, and
beats the superseded slope on 11 of 13 with the gap concentrated on the hard units —
50091 0.4835 → 0.8131, 500193 0.8713 → 0.9260, 50062
0.8232 → 0.8700. On well-behaved units the two tie. The estimator does not make good
sweeps better; it stops bad ones from being wrong.
Practical consequence: a new sensor still needs a dilution ladder before its readings mean anything in ppb, and a repaired or re-capped one needs a fresh one. What has changed is that the calibration can now be derived on a better-conditioned feature.
5Which units look like hardware, not modelling
Every unit in the barrel is fitted and reported — none is dropped for fitting badly. But some of what is left over is the sensor, not the model, and those units need bench attention rather than a better correction. This section separates the two.
The axes, and what each one implicates
A unit is listed when it exceeds 4 MAD above its own batch on an axis, or trips an absolute fault rule. Tier is by convergent evidence: three or more independent axes agreeing is rework, two is investigate, one is watch. Independence is the point — one extreme number is usually a window artefact, three pointing at the same unit is not.
What is still open
- Does drift depend on sensor age? It looks like it — new units drifted −2.1 %/day against −0.1 for 13-day-old ones — but holding optics fixed collapses the gap to −0.49 %/day on five units spanning −4.6 to +2.0. Not established.
- Do FDOM and Chl-a drift more than TLF? Medians say −2.10, −3.10 and −0.62 %/day, but all pairwise tests come back non-significant (p = 0.10–0.54) and none of the three windows passes the split-half check.
- Is there any drift in air at all? The burn-in fleet’s 44 h air IQR spans zero. Air is too thermally quiet and too short to prove presence or absence.
- What moves sTLF without moving ToF? Seven units in the barrel shake and 500146 on the degassing rig moved on the signal channel alone. Not a window effect; unexplained.
- What happens past two weeks? Every reliable measurement comes from days 10–12 of one campaign. Any 50-day projection is a four-fold extrapolation beyond the data.
- Are the other four new-TLF units sound? With 500156 pulled for rework, the remaining four still sit above every other batch on corrected residual. That could be the build, or it could be that a 27 °C ramp is a hostile window to fit in. A steady hold separates them; nothing available now does.
- Can the Chl-a batch be assessed at all in DI? Those ten units return no amplitude on about 71 % of readings because DI water has no chlorophyll in it. That is the correct answer to the question they were asked, and it means this campaign cannot clear or condemn them. They need water containing the analyte they measure.
All five are limited by the same thing: temperature and elapsed time have moved together in almost every window we have. The ramp-and-hold now running on the barrel is the first design that separates them — and how long the hold runs decides whether it measures drift or the settling transient on top of it.