Skip to content

Intervention log

An append-only record of every occasion on which a human had to intervene in the running service. This is the artefact that the T1.1 operability test is scored from.

It exists because the T1.1 operability test is the only test on the engineering hypotheses page that cannot be measured retrospectively. Every other test can be reconstructed later — from the scenario suite, from MLflow timestamps, from the runbooks, or from billing history. "How many times did a human have to intervene, and why?" is unrecoverable unless it is written down as it happens.

How to add an entry

Append a row to the log below. One row per intervention, newest last. The columns are deliberately few, so that logging an intervention is never the reason an intervention goes unlogged:

Column What goes in it
Date YYYY-MM-DD HH:MM UTC of the intervention, not of the underlying fault. The time of day matters: the T1.1 operability test scores out-of-hours interventions separately, and a date alone cannot be scored for that
Trigger What alerted us — a Sentry alarm, a missed check-in, an NGED email, a routine glance at the Dagster UI
Cause One of the cause categories below
Minutes Human-minutes spent, start to finish, rounded to the nearest 5
Runbook? Did Operating the live service already cover it? yes / partial / no
Notes One sentence. Link the issue or PR if there is one

An intervention is any occasion a human had to intervene in the running service, including occasions that turn out to be trivial. Feature work and deliberate upgrades are not interventions. Unglamorous keep-it-running chores are interventions and belong in routine-ops: a credential rotation, a certificate, or a dependency bump forced by an upstream deprecation.

A run that failed and then recovered on its own retry is not an intervention, but log it anyway, with Minutes = 0 and Cause = self-recovered. Self-recovery is evidence for the design rather than against it, and it is only evidence if somebody wrote it down.

"Runbook?" is not bookkeeping. A gap in operations.md is itself a finding: the T1.4 runbooks-alone operability test claims that staff who did not develop the code can run this service from the runbooks alone, and every no is a point against that claim.

Cause taxonomy

The taxonomy is the substance of the test. The T1.1 operability test predicts that at least 90% of entries fall into upstream-contract — that essentially the only cause that should ever need a human is an upstream provider changing the shape of what they publish. Every entry in another category counts against that 90%, which is exactly why the categories are recorded separately rather than lumped together. The threshold is deliberately not 100%, so a single fluke cannot falsify the claim on its own.

Category Meaning
upstream-contract An upstream provider changed a format, schema, column, unit, or file layout. The one category the T1.1 operability test predicts.
upstream-outage Upstream data absent or stuck for long enough that a human had to act, without the contract itself changing
infrastructure The host, container, scheduler, network, or cloud account — anything below our own code
our-bug A defect in this codebase
model Forecast quality required a human decision — a promotion, a rollback, a retrain
routine-ops Keep-it-running work that needed a human without anything having failed: credential rotation, a certificate, a dependency bump forced by an upstream deprecation
self-recovered Not an intervention at all — a run that failed and recovered on its own retry, logged with Minutes = 0 because self-recovery is evidence for the design

Categories are append-only, like the hypothesis labels themselves. If a cause genuinely does not fit, add a category rather than stretching an existing one — a taxonomy bent to fit the data measures nothing.

The scoring window opens at v1.0

Log everything from day one, but score the T1.1 operability test only from v1.0 onward.

While the system is being actively rebuilt through v0.2–v0.9, a good deal of what looks like an intervention is really development churn, and counting it would spuriously falsify the ≥90% threshold. The pre-v1.0 entries are still worth having: they are what the cause taxonomy is built from, and they are the honest record of what the service actually demanded of us on the way up.

The reverse also holds, and matters more. A quiet pre-v1.0 stretch does not score in favour of H1, a service that mostly runs itself either. It would be the plainest possible case of the selective reading these hypotheses exist to prevent: counting the good weeks of an excluded window while discounting the bad weeks.

The log

Date Trigger Cause Minutes Runbook? Notes
2026-08-13 19:56 Sentry alarm — but from the pre-v0.2 code on a laptop, not from AWS upstream-outage <5 yes Dynamical.org first published the 2026-08-09 00Z ECMWF run with a variable wholly missing, and v0.1 treated that as fatal, so the partition was re-materialised by hand. Four forecast slots ran on the previous day's run in the meantime. #493 added the retry that covers it
2026-09-23 19:00 Building the v0.2.1 image failed our-bug 20 no data/ on the workstation is a symlink to another disk, and the build script's COPY data/production_model/ sent the symlink itself rather than its target, so the build failed with "not found". Worked around with a worktree holding a real copy of the model, then fixed properly in #864, which passes the model to the build as a named build context resolved with realpath
2026-09-23 19:10 apt full-upgrade on the control-plane box lost its SSH session mid-run infrastructure 15 no Upgrading Tailscale restarted tailscaled, which carried the SSH session away and left apt waiting on a needrestart prompt with nobody to answer it. Killed the stuck needrestart process, then finished with dpkg --configure -a and a NEEDRESTART_MODE=a apt full-upgrade inside tmux, so a dropped connection can no longer strand the prompt

Periods covered

An empty table means something only if the periods are recorded, not only the entries. An empty log with no stated period is indistinguishable from a log nobody kept.

Period Version Scope Interventions Scores T1.1, operability?
2026-07-15 18:00 UTC → 2026-08-14 00:00 UTC v0.1 28 time series, 6-hourly live_forecasts on AWS 1 No — pre-v1.0
2026-08-14 00:00 UTC → 2026-09-23 18:00 UTC v0.2 31 time series, 6-hourly live_forecasts on AWS, with live_forecasts_are_healthy reporting on each slot 0 No — pre-v1.0
2026-09-23 18:00 UTC → ongoing v0.2.1 31 time series, 6-hourly live_forecasts on AWS, under a champion retrained on corrected NGED timestamps 2 No — pre-v1.0

Figures for the v0.2 period below are stated as of 08:00 UTC on 28 August 2026, after that day's 06:00 UTC slot. Every count in this section moves within the day, so the as-of instant is part of the measurement rather than a formality.

v0.1 on AWS, 2026-07-15 to 2026-08-13

The first live_forecasts run on AWS was the 18:00 UTC slot on 15 July 2026, and the last was the 18:00 UTC slot on 13 August 2026, an hour before the v0.1 stack was retired at roughly 19:00 UTC that evening to make way for v0.2. Over that window the schedule called for 117 consecutive 6-hourly forecast slots, and every one of them produced a forecast for all 28 time series. One ECMWF run was lost and one human intervention was needed, both described below.

The VM was deployed once, on 15 July, and no code was pushed to AWS until it was retired. The one operator action in the window was the NWP backfill logged above, so the period is close to unattended but not entirely so.

Verified by counting distinct power_fcst_init_time values with fold_id = "live" and experiment_name = "xgboost_cv_0001" — v0.1's promoted model — in the power_forecasts Delta table on S3. All 117 scheduled slots are present, every consecutive pair is exactly 6 hours apart, and all 28 time series appear in every one of the 117.

The 9 August ECMWF run, and what it shows

Dynamical.org publishes each ECMWF run as roughly 40 separate Icechunk commits, so a run can be readable and incomplete at the same time. The 2026-08-09 00Z run was first published with a weather variable wholly missing, and repaired 3 hours 25 minutes later. v0.1 treated a wholly-missing variable as fatal, so its ecmwf_ens run failed and no partition was written for 9 August. The partition was re-materialised by hand on 13 August at 19:56 UTC.

The forecast did not stop. Four slots — 2026-08-09 12:00 UTC through 2026-08-10 06:00 UTC — ran on the 8 August 00Z run instead, at 36, 42, 48 and 54 hours old against the 12–30 hours a healthy slot sees. Every other slot in the window used NWP no older than 30 hours.

That is Principle 1 — the power forecast never stops — working in production rather than on paper, and it is the more interesting result on this page. A missing input degraded the forecast instead of stopping it, and the degradation was bounded and visible after the fact. v0.2 closes the gap that made the intervention necessary at all: a wholly-missing variable is now a retryable "not ready yet" rather than a fatal error, with a 4-hour retry budget that covers the 3h25m this republication took.

Three caveats, without which the window would be worth more than it is:

  • The deployment had no telemetry at all, so nothing on AWS could have alerted. The earliest Sentry commit of any kind — the failure hook, the check-in and the freshness warning all arrived in the same week — lands on main on 21 July, 6 days after the box was deployed and never updated. Whatever was running there therefore predates Sentry entirely: the missed-check-in monitor never existed on that box, and the alarm that did surface the missed run came from newer code running on a laptop. live_forecasts_are_healthy landed later still: that check reads each slot's rows back and counts missed NWP runs. The four degraded slots were reconstructable only because nwp_init_time travels on every forecast row: the degradation was recoverable from the data, but nothing in the deployment announced it.
  • A month is short, and this is the easy case. v0.1 is 28 time series and one ECMWF run per day. An upstream contract change — the dominant cause the T1.1 operability test predicts — did not happen in a window this short; a partial publication is a milder fault than a changed schema.
  • It does not score. The window opens at v1.0, as above.

So what this is, stated plainly: weak, non-scoring evidence for H1, a service that mostly runs itself, drawn from a window the scoring rule excludes. The deployed stack served all 117 scheduled slots over 29 days, absorbed one lost NWP run by degrading rather than stopping, and cost a human about a minute. That is worth recording, and it is not worth more than that.

v0.2 on AWS, from 2026-08-14

v0.2 was deployed on the evening of 13 August 2026, replacing the v0.1 stack at roughly 19:00 UTC. Its first live_forecasts run was the 00:00 UTC slot on 14 August 2026, which is where this period starts. The deployment itself is a deliberate upgrade, so it is not an intervention and has no row in the log.

v0.2 forecasts 31 time series, three more than v0.1, under the promoted model xgboost_cv_0003. From the 00:00 UTC slot on 14 August 2026 to the 06:00 UTC slot on 28 August 2026 the schedule called for 58 consecutive 6-hourly slots. Every one of those slots produced a forecast for all 31 time series. No NWP run was missed: every slot in the window forecast from NWP between 12 and 30 hours old, the healthy band for a once-daily ECMWF run.

Verified by counting distinct power_fcst_init_time values with fold_id = "live" and experiment_name = "xgboost_cv_0003" — v0.2's promoted model — in the power_forecasts Delta table on S3, the same query that verified the v0.1 window. All 58 scheduled slots are present, every consecutive pair is exactly 6 hours apart, all 31 time series appear in every one of the 58, and no slot's nwp_init_time is more than 30 hours before its power_fcst_init_time.

Three changes make the next stretch better evidence than the last. live_forecasts_are_healthy reads each succeeding slot's rows back and reports missed NWP runs, so a slot forecasting from stale inputs is recorded as degraded rather than passing unremarked. A wholly-missing NWP variable is now retried for 4 hours instead of failing, which is what would have made the 9 August intervention unnecessary. And the deployment carries Sentry, which v0.1's never did, so an ecmwf_ens run that does exhaust its retries reports itself from AWS rather than waiting for somebody to run the code on a laptop.

What still would not reach us is the degraded slot, where nothing failed. Dagster runs no check for an asset that raised, so this is the succeeding-run case: live_forecasts_are_healthy returns its warning to the Checks view and sends nothing to Sentry, and the slot's check-in reports the service healthy regardless, so a degraded run looks like a good one from outside — the gap at #501.

v0.2.1 on AWS, from 2026-09-23

v0.2.1 was deployed on the evening of 23 September 2026. Its first live_forecasts run was the 18:00 UTC slot that day, materialised by hand once the power table rebuild below had finished; the next two slots, 00:00 and 06:00 UTC on 24 September, ran unattended on live_forecasts_schedule.

v0.2.1 corrects NGED's power timestamps, which were 30 minutes late before 26 March 2026 (see correcting late power stamps). The fix only takes effect on a rebuild, so power_time_series.delta on S3 was moved aside and power_time_series_and_metadata re-materialised, re-downloading NGED's full history under the corrected code. The champion was retrained on that corrected history — xgboost_baseline_retrain_790 replaces xgboost_cv_0003 — because every model trained before the fix had learnt the 30-minute offset. NGED have no plans to correct their own history at source, so the rebuild carries no risk of double-correcting it.

Verified by reading the power_forecasts Delta table on S3 for fold_id = "live" since 2026-09-22: each of the three v0.2.1 slots forecasts all 31 time series with all 51 ensemble members and no null or NaN values, from NWP between 18 and 30 hours old. No primary key repeats. Forecast values, roughly −97 MW to 400 MW across old and new models alike, sit in the expected range for these feeders. Ensemble-mean error against the actuals received so far is comparable between the outgoing and incoming champions, as expected this soon after a retrain: too few actuals have arrived yet to say whether the retrain improved accuracy.

Two script and box issues surfaced while deploying, both logged in the log above rather than repeated here. Neither reached the running service: both were caught and fixed before the image was pushed or the box was updated.

One stale series was noticed during verification, not logged as an intervention because nothing failed: time series 29's last reading was from 17:30 UTC on 23 September, roughly 15 hours before the check, while every other series was current to within a couple of hours. power_fcst handles this the way it handles any stalled meter — as missing input, degrading rather than raising — so no forecast slot was affected. Worth confirming with NGED whether that gap is expected.

See also