Intervention log
An append-only record of every occasion on which a human had to intervene in the running service. This is the artefact that the T1.1 operability test is scored from.
It exists because the T1.1 operability test is the only test on the engineering hypotheses page that cannot be measured retrospectively. Every other test can be reconstructed later — from the scenario suite, from MLflow timestamps, from the runbooks, or from billing history. "How many times did a human have to intervene, and why?" is unrecoverable unless it is written down as it happens.
How to add an entry
Append a row to the log below. One row per intervention, newest last. The columns are deliberately few, so that logging an intervention is never the reason an intervention goes unlogged:
| Column | What goes in it |
|---|---|
| Date | YYYY-MM-DD HH:MM UTC of the intervention, not of the underlying fault. The time of day matters: the T1.1 operability test scores out-of-hours interventions separately, and a date alone cannot be scored for that |
| Trigger | What alerted us — a Sentry alarm, a missed check-in, an NGED email, a routine glance at the Dagster UI |
| Cause | One of the cause categories below |
| Minutes | Human-minutes spent, start to finish, rounded to the nearest 5 |
| Runbook? | Did Operating the live service already cover it? yes / partial / no |
| Notes | One sentence. Link the issue or PR if there is one |
An intervention is any occasion a human had to intervene in the running service, including
occasions that turn out to be trivial. Feature work and deliberate upgrades are not interventions.
Unglamorous keep-it-running chores are interventions and belong in routine-ops: a credential
rotation, a certificate, or a dependency bump forced by an upstream deprecation.
A run that failed and then recovered on its own retry is not an intervention, but log it anyway,
with Minutes = 0 and Cause = self-recovered. Self-recovery is evidence for the design rather
than against it, and it is only evidence if somebody wrote it down.
"Runbook?" is not bookkeeping. A gap in operations.md is itself a finding: the
T1.4 runbooks-alone operability
test claims
that staff who did not develop the code can run this service from the runbooks alone, and every no
is a point against that claim.
Cause taxonomy
The taxonomy is the substance of the test. The T1.1 operability
test predicts
that at least 90% of entries fall into upstream-contract — that essentially the only cause
that should ever need a human is an upstream provider changing the shape of what they publish. Every
entry in another category counts against that 90%, which is exactly why the categories are recorded
separately rather than lumped together. The threshold is deliberately not 100%, so a single fluke
cannot falsify the claim on its own.
| Category | Meaning |
|---|---|
upstream-contract |
An upstream provider changed a format, schema, column, unit, or file layout. The one category the T1.1 operability test predicts. |
upstream-outage |
Upstream data absent or stuck for long enough that a human had to act, without the contract itself changing |
infrastructure |
The host, container, scheduler, network, or cloud account — anything below our own code |
our-bug |
A defect in this codebase |
model |
Forecast quality required a human decision — a promotion, a rollback, a retrain |
routine-ops |
Keep-it-running work that needed a human without anything having failed: credential rotation, a certificate, a dependency bump forced by an upstream deprecation |
self-recovered |
Not an intervention at all — a run that failed and recovered on its own retry, logged with Minutes = 0 because self-recovery is evidence for the design |
Categories are append-only, like the hypothesis labels themselves. If a cause genuinely does not fit, add a category rather than stretching an existing one — a taxonomy bent to fit the data measures nothing.
The scoring window opens at v1.0
Log everything from day one, but score the T1.1 operability test only from v1.0 onward.
While the system is being actively rebuilt through v0.2–v0.9, a good deal of what looks like an intervention is really development churn, and counting it would spuriously falsify the ≥90% threshold. The pre-v1.0 entries are still worth having: they are what the cause taxonomy is built from, and they are the honest record of what the service actually demanded of us on the way up.
The reverse also holds, and matters more. A quiet pre-v1.0 stretch does not score in favour of H1, a service that mostly runs itself either. It would be the plainest possible case of the selective reading these hypotheses exist to prevent: counting the good weeks of an excluded window while discounting the bad weeks.
The log
| Date | Trigger | Cause | Minutes | Runbook? | Notes |
|---|---|---|---|---|---|
| 2026-08-13 19:56 | Sentry alarm — but from the pre-v0.2 code on a laptop, not from AWS | upstream-outage |
<5 | yes | Dynamical.org first published the 2026-08-09 00Z ECMWF run with a variable wholly missing, and v0.1 treated that as fatal, so the partition was re-materialised by hand. Four forecast slots ran on the previous day's run in the meantime. #493 added the retry that covers it |
| 2026-09-23 19:00 | Building the v0.2.1 image failed | our-bug |
20 | no | data/ on the workstation is a symlink to another disk, and the build script's COPY data/production_model/ sent the symlink itself rather than its target, so the build failed with "not found". Worked around with a worktree holding a real copy of the model, then fixed properly in #864, which passes the model to the build as a named build context resolved with realpath |
| 2026-09-23 19:10 | apt full-upgrade on the control-plane box lost its SSH session mid-run |
infrastructure |
15 | no | Upgrading Tailscale restarted tailscaled, which carried the SSH session away and left apt waiting on a needrestart prompt with nobody to answer it. Killed the stuck needrestart process, then finished with dpkg --configure -a and a NEEDRESTART_MODE=a apt full-upgrade inside tmux, so a dropped connection can no longer strand the prompt |
Periods covered
An empty table means something only if the periods are recorded, not only the entries. An empty log with no stated period is indistinguishable from a log nobody kept.
| Period | Version | Scope | Interventions | Scores T1.1, operability? |
|---|---|---|---|---|
| 2026-07-15 18:00 UTC → 2026-08-14 00:00 UTC | v0.1 | 28 time series, 6-hourly live_forecasts on AWS |
1 | No — pre-v1.0 |
| 2026-08-14 00:00 UTC → 2026-09-23 18:00 UTC | v0.2 | 31 time series, 6-hourly live_forecasts on AWS, with live_forecasts_are_healthy reporting on each slot |
0 | No — pre-v1.0 |
| 2026-09-23 18:00 UTC → ongoing | v0.2.1 | 31 time series, 6-hourly live_forecasts on AWS, under a champion retrained on corrected NGED timestamps |
2 | No — pre-v1.0 |
Figures for the v0.2 period below are stated as of 08:00 UTC on 28 August 2026, after that day's 06:00 UTC slot. Every count in this section moves within the day, so the as-of instant is part of the measurement rather than a formality.
v0.1 on AWS, 2026-07-15 to 2026-08-13
The first live_forecasts run on AWS was the 18:00 UTC slot on 15 July 2026, and the last was the
18:00 UTC slot on 13 August 2026, an hour before the v0.1 stack was retired at roughly 19:00 UTC
that evening to make way for v0.2. Over that window the schedule called for 117 consecutive 6-hourly
forecast slots, and every one of them produced a forecast for all 28 time series. One ECMWF run
was lost and one human intervention was needed, both described below.
The VM was deployed once, on 15 July, and no code was pushed to AWS until it was retired. The one operator action in the window was the NWP backfill logged above, so the period is close to unattended but not entirely so.
Verified by counting distinct power_fcst_init_time values with fold_id = "live" and
experiment_name = "xgboost_cv_0001" — v0.1's promoted model — in the power_forecasts Delta table
on S3. All 117 scheduled slots are present, every consecutive pair is exactly 6 hours apart, and all
28 time series appear in every one of the 117.
The 9 August ECMWF run, and what it shows
Dynamical.org publishes each ECMWF run as roughly 40 separate Icechunk commits, so a run can be
readable and incomplete at the same time. The 2026-08-09 00Z run was first published with a weather
variable wholly missing, and repaired 3 hours 25 minutes later. v0.1 treated a wholly-missing
variable as fatal, so its ecmwf_ens run failed and no partition was written for 9 August. The
partition was re-materialised by hand on 13 August at 19:56 UTC.
The forecast did not stop. Four slots — 2026-08-09 12:00 UTC through 2026-08-10 06:00 UTC — ran on the 8 August 00Z run instead, at 36, 42, 48 and 54 hours old against the 12–30 hours a healthy slot sees. Every other slot in the window used NWP no older than 30 hours.
That is Principle 1 — the power forecast never stops — working in production rather than on paper, and it is the more interesting result on this page. A missing input degraded the forecast instead of stopping it, and the degradation was bounded and visible after the fact. v0.2 closes the gap that made the intervention necessary at all: a wholly-missing variable is now a retryable "not ready yet" rather than a fatal error, with a 4-hour retry budget that covers the 3h25m this republication took.
Three caveats, without which the window would be worth more than it is:
- The deployment had no telemetry at all, so nothing on AWS could have alerted. The earliest
Sentry commit of any kind — the failure hook, the check-in and the freshness warning all arrived
in the same week — lands on
mainon 21 July, 6 days after the box was deployed and never updated. Whatever was running there therefore predates Sentry entirely: the missed-check-in monitor never existed on that box, and the alarm that did surface the missed run came from newer code running on a laptop.live_forecasts_are_healthylanded later still: that check reads each slot's rows back and counts missed NWP runs. The four degraded slots were reconstructable only becausenwp_init_timetravels on every forecast row: the degradation was recoverable from the data, but nothing in the deployment announced it. - A month is short, and this is the easy case. v0.1 is 28 time series and one ECMWF run per day. An upstream contract change — the dominant cause the T1.1 operability test predicts — did not happen in a window this short; a partial publication is a milder fault than a changed schema.
- It does not score. The window opens at v1.0, as above.
So what this is, stated plainly: weak, non-scoring evidence for H1, a service that mostly runs itself, drawn from a window the scoring rule excludes. The deployed stack served all 117 scheduled slots over 29 days, absorbed one lost NWP run by degrading rather than stopping, and cost a human about a minute. That is worth recording, and it is not worth more than that.
v0.2 on AWS, from 2026-08-14
v0.2 was deployed on the evening of 13 August 2026, replacing the v0.1 stack at roughly 19:00 UTC.
Its first live_forecasts run was the 00:00 UTC slot on 14 August 2026, which is where this period
starts. The deployment itself is a deliberate upgrade, so it is not an intervention and has no row
in the log.
v0.2 forecasts 31 time series, three more than v0.1, under the promoted model xgboost_cv_0003.
From the 00:00 UTC slot on 14 August 2026 to the 06:00 UTC slot on 28 August 2026 the schedule
called for 58 consecutive 6-hourly slots. Every one of those slots produced a forecast for all 31
time series. No NWP run was missed: every slot in the window forecast from NWP between 12 and 30
hours old, the healthy band for a once-daily ECMWF run.
Verified by counting distinct power_fcst_init_time values with fold_id = "live" and
experiment_name = "xgboost_cv_0003" — v0.2's promoted model — in the power_forecasts Delta table
on S3, the same query that verified the v0.1 window. All 58 scheduled slots are present, every
consecutive pair is exactly 6 hours apart, all 31 time series appear in every one of the 58, and no
slot's nwp_init_time is more than 30 hours before its power_fcst_init_time.
Three changes make the next stretch better evidence than the last. live_forecasts_are_healthy
reads each succeeding slot's rows back and reports missed NWP runs, so a slot forecasting from stale
inputs is recorded as degraded rather than passing unremarked. A wholly-missing NWP variable is now
retried for 4 hours instead of failing, which is what would have made the 9 August intervention
unnecessary. And the deployment carries Sentry, which v0.1's never did, so an ecmwf_ens run that
does exhaust its retries reports itself from AWS rather than waiting for somebody to run the code
on a laptop.
What still would not reach us is the degraded slot, where nothing failed. Dagster runs no check
for an asset that raised, so this is the succeeding-run case: live_forecasts_are_healthy returns
its warning to the Checks view and sends nothing to Sentry, and the slot's check-in reports the
service healthy regardless, so a degraded run looks like a good one from outside — the gap at
#501.
v0.2.1 on AWS, from 2026-09-23
v0.2.1 was deployed on the evening of 23 September 2026. Its first live_forecasts run was the
18:00 UTC slot that day, materialised by hand once the power table rebuild below had finished; the
next two slots, 00:00 and 06:00 UTC on 24 September, ran unattended on live_forecasts_schedule.
v0.2.1 corrects NGED's power timestamps, which were 30 minutes late before 26 March 2026 (see
correcting late power stamps). The fix only takes effect on a
rebuild, so power_time_series.delta on S3 was moved aside and power_time_series_and_metadata
re-materialised, re-downloading NGED's full history under the corrected code. The champion was
retrained on that corrected history — xgboost_baseline_retrain_790 replaces xgboost_cv_0003 —
because every model trained before the fix had learnt the 30-minute offset. NGED have no plans to
correct their own history at source, so the rebuild carries no risk of double-correcting it.
Verified by reading the power_forecasts Delta table on S3 for fold_id = "live" since
2026-09-22: each of the three v0.2.1 slots forecasts all 31 time series with all 51 ensemble
members and no null or NaN values, from NWP between 18 and 30 hours old. No primary key repeats.
Forecast values, roughly −97 MW to 400 MW across old and new models alike, sit in the expected
range for these feeders. Ensemble-mean error against the actuals received so far is comparable
between the outgoing and incoming champions, as expected this soon after a retrain: too few
actuals have arrived yet to say whether the retrain improved accuracy.
Two script and box issues surfaced while deploying, both logged in the log above rather than repeated here. Neither reached the running service: both were caught and fixed before the image was pushed or the box was updated.
One stale series was noticed during verification, not logged as an intervention because nothing
failed: time series 29's last reading was from 17:30 UTC on 23 September, roughly 15 hours before
the check, while every other series was current to within a couple of hours. power_fcst handles
this the way it handles any stalled meter — as missing input, degrading rather than raising — so
no forecast slot was affected. Worth confirming with NGED whether that gap is expected.
See also
- Engineering Hypotheses → H1, a service that mostly runs itself — the hypothesis this log scores, and the other three tests that sit alongside the T1.1 operability test.
- Operating the live service — the runbooks whose coverage the
Runbook?column measures. - Inherent Stability — the design that is meant to keep this log short.