Extending the training history with ERA5
Status: ๐ง Planned (v0.5). Epic: #145. Ingest: #143; pre-training experiments: #167.
Our ECMWF ENS archive starts 2024-04-01; most trial-area power series go back to late 2019. ERA5 covers the gap, and shares the ENS's IFS lineage, so the reanalysis-to-forecast domain shift is smaller than a different-model reanalysis would give.
Almost all the cost is in the data layer. Once ERA5 and the paired residual statistics exist, moving between the variants below is mostly configuration โ so they are leaderboard arms, not decisions to make up front.
Scope the ingest to include the 2024+ overlap
Fetch ERA5 across the 2024+ ENS overlap as well as the 2020โ2023 gap. This is the one design decision that changes what the ingest builds, and it exists to avoid era confounding: a feature value that occurs in only one time period is learned as a proxy for that period, not for the underlying quantity we mean by it. If NaN-NWP appears only before 2024, "weather is missing" becomes a perfect proxy for "2020โ2023" โ a demand regime carrying COVID distortions, less embedded PV, and lower EV and heat-pump penetration. So a 2027 feed failure would be forecast from a 2021 regime. The same trap applies to a source flag, to a lead-time-zero encoding, and to the ENS-only spread and quantile columns, which have no ERA5 equivalent.
The general rule: an era covariate is safe exactly when the value production will see is well-represented in the modern era. So, alongside the overlap fetch, randomly mask NWP features and spread/quantile columns on a subset of 2024+ rows, as a configurable augmentation step.
The overlap has a second payoff: paired ERA5 and ENS on identical target times, which turns the reconciliation question below into estimation rather than guesswork.
Reconciling ERA5 with ENS
-
Lead-time-zero framing. Do not degrade; treat ERA5 as a forecast at lead zero and let the lead-time feature carry the discounting. That framing separates the physical weather-to-power response (genuinely lead-time-invariant) from how far to trust it as forecast error grows. Cheapest arm and the right first run, but note the tension: pre-2024 rows then carry only lead zero, so every split on lead time partitions the modern rows off beneath it, and the extra history reaches the 3โ10 day band only through the structure above such splits and in the trees that never make one. Boosting shares more across trees than that phrasing might suggest, so this is a weakening rather than a wall โ but expect the win at short leads unless the invariance assumption holds strongly.
-
Degrade ERA5 towards ENS error statistics. Fit the
ENS โ ERA5residual distribution per variable, per lead time, per season on the overlap, then sample from it when synthesising pre-2024 features. Quantile mapping per horizon is the cheap version and is probably enough for temperature. Degrading towards ENS error statistics is not merely an alternative to lead-zero framing: it is what makes the extra history populate the long leads at all. -
ENS reforecasts โ considered and rejected. Under Cycle 49r1 the medium-range reforecasts run over the past 20 years with an 11-member ensemble, so they are real forecasts with real lead-time error and the mismatch would largely disappear rather than needing correction. We are not going to do this: the only access is MARS, and the download would take far too long. Recorded so it is not re-litigated.
ERA5 splits one horizon into two
Today nwp_lead_time_hours (how old the weather is) and the forecast horizon (which power lags are
available) differ only by the constant NWP_PUBLICATION_DELAY_HOURS, so one column carries both.
ERA5 decouples them: weather age is zero. But power-lag availability must still mirror production,
or pre-training teaches the model to lean on lags that vanish at serve time. Pre-training rows
therefore need a sampled pseudo-horizon driving _nullify_leaky_lags, carried separately from
weather age.
Single pool vs two-phase warm start
-
Single pool, with per-era sample weights. All rows in one training set, and the weight on pre-2024 rows becomes a tunable hyperparameter rather than a yes/no decision. Run this first: it is the cheap form of the mixed warm start below, and recency sample weights are already a Tier-1 item on the XGBoost improvements page.
-
Two-phase warm start. Train on the ERA5 history, then continue boosting on 2024+ ENS data. Warm start only adds trees, so the correcting trees see roughly 2 years and very few examples of each season. And if phase one over-trusts weather, shrinking an over-confident component additively is harder than never building it. Mixed phase two โ keeping down-weighted (and possibly degraded) ERA5 rows in phase two โ is the middle path.
-
The source flag follows from that choice, not the other way round. In a pure phase two the flag has zero variance and XGBoost can never split on it, so drop it there. Under a single pool or a mixed phase two it has variance and earns its place โ but only because the overlap fetch decorrelates it from date.
Era covariates
The 2020โ2026 span contains regime changes the weather cannot explain, in two shapes.
Smooth trends โ EV and heat-pump uptake, embedded PV build-out, the 2022 price shock and the Demand Flexibility Service. Handle these with recency sample weights, not a date ordinal: trees extrapolate flat, so a date feature always sits beyond its training range at inference. The init-time-anchored features absorb level drift for the same reason.
COVID lockdowns are a pulse, and the case for a dedicated feature:
-
A lockdown scalar in \([0, 1]\) passes the era-covariate safety rule by construction: production always sees 0, and 0 is abundant in the modern era. That safety pattern is the opposite of the NaN-NWP case, and the confounding is benign โ the feature is the mechanism by which the model quarantines the anomalous period.
-
Source it rather than hand-coding dates, and prefer mobility to stringency. The Oxford COVID-19 Government Response Tracker publishes a daily UK stringency index (0โ100); Google's COVID-19 Community Mobility Reports sit closer to the causal driver of substation demand, and capture both the voluntary March-2020 withdrawal and the slow return through 2021โ22. Both series ended in 2022, so check they are still downloadable โ though a scalar that reads 0 for every future forecast makes a dead source a back-fill problem, not a serving problem. The measured evidence favours mobility, thinly: Chen et al. (2020) take UK national mean absolute percentage error from 10.11% to 8.74% by feeding mobility data into a day-ahead neural network, on a two-week test window, in an arXiv preprint. Retraining on pandemic data without mobility, though, made the UK figure worse, at 13.78% โ so retraining across the lockdown is not safe on its own. The stringency index turned up in our search only in explanatory econometrics, and there Berezvai et al. (2022) needed a quadratic specification, which one linear scalar cannot express.
-
Check whether simply adapting faster does the same job, before building the covariate. de Vilmarest and Goude (2021) compare a Kalman filter handed the break date against one merely allowed to adapt faster everywhere with no break date at all. The no-break version won on three of four model families, and learning the variances rather than fixing them matched it. The recency sample weights above are the same idea. Run that arm first: the covariate has to beat faster adaptation, not merely beat doing nothing.
-
The pre-lockdown regime may never come back, and a scalar that returns to 0 says it does. Prabowo et al. (2023), on 13 building complexes in Melbourne, report distribution shifts during lockdown "which do not fully revert to their pre-lockdown state even after restrictions are lifted". If GB substation demand behaved the same way โ and permanent home-working makes that plausible โ the post-lockdown era is a third regime rather than a return to the first. A lockdown scalar cannot say so, because it reads 0 both before 2020 and after 2021, for two different worlds. Recency sample weights can, which is a second reason to run them as the control arm.
-
One published result supports the plan above: treat the lockdown as a labelled example rather than as data to discard. Abรฉlรจs et al. (2024) calibrate a slow process-noise variance on pre-COVID data and a fast one on 2020, then let a Markov switch choose between them at run time; tested on French national demand after the lockdowns, the switching version beat both fixed-variance filters. The lockdown label justifies itself by calibrating the fast regime, and nothing about a stringency or mobility series is needed at inference.
-
All of this evidence is national or building-level, none of it a distribution substation. A de Vilmarest, Abรฉlรจs, and Berezvai results are national transmission demand, and the only UK-level figure anywhere in this set is Chen et al.'s national mean absolute percentage error. A primary serves a few thousand customers with a strongly non-average mix, so its lockdown response could be far larger or far smaller than the national one, depending on whether it feeds a city centre or a dormitory estate. A search of OpenAlex for COVID-19 load forecasting at distribution substations returned nothing, so the magnitude is a per-substation question we will have to answer from NGED's own history.
-
Keep exclusion as an ablation arm. Dropping 2020-03 to 2021-07 costs roughly 1.3 of about 5.5 winters. Probably the wrong trade โ lockdown distorts the occupancy and calendar response far more than the weather-to-power response, which is what the extra history is for โ but it is one config flag, so measure it rather than assuming.
-
It breaks under the global model. A per-series booster learns its own sign and magnitude for a national scalar, which handles customer mix for free: an industrial-estate primary and a residential one moved in opposite directions. A single global booster per
time_series_typecannot, without a customer-mix covariate we do not have. -
It is a v0.6 requirement, and a v0.6 test case. An unmodelled 16-month regime is the largest phantom event the stage-1 switching baseline could face, so the covariate is a requirement there rather than a nicety. Conversely, COVID is a free labelled test case for that milestone's self-resetting residual accumulators, which detect regime shifts with no hand-coded dates.
Two things to check in the pre-2024 power data before trusting it. NGED's switching logs go back to at least 2019, so the gap years contain real switching events and need whatever masking the modern data gets. And the primaries' "Disaggregated Demand" depends on which embedded generators were metered at the time, so a meter coming online mid-history silently redefines that series โ the Embedded Capacity Register and MPAN-to-substation ingests carry the connection dates needed to check.
Evaluation
-
Scoring against ERA5 is a diagnostic, never the promotion criterion. It decomposes total error into the weather-to-power response โ the part we can actually improve, since NWP error is exogenous to us โ and the implicit hedging against forecast error. Expect the two rankings to disagree: under perfect weather the best model leans hard on weather features, so a large divergence is information about how much hedging a model does, not a bug. The same scope carries the perfect-weather ceiling, which sizes how much of our error is the weather forecast's fault and so gates how much to invest in the weather input at all. It lands as a new
evaluation_scope, not as a new fold, so leaderboard folds stay ENS-only and both principle 8 and the rejection of reanalysis-backed validation folds stand. -
Validate the no-NWP fallback on held-out 2024+ rows with NWP artificially removed, never on pre-2024 rows โ otherwise we measure fallback skill in a demand regime we will never forecast again.
-
Decide the evaluation protocol before sweeping the variant grid. The cells are not independent, and a dozen runs against a single held-out period produce a winner whether or not there is a real difference โ particularly since these questions hinge on seasonal behaviour and we have only two ENS winters to evaluate against. We need a stated test separating a genuine improvement from run-to-run variance.
-
ERA5 is a single frozen IFS cycle across the whole archive, so year-over-year comparison within it is not contaminated by NWP system upgrades. The flipside: our ENS archive spans cycle changes, so some apparent drift there is the weather model changing rather than the electricity network.
A staged-GRIB route fills three of the missing years without waiting for the Zarr backfill
Dynamical.org are backfilling the operational IFS ENS archive as a queryable Zarr store โ the real forecasts as they were issued, not reforecasts โ from ECMWF's MARS tape archive: 2016-03-08 to 2024-04-01, 51 members, 0.25ยฐ, 00Z initialisations only (dynamical-org/reformatters#446). Honest multi-year folds from that would be strictly better than pre-training and would make most of the variant grid unnecessary, but as of 2026-05 the Zarr estimate was ~November 2027, MARS-bound at roughly 0.8 TB/day against ~446 TB remaining โ well after v1.0.
Dynamical.org also stage the same MARS files as raw GRIB1 on Source Cooperative, ahead of turning
them into Zarr, and a pilot proved we can decode those files ourselves. The staged bucket carries
complete dates from 2021-03-21 to 2024-03-31 today โ about three of the missing years โ readable
anonymously. The pilot (#951,
merged) fetched the control member for 23 dates across that range and checked the result six ways:
317 of 317 sampled messages decoded bit-exact against ecCodes, the idx offset chain had no gaps, 20
of 20 re-fetched byte ranges hashed identically to the first fetch, 2t and 2d confirmed in
kelvin, and de-accumulation matched Dynamical.org's own clipping rule. A follow-up listing check
confirmed every one of the 23 pilot dates also has the full 51-member surface and pressure-level
files, not just the control member, so control-first fetching is not a separate question to raise
with Dynamical.org โ the staged bucket already orders nothing, and we can fetch either the control
member alone or all 51 members as needed.
The wider fetch is on hold. Dynamical.org has indicated they may be able to materialise their Zarr backfill over this same range sooner than the ~November 2027 estimate above, which would replace the hand-rolled GRIB decode with a plain Zarr read. Issue #959 tracks the wider staged-GRIB fetch and is paused pending their reply, rather than committing to a fetch effort estimated at ~4-5 hours and 530 GB for the control member alone, or up to 6-14 days on the workstation (or an estimated $10-30 on a cloud machine, in a few hours) for all 51 members, if Dynamical.org's own Zarr route lands first.
Two details still worth tracking regardless of which route lands:
-
00Z only, which runs against #350's move to the live service's four daily inits.
-
The backfilled span crosses further ENS resolution upgrades: 41r2 in 2016-03 (32โ18 km), 48r1 in 2023-06 (18โ9 km, within the staged bucket's complete-date range), and 49r1 in 2024-11. Each is an era boundary under the
studyskill.
The ERA5 ingest is unconditional either way: capacity estimation, the weather-abnormality climatology, and the ERA5 diagnostic scope all need it regardless of which ENS backfill route lands.
Implementation details (deleted when this ships)
Ordered, and deliberately not one PR. Steps 1โ2 are the data layer; the rest are experiments.
- Ingest ERA5 for 2020 to present, gap and overlap (#143).
- Compute paired
ENS โ ERA5residual statistics on the overlap, per variable, per lead time, per season. - Build the masking augmentation (NWP features, and spread/quantile columns separately) as a configurable step, plus the sampled pseudo-horizon for lag nullification.
- Add the lockdown covariate and the per-era sample weights.
- Define the evaluation protocol and the variance-versus-improvement test, and add the ERA5
diagnostic
evaluation_scope. - Sweep the variant grid (#167): reconciliation method ร single-pool/two-phase ร pure/mixed phase two ร flag on/off ร spread columns present/masked.
Ordering against the rest of v0.5. The Tier-1 and Tier-2 config wins on XGBoost improvements do not wait for any of this, and one of them, the lead-time feature, is a prerequisite for the lead-time-zero framing. The data-hungry structural items (batched training, ensemble-member training, the global model) are worth running after the history lands, since that is where four extra years change the answer most.