Skip to content

The current state of the art in energy forecasting

Flexpectation is National Grid Electricity Distribution's (NGED's) project to forecast net demand at its substations, delivered with Open Climate Fix. This review sets Flexpectation's plan against what the energy forecasting literature has published.

Summary

No honest review of the energy forecasting literature can name a canonical state of the art. Energy forecasting papers measure performance in different ways against different datasets, so the literature cannot rank the approaches it contains. The literature is like an international football tournament where every team plays by different rules, with different size goals. Energy forecasting researchers have done great work over the years. The lack of comparability is nobody's fault, but a systemic failure. The industry is already aware of the failure, and people are trying to fix the failure. We review several substantial efforts to compare forecasting approaches fairly.

What the literature does show is that the machine learning approach Flexpectation version 1 uses — a gradient-boosted tree, which builds its forecast from hundreds of small decision trees, each fitted to the error the trees before it left behind — is a sensible place to start. The literature we reviewed provides no conclusive evidence that a more sophisticated model family delivers a large, dependable improvement over a gradient-boosted tree at substation level. NGED's own Electricity Flexibility and Forecasting System reached the same choice independently in 2019.

In terms of machine learning research, Flexpectation is ambitious: several of our research ideas have no precedent in the literature we reviewed. For example, we found no paper that detects a switching event by checking that the power leaving one substation arrives at its neighbours; none turning a substation's switching-contaminated history into a useful input rather than deleting that history, rewriting that history, or absorbing the accuracy loss of leaving that history in; no published model that recovers a latent normal-running-arrangement demand for a distribution substation; none driving a probabilistic substation forecast from a weather ensemble across a 14-day horizon; none reading a substation forecast off a pre-trained weather encoder; none aggregating building thermal physics up to a substation and putting that physics inside a probabilistic forecast; no capacity estimator run across a mixed fleet of individually metered generators at one distribution network; and none putting unmetered generation inside a probabilistic substation forecast over a multi-day horizon. Most striking of all, almost every study we reviewed that touches more than one of these challenges solves the challenges as a pipeline. Each stage's output is frozen before the next stage sees it.

At least some of Flexpectation's ambitious research is likely to fail. Each absence above says that we did not find prior work, not that the approach will succeed. Some of these ideas will turn out to be worse than the gradient-boosted tree Flexpectation version 1 starts from. A negative result, published clearly, is a real outcome of the project rather than a failure of it. What makes the ambition worth attempting is that the nine challenges surface in the same place — as a discrepancy between what a substation metered and what the weather and the calendar say the substation should have metered. As a result, one model reasoning about several of the challenges at once has information that a serial pipeline throws away. None of that risk falls on the forecast NGED receives: version 1's gradient-boosted tree is the deliverable. Every idea above has to beat it on held-out data before it goes anywhere near an operational forecast.

Northern Powergrid's Artificial Forecasting project has run operationally through a full winter flexibility procurement cycle. That operational run is among the clearest evidence we found that a forecast of this kind changes what a network operator does.

Flexpectation's uncertainty comes from a weather ensemble. Instead of one weather forecast, the European Centre for Medium-Range Weather Forecasts (ECMWF) runs 51 of them from slightly different starting conditions. The spread between those 51 members is what tells us how uncertain the weather is: where the members agree the forecast can be narrow, and where the members disagree the forecast has to be wide.

The value NGED gets from the forecast sits in both tails of the distribution: the upper tail, where flexibility procurement holds demand under a limit, and the lower tail, where curtailment holds export under that same limit. Yet, most energy forecasting research is focused on the middle of the distribution. The limit is not always the rating of the substation itself. NGED also derives upstream limits, at bulk supply points and grid supply points, by combining several substation forecasts in a power-flow model.

Standard accuracy measures can reward the wrong forecast and hide badly calibrated uncertainty, so measuring a power forecast well takes more than one score. Mean absolute error rewards flat forecasts that are of little use for either flexibility or curtailment decisions: a peak predicted an hour late is penalised twice, once for the peak that did not happen and once for the peak that was missed. An overly smooth forecast avoids both penalties. Ranking well on one measure also says little about other measures. Across 200 German low-voltage feeders (Kaas et al. (2026)), the two models that came first and second on consumer peaks in the quantile version of an overload-decision metric stated their own uncertainty badly. Their 90% ranges contained the true value less than half the time at those consumer peaks.

Three published results point against parts of Flexpectation's plan, and we intend to test all three rather than avoid them. More detailed weather data has not always improved performance; weather data has improved performance less than expected at low voltage in the past; and a pre-trained machine-learning model trained on none of NGED's data may match models trained on all of it.

Whilst the literature we found does not tell us exactly which algorithms provide the best forecasting performance, the literature is clear on how to research and develop a state of the art forecast. There's no magic. Machine learning is an empirical science, and most research ideas fail. John Jumper, who shared the 2024 Nobel Prize in Chemistry for his work on AlphaFold, puts the share of research ideas that fail at around 90%, and treats that rate as an ordinary and necessary feature of doing research rather than as evidence of doing it badly (Nobel Week interview, 6 December 2024, from 14:12). So progress comes largely from being able to quickly test many ideas under identical conditions and carefully measure performance. We have built a machine-learning operations (MLOps) framework that should allow us to test research ideas as efficiently as possible.

The platform for running those experiments is already built, and built for speed. Flexpectation has hundreds of machine learning (ML) ideas to test. The software platform is therefore designed to make each experiment quick to run, quick to validate, and directly comparable with every experiment before it. That speed makes failure fast and low-effort. Fast, low-effort failure is what makes trying hundreds of ideas realistic.

The fact that the industry doesn't yet know the state of the art is a huge opportunity for the Flexpectation project. We are in a very privileged position where we can try hundreds of ideas, and test the best ideas in the real world. We have an opportunity to make a significant contribution to the energy forecasting community by publishing leaderboards of ML experiments, and hence help the industry as a whole to better understand how multiple approaches perform.

AI disclosure

The bulk of the ideas in this literature review are "human". The structure of this literature review is human; the research questions are human; the text was either written manually or drafted by Claude and heavily reviewed and edited manually.

Claude Code did the mechanical work of the review. We used Claude Code as a "research assistant": it searches the literature, downloads PDFs, creates tables summarising papers, traces citations forwards and backwards (to find updates on results published a few years ago), adversarially reviews its own text to check claims against the source PDFs, finds gaps in the literature, writes little Python scripts to download data published in the literature to confirm results, etc.

We're confident the facts in Claude's "research notes" are accurate because we configured Claude to check against the downloaded PDFs (rather than half-remembering information embedded in the large language model's weights), and because we ran on the order of 100 rounds of agentic adversarial review and hundreds of manual fact checks. (The "literature review" process we developed is written up as the literature-review Claude Code skill).

But — to our tastes — Claude struggles to write readable prose, so the text has been heavily re-written and cut down by hand.

What the literature says about the nine challenges Flexpectation aims to solve

The nine challenges Flexpectation's specification breaks into get very uneven coverage in the literature. This section takes each in turn: what the challenge is, what the literature says, and what that means for Flexpectation. The first challenge (probabilistic forecasts of net demand at substations) has a large body of literature, and the second challenge (forecasting metered generators) is the most mature field on the list. The eighth challenge (disaggregating unmetered solar and wind) needs the longest review of the literature of the nine, because the published work sits either side of the aggregation level NGED meters at and the review borrows from four fields outside energy forecasting. Forecasting net demand is the highest priority of the nine challenges. The other eight challenges exist mainly to improve our net-demand forecast.

No single published project we reviewed answers all nine challenges, so each challenge has its own closest precedent. The table below names that precedent for each challenge, alongside what the precedent means for Flexpectation. The sections that follow give the evidence behind each row.

Challenge Closest published precedent What this means for Flexpectation
1. Probabilistic net-demand forecasts at substations Artificial Forecasting at 551 primary substations, Pinheiro et al. (2023) at 96,989 Portuguese secondary substations, Scottish and Southern Electricity Networks' (SSEN) TRANSITION at 13 A gradient-boosted tree (GBT) is a defensible default for Flexpectation version 1, but the literature paints GBTs as a sensible starting point rather than a proven winner
2. Forecasting metered generators Dantas and Browell (2026) on 73 wind farms in Great Britain (GB) from the European Centre for Medium-Range Weather Forecasts (ECMWF) ensemble, the Hybrid Energy Forecasting and Trading Competition (HEFTCom)'s day-ahead portfolio forecast, and Nguyen and Müsgens (2026)'s meta-analysis of 4,687 skill scores from 188 solar forecasting papers Gradient-boosted trees fitted separately for each kind of generator are the approach the papers we read reach for most often, and what won when teams were scored against each other on the same data. A higher-resolution deterministic forecast beat the ensemble at short lead times
3. Estimating the effective capacity of metered generators Viotti et al. (2026), fitting a wind farm's capacity against a capacity factor simulated from reanalysis weather, and Dantas and Browell (2026), ratcheting a running maximum of the farm's own metered production. Every method we found covers one generation technology, and most work from a revenue meter alone Flexpectation version 1 needs an estimator that can track effective capacity downwards, which is exactly where the two published wind methods differ
4. Detecting switching events Bouman et al. (2024) at 180 Dutch primary substations, using a second load estimate built from smart meters; a Korean series of four papers, three on one feeder and one on two; ATLAS on GB substations in 2016 The one published result we found scoring both precision and recall reports F1.5 scores (a blend of precision — the share of flagged points that really were switching — and recall — the share of switched points the detector flagged — weighted towards recall, 0 for a useless detector and 1 for a perfect one) between about 0.2 and 0.5, from different detectors at different event lengths, and achieved with a second load estimate not available to this project, so Flexpectation should expect worse rather than better
5. Forecasting a substation as if it were always in its normal running arrangement Three published responses: leave the level shifts in (Huyghues-Beaufond et al. (2020)), rewrite the history (Paredes and Vargas (2017)), or adapt to the new level (de Vilmarest et al. (2024)) Every published solution we found throws information away. In contrast, Flexpectation version 1 makes the abnormal periods an input to the ML model, and drops the abnormal periods from the training target
6. Detecting faulty metering Bouman et al. (2024)'s Dutch dataset which merges metering faults and switching into a single class None of the three GB projects we checked publishes labels or an accuracy figure, and Flexpectation is not labelling NGED's telemetry either, so a precision and a recall are out of reach. Flexpectation judges its cleaning rules downstream instead, by whether excluding the periods a rule flags improves the forecast on held-out data
7. Recovering signed power from apparent-power meters Bouman et al. (2024) and Western Power Distribution's 2017 Time Series Data Quality both recover the sign from a second measurement of the same power; SSEN's TRANSITION instead uses a meter's own 4-year average net demand together with a model of the generation behind that meter Flexpectation version 1 forecasts the affected series in apparent power and flags those series to NGED; version 2 puts the magnitude inside a differentiable-physics forward model — the phase-retrieval formulation — and breaks the sign ambiguity with weather and with the persistence of flow direction
8. Disaggregating unmetered solar and wind Teng et al. (2023) transferring from fully-metered Dutch substations, and UK Power Networks' Power Flow to Solar Capacity, this work's direct predecessor UK Power Networks' Power Flow to Solar Capacity attacked the same problem on the same kind of GB primary-substation data, and Open Climate Fix delivered that project too
9. Disaggregating heat pumps, chargers, and batteries (stretch goal) Ostermann and Haug (2024) on aggregated charging demand day-ahead Heat pumps, chargers, and batteries stay inside net demand in Flexpectation version 1 rather than being forecast separately

1. Producing probabilistic forecasts of net demand at substations

The challenge

Net demand is gross demand minus whatever generation sits behind the substation. Flexpectation version 1 forecasts the 20 substations among the 32 series in NGED's trial area — 16 primary substations, 2 grid supply points, and 2 bulk supply points. Version 2 extends that to net demand at every grid supply point, bulk supply point, and primary substation in NGED's licence areas. Our forecasts will be half-hourly, 14 days ahead, updated every 6 hours, and probabilistic. NGED mostly acts on the forecast 1 to 10 days ahead. The question NGED asks of the forecast is "how likely is net demand to run outside the substation's firm capacity?" rather than "what is the most likely net demand?". A substation's firm capacity is the load it can carry safely with its largest transformer out of service.

Two costs hang on the answer, and both costs sit in the tails of the forecast distribution rather than at the distribution's centre. Calibrated extreme quantiles at both ends therefore reduce both costs. The first cost is what NGED spends procuring flexibility to hold demand under the substation's capacity. The second is what curtailing embedded generators costs to hold export under a substation's export capacity. A quantile is a level the forecast says net demand will stay below a stated fraction of the time. A calibrated quantile is a quantile the outturn crosses exactly that often: the level given as the 99th percentile is exceeded 1 time in 100, no more and no less.

A substation's firm capacity is not a single number. A transformer's safe rating rises as the air gets colder and as wind carries heat away from the transformer. The same plant therefore carries more on a windy January night than on a still August afternoon. And because the plant has thermal mass, the plant can take a large overload for a short period without damage. How long an exceedance lasts therefore matters as much as how far above the rating the load goes. A single firm capacity is a planning convention laid over a limit that moves.

What the literature says

A large literature forecasts substation load, but almost none of it can be compared study to study, and we found no paper driving a probabilistic substation forecast from a weather ensemble across a 14-day horizon.

Papers reviewed

The 10 papers below span national demand down to individual low-voltage feeders. Only one of the 10 papers runs in live production at national scale. Each entry gives what was forecast and at what scale, the horizon, the result and the baseline the result was measured against, and the weather input.

  • Kaas et al. (2026) — net load at 200 low-voltage feeders — the lines running from a substation out to clusters of customers — in Germany, 4 days ahead. A general-purpose foundation timeseries model (Chronos-2) that was not trained on the authors' data beat every purpose-trained model on mean absolute error, 3.8 kW against 4.2 kW. Weather: 1–3 h forecasts, so effectively after the fact at the 4-day horizon.
  • Hertel et al. (2026) — load in Germany and Portugal, at transmission level, 200 low-voltage feeders, and 287 individual customers, 4 days ahead. Their best model beat a day-type persistence forecast by 59.6% at transmission level, 42.3% at low-voltage feeders, and 23.3% at individual customers. Weather: 1–3 h forecasts at the feeders, reanalysis (a modelled reconstruction of past weather) elsewhere.
  • Browell and Fasiolo (2021) — regional net load at 14 grid supply point groups in GB, day-ahead. Their forecast held the same risk with up to 24.6% less upward reserve than a fixed-tail alternative, falling to 3.2% at the least extreme risk level tested. Weather: real forecasts.
  • Pinheiro et al. (2023) — load at 96,989 secondary substations in Portugal, day-ahead. Their forecast was 42–47% better than the reference benchmark at system level, and at substation level beat a naive forecast on 83–87% of operator-owned and 66–70% of customer-owned sites (the paper's body text and the caption of a figure on the next page give different pairs of numbers for that statistic, so the ranges span both). Weather: real forecasts, 7–8 h old. Pinheiro et al. is the only study in this review running in live production at national scale.
  • Faustine et al. (2025) — net load at Stentaway, a primary substation in NGED's South West licence area, and at a low-voltage substation on Madeira serving about 100 consumers; day-ahead. A multi-layer perceptron trained by quantile regression matched or beat N-BEATS, N-HiTS, and a long short-term memory neural network at both sites, reaching a normalised root-mean-square error of 0.08 and 0.07 against each substation's installed capacity. Every model was held to a comparable parameter count rather than tuned individually. Weather: observed rather than forecast.
  • Gilbert et al. (2023) — load at four levels of a hypothetical distribution network in GB assembled from London smart meters, primary substation down to household, day-ahead. Combining forecasts gained 0.0–0.4% averaged over all periods, but 5.7–9.0% when restricted to peaks. Weather: none at all.
  • SSEN TRANSITION 2021 — net load in Oxfordshire at 13 primary substations, plus their bulk supply points and their 33 kV and 11 kV feeders, 30 minutes to 10 days ahead. The project reported 11 of 13 primary substation models below 10% mean absolute percentage error when fitted. The two that missed 10% reached 13.4% and 19.7%. Of the 11 kV feeders the project built models for, 94% came in below 20%. Weather: 40-member ICON-EU ensemble to 4 days, then one deterministic forecast to 10 days.
  • Artificial Forecasting (Northern Powergrid) — demand and export at 551 primary substations with export data, 171 of those substations modelled, and active power at 729 secondary substations; day-ahead to week-ahead at the primary substations, evaluated to 11 days, and week- to month-ahead at the secondary substations. The published results give about 8% lower mean absolute error of utilisation rate than Northern Powergrid's existing method. Artificial Forecasting also captured 83% of the top 10% of demand values inside its 5th-to-95th-percentile band, short of the 90% that band nominally claims, so the forecast was under-covered at exactly the peaks a network operator acts on. Artificial Forecasting beat its comparison benchmarks at all eight of the near-capacity substations it was evaluated on. Weather: real forecasts at the primary substations; none in the published secondary-substation results.
  • Ruhhütl et al. (2023) — load and generation at Austrian primary substations, count not stated, day-ahead. The paper reports 3–8% mean absolute percentage error for load, against no baseline the paper states, so not a target. The error varied with how industrial and how large the supplied area was. Generation is forecast per technology: photovoltaic to 1–5% of installed power, run-of-river and biomass to 5–15% mean absolute percentage error. Linear and Gaussian regression were preferred over tree regression and a neural network. Weather: real forecasts of global radiation, temperature, and precipitation, from a weather station chosen per substation.
  • Mesarcik et al. (2025) — active power in the medium-voltage grid in the Netherlands, trained on 312 Alliander substations over 10 years and tested on six chosen for difficult forecasting behaviour, 2 days ahead. Their model reached a mean relative mean absolute error of 0.07 at the 50th quantile, against 0.08 for a gradient-boosted machine and 0.09 for a linear model — both OpenSTEF models already in production at Alliander. Error scaled by the signal's own 1st and 99th percentiles, not by a rating. Weather: Open-Meteo, four variables; their model trained on actual weather where the two baselines trained on 1-hour-ahead forecasts.

What this means for Flexpectation

Model family choice

Building Flexpectation version 1 on a GBT such as XGBoost is defensible, but the literature paints GBTs as a sensible default rather than a proven winner. NGED's own Electricity Flexibility and Forecasting System (EFFS) project picked XGBoost, which gave the best results of the three methods the project tested and was also easy to automate. No study we read shows a large, dependable margin for a model family more sophisticated than XGBoost at substation level.

Both deployments by network operators that actually tried boosted trees kept a simpler model instead. Pinheiro et al. (2023), running a live system forecasting 96,989 Portuguese secondary substations, scored 199 MW root-mean-square error at system level with a tuned gradient-boosted tree against 191 MW for a generalised additive model, the boosted tree 4% worse. Pinheiro et al. rejected the boosted tree on the effort of tuning it and on the interpretability given up with it. Artificial Forecasting kept the simpler model when forecasting customer export at primary substations. Measured against the Bayesian ridge regression Artificial Forecasting went on to adopt (a linear model that shrinks its coefficients and reports uncertainty on them), boosted trees "helped some substations but harmed others".

Neither end of the sophistication scale is a safe bet. Mesarcik et al. (2025) caution about the uncertainty a boosted tree reports rather than the accuracy it reaches. On the one substation whose calibration they plot, their gradient-boosted machine's 95th percentile forecast corresponded to the 80th percentile of the measured data. A structured state space model and a linear quantile model both tracked the ideal calibration line closely. Hertel et al. (2026) make the same point from the other end of the sophistication scale. Their purpose-built Transformer variant — the neural-network architecture, not the electrical kind — lost to a standard encoder-decoder Transformer on all three of their datasets. Faustine et al. (2025) reach the same conclusion at Stentaway substation in Plymouth, a primary substation in NGED's own South West licence area. A multi-layer perceptron trained by quantile regression, the plainest feed-forward neural network in their comparison, matched or beat N-BEATS, N-HiTS, and a long short-term memory neural network at Stentaway and at a low-voltage substation on Madeira serving about 100 consumers, reaching a normalised root-mean-square error of 0.08 and 0.07 day-ahead against each substation's installed capacity. Every model in that comparison was held to a comparable parameter count rather than tuned individually. The weather covariates are observed rather than forecast, and the 7-day figures come from feeding the day-ahead model its own output. The margins therefore bound how the architectures rank rather than what Flexpectation should expect.

What did help was refitting the model every month. On both datasets where Hertel et al. (2026) tried refitting, the retrained model beat the static model. So, for Flexpectation, the literature suggests that the choice of model family may matter less than the data, the feature engineering, and how often the model is refitted.

The simplest positive result we found came from specialising the same model to periods of the year rather than from changing the model family. Pinheiro et al. (2023) built a master model from a general-purpose generalised additive model plus copies of that model fitted to weekends, August, public holidays, Easter, Carnival, Christmas and New Year, and the remaining seasons. The weights combining those copies were updated online as new data arrived. Adding those copies in turn cut system-level root-mean-square error from 203 MW to 154 MW, a 24% reduction. The largest single step, to 180 MW, came instead from giving the general-purpose model a covariate for the demand one week earlier rather than from any period-specific copy.

Read those results knowing that when a paper says "XGBoost" it usually means a model with considerably less feature engineering than what we plan to implement. Kaas et al. (2026) give their ML model lagged power, weather, time, and six columns describing each low-voltage feeder — among them how many housing units, industrial and commercial units, and photovoltaic systems the feeder serves — and nothing beyond that: no clear-sky index, no wind power curve, no monotone constraints. Pinheiro et al. (2023) ran one of the two comparisons on equal terms we found: their GBT and their generalised additive model — a regression that fits a separate smooth curve for each input and adds the curves together — received the same features. But that shared feature set was itself short: a linear trend, load lagged 24 hours and 1 week, time of day, 9 day types, the named public holidays, day of year, and temperature interacted with time of day and with day of year. That shared feature set carried no irradiance and no wind. Faustine et al. (2025) ran the other, giving CatBoost and a random forest the same lagged net load, irradiance, temperature, and calendar features as their neural networks, and holding every model to a comparable parameter count rather than tuning each model. That shared feature set carried no wind. So no published head-to-head we found gives a GBT the feature engineering we plan to implement in Flexpectation version 1.

One published result says a non-weather input helped more than a better weather forecast, at the substations where Flexpectation would least expect it. Artificial Forecasting's Alpha work added the National Energy System Operator's national demand and operational-margin data to its substation models. The project reports that the operational-margin feature "generator availability" was "almost universally heavily used as a feature in the model and almost universally substantially improved results" at primary substations connected to wind generation. Northern Powergrid call that result surprising, because the models the feature improved on already carried a forecast wind-speed feature. No figure isolating that one feature from the rest of the inputs added in the same round is published. The finding therefore says which input is worth trying rather than what the input is worth.

Limits on the published numbers

None of the numbers above is a target for Flexpectation, because the studies cannot be compared even with each other. Kaas et al. (2026) and Hertel et al. (2026) name different models as best, even though they use data from the same 200 low-voltage feeders in Germany. Inside Kaas et al. (2026), mean absolute error and an overload-decision metric name different winners again. Neither disagreement is a mistake. The two papers test different sets of models at different time resolutions, and the two metrics answer different questions.

Accuracy got worse further down the electricity network in every study that forecast more than one voltage level. What shrank is the headroom above a naive rule rather than the usefulness of the forecast. Hertel et al. (2026) ran the same models against a day-type persistence baseline on three datasets: a German transmission control area, 200 German low-voltage feeders, and 287 individual Portuguese clients. The margin over that baseline shrank from 59.6% to 42.3% to 23.3% as aggregation fell. Their own gloss is that it is easier to beat a simple approach on highly aggregated data than on volatile feeder-level and client-level data.

The one study we found reporting results substation by substation at scale shows how much skill the shrinking headroom takes away at an individual site. Pinheiro et al. (2023)'s model beat a "same time yesterday" forecast at 83 to 87% of operator-owned secondary substations but at 66 to 70% of customer-owned secondary substations.

NGED's primary substations may not behave the same way, because a primary substation aggregates far more customers than a Portuguese secondary substation does. A forecast at a primary substation may also carry a larger percentage error than a forecast at a grid supply point and still support flexibility procurement and curtailment decisions just as well. What NGED needs from the forecast is a reliable answer to "will this substation run outside its firm capacity?". This project can measure whether decision-usefulness really is flat across voltage levels, and we intend to.

Horizon, ensembles, and tails

On the rest of the Flexpectation specification — a weather ensemble driving substation-level uncertainty all the way out to 14 days — we found no published result to lean on, for or against. Two published papers ask for exactly that combination. Haben et al. (2021) end their review of 221 low-voltage papers by naming "post-processed weather ensemble predictions to generate multi-step probabilistic forecasts of load at different levels of the LV [low-voltage] hierarchy" as an avenue of future research. Of those 221 papers, 3 used a weather forecast and none used an ensemble of weather forecasts. Pinheiro et al. (2023), published after that review closed, is a fourth paper using a real weather forecast. But Pinheiro et al.'s inputs are single point forecasts rather than an ensemble. So even the largest deployment in this review used no weather ensemble.

One paper has built exactly that, but for the whole of Great Britain rather than for a substation. Ludwig et al. (2023) drive a multi-step probabilistic load forecast for Great Britain's national demand 1 to 6 days ahead from a post-processed weather ensemble, the same 51-member ECMWF ensemble Flexpectation uses. Ludwig et al. also ask for the method to be pushed down "to different layers of the energy hierarchy, including the low voltage level".

Flexpectation's 14-day horizon sits near the edge of what a weather ensemble can reliably forecast. Buizza and Leutbecher (2015) found that the lead time beyond which a weather ensemble stops beating a climatological distribution is 16 to 23 days ahead. Buizza and Leutbecher measured that lead time on upper-air variables, not on the near-surface temperature and irradiance that drive substation load.

Almost every substation-load study we found optimises average accuracy, but NGED's question is about both ends of the distribution. The literature gives a direct warning about the tails. Browell and Fasiolo (2021) is the only study we found that models the tails explicitly, and they find that "below 1% and above 99% the forecasts based on quantile regression only are not calibrated at any GSP [grid supply point] Group. Therefore, these quantiles are not suitable for use in decision-making". Browell and Fasiolo reached that finding with 5 years of half-hourly data, across regions far larger than a substation. Browell and Fasiolo work across risk levels from 0.01% to 0.25%. The 0.05% level, or one part in 2,000, corresponds to reserve being sufficient in all but about 4 hours a year. Outside those percentiles Browell and Fasiolo switch to a fitted parametric tail at each end. Flexpectation plans to follow Browell and Fasiolo and fit parametric tails rather than reading extreme quantiles straight off the model. The lower tail is the tail curtailment turns on, because a substation runs closest to its export limit when embedded generation is high and demand is low.

Model families for Flexpectation version 2

All the text above is a verdict on Flexpectation version 1. Flexpectation plans to research three more sophisticated ML model families in 2027, each planned to simultaneously reason about multiple sources of variation in substation power flow: pre-trained encoders (a model trained once on a large body of data, then frozen and reused across several later tasks), connectivity-map models, and differentiable physics (explicitly building the known behaviour of a solar panel, a wind turbine, or a building into the model, so the model has to learn only the physical parameters, not the equations). The pipelines of separate models in this literature cannot reason about several sources at once. The closing section of this review sets out the case for the work we plan in Flexpectation version 2.

The evidence behind those three ML model families is uneven.

Pre-trained models have the best support of the three. But that support comes from computer vision and Earth observation rather than from energy forecasting. DINOv3 and AlphaEarth Foundations are each a single frozen encoder read by many downstream tasks. Flexpectation plans the same arrangement, set out under "Pre-trained encoders" below. The one energy-forecasting result we found pointing the same way measures a different arrangement, a foundation model used on its own rather than an encoder reused across tasks. The general-purpose model Kaas et al. (2026) tested was pre-trained on time series, had never been trained on their data, and still beat every purpose-trained competitor across 200 German low-voltage feeders. The margin is a single median figure that the same paper's covariate ablation reverses.

Connectivity-map models have been measured on NGED's own published data and won. But the graphs those models won with were built from geographic distance and from correlation rather than from electrical connectivity. Campagne et al. (2025) compare eight graph neural network architectures against feed-forward, persistence, and foundation-model baselines on French regional load and on the GB distribution networks' open smart-meter feed — about 2 million meters and 50,000 substations across the areas of NGED and SSEN. The graph-aware models won on both datasets. But their graphs are built from geographic distance or from correlation between series, never from electrical connectivity. Whether NGED's own connectivity map improves a forecast is therefore a question we found no published answer to.

Topology otherwise enters this literature only as the summation constraint of hierarchical forecast reconciliation. Nespoli et al. (2020) apply that constraint to real secondary substations and cabinets in a Swiss distribution grid and gain up to 10% in root-mean-square error at the upper levels of the hierarchy, and under 1% at the bottom. A summation constraint carries no information about which substation neighbours which, and the constraint stops holding the moment the distribution network is switched into an abnormal running arrangement (challenge 4 above). A summation constraint is therefore not enough for Flexpectation. The nearest exception we found feeds which busbar connects to which — a busbar being the conductor inside a substation where several circuits meet — into a graph neural network. Jung et al. (2026) build one node per bus and one edge per physical line, taken from a real feeder in Gyeonggi-do, South Korea, and forecast voltage magnitude and phase angle directly rather than load. Their model beat the conventional pipeline — a long short-term memory load forecast followed by a Newton-Raphson power flow — only under rapidly increasing load, the difference under steady load being "not statistically significant". Two features of that study limit how far the result carries: the voltages Jung et al. train and score against are computed by power flow rather than metered, on a 10-bus radial feeder they call modified and simplified. Jung et al. also assume "the network topology remains fixed throughout the analysis" while noting that in practice "the topology may change over time due to switching operations, fault restoration procedures, or maintenance activities" — which is the case challenge 4 is about.

Differentiable physics is established for a generator's own output, so what would be new for Flexpectation is the substation rather than the method. Gijón et al. (2025) fit a turbine model to a wind farm's metered production, and Pierrot and Pinson (2024) fit a wind farm's capacity jointly with the forecast. A search for differentiable physics applied to substation demand forecasting produced no strong result. We found nobody aggregating building thermal physics up to a substation and putting that physics inside a probabilistic forecast, though the ingredients exist separately.

Pre-trained encoders

The case for pre-training an encoder rests on results from computer vision and Earth observation rather than from energy forecasting. The idea is to train one model (an "encoder") on a very large body of data until the encoder can turn a raw input into a compact numerical summary that keeps what matters and throws the rest away, and then to freeze the encoder's weights. Every later job reads the frozen summary instead of the raw input, including jobs nobody had in mind while the encoder was being trained. Each job needs only a small model of its own and a modest amount of its own data. The heavy computation happens once, when the encoder is trained, and is then shared, instead of being repeated from scratch by every model that needs it.

Two recent models show how well the arrangement works. Siméoni et al. (2025) describe DINOv3, a 7-billion-parameter vision model trained on unlabelled images. Siméoni et al. keep DINOv3 frozen throughout their evaluation and read every task off its representations. Siméoni et al. report that fine-tuning of the encoder "is not necessary to obtain strong performance" on tasks as different as segmentation, depth estimation, and object detection. Brown et al. (2025) describe AlphaEarth Foundations, which encodes satellite and other Earth-observation data into one 64-byte embedding per 10-metre cell per year. Brown et al. report that the embeddings cut error magnitude by about 24% on average against a representative sample of other featurisation methods, across a broad set of sparse-data mapping evaluations, without re-training on any of those evaluations.

Flexpectation plans one frozen weather encoder serving several tasks rather than one encoder per task. Both DINOv3 and AlphaEarth Foundations actually demonstrate that arrangement. The breadth matters as much as the freezing: DINOv3 and AlphaEarth Foundations are each a single encoder serving many different tasks rather than one encoder per task. A single encoder serving many tasks is what Flexpectation plans for its own weather encoder — the same frozen representation feeding the substation net-demand forecast, the metered-generator forecasts, and the disaggregation of unmetered generation.

Neither result promises that a pre-trained encoder beats hand-designed features. Brown et al. report that learned featurisations "don't always outperform designed featurization methods in scarce data regimes", and present AlphaEarth Foundations as the exception on their own evidence: the one learned featurisation in their comparison that consistently beat the alternatives they tested. The gradient-boosted tree on hand-designed features stays the baseline the encoders have to beat.

The encoders Flexpectation plans to pre-train cover weather and time, and possibly a third for place. We plan to research a neural network that turns the raw ECMWF ensemble into a calibrated probabilistic weather forecast in physical units, which a substation model then reads. Alongside that weather encoder we plan a time encoder that learns how people use the calendar — e.g. that Christmas is not an ordinary day — and possibly a space encoder holding the standing geographic context of each substation.

The weather encoder Flexpectation plans splits into two halves: turning a raw ensemble into a calibrated forecast, then freezing that forecast as a representation other models read. Each half has already been built separately, by different authors. Rasp and Lerch (2018) built a neural network that post-processes a 50-member ECMWF ensemble into calibrated probabilistic 2-metre temperature at 537 German weather stations 48 hours ahead. That neural network cut mean continuous ranked probability score — a single number scoring a whole forecast distribution against what actually happened, where lower is better — from 1.16 for the raw ensemble to 0.78, with a learned per-station embedding one of the two components the authors credit for the gain. Mitra and Ramavajjala (2023) built the second half: they freeze a weather autoencoder and train small models on the frozen representation alone, at accuracy comparable to purpose-built models. But the targets Mitra and Ramavajjala predict are further weather variables rather than a quantity measured on an electricity network.

A network operator has already fine-tuned a pre-trained weather model on its own sensors. Bodnar et al. (2025) post-train Silurian AI's 1.5-billion-parameter Generative Forecasting Transformer on Hydro-Québec's transmission-line weather stations, wind-farm met masts, and icing sensors. The post-trained model cut mean absolute error against numerical weather prediction benchmarks by 15% for temperature, 35% for total precipitation, and 15% for hub-height wind speed at 6 to 72 hours ahead. But the forecasts are of weather at the assets rather than of power. Only the two icing indicators — wind-turbine icing risk and rime-ice accretion on overhead conductors — are probabilistic.

The nearest we found anyone joining the two is one entrant in HEFTCom, a competition to forecast a GB wind-and-solar portfolio day-ahead. Browell et al. (2026) report that team Rnt fed embeddings from their own AI weather models into downstream neural networks and finished third of the ranked entrants. What we found nobody doing is pre-training a weather encoder against observations and then reading a substation's probabilistic load forecast off it, or using a differentiable model of a solar or wind farm to strip out the variance the engineering explains so that the weather encoder trains on a clean weather signal.

2. Forecasting metered generators

The challenge

Of the 32 series in the trial area, 12 are individually metered generators — 6 solar farms, 3 wind farms, a biofuel plant, a battery, and a gas generator. Each needs the same probabilistic, half-hourly, 14-day forecast as a substation. Solar and wind are driven by weather the ensemble supplies directly. The battery, the gas generator, and the biofuel plant are probably dispatched on market prices and operator decisions.

What the literature says

Forecasting wind and solar from a weather forecast is a well-studied area, and one paper matches Flexpectation's challenge closely. No paper we found forecasts a distribution-connected battery or a distribution-connected gas generator. The closest cases we found are both at Austrian primary substations, where Ruhhütl et al. (2023) forecast biomass generation from the previous day's output, and forecast market-dispatched pumped-storage hydro from the generation schedule its operator is obliged to provide.

What this means for Flexpectation

Model choice for wind and solar

Whether weather-forecast error or weather-to-power error dominates a wind forecast flips with lead time. Handling both kinds of error in one model removes a seam Flexpectation would otherwise carry. Dantas and Browell (2026) forecast 73 wind farms in GB — 34 onshore, 39 offshore — from the ECMWF ensemble, seamlessly from 6 to 162 hours (6.75 days) ahead. Two of their conclusions bear on Flexpectation. Whether weather-forecast error or weather-to-power conversion error dominates flips with lead time. Weather-to-power uncertainty dominates the short term, and weather-forecast uncertainty dominates the mid-term. The transition between the two typically falls 2 to 3 days ahead, arrives earlier for offshore farms than for onshore farms, and varies dramatically between farms. Handling both kinds of error in one model is what lets that model cover 6 to 162 hours, where the field had previously used a short-term model and a separate mid-term model. Flexpectation faces the same seam over its 14-day horizon. Dantas and Browell are evidence that the seam can be removed rather than managed.

A second conclusion is more uncomfortable for a project built on an ensemble: a deterministic forecast at higher resolution beat the ensemble at short lead times, on unequal terms. Dantas and Browell (2026)'s short-term reference method uses ECMWF's deterministic High Resolution forecast (HRES) at 0.1° and hourly steps, while their own method uses the ensemble at 0.5° and 6-hourly steps. The archive Dantas and Browell drew on carries no 100 m wind and no finer ensemble. On those unequal terms "the short-term method is better than the proposed method for horizons up to the day ahead", although it "cannot outperform the proposed method for horizons beyond 1 day ahead". Match the two on time step and on variables, and their own method matches the short-term method on the first day and beats it from 2 days ahead. Even then the deterministic reference keeps its finer 0.1° grid. The lesson for Flexpectation is that a comparison of ensemble against deterministic measures the resolution difference unless the resolution is equalised first, and that equalising resolution fully is harder than it sounds.

Restricted to the periods when the ensemble members disagreed most, their method showed a real gain that the full-year average score concealed. That gain is the argument for running a weather ensemble at all. Averaged over a full test year, Dantas and Browell (2026)'s method showed no gain over the state of the art at day 0 and day 1. Restricted to the periods when the ensemble members disagreed most — frontal passages and the like — their method showed a real gain even at those short lead times in the continuous ranked probability score (CRPS), "which was not evident in the long-run average CRPS", because a deterministic method "is not able to discriminate between high/low weather uncertainty".

Dantas and Browell's method lacks two capabilities that Flexpectation needs. Dantas and Browell fit a separate model per wind farm rather than one model across all 73 farms. They list as future work a "member-by-member correction to retain spatio-temporal structure in ensemble members", which "would allow for spatio-temporal coherence between forecasts from different wind farms" — a plain signal that the forecasts as published carry no such coherence. A net-demand forecast that adds several generators and a substation together needs precisely that coherence, and cannot take it from this paper.

Gradient-boosted trees, fitted separately for each kind of generator, are the approach the papers we read reach for most often, and what won when teams were scored against each other on the same data. Dantas and Browell (2026) model the weather-to-power relationship with quantile regression on gradient-boosted trees, fitting a separate model for each quantile. In HEFTCom — where every team forecast the combined day-ahead output of one GB portfolio, the 1.2 GW Hornsea 1 offshore wind farm plus the aggregate solar capacity of East England, about 3.6 GW together — the winning team fitted gradient-boosted trees separately for wind and for solar and separately for each weather source. The winning team scored a mean pinball loss — the score for a forecast that states a range rather than a single number, penalising it more heavily for missing on the side it claimed was unlikely — of 22.18 MWh against the organisers' starter benchmark of 53.58. The next two teams scored 23.18 and 24.64, and the organisers' own more competitive reference, entered unranked, scored 25.38. Of the top 10 teams, 9 forecast wind and solar separately before combining the two forecasts. And Browell et al. (2026) conclude that gradient-boosted trees remain competitive for day-ahead wind and solar forecasting, with performance depending heavily on implementation. NGED's own EFFS project selected XGBoost when the project evaluated model families.

One result cuts the other way, though team Rnt's route is not an argument against trees. Rnt finished third in HEFTCom's forecasting track using no tree-based model at all, feeding embeddings from machine-learned weather-forecasting models they built in-house into downstream neural networks that predicted wind and solar generation. That route rests on building and running a weather model, not on a different downstream model family.

The solar-forecasting meta-analysis

The largest meta-analysis of solar forecasting we found puts individual machine-learning models level with classic statistical models at the range NGED acts on, and only combinations ahead. Nguyen and Müsgens (2026) meta-analyse 4,687 skill scores extracted from 188 solar forecasting papers, fitting a separate regression for each horizon band. Their baseline class is classic statistical time-series models — the autoregressive integrated moving average (ARIMA) family, exponential smoothing (ETS, for error, trend, and seasonality), and multivariate relatives such as autoregressive models with exogenous inputs (ARX). Every figure in the table is percentage points of skill score against that baseline. In that table "ensemble" means a combination of forecasting models, not a weather ensemble.

Model class Intra-hour (up to 1 hour) Intra-day (1 to 6 hours) Day-ahead (over 6 hours)
Ensemble-hybrid: average several models, and chain one model's output into the next as an input +12.8 +21.2 +7.0
Pure ensemble: aggregate several models, without the chaining not significant −7.0 +8.3
Hybrid: the chaining without the aggregating +8.6 −19.3 −11.3
Image-based: sky or satellite imagery not significant +10.3 not significant
Individual machine learning, including gradient-boosted trees −3.1 not significant not significant
Regression −11.0 −5.3 not significant
The weather model's own irradiance field, used directly as the forecast not significant −17.4 −14.3

Read the table above with two qualifications about the source itself. The figures are the regression coefficients from the paper's Table 3. The paper's own summary text gives the ensemble-hybrid intra-hour and intra-day figures the other way round, and rounds the day-ahead ensemble gain to 8.5. Nguyen and Müsgens's model classes follow each surveyed paper's own nomenclature, so the boundary between an ensemble and a hybrid is fuzzy by their own account.

Read the model-class table above as the effect of the model class alone, not of the model plus its data. Model class and input are separate variables in the same regression, so each model-class figure is estimated with the inputs held constant. And their "classical time-series" class is not weather-blind: the class explicitly includes autoregressive models with exogenous inputs and vector autoregressive models. Those models are where a weather forecast enters a classical model. The comparison is therefore between a time-series model and a machine-learning model given the same data. At day-ahead range the machine-learning model wins nothing. The weather forecast itself is worth far more than the model wrapped around it, as the next table shows.

Two limits complicate that reading. The regression cannot detect whether a machine-learning model exploits a weather forecast better than a classical model does, and most of the evidence behind the regression comes from models carrying no weather forecast at all. Their regression carries no interaction between model class and input, so it cannot detect whether a machine-learning model exploits a weather forecast better than an autoregressive model does. That question matters for Flexpectation. And only 19% of the 4,687 skill scores in Nguyen and Müsgens (2026) use numerical weather prediction as an input at all, against 91% that use lagged power. So most of the evidence separating the model classes comes from models with no weather forecast in them.

The bottom row is the weather model used raw, and for most of that sample no power curve is involved. The class represented by the bottom row in the table above is the numerical weather prediction irradiance field itself — usually global horizontal irradiance, at most post-processed or averaged across several weather models — used as the forecast rather than fed as an input to a fitted model. Of the 188 papers surveyed by Nguyen and Müsgens (2026), 118 forecast irradiance rather than photovoltaic (PV) power output. Only 70 papers in the survey forecast PV power output. So for most of the sample the weather model's irradiance field is directly comparable to the irradiance those papers forecast. The authors' regression separates the model class and the forecast target as separate variables, so the 14.3-point penalty is estimated with the target held constant. But the authors never report which targets the numerical-weather-prediction papers were forecasting. Nguyen and Müsgens's advice is to exhaust the simple models first, because classical statistical time-series methods "still have very good performance compared to more complex methods such as individual ML models".

Most of NGED's metered generators are solar, and the largest meta-analysis of solar forecasting we found confirms the importance of numerical weather prediction (NWP) inputs at the lead times Flexpectation cares about. Numerical weather prediction is the largest input effect Nguyen and Müsgens (2026) measure. The inputs that improve skill at short range carry the opposite sign at day-ahead range. The table below shows percentage points of skill score improvement over the classical statistical baseline:

Input Intra-hour (up to 1 hour) Intra-day (1 to 6 hours) Day-ahead (over 6 hours)
Numerical weather prediction −9.0 −2.3 +11.6
Locally measured weather not significant +9.1 +5.1
Lagged solar power +5.7 +8.2 −6.4
Data from neighbouring sites +3.6 +3.9 −5.5

These input figures say which inputs are worth carrying at long range in general, not how much any one input improves the forecast at NGED's day-10 horizon. Each input is a yes-or-no variable rather than a choice between alternatives, so one model can carry several inputs. Nguyen and Müsgens's sample is deterministic forecasting of irradiance or plant output rather than probabilistic substation net demand. Their day-ahead band lumps the whole of NGED's 1-to-14-day window into a single category. The figures therefore say which inputs are worth carrying at long range rather than how much those inputs improve the forecast at day 10.

A GB project should expect the skill scores it can reach to sit below the skill scores a typical solar paper reports. A skill score is meant to normalise away location. But Nguyen and Müsgens (2026) find that a skill score does not normalise away location. Their regressions use the warm-temperate Köppen-Geiger zone C, which is GB's, as the baseline. The equatorial and arid zones score 2.1 to 6.6 percentage points higher at every horizon, and the snow zone 8.0 points higher intra-hour, with no significant difference at longer horizons. Nguyen and Müsgens read that as the reference model doing relatively worse where forecasting is harder, which inflates the skill score rather than reflecting a better forecast. Nguyen and Müsgens conclude that findings have to be transferred between climate zones carefully.

Differentiable physics for generators

For generators, the benefit from better weather-to-power physics is largest at short lead times. Differentiable physics (DP) attacks the weather-to-power half of the error. So on Dantas and Browell (2026)'s measurement DP has most to offer inside the first 2 to 3 days of the 1-to-10-day window NGED acts on. Beyond 3 days DP has less to offer, because the weather forecast itself is the largest source of error.

Adding a learned residual to a physical generator model is established practice, and the physical model can be fitted to the power data rather than read off a specification sheet. Gijón et al. (2025) write the actuator-disc equation for a turbine's power output, P = ½·Cp·ρ·A·v³, into a TensorFlow model, and treat the air density ρ and the area A swept by the blades as known. The power coefficient Cp — the aerodynamic term the equation does not fix — is estimated from wind speed, pitch angle, and rotor speed by a neural network whose sigmoid output layer holds Cp below the Betz limit of 0.5926. That neural network is trained against the measured power of a wind farm of four turbines. The gradient of the power error therefore passes back through the physical equation itself. A second neural network is then trained on the residual, cutting the physics model's mean absolute percentage error by 37% and its mean absolute error by 28%. Conformalised quantile regression supplies the uncertainty. Gijón et al. also compare their hybrid model against a purely data-driven model given the same eight inputs, and report that the hybrid model "essentially matches" the data-driven model rather than beating it. Adding the physics model therefore made the forecast interpretable without making it less accurate.

But Gijón et al. predict power from measured wind rather than forecasting it days ahead. Their inputs are the turbine's own measurements at the moment being predicted. Their accuracy therefore says how well a fitted turbine model turns a known wind speed into power, not how well a forecast of that wind speed turns into a forecast of power days ahead. We found nobody putting a differentiable model of a generator inside a distribution network's probabilistic net-demand forecast.

Inferring engineering parameters

A second reason to try differentiable physics on generators, beyond the accuracy gain above, is to infer the engineering parameters that distribution network operators' registers do not record: the capacity a site can actually export today, a solar array's tilt and azimuth, a turbine's power curve. The generation forecasts in this literature are handed those engineering parameters: Teng et al. (2023) are given each site's capacity. HEFTCom's portfolio was the 1.2 GW Hornsea 1 offshore wind farm plus the solar capacity of a region. When an export-cable fault cut that wind farm's available capacity mid-competition, the winning team clipped its quantiles to the capacity implied by the outage notices the farm is obliged to publish. The organisers' benchmark ignored the fault and, in Browell et al. (2026)'s words, "performed extremely poorly as a result". NGED's embedded generators publish no outage notices of that kind. Estimating each generator's available capacity from its own metered output instead is challenge 3 below, which sets out the published methods in detail.

NGED's Embedded Capacity Register (ECR) gives a registered capacity for generation of 50 kW and above, but provides no other engineering parameters. NGED's August 2026 ECR lists 5,598 connected generators totalling 11,456 MW, of which 4,202 generators and 5,958 MW are solar. But a registered capacity is contractual rather than operational — the export limit is the limit "permitted as per the connection agreement". And the register carries no panel tilt, panel azimuth, or ratio of direct-current to alternating-current rating. Hence Flexpectation plans to infer those engineering parameters from the power data, using differentiable physics.

A differentiable model could infer both the operational capacity and the panel orientation of each generator, and each of those two inferences has been made to work on its own. Pierrot and Pinson (2024) treat a wind farm's capacity as a time-varying bound fitted jointly with the forecast, and beat probabilistic persistence by 34.2% on continuous ranked probability score over a 5-month test period, drawn from 14 months of data, at the Anholt offshore wind farm. Their one clean test of tracking the bound on its own gained 2.43%. Meng et al. (2020) infer the tilt and azimuth of 13 roof photovoltaic systems in the Netherlands to mean absolute errors of 4.3° and 4.5°. Meng et al. match the shape of each system's hourly output against plane-of-array irradiance from a weather station up to 195 km away. Because both curves are normalised before matching, their method needs no nameplate rating. But Meng et al. do not forecast PV power. They only infer tilt and azimuth.

We have not found any evidence in the literature to tell us how much a PV power forecast would be improved by inferring tilt and azimuth. Flexpectation therefore treats the gain as a hypothesis to test rather than a settled prize. Meng et al. (2020) and Saint-Drenan et al. (2015) both recover a system's tilt and azimuth from its metered alternating-current power output paired with an irradiance series measured somewhere else — a weather station for Meng et al., the HelioClim-3 satellite database for Saint-Drenan et al. — and land within a few degrees. But both papers report their accuracy in degrees alone. Saint-Drenan et al. also found that an azimuth fitted 5° from the true azimuth gave better simulations than the true value, because the fit balances the systematic error of the physical model. So accuracy in degrees is the wrong target. What matters is an effective tilt and azimuth that make the forecast right.

Inferring a metered generator's engineering parameters is a different problem from disaggregating unmetered generators out of a substation's power flow. For a single metered site, fitting tilt, azimuth, and the effective direct- and alternating-current capacities is the plan for Flexpectation — by gradient descent inside the forecast rather than by grid search, so that the fit is joint and probabilistic. Challenge 3 below sets out how that effective capacity would be estimated. For unmetered solar behind a substation, Saint-Drenan et al.'s algorithm "performs poorly", because it assumes one orientation per plant where the series is "the aggregated production of modules with different orientations".

Choosing the wrong combination of solar physical models raises mean absolute error by as much as 13% even when a plant's geometry is known exactly. Mayer and Gróf (2021) score all 32,400 combinations of 9 irradiance-separation, 10 transposition, 3 reflection-loss, 5 cell-temperature, 4 module, 2 shading, and 3 inverter models against a year of 15-minute metered output from 16 ground-mounted plants at 14 sites in Hungary. Mayer and Gróf find that the best model chain has 13% lower mean absolute error than the worst. Mayer and Gróf name irradiance separation and transposition — the step that projects horizontal irradiance onto the plane of the array, and so depends on the array's tilt and azimuth — as the two steps whose model choice matters most. Two limits ride along with that 13%. The 13% is the gap between the extremes of 32,400 chains rather than a typical penalty, and every plant's tilt and azimuth came from its design documentation and stayed fixed. So Mayer and Gróf bound how much accuracy the wrong physical model loses rather than how much not knowing a system's orientation loses. Flexpectation therefore plans to implement a fleet model representing the aggregate as a learned mixture of east-, south-, and west-facing basis shapes, with a soft clip standing in for many differently-sized inverters saturating in turn. Challenge 8, below, discusses disaggregation in more detail.

The battery, gas generator, and biofuel plant

The trial area's battery, gas generator, and biofuel plant each need a method, and the literature supplies a method to borrow for the battery, two methods for the gas generator — one method needing the operator's generation schedule, the other never yet fitted to an embedded generator — and a partial method for the biofuel plant. For the battery, Bian et al. (2024) recover a price-taking storage operator's own optimisation parameters from historical prices and observed dispatch.

For the gas generator, the closest published case we found forecasts a market-dispatched plant from the schedule its operator provides, not from weather or from the plant's own history. Ruhhütl et al. (2023) call predicting pumped-storage hydro "almost impossible", because the plant follows continuously changing market prices and the operator's own strategy. Ruhhütl et al. forecast pumped-storage hydro instead by linear regression on the generation schedule its operator is obliged to provide, together with temperature. Ruhhütl et al. report no accuracy figure for that class of plant, saying only that pumped-storage plants "depend highly on the accuracy of the provided schedule". So what the method needs is a schedule rather than a better model.

The second route models how the gas generator picks its own output, and nobody we found has fitted a model of that kind to an embedded generator's metered output. Fitting a model of that shape to an embedded generator's metered output is what Bian et al. do for storage. We found nobody doing the same for a gas generator. Short et al. (2017) model how a decentralised combined heat and power plant picks its own output, as a mixed-integer linear program over a piecewise-linear fuel cost, ramp limits, and a start-up cost, maximising profit against day-ahead and intra-day prices. What could transfer to Flexpectation's differentiable-physics work is the structure rather than the solver: Short et al. approximate the fuel cost by three affine pieces chosen "to ensure convexity". The economic-dispatch half of their model is therefore a linear program a gradient can pass back through. The on/off unit-commitment decisions would have to be smoothed first, though.

A GB gas-network project has tested the declared-schedule route on embedded gas generators, and what stopped it was data rather than modelling. SGN and Northern Gas Networks' Forecaster for Embedded Generation (FEmGE) 2026 Network Innovation Allowance (NIA) project reconstructed gas generators' electricity output from the Physical Notifications each plant gives the National Energy System Operator (NESO), plus the balancing bids and offers NESO accepts. Plants that self-dispatch rather than trade through the Balancing Mechanism were placed out of scope as harder still. FEmGE found no public record matching a plant's electricity meter number to its NESO unit identifier. Many small plants also sit inside aggregated units whose composition is unpublished. Forecasting performance also "reduced significantly" when transmission-connected plant was excluded, because distribution-connected generators are a small part of a zonal total. NGED does not have that problem, because NGED meters its gas generator directly. FEmGE published no accuracy figure, and concluded that more complex modelling would not improve accuracy without wider access to embedded generators' own data.

For the biofuel plant the same paper supplies a partial method. Ruhhütl et al. (2023) forecast biomass generation behind each Austrian primary substation from the previous day's generation, scaled to installed power and spread across the day as a constant band, to a mean absolute percentage error of 5 to 15%. The problem has the same shape, though a biomass station burning solid fuel is not the same plant as a biofuel generator.

3. Estimating the effective capacity of metered generators

The challenge

We call the amount of generation actually available at a metered site its effective capacity: the power output a generator could produce right now if the weather allowed, as opposed to its nameplate rating. The meter on a metered generator records that generator's output alone — not a substation's load, and not a second generator. That single-generator scope is what separates challenge 3 from challenge 8, where unmetered generation has to be recovered from a substation's net flow. Turbines go out for repair, inverters degrade, and sites are curtailed — told by the network operator to generate less than they could. A 20 MW wind farm that has been limited to 14 MW for a month is, for forecasting purposes, a different wind farm. A model trained on its nameplate rating cannot see the difference.

What the literature says

A method exists for each generation technology separately, but we found none run across a mixed fleet of individually metered generators at a distribution network operator. Only one study we found, Viotti et al. (2026), measures what estimating capacity is worth downstream, for wind alone and at the scale of a bidding zone. Dantas and Browell (2026) estimate a wind farm's capacity but treat the estimate as a known input rather than measuring what getting it right is worth.

The clearest published demonstration we found of why effective capacity matters is incidental, and it happened inside a competition. Hornsea 1's export cable faulted on 19 January 2024, about 2 weeks before HEFTCom's competition period was due to start. Neither the organisers nor the participants were monitoring the market-transparency feed carrying the outage messages. The problem was identified only after the competition had initially started on 1 February. The organisers call not anticipating the fault "an oversight", and restarted the competition on 20 February, a month after the fault. Many teams still struggled in the weeks that followed. Teams forecasting wind and solar separately could post-process their wind forecast for the new export limit. Teams forecasting the combined total "found it harder to adapt".

What this means for Flexpectation

Flexpectation version 1 needs an estimator that can track effective capacity downwards, and that is exactly where the two published wind methods differ. Dantas and Browell (2026) needed available capacity for the same reason we do. Rather than use a nameplate rating, Dantas and Browell estimate a time series of available capacity for each farm from that farm's own metered production, needing no capacity register and no outage messages. Dantas and Browell did use one data source Flexpectation will not have in the same form: they excluded curtailed half-hours using published bid-acceptance volumes. Those volumes exist for transmission-connected wind farms and not for NGED's embedded generators. NGED's active network management system records curtailment for each of NGED's generator customers. Like any operational log, the active network management record is a noisy label: curtailment can happen with no matching log entry, and a logged event may differ from the generator's actual output. The active network management record therefore cannot simply be dropped in where Dantas and Browell used bid-acceptance volumes (see effective-capacity estimation). The general shape of that capacity-estimation rule is a running maximum of production, which ratchets upwards and never comes back down. In contrast, Viotti et al. (2026) fit the most likely capacity time series instead, by quadratic optimisation against a capacity factor simulated from reanalysis weather and a power curve. Viotti et al. publish a monotonic variant alongside a non-monotonic variant. The direction of travel is what matters for NGED: a turbine out for repair for a month makes effective capacity fall. A ratchet cannot follow it down. Flexpectation version 1 will therefore implement estimators that can fall as well as rise.

The published numbers favour fitting over ratcheting, on hourly region-aggregated data. Viotti et al. (2026) say that estimating capacity using a running maximum "requires monotonically increasing capacity and relies on frequent high wind events". Viotti et al. publish a non-monotonic capacity estimator, which can follow capacity down when a turbine goes out for repair. The non-monotonic variant produced the lowest day-ahead forecast error across Sweden as a whole, 2.0% below a model normalised by the running maximum on mean absolute error and 2.3% below that model on root-mean-square error. The authors say the non-monotonic variant yields the best forecasts "possibly because it captures real changes in available capacity or corrects seasonal wind-speed biases". Viotti et al. also caution that the difference in forecast error may not reflect the quality of the normalisation at all.

Which variant produced which figure matters to NGED, because the two point opposite ways. Viotti et al. also report 27.2% lower normalised mean absolute error than the running maximum at quantifying capacity after a new wind farm connects. But that 27.2% is scored by the monotonic variant, the variant that still assumes capacity only rises, and on that same test the monotonic variant's error is 31% below that of the non-monotonic variant NGED would need. The 2.0% forecast figure above belongs to the simplified non-monotonic parameterisation instead. So the variant that estimates capacity best is not the variant that forecasts best, and the variant NGED needs wins only the second contest.

Two caveats temper both figures for NGED. Viotti et al.'s target is a Swedish bidding zone rather than a single farm. Viotti et al. report that at 5-minute resolution the running maximum is already a robust estimate of one farm's installed capacity. The fitting therefore shows its advantage on hourly, region-aggregated data. Viotti et al. do test the de-rating case, by suppressing production in the 30 days after a step. There both their method and the running maximum get worse, with no comparable improvement figure to report. So the case NGED cares about most is the one their paper answers least well. Whichever estimator wins, normalising by effective capacity stays a hypothesis to test rather than a settled preprocessing step, because no study we found has measured whether normalising improves the forecast NGED acts on.

A competition on Norwegian wind shows that its entrants reached for normalising by capacity, and also that those entrants were handed the capacity rather than having to estimate it. WindAI asked for the hourly wind power of four Norwegian bidding zones 2 days ahead, and supplied wind park metadata — including each park's installed capacity — alongside the weather and production data. Authen et al. (2026) report that Statnett scored 5% of its weighted assessment on "robustness to changes in installed wind power capacity, evolving weather patterns, long-term climate variability". The two highest-placed teams both predicted capacity factor rather than absolute production, which the winning team did "to account for maintenance events and future capacity expansions". Both teams used Nord Pool unavailability messages to cover planned maintenance and outages. A team given an honourable mention fitted a physical power curve for each wind park under sequential Bayesian updating over a sliding window, to absorb "capacity changes or the commissioning of new wind parks", and initialised a new park's parameters from a prior built on the existing fleet. That prior is an answer to the cold-start problem an estimator faces at a generator that has just connected. Authen et al. decline to credit any of that with the differences in accuracy, because unavailability messages covered between 1% and 13% of timestamps depending on the bidding zone. Downtime events therefore "represent only a limited fraction of the full dataset". Two conclusions follow for Flexpectation. Independent teams converged on capacity factor as the target. That convergence is evidence about what practitioners believe, not a measurement of what the belief is worth. The hypothesis in the paragraph above therefore stands unaltered. And the part WindAI could skip is the part any GB distribution network operator cannot skip: the Embedded Capacity Register records the export limit permitted by a site's connection agreement rather than what the site can generate. So Flexpectation has to estimate the effective capacity that WindAI's entrants were given.

For solar, the equivalent estimate can be made from the power signal and nothing else, which matters because half of the trial area's metered generators are solar farms. The tool most often cited in the papers we read, the open-source RdTools, does need site irradiance to pick out the clear-sky periods it analyses. RdTools' own documentation warns that a satellite substitute gives less stable results. Meyers et al. (2020) removed that requirement: their unsupervised signal-processing approach "only requires a measured power signal as an input — no irradiance data, temperature data, or system configuration information". Meyers et al. validate the approach against RdTools on the same dataset, reporting greater robustness to data anomalies. That approach ships as the open-source StatisticalClearSky library. Meyers et al. automated the approach's data cleaning and preprocessing with a second open-source library, Solar Data Tools. Solar Data Tools now carries a pipeline of its own, which detects capacity changes and clipping, and estimates degradation with a Monte Carlo step that returns a distribution rather than a point estimate.

Estimating capacity jointly with the forecast, rather than in two stages, has also been published — and reading its headline figure carefully matters. Pierrot and Pinson (2024) treat a wind farm's available capacity as the unknown, time-varying upper bound of a generalised logit-normal distribution and track that bound online by normalised gradient descent. That method improved the 10-minute-ahead continuous ranked probability score by 34.2% over probabilistic persistence and 17.9% over a benchmark that holds the bound fixed. The fixed-bound benchmark is their earlier rolling maximum-likelihood method rather than the same gradient-descent model with its bound frozen. The 17.9% therefore mixes the gain from tracking the bound with the gain from changing the fitting method. Their one clean test of tracking on its own pairs the rolling maximum-likelihood method with a varying bound against the identical method with a fixed bound. That test gained 2.43%, which Pierrot and Pinson report as no "significant improvement when compared to its equivalent with a fixed bound". Tracking a varying bound is worth having, then, but this paper does not show it is worth 17.9% by itself.

Flexpectation plans to attempt the mixed-fleet combination two ways, neither of which starts from scratch. The first is the two-stage route: estimate a capacity time series from the meter, then normalise by that series before training. Flexpectation would run the quadratic-optimisation method of Viotti et al. (2026) and the Solar Data Tools pipeline against each other on our own sites, with the running maximum of Dantas and Browell (2026) as the reference the published numbers are quoted against rather than as a candidate, because a ratchet cannot follow a de-rating downwards. The second is joint estimation, of which Pierrot and Pinson (2024) are the published precedent: a differentiable-physics model of each generator in which the physical parameters — including the plant's direct-current and alternating-current capacity — are fitted as probability distributions rather than as single numbers. Capacity is then recovered with its own uncertainty attached, and the forecast inherits that uncertainty instead of treating capacity as known.

One published result suggests the normalisation may not earn its place at all, which is why Flexpectation intends to measure it rather than assume it. NGED's specification asks that effective capacity be tracked over time and, optionally, combined with the forecast into a "prevailing conditions" view. de Vilmarest et al. (2024) removed the embedded wind and solar capacities from their model of GB regional net load. A Kalman filter tracking the coefficients absorbed the loss completely: error rose by more than 10% for the same model fitted offline, and fell by 0.4% for the adaptive model. The capacities de Vilmarest et al. removed are the aggregate installed capacity of a whole region's embedded generation rather than one metered generator's effective capacity. The result is therefore a caution about normalisation rather than a like-for-like test of Flexpectation's normalisation.

4. Detecting switching events

The challenge

When a cable fault or planned maintenance moves part of a distribution network from one substation to another, the load the first substation meters steps down. Each substation that picks up part of that transferred load records a rise, with no change in the underlying demand. The pick-up is usually shared across two or three neighbouring substations. Usually only part of a substation's load moves — a continuous fraction, with no minimum size — rather than a whole subgrid. NGED's substations spend roughly a tenth of their operating time in an abnormal running arrangement. Switching records have been extracted into labels only for the Flexpectation trial area. Any method meant to scale beyond the trial area therefore has to work from power measurements alone.

What the literature says

We found several papers on detecting switching events from metered load, but all these approaches only consider one substation at a time. Bouman et al. (2024) detect switching at an actual distribution network operator, but detect it in the gap between the substation's own meter and a second estimate of the same load, built from smart-meter and bulk-customer readings taken below the substation. A Korean series of four papers detects load transfers on a distribution feeder from that feeder's own load alone. All four Korean papers are open access. All four score against the same nine logged transfers on the Kimhwa distribution feeder in Gangwon province, measured hourly through 2019, and one of the four papers scores against a second feeder as well.

Paper Method Logged transfers found
Kim et al. (2020) Long short-term memory (LSTM) neural network, flagging where measured load departs from its prediction 7 of 9
Kim et al. (2022) Polynomial and standard-pattern preprocessing 7 of 9, and 7 of 7 on a second feeder
Kim (2024) A moving average and a moving standard deviation, thresholding the residual of a seasonal-trend decomposition 8 of 9
Kim (2025) Robust seasonal-trend decomposition, a Haar wavelet transform of the residual, then Pruned Exact Linear Time changepoints, then an isolation forest over each candidate 7 of 9

Every count is the share of logged events found. No paper in the series reports a false-alarm rate. Kim (2024) states that the method flagged more transfers than the nine that were logged, and attributes the surplus to unplanned operational switching the log does not record rather than counting the surplus as false positives. Kim (2025) reports that the isolation forest's probability score separated true positives from false positives. That report concedes that false positives existed but puts no number on them. Both papers explain their misses the same way: the transfers they did not catch moved too little load to show up. Both papers argue that a transfer carrying no material load change matters less. The scores do not track how elaborate the method is: the simplest of the four, a threshold on a decomposition residual, found the most events. Kim (2025)'s pipeline — the closest of the four to what Flexpectation plans — found 7 of the 9, an average detection rate of 78%.

Electricity North West's ATLAS project sorted step changes into erroneous data and distribution network reconfigurations on substations in GB in 2016, from power measurements alone. ATLAS published no precision or recall for either rule. ATLAS processed 5 years of half-hourly demand for "over 70 BSPs [bulk supply points] and 380 primary substations" — a GB fleet more than 10 times the size of Flexpectation's trial area — in two stages. The first stage flags any abrupt change, firing where the half-hourly change in demand exceeds ±80% of the standard deviation of the demand series. The second stage decides what kind of change it was, and is the part that matters here. One rule handles blocks of "unreasonably zero or negative demand", and a separate rule handles "switching operations and network reconfigurations". So the distinction between a broken meter and a reconfigured distribution network was drawn on GB primary substations, on power alone, without a bottom-up reference series. ATLAS was a data-preparation project rather than a detector-benchmarking project, which is why ATLAS reports no precision or recall. The project pairs its rules with "the importance of visual sense checks of the obtained processed demand data".

What this means for Flexpectation

Bouman et al. (2024) detect switching events, but do not forecast. Bouman et al., working with the Dutch distribution network operator Alliander, study 180 primary substations at 15-minute resolution over roughly a year. The events they detect run from a few minutes to several months. Alliander's purpose is capacity planning. A switch pushes the maximum and minimum load a substation records to the wrong value, and those two extremes decide whether the substation needs a bigger transformer. The detected periods are therefore cut out of the history before the extremes are read off. In contrast, Flexpectation needs a forecast that keeps running through a switching event.

Only one published result we found scores switching detection on both precision and recall, and its scores are low: about 0.2 on events shorter than 3 days, and about 0.5 on events of 42 days or longer. Bouman et al. (2024) score their detectors with the F1.5 score, which blends precision — the share of flagged points that really were switching — with recall — the share of switched points the detector flagged — weighting recall the more heavily of the two. Bouman et al. weight recall more heavily "to give a higher importance to the recall term, as the potential impact of a false negative is higher than that of a false positive in power grid expansion planning". An F1.5 score of 1 marks a perfect detector and 0 marks a useless detector, so higher is better. Those two scores come from different detectors, because no single method they tried wins across the range. Bouman et al. report the score separately for four event lengths — 15 minutes to 6 hours, 6 hours to 3 days, 3 to 42 days, and 42 days or longer. On the two shortest bands the best detectors, statistical process control and an isolation forest, reach about 0.2. Binary segmentation scores near what random guessing would give. On the longest band, binary segmentation reaches nearly 0.5. Combining the detectors, by flagging a point if any of them fired, raised recall but added enough false positives that on the two shortest bands the combination scored only marginally better than binary segmentation alone. Both figures were achieved on a Dutch distribution network, with the help of a second load estimate constructed bottom-up from smart meter data.

Flexpectation will model its own reference time series for each substation. Alliander's bottom-up estimate gives Bouman et al. (2024) a second opinion on what each substation's power should have been. Bouman et al. fit and rescale that bottom-up estimate to the measured series, then hunt for step changes in the difference between the estimate and the measurement, so that normal daily and seasonal variation largely cancels and leaves a much cleaner signal. A bottom-up estimate of substation load is not available to this project. Building that estimate is out of Flexpectation's scope, because the project uses no telemetry from below primary substation level. Flexpectation plans to produce that second opinion from the substation's own meter plus weather and the calendar. The first attempt is classical: a multiple seasonal-trend decomposition of each series into a trend and daily, weekly, and annual cycles, leaving a remainder in which a switch shows up as a sustained level shift. The second attempt uses the project's existing XGBoost training pipeline, trained with no power-lag features, so that an earlier switching event cannot contaminate the expected-power estimate. Neither route needs metering from below the substation.

Flexpectation also plans to investigate using a signal that Bouman et al.'s one-substation-at-a-time method cannot see: the power has to go somewhere. Bouman et al. (2024) score each substation against its own history — "the current analysis considers one year of measurements for one station at a time". So no step in their method asks whether the power that left one substation turned up at another. When one substation's metered power drops, the substations that picked the load up should rise at the same moment, and their rises should sum to the drop. A step whose rise and drop fail to balance is more likely a meter fault or a one-off than a switch. That mismatch is where a per-substation detector's false positives come from. The catch is that an NGED transfer usually fans out across two or three neighbours, so the search runs over subsets of neighbours rather than over pairs. The balance holds only approximately.

We looked for a method that checks both sides and found none, and the closest published precedent is a 1984 regression written for long-range planning. The search ran to 40 title-and-abstract queries and 10 full-text queries across OpenAlex, Semantic Scholar, Crossref, and arXiv; the works citing Bouman et al. (2024) under both its journal and its arXiv identifier, four citing works, all read at abstract level; and the titles of all 3,160 projects on the Energy Networks Association's Smarter Networks Portal, which publishes no abstracts to search. Willis et al. (1984) correct annual peak-load curve fits rather than detecting an event at a point in time. Their regression needs neither the size nor the direction of a transfer as an input. The title names a "load transfer coupling" regression, which suggests the fit couples the substations that exchange load. That coupling is the feature that would make Willis et al. the closest precedent. But we could not obtain the full text to check, and the abstract does not say.

NGED's switches are usually partial and fan out to two or three substations, so we should expect worse F1.5 scores than Bouman et al.'s 0.2 to 0.5, not better, even with the two additions above. Do not judge the difficulty from how obvious a switch looks on a chart. A negative result is worth having here, because evidence that switching cannot be recovered from power measurements alone would justify extracting switching labels from NGED's operational systems instead of continuing to infer them.

An F-score is the wrong shape for this problem on its own, because it counts every event once regardless of how much power moved. What degrades a forecast is megawatts: a switch skews the forecast roughly in proportion to the load it transfers. A missed transfer too small for an expert to see by eye barely degrades the forecast, while a missed large transfer contaminates the history the model trains on. Ranking a detector by an unweighted F-score therefore treats those two misses as equally bad. The argument about whether a false positive or a false negative is worse therefore has no general answer: the answer depends on the size of the event.

The Korean papers reach for that defence without measuring it, with one exception. All four of the Korean papers above explain their misses by saying the transfers they did not catch moved too little load to show up, and argue that a transfer carrying no material load change matters less. But only Kim (2024) publishes the megawatt step and the percentage load change of each of the nine logged transfers. On that feeder every transfer above the mean load-change rate of 37.5% was detected, and the single transfer the method missed moved 22.7%. Those figures are the closest to a measured detection floor we found in this literature.

Flexpectation therefore reports an F-score banded by event duration and pairs it with a detection sensitivity floor. The F-score makes our detector comparable with the one published measurement we found: F1 as the headline number, with Bouman et al.'s recall-weighted F1.5 alongside it wherever the head-to-head is the point. Flexpectation bands the score by duration the way Bouman et al. band their F1.5 score, because a single fleet-wide score hides the fact that short events are the hard events. The floor answers the question the F-score cannot: how large a transfer has to be before we can see it. The floor is not one megawatt figure but a frontier in transferred magnitude against event duration, reported per series relative to that series' own residual noise, because a 2 MW step is obvious at a quiet substation and invisible at a busy substation. Pairing the frontier with the forecast error that missed events below it actually cause is what turns a detection score into a statement about the forecast.

5. Forecasting a substation as if it were always in its normal running arrangement

The challenge

NGED plans its distribution network against what each substation would carry under its normal running arrangement. Flexpectation therefore aims to forecast substations as if they were always in their normal running arrangement, including a substation that has been sitting in an abnormal arrangement for weeks. Forecasting through an abnormal arrangement is a weaker requirement, and not the requirement NGED has: a model can take lagged power inputs from inside an abnormal period and stay well-behaved anyway, yet still report what the substation will carry rather than what the substation would have carried under its normal arrangement. Predicting the power flow under the normal running arrangement means the forecasting target goes unmeasured during periods of abnormal running, and leaves the training history contaminated: past readings taken while the distribution network was abnormally configured describe a different scenario from the scenario being forecast.

What the literature says

The literature answers an abnormal running arrangement three ways: leave the level shift in the data, rewrite the history to remove it, or let the model adapt to the new level. Leaving the level shifts in, as Huyghues-Beaufond et al. (2020) do, running change-point detection across 342 medium-voltage feeders in the UK and using the change-points to bound the segments within which they remove outliers; rewriting the history, as Paredes and Vargas (2017) do; or adapting to the new level, as de Vilmarest et al. (2024) do.

One published system chose its model on robustness to switching rather than on accuracy. Ruhhütl et al. (2023) compared linear, tree, Gaussian, and neural network regressions for day-ahead load at Austrian primary substations, and report that the Gaussian model "has the lowest MAPE [mean absolute percentage error] of all regression models" but "is barely able to calculate predictions when there is a major deviation from the normal switching status", while linear regression "is a little less accurate but is very flexible in terms of deviations from the switching status".

Ruhhütl et al. also clean "major deviations of the normal switching status" out of the training data before fitting, which removes those periods from the training set rather than correcting them to what the normal arrangement would have carried. Neither the size of the accuracy sacrifice nor the size of the switching failure is quantified. The paper therefore shows that an operator traded accuracy for switching robustness without saying how much accuracy the trade gave up.

What this means for Flexpectation

Every published solution we found throws information away. Leaving the level shifts in the data hurts performance. Rewriting history erases the level shifts. Adapting to the new level forgets that a switch happened. Adapting is disqualifying here, because the quantity NGED needs is what the substation would have carried under its normal arrangement. Flexpectation version 1 will therefore detect the abnormal periods automatically, flag the lagged power inputs that fall inside an abnormal period, and drop those periods from the training target — a combination no published method we found uses.

Rewriting the history is the fallback, because among the three published responses rewriting the history is the only response that targets the quantity NGED needs and reports a measured benefit for doing so. Paredes and Vargas (2017) rewrite the history to the level it would have had if the switch had never happened, across 169 real feeders, and report better medium-term forecasts for it. Northern Powergrid's Artificial Forecasting project rewrites its history too, in step 6 of the data-preparation pipeline set out in its Alpha deliverable WP2-D2 Results Scope Item 2. That pipeline rescales a block of older readings to align its median with the median of the most recent block whenever the older block's median falls outside the 10th-to-90th-percentile range of the most recent block. Northern Powergrid hold no readily accessible record of their own distribution network's configuration changes. So that pipeline hypothesises the timestamps from the load itself and confirms them with the control room. Flexpectation faces the same situation outside the trial area, where switching records are not available to the project in machine-readable form.

The fix is a level shift applied to the older half of each series. Paredes and Vargas measure how far average demand moved across the step and add that difference to every reading before the step. The variant they recommend uses a separate difference for each hour of the day and each day of the week rather than one number for the whole series. Paredes and Vargas take the event times from expert identification rather than from a detector, since detection was not their subject. Adaptive models are the live alternative — they track a new level once it arrives, including a level that arrives abruptly. de Vilmarest et al. (2024) let a Kalman filter track the drift on the 14-region GB dataset of Browell and Fasiolo (2021) instead of correcting the history, cutting error by about 4% in 2019, 7% in 2020, and 8% in 2021 against the same model refitted every day. But a switching event is a step rather than a drift. A model that simply adapts to a new load level never records that a switch happened, so it cannot report what the substation would have carried under its normal arrangement, which is the quantity NGED needs.

Flexpectation version 1 feeds the model its switching-contaminated history deliberately: the abnormal periods are an input to the ML model but are removed from the training target. First, take each substation's abnormal running arrangements from the detector of challenge 4 rather than from an operational log, and hand those periods to the model as a flag on each lagged power input, so the model can read a lag that falls inside an abnormal period correctly. Second, drop the abnormal half-hours from the training target, so the model is never asked to predict an abnormal arrangement. An alternative worth testing early is to skip the flag and give the model challenge 4's reference time series alongside the lagged power, leaving the model to notice for itself where a lagged reading departs from what the reference series expected. That reference-series difference plays the same role as the residual Bouman et al. (2024) detect on, but is built the other way round. Bouman et al.'s residual is metered load minus a topology-informed reconstruction, which goes stale the moment the distribution network is switched, whereas Flexpectation's residual would be metered load minus a model that never sees topology at all.

The closest published precedent we found for each half of the plan sits outside the problem NGED has. For the first half, Liu et al. (2019) fit a separate regression per substation operating condition, though their switching moves load between transformers inside one substation, so the substation total stays metered throughout. For the second half, Salinas et al. (2020) state the mechanism for a probabilistic forecaster, motivated by retail stock-outs, and say they omitted the experiments for that mechanism. Searching OpenAlex, Crossref, and arXiv for sample masking, zero sample weights, gappy targets, and the exclusion of anomalous periods, we found no load-forecasting study reporting how much dropping contaminated periods from the training target improves a forecast. So Flexpectation will have to measure that itself.

Flexpectation version 2 plans to go further and treat the normal-arrangement demand as a latent variable to be inferred for every metered substation, rather than a series to be repaired first, through a differentiable-physics model of each substation with separate photovoltaic, wind, and demand components. Recovering a demand the meter never saw is mature where demand is censored — airline revenue management calls that recovery unconstraining, and retail and electric-vehicle-charging work calls the same recovery censored-demand recovery, as in Hüttel et al. (2023). Estimating what a curtailed wind farm would have produced is the closest analogue we found inside the energy sector. But censoring is one-sided, so the observed value bounds the latent demand from below, whereas an abnormal running arrangement substitutes a different set of customers and can read either side of the normal-arrangement demand. Searching the same three indexes for latent-demand, censored-demand, counterfactual, synthetic-control, differentiable-physics, and physics-informed formulations applied to substation demand, we found no published model that recovers a latent normal-running-arrangement demand for a primary substation.

6. Detecting faulty metering

The challenge

Distribution telemetry, NGED's included, carries stuck values that repeat unchanged for hours or days, zeros that mean "no reading" rather than "no load", physically impossible values, and gaps running from a single half-hour to several months. A model trained on uncleaned data learns the fault. A forecast that fails silently because the series' recent history was stuck is worse than a forecast that reports itself degraded.

What the literature says

Faulty metering is usually a data-cleaning step mentioned in passing rather than a research problem in its own right. The only public dataset with labelled faults we found is Bouman et al. (2024)'s 180 primary substations, labelled at 15-minute resolution and released as "STORM onderstation" on the open data portal of Liander, Alliander's distribution network operator, explicitly so that others can train and validate algorithms against that dataset. The Liander dataset is the one place in this review where the evaluation data for a challenge is already public. Nearly 4% of its timestamps are labelled as the labeller being unsure — a figure we counted from the released dataset, because the paper does not report that figure.

What this means for Flexpectation

The literature offers two shapes of detector — test a reading against a redundant measurement of the same power, or against a forecast of what that reading should have been — and NGED's primary-substation telemetry rarely carries the redundant measurement, which leaves the forecast route. One family tests a measurement against a physical relationship the measurement has to satisfy. UK Power Networks' Distribution Network Visibility checked 377 remote terminal units against the physics their readings have to obey rather than against a forecast, and found 95% of those units obeyed the expected logic within 15 kVA. Bouman et al. (2024) do the same with a second estimate of a substation's load, built from smart meters. The other family tests the measurement against a forecast of what that measurement should have read, the route Moriano et al. (2016) and Martín et al. (2018) take to find calibration drift in secondary-substation monitoring equipment.

The published method that fits NGED's telemetry most closely treats faulty metering and switching as one challenge, and merging the two faults is exactly what stops the Dutch labels separating faulty metering from switching. Bouman et al. (2024) treat measurement errors and switch events as the two contaminants that must be filtered out before substation measurements can be used, and detect both on the same residual. Detecting both on one residual is also what merges the two classes in the Dutch labels. The Dutch dataset can therefore train a detector but cannot settle whether a flag is a stuck meter or a distribution network reconfiguration — the separation challenges 4 and 6 exist to make.

The faults that dominate distribution telemetry are not the faults the model-based detectors were built for, and the GB projects that met those faults used threshold rules. Moriano et al. and Martín et al. score calibration gain and offset drift plus outliers, injected into clean data rather than found in the wild, whereas distribution telemetry, NGED's included, carries stuck values, false zeros, and multi-month gaps. Western Power Distribution, NGED's predecessor company, ran the 2017 Time Series Data Quality project to tackle exactly this problem. The project searched for zeros, for "non-varying non-zero values, perhaps indicating a 'stuck' or incorrectly configured sensor", and for gaps, and found metering defects common rather than exceptional in Western Power Distribution's data at the time. In 2017, 13.8% of analogues in the South West licence area recorded only zeros, and 20.7% company-wide — with the caveat that "many of these may be valid open circuit values, however some will reflect incorrect values". In 2017, between 1% of the distribution network's control-system data points in the South West and 36% in the Midlands were unavailable to planners. In the same year, 63% of new solar sites' analogues had not been commissioned correctly. A fault detector for distribution telemetry therefore must not assume that faults are rare.

None of the three GB projects reports how often its checks are right. Electricity North West's ATLAS, UK Power Networks' Distribution Network Visibility, and Western Power Distribution's Time Series Data Quality all tackled faulty metering substantively. None of the three published a figure for how often a flagged reading really was faulty, nor a label set to measure that against. Distribution Network Visibility's 95% is the share of units whose readings obeyed the expected logic, not a detection accuracy. The GB record therefore tells us what to look for rather than how well the approaches worked. What Distribution Network Visibility did publish is the shape of the output: a daily health report ranking units for maintenance. We found no GB labelled set with a taxonomy separating metering faults from switching, and Flexpectation is not producing a labelled set either. So Flexpectation's cleaning rules are judged by whether excluding the periods they flag improves forecast accuracy on held-out data rather than by a precision and a recall. A run of implausible values is a fault to a forecaster and a real event to a control engineer. Only the purpose of the analysis settles which.

7. Recovering signed power from apparent-power meters

The challenge

Of the trial area's 16 primary substations, 8 are metered in apparent power (MVA) rather than real power (MW), as are two of the three wind farms. An apparent-power meter reports a magnitude with no direction, so when a solar farm behind the substation exports more than the substation's customers are drawing, the substation's reading rises instead of going negative: the trace "bounces" off zero. A midday export reads as a midday peak. NGED report that one of those 10 apparent-power sites has shown the bounce on sunny days, with two more suspected. The meter is not faulty. The difficulty is that the quantity NGED needs forecast is signed net demand, while an apparent-power meter reports the absolute value of signed net demand — and reports even the absolute value only approximately.

What the literature says

A magnitude-only measurement leaves more than one state of the electricity network consistent with the reading, a result power-system state estimation has worked with since the 1990s. Abur and Expósito (1997) showed that a measurement set containing current magnitudes can admit multiple solutions. Ju et al. (2018) carry the result into distribution networks with the remedy: where branch current-magnitude measurements "are the only ones to make the branch observable" the solution "is not unique". A current-magnitude measurement can therefore only sharpen an estimate that other measurements have already pinned down.

Two of the three published attempts we found rest on a second measurement of the same power. Bouman et al. (2024)'s Dutch substations carry the same limitation as NGED's MVA-metered substations, measuring only the absolute current. Bouman et al. recover the sign from a bottom-up load estimate built from smart meters, wherever the substation meter reads non-negative throughout while the bottom-up estimate goes negative. Western Power Distribution, NGED's predecessor, set out in the 2017 Time Series Data Quality NIA project to "first detect then assign directions to power flows where absent", and piloted a tool reconciling summed current at a substation's transformers against summed current along its feeders. The tool flipped a candidate feeder's direction where the two current sums disagreed by more than a threshold. In 2017, Time Series Data Quality also counted the circuits at stake across Western Power Distribution's licence areas: 204 in the South West licence area and 326 company-wide "experience reverse flows which are not apparent from the existing analogue values".

SSEN's TRANSITION, the third attempt and the closest to NGED's position, uses the meter's own history together with a model of the generation behind the meter. SSEN's TRANSITION met feeders metered in amperes only, where "the direction of the flow cannot be captured by the analogues", and settled the direction in three steps in the project's Load Forecasting Solution report: take the reading at face value where modelled generation is too small to have pushed the flow negative; flip the sign where modelled generation exceeds the average net demand that meter recorded over 4 years; then flip back wherever the recovered underlying demand comes out greater than net demand. The report gives no accuracy figure for the direction step, calling the result "a satisfying initial level of automated computation".

What this means for Flexpectation

Flexpectation version 1 does not attempt the recovery: version 1 forecasts each series in the unit its own meter reports, real power where the meter is directional and apparent power where the meter is not, and flags the affected series to NGED alongside the forecast. Forecasting the bounced trace is honest about what was measured. At a substation whose generation never reverses the flow, the apparent-power trace and the signed trace are identical. What forecasting the bounced trace cannot do is tell NGED whether a peak forecast at midday represents demand approaching the substation's import capacity or export approaching the substation's export capacity.

Flexpectation version 2 puts the meter's behaviour inside the model rather than repairing the series first. The differentiable-physics forward model reconstructs a substation's signed net flow from gross demand, metered generation, and unmetered generation, and compares the magnitude of the reconstruction against the apparent-power reading. The bounce is therefore predicted rather than removed. Recovering a signal from the magnitude of a transform of that signal is the phase-retrieval problem. Dong et al. (2023) describe phase retrieval as non-convex, because a signal satisfying the magnitude equation is always one of a family of solutions. An apparent-power meter takes the magnitude half-hour by half-hour, so nothing in the measurement couples one half-hour's sign to the next. The family therefore holds one member for every assignment of signs across the window rather than two members in all. Dong et al.'s uniqueness results do not rescue the problem either: those results turn on how far the number of measurements exceeds the number of unknowns. An apparent-power meter gives exactly one reading per unknown — the ratio at which Dong et al. call Fourier phase retrieval "fundamentally ill-posed as we only know amplitudes". Dong et al.'s own prescription for that regime is the prescription Flexpectation is following, to "leverage a priori information on the object". So what Flexpectation adds is not the formulation but the information that breaks the ambiguity — and that information carries the whole weight. The reconstruction's solar module has to track irradiance. A prior holds the direction of flow to persist for hours rather than flickering from one half-hour to the next.

Apparent power is the magnitude of real power only near unity power factor, and the approximation is weakest exactly at the bounce the reconstruction is trying to explain. As real power passes through zero, reactive power dominates the measured magnitude. The apparent-power trace therefore has a soft floor above zero rather than a clean reflection of the signed flow. The reconstruction will therefore under-fit the bottom of the bounce. The failure mode to design against is an optimiser that explains the soft floor with demand that was never there.

8. Disaggregating unmetered solar and wind from a substation's net flow

The challenge

Rooftop solar panels and small wind turbines appear only as a dent in a substation's net power flow. Recovering both the half-hourly output of that unmetered generation and its installed capacity, from the substation's net flow alone, is what we call disaggregation. Disaggregation is a different task from estimating how much of a metered generator's capacity is available today, which is challenge 3.

Distribution network operators (DNOs) do not know exactly how much capacity is installed. Ofgem's December 2025 consultation on asset visibility estimates that distribution network operators "are aware of less than half" of the consumer and distributed energy resources on their distribution networks. Ofgem attributes that figure to the Department for Energy Security and Net Zero's engagement with the operators rather than to a measurement, and that department's own footnote traces the figure to estimates the operators derived from other datasets, from sales volumes, and from grants processed. Each DNO's Embedded Capacity Register records generation of 50 kW and above and names the primary substation each site sits behind. But the capacity recorded is the export limit "permitted as per the connection agreement" rather than what a site can generate. Below 50 kW the register is silent, and that is where most of the panels are: of the 22,560 MW of solar photovoltaic capacity installed in GB by the end of July 2026, 8,503 MW — 38% of the total — sits in arrays smaller than 50 kW, spread across 2,058,822 of the 2,068,186 installations, according to the Department for Energy Security and Net Zero's solar deployment statistics.

None of the other registers fills the Embedded Capacity Register's gap either. Other registers exist, but none provides a complete picture. The Renewable Energy Planning Database tracks projects through the planning system and starts at 150 kW, a threshold lowered from 1 MW only in 2021. So smaller projects that cleared planning before 2021 may be absent. The Feed-In Tariff register closed to new applicants on 1 April 2019. And a domestic array reaches NGED only when the installer notifies NGED, as installers are required to do. None of these registers records the panel tilt, the panel azimuth, or the ratio of direct-current to alternating-current rating.

What the literature says

Splitting generation out of a substation's net flow has been done at GB primary substations, but in the GB projects we found that have published a result the generation was either metered or its capacity read from a register, rather than inferred from the net flow. Northern Powergrid's Artificial Forecasting models gross demand and customer export independently at primary substations, but that customer export is metered. The baseline Artificial Forecasting measures its customer-export models against is an extrapolation from Northern Powergrid's own Distribution Future Energy Scenarios, not a capacity inferred from the net flow. SSEN's TRANSITION split net load into demand and generation, forecast demand and generation separately, and recombined the two forecasts. TRANSITION's rooftop solar is not metered, but TRANSITION read each installation's capacity from a list of Feed-In Tariff installations. No register would carry Flexpectation as far, for the reasons set out under "The challenge" above.

Flexibility Market Asset Registration, the register now being built, sounds as though it should close this gap. The register will record the assets that trade flexibility, close to the complement of the arrays this challenge has to find. Ofgem appointed Elexon in 2025 to deliver Flexibility Market Asset Registration, digital infrastructure due by the third quarter of 2027 that will collect, store, and share data on assets participating in flexibility markets, aimed first at assets under 1 MW.

A rooftop array that never trades flexibility will never enter Flexibility Market Asset Registration, and those are exactly the arrays this challenge has to find. Ofgem's asset visibility consultation says the register collects data on assets "when they are first registered into a DSO or NESO flexibility market" — the market of a distribution system operator, or of the National Energy System Operator — so a rooftop array that never trades flexibility never enters the register, and the arrays this challenge has to find are the arrays that appear in none of the registers above. Where an asset does enter a flexibility market that NGED itself runs, NGED is the counterparty and already holds the data. What Flexibility Market Asset Registration adds there is one standardised record across markets rather than a generator NGED could not previously see.

We found published benchmarks for inferring capacity from the net flow, but those benchmarks work on individually metered premises, sit at a voltage level below NGED's, or do not say what aggregation they used. The one GB project we found doing the same at primary substations has not yet published a result. Gouveia et al. (2026) benchmark that inference at low-voltage substations serving 10 to 100 customers rather than at a primary. UK Power Networks' Power Flow to Solar Capacity project (with Open Climate Fix) infers solar photovoltaic capacity behind UK Power Networks' primary substations. Kanchana et al. (2026) separate load, photovoltaic generation, and energy storage from one aggregated net-load series, and report doing so "without requiring capital-intensive customer-level metering", which is NGED's position exactly. We hold the publisher's landing page for Kanchana et al. rather than the full text. That page names no customer count, no country, no time resolution, and no comparison method. So how far the reported errors of 8.14% for load, 5.12% for photovoltaic generation, and 11.51% for storage would carry to a GB primary substation cannot be judged from what we have read. The same page says a generative adversarial network fills gaps in the load measurements while a variational autoencoder generates synthetic photovoltaic profiles, and that "observed net-load profiles are assembled to create validation scenarios". So whether the mixture being separated is a mixture a meter recorded is a question the page leaves open.

Only one result we found separated solar from demand at a real primary substation without being told the installed capacity, and it used that substation's own reactive power. Kara et al. (2018) estimate the solar generation downstream of a substation in Riverside, California, from the substation's active and reactive power, and report a root-mean-square error of 6% of installed capacity across all sky conditions. The estimator is given neither the installed capacity nor the panel geometry. But the estimator does need one input NGED would struggle to supply at most substations: the metered output of a second photovoltaic plant, 4 miles away on a different feeder and low-pass filtered, standing in for irradiance. Regressing the load's active power on the measured reactive power is what makes the separation work: the simpler alternative, assuming the power factor measured at night holds through the day, is broken by the solar plant's own reactive power consumption. The preprint version of Kara et al. found that reactive power consumption responsible for about 25% of the overestimation at its peak.

Four features of Kara et al.'s setup limit how far the result carries to a GB primary substation. The generation behind that substation is a single 7.5 MW solar site, "the only generation asset located at this substation", rather than a fleet of differently-oriented rooftops. The ground truth comes from a second measurement device at that site's own point of interconnection. Kara et al. also had to detect and compensate capacitor-bank switching before the reactive power was usable. Kara et al.'s accuracy was still improving as the sampling rate rose to one sample every 5 minutes. A 5-minute sampling rate puts NGED's half-hourly data below the rate at which the errors settled. Kara et al. name the amount of photovoltaic capacity behind the substation, its volatility, and its spatial spread as factors whose effect on their method they had not studied — the three respects in which a GB primary substation differs most from their test case.

Uncertainty and a multi-day horizon each appear in the disaggregation work we found, but not in the same forecast. Zhang et al. (2022) attach uncertainty, disaggregating rooftop solar out of net load at grid supply point and feeder level with a multi-quantile recurrent neural network scored on reliability and sharpness. NESO's embedded wind and solar forecasts — "embedded" meaning generation sitting on the distribution network with no transmission metering, which NESO's own field definition describes as "invisible to the National Energy System Operator (NESO)" — are half-hourly to 14 days ahead, as a single number per half-hour with no uncertainty attached. A survey of behind-the-meter solar forecasting whose 162 references reach 2021, Erdener et al. (2022), judged that "the literature explicitly focused on uncertainty quantification within BTM [behind-the-meter] systems is immature", and recommended probabilistic approaches as the way to represent that uncertainty.

A second GB forecast of the same unmetered generation runs inside NESO's control room, and Open Climate Fix built the model covering its first 6 hours, so we have an interest to declare. NESO's Solar NowCasting project, run with Open Climate Fix under the Network Innovation Allowance, reports that the forecast it produced "was 2.8 times better than our previous Photo Voltaic (PV) forecast (for forecasts up to two hours ahead)", and that the first fully operational service reached NESO's control room in December 2022. That operational service runs PVNet, Open Climate Fix's model, at 0 to 6 hours ahead and blends PVNet with other models beyond 6 hours. Open Climate Fix now supplies the same forecast as its Quartz Solar product, which NESO uses in its control room. The project's own record on the Smarter Networks Portal reports "Accuracy improvement over the previous model by approximately 30% for the GSP and National forecasts (4-8 hours)" and lists "Probabilistic forecasts for all horizons" among its outcomes.

This challenge says one combination is missing: unmetered generation, forecast probabilistically, at a spatial level below the country. The GB work we found has built that combination once — at grid supply point level rather than at primary substations, and for solar rather than for net demand. How NESO builds the embedded solar forecast it publishes is a separate question we cannot answer: NESO runs more than one solar forecast, and the published series does not name the model behind it.

Two figures are quoted for what that national solar forecast is worth, and both are rough approximations rather than audited results. National Energy System Operator states that a better solar forecast avoids around £30 million a year in imbalance cost, rising to as much as £150 million a year by 2035 at the government's solar target, and Open Climate Fix states that the same forecast avoids around 300,000 tonnes of carbon dioxide a year. Both figures describe one system and one forecast — Great Britain's national electricity system, and the Quartz Solar forecast Open Climate Fix supplies and NESO runs in its control room — and both are annual rates resting on the same halving of NESO's solar forecast error. Neither figure has an independent audit published, and neither figure is a substation-level result. So both belong here as an indication of the scale of the prize rather than as a measurement Flexpectation can build on.

The model behind that service is published and open source, which is unusual in this literature, and its limits are what keep Flexpectation's challenge open. PVNet is released under the MIT licence and described by its authors as "a multi-modal late-fusion model for predicting renewable energy generation from weather data", combining numerical weather prediction with satellite imagery, recent generation, and the sun's position. The accompanying paper, Fulton et al. (2024), makes "0-8 hour lead time forecasts for grid regions across Great Britain" and limits the model's inputs "to be reflective of those available in a live production system". Fulton et al. is a workshop paper rather than a peer-reviewed paper, which we flag because the rest of this review holds its sources to that standard, and one of the paper's authors is an author of this review. PVNet forecasts generation at grid supply points where the generation is the whole signal, whereas Flexpectation has to separate unmetered generation from demand inside a single net-flow measurement at a primary substation. PVNet's horizon is hours rather than the 14 days NGED needs. And the grid supply point regions PVNet forecasts are far larger than a primary substation, so PVNet's accuracy figures say nothing about how the same approach would perform at Flexpectation's scale.

An uncertainty estimate is useful only if the estimate widens where the answer gets worse. Only one substation-level disaggregation we found tested for that widening, and that study reports the widening holding — until the generation pattern is unlike any pattern in the training data. Yi and Wang (2022) summarise their two journal papers on disaggregating behind-the-meter solar at substations, and pose the problem as one of partial labels: for some aggregate measurements the operator knows which load types are present, but never their individual values. Yi and Wang's Bayesian dictionary-learning estimator reaches a total error rate of 8.97%, against 20.61% and 37.12% for two methods that need fully labelled training data. The estimator's error weighted so that the estimates the method is unsure about count less — 0.13 to 0.16 — comes out far below the estimator's unweighted root-mean-square error of 5.19 to 6.20. Yi and Wang read that gap as showing that the estimates carrying the largest errors are also the estimates carrying the largest uncertainty. Where the test period's solar pattern is unlike any pattern in the training data, however, Yi and Wang report that the true load may fall outside the 99.7% confidence interval. A generation pattern unlike any pattern in the training data is the failure mode that matters most to Flexpectation, whose substations will carry generation mixes that no training substation had. The validation runs on 360 generated training samples covering two industrial loads and one solar generation, not on measurements from a real substation, and carries no forecast horizon.

The survey of behind-the-meter solar forecasting by Erdener et al. tabulates net-load disaggregation studies that run either at individually metered premises or at a whole balancing area, with no study at the aggregation level of a primary substation. Erdener et al. (2022) tabulate eight studies that recover photovoltaic capacity, panel tilt, or panel azimuth by disaggregating net load. Seven work on individually metered premises — two photovoltaic plants in Switzerland, and customer sets of 40, 100, 183, 197, 300, and 1,300 — and the eighth works on the zone of Independent System Operator New England that covers the state of Maine. A GB primary substation sits between those two aggregation levels, at a level Erdener et al.'s table does not cover.

The smart-meter literature on disaggregating rooftop solar is larger than the substation literature, but the smart-meter work sits at individual premises rather than at a substation, and leans on a neighbouring-customer comparison that substation telemetry cannot support. Cheung et al. (2023) use the consumption patterns of neighbouring customers known to have no panels, which substation telemetry cannot observe. Cheung et al. are also the one study we found that varies the aggregation count on measured household data: across 5, 10, and 20 Australian customers per aggregated series, their own method's solar mean absolute scaled error stays between 1.02 and 1.28 — around the average change between consecutive readings — with solar mean absolute percentage error of 21 to 25%. Cheung et al. report that both measures stayed "almost the same as aggregation level varied". The Kara-derived baseline they re-implemented for comparison degraded instead, from a mean absolute scaled error of 1.47 at 5 customers to 2.20 at 20, and from 28% to 43% mean absolute percentage error. Results reported elsewhere at an aggregate level are usually sums of individually metered households rather than a measurement taken at a real aggregation point. The smart-meter literature therefore stops far below the thousands of customers behind a GB primary substation.

It is unsettled whether more customers behind a substation makes the estimate easier or harder. Only one study we found varies the aggregation count on a simulated feeder, and that study does not settle the question. Tang et al. (2024) estimate installed photovoltaic capacity from 24-hour net-load curves for feeders of 20 to 80 London households, and report "a general trend of increasing RMSE [root-mean-square error] values as the number of households increases". The rising root-mean-square error is weaker evidence than the trend first appears: the error is in kilowatts against a total capacity that itself rises with the household count, the percentage error moves the other way, and the trend reverses sharply between 70 and 80 households. The load and the household count are real, but the solar is simulated at three azimuths, 45°, 0°, and −45°, all of them southerly. So the study has none of the north- and east-west-facing roofs a real street would carry.

What this means for Flexpectation

UK Power Networks' "Power Flow to Solar Capacity" project is highly relevant to Flexpectation: the project works on the same kind of GB data, and Open Climate Fix is a partner in both projects. UK Power Networks' Power Flow to Solar Capacity (2024 to 2026, £0.4 million) infers the capacity of unmetered solar sitting behind each primary substation from half-hourly substation load and weather, then forecasts that generation. Open Climate Fix is a partner in both Power Flow to Solar Capacity and Flexpectation. So what Power Flow to Solar Capacity found about inferring solar capacity from GB primary-substation data reaches Flexpectation directly rather than only through what has been published.

The nearest published method we found splits unmetered wind and solar out of substation measurements, but needs each site's installed capacity. Teng et al. (2023) train on 10 Dutch substations that carry complete renewable metering, then predict solar and wind power separately at substations with none, from the substation's measured total load, weather, geospatial position, and each site's known renewable capacity, at 15-minute resolution. Teng et al. report a root-mean-square error of 0.07 against 0.70 for a default transfer-learning model, on a min-max-scaled target. The 0.07 reads as 7% only if the scale runs from 0 to 1, and Teng et al. do not say what the scaling divides by, so the figure does not transfer to another dataset. The 0.07 should not be read as achievable here: Teng et al. are told each site's capacity, whereas inferring the capacity is half of what Flexpectation plans to achieve. Teng et al. call their method domain adaptation for zero-shot learning in sequence, or DAZLS, and benchmark DAZLS against "the energy splitting model in the OpenSTEF software package" on the same data, reporting that DAZLS "significantly outperforms" the OpenSTEF splitter. That comparison is worth noting because two of the authors work at Alliander, the operator that builds OpenSTEF. The comparison is therefore a like-for-like statement from the team that maintains both DAZLS and OpenSTEF. The OpenSTEF splitter is the operational relative of this challenge, and a far simpler splitter than DAZLS.

Inferring the capacity from the net flow alone has been measured, at a smaller scale than NGED's. Gouveia et al. (2026) benchmark data-driven against model-based estimators of the photovoltaic capacity installed behind a low-voltage substation, working from the net load and irradiance series. Gouveia et al.'s substations serve 10 to 100 customers, against the thousands behind a GB primary.

Gouveia et al.'s two transferable results are that data-driven estimators beat model-based estimators on noisy data, and that a model trained in one country held under 5% mean absolute percentage error in two others. The data-driven estimators matched the model-based estimators on clean data, and beat the model-based estimators clearly on noisy data — the condition distribution telemetry, NGED's included, is often in. And models trained on a Belgian dataset, then applied unseen to American and Australian datasets with only approximate irradiance, stayed under 5% mean absolute percentage error once the linear models were regularised. What Gouveia et al.'s estimators produce is a capacity figure rather than a forecast. What Flexpectation adds is therefore putting the capacity estimate inside a probabilistic multi-day forecast, and disaggregating the full shape of the unmetered generation.

GB already has an operational forecast of unmetered generation, but only at national scale and without uncertainty. NESO's embedded wind and solar forecasts, described under "What the literature says" above, match the resolution and horizon Flexpectation is specified to deliver, but cover GB as one region rather than substation by substation.

Observational cosmology and systems biology have separated superposed signals for decades, and both give the same warning: a small residual against the measured total is not evidence that the components were separated correctly. Fitting a sum of physically parameterised components to one measurement is routine in observational cosmology, where the technique is called component separation. Hensley and Bull (2018) show that giving the nuisance component too simple a model biases the component of interest, and that the fit does not announce the bias: "models that are strongly biased but still yield low χ² values are the most dangerous". Two consequences follow for Flexpectation. Effort spent making the demand model richer is justified on separation grounds even where the richer demand model does not improve the fit to the substation's net flow. And the diagnostic to watch is the joint distribution over the components rather than the residual. Working in systems biology, Wieland et al. (2021) add the matching warning about uncertainty: a differentiable model gives these confidence intervals, read off the curvature at the optimum, with little extra computation. The intervals are "insensitive to practical non-identifiabilities" and can look reassuringly finite for a parameter the data do not constrain at all.

Fitting a differentiable physical forward model to measurements is routine in exploration geophysics, and that field reports that the order in which the fit admits fine detail decides whether the fit converges at all. Full-waveform inversion recovers the properties of the rock beneath a seismic survey — chiefly the speed at which sound travels through each point of the subsurface — by simulating the seismograms those properties would produce and adjusting the properties until the simulation matches the recording. Full-waveform inversion is the procedure Flexpectation version 2 applies to a substation's net flow. Virieux and Operto (2009) report a failure mode the field calls cycle skipping: because a seismogram oscillates, a starting model that mis-predicts an arrival by more than half a period leads the optimiser to match the wrong cycle, and "the so-called cycle-skipping artifacts will lead to convergence toward a local minimum". The remedy Virieux and Operto report as standard practice is a multi-scale schedule that inverts the low frequencies first, "because low frequencies are less sensitive to cycle-skipping artifacts", then admits successively higher frequencies, each stage starting from the model the previous stage produced. The arithmetic relating frequency to recovered detail belongs to wavefields and does not carry to a half-hourly power series, but the local-minimum mechanism does. A substation's net flow is periodic on a daily cycle, so fitting the slowest-varying structure of each component before admitting half-hourly detail is the transferable precaution.

Whether two simultaneously fitted components can be told apart is a property of the measurements rather than of the optimiser, and exploration geophysics tests for that coupling before fitting and orders the fit around the answer. Virieux and Operto (2009) report that adding a second class of physical parameter makes the problem more ill-posed, because "more degrees of freedom are considered in the parameterization" and because "the sensitivity of the inversion can change significantly from one parameter class to the next", and that "different parameter classes can be more or less coupled as a function of the aperture angle" — the angle at which a source and a receiver view the same point in the rock. Coupling of that kind "can be assessed by plotting the radiation pattern of each parameter class", so the field tests separability in advance rather than discovering it afterwards. Where the speed of sound and the rock's density carry the same signature at short apertures, "these two parameters are difficult to reconstruct from short-offset data", and Virieux and Operto cite a study concluding that the speed of sound and the rock's absorption "cannot be imaged simultaneously from short-aperture data". The response Virieux and Operto report is to order the parameter classes rather than run the joint fit for longer: one study they cite recommends fitting the speed of sound, denoted VP, alone first and fitting the speed of sound with the absorption jointly second, "because the reliability of the attenuation reconstruction strongly depends on the accuracy of the starting VP model". One limit rides along: a seismic survey chooses where to put its sources and its receivers, and the long offsets and wide apertures that break the coupling are a design choice, whereas Flexpectation takes the telemetry NGED already collects. What transfers is therefore the test for coupling and the fitting order rather than the survey design.

Hyperspectral unmixing leans for identifiability on each component being observed alone somewhere, and Flexpectation can test whether that condition holds before fitting a forward model. Hyperspectral unmixing, which splits one mixed image pixel into the spectra of the materials composing it, calls the condition the pure-pixel assumption: for each component there is at least one observation containing only that component. Bioucas-Dias et al. (2012) set out a weaker sufficient condition too, and show what happens when neither holds — on a highly mixed data set with no observations near the extremes, the fitted simplex comes out smaller than the true simplex, so the recovered components are biased rather than merely uncertain. Flexpectation's pure observations are the half-hours when nature switches one component off — night for solar, calm hours for wind, and the substations carrying no embedded generation at all. Whether a given substation has those half-hours is a question the telemetry can answer on its own.

9. Disaggregating other distributed energy resources: heat pumps, electric-vehicle chargers, and batteries

The challenge

Heat pumps, electric-vehicle chargers, and price-sensitive domestic batteries change the shape of a substation's load in ways a model trained on history cannot anticipate, because the number installed behind any given substation is growing quickly. A Flexpectation stretch goal is to disaggregate and forecast heat pumps, electric-vehicle (EV) chargers, and batteries separately rather than leaving them inside net demand.

The number of each installed grows fast enough to matter within Flexpectation's own lifetime. Every figure below is the Holistic Transition pathway of NESO's Future Energy Scenarios (FES):

What 2024 2030
Battery-electric cars on the road in GB 1.4 million 8.2 million
Heat pump stock in GB 361,000 2.3 million
Battery storage below 1 MW in GB 191 MW 975 MW
The same class in the 2024 scenarios, allocated to the four grid supply point groups NGED serves 49 MW 308 MW

Electric cars and heat pumps each rise roughly sixfold over those 6 years, and the battery figure cannot be checked against the Department for Energy Security and Net Zero's own installation count. "Below 1 MW" is NESO's own class boundary, and NESO defines the class as generation and storage under 1 MW that therefore "includes some larger commercial installations". So the class is wider than the batteries that fall below the Embedded Capacity Register's 50 kW floor. How much of the 191 MW sits below that floor is a question the FES scenarios cannot answer. Ofgem's asset visibility consultation says the Microgeneration Certification Scheme's (MCS) installation database certifies battery storage up to 50 kW. The Department for Energy Security and Net Zero (DESNZ) publishes counts drawn from the MCS database: 73,987 domestic retrofit battery installations between September 2023, when the series starts, and March 2026, rising from 24,242 in the 2024/25 financial year to 44,033 in 2025/26, with the 72,459 installations that passed outlier testing holding 666,880 kWh between them. The DESNZ figures cannot be subtracted from the FES figure of 191 MW, because the DESNZ figures measure stored energy rather than power, and count domestic retrofits rather than every installation below 50 kW, and accumulate from September 2023 rather than reporting a stock. The DESNZ series carries no projection, so the 2030 figure has no counterpart outside the FES scenarios.

Scaling the GB figures to NGED is only approximate, because NESO holds NGED's share of GB battery capacity fixed at 36% across every forecast year. NESO's regional breakdown of the 2024 scenarios allocates 36% of GB's sub-1 MW battery capacity to the four grid supply point groups NGED serves, and holds that 36% constant across every forecast year. The regional number is therefore an allocation of the GB total rather than a forecast of NGED's own area. The 2024 regional breakdown's own GB totals are lower than the figures in the table above, at 138 MW for 2024 and 863 MW for 2030. So the 49 MW and the 308 MW are 36% of those totals rather than of the 191 MW and the 975 MW. NGED serves nearly 8 million customers, a little over a quarter of the GB total.

What the literature says

Heat pumps, EV chargers, and batteries have the fewest directly relevant papers we could find of the nine challenges. The one study we found that measures EV charger forecast skill against aggregation, Ostermann and Haug (2024), forecasts charging demand over a 24-hour horizon at 15-minute resolution from 350,000 charging processes at more than 500 locations across Germany, and repeats the exercise at four aggregation levels: the individual site, the postal code, the transmission system operator's zone, and the whole portfolio. Eight machine-learning and deep-learning models are set against a naive benchmark that predicts the average of the same quarter-hour on the same weekday. Of the five individual sites, only the site with 145 charge points beat that benchmark by a clear margin. At the site with 8 charge points some models beat the benchmark, and at the sites with 3, 4, and 14 charge points none did. We found no papers discussing whether domestic batteries responding to a common price signal average out as more are added. On the topic of heat pumps, we found no measurement of heat-pump diversity in cold weather. The volume of work on these three resources is easy to demonstrate. Our searches were framed around substation and generation forecasting rather than around electrification. Of the 305 papers accepted for the Brussels workshop of June 2026 held by the International Conference on Electricity Distribution (CIRED), 28 have a title naming electric vehicles, chargers, heat pumps, or batteries — more than the 23 whose titles name forecasting or prediction at all.

Charger errors cancel rather than compound as more chargers are added, and the measurement is NGED's own. The Electric Nation trial, run by NGED under its former name with 673 participants and over 130,000 charging events, fits the demand of a group of chargers as Group Demand = N·P + Q√N, where P is the mean demand per charger and Q the deviation. The mean scales with the number of chargers and the deviation only with its square root, so relative uncertainty falls as more chargers are added. Bollerslev et al. (2022) simulate Danish driving and plug-in behaviour on synthetic feeders and fit the exponent at between 0.42 and 0.51 across battery sizes and plug-in behaviours at an 11 kW charger, against the 0.5 that complete independence would give.

What makes electric-vehicle charging the harder problem for a distribution network is when the charging lands, and an automated tariff can re-synchronise a population that had diversified. In Electric Nation's third trial, with a time-of-use tariff, the share of charging events starting in the 22:00 hour rose from 5.8% without the tariff to 24.7% with the tariff. Among participants using the smart-charging app the share reached 37.6%, against 5.5% for participants on the same tariff who did not use the app. Nothing we found tests whether that re-synchronised peak survives at the aggregation a primary substation carries.

Heat pumps diversify in an average winter, but whether that diversity survives the cold weather that actually matters is untested in the work we found. Love et al. (2017) measured around 700 domestic heat pumps in GB and found demand per heat pump falling from 4.0 kW for a single unit to 1.7 kW once 275 are aggregated, with the spread between samples falling from 1.5 kW to 0.1 kW. But a heat pump sized small relative to a house's heat demand runs flat out for hours in cold weather. Northern Powergrid's code of practice for the economic development of the low-voltage system notes in a footnote that "further research is required to examine whether the increase in duty cycle (and hence average demand) with lower than average winter ambient temperatures is material when designing a LV system" — which is precisely the condition under which a substation approaches its limit.

No diversity factor helps for domestic batteries, and a GB network operator's own code of practice agrees. That same code of practice fits diversity curves to measured trial data for general domestic load, heat pumps, and chargers alike, and then states that diversity "should not be applied when considering a BESS device" — a battery energy storage system — a diversity factor of exactly one.

We ran a targeted literature search for disaggregating heat pumps, chargers, and batteries from substation measurements, and found the work split by asset, with "substation level" used for aggregations far smaller than a GB primary substation. Gao et al. (2024) disaggregate thermostatically controlled loads — air conditioners, heating and ventilation units, and furnaces — from an aggregated residential load by contrastive sequence-to-point learning, and generalise the same model to photovoltaic generation and electric-vehicle charging. The 8.78% Gao et al. report is a mean absolute percentage error between the estimated thermostatically controlled load and the metered thermostatically controlled load. The denominator is therefore the appliance load being recovered rather than the aggregate the load was pulled out of. Gao et al. give 8.78% as a best case for the bi-directional model structure and 11.26% as a best case for the unidirectional structure, and name no baseline either structure is measured against. Gao et al.'s aggregate is the sum of the Pecan Street dataset's 25 individually metered homes in Austin, Texas, and 25 in New York — two orders of magnitude below the thousands of customers behind a GB primary substation, and a sum of household meters rather than a measurement taken at a real substation.

The electric-vehicle and battery papers we found each needed the individual assets metered live, and neither ran on a feeder-head load a meter actually recorded. Ebrahimi et al. (2022) split electric-vehicle charging out of a feeder-head load hour by hour, and needed the charging power and stored energy of 19 vehicles metered live, alongside the hourly energy price and the ambient temperature, to do it. The feeder-head load Ebrahimi et al. worked on was assembled rather than measured: hourly demand published by Independent System Operator New England for its Connecticut zone, peaking at about 1,455 kW, added to a simulated charging schedule built from the plug-in and plug-out records of 201 Nissan LEAFs in the My Electric Avenue trial. Wang et al. (2022) separate behind-the-meter photovoltaic generation and battery charging jointly by contextually supervised source separation — the method family Kara et al. extended for solar under challenge 8. That shared method family suggests the battery problem and the solar problem are the same problem with another component added. Of the Gao et al. and Wang et al. papers we hold Gao et al.'s abstract, introduction, and dataset description from the publisher's landing page, and Wang et al.'s abstract. The Wang et al. citation therefore carries no more weight than the existence of the work. Both the Gao et al. and the Wang et al. full texts are closed.

The two heat-pump disaggregation papers we obtained separate a heat pump from the total load of the single premises the heat pump sits in, and the widest aggregate either paper reports is five households. Brudermueller et al. (2023) estimated the heat pump's own 15-minute energy in 363 Swiss single-family houses, each fitted with a second meter on the heat pump and none fitted with photovoltaics. Brudermueller et al. explained 83% of the variance in that second meter's readings across households held out of training, against 63% for the better of the two published baseline algorithms in the comparison. Gisiger et al. (2026) ran the same task over 7,021 Swiss premises with heat pumps through one heating season of 15-minute readings. Gisiger et al. summed the estimated heat-pump load of five of those premises, drawn at random from the dataset and treated as sharing one transformer, to within 6% of the metered total over an evening peak of 17:00 to 21:00. Gisiger et al.'s error, normalised by the mean metered heat-pump load, rose from 0.69 on the Swiss data the model was trained on to 0.78 on Brudermueller et al.'s separate Swiss dataset, and to 1.24 on a German dataset of single-family houses with heat pumps. Gisiger et al. attribute that rise to differences in heat pump types, building stock, occupancy patterns, and data collection methods. The rise is measured evidence that a heat-pump model does not survive a change of dataset unaltered, with a change of country degrading the model further still.

What this means for Flexpectation

For Flexpectation version 1: heat pumps, chargers, and batteries stay inside net demand rather than being forecast separately. In the one measurement we found, the only site size that clearly beat a naive benchmark 24 hours ahead was 145 charge points. Forecast uncertainty grows with lead time. So over the 14 days NGED needs, a site would probably have to be larger than 145 charge points before a separate charger forecast was worth making.

Compared to the literature we found, Flexpectation version 2 plans to invert which half of the problem is learned: the disaggregation methods in this challenge's literature learn each resource's signature from premises where that resource is separately metered, whereas we plan to use a differentiable physical model of each distributed energy resource where we write the physics in code and learn only the parameters. Gao et al. train on the Pecan Street homes' individually metered appliances, Brudermueller et al. on 363 Swiss houses each fitted with a second meter on the heat pump, and Gisiger et al. on 7,021 Swiss premises with heat pumps, while Ebrahimi et al. need the charging power and stored energy of 19 vehicles metered live. Exogenous inputs already appear in that work: Ebrahimi et al. take the hourly energy price and the ambient temperature, and Gisiger et al. find detection easier in colder weeks. But each paper uses those inputs as features feeding a mapping learned from metered examples. So what the model knows about a heat pump is what heat pumps looked like in the training set. A differentiable physical model states the relationship between outdoor temperature and heat-pump electrical demand as an equation instead, and fits the building's parameters. The measured argument for that inversion sits in this challenge's own literature: Gisiger et al.'s error, normalised by the mean metered heat-pump load, rose from 0.69 on the data the model was trained on to 1.24 on a German dataset, which Gisiger et al. attribute to differences in heat pump types, building stock, occupancy patterns, and data collection methods — differences that change a learned signature but not the equations behind that signature.

The inversion also supplies the separability an aggregate needs, because each distributed energy resource answers to a different driver: outdoor temperature moves the heat pumps, irradiance and panel orientation move the solar, and price moves the batteries. Whether a given substation's mixture is separable in practice is a property of the measurements rather than of the optimiser, which is the test challenge 8 borrows from exploration geophysics and from hyperspectral unmixing. Two limits ride along: the thermal physics of the thousands of premises behind a primary substation is not one building's physics repeated. And a differentiable physical model removes the need for submetered training examples at every substation without removing the need for submetered ground truth to check the answer against.

The spiky, synchronised charging that makes electric-vehicle load hard to forecast is what makes that load easy to detect in aggregate, while heat pumps are hard to detect at all. Northern Powergrid's smart-meter detection trial, on 1,500 monitored premises, found that "EV [electric vehicle] identification at premises level was found to be relatively straightforward", though "a lack of ground truth, such as registered charging points, precluded formal validation". The trial also found that "aggregation does mask some signals, although EV usage is still clearly identifiable at feeder and substation level". The same trial found that "the detection of ASHP [air-source heat pumps] is frustrated by the low levels of adoption (<1% of premises) and differences in operation (low-slow vs high-fast)". Gisiger et al. (2026) detected a heat pump at a single premises from one week of 15-minute readings with a precision of 0.896 by a rule counting sharp rises in power and 0.953 by a convolutional neural network. Gisiger et al. also found detection easier in colder weeks, when heat pumps run more. Gisiger et al. assembled that evaluation set around premises known to have heat pumps without reporting how many premises without a heat pump the set held. The precision therefore does not carry to the fewer than one premises in a hundred Northern Powergrid was searching.

Three published results that point against this project's plan

Three results in this literature point against Flexpectation's plan. We intend to test all three rather than avoid them.

More detailed weather data has not always improved the forecast

Adding the spread of the weather across a region, on top of that region's average weather, improved forecast skill significantly in only 2 of GB's 14 grid supply point group regions, and made forecast skill significantly worse in 3. Browell and Fasiolo (2021) forecast day-ahead net load — demand minus embedded generation — half-hourly for each of GB's 14 grid supply point groups over 2014 to 2018, from the European Centre for Medium-Range Weather Forecasts' high-resolution run issued at midnight and available around 06:00 UTC. The model Browell and Fasiolo were trying to improve already used the weather across the whole of each region: wind speed and solar irradiance averaged over the numerical weather prediction cells covering the region, alongside a temperature forecast taken from the single cell of highest population density. What Browell and Fasiolo added on top was the spread of the weather across those cells — the spatial standard deviation, minimum, and maximum of the gridded fields. Measured by the Diebold-Mariano test against the same model without the spread features, adding the spread improved the pinball score significantly in 2 of the 14 regions, worsened the pinball score significantly in 3, and made no significant difference in the remaining 9. Browell and Fasiolo report that cross-validation had suggested a small gain which "is not consistently reproduced on test data and therefore inconclusive". Browell and Fasiolo conclude that gridded numerical weather prediction "does not appear to add significant value to deterministic and probabilistic net-load forecasts in the present framework" — while allowing that "it is possible that other forecasting methods would be able to extract value from this data by constructing different features".

The question this result puts to Flexpectation is not whether weather matters, but whether the spread of the weather across a region does. Weather itself mattered a great deal to Browell and Fasiolo's model: adding the regionally-averaged wind and irradiance to a model carrying only calendar features and the point temperature cut the pinball score — the single-quantile equivalent of the continuous ranked probability score — by 40% overall, by 60% in North Scotland, where embedded wind capacity exceeds peak load, and by 10% in Greater London, where there is little embedded generation.

Northern Powergrid's Artificial Forecasting project found the same pattern with finer weather data. Artificial Forecasting obtained postcode-level weather forecasts for two wind-connected primary substations after the deliverable reported that the project's wind-connected models had performed poorly. Artificial Forecasting found that the postcode-level forecasts "did not notably improve model performance". The deliverable nonetheless names better weather data as a next step, without saying what would be better than postcode-level.

One published result points the other way, and the result separates the ensemble from the resolution rather than confirming that finer weather data helps. Dantas and Browell (2026) forecast 73 wind farms in GB from the ECMWF deterministic model, taken from the operational archive at full spatial and temporal resolution, and from the ECMWF ensemble, taken from the THORPEX Interactive Grand Global Ensemble (TIGGE) archive, which "only stores a limited set of atmospheric variables at 6-hour/0.5° resolution". The deterministic-driven model beat the ensemble-driven model "for all forecast horizons and case studies". Dantas and Browell then constrained the deterministic model to the ensemble's 6-hourly steps and single 10 m wind level, leaving the horizontal grid untouched. The advantage held at the offshore farms and at only some of the onshore farms. Dantas and Browell read that pattern as evidence that "all ENS-based methods presented here could be improved with access to the full-resolution model data". Two limits govern how far the result carries to Flexpectation. The comparison never matched the horizontal grids. So the comparison does not separate the ensemble from the resolution as cleanly as the temporal constraint does. And once the ensemble was calibrated, the ensemble-driven method overtook the deterministic method "from around 4 days ahead" at onshore farms and "from 3 days ahead" offshore, with some offshore farms best served by the ensemble at every horizon. Flexpectation forecasts 14 days out.

Weather improved low-voltage forecasts less than expected in the past

Adding a temperature input made little difference to low-voltage feeder forecasts, and made three of the tested methods worse. Haben et al. (2019) forecast the demand of 100 real low-voltage feeders in Bracknell up to 4 days ahead at hourly resolution, and ran each of their methods twice, once without a temperature input and once with. Both the observed temperature and a temperature forecast were tried, taken from a weather station about 16 km away. Adding temperature had "minimal effect on the forecast accuracy", and for three of the methods adding temperature made the forecast worse. Haben et al. suggest that temperature is correlated with seasonality, so a model can end up "erroneously training on the temperature as a surrogate for seasonality". The feeders carried an average of 45 households each, and the data were collected in 2014 and 2015. We expect how much weather matters at a substation to be changing quickly, because embedded solar generation and heat pumps make a substation considerably more weather-dependent. There are far more of both on the distribution network now, in 2026, than there were in 2015. That expectation is a prediction, though, not a measurement. And the Scottish primary-substation sensitivities of Fox et al. (2018), measured on the 10 years of weather and substation data before its publication, say weather was already moving primary substation demand well before the mid-2010s.

A model trained on none of NGED's data may match a model trained on all of it

A general-purpose model that had never seen the German feeder data beat every model trained on 160 of those feeders, 3.8 kW against 4.2 kW on mean absolute error. Kaas et al. (2026) tested Chronos-2, a general-purpose time-series model that had never seen the German feeder data, against models trained on the first 160 of their 200 German low-voltage feeders and scored, like Chronos-2, on all 200 feeders. Chronos-2 beat every purpose-trained competitor on mean absolute error, 3.8 kW against 4.2 kW. The authors describe the purpose-trained models as lightly engineered, and challenge 1 above found only a modest return to model sophistication. But the comparison still tells us how much any programme of heavy engineering is likely to improve accuracy: a model trained on the feeders' own history was beaten by a model trained on none of that history.

Two limits keep that one result from settling the question. The margin is a single number: Kaas et al. report the median across the 200 feeders, 3.839 kW against 4.184 kW for the best purpose-trained model, with no confidence interval, no repeated runs, and no test of whether a gap of that size could have arisen by chance. And the margin rests on the weather covariates as much as on the pre-training, because the same paper's ablation, which strips the covariates from the foundation models but not from the purpose-trained models, puts Chronos-2 at 4.813 kW — behind the purpose-trained model that keeps the covariates.

If the result does hold on NGED's substations, Flexpectation still delivers, and the finding is worth having independently. A forecast is one component of what this project builds: the ingest, the contracts, the degradation ladder, the leaderboard, the delivery tables, and the live service all stand whichever model wins. A pre-trained model that beat a purpose-trained model would simply be the model the leaderboard promoted. Establishing which of the two is better on a distribution network, measured against a common protocol on a real operator's telemetry, is a research result no study we read has published. A network operator deciding whether to train its own models would want the answer. Flexpectation is an innovation project, and a well-measured negative result is one of the outcomes Flexpectation exists to produce.

Evaluating the performance of power forecasts

The nine challenges above need three different kinds of evaluation: scoring a forecast, checking an estimate of a quantity NGED does not meter, and scoring the detection of a rare event. The literature is far stronger on scoring forecasts than on the other two kinds of evaluation. Scoring a forecast has settled practice Flexpectation can adopt. Checking an estimate of a quantity NGED does not meter — the effective capacity of a metered generator, the half-hourly output of unmetered solar, and the direction of flow behind an apparent-power meter — has no ground truth to score against. Each quantity needs its own basket of substitutes. For disaggregating unmetered solar we identified six possible substitutes, of which the disaggregation papers we read use three. Scoring the detection of a rare event, such as a switching event or a metering fault, has good academic practice. None of the GB projects we checked published a number to compare against.

Mean absolute error rewards a flat forecast that would be of little use for either flexibility procurement or curtailment decisions, so a peak-aware score belongs alongside a proper score rather than instead of a proper score. A forecast that predicts the right peak at the wrong time is penalised twice by mean absolute error — once for the peak it predicted that did not happen, and once for the peak that did happen and the forecast missed. A flat, featureless forecast avoids both penalties. Meteorologists named that effect the "double penalty". The meteorologists' conclusion transfers to substation forecasting: a score that forgives a peak predicted an hour late is generally no longer a proper scoring rule — a score whose expected value is optimised when the forecaster publishes the predictive distribution the forecaster actually believes, so that no hedged forecast, flatter or later-peaking, scores better on average. The same argument runs at the other end of the distribution: curtailment turns on the half-hours of deepest export, and a flat forecast hides the deepest export half-hours too.

Two teams independently concluded that mean absolute error was the wrong measure for peaks. Pinheiro et al. (2023) adopted a peak-aware error measure for exactly this reason. Artificial Forecasting built a metric over the top 10% of demand values, made that metric its primary measure for comparing models, and reported the metric both against actual demand and normalised to transformer rating.

Only one metric in the work we found holds risk constant and prices the forecast in money at distribution level, and that metric was published on a synthetic distribution network rather than a real one. Bernecker et al. (2025) fix at 95% the confidence level at which a network operator acts, and compare what two forecasts cost that operator in congestion management: 3,102 euros a year using standard load profiles against 86 euros using a smart-meter-informed forecast, a 97% reduction, alongside a 90% fall in the number of voltage violations. Bernecker et al. also give the sensitivity NGED would want: a 1% cut in the standard deviation of forecast error is worth about 1.4% of congestion-management cost on average across rollout levels. The saving varies between rollout levels, though, and is negative at some of them. We read the sections of that paper bearing on the cost calculation rather than the whole of it. Two features of the study keep the gap open: the modelled distribution network is a modified IEEE 33-node test system rather than a real distribution network, and what Bernecker et al. compare is two information levels rather than two forecasting models. We found no case of the metric being used to rank one forecast against another at a real substation.

The rest of that decision metric exists in pieces, and the piece still missing is the price on a real distribution network. Browell and Fasiolo (2021) fix a risk appetite, compute the reserve volume each forecast would need to hold that risk, and compare — the harder half of the job, done across whole grid supply point groups. Angus et al. (2027) bring the same idea down to individual assets, forecasting day-ahead how hard each of 644 low-voltage transformers in the UK can safely be pushed, and winning 10 to 12% more capacity than a fixed setting while the risk of overheating came out at whatever percentile they asked for. We read Angus et al.'s preprint rather than the published paper. Meteorology has priced forecast decisions this way for decades: Richardson (2000) computed the relative economic value of the ECMWF ensemble across the whole range of ratios between the cost of acting on a forecast and the loss avoided by acting. Every published version of that curve we found on a real distribution network, though, is denominated in energy volumes or in spare capacity rather than in money.

A forecast can state its own uncertainty badly without any accuracy score revealing the fault. Kaas et al. (2026) scored models on 200 German low-voltage feeders with an overload-decision metric evaluated at each model's 95th percentile for consumer peaks and each model's 5th percentile for producer peaks. The two models that came first and second on consumer peaks in the quantile variant of the overload-decision metric — Chronos-Bolt, a time-series foundation model, and a weekly-naive baseline — turned out to have 90% ranges containing the true value only 62% and 58% of the time across the series as a whole, and 43% and 49% of the time at the consumer peaks themselves. In Kaas et al. (2026)'s results, a model that understates its uncertainty raises fewer false alarms. That model scores well on a threshold-crossing test while being exactly the model an operator should not trust near a capacity limit.

A cross-validation fold shorter than a year cannot show whether a model handles both ends of the year, which is one length rule worth adopting outright. Pinheiro et al. (2023) held out the whole of 2019 and note that "one year is the minimum acceptable to test a forecasting model whose target value shows annual seasonality". Substation load shows exactly that seasonality, so any cross-validation fold — one train-then-test slice of the history — shorter than a year cannot tell us whether a model handles both ends of the year. NGED needs both: winter is when NGED buys flexibility, and summer, when embedded solar output is highest against the lowest demand, is when export constraints bind and generators are curtailed.

Every forecasting paper we read that describes its split keeps most training data out of the future of its test data, and the training window usually grows rather than slides. Flexpectation's own protocol matches this literature where the literature has settled: an expanding training window with the validation window lying strictly after it, which is what most of the papers above do — Pinheiro et al. (2023) slide a fixed 3-year window rather than expanding it, and Gilbert et al. (2023) interleave their test blocks with training data from later in the same year. Our validation window is a complete year, which meets Pinheiro et al.'s minimum above, and power lag features shorter than the lead time are nullified, so a forecast can never see the load it is predicting.

None of the papers we read addresses the leakage a frequently reissued forecast creates, and Flexpectation reissues its forecast often enough for the leakage to matter. When a forecast covering 14 days is reissued every 6 hours, every target half-hour is covered by 56 separate forecasts. For example, with runs at midnight, 06:00, noon, and 18:00 each day, the half-hour from 10:30 to 11:00 on 15 September is covered by every run from noon on 1 September to 06:00 on 15 September. The literature describes two traps. If we were to count the 56 forecasts as independent, a significance test would report a confidence the data does not support. If we were to let a target half-hour fall on both sides of a train-test boundary, the test set would be contaminated outright.

Only one paper in this review even partly addresses the leakage, and a second avoids the leakage only by accident. We searched every paper in this review that reissues a forecast more often than its horizon is long — looking for a gap or buffer between training and test, a block bootstrap, or a correction to the number of independent observations — and found one partial treatment. Hertel et al. (2026) compare models with Diebold-Mariano tests implemented after the R forecast package, whose variance estimator corrects for serial correlation in the loss differential when the estimator is told the forecast horizon. The paper does not say which horizon Hertel et al. pass, so we cannot tell whether the correction was applied. The one paper the problem cannot reach is Kaas et al. (2026), whose stride equals their horizon: their forecasts run 4 days each, so no two share a target. That is our inference from their design rather than a claim they make, and they give a different reason for wanting a shorter stride, that it would "provide more insights", while describing exactly what a shorter stride would create — "each data point in the dataset covered by multiple forecasts, as opposed to a single forecast per data point in the used configuration". Flexpectation will treat the leakage as an open methodological question rather than a solved question, and will report what it did about the leakage rather than leave it implicit.

We have no ground truth for the unmetered solar and wind behind a substation's meter. The disaggregation papers we read use three substitutes for truth, each of which fails differently. Three further substitutes appear in none of those papers. The three in use are to hold out a generator that is metered, treating the generator as if it were unmetered, recovering the generator's output from a real substation's net flow and comparing against the meter; to sum individually metered generators into a synthetic substation and score the disaggregation against the components that went into the sum; and to compare against an independent tool rather than against truth. The three that appear in none of them are to measure whether the estimate improves the forecast it was built to improve, which the capacity-estimation literature does use; to check an estimate against a physical model; and to use a substation where every feeder and every embedded generator is metered, purely as validation. The sixth substitute stays out of reach, because a fully metered substation is a field deployment rather than an analysis.

The physical checks are concrete and require little effort: disaggregated components must sum to the measured net flow; disaggregated solar must be zero at night and must sit under the clear-sky envelope; disaggregated wind must track wind speed rather than irradiance; and an inferred rooftop-solar capacity must be plausible for the area a substation serves. None of those checks needs a label, and a violation is a detectable error whatever the truth turns out to be. Using physical consistency to score an estimate, rather than to shape it, is close to absent from the papers we read. Physical-consistency scoring is the least effort of any evaluation on the list.

No one substitute for ground truth is trustworthy alone, so the best approach may be to run multiple proxy tests and report where they disagree. Each test fails in a different way, so an estimate that survives multiple tests is better supported than an estimate from the single best substitute.

The five are not five attempts at the same measurement. The hold-out is biased towards the sites that happen to be metered. Synthetic aggregation systematically flatters, because a clean sum of metered sources has no switching events, no false zeros, and no unmetered load. A score from synthetic aggregation should therefore be reported as performance under idealised aggregation rather than as real-world skill. The remaining three each answer a narrower question than they appear to: the independent-tool comparison says only whether we agree with an existing method; the physics checks find wrongness but never confirm rightness; and the downstream test measures whether the estimate is useful, which is not the same as whether the estimate is right, because an estimate that is wrong in a way the forecast does not care about will score well. Every number we publish will name the substitute behind that number.

The effective capacity of a metered generator has no ground truth either, and most of the six substitutes above cannot be applied to a single generator's meter. A generator's own meter is the input a capacity estimator works from rather than a label to score the estimator against, so there is nothing to hold out and nothing to aggregate synthetically. Of the physical checks above, only the night-zero and clear-sky bounds carry over to a single generator's output, because the rest test a substation's net flow. The same multiple-test logic therefore produces a different basket for challenge 3: a head-to-head contest between candidate estimators in which downstream forecast skill decides, supported by synthetic fault injection — scaling a known period of a healthy generator's output down by a known factor, then asking whether the estimator finds the change, at the right date and the right size — together with the precision and recall of the fitted change dates against NGED's maintenance records, robustness to unlogged curtailment, and the calibration of each estimator's stated uncertainty.

Scoring a rare-event detector is a class-imbalance problem before it is anything else, and the best-worked example in this review chose its answer deliberately. Score by timestamp and the long events decide the number; score by event and short events do. Bouman et al. (2024) split events into four duration bands, score each band separately and average, and exclude from the scoring entirely the timestamps a labeller marked uncertain. Bouman et al. set the detection threshold by maximising that averaged score rather than by the conventional two- or three-standard-deviation control limit, and resample the test stations 10,000 times to put an uncertainty on the result. Their own verdict on how well the detection worked is that performance "is relatively low across the board, even on the train data. This indicates that the problem is hard to learn, though it generalizes fairly well".

Two further choices in this literature are worth copying, because both make a flag defensible to the engineer whose substation it lands on. Perry and Muller (2022), detecting step changes across 101 manually labelled photovoltaic power and irradiance streams, score a detection as correct if it lands "within 30 days of their labelled occurrence" — the right shape for a problem where the exact timestamp of a gradual shift is not knowable, and a tolerance that has to be stated, because the score means nothing without it. Martín et al. (2018) set their detection threshold from instrument physics rather than from the data: transformers contribute up to ±1% error and the measurement equipment ±0.5% to ±1%, so ±2% is the inherent floor. Martín et al. set the threshold at ±4% "to avoid detection of false gain and offset errors".

None of the three GB projects offers a number to compare a detector against, which we checked rather than assumed. Across Electricity North West's ATLAS — both its 2016 methodology and its 2018 close-down report — UK Power Networks' Distribution Network Visibility, and NGED's own Time Series Data Quality, the words precision, recall, F-score, true positive, and false positive do not appear at all. That absence is a gap in the published GB record rather than a target Flexpectation is setting itself.

Leaderboards of machine learning results

Flexpectation is building a leaderboard, not a public ML competition, and the distinction changes which published lessons apply. Our leaderboards carry our own experiments, one per class of time series. For example, solar farms, wind farms, batteries, and the demand at primary substations each get their own leaderboard. Results for individual generators are published without the generator's name or ID. The leaderboards are public to view and reproducible, but we are not inviting other teams to submit entries. Anyone who wants to benchmark against us can rerun the setup for themselves. Not inviting outside entries means the literature's lessons about attracting entrants, prize pots, and qualifying rounds do not apply to us. The lessons about protocol — what makes a comparison trustworthy — apply with more force, because rival entrants give a competition some of its integrity by wanting to catch each other out.

Energy forecasting has run competitions on common data for over a decade, and only the second track of the Global Energy Forecasting Competition 2017 (GEFCom2017) forecast at anything like the level NGED acts on. The last row of the table below describes what Flexpectation is building. The two columns that decide whether a precedent exists are the aggregation level and whether the leaderboard is still open.

Leaderboard What entrants forecast Aggregation level, set against a primary substation Take-up Standing or closed
Global Energy Forecasting Competitions 2012, 2014, and 2017 (Hong et al. (2020)) Hierarchical load, price, wind, and solar, with the data published alongside the papers introducing each competition Varies, up to national Hundreds of contestants from more than 60 countries Closed
The second track of GEFCom2017 (Hyndman (2020)) Probabilistic load 183 delivery-point meters of a US utility — the closest of the leaderboards in this table to a distribution network's aggregation level 177 entrants across both tracks Closed
BigDEAL Challenge 2022 (Shukla and Hong (2024)) The timing of peak demand rather than its size; the final match asked for the magnitude, timing, and shape of daily peak load Three neighbouring local distribution companies — whole utilities, well above a primary substation 78 teams from 27 countries Closed
HEFTCom (Browell et al. (2026)) The combined day-ahead output of one GB wind-and-solar portfolio A single 3.6 GW portfolio: the 1.2 GW Hornsea 1 offshore wind farm plus a regional solar aggregate — the generation mix closest to NGED's, though at portfolio rather than substation level Over 170 teams registered, 66 submitted, 24 completed Closed; the competition period was 3 months
Three competitions NGED funded with Energy Systems Catapult (McSweeney et al. (2023)) 1-minute peaks inside half-hourly averages; the daily peak a hidden population of electric-vehicle chargers added; and missing values. None was a load forecast NGED's own grid supply point, bulk supply points, and primary-substation feeders 37 teams, over 2,500 submissions Closed, though the pages and data are still readable on CodaLab
WindAI (Authen et al. (2026)) Hourly wind power for the whole of a target day two days ahead, submitted daily against an outturn that had not yet happened Four bidding zones of a transmission network — regions far above a primary substation 9 teams carry an average score in the competition summary Closed; the live evaluation ran over 10 working days in autumn 2025
Energy-Arena (Kleinebrahm et al. (2026)) The paper describes deterministic day-ahead tasks; the running platform carried 24 challenges across prices, load, wind, and solar in September 2026 — 8 scored as point forecasts, 8 as quantiles, and 8 as ensembles Not a distribution network Not stated in what we read Standing
TS-Arena (Meyer et al. (2026)) 186 live energy series Not a distribution network 13 foundation models and 3 statistical baselines run by the platform team, plus outside entries Standing
Predico (Elia Group) Quarter-hourly probabilistic generation: Belgian solar out to 10 days ahead, and the German wind and solar markets that 50Hertz runs day-ahead National generation totals of two transmission networks Forecasters join by application; the number taking part is not published Standing
Flexpectation's leaderboards Net demand at substations, and output at metered generators One board per class of time series Public to view and reproducible; outside entries not invited Standing

No leaderboard in the table combines the three properties of Flexpectation's leaderboards: they keep running rather than closing after a fixed period, they forecast at substation level, and they score methods on NGED's own data. We found no example of a standing leaderboard for substation forecasting — a leaderboard that keeps accepting entries after its competition closes. Two of the three competitions NGED funded sat at exactly the levels NGED forecasts. The gap is therefore scoped to forecasting rather than to the voltage level. McSweeney et al. (2023) draw the same conclusion this review does, writing that "many solutions are only tested on private data using a single method only compared (if at all) to simple, non-competitive benchmarks", which "limits the reproducibility and usefulness of the outputs", and pairing their own results with the caveat that those results came "despite the necessary reduction in realism" of a curated competition dataset. What they recommend keeping open is the unranked practice phase, "as it allows teams to continue experimenting within the platform". Flexpectation's leaderboards are meant to fill that gap, though we would be glad to be pointed at a counter-example.

WindAI is the closest of these competitions to challenge 3's problem of a generator whose capacity keeps changing, because robustness to that change was a scored criterion rather than an afterthought. Statnett, Norway's transmission system operator, asked entrants for the hourly wind power of each of four Norwegian bidding zones two days ahead. Authen et al. (2026) report a weighted assessment giving 65% to accuracy, 20% to trustworthiness and explainability, 10% to implementation and presentation, and 5% to "robustness to changes in installed wind power capacity, evolving weather patterns, long-term climate variability". What the entrants did with that 5% of the assessment, and why a GB distribution network operator cannot copy them, is set out under challenge 3 above. Two further results transfer. The top three entries all used gradient-boosted decision trees. Authen et al. conclude that the more complex deep-learning architectures' "additional complexity did not translate into superior performance". And the placings did not follow the accuracy order: WindSight recorded a lower average root mean square error than Knowit, 216.22 MW against 217.57 MW, and Knowit still took second place. The other 35% of the assessment is what produced that reordering.

Predico is a standing leaderboard that pays its entrants, which is one mechanism Flexpectation's leaderboards deliberately do without. Elia Group describes Predico as "a collaborative forecasting market platform enabling entities with common interests to procure and sell forecasts", where buyers receive "skill-weighted aggregate market forecasts" and forecasters are remunerated on performance: the Belgian solar market carries €7,000 a month, split by accuracy rank and information contribution, with the best forecasters earning about €1,500 to €2,000 a month. Forecasts are quarter-hourly and probabilistic, given as the 10th, 50th, and 90th percentiles, out to 10 days ahead for Belgian solar generation and day-ahead for the German wind and solar markets 50Hertz runs. The platform documentation scores the median submission by root mean square error and the 10th-to-90th percentile pair by the mean Winkler interval, and ranks forecasters monthly. Predico forecasts the national generation totals of two transmission networks rather than any asset on a distribution network. Elia Group describes the platform as a proof of concept in which participants cannot yet create their own markets. So what Predico offers Flexpectation is a worked example of a standing, publicly ranked board rather than a precedent at NGED's aggregation level.

One mechanism that makes a leaderboard trustworthy is time rather than policing. The central idea of Meyer et al. (2026)'s TS-Arena is that a forecast is submitted before the outturn it will be scored against physically exists, which "makes test-set contamination impossible by design". HEFTCom made the same argument from experience: because the competition ran on the real, unknown future, "data leakage, accidental or deliberate, was impossible". A half-hourly forecasting service meets that condition easily: every day supplies 48 fresh evaluation points that can never be reused. The condition that the answer did not exist when the model was frozen holds automatically. The corollary is uncomfortable for anyone relying on a fixed hold-out set. TS-Arena states the corollary plainly: "leveraging any fixed dataset that is not evolving over time and directed into the future — regardless of how carefully curated — can eventually lead to information leakage". Hong et al. (2020) name the same failure from the other end, that "some datasets have been studied so well that the researchers may use some of the future information to give unfair advantage of their proposed methods".

TS-Arena's leaderboard is populated almost entirely by models its own operators run, which is near enough to Flexpectation's position that its self-imposed rules transfer. TS-Arena does invite outside entries, where Flexpectation's leaderboards do not, but TS-Arena's reference models "act as neutral participants, autonomously requesting context from the API Portal and submitting forecasts to it", so that those models "operate under the exact same constraints (e.g., submission windows, data access) as other (external) participants". Each foundation model is run from its authors' own repository at its authors' recommended defaults, with no domain-specific tuning. All three of those rules are available to a single team, and Flexpectation intends to adopt them: our own models go through the same evaluation interface as any baseline, and a baseline is run as its authors published it.

One way a single team fools itself is not fabrication but running the baseline badly. Kleinebrahm et al. (2026) describe running the baseline badly as a general problem with published comparisons: competing methods "are not always implemented or optimized with equal care", so reported differences "may reflect differences in implementation quality rather than inherent methodological advantages". Hong et al. (2020) put the same problem more bluntly, that "sometimes the parameters are manipulated, so that the competing models are being dominated by the proposed ones", alongside two related habits — picking the error measure that favours the proposed method, and skipping comparison with naive models altogether. A team that runs every entry on its own leaderboard is exposed to all three by construction. That is why the authors'-code-and-authors'-defaults rule above matters more for Flexpectation than it does for a competition. That exposure is also a large part of the reason why, in Flexpectation version 1, we are putting effort into optimising our XGBoost forecasts before trying more novel approaches.

The submission deadline, not a rule about which features are allowed, is what defines a fair information set. Kleinebrahm et al. (2026) give a worked example of the trap: several published papers use the day-ahead wind and solar forecasts that the European Network of Transmission System Operators for Electricity publishes as inputs to day-ahead price models. But those forecasts are "released only after 18:00 on the day before delivery, whereas the day-ahead market already closes at 12:00 on that day", so the feature did not exist when the forecast had to be made. Kleinebrahm et al.'s fix is structural rather than procedural, in that each competition "implicitly defines an operational information set through the submission deadline". Flexpectation has the same hazard in the delay between an ECMWF run and its arrival. The same fix is available: score against the data that had actually landed at the forecast's issue time.

Carry two baselines, a lower baseline below the achievable skill and an upper baseline at the achievable skill, rather than a single baseline. Doubleday et al. (2020) distinguish the two jobs a benchmark does: a yardstick, which need not be a good forecast, and what they call a point on the yardstick — a target for a new method to beat, which "should be close to the state of the art". Doubleday et al. recommend carrying both kinds of baseline, so that a new method can be positioned between the two baselines rather than declared better than a single baseline. Flexpectation's leaderboards carry both kinds of baseline: persistence (tomorrow resembles a comparable recent day) and climatology (tomorrow resembles the historical average for the time of year) as the naive yardstick, and the manual heuristic as the point on the yardstick a new model has to reach.

Flexpectation's leaderboard today reuses one fold for both model selection and the published result, so the winner's reported skill is optimistically biased. The fold that Flexpectation currently reports serves as both the model-selection set and the reported result. Every hyperparameter choice and feature ablation is therefore adjudicated on the same 12 months the leaderboard publishes. With hundreds of experiments planned, that bias will grow. A leaderboard wears out through repeated use. Hyndman (2020), who has co-organised a forecasting competition, expects that wear: "over-study of a single benchmark data set means that methods will eventually over-fit the published test data. I suspect this has happened with the M3 data over the past 20 years, and it is likely to happen with the M4 data, despite its much larger size. Therefore, a wider range of benchmarks is desirable, and these need to be updated regularly. Consequently, there can never be a 'final forecasting competition'." Our own fold is small in effective sample size rather than in row count, because consecutive half-hours are strongly correlated. Strong correlation between consecutive half-hours shrinks the evidence a fold carries just as a small row count would. The structural fix is a final-test window that no model selection is allowed to touch, and that final-test window is scheduled. Until the final-test window lands, three limits hold: leaderboard numbers are selection metrics rather than estimates of future skill, differences smaller than fold-level noise should not drive decisions, and the number of experiments run against a fold is itself a statistic worth publishing beside the fold's results.

Rankings travel better than absolute numbers do. Where a benchmark has enough data behind it, the ordering of models survives a change of test set even when the accuracy level does not. The survival of the ordering decides what a leaderboard should report as its headline. Recht et al. (2019) found the ordering of models preserved on a freshly collected test set while the accuracy level moved by "approximately five years of progress in a highly active period of machine learning research". Fildes (2020), reviewing the M4 competition, compared its daily micro series against a real retail forecasting problem and found the same method scoring 1.665% on the M4 daily micro series and 11.1% on the retail problem. Fildes's conclusion points the same way Flexpectation is going: "each organization needs to organize its own forecasting competition for its own forecasting problems, and should not rely on even large benchmark data sets", with the published competition useful for narrowing "the pool of methods to be considered" rather than for predicting your own error. So a leaderboard should lead with ranks and with margins over a stated baseline, and treat an absolute skill number as valid only on the distribution it was measured on.

A finite evaluation window can rank the wrong model first, and several months is not obviously enough. Messner et al. (2020) produce a wrong ranking rather than merely asserting that a wrong ranking is possible: they fit three forecasting models with three different loss functions, so that each model is optimal for one metric by construction, then score all three on the first 200 time steps. The model built for the quadratic loss wins all three metrics, while the two models built to win on mean absolute error and on the quantile score both lose on their own metric. Their conclusion is the sharpest warning we found about reading a leaderboard: "evaluation results based on a finite data set are always subject to some degree of uncertainty and the best ranked forecast does not necessarily have to be the truly best one. Depending on the actual setup, e.g., in a benchmarking exercise to hire a forecaster, it should be remembered that even periods of several months may still yield uncertainty in terms of who the best forecaster truly is." HEFTCom's own competition period was 3 months. The practical response, which TS-Arena adopts, is to publish an interval on the ranking rather than the ranking alone, so that a new entry near the top is visibly provisional. TS-Arena's interval comes from replaying the round order in random permutations, so it widens for models with few rounds rather than measuring sampling error over a finite window. Meyer et al. (2026) warn against "treating short-term success as proven superiority", noting that the confidence intervals of their own top models overlap.

A leaderboard with no outside entrants cannot support two kinds of claim the competitions above can, and Flexpectation should not make them. The Critical Assessment of Structure Prediction (CASP) competition's 14-year plateau in one of its scored categories (Kryshtafovych et al. (2021)) is evidence about the difficulty of protein structure prediction only because dozens of groups were attacking the problem independently. A plateau on our leaderboard would be ambiguous between a hard problem and a team that did not think of the right idea.

CASP is a recurring competition rather than a standing benchmark, and Flexpectation is proposing the division of labour that produced AlphaFold2. Stating CASP precisely matters, because the precise version supports Flexpectation's design better than the loose version. CASP gathers every two years proteins "for which the experimental structure is about to be solved or is solved but still not public" and gives the sequences to entrants, so each target is single-use once the structure is published. The standing benchmark in that field is CAMEO (Continuous Automated Model EvaluatiOn), which Robin et al. (2021) describe as complementing CASP by running "fully automated blind evaluations" against the weekly pre-release of forthcoming structures. AlphaFold2 was developed against neither CASP nor CAMEO, but against a temporal hold-out of its own: Jumper et al. (2021) score it on structures "deposited in the PDB after our training data cut-off", which is what a live forecasting service gets for free. The blind competition was the audit; the temporal hold-out was a check the team could run for itself, on data no rival had to supply. Splitting the audit from the self-run check is the same division of labour Flexpectation is proposing, and that split is the reason a leaderboard without entrants is a coherent design.

The same limit applies to a claim about a whole class of forecasting method. The first M-competition's conclusions about whole classes of method — that statistically sophisticated methods do not typically forecast more accurately than simpler methods, which the M3 competition did not go on to support, and that a combination of several methods forecasts more accurately, on average, than the individual methods going into the combination (Hyndman (2020)) — describe what many independent people chose to try. No single team's leaderboard can support a conclusion about a whole class of method.

What our leaderboard can do is narrower and still worth having: show which approaches beat a stated baseline on NGED's own data, under one protocol, with the metric definitions, the code, and the forecasts published — anonymised for individual generators — so that anyone can check the arithmetic or rerun the comparison themselves.

A single-team leaderboard cannot take its credibility from rivals, so it has to earn credibility by declaring its own gaps, and Flexpectation's gap is known in advance: the leaderboard runs on the 32 trial-area series while the service is meant to reach the whole of NGED's distribution network, so we should expect our published numbers to flatter what happens at scale, and should say so each time we publish them. A benchmark of 32 series is also small enough that the constraint on what can be learned from it is likely to be its size. That is an argument for extending the leaderboard to the wider distribution network as soon as the data allows rather than for running more experiments against the trial area.

Publishing results that others can compare against

Energy forecasting's own senior figures say that published results in the field cannot be compared with each other, which is the problem this review ran into at every one of the nine challenges. Hong et al. (2020), a review written by six widely cited authors in the field, concludes that "most papers can never be replicated, because the data have never been published". Flexpectation publishes the evaluation protocol, the metric definitions, and the code that computes the metrics, so that someone outside the project can check how the results were produced rather than take the results on trust. The telemetry itself is shared only where NGED's data policy allows. A metered generator's time series is never published with the generator's name or ID, because a single site's output can be commercially sensitive.

Flexpectation commits to nine practices, from correcting for ensemble size to publishing negative results, that let an outsider check its published numbers rather than take them on trust. Two of the nine are already argued above and appear here only as pointers, so that the list can be read as a whole.

  • Every ratio comes with its reference forecast, the population it was scored on, and the number of ensemble members that produced it. Weigel et al. (2007) show that a ranked probability skill score is biased downwards by an amount that depends on ensemble size. A score from our 51 ensemble members is therefore not comparable with a score from a study using 10 ensemble members until Weigel et al.'s correction is applied. We apply that correction.
  • Accuracy is reported separately for each class of asset — grid supply points, bulk supply points, primary substations, and metered generators — each against its own stated naive baseline, because a single project-wide accuracy target would set a different level of difficulty for each class of asset.
  • The fraction of series that beat their naive baseline is published alongside the average error, never the average alone. An average error across a population can improve while the model gets worse at a substantial minority of series. That minority is what an operator notices.
  • The battery, the gas generator, and the biofuel plant are reported separately from the wind and solar sites, because those three assets are dispatched on market signals that no weather forecast contains.
  • A peak-aware score is reported alongside a proper scoring rule, never instead of a proper scoring rule, for the reason set out under "Evaluating the performance of power forecasts" above.
  • The tail is scored with a threshold-weighted continuous ranked probability score, weighted above a fixed per-series threshold set at the 99th percentile of that series' own measured history, rather than by selecting the periods in which an exceedance happened. The obvious alternative — keep only the periods in which net demand crossed the limit, and score those — is not merely noisy but biased: Lerch et al. (2017) show that choosing which periods to score on the basis of what happened rewards a forecaster who over-predicts extremes, and can rank a deliberately biased forecast above an honest forecast. Gneiting and Ranjan (2011)'s threshold-weighted score puts the emphasis inside the score instead, and stays a proper scoring rule while doing it. A GB distribution network has already been scored this way: Maia et al. (2026) compare fault-count forecasts for SP Energy Networks against a quantile-regression baseline on the threshold-weighted score, because an unweighted score "would place substantial emphasis on parts of the predictive distribution where the two models are identical".
  • Coverage — how often reality fell inside the range the forecast claimed — is broken down by season, by forecast lead time, and by how heavily loaded the substation was. A coverage figure averaged over a year can read as a healthy 90% while being 99% in the quiet months and 70% at the winter peaks. The winter peaks are the periods when NGED buys most flexibility. Conformal prediction does not remove the need for the breakdown: Foygel Barber et al. (2020) prove that a distribution-free guarantee holds only on average across all conditions, never separately for the conditions that matter. A conformal forecast can therefore promise 90% coverage overall while failing at the peaks.
  • Each metered generator's series is normalised by its estimated effective capacity before training — unless the comparison described under challenge 3 above shows the normalisation is not needed — and that estimate is tracked as it changes.
  • Negative results are published too, including whether an off-the-shelf model given none of our data matches our own, and whether sustained experimentation stops yielding improvements.

What the literature says about machine-learning operations (MLOps)

Machine-learning operations — building, testing, deploying, and monitoring machine learning as production software — is a core aim of Flexpectation, so what the literature does and does not establish about the practice matters to this project.

The field is a large body of description and almost no measurement. The field has a settled definition, a vocabulary for the failure modes the practice exists to prevent, surveys of the available tools, and maturity models — but across the six reviews of machine-learning operations read for this section, nobody has published a metric showing what adopting the practice delivers. Energy forecasting has no separate body of findings to fall back on: what exists is a handful of platform descriptions, no mature energy-specific platform among the platforms Zhao et al. screened, and no paper we read giving a retraining cadence a network operator could act on. Yet the combined solar-and-wind forecast error at Europe's transmission operators has been measured roughly doubling over 5 years.

That absence sets the terms for what Flexpectation can claim for its own experiment framework. The practice this project is betting on is fast, comparable iteration. The case for fast, comparable iteration rests on a structural argument about how fields make progress, together with testimony from senior practitioners, rather than on a controlled measurement. And the documentation needed to run a controlled measurement is itself largely missing from published work. The better-documented precedent lies outside machine learning: operational meteorology has tied production-model changes to measured changes in forecast skill for decades. One finding cuts against this project directly, and the section below sets the finding out rather than quoting around the finding: the same structural argument predicts that fields which cannot share their data will fall behind in their rate of progress, and the substation telemetry this project uses is not published.

Four papers to start learning about machine-learning operations. Start with Kreuzberger et al. (2023), which defines the term "MLOps", derives nine principles, and draws the architecture and the roles that go with those principles. Read Eken et al. (2025) next, the broadest synthesis among the reviews we read, which reads grey literature alongside journals. That breadth lets Eken et al. capture the practice practitioners write down outside the academic record. Then read Zhao et al. (2026), the only paper of the four written for energy forecasting, which maps platform capabilities against an energy-forecasting lifecycle rather than a generic lifecycle. Sculley et al. (2015) is worth adding for the terms alone. Much of the field argues in the vocabulary Sculley et al. established.

MLOps research describes good practice but does not measure what the practice improves

Kreuzberger et al. (2023) give the definition most of the field now uses. The definition rests on a structured review that narrowed 1,864 retrieved articles to the 194 read in detail and then to 27 peer-reviewed articles, a review of the available tools, and eight interviews with practitioners. Kreuzberger et al. define machine-learning operations as "a paradigm, including aspects like best practices, sets of concepts, as well as a development culture when it comes to the end-to-end conceptualization, implementation, monitoring, deployment, and scalability of machine learning products", drawing on machine learning, software engineering, and data engineering together. From that evidence, Kreuzberger et al. derive nine principles. What Kreuzberger et al. do not do, and do not claim to do, is measure what adopting the nine principles changes.

The failure modes the practice exists to prevent were named from experience rather than from measurement, and the naming is the contribution. Sculley et al. (2015) are explicit about the standing of their own paper, which "does not offer novel ML algorithms, but instead seeks to increase the community's awareness of the difficult tradeoffs that must be considered in practice over the long term". Sculley et al.'s paper also rests on what the acknowledgements call "accumulated folk wisdom" from running machine learning at Google. The paper reports no experiment and no number.

What the paper contributes is a vocabulary much of the field now uses. That vocabulary covers entanglement, where mixing signals together makes any one improvement impossible to isolate, along with correction cascades, where a model learned on top of another model's output makes the model underneath hard to improve, undeclared consumers, unstable and underutilised data dependencies, direct and hidden feedback loops, glue code, pipeline jungles, dead experimental codepaths, configuration debt, reproducibility debt, and process management debt. The vocabulary also covers the principle Sculley et al. abbreviate to CACE, "Changing Anything Changes Everything": no input to a model is ever really independent of the others, so changing one feature shifts the weight the model puts on the rest, and the same principle holds for a hyperparameter or a sampling method. One widely used term is not Sculley et al.'s: "training-serving skew" is later vocabulary. The words "skew" and "serving" appear nowhere in the paper.

Four further reviews agree that the field is largely conceptual. Only one review we found went looking for a measure of effectiveness, and that review reported finding none. Woźniak et al. (2025) screened 2,615 records returned by their database searches down to the 135 publications that passed a title-and-abstract screen and then to the 41 publications kept after a full-text read. Woźniak et al. asked as one of their four research questions what metrics measure the effectiveness of a machine-learning-operations implementation. Woźniak et al. answer that "None of the reviewed articles presented metrics that could measure the effectiveness of MLOps implementation in an organization", an absence Woźniak et al. call "unexpected" and attribute to the immaturity of the area. Eken et al. (2025) cast the widest net of the six reviews read for this section, a multivocal review analysing "a corpus of 150 peer-reviewed and 48 grey literature" precisely because so much of what the field knows sits outside the journals. Eken et al. reach the same place from the other direction: an "impact analysis framework needs to be created" so that practitioners can "assess benefits and drawbacks using quantifiable metrics", which Eken et al. list as future work rather than as a metric the literature already offers.

The other two reviews reach the same verdict from different corners of the field. Lima et al. (2022) screened 1,905 articles down to the 30 articles they kept, and concluded that machine-learning operations "is still in its initial stage". Rajashekarappa et al. (2026), narrowing 186 database records down to 12 studies of manufacturing specifically, report that "fully automated MLOps frameworks remain underdeveloped".

The largest empirical study we found does not break the pattern. John et al. (2025) interviewed practitioners at 14 companies and built a framework, a maturity model, and a taxonomy from what those practitioners described. John et al. measured no outcome. The benefits John et al.'s paper lists are benefits the interviewees and the prior literature claim, not effects John et al. measured.

MLOps in energy forecasting

What exists for energy forecasting specifically is a handful of platform descriptions rather than a body of findings that agree or disagree with each other. Zhao et al. (2026) screened 256 candidate documents — vendor documentation, open-source repositories, and academic papers — down to the 31 they kept. Zhao et al. mapped the 13 general-purpose machine-learning-operations platforms those 31 documents describe against an energy-forecasting lifecycle, scoring each platform capability as native, partial, or not clear from the platforms' own documentation. Zhao et al.'s first finding is the shape of the field rather than a ranking: "No energy-specific mature MLOps platforms were identified within the screened sources". As a result, energy forecasting adapts general-purpose platforms to the domain. Zhao et al. are explicit that their mapping "does not perform hands-on deployments, runtime benchmarking, cost comparisons, or empirical evaluation of forecasting accuracy". Zhao et al. close by naming the study that does not yet exist: "A natural next step is a hands-on empirical benchmark that evaluates the actual implementation complexity and operational performance of platforms."

The individual platform descriptions supply worked examples and no comparison between platforms. Subramanya et al. (2022) build and run a pipeline for day-ahead price forecasting in the Finnish reserve market. But Subramanya et al. report no accuracy figure and no measurement of the engineering effort the pipeline saved. Pelekis et al. (2024) go further than Subramanya et al. towards a worked example with DeepTSF, an open-source platform that orchestrates its pipeline with Dagster and tracks experiments with MLflow. Pelekis et al. tune a deep-learning model, neural basis expansion analysis (N-BEATS), over 100 hyperparameter trials on a day-ahead forecast of Italy's national electricity load, then backtest the winner on a held-out year. What no platform description in this section supplies is a comparison between platforms: DeepTSF is measured against no baseline platform and no second orchestrator. Pelekis et al. report that deployments in the I-NERGY project have "already proven DeepTSF's efficacy in DL-based load forecasting" without attaching a number to that claim.

The one paper we found that argues for machine-learning operations from inside power-systems forecasting makes a different point altogether. Gürses-Tran and Monti (2022) find that forecast developers "predominantly assess residuals and error statistics when tuning the targeted model's quality", so that "eventual cost or rewards of the underlying business application are typically not considered in the model development phase".

Forecast error at Europe's transmission operators grew measurably over 5 years, yet no paper we read gives a retraining cadence for an energy forecast in production. Kazmi and Tao (2022) analysed 5 years of day-ahead forecasts published by 16 European transmission system operators and found that "the combined forecast error due to solar and wind has roughly doubled during just the last five years", with the errors "highly autocorrelated". That autocorrelation means structure remains that a better model could exploit. Heidrich et al. (2022) tackle the resulting problem by cutting the effort retraining takes, observing that "Most methods for coping with such concept drifts rely on computationally expensive retraining", and updating a lightweight profile instead of retraining the whole model.

What none of these papers supplies is a number a network operator could act on. The retraining triggers the papers state are qualitative — Subramanya et al. update their pipelines "if the performance has gone down", and Gürses-Tran and Monti say of their own ProLoaF model that training "is performed once and does not require re-training, as long as the used training dataset is still representative of the system under study". So how often a substation forecast must be retrained is a question Flexpectation will have to answer from its own data.

The literature settles which orchestrator an energy-forecasting platform should run on no better than it settles the retraining cadence. The platforms Zhao et al. map are general-purpose machine-learning platforms rather than orchestrators — Kubeflow, ZenML, ClearML, Polyaxon, Metaflow, Domino, Databricks, SageMaker, and Vertex AI among the 13. Airflow enters that mapping only as a tool Metaflow integrates with. Dagster is named in none of the six machine-learning-operations reviews cited above (Kreuzberger et al., Lima et al., Eken et al., Woźniak et al., Rajashekarappa et al., and Zhao et al.). DeepTSF, the one energy-forecasting platform we found that is built on Dagster, benchmarks no orchestrator at all. The published evidence therefore shows Dagster to be a workable foundation for an energy-forecasting pipeline, and says nothing about whether Dagster is the better of the two tools Flexpectation weighed. The reasoning behind that choice is set out in Why Dagster, not Airflow.

Operational meteorology has tied production changes to measured skill for decades

Operational meteorology has been running continuous verification of production forecasts for decades, and has documented that practice far more thoroughly than the machine-learning-operations literature has documented its own. Brown et al. (2021) describe the Model Evaluation Tools, verification software built since 2007 and used operationally by the United States National Weather Service and others, noting that "Forecast verification/evaluation has been a subject of research and also applied to operational forecasts for more than a century". Brown et al. report a user community of "more than 3,700 researchers and operational users from 124 countries". Hoffman et al. (2018) show what continuous verification delivers: tracking the skill of three operational forecasting centres continuously, Hoffman et al. attribute a "7.37% increase in the probability of improved skill" to a single, named model upgrade made in 2016. Tying a specific production change to a measured change in skill is what the machine-learning-operations literature we read does not do. Meteorology has been tying production changes to measured skill routinely.

Fast, comparable iteration is argued for, not measured

Fast, comparable iteration is the practice within machine-learning operations that Flexpectation is betting on, and the case for that practice rests on a structural argument and on practitioner testimony rather than on a controlled measurement. Donoho (2024) makes the structural argument. Donoho identifies three practices, labelled the frictionless-reproducibility triad — data sharing, the ability to re-execute another researcher's workflow exactly, and challenge problems with "a shared public dataset, a prescribed and quantified task performance metric, a set of enrolled competitors seeking to outperform each other on the task, and a public leaderboard". Donoho argues that fields adopting all three "commonly benefit from very high velocity of progress", because frictionless reproducibility "spontaneously spawns groups of inspired researchers to a tight loop of iterative experimental modification and improvement". Donoho hedges the claim in the same sentence: "Of course, not every field works this way." Donoho offers historical case narrative rather than a measurement. Donoho's paper is therefore the strongest argument for the mechanism among the sources this review found, and is not evidence of an effect size.

The practitioner testimony reaches the same conclusion as Donoho's structural argument, and is careful to say that speed comes from the protocol rather than from haste. John Jumper, whose estimate that around 90% of research ideas fail opens this review, elsewhere credits a prototype that "would give wrong answers at incredible speed", which "made it easy to start becoming very adventurous with the ideas you try" (MIT Technology Review (2025), 24 November 2025). Ng (2018) writes that researchers "will usually try out many dozens of ideas before they discover something satisfactory". Ng also writes that a development set with "a single-number evaluation metric helps you quickly evaluate algorithms, and therefore iterate faster". Godbole et al. (2023) recommend "running a larger number of shorter experiments and reserving the longest 'production length' runs for the models we hope to launch". Andrej Karpathy's autoresearch fixes each training run at 5 minutes so that "you can expect approx 12 experiments/hour and approx 100 experiments while you sleep". Karpathy states the reason for the fixed budget plainly: the fixed budget "makes experiments directly comparable regardless of what the agent changes".

Every one of those accounts describes fast iteration under a fixed and comparable protocol, which is a different claim from going fast. Karpathy warns against reading the case for fast iteration as licence to hurry, writing that "a 'fast and furious' approach to training neural networks does not work and only leads to suffering" (Karpathy (2019)).

Two limits on that testimony bound what Flexpectation can claim for its own experiment framework. The first limit is that the accounts above, though the accounts come from senior practitioners across several organisations whose results can be checked independently, are testimony rather than measurement. Each account describes a different quantity — a rate at which ideas fail, a count of experiments per hour, the turnaround time of a tuning trial — rather than one shared metric. The second limit is that the documentation needed to measure any effect is itself largely missing. Gundersen and Kjensmo (2018) surveyed 400 papers drawn from four instalments of the International Joint Conference on Artificial Intelligence (IJCAI) and the Association for the Advancement of Artificial Intelligence (AAAI) conference series, scored each against 16 documentation variables grouped into three factors, and found that "between 20% and 30% of the variables for each factor are documented", with no paper documenting all of the variables. A field that records so little about how its experiments were run cannot easily measure whether a change to how the experiments are run helped.

Fields that cannot share their data are predicted to fall behind

Donoho's account of fields that cannot share their data describes Flexpectation's position, and that account is the part of his argument this project has to answer rather than quote selectively. Donoho predicts that fields with "inhibitions against data sharing, for example, because of confidentiality restrictions" will not make the transition he describes and "will be noticeably lagging behind in rate of progress". The substation telemetry this project uses is not (yet) published. Donoho also names the arrangement a field with data-sharing restrictions can still reach, which he calls a bring-your-own-data challenge: a shared task and shared code over data that "is private and only a few credentialed researchers ever get to see", as happens in clinical research. The leaderboard set out under "Leaderboards of machine learning results" above sits in that category — public to view and reproducible in method, with the underlying telemetry restricted. The honest reading of Donoho is that the arrangement recovers part of the benefit of an open challenge rather than all of the benefit.

What network operators have already built

We found nine projects run by electricity network operators that have already built a forecasting capability overlapping Flexpectation's. The last row of the table below is Flexpectation itself, so the comparison is direct. Where a project's published deliverables do not answer a column, the cell says so rather than being left blank. Flexpectation's own registration on the Smarter Networks Portal records a budget of £841,733 and a January 2026 to March 2028 delivery window.

Project What the project forecasts Scale Horizon Uncertainty published
Artificial Forecasting (Northern Powergrid) Demand and customer export at primary substations; active power at secondary 551 primary substations with export data, 171 modelled; 729 secondary substations Day-ahead to week-ahead at primary, evaluated to 11 days; week- to month-ahead at secondary Half-hourly, with 5th-to-95th-percentile bands
SSEN TRANSITION Net load, split into demand and generation, then recombined 13 primary substations, their bulk supply points, and their 33 kV and 11 kV feeders 30 minutes to 10 days A 40-member ICON-EU ensemble to 4 days, one deterministic forecast after that
SSEN FastTrack How the connections queue, around 180 GW, will load the distribution network Primary substations up to the grid supply point A planning horizon rather than an operational one A probability that a queued connection becomes real load
NGED's EFFS Grid supply points, bulk supply points, primary substation transformers, and generation sites Across NGED's whole distribution network 1 hour to 6 months None
UK Power Networks' Power Flow to Solar Capacity The capacity of unmetered solar behind each primary substation, then that solar's generation Not stated in what we read Not stated in what we read Not stated in what we read
SP Energy Networks' Predict4Resilience Electricity network faults, not load Per district Up to 4 days at 24-hour resolution in the project's published method paper; the project's registration document states up to 7 days A probability distribution driven by a 50-member ECMWF ensemble
Fox et al. (2018) (SP Energy Networks) The effect of weather on past peak demand, not a forward forecast 13 primary substations in the proof of concept, almost 400 in production Backwards over 10 years None
OpenSTEF (Alliander, the Netherlands) Net load, with a splitter into solar, wind, and residual parts Thousands of grid connection points To 48 hours Yes; the framework is built for probabilistic forecasting
Cordier et al. (2024) (Enedis, France) Consumption and generation at the substation since 2015; the finer-grid method the paper describes covers consumption, not generation All 2,300 high-voltage-to-medium-voltage substations, extending to 3,678 of the more than 5,000 transformers inside them, and towards 750,000 medium-to-low-voltage substations Not stated in the paper; the forecasts run at 10- or 30-minute resolution None stated in the paper
Flexpectation Net demand, with unmetered generation inferred 32 series in the trial area; 52 grid supply points, 271 bulk supply points, and 1,161 primary substations across NGED's whole distribution network from 2027 14 days, updated every 6 hours A 51-member ECMWF ensemble across the whole horizon

SSEN TRANSITION (2018 - 2023; £12.6 million in the project's close-down report, £14.5 million on SSEN's own project page) is the closest precedent we found for Flexpectation's method. TRANSITION split each substation's net load — demand minus whatever generation behind that substation happened to produce — into demand and generation, forecast demand and generation separately, then recombined the two forecasts. Flexpectation adds an ensemble that spans the whole 14-day horizon, and deployment across a whole distribution network. TRANSITION set out to build neither the full-horizon ensemble nor the whole-distribution-network deployment. TRANSITION's ensemble covered the first 4 days, so from day 4 to day 10 a single deterministic forecast was all TRANSITION had. Flexpectation's forecast horizon runs to 14 days. And TRANSITION was a 13-substation trial rather than a deployment across a whole distribution network. TRANSITION also used the distribution network's connectivity map — the record of which substation feeds which — throughout, and ranks "historical network connectivity data availability" as "just as important as historical net demand and generation measurements". That ranking is a GB operator's own verdict on the connectivity-map input Flexpectation plans to use explicitly. The rest of TRANSITION's published design matches what Flexpectation is building.

NGED's own Electricity Flexibility and Forecasting System independently selected XGBoost, which the system's evaluation reported as the most accurate of the three methods tested and as easy to automate. The project compared XGBoost against a long short-term memory (LSTM) neural network and against ARIMA. The evaluation report says XGBoost "provided the best results of the three methods tested, closely followed by LSTM", recommending XGBoost because XGBoost also allows simplified testing of features and can be easily automated. The report caveats that the LSTM could not be fully explored for want of graphics processing units, and expects that more testing would have brought the LSTM level with XGBoost rather than past XGBoost. Selecting XGBoost is the same starting point Flexpectation uses. EFFS ran from 2018 to 2021 as a Network Innovation Competition project costing £3.3 million, and its forecasts were deterministic. Publishing uncertainty bands is the step Flexpectation adds.

UK Power Networks' Power Flow to Solar Capacity is the direct predecessor of Flexpectation's unmetered-solar work, as challenge 8 above sets out.

SSEN FastTrack and SP Energy Networks' Predict4Resilience are both probabilistic, but they aim at different questions. FastTrack puts a probability on how much of the connections queue turns into real load and how that load behaves, which is a planning question rather than the operational question Flexpectation asks. Predict4Resilience drives a probability distribution of electricity network faults per district from an ensemble weather forecast, in a tool built with control-room engineers — the GB precedent we found for putting ensemble-derived distributions in front of network operators. Maia et al. (2026) publish the method: additive quantile regressions for ordinary fault counts, a discrete generalised Pareto distribution for the extremes, and a 50-member ECMWF ensemble carrying the weather uncertainty into both. Engineers at SP Energy Networks assessed the resulting forecasts in a trial running from October 2024 to March 2025, and found the forecasts "sufficiently reliable to inform decision-making".

SP Energy Networks has also published at Flexpectation's own voltage level, and the study is the GB precedent we found for putting gridded weather onto individual primary substations. Fox et al. (2018) ran a numerical weather prediction model over Scotland at 1 km resolution for 10 years, mapped that model onto each primary substation weighted by customer density, and used the model to separate the effect of weather on peak demand from the effect of everything else. Demand fell by between 1.4% and 4.8% for each degree Celsius of effective temperature, differing substation by substation with the mix of customers behind each substation. Every one of the 13 sensitivities was negative. Fox et al.'s method corrects history for planning rather than forecasting forward.

Two of the nine projects in the table are outside GB: OpenSTEF in the Netherlands and Enedis in France. OpenSTEF is also the only operational forecasting system run by a network operator in this review whose code can be read rather than inferred from a deliverable. OpenSTEF ships a component splitter that breaks a net-load forecast into solar, wind, and residual parts — the operational relative of challenge 8.

Enedis has forecast all 2,300 of its high-voltage-to-medium-voltage substations since 2015, and is now extending the forecast below the substation (Cordier et al. (2024)). The extension reaches 3,678 of the more than 5,000 transformers inside those substations, and is heading towards the 750,000 medium-to-low-voltage substations beyond those transformers.

Fitting a model to each transformer beat the method Enedis runs in production, which shares one substation forecast out across its transformers by fixed coefficients. The per-transformer models scored 6.0% mean absolute percentage error against 9.3% on the day those coefficients were refreshed, and 8.1% against 13.0% across the whole test period. That second comparison counts only the transformers whose coefficient then moved by less than 2.5%, and on that comparison 84% of transformers were more accurate under their own model. Cordier et al. chose both comparisons deliberately, as the cases where the fixed-coefficient method is "the most relevant and the most difficult to outperform". Cordier et al. do not say what their percentage error is normalised by, and report that the complete pipeline has not yet been evaluated end to end. Cordier et al.'s medium-to-low-voltage step was tested on about 100 substations using measured rather than forecast inputs. So the test measures the disaggregation rather than the forecast.

Northern Powergrid's Artificial Forecasting

Northern Powergrid's Artificial Forecasting is the closest concurrent project we found to Flexpectation. Artificial Forecasting is an Ofgem Strategic Innovation Fund (SIF) programme, with about £3.9 million of grant across its three phases, run by Northern Powergrid with Faculty, EV.energy, and Oaktree Power, the final Beta phase running to February 2027. The Beta deliverables that the rest of this section draws on sit under a separate project registration from the Alpha deliverables. Artificial Forecasting does much of what Flexpectation does at primary substations, and also covers secondary substations, which Flexpectation does not.

Artificial Forecasting has run operationally through a full winter flexibility procurement cycle. A forecasting service for primary substations is deployed and has passed Northern Powergrid's architecture review board, data governance, and information security checks for its current deployment. Northern Powergrid's System Forecasting team used the service operationally through a full winter flexibility procurement cycle to support week-ahead dispatch decisions. The service produces half-hourly probabilistic forecasts with 5th-to-95th-percentile bands, flags forecast exceedances of firm capacity, and is benchmarked against Northern Powergrid's existing growth-based and persistence methods and a rolling 4-week baseline. The deliverable states that performance did not materially degrade on average across the 11-day horizon.

Artificial Forecasting's value case puts whole-life net present value at around £60 million for one distribution network operator, or £250 million if three further operators adopt Artificial Forecasting. The net present value comes from a 3% reduction in spending on reinforcement — building bigger transformers and cables — in the current price-control period, rising to 6% in the next, and from a 25% improvement in the cost-effectiveness of contracted flexibility. Curtailment is not included in Artificial Forecasting's four benefit categories. The forecast covers customer export at primary substations. But the one published value case in this review puts no money on curtailment, which is in Flexpectation's scope alongside flexibility procurement. The Artificial Forecasting project pairs those figures with a direct caveat: Artificial Forecasting reports early Beta evidence, from one winter procurement cycle, supporting the performance assumptions behind the value case, which "remains appropriate, subject to further validation".

Artificial Forecasting is independent evidence that short-term substation forecasting is operationally useful, that a network operator will change its procurement process around a half-hourly probabilistic substation forecast, and that a benefits case has been made and accepted. Artificial Forecasting's core intellectual property is to be made available royalty-free to other GB distribution network operators.

Flexpectation is nonetheless attempting more than Artificial Forecasting's published deliverables describe. The two projects overlap on forecasting net demand at primary substations and on forecasting metered generation. Artificial Forecasting's Beta registration also lists load disaggregation among the project's innovations, describing "a novel approach to forecasting HV [high-voltage] load, separately modelling gross demand and distributed generation". The two series that approach separates are each already measured rather than inferred, which is a different task from the task Flexpectation takes on. The Beta annual progress report produces net demand "by independently modelling customer export data", the Alpha technical report covers "all 160 substations where both gross demand and customer export data were available", and the Embedded Capacity Register enters the model as an input feature, listing what is registered rather than estimating what is not. In contrast, Flexpectation's challenges 8 and 9 are the different problem of inferring an unmetered generator's half-hourly output from a substation's net flow, which is blind source separation.

Two more of Flexpectation's challenges do have a counterpart in Artificial Forecasting's deliverables. The Artificial Forecasting Beta annual progress report describes automated health checks and dashboards that "highlight substations where input data is degraded (e.g. faulty sensors, frozen or anomalous values)" and an extract-transform-load (ETL) pipeline that "flags frozen/spiky SCADA [supervisory control and data acquisition] data before modelling", which is Flexpectation's challenge 6. The Alpha user research treats planned and unplanned outages as data worth bringing in and as a reason to widen the error margin, which is a different response to challenge 4's problem rather than no response.

Five of Flexpectation's nine challenges have no counterpart we could find in Artificial Forecasting's published deliverables: tracking the effective capacity of metered generators; forecasting a substation as if it were always in its normal running arrangement, rather than dropping the periods when it was not; recovering signed net demand from an apparent-power meter; inferring unmetered solar and wind from a substation's net flow; and doing the same for heat pumps, chargers, and batteries. Across every Artificial Forecasting deliverable published on the Smarter Networks Portal — Discovery, Alpha, and Beta — searches for "abnormal", "unmetered", "apparent power", "non-directional", "blind source", and "source separation" return nothing at all. "Capacity" appears 123 times but never as an effective or derated capacity. And the five occurrences of a "switch" stem are generators switching on or off, switchgear asset types, and switching over a data feed. Heat pumps and electric vehicles do appear, as drivers of demand growth and as model features rather than as quantities separated out of a net flow. Flexpectation also delivers 1st and 99th percentiles where Artificial Forecasting's published bands run from the 5th to the 95th. The curtailment decisions in Flexpectation's scope turn on those outer percentiles.

Why we think this ambitious plan can be done

Measured against the studies we found, the plan for Flexpectation sits outside the published literature in nine ways at once. The nine absences are itemised in the summary above. The distance between Flexpectation's plan and the published literature says more about where our search fell short than about the quality of the work that fills the rest of the field. On the switching-events absences the nearest precedent we found is Liu et al. (2019), which conditions a forecast on an operating-state label — but for switching between transformers inside one substation, where the substation total stays metered throughout.

Flexpectation attempts all nine challenges above, across four families of model:

  • a heavily-tuned version of the gradient-boosting approach that won the tabular forecasting competitions reviewed above, and which NGED's own EFFS project independently selected;
  • weather and time encoders pre-trained on large datasets, so that a model for one substation can start from what has been learned across all substations;
  • models that use the connectivity map explicitly;
  • differentiable physics — building known physical behaviour directly into the model, so that the model has to learn only what the physics cannot supply: the response of a solar panel and of a wind turbine on the generation side, and the thermal response of buildings on the demand side. Gijón et al. (2025) fit a model of that kind to a single wind farm.

By the standard of scope in this literature, each of the four strands is a separate piece of work. Almost every study reviewed above takes on one of the nine challenges, at one voltage level, with one family of model. The few studies that touch more than one challenge almost all solve those challenges as a pipeline rather than together. Pre-training weather and time encoders and then reading a substation's probabilistic forecast off the encoders would be a full paper by that standard. So would each of the other three strands. Sizing the four strands as separate papers scopes the work rather than promising an output. How many of the strands survive contact with the data is exactly what the project has to find out.

Only the heavily-tuned gradient-boosting model, the first of the four strands, is in scope for Flexpectation version 1. The other three strands belong to the scale-up across NGED's whole distribution network from 2027, as does the disaggregation of unmetered generation. That scale-up is itself a falsifiable claim the project has written down: the architecture goes from 32 to about 2,500 time series without structural change (H5).

The main reason for attempting all nine challenges at once is that the nine may be one challenge rather than nine. A switching event, a turbine out for repair, and a stuck meter all surface in the same place: as a discrepancy between what a substation metered and what the weather and the calendar say the substation should have metered. Almost every study reviewed above that touches more than one of the nine challenges solves those challenges as a pipeline. The exception we found, Pierrot and Pinson (2024), fits one wind farm's time-varying capacity jointly with its probabilistic forecast rather than a substation's several challenges together. In the pipelines one stage's output is frozen before the next stage sees it. So an error made early cannot be corrected later, and the forecast error never gets to tell the capacity estimator that the estimate was wrong.

The question we want to answer is whether one model that estimates capacity, switching state, and demand together beats the serial pipeline every study we read used. NGED's specification leaves room for that combined approach, asking that capacity, switching state, and demand be handled, rather than that each be handled explicitly. The one published result we found that bears on the question points the joint way: de Vilmarest et al. (2024), described under challenge 3, removed the embedded wind and solar capacities from their model of GB regional net load. The adaptive version got 0.4% better, absorbing into its own coefficients what the explicit capacity figure had been supplying, while the offline, non-adaptive version got more than 10% worse. The de Vilmarest et al. finding is one result, on regions far larger than a substation, for one phenomenon out of several. There are reasons to doubt the finding generalises: we expect a gradient-boosted tree to do badly at the subtraction a two-stage residual hands the tree precomputed, and tens of thousands of training rows per series is a small sample in which to hope a model discovers an implicit baseline for itself. Neither expectation is measured here. We expect the answer to differ by model family, which is part of why the differentiable-physics strand matters: the differentiable-physics strand is the one family in which capacity, weather response, and demand are estimated jointly by construction.

One reason for confidence is that one more experiment takes compute time rather than staff time. The core forecast already exists and runs today, on an experiment framework that makes one more experiment take compute time rather than staff time. That low marginal effort is what makes it realistic to run on the order of hundreds of machine-learning experiments a month. The introduction to this review makes the same argument. The project states the claim as a falsifiable hypothesis: when experimentation is the active workstream, one person can register at least 100 leaderboard experiments in a month (H2).

Expecting several of the four model families to fail is what makes those model families research directions rather than engineering tasks, and both NGED and this project count a negative result as a real outcome. The honest expectation is that some deliver clearly, some produce a negative result worth publishing, and some are abandoned. Both NGED and this project count a negative result as an outcome: evidence that switching cannot be recovered from power data alone, for instance, would be worth having, because that evidence would justify extracting switching labels from operational systems instead of continuing to look.

References

Every source cited above, in alphabetical order by first author.