Skip to content

Net-Demand Disaggregation β€” Approach & v2 Research Roadmap

Status: πŸ”¬ v2 research. This page is the canonical home for the v2 disaggregation arc: recovering latent demand and unmetered DER generation from net substation power, across the whole NGED network. It covers the plan and architecture; the methods it builds on are explained in the techniques pages β€” Differentiable Physics for the forward models and Convex Optimisation for the convex machinery. The sibling v2 arc β€” abnormal running arrangements and latent-demand recovery under switching β€” has its own canonical doc, Switching events & latent demand. How progress is measured lives in Evaluating disaggregation. Estimating unmetered capacity is roadmap v2.0, and the graph-structured engine is post-v2 research; both build on the metered-generator capacity estimates from v0.7. The Python in this document is illustrative sketch code, not the implementation. See the roadmap index for status conventions.

This page is the deep-dive behind two of the "innovative and unique" capabilities the Milestone 1 report highlights: natively handling unmetered generation and apparent-power (MVA) metering. (The third β€” dynamically-changing effective capacity of metered generators β€” is the v0.7 deliverable, planned in Capacity estimation; v2 builds on its output.)

The problem: net power is not demand

What a substation meter records is not demand. The meter reading is net power β€” the sum of true underlying demand minus behind-the-meter generation (rooftop PV, small wind, battery discharge) plus any unregistered or poorly-metered embedded generation. The "latent, unobserved demand" is the load that would be seen at the meter if all distributed energy resources (DERs) were removed. Recovering that latent signal is the disaggregation problem.

Compounding this, each primary substation spends roughly 10% of its operating time in an abnormal running arrangement (ARA) β€” a state in which switching events reroute a block of load from its normal parent substation to a neighbour. The metered signal is therefore structurally different from what it would be under normal topology. NGED requires forecasts expressed as if the network is always in its normal running arrangement β€” the latent demand under nominal topology, which is precisely the quantity network planners need. The full problem statement lives in the background docs (switching events, NGED's network); forecast building blocks covers how the "normal running arrangement" target is delivered.

The engine below attacks this by inversion through a differentiable forward model: model each substation's meter reading as the physical sum of its latent parts, then run the model backwards. The key product is not just the DER estimates but the latent demand signal itself β€” which can then be used directly as the target for a standard probabilistic forecasting model, free from the confounding effect of DERs.

DER tractability ranking

Disaggregation works best where there is a common observable exogenous driver and homogeneous behaviour across sites β€” conditions that let errors average out rather than compound. The tractability ranking across DER types is:

DER type Tractability Key reason
PV Excellent Irradiance-driven; panel behaviour is near-identical across sites; errors average out at fleet level
Wind Good Wind-speed-driven via a learnable power curve; more spatial heterogeneity than PV but still exogenous
Heat pumps Intermediate Temperature-driven with coefficient-of-performance (COP) rolloff; heterogeneity partly averages out at substation aggregate level
EVs Poor No clean exogenous driver; behaviour is synchronised (school-run, cheap-rate charging), so errors compound rather than cancel; synchronised peaks are exactly what matters to the grid
Batteries Very poor Pure latent control β€” tariff/market-driven with no physical exogenous signal; two identical batteries sitting next to each other can dispatch in opposite directions simultaneously

Practical conclusion for v2 scope: disaggregation targets PV (primary), wind (secondary), and heat pumps (worth attempting at substation-aggregate level). For batteries, the right approach is price-driven behavioural-clustering methods β€” the kind targeted by OCF's NESO "EDGE" project proposal (currently blocked waiting for data, as of June 2026) β€” rather than physics-based disaggregation. For EVs, honest publication requires wide uncertainty intervals and a clear caveat that the synchronised-peak regime (precisely the regime NGED cares about most) is the hardest case.

The literature backs both calls and adds a caveat for heat pumps, as the energy-forecasting review sets out: a Northern Powergrid code of practice puts batteries at a diversity factor of exactly one; NGED's own Electric Nation trial found a time-of-use tariff re-synchronising EV charging into the 22:00 hour, the failure mode the caveat above covers; and the one heat-pump diversity measurement the review found holds only in an average winter, untested in the cold snaps when a substation is under most strain.

The forward model

The substation meter reading is treated as the output of a forward model over the latent parts:

observed_power(t) = latent_demand(t) βˆ’ pv_generation(t) βˆ’ wind_generation(t) βˆ’ battery_net(t) + losses(t)

Each right-hand-side term is modelled explicitly:

  • latent_demand(t) β€” what we want: a smooth, weather-driven, time-of-week-structured signal representing the true underlying load.
  • pv_generation(t) β€” estimated from irradiance (from NWP and/or satellite) via a differentiable physics model of panel conversion efficiency (temperature and spectral correction, clipping at inverter limits). Panel capacity is a latent parameter estimated jointly.
  • wind_generation(t) β€” estimated from wind speed via a differentiable power curve.
  • battery_net(t) β€” handled via a state-space component with charge/discharge dynamics. (A later / stretch component β€” battery disaggregation is a v2 stretch goal in the roadmap.)
  • losses(t) β€” approximated as a smooth function of load level. (Also a later refinement.)

The live ECMWF ENS feed carries global short-wave irradiance only, so the beam/diffuse split the PV physics needs must come either from a decomposition model or from a second forecast source. ECMWF ENS from Dynamical.org publishes the global horizontal component alone, and the CAMS Radiation Service supplies the split over history but issues no forecast. The decomposition route means a differentiable model that splits global horizontal irradiance into its direct and diffuse parts. The second-source route has candidates already on the shortlist: ICON-EU (deterministic, 120 hours), the Met Office's UKV and MOGREPS-UK (126 hours or less), and WeatherNext 3 all carry a direct component. Of the four, only WeatherNext 3 reaches NGED's 14-day horizon. Which feed carries a direct beam sets out what asking for one would cost, and Is the beam/diffuse split worth having? sets out what the published evidence says.

Metered vs. unmetered DERs

A crucial distinction runs through the whole project: each generation term above is really the sum of a metered and an unmetered part. NGED meters some large DERs directly β€” utility-scale solar PV farms, large wind farms, grid-scale batteries β€” while a long tail of small DERs is unmetered behind the substation: domestic rooftop PV, small distributed wind, home batteries (and, strictly, EV chargers and heat pumps on the demand side). Written out in full, the forward model is closer to:

observed = latent_demand
           βˆ’ (pv_metered + pv_unmetered)
           βˆ’ (wind_metered + wind_unmetered)
           βˆ’ (battery_metered + battery_unmetered)
           + losses

We keep the compact equation above for readability, but the model treats the two classes differently:

  • Metered DERs are modelled per asset. Because we know the site exists and have its own generation meter, we can fit explicit, physically-interpretable parameters for it β€” for a metered PV farm, that single site's panel tilt, azimuth and effective capacity (see DifferentiableSolarPlant). This is the v0.7 deliverable β€” see Capacity estimation β€” and v2 consumes its output: verified, accurately-tracked metered assets are what anchor the harder unmetered inference.
  • Unmetered DER fleets cannot be modelled as a single asset β€” a primary substation may sit above hundreds or thousands of rooftops with a mishmash of orientations. These are modelled as an aggregate fleet node via the physics-informed basis expansion in UniversalSolarFleetNode. Estimating and disaggregating the unmetered DERs is the harder, v2 goal.

This is also where the project graduates from estimating capacity to directly forecasting power with the physics models (including for MVA-metered sites, below), and where the latent-demand inversion of the forward model is realised in full.

Is the beam/diffuse split worth having?

A throw-away experiment on the trial area's six metered solar farms has since run that comparison, and the answer depends on the product's resolution. On a 5 km satellite retrieval, giving the model the product's own beam field rather than a separation model's estimate from the same product's global irradiance cut error by about 1.5% relative. On a 31 km reanalysis it added nothing detectable. Both claims are about those two products on those six sites. The write-up is Does a weather product's beam/diffuse split help a PV forecast?, and the code sits in a pull request kept for reference rather than merged, #785.

Which product feeds the model matters far more than which split it sees. Swapping the 31 km reanalysis for the 5 km retrieval moved mean absolute error by 4.29 points of P99 output, against 0.126 points for the largest split contrast anywhere in that experiment. Any effort spent on the beam field is worth weighing against effort spent on the retrieval that carries it.

A 30-minute timestamp error is absorbed into a physical model's fitted azimuth, which is how a model-chain comparison can silently answer a different question. In that experiment the fitted azimuths move by about 35 degrees when the timestamps are shifted by one half-hour β€” whether the timestamp marks the start or the end of the averaging window. Settle the timestamp convention before comparing chains that fit orientation.

We found no study that runs the clean comparison β€” one NWP, one PV model chain, one arm fed the model's own direct beam and the other fed a separation model's estimate from the same model's global irradiance. We searched the terms "separation model", "decomposition model", "direct irradiance", "fdir", and "model chain" on the open web, in the project's literature/ library, and in the reference lists of the model-chain papers below. The absence is a statement about that search, not about the solar-forecasting literature as a whole.

In the one systematic model-chain study we found, separation was among the two steps that moved the error most. Mayer and GrΓ³f (2021) built 32,400 model chains from every combination of 9 separation models, 10 transposition models, 3 reflection-loss models, 5 cell temperature models, 4 performance models, 2 shading models, and 3 inverter models, and verified every chain against a year of 15-minute production data from 16 Hungarian photovoltaic plants, at day-ahead and intraday horizons. The gap between the best chain and the worst is 13% in mean absolute error and 12% in root-mean-square error. They name separation and transposition as the two steps that move the error most, and the inverter model as the step that moves it least. Those figures bound how much accuracy a well-chosen separation model adds. They do not measure how much accuracy skipping separation would cost, because every chain the study tested contains a separation model.

In the largest separation-model validation we found, every model's error grew under cloud enhancement and over bright ground. Gueymard and Ruiz-Arias (2016) validated 140 separation models against 1-minute measurements from 54 research-class radiometric stations across seven continents and four climate zones, 49 of the 54 from the Baseline Surface Radiation Network. They report that every model's error grows under cloud enhancement β€” bright cloud edges reflecting extra light onto the ground β€” and over high-albedo surfaces. Broken cloud matters for GB substations, because that is when a substation's solar output moves fastest.

The one recent post-processing study we found that builds on ECMWF still uses global irradiance plus a separation model. Horat, Klerings and Lerch (2024) post-process ECMWF ensemble global horizontal irradiance and put a separation step inside their model chain. Departing from that practice is worth measuring rather than assuming it gains anything.

Three caveats cut against the direct beam being an easy win. A forecast's direct beam is not ground truth, because the beam carries that model's own cloud errors. Over GB the diffuse fraction is high, so there is less direct beam for a better estimate to improve. And the differentiable PV model would learn its own corrections either way.

The physics chain requires the split, so the question is only which source should supply it. The test is the one this page already sets: held-out metered photovoltaic output, which is what the experiment above ran on six sites.

A gradient-boosted tree settles the question for the tree path alone. A tree fed the split as extra features can ignore those features, where the physics chain cannot proceed without the split. A null result from a tree is therefore evidence about the tree path, and does not justify dropping the direct beam from the physics plan. The experiment above used a tree as its primary instrument and a fitted physical model as a second one, and the two disagreed on the sign β€” the physical model divides the horizontal beam by the cosine of the solar zenith angle, which magnifies a beam error without limit near the horizon, so the physical model answers a different question rather than confirming the tree's. The write-up's comparison of the two instruments sets out how the fitted geometry was ruled out as the cause.

The graph-structured engine

The distribution network is fundamentally a topological graph, and we model it as one. The graph is a data structure: a fixed map of which substations can exchange load with which neighbours, and which sites see which weather. Each substation is reconstructed as the sum of its own differentiable-physics modules (gross demand, metered/unmetered PV, metered/unmetered wind), whose latent parameters β€” most importantly each module's capacity β€” are inferred directly from that substation's metered power and the local weather. The components are separable because each has a distinct exogenous driver and temporal signature (PV tracks irradiance, wind tracks wind speed, demand tracks time-of-week and temperature), so each substation's fit can pull them apart locally. The graph carries the structural prior β€” who can exchange load with whom β€” and a hard Kirchhoff balance closes the books. This mirrors the schematic in the Milestone 1 report (Fig. 10):

Schematic of a possible implementation of a graph-structured model, capturing electrical and
spatial relationship of different grid components

The generation and demand halves of this differentiable-physics engine stand on different evidence. The energy-forecasting review found differentiable physics established for a generator's own output β€” GijΓ³n et al. (2025) fit a turbine model to a wind farm's metered production β€” which is the precedent the photovoltaic and wind nodes below build on. For the gross-demand node the review found no comparable precedent: a search for differentiable physics applied to substation demand forecasting produced no strong result. The review also found nobody aggregating building thermal physics up to a substation and putting it inside a probabilistic forecast, though the ingredients exist separately.

Node definitions

Following the report's schematic, the graph uses the following node types:

  1. Substation nodes β€” the measured, net blended power flow at a primary substation (the main target constraint).
  2. Metered load nodes β€” demand that NGED meters directly, where such metering exists.
  3. Gross demand nodes β€” the underlying, unmetered consumer load, inferred by the model. Implemented as a BasisLoadNode: a shared MLP learns a small set of universal demand-profile shapes (e.g. residential, commercial, light-industrial) as functions of time-of-day, day-of-week, and temperature; each substation carries a local "style vector" of mixing weights that describes its particular customer mix. The universal basis curves are shared across all substations; only the style vector is site-specific, so the model can distinguish a residential suburb from an industrial estate without re-learning basic human demand patterns from scratch at each site. (Learning the shared shapes is non-convex PyTorch work β€” but once the dictionary is frozen, fitting a new substation's style vector against it is a small convex problem: cheap onboarding of new substations without retraining, and warm starts for the joint fit. See the fixed-shapes pattern.)
  4. Metered PV / Wind nodes β€” generators with dedicated, live generation metering.
  5. Unmetered PV / Wind fleet nodes β€” aggregated behind-the-meter (BTM) solar and distributed wind, grouped by location, with no direct metering. PV fleets are each implemented as a UniversalSolarFleetNode; wind fleets get their own analogous node type, with a shared learnable aggregate power curve in place of the orientation-mix bases (a turbine fleet has no tilt/azimuth to mix).
  6. Heat pump nodes (v2 stretch) β€” heat pump demand exhibits a distinctive J-curve: as temperatures fall, heating demand rises, but aggregate COP also falls, so electricity draw grows super-linearly with cold. This non-linearity cannot be captured by treating temperature as a plain regression feature β€” it requires an explicit COP-rolloff function. A HeatPumpNode models this: it takes ambient temperature, applies a learned COP curve, and outputs the net electricity demand attributable to heat pumps. Contributes to gross demand at substations with significant residential or commercial heat pump penetration.

Each generation node feeds the substation through a curtailment gate: a separate multiplicative factor, driven by NGED's Active Network Management (ANM) curtailment data feed, that represents network-enforced reductions. Keeping curtailment in its own gate (rather than inside the capacity parameter) is what lets the effective-capacity estimate stay a clean measure of physical availability β€” see Capacity estimation (including the caveat there that the ANM feed is itself imperfect) and Fig. 10.

The fusion mechanism

Spatial weather correlations (e.g. if it is raining at Substation A, the adjacent Unmetered PV Fleet B is probably cloudy too) are already supplied by the gridded NWP each node consumes. Where cross-site information genuinely helps the under-determined per-site fit, it enters as hierarchical parameter sharing: the unmetered-fleet and demand nodes share a small set of universal basis shapes (UniversalSolarFleetNode; BasisLoadNode above), with only a per-site style vector learned locally. Each node's physics modules compute explicit physical generation. A hard Kirchhoff balance node then aggregates the elements:

\[\text{Net substation flow} = \text{Gross demand} - \gamma_{\text{PV}}\,(\text{PV}_{\text{metered}} + \text{PV}_{\text{unmetered}}) - \gamma_{\text{wind}}\,(\text{Wind}_{\text{metered}} + \text{Wind}_{\text{unmetered}})\]

where the \(\gamma\) terms are the per-asset curtailment gates. The error between predicted and measured substation flow produces a gradient that flows back through the shared and per-site parameters, optimising them and the physical parameter posteriors simultaneously.

To be explicit about a design boundary: we do not currently plan to use message-passing graph neural networks (GNNs) anywhere in this project. The graph does its work as a data structure, and cross-site statistical strength arrives through the shared parameters and the hard flow balance above. We would not completely rule a GNN out towards the very end of the project β€” if measured residuals ever showed spatial structure that the gridded NWP and hierarchical sharing demonstrably miss β€” but nothing in the design depends on one. The same graph also underpins switching-event handling, where it is likewise used only as a data structure β€” see Switching events & latent demand, Part 2.

Unmetered installed capacity grows monotonically

The installed capacity of an unmetered fleet essentially only ever grows, as more households and businesses fit panels β€” unlike the effective capacity of a metered generator, which moves in both directions with faults and repairs. A monotonic representation is therefore the right prior here. We model capacity as a cumulative sum of per-week increments, each constrained to be non-negative, with an L1 (sparsity) penalty pushing most weekly increments to exactly zero β€” because installs happen in occasional bursts, not every week. The running total is then non-decreasing by construction. (This is the representation UniversalSolarFleetNode implements; in a convex host the same prior is a hard constraint plus an \(\ell_1\) penalty with exact zeros.)

Combining the physics with the weather encoder

+-------------------+
| Learnt parameters |
|     per site:     |
|                   |
|  β€’ PV tilt        |                   +---------------+
|  β€’ PV azimuth     |     <=======>     | pvlib-pytorch |----------+
|  β€’ AC capacity    |                   +-------+-------+          |
|  β€’ DC capacity    |                           ^                  |
|  β€’ etc.           |                           |                  v
+-------------------+                           |           +-------------+
                                         +------+------+    |  multi-seq  |
                                         | Irradiance  |    |  alignment  |
                                         | Temperature |    |  with axial |---> [ pΜ‚ ]
                                         +------+------+    |  attention  |
                                                ^           +-------------+
+--------------+     +-------------------+      |                  ^
| Weather data |---->|  weather encoder  |------+                  |
+--------------+     +-------------------+                  +------+------+
                                                            |   History   |
                                                            +-------------+
  • NWP bias is handled by the weather encoder: "the weather model says it's cloudy, but historically this specific pressure pattern at this location means it's actually clear" (feature-level correction).
  • Physical constraints are handled by the differentiable physics: "based on the corrected weather, the geometry of the sun and panel dictates \(X\) power" (first-principles baseline).
  • Systematic / local anomalies (the "unknown unknowns") are handled by the retrieval / alignment module: "on days that looked exactly like this in the past, the physics model consistently over-predicted the evening ramp-down by 5% because of that one tree on the horizon" (residual correction).

See Learned encoders for the encoder modules themselves, and pvlib-pytorch for the planned differentiable PV library in the diagram.

The convex dictionary baseline

Before (and alongside) the full engine, there is a much simpler disaggregator worth building β€” one that uses no neural networks and no gradient descent at all. The trick: replace learning continuous physics parameters with convex selection from a discrete menu.

Precompute per-unit output curves for a menu of candidate systems, all driven by the actual local weather: a few dozen PV orientations (tilt Γ— azimuth combinations), a few wind-turbine classes, a couple of heat-pump temperature responses. Each menu item is then a known signal β€” a fixed shape. Model the substation's net demand as an unknown residual demand shape minus an unknown, non-negative, slowly-growing amount of each menu item:

\[ \text{net}(t) \;=\; \text{demand}(t) \;-\; \sum_{k \in \text{menu}} c_k(t) \, u_k(t) \]

where \(u_k(t)\) is menu item \(k\)'s known per-unit output and \(c_k(t)\) its unknown installed amount. This is linear in the unknowns, hence convex β€” the fixed-shapes pattern end to end, with every prior available in its exact convex form: a sparsity penalty selects the few menu items genuinely present behind each substation (exact zeros, so the selection is literal); the monotone-growth constraint encodes that fleets don't shrink; and fitting many substations jointly against shared weather sharpens the identification. There is respectable precedent: Wytock and Kolter (2014)'s contextually supervised source separation did convex energy disaggregation in exactly this spirit.

Hard limits of the convex-only route β€” stated up front, because they define its role:

  • It cannot refine the menu. If reality sits between two menu items, the fit returns a blend; systematic error in the physics curves becomes bias, not a shape the model can learn to correct. (The full engine, which learns shapes, can correct this.)
  • It cannot do behaviour. EV plugging and battery arbitrage are not weather-shaped dictionary atoms β€” see the tractability ranking; batteries are already ceded to price-driven methods regardless of estimator.
  • No posteriors β€” the standard limitation of the convex route.
  • It cannot beat collinearity. Heat-pump load versus ordinary cold-weather heating: where the data cannot distinguish two stories, a convex model honestly refuses to β€” which is a feature for trustworthiness, but a ceiling on what it can resolve.

Its role: a transparent early disaggregator, and permanently the baseline on the disaggregation leaderboard β€” simple, reproducible, and embarrassing to any fancier model that cannot outperform it. The full engine is worth its added complexity only if it beats this baseline.

Apparent-power (MVA) metering

Some substations are metered only in apparent power (MVA), which reports the absolute value of flow and so cannot distinguish import from export. When embedded generation pushes power back into the grid, an MVA trace "bounces" off zero instead of going negative. Because the forward model reconstructs signed demand and generation explicitly, it handles this natively: we compare the measured MVA reading against the magnitude of the reconstructed net flow,

\[\text{MVA}_{\text{measured}} \approx \bigl|\,\text{Net substation flow}\,\bigr|\]

(assuming near-unity power factor). The physics grounds the model so that a sunny-day "bounce" is correctly attributed to reverse power flow from generation, not to a spike in demand. This MVA-magnitude reconstruction is one of the two capabilities the Milestone 1 report highlights for this engine β€” the other being unmetered disaggregation. (Note this reconstruction is intrinsically non-convex β€” the sign ambiguity means two valleys by construction β€” so it belongs to the PyTorch side of the tooling rule.)

Two implementation cautions:

  • The magnitude loss needs smoothing. \(|x|\) is non-differentiable at zero and its gradient flips sign there β€” exactly where the bounce lives. Compare against a smoothed magnitude, e.g. \(\sqrt{x^2 + \epsilon}\), and add a temporal-continuity prior on the sign of the reconstructed flow: flow direction persists for hours, it does not flicker half-hour to half-hour.
  • The near-unity power-factor assumption is weakest precisely at the bounce. As real power passes through zero, reactive power dominates the measured magnitude, so the MVA trace has a soft floor above zero rather than a clean reflection. Expect the reconstruction to under-fit the bottom of the bounce, and do not let the optimiser explain the floor with phantom demand. The energy-forecasting review confirms both cautions: a magnitude-only reading leaves more than one state of the network consistent with it, a result power-system state estimation has worked with since the 1990s. And apparent power is the magnitude of real power only near unity power factor, so the approximation is weakest exactly at the bounce. SSEN's TRANSITION, the closest published attempt to NGED's position we found, resolves the ambiguity using the meter's own history together with a model of the generation behind the meter, rather than a second independent measurement.

Handling abnormal running arrangements

Abnormal running arrangements (ARAs) β€” where switching events reroute load between substations, so the metered signal no longer reflects the normal running arrangement β€” are covered in their own canonical doc: Switching events & latent demand.

In brief, the v0.6 stage detects switching events with unsupervised statistics on the power series. The v2 stages reconstruct the latent demand each substation would have metered under the normal running arrangement, using a time-varying mixture over the neighbourhood graph (optionally type-resolved into demand / PV / wind, each a physics module as in the engine above). Two points matter for consistency with the rest of this document:

  • The graph is a data structure β€” who can exchange load with whom.
  • Conservation is a node-level flow balance across the 2–3-way fan-out observed in the trial area (a source's loss absorbed by a subset of neighbours whose pickups sum to it), not a pairwise equal-and-opposite transfer.

The alternative formulation β€” a discrete "switching state-space model" over per-feeder load blocks β€” is rejected: NGED's network is meshed and run radially with movable cut points, so there is no stable, re-identifiable feeder unit to discover and route (see switching-events.md, Part 4). The output β€” topology-normalised latent demand β€” remains the NGED-required target variable.

What already exists (prior art)

The component ideas each have precedent, which is important for calibrating the novelty claim.

Behind-the-meter PV and load disaggregation is a mature subfield. There is substantial published work on separating net load into behind-the-meter PV and native demand, including spatiotemporal GNN approaches where nodes are net-load measurements at neighbouring units and message passing encodes spatial correlation. Unsupervised methods that leverage the irradiance–PV correlation without any physical model also exist. "GNN over neighbouring nodes for net-load β†’ PV + load disaggregation" is, by itself, a known approach. Convex disaggregation also has precedent: Wytock and Kolter (2014)'s contextually supervised source separation is the direct ancestor of the dictionary baseline above.

The nearest GB precedent we found is a sibling Open Climate Fix project on the same problem β€” see the energy-forecasting review's assessment of UK Power Networks' Power Flow to Solar Capacity. The review found no published benchmark of inferring capacity from the net flow at primary-substation aggregation. The nearest published method at a comparable scale, Teng et al. (2023)'s DAZLS, splits unmetered wind and solar out of Dutch substation measurements but needs each site's installed capacity as an input β€” in contrast Flexpectation infers the capacity from the substation powerflow. The one result the review found that separated solar from demand at a real distribution substation without being told the installed capacity, Kara et al. (2018), needed the substation's own reactive power and a nearby solar plant's output standing in for irradiance, neither of which NGED's primary substations routinely supply.

Switching state-space machinery exists off the shelf. Recurrent switching linear dynamical systems (rSLDS; Linderman et al. (2017)) and explicit-duration variants (RED-SDS) are standard tools for unsupervised segmentation of multivariate time series into discrete latent modes. We considered this machinery for ARA handling but do not plan to adopt it: it presumes a discrete, re-identifiable switching unit (a per-feeder "block") that NGED's meshed, radially-run network with movable cut points does not possess (see switching-events.md, Part 4). Our chosen formulation is the continuous neighbourhood mixture described there.

Topology and switch-state identification has been studied, but overwhelmingly using voltage measurements. Voltage at primary substations is not part of this project's data feed. At half-hourly resolution, tap-changer movements would blur any topology signal in voltage anyway. Tap-changer movements could themselves reveal topology, but only in data sampled at around 1 Hz.

Nguyen et al. (2026) make evaluating one candidate switch configuration fast on a large electricity network, but they do not search over configurations. Their Sherman-Morrison-Woodbury update refreshes the inverse of the admittance matrix when switches move, instead of re-inverting that matrix from scratch. The speed-up scales with the size of the matrix. On their largest, 8,500-node feeder the update is 28 times faster than re-inversion; on their 13-bus feeder, only 1.1 times. On that small feeder the update is slower than re-inversion once the iterative refinement is switched on, and the paper recommends that refinement for small systems. The paper also reports the update is worst conditioned under "ill-conditioned switch reconfigurations with high-impedance tie switches or near-parallel paths". A search over a meshed electricity network with movable cut points would meet that case routinely. Nguyen et al. take the switch positions as a known input. So the update would accelerate the inner loop of a topology search that this project would still have to write. The update is a candidate building block, not a plan.

Where this work is novel

The novelty lies in the combination and problem framing, not in any single component:

1. Switching events as the primary disaggregation target, not an afterthought. Existing disaggregation literature treats the electricity network's topology as fixed and known. The ARA problem β€” where the topology itself is a latent variable that flips over timescales of minutes to months β€” has not been addressed in the disaggregation literature we reviewed. This is not a minor extension; it changes the structure of the inference problem fundamentally. The energy-forecasting review reports the nearest precedent it found as Liu et al. (2019), who condition a forecast on an operating-state label. But that precedent is for switching between transformers inside one substation, where the substation total stays metered throughout.

2. Power conservation as the cross-node inference signal. Prior spatial-disaggregation work uses spatial correlation as a soft prior. Here the graph edges carry a hard physical constraint: rerouted power is conserved across the affected neighbourhood as a node-level flow balance (a source's loss is absorbed by a subset of neighbours whose pickups sum to it). This is a stronger and more principled basis for cross-node inference than learned message passing.

3. Joint estimation of latent demand, DER parameters, and routing state. Existing approaches treat the topology as known, or the DER parameters as known, or the load as known, and estimate one unknown from the others. The joint inference problem β€” all three unknowns simultaneously, end-to-end differentiable β€” has not been cleanly tackled at this level of the distribution network.

4. The target variable is operationally defined by the DNO's requirement. Framing the output as "demand under normal running arrangement" is not just a modelling convenience β€” it is the variable that network operators actually need for planning and forecasting. This operational grounding, combined with the open evaluation protocol (rigorous, reproducible, multi-DNO leaderboards), is the publishable contribution that distinguishes OCF's approach from prior academic work.

5. Application at primary substation resolution with half-hourly data. The bulk of prior work operates at GSP/DNO-region scale (e.g. Sheffield Solar's PV Live) or at individual household level (NILM). The primary substation level β€” aggregating hundreds of customers, but below the GSP β€” is the level at which DER invisibility is operationally critical, and it is the level at which NGED's data exists. Systematic, open benchmarking at this resolution does not yet exist, as far as the published work we reviewed shows.

6. Real-power-only inference β€” the "no-voltage" constraint as a novelty claim, not just a limitation. As the prior art review notes, existing topology and switch-state identification work relies overwhelmingly on voltage measurements. Voltage at primary substations is not part of this project's data feed. At half-hourly resolution, tap-changer movements would blur any topology signal in voltage anyway. This work therefore demonstrates that the switching inference problem is solvable from real-power balance alone. Framing real-power-only inference as a deliberate design choice inverts the standard assumption and is itself a publishable contribution.

Technical architecture summary

Layer Component Role
Graph Primary substations as nodes; reconfigurable boundaries as edges (a plain data structure) Structural prior on which substations can exchange load
Forward model Differentiable physics (irradiance β†’ PV, wind speed β†’ wind power, state-space battery) Converts latent demand + DER params β†’ predicted meter reading
Reconstruction loss Squared residual, summed over nodes and time Drives joint inversion of latent demand and DER parameters
Cross-site coupling Hierarchical parameter sharing (shared basis + per-site style vector) + hard Kirchhoff balance Borrows statistical strength across sites
Transparent baseline Convex dictionary disaggregator The reproducible floor the engine must beat on the leaderboard
ARA handling Time-varying neighbourhood mixture with node-level flow balance β€” see switching-events.md Reconstructs latent demand under the normal running arrangement
Output Latent demand under nominal topology, per substation, per half-hour Target variable for downstream probabilistic forecasting
Forecast layer XGBoost or neural sequence model on cleaned latent demand Produces 14-day probabilistic forecasts in NGED-required format

Correcting satellite irradiance over Great Britain

PV is the DER we have the best chance of disaggregating well. Error we leave in the PV estimate does not stay in the PV estimate β€” the framework absorbs that error into gross demand or into another DER. So the irradiance input is worth improving.

Per-timestep uncertainty would help the disaggregation more than a corrected mean would. A variance that widens under broken cloud and narrows under clear skies feeds a heteroscedastic likelihood in the reconstruction loss. That is the probabilistic treatment the rest of the system already assumes. A corrected mean carrying no uncertainty tells the optimiser nothing about which timestamps to trust.

Check whether the Copernicus Atmosphere Monitoring Service (CAMS) has published per-timestep uncertainty for its Radiation Service before building anything. Lezaca Galeano et al. (2025) describe a concept study rather than an operational product: a look-up table conditioning the distribution of Radiation Service deviations on cloud probability, clear-sky index, and solar zenith angle. They report an average continuous ranked probability score of 50 W/mΒ² for global horizontal irradiance, with per-location values between 40 and 60 W/mΒ², and little change when the set of characterisation stations changes. Both British research-grade stations sit in their 66-station reference database β€” Camborne among the 40 stations that built the look-up table, Lerwick among the 26 held back to test it β€” so Great Britain is represented at both ends of its own latitude range. A conference talk reports work on delivering the model to users, but the request form still exposes no uncertainty variable.

If the uncertainty model is still unpublished when v2 starts, ask the paper's corresponding author whether they can share the look-up table, even as a one-off file transfer. A look-up table is small.

Do not start by correcting aerosols. Lezaca Galeano et al. did not condition their look-up table on aerosol optical depth or surface albedo, because the SHAP contribution of both ranked below cloud properties and irradiance level β€” testing them is named as future work. The Radiation Service already draws on 3-hourly aerosol analyses, so the aerosol term is both smaller than the cloud term and largely handled.

Do not interpolate raw residuals between weather stations. Perez et al. (1997) put the distance at which satellite irradiance overtakes irradiance interpolated from a ground station at roughly 34 km for hourly data. Zelenka et al. (1999) cite 20 to 30 km and go further, recommending satellite estimates even close to a measuring site, because irradiances 5 km apart already differ by about 15%. The Met Office's MIDAS network of weather stations has of the order of 50 to 80 British stations reporting hourly radiation across 229,000 kmΒ², putting the average point 27 to 34 km from its nearest station β€” at or inside the distance where interpolating stops helping. Cloud residuals decorrelate over tens of kilometres, so interpolating them mostly spreads one station's local cloud-timing error across a location 40 km away.

Some of the satellite-minus-station residual should deliberately not be corrected out. Lezaca Galeano et al. are explicit that their modelled deviation is not a difference from ground truth, but the expected spread between a spatial average over several kmΒ² and a point measurement. Which of the two we want flips between milestones: a single metered solar farm in v0.7 behaves like a point, whereas an unmetered domestic fleet spread across a substation's catchment is closer to the area average the satellite already reports. Correcting the satellite towards a pyranometer would improve the estimate for the solar farm and degrade the estimate for the domestic fleet.

If we publish any of this irradiance-correction work, publish a dataset with an evaluation report, not a method. Merging a network of ground weather stations into a satellite irradiance field at national scale was done for Belgium by JournΓ©e and Bertrand (2010), and site adaptation is routine in the solar-resource literature we read, so a novelty claim would not survive review. Validation would have to be leave-one-region-out rather than leave-one-station-out, hold Camborne and Lerwick out entirely, and beat a per-station monthly scale factor. The end-to-end test is held-out metered PV, because the disaggregation's own residual fits by construction.

Ideas for after Flexpectation

Two ideas outlive Flexpectation, and a map of installed DER capacity is the more valuable. The other idea is an irradiance nowcast. Irradiance recovered from metered generators' output, by running each generator's calibrated physics model backwards, is mostly a check on the capacity estimates rather than a second product. The irradiance nowcast is the one use of recovered irradiance that would stand on its own.

Combining the two ideas would give an estimate of live PV generation, which Sheffield Solar's PV Live already publishes for Great Britain. Flexpectation's capacity estimates would therefore be better offered to PV Live than published as a second estimate.

Publish a map of installed DER capacity across GB

A substation-level map of installed DER capacity is the more valuable of the two, and PV is the right place to start. No public source records where unmetered distributed PV sits at substation granularity: the Embedded Capacity Register covers registered connections, MCS covers certified installations at postcode-district resolution, and Sheffield Solar's PV Live estimates output at grid-supply-point level rather than capacity at substation level. NESO, the network operators, Ofgem, DESNZ, and local authorities planning heat and EV rollout each substitute a proxy for a number that drives connection decisions, flexibility procurement, and reverse-power-flow risk.

Three constraints would shape what could actually be published:

  • Data access binds harder than method. A map covering all of Great Britain needs telemetry from more than one network operator. An NGED-only map covers roughly a quarter of the country and is still worth publishing β€” but it should be described that way from the start, rather than implying national coverage.
  • Validation deserves at least as much effort as the disaggregation itself. Evaluating disaggregation owns the protocol. The spokes that bear on a published map are the two that read no register: synthetic aggregation and a manual capacity survey from aerial imagery. Corroborating against the Embedded Capacity Register and the Microgeneration Certification Scheme is weaker than it looks, because the capacity work plans to use those registers as priors. Publishing adds one difficulty the protocol page does not carry β€” the catchment boundaries are themselves uncertain, because which property sits on which low-voltage feeder is not public.
  • A monthly time series beats a snapshot. The fleet grows fast enough that one release dates quickly, and a monthly series answers the question network planning actually asks β€” where capacity is being added, not just where it now sits. Publishing monthly is a standing commitment and should be costed as one.

GB-wide inverse irradiance mapping

Once the architecture has calibrated, parameter-verified "virtual sensors" across the metered fleet, we can run the inversion trick at scale. Freezing the calibrated asset parameters and running gradient descent backward through the physics modules β€” from measured generation to the weather inputs β€” recovers a surface-irradiance estimate (and, for wind, a wind-speed estimate) at each metered site. These point estimates are sparse virtual observations; a spatial interpolation step (e.g. graph-based or geostatistical) then fills in a denser field across Great Britain. The result would be a half-hourly, physics-validated weather product, independent of the NWP, useful as a cross-check for real-time grid balancing. This is a research aspiration well beyond v2, and the density of the recovered field is fundamentally limited by the spatial coverage of the metered fleet.

Is a published irradiance dataset worth it?

Publish the recovered irradiance field as an artefact of the work that produced it, rather than making the field a goal in itself. Three arguments against targeting it:

  • The market is not underserved. CAMS, SARAH-3, PVGIS, and several commercial services already publish irradiance covering Great Britain, most maintained by funded teams. A new entrant has to prove it is better, and the users who care most β€” yield assessors β€” need a long, stable, documented record that a research product cannot offer for years.
  • The provenance is circular for the largest use case. A field inverted from PV generation is not independent of PV. Anyone wanting irradiance in order to model PV has been handed PV data in another coordinate system, and the obvious reviewer question has no clean answer.
  • Maintenance is a standing commitment. Versioning, reprocessing, a digital object identifier, and user support do not stop. A dataset that stops updating stops being used, and an abandoned dataset costs more credibility than never publishing.

The inversion earns its place as a diagnostic instead. A calibrated fleet implying an irradiance field that disagrees with CAMS systematically β€” in one region, or one season β€” is evidence of a fault either in the capacity estimates or in CAMS. That is the cheapest independent check available, and a strong figure in a paper.

An irradiance nowcast would be a more useful product

Latency is the one route by which a published irradiance field would stand on its own. Both archive products are far too slow to nowcast: the CAMS point service runs to yesterday, and the SARAH-3 Interim Climate Data Record lands 2 to 5 days behind. Going direct to EUMETSAT does not close the gap either, because Meteosat Third Generation's Flexible Combined Imager scans the full disc every 10 minutes, with a 2.5-minute rapid scan over Europe, against Meteosat Second Generation's 15 minutes β€” and a retrieval still has to run on top. Live PV telemetry arrives in minutes and is denser over Great Britain than any satellite retrieval, so an irradiance nowcast inverted from the fleet would occupy a niche the archives do not compete for. That is a different product from the historical field described above, and worth testing cheaply before committing to either.

Offer PV Live our capacity estimates, not a second live PV estimate

Multiplying the capacity map by irradiance would give the PV output behind every substation the map reaches. Installed PV capacity per substation, multiplied by an irradiance-driven model of the yield per kilowatt peak (kWp) of panels, gives PV output for any half-hour the irradiance covers. The recovered irradiance field covers past half-hours, and the irradiance nowcast would bring the estimate up to the present.

Sheffield Solar's PV Live already publishes an estimate of live PV output for Great Britain, at grid-supply-point and national level, and the National Energy System Operator (NESO) depends on PV Live. NESO's predecessor, National Grid's Electricity System Operator, funded Sheffield Solar's original PV Live methodology. The Electricity System Operator and then NESO have procured PV Live as a commercial service since 2021. PV Live models the yield per kWp from a live sample of reporting PV systems. PV Live then multiplies that yield by an estimate of installed PV capacity, compiled from national registers, for each grid supply point's area and for Great Britain as a whole. A second estimate of the same generation would duplicate that service. Wherever the second estimate disagreed with PV Live's estimate, every user would have to arbitrate between the two estimates.

Nationally, most of PV Live's error comes from its estimate of installed capacity, not from the statistical error of its yield model. Huxley et al. (2022) decompose the error in the national estimate for Great Britain. Huxley et al. put the capacity error at Β±5% and the statistical error in the yield model below Β±1%, and those two errors combine by root-sum-square to Β±5.1% overall. All three figures are three-standard-deviation bounds, expressed as a percentage of the national estimate, computed for the 12.86 GW of direct-current (DC) PV capacity installed in Great Britain in January 2020. Huxley et al. found no significant national bias from the makeup of the sample. Removing the yield model's statistical error entirely would therefore narrow the national error only from Β±5.1% to Β±5.0%.

Huxley et al. name three routes beyond static capacity registers, and Flexpectation's planned method for estimating PV capacity from substation power flows draws on two of those routes. Huxley et al. call for a move from static registers towards an estimate of operational grid-connected capacity: the capacity actually connected and generating, net of systems that are offline, faulty, or decommissioned. Huxley et al. suggest that estimate could come from power flows on the electricity network, from satellite imagery, or β€” the route they judge most likely to succeed β€” from a combination of complementary datasets. Flexpectation's graph-structured engine fits the capacity of unmetered DER, PV included, to power flows on the electricity network. Fitting to power flows is also the route of the nearest GB precedent we found, Power Flow to Solar Capacity, an Open Climate Fix project for UK Power Networks, separate from Flexpectation. The plan for estimating capacity also folds the registered capacity in NGED's Embedded Capacity Register into the fit as a prior β€” a starting value the fit is penalised for moving away from. Flexpectation's method therefore draws on Huxley et al.'s combination route too.

Regionally, Huxley found PV Live's yield biased for solar farms, because PV Live's sample held only domestic roofs. In the version of PV Live that Huxley (2021) analysed in the PhD thesis behind the paper, the yield model drew on about 20,000 domestic systems and on no commercial or utility-scale systems. Roughly half of Great Britain's PV capacity is commercial or utility-scale. Huxley tested the domestic yield model against the export-meter readings of about 700 solar farms. Huxley kept only systems above 3 MW, which are likely to be ground-mounted solar farms with little on-site consumption. At grid-supply-point level, the modelled yield sat slightly above the solar farms' measured yield under cloud, and up to 20% below that measured yield under clear skies. Clear skies rarely cover all of Great Britain at once, so the bias averaged out nationally.

Huxley attributes the clear-sky gap to hotter cells on domestic roofs, but other differences between solar farms and roofs probably contribute. A roof-mounted panel is insulated on one side, while a ground-mounted array has air flowing around it. Huxley's own arithmetic, though, puts a 20% loss from temperature at a 50 Β°C difference in cell temperature. The mounting coefficients of King, Boyson and Kratochvil (2004) put a glass module on a close roof mount about 17 Β°C hotter than the same module on an open rack, at 1,000 W/mΒ² and a 1 m/s wind. Huxley also lists optimised orientation and tilt, ground mounting, and cheaper panels among the ways commercial and utility-scale systems differ from domestic roofs, without saying which way each difference moves the yield. Huxley asks for directly metered output from a set of solar farms, to rule out on-site consumption and other on-site generation behind the solar farms' export meters.

PV Live applies a yield per kWp measured on domestic roofs to all the PV capacity in a region, so PV Live models a solar farm's output as reaching the inverter's limit no more often than a domestic roof's output does. An inverter caps a PV system's output at the inverter's alternating-current (AC) rating, which is called clipping. The higher a system's ratio of DC panel rating to AC inverter rating, the more often the system's output is clipped. Huxley explains why PV Live normalises yield by DC capacity: the DC rating is comparable across systems, whereas the size of each inverter relative to its panels varies between installations. The rules of thumb that consultancies and inverter vendors quote put the DC:AC ratio at 1.25 to 1.50 for utility-scale plants against 1.1 to 1.25 for domestic and small commercial systems. None of those rules of thumb was measured in Britain. On those ratios, a solar farm would reach its AC limit on more clear days than a domestic roof would.

Clipping lowers a solar farm's clear-sky yield, so the net 20% gap understates the differences that raise a solar farm's yield above a roof's. Splitting the net 20% into its parts needs metered solar-farm output, which the domestic sample does not contain.

The reporting systems Sheffield Solar has published are also almost all small domestic roofs, and their records carry no inverter ratings. Open Climate Fix publishes Sheffield Solar's reporting PV systems, with records from 2010 to 2025, as the uk_pv dataset. Of the 30,757 systems in the dataset, 98.6% are rated at 4 kWp or less, the median rating is 2.85 kWp, and 13 systems exceed 50 kWp. The dataset records each system's DC rating in kWp, its tilt, and its orientation, but not its inverter's AC rating. The records therefore cannot show how close each inverter sits to its panels' DC rating.

Fitting NGED's metered solar farms would give each farm's ratio of effective DC capacity to its AC ceiling, which a domestic sample cannot give. Both candidate estimators in Flexpectation's plan for estimating the capacity of metered generators fit a metered solar farm's panel orientation, and fit the farm's DC capacity and AC capacity separately. The DC capacity comes from unclipped half-hours, scaled against irradiance from the CAMS Radiation Service, and the flat plateau on a clear day measures the AC ceiling. Pooled across NGED's metered solar farms, the fitted ratios would give a DC:AC ratio by plant size for solar farms in NGED's licence areas. None of the public sources we checked publishes that ratio. NGED's metered solar farms are also the kind of solar-farm output Huxley asks for.

The fitted ratio has three limits: the DC capacity is effective rather than nameplate, the AC ceiling can be an export limit, and a farm that rarely clips leaves the ceiling unmeasured. The fitted DC capacity folds in soiling, degradation, and shading. The AC ceiling is the inverter rating or an export limit, whichever is lower. Flexpectation's capacity estimators are designed to keep curtailment out of the fit, so that a curtailment instruction under a flexible connection is not mistaken for the ceiling. A solar farm that rarely clips in Britain's climate leaves the ceiling unmeasured.

The orientation mix and the aggregate clipping of the unmetered PV fleet behind each substation would be weaker offers, because orientation and clipping are hard to tell apart in a substation's net flow. The graph-structured engine's model of an unmetered PV fleet fits the shares of the fleet facing east, facing south, facing west, and mounted on single-axis trackers. The fleet model also fits an aggregate AC limit and the sharpness of the shoulder where the fleet's inverters saturate one after another. An east-and-west orientation mix and inverter clipping both flatten the midday peak. Clipping shows only on the brightest clear days, while the orientation mix shapes every clear day and shifts with the season. Separating clipping from the orientation mix therefore needs many clear days spread across the year. The net flow at a substation also mixes the daily shape of PV generation with the daily shape of demand. The tilts and orientations in the uk_pv dataset could give a regional prior for the domestic orientation mix.

We know of no per-substation record of panel orientation, so a fitted orientation mix can be scored only on synthetic substations. Synthetic aggregation sums individually metered generators into a synthetic substation whose components are known, and scores the fit against those components.

Offer the capacity map first: a substation-level estimate of installed PV capacity is inferred from power flows, not from the registers PV Live relies on. PV Live derives its capacity figure from the Feed-in Tariff database, the Microgeneration Certification Scheme, the Renewable Energy Planning Database, and Solar Media's market data, and publishes that figure quarterly at national and regional level. A capacity map inferred from substation power flows would be independent of all four sources. Flexpectation's plan for estimating capacity folds in NGED's own Embedded Capacity Register (ECR) as a prior. The ECR lists almost no installations below 50 kW. Most of the domestic capacity behind a substation would therefore rest on the power flows alone. The public sources we checked record no unmetered PV capacity at substation granularity.

PV Live would receive the capacity map summed to each grid supply point in NGED's licence areas, not a map of Great Britain. PV Live works at grid-supply-point level. A map built from NGED's power flows covers only NGED's licence areas.

Offer the fitted DC:AC ratios second, because PV Live could use a capacity figure within its existing method but could use a ratio only by changing that method. A capacity figure is a new input to PV Live's existing multiplication of yield by capacity, once the soiling, degradation, and shading the fit folds into effective capacity are added back, using an estimate of those losses. The add-back matters because PV Live's yield per kWp already carries a domestic system's typical losses, so multiplying that yield by an effective capacity would count those losses twice. A DC:AC ratio changes PV Live's estimate only if PV Live models solar-farm capacity apart from domestic roofs. The orientation mixes and aggregate clipping of the unmetered PV fleets come last.

The nowcast is a stronger offer for PV Live's near-real-time estimate

PV Live's near-real-time estimate runs on a far smaller sample than its day-plus-one estimate, and that gap is a distinct source of error from the capacity and yield errors above. In the version Huxley (2021) analysed, the near-real-time sample was "of the order of 1000 systems," against the roughly 20,000 domestic systems the day-plus-one estimate draws on. Huxley found the near-real-time estimate carried a root-mean-square error of 640 MW against the day-plus-one estimate. Correcting for a lag in capacity reporting brought that error down to 293 MW. Huxley attributed the remaining 293 MW to the smaller near-real-time sample itself. Sheffield Solar's current documentation describes the same two-stage structure: PV Live recomputes each half-hour's outturn on day-plus-one, "to make use of sample data which only becomes available on day+1."

That near-real-time sample spreads thin across Great Britain's grid supply points, so a thin or unrepresentative sample at one grid supply point would show up in that grid supply point's regional estimate. NGED's licence areas alone contain 52 grid supply points, and Great Britain has several times that number. Splitting a sample "of the order of 1000 systems" across that many grid supply points leaves only a handful of reporting systems per grid supply point on average. We found no published breakdown of the near-real-time sample by grid supply point, so that average is arithmetic from Huxley's national figure, not a number Sheffield Solar has confirmed regionally β€” some grid supply points plausibly hold far fewer live-reporting systems than the average, and some plausibly hold more.

The nowcast could close that regional gap without the circularity that rules out publishing a general-purpose irradiance dataset. Multiplying the capacity map by irradiance uses each metered generator's own fitted orientation and DC:AC ratio, not a curve fitted to domestic roofs, so the circularity that applies to a domestic yield model does not apply here. The estimate would still cover only NGED's licence areas, not every grid supply point where PV Live's near-real-time sample runs thin, and testing it needs a live comparison against PV Live's own near-real-time output rather than the historical validation the disaggregation protocol already covers.

Offering PV Live this nowcast would not mean handing over any customer's raw PV power data. The inversion holds each generator's own fitted capacity, orientation, and DC:AC ratio fixed and inverts that generator's measured power into an estimate of the irradiance falling on it. Those fitted parameters are exactly what make a generator's power curve identifiable, and inverting for irradiance factors them back out, so the recovered value carries the local weather rather than that generator's output. Aggregating the recovered irradiance across every generator behind a grid supply point into one estimate per grid supply point per half-hour, before anything leaves NGED's own systems, removes what a single generator's estimate could still reveal on its own.

Evaluating disaggregation

There is no single clean ground truth for disaggregation, so progress is measured with a multi-pronged protocol β€” see Evaluating disaggregation.