Effective-capacity estimation
Status: 🚧 Planned (v0.7). Epic: #141; issues: #157 (solar), #158 (wind). This page is the plan for estimating the time-varying effective capacity of metered generators (roadmap v0.7) — run as a head-to-head between candidate estimators, with the winner shipping in v1. The methods the candidates build on are explained in the techniques pages: Convex Optimisation and Differentiable Physics. The v2 work this feeds — full disaggregation of unmetered DERs — has its own canonical page, Net-demand disaggregation. The Python in this document is illustrative sketch code, not the implementation. See the roadmap index for status conventions.
The problem
A generator's effective capacity drifts over time — turbines fail, inverters drop out, panels soil
and degrade (or are cleaned and replaced), arrays are extended. A static nameplate value introduces
large downstream errors. The v0.7 deliverable is a time-varying effective-capacity series per
metered generator, feeding the
effective_capacity delivery table. Because this
turns v0.1's single scalar-per-series capacity into a time-varying series, the metrics pipeline must
also swap its time_series_id-only NMAE-denominator join for a temporal as-of join — see
Effective-capacity normalisation, and the v0.7 upgrade to
time-varying.
How capacity feeds the forecast: a two-pass approach. The first pass estimates effective capacity (this page); the second normalises each generator's time series by its effective capacity before the power-forecast model trains on it. XGBoost continues to do the actual power forecasting in v1; the estimators here are responsible only for estimating capacity of metered generators. None of the "clever" latent-demand or switching inversion happens here — that is v2 research.
The changepoints in the fitted capacity series are a deliverable in their own right: sudden, sustained drops are exactly the generator-fault signal worth surfacing to NGED, and a capacity series that names its change dates doubles as a real-time health and availability monitor.
Several estimators, one winner
There is more than one credible way to estimate effective capacity. The candidates occupy genuinely different corners of the tooling space. v0.7 aims to compare several candidates head-to-head on the same data and the same judging criteria, and the winner ships in v1. The losers do not disappear: they stay on the leaderboard as permanent baselines and honesty checks.
No published method we found already solves this across a mixed fleet, which is why we cannot simply adopt one paper's method. The energy-forecasting review found a method for each generation technology separately. But none run across a mixed fleet of individually metered generators at a distribution network operator — which is NGED's position, with solar, wind, a battery, a gas generator, and a biofuel plant behind one set of primaries.
The contest has a second, deliberate purpose beyond picking the best estimator: building hands-on experience with convex optimisation (CVXPY) during v1, so that its fit for the v2 problems — and our advice to NGED about tooling — rests on first-hand evidence rather than paper argument. The Convex Optimisation page makes a strong theoretical case; v0.7 is where that case meets NGED data.
The candidates:
- Candidate A — the convex estimator: a censored quantile-envelope fit with fused-lasso changepoints, solved exactly by CVXPY, with panel orientation found by grid search.
- Candidate B — the differentiable-physics estimator: the variational PyTorch model, fitting orientation and capacities as posteriors.
- Cheap baselines: a rolling quantile of clear-sky-normalised output, and SLAC's off-the-shelf convex capacity-change detection. Most simply, we may start with a simple rolling p99.
Despite coming from different toolchains, the two serious candidates are the same kind of method: both are inverse modelling — write down a forward model mapping unknown parameters (capacities, orientation) to predicted power, then invert it against the observed power to recover the parameters. Neither is "more physical" than the other; a fixed pvlib per-unit curve inside a convex problem is exactly as much a physics model as a transposition calculation inside a PyTorch module. What genuinely separates them is expressiveness versus guarantees: Candidate A restricts its forward model to what convexity can certify and gets the exact, reproducible global optimum in return. Candidate B may write down any differentiable physics it likes — and gives up the certificates. The head-to-head is therefore also a live measurement of that trade-off on NGED data; the framing is developed fully in Two routes to the same inverse problem.
There is a testable prediction to settle, made when this contest was designed: the convex estimator wins per-site, and the PyTorch route only pulls ahead when pooling many sites to learn shared physics corrections (soiling, snow, systematic irradiance-model bias) — which is fleet-scale work beyond v0.7. The head-to-head protocol is shaped to confirm or kill that prediction.
What every candidate must get right
The requirements in this section belong to the problem, not to any particular estimator. Each candidate meets them with its own machinery, and the judging checks each one.
What effective capacity must exclude
Active Network Management (ANM) curtailment is a deliberate, network-driven reduction, not a loss of physical capability. Folding curtailment into capacity would corrupt exactly the signal NGED needs. We identify curtailed periods from NGED's curtailment/ANM feed and keep them out of the capacity estimate — in the physics-model formulation this is a separate multiplicative curtailment gate on the generator's output (see the v2 engine's node definitions for where the same gate reappears at scale); in the convex formulation it amounts to masking or down-weighting flagged periods.
The ANM feed is imperfect, in both directions. Like any operational log, the ANM feed is an imperfect label: curtailment can happen with no matching log entry (for example, a generator's economic self-curtailment). A logged event may also differ from the generator's actual output. So the feed is a noisy label, not ground truth — use it, but do not lean on it:
- Unlogged curtailment (including economic self-curtailment during negative-price periods, which never appears in the ANM feed) would read as capacity loss to any estimator that fits the middle of the data. The structural defence is to fit an upper envelope instead — an asymmetric, quantile-flavoured loss under which curtailed samples fall below the fitted capacity without dragging it down. This defence matters for every candidate: it is native to the convex estimator's pinball loss, and the differentiable-physics candidate should adopt the same asymmetric loss in place of a plain Gaussian likelihood for the same reason. Masking periods flagged as negative-price from public market data is a complementary mitigation.
- Spuriously-logged curtailment merely discards good samples (the period is masked though the generator ran freely) — a milder failure, but one more reason the estimate should not depend on the feed being right.
Ask for the raw export cap rather than a derived curtailment volume, and treat the cap as a
constraint rather than as a label. Joining both to six metered solar farms in the beam/diffuse
experiment
found one of the six under active network management. Its cap sits at the connection limit for
80.7% of the half-hours since the scheme went live, and on the bright hours where the cap never
moved the site's yield matches the other five
to within half a percent, which is what says the cap is read correctly. A physics-based estimator
can also consume the cap directly, because it enters as min(what the weather allowed, the cap)
— an upper bound on export rather than a volume to subtract, so it changes nothing on the hours
it does not bind.
The setpoint record covers one generator from August 2024, so the estimator cannot depend on a curtailment label existing. That generator's setpoint history reaches back to the day after its telemetry begins, but the scheme does not enforce anything until 6 August 2024 and the cap reads a flat zero until then, so the usable record starts there. The other five generators have no curtailment record at all. So an estimator cannot depend on a label existing, and the upper-envelope loss remains the structural defence for every period and every generator the record does not reach.
The regularisation prior: piecewise-constant capacity
The capacity series must not be free to bounce around at the data's sampling rate, or it will simply soak up whatever noise the rest of the model cannot explain.
Metered effective capacity (can go up or down)
The effective capacity of a metered generator changes in both directions: it drops when turbines fail or inverters trip, and recovers when they are repaired. The right prior is piecewise-constant: capacity holds a level, then steps to a new one — which is expressed as a penalty on step-to-step change (a total-variation penalty, equivalently the fused lasso: an \(\ell_1\) penalty on successive differences). This lets capacity track genuine, persistent changes (a turbine offline for a fortnight) while refusing to chase half-hourly noise.
Published wind-capacity estimators split on exactly this direction of travel, and the numbers favour fitting over ratcheting. The energy-forecasting review records that Dantas and Browell (2026) estimate a wind farm's available capacity as a running maximum of its own metered production, a ratchet that can only rise. Viotti et al. (2026), by contrast, fit a piecewise capacity series by quadratic optimisation and publish both a monotonic and a non-monotonic variant. On hourly, region-aggregated Swedish data the non-monotonic variant gave the lowest day-ahead forecast error, 2.0% below the running-maximum normalisation on mean absolute error. But neither method improved clearly on Viotti et al.'s own de-rating test, which suppressed production for 30 days to simulate a fault. A ratchet cannot follow capacity down at all, which is why both candidates here are built to fall as well as rise.
The prior is shared; how exactly each candidate realises it is part of the contest. A proximal convex solver produces exactly zero change on most days — so the nonzero steps are a literal event log — while gradient descent produces approximately-zero changes that need a threshold to read as events (see the exact-zeros property).
(The installed capacity of an unmetered fleet behaves differently — it essentially only grows — and its monotone prior is documented with the v2 work: Unmetered installed capacity grows monotonically.)
Identifiability: the data goes silent at night
At night, and deep in winter gloom, the observed power says almost nothing about capacity — a 5 MW site and a 50 MW site both meter ≈0. Fitting those samples adds noise, not information. Every candidate should weight the fit by the per-unit physics output (or drop low-irradiance / low-wind samples outright), and let the piecewise-constant prior carry the estimate across the uninformative gaps — that is precisely what the prior is for.
Robustness to missing inputs
Both candidates ingest metered generation that really does have gaps — stalled telemetry, missed NWP runs, a wholesale-absent weather variable — and the winner's capacity estimate feeds v1.0 forecasting, so an estimator that mis-estimates under an outage propagates the error downstream. The missing-data argument here is the same reasoning as the section above: the data going silent at night is missingness with a known cause, and an outage is missingness with an unknown one.
So missingness robustness is a head-to-head judging criterion, scored against the same failure-scenario vocabulary the forecasting leaderboard uses (Metrics & Leaderboard). Two things are checked: that the estimator still returns an estimate at all under each scenario, and that its uncertainty widens honestly when it does — an estimator that quietly returns a confident number from half the data is worse than one that returns a wide interval.
The differentiable-physics candidate has a structural advantage here, and it should count in the judging. A physical forward model degrades most gracefully of all — an absent input is replaced with a prior or a physical bound, with no branching and no fallback path. The wider principle is Inherent Stability.
Keeping weather bias out of capacity
The irradiance driving any physics-based estimator is itself biased: NWP and satellite products carry regional, seasonal error (satellite retrievals degrade at low UK winter sun angles, for example). Per site, that bias is indistinguishable from slow capacity drift — an unconstrained fit will alias it into exactly the signal we deliver to NGED. The mitigation exploits structure: every site in a region sees the same weather bias, while genuine capacity changes are site-specific. Fitting a shared regional irradiance-bias term jointly across the metered fleet separates the two. This is the v0.7-sized version of what the weather encoder does in v2.
How naturally each candidate accommodates this term differs sharply — it is a genuine structural advantage of the differentiable-physics route and a documented caveat of the convex route.
Causal vs smoothed capacity — a lookahead trap in the two-pass scheme
A capacity series regularised over the whole record is a smoother: it uses future observations, so
the estimate for a given day changes once a later fault is seen. The smoothed estimate is correct
for the historical effective_capacity table and
the NMAE denominator — but the capacity used to normalise at forecast init time, in live running
and in backtests, must be the causal (filtered) estimate available at that init time, or
backtest skill is quietly inflated by lookahead. This causal-estimate requirement is the same
no-lookahead invariant the feature pipeline enforces for power lags. It binds every candidate
equally.
Candidate A — the convex estimator (CVXPY)
At first sight capacity estimation looks non-convex: predicted power is capacity × per-unit output, and the per-unit output depends non-linearly on panel tilt and azimuth. But the non-convexity lives entirely in two angles. Conditional on orientation, the whole problem is convex — and two angles are cheap to enumerate. That decomposition is the whole design:
- Outer loop: grid-search tilt × azimuth (~30 tilts × ~70 azimuths), computing the per-unit
output curve \(u_t\) with plain pvlib at each grid point.
(No PyTorch, and no
pvlib-pytorch, anywhere in this candidate.) - Inner solve: given \(u_t\), fit the capacity traces by solving one convex problem — certified global optimum, deterministic, warm-started from the previous grid point.
An exhaustive outer search wrapped around a certified inner solve is fully deterministic and, up to the grid resolution, leaves no valley unvisited — see the tooling page for where this trick sits in the convex/non-convex landscape.
The model
Two piecewise-constant traces per site, in daily blocks: DC effective capacity \(c^{\text{dc}}_{d}\) (the array side) and AC effective capacity \(c^{\text{ac}}_{d}\) (the inverter side). Predicted power is the DC output hard-clipped at the AC limit:
where \(u_t\) is the fixed per-unit physics output and \(d(t)\) maps each half-hour to its day. Daily blocks keep the problem small — roughly 730 unknowns per trace per year, which solves in well under a second.
The clip is censoring, not an obstacle
The hard \(\min\) looks like it breaks convexity, but clipped samples are observable in the data — flat plateaus at a common ceiling on bright days. Pre-classify each timestep as clipped/unclipped and the problem splits into convex pieces, exactly the censoring pattern from the tooling page:
- Unclipped samples: observed power ≈ \(c^{\text{dc}}_{d(t)} u_t\) — linear in the unknowns.
- Clipped samples: the plateau level is a direct noisy measurement of \(c^{\text{ac}}\); and the DC-side output must have exceeded the AC limit for clipping to occur — a linear inequality, added as a hinge penalty. Even censored points carry information.
This censoring treatment is Tobit regression in energy clothing — and the plateau classifier is where the formulation is weakest; see the caveats.
Loss and penalties
- Fit: a quantile envelope, not least squares. Curtailment appears as excursions below the physics prediction; least squares would drag capacity down. A pinball loss at a high quantile \(\tau\) fits the envelope: capacity = "what the site delivers when nothing holds it back", with curtailed samples falling below the envelope without biasing it. \(\tau\) is a real tuning choice (too low: curtailment leaks in; too high: the fit chases meter glitches) and should be validated on synthetic injections.
- Penalty: fused lasso on both traces. \(\lambda \bigl( \lVert \Delta c^{\text{dc}} \rVert_1 + \lVert \Delta c^{\text{ac}} \rVert_1 \bigr)\) gives sparse, exactly-zero day-to-day changes. The payoff is the deliverable itself: the fitted changepoints — each with a date and a magnitude — are the site's event log (inverter trips, feeder outages, array expansions, deratings), directly cross-checkable against NGED's maintenance records.
- Priors, where records exist. A registered capacity, a previous year's fit, or a connection record enters as one more convex penalty (priors as penalties) — including asymmetric priors and timing priors (a known March expansion makes jumps cheap at that date, expensive elsewhere).
- Which way an asymmetric prior leans depends on which capacity is being estimated, and it flips between milestones. For a metered generator — the v0.7 quantity — the register names an asset we can see, and effective capacity sits below that nameplate as soiling, degradation, shading and derating accumulate, so it is cheap to fall below the registered value and expensive to exceed it. For an unmetered fleet behind a substation — the v2 quantity — the register is close to a lower bound instead, because the domestic installations missing from it add capacity on top of what it lists, so the penalty should lean the other way.
- The register constrains \(c^{\text{ac}}\) only — it carries no direct-current rating. Every
capacity column in NGED's Embedded Capacity
Register is MW or
MVA. Checking the August 2026 release (7,211 rows, 5,236 of them solar): use
energy_source_&_conversion_tech_1_reg_capacity_mw, the only capacity field populated for every solar row, rather thanalready_connected_registered_capacity(mw)at 78% orconnected_maximum_export_capacity(mw)at 73%. The MW and MVA pair is one number rather than two, because MW is MVA × 0.95 for 99% of rows — an assumed power factor, not a measurement. The registered capacity equals the export MVA for 62% of solar rows and exceeds it for the rest, which is genuine export limitation. So \(c^{\text{dc}}\) has to come from the fit or from the assumed direct-to-alternating-current ratio; the register cannot supply it. - The register also cannot see domestic rooftop PV, which is why it is a lower bound for a substation's fleet. In the same release only 1.2% of solar entries sit below 50 kW and 64% fall between 50 and 250 kW, so the 3–5 kW domestic installations that make up the unmetered fleet are absent by design.
Where a DC:AC ratio would have to come from, and why we should measure ours rather than borrow one. Great Britain's registers split across the two units, so which source a capacity came from determines what it means:
| Source | Capacity reported |
|---|---|
| NGED's Embedded Capacity Register | AC only |
| Sheffield Solar's capacity report | DC only, by size band |
| Microgeneration Certification Scheme | both, per installation |
The Microgeneration Certification Scheme is therefore the one source that could yield a Great Britain DC:AC ratio broken down by size and by installation year, because it records both numbers for the same installation. That calculation needs record-level access we do not have. The scheme's data dashboard is free but serves only aggregates — counts and capacity by month, location, and technology — and a ratio needs both numbers on the same installation, so it would take a data request with terms we have not seen. Treat asking as a task worth trying rather than a source we can plan on, alongside asking for the CAMS uncertainty look-up table. Two further cautions before trusting such a calculation. The Department for Energy Security and Net Zero report that the scheme's DC field — which they call total installed capacity, against declared net capacity for AC — was often left empty in the early years of the Feed-in Tariff, so the DC side is sparsest over 2010 to 2014, exactly the period any time trend leans on. And the same department notes only that the gap between the two is widening for commercial solar farms; they publish no ratio.
The published ratios are rules of thumb, not measurements, and none of them is British. Solar consultancies and inverter vendors quote roughly 1.25 to 1.50 for utility-scale plants and 1.1 to 1.25 for domestic and small commercial (SLR, Solargis, RatedPower). The one well-documented trend comes from Lawrence Berkeley National Laboratory's Utility-Scale Solar series, which puts the American inverter loading ratio near 1.2 in 2010 and above 1.3 by 2017. Borrowing that trend for Great Britain would understate it if anything: lower irradiance means a given ratio clips away less energy, so oversizing is cheaper here than in the United States and the economics point to higher ratios rather than lower ones. Treat every figure in this paragraph as a sanity check on a fitted value, never as a prior.
Wind: same structure, simpler
A turbine has no separate inverter limit, but its power curve's rated plateau plays the identical censoring role: saturated high-wind samples measure capacity directly, and the same envelope + fused-lasso machinery applies with \(u_t\) from a reference power curve. Curtailment is a larger fraction of life for wind than for solar, which makes the quantile-envelope loss even more clearly right there.
Sketch
Illustrative, untested sketch, not the implementation — one site, one candidate orientation (the inner solve of the grid search):
import cvxpy as cp
# Inputs for one site at one candidate orientation:
# u: (T,) per-unit output from pvlib at that orientation
# y: (T,) observed power (MW)
# clipped: (T,) bool — the plateau classifier's verdict (see caveats)
# day: (T,) int — maps each half-hour to its daily capacity block
c_dc = cp.Variable(n_days, nonneg=True) # DC effective capacity (MW)
c_ac = cp.Variable(n_days, nonneg=True) # AC (inverter) effective capacity (MW)
# Unclipped, informative samples: pinball loss at a high quantile fits the envelope.
informative = ~clipped & (u > 0.15)
r = y[informative] - cp.multiply(u[informative], c_dc[day[informative]])
fit_dc = cp.sum(TAU * cp.pos(r) + (1 - TAU) * cp.pos(-r))
# Clipped samples: the plateau level directly measures AC capacity...
fit_ac = cp.sum_squares(y[clipped] - c_ac[day[clipped]])
# ...and the DC-side output must have *exceeded* the AC limit for clipping to occur.
censor = cp.sum(cp.pos(c_ac[day[clipped]] - cp.multiply(u[clipped], c_dc[day[clipped]])))
# Fused lasso: sparse, exactly-zero day-to-day changes -> piecewise-constant traces.
fused = cp.norm1(cp.diff(c_dc)) + cp.norm1(cp.diff(c_ac))
problem = cp.Problem(cp.Minimize(fit_dc + fit_ac + mu * censor + lam * fused))
problem.solve(solver=cp.CLARABEL) # certified global optimum, warm-startable
Honest caveats of the convex route
- Effective capacity, not nameplate truth. Any systematic bias in the pvlib per-unit model — soiling, unmodelled horizon shading — is absorbed into the capacity estimate. For forecasting that is a feature (it is effective capacity), but the fitted MW must not be read as nameplate truth against capacity registers.
- The plateau classifier sits outside the guarantee. The censoring split is a heuristic pre-step, and at half-hourly resolution with real meter noise it is fiddly: misclassified samples poison both the AC fit and the censoring constraint. It needs its own validation (synthetic clipping injected into unclipped sites is the obvious test).
- The shared regional irradiance-bias term breaks the one-shot convexity. A multiplicative regional bias × per-site capacity is bilinear — the same structure that makes the mixture model alternation-only. The convex route must then choose between: alternation (bias given capacities, capacities given bias — each step convex, but the certified-global-optimum headline no longer applies to the joint answer); a plug-in pre-estimate (e.g. the fleet-median clear-sky-normalised residual per region, fixed before the capacity solve — every solve stays certified but the two-step answer has no joint certificate); or an additive-only bias (jointly convex, but a physically weaker model of what irradiance error does). None is free. The differentiable-physics candidate fits the shared bias jointly without contortion — a genuine structural advantage to weigh in the head-to-head.
- Uncertainty is a bolt-on. The solve returns a point estimate; error bars come from bootstrap or sensitivity re-solves, with known weaknesses. See the uncertainty criterion below.
- Grid-search compute grows with the fleet. ~2,000 warm-started sub-second solves per site is trivial for the trial area's metered generators; at V2 scale (~2,500 series, though far fewer metered generators) it wants revisiting — e.g. a coarse-to-fine grid.
Candidate B — the differentiable-physics estimator
The variational single-site model
(DifferentiableSolarPlant
and its wind analogue): site coordinates locked, live weather passed through the physics, and tilt,
azimuth, and DC and AC capacity fitted as mean-field posteriors by gradient descent on an evidence
lower bound (ELBO). Its distinct strengths in this contest:
- Native posteriors. Every parameter carries a fitted spread, so capacity uncertainty comes out of the estimator rather than being bolted on — directly relevant to the uncertainty criterion.
- Joint fleet-wide corrections without contortion. The shared regional irradiance-bias term — and, later, learned corrections for soiling, snow, or systematic pvlib bias — drop into the same training loop. This is the "fleet-wide-with-learning" regime where the crossover prediction expects PyTorch to win.
- Continuity with v2. The fitted modules and the experience of training them carry straight into the v2 engine, where PyTorch is unavoidable.
Candidate B sits on the well-precedented half of the differentiable-physics strand. The energy-forecasting review found differentiable physics established for a generator's own output: Gijón et al. (2025) fit a turbine model to a wind farm's metered production, and Pierrot and Pinson (2024) fit a wind farm's capacity as a probability distribution jointly with its forecast, which is the shape Candidate B uses. The review found no comparable precedent for the demand-side half of the same strand — aggregating the thermal response of building stock up to a substation inside a probabilistic forecast — which is the half the v2 engine leans on instead.
And its costs, mirror-images of Candidate A's strengths: gradient descent brings learning rates, schedules, seeds, and stopping criteria for a per-site problem the convex route solves exactly; the fused-lasso-style penalty yields approximately-zero changes, so reading changepoints off the capacity trace needs a threshold; and there is no global-optimum certificate. Candidate B should also adopt the envelope-flavoured asymmetric reconstruction loss (above) — a plain Gaussian likelihood fits the middle of the data and inherits the curtailment bias.
Cheap baselines to beat
Both headline candidates must beat deliberately simple baselines on the same leaderboard:
- a rolling quantile of clear-sky-normalised output (observed power ÷ physics-predicted clear-sky power);
- the convex capacity-change detection in SLAC's solar-data-tools.
If neither candidate can beat the cheap estimators, that is a finding worth surfacing early — and the baselines slot into the same leaderboard discipline used for the forecasting models.
Uncertainty — a first-class judging criterion
Our working hypothesis (a hunch, stated so the contest can test it): capacity-estimate error will contribute a share of total energy-forecast error comparable to every other source combined. If that is even half right, a capacity estimate without honest uncertainty quietly launders one of the largest error sources in the system into numbers that look exact.
Two published results temper that hypothesis in opposite directions, which is itself a reason to measure it rather than assume it. Pierrot and Pinson (2024), the direct precedent for Candidate B's native posteriors, improved continuous ranked probability score by 34.2% over probabilistic persistence. But their one clean test isolating a varying capacity bound from every other change in their method gained 2.43%, which they call no significant improvement. de Vilmarest et al. (2024) removed embedded wind and solar capacity from an adaptive model of GB regional net load. They found error fell by 0.4%, against a rise of more than 10% for the same model fitted offline — evidence that an adaptive model can absorb a missing capacity signal rather than needing it, at the regional scale that result was measured on. Neither is a like-for-like test of a metered generator's effective capacity, as the energy-forecasting review sets out. But both are reasons the hypothesis above stays a hypothesis.
The hypothesis is cheap to test, and the contest should: perturb the capacity series by its plausible error band, run the perturbed series through the two-pass normalisation, and measure the change in downstream forecast NMAE. That experiment directly quantifies how much forecast error flows through the capacity channel.
What we ideally want from the winning estimator, in order:
- Uncertainty captured in the capacity estimate itself — a spread, not a point.
- Propagated through the two-pass scheme — run the normalisation per capacity-sample (posterior draws for Candidate B; bootstrap ensemble for Candidate A), so the forecast inherits the capacity uncertainty rather than ignoring it.
- Decomposable — the final forecast's spread should be attributable: this much from capacity, this much from weather, this much from the model. Per-sample propagation is what makes that attribution possible at all.
The judging stance, stated plainly: calibrated, decomposable uncertainty is a first-class criterion, not a hard shipping gate — a point-estimate-only winner may still ship in v1 with uncertainty deferred. But between candidates of comparable point accuracy, the one with honest, decomposable uncertainty wins — even at a slight cost in mean accuracy. This criterion structurally favours Candidate B's native posteriors over Candidate A's bootstrap; the docs say so openly rather than pretending the criteria are neutral, and Candidate A can close the gap by demonstrating that its bootstrap intervals are actually calibrated (see the evaluation hooks below).
The head-to-head protocol
All candidates run on the same sites, the same weather inputs, and the same folds. There is no direct ground truth for capacity, so judging combines proxies, each targeting a claim a candidate makes.
Run downstream forecast skill first, and let it decide. Each candidate's capacity series feeds the two-pass normalisation, and the resulting forecasts compete on the existing leaderboard (NMAE and pinball, per metrics-and-leaderboard). That is the deciding metric, and it is cheap because the leaderboard already exists.
The five checks below are diagnostics, built out for whichever candidate survives that first comparison — and for a losing candidate only where the margin was close enough that the diagnosis changes what we do next. Running all six on every candidate up front costs more than the contest is worth, and none of the five can overturn a clear result on forecast skill:
- Synthetic fault injection — scale a known period of a healthy site's output down by a known factor and check: does the estimator find the changepoint, at the right date, with the right magnitude? Does its uncertainty interval cover the truth? (The same injection discipline the switching detector uses.)
- Event-log quality — precision/recall of fitted changepoints against NGED maintenance and outage records, where records exist.
- Robustness to curtailment label noise — inject unlogged synthetic curtailment and measure how far each estimate is dragged.
- Uncertainty calibration — coverage of the stated intervals under fault injection and on held-out periods, per the criterion above.
- Runtime and operability — wall-clock per site, determinism across re-runs, and the count of tunable knobs that had to be tuned.
The winner ships in v1 and populates the
effective_capacity table; the rest remain on the
leaderboard as standing baselines. Whichever way it goes, the result settles the crossover
prediction with evidence — which is itself a deliverable of the
contest, feeding the v2 tooling choice and our advice to NGED.
Irradiance inputs
The beam/diffuse decomposition the physics needs (for either candidate — pvlib's transposition wants
the same inputs as the differentiable
model) is
covered by the weather ingests: the CAMS Radiation Service as the primary input, with two of its
15-minute values summed to the 30-minute window the meter averages over, and ERA5's near-real-time
ERA5T stream for the capacity estimate's freshness — see Data sources → Weather
data for both specs, why CAMS is preferred to CM SAF SARAH-3, and why
ERA5 beats CERRA here. The live ECMWF ENS feed carries only GHI — fine for v0.7, but v2 physics
forecasting of PV needs a differentiable GHI → DNI/DHI decomposition model, or fdir from
another source — see the forward model for both routes
and the sources that take them.
The shared irradiance-bias term has an expected sign, which gives it a prior. The CAMS Radiation Service reads high in clear conditions and low in cloudy ones (Lezaca Galeano et al. (2025), from an inspection run at two continental stations, and attributed more broadly to the Radiation Service's own validation reports). Capacity is identified mostly from clear periods, so an irradiance input reading high there pushes the fitted capacity down. That sign gives the regional bias term a prior direction rather than only a functional form. It is a weak test rather than a clean one: soiling and unmodelled shading push the fit the same way, as the caveats below note, so only an opposite-sign result is informative.
Fit capacity on clear periods, not on the whole record. A metered solar farm is small enough to behave like a point, while the irradiance estimate represents an area of several km², so every reading compares the two. That mismatch is smallest under clear skies, because a clear-sky irradiance field varies little over such an area, and clear periods are also where capacity is most identifiable. Restricting the fit also suits the plug-in pre-estimate, whose fleet-median residual is already clear-sky-normalised. Correcting the irradiance itself is v2 work on satellite irradiance over Great Britain.
Design caveat — should ERA5 stay offline? Feeding ERA5 into the live system adds a new near-real-time data dependency: another external feed to ingest on a daily-ish cadence, monitor, and recover when it lags. Because effective capacity moves slowly (daily blocks), a tempting alternative is to run the ERA5-based capacity estimation offline on a periodic job that refreshes the
effective_capacitytable, and keep the production forecast path dependent only on ECMWF ENS (plus the power feed) — no new real-time dependency, and the live forecast just reads the slowly-updated capacity table. The cost is that ECMWF-only would almost certainly give a slightly worse capacity estimate than ERA5 would. Worth weighing before we commit ERA5 to the real-time critical path; it shapes the live-service cadence.
What the beam/diffuse experiment measured on this fleet
A throw-away experiment on the six metered solar farms in the trial area fitted a five-parameter physical model per site, and several of its by-products bear directly on this plan. The write-up is Does a weather product's beam/diffuse split help a PV forecast?; the code is in a pull request kept open for reference rather than merged, #785, answering [issue
784](https://github.com/openclimatefix/nged-substation-forecast/issues/784). None of it is a
capacity estimator, and none of it settles which of the candidates above to build. What follows is what it established and the traps it hit.
Five parameters per site are identifiable from a meter and an irradiance series alone, and the fitted values are physically plausible. Panel tilt, panel azimuth, capacity, an inverter clipping limit, and a temperature coefficient were fitted per site on each training fold by a Powell optimiser, from one fixed starting vector and seven random ones. On the corrected timestamps the fitted tilts land between 14 and 27 degrees and the azimuths within 5 degrees of due south, which is what these arrays plausibly are. That is the core feasibility question behind candidate B, answered for solar on this fleet.
Fitting those five parameters needs no gradients, so candidate B's case has to rest on something other than the fit being hard. The objective is not convex: refitting all 30 site-and-arm combinations from 64 independent random starts each, a median of only 3 of the 64 reach the lowest loss found, and the worst start lands at 3.5 times it. What makes a derivative-free optimiser enough is the low dimension rather than a single basin — 5 parameters, and a fixed starting vector that finds the lowest loss in 19 of the 30 fits and is the only start to find it in 6. Refitting from the random starts alone moves every arm-to-arm difference by at most 0.0001 MW. So the arguments for candidate B are the ones already listed there — fitted posteriors, fleet-wide shared terms, and continuity with v2 — and not that a single site's PV parameters are hard to recover.
A half-hour timestamp error is absorbed into the fitted azimuth, so pin the timestamp convention before trusting a fitted orientation or the capacity that comes with it. The same model fitted on the same rows moves its fitted azimuth by about 35 degrees when the timestamps are shifted by one half-hour, which is a quarter of the range a GB array's orientation can plausibly occupy, from a 30-minute change in what the timestamp means. The power timestamps on this feed were half an hour late until 08:30 UTC on 26 March 2026. Any estimator that fits orientation has the same exposure, and an orientation error feeds straight into the capacity it reports.
What limits the fitted physical model is its specification, not its optimiser, and the term that hurts is the transposition. Given the same rows, the physical model and the tree disagree about which beam field is better, and the physical model divides the horizontal beam by the cosine of the solar zenith angle to recover the direct normal irradiance. That division magnifies a beam error without limit as the sun approaches the horizon: below 10 degrees of elevation the physical model's arm ordering reaches +0.56 percentage points against +0.00 to +0.17 in the bands above it. Holding tilt and azimuth equal across the arms makes the disagreement larger rather than smaller, from +0.128 to +0.147 points, so the fitted geometry is not the cause. An estimator built on the same model chain inherits the same sensitivity, and a floor on the zenith cosine is the cheapest guard against it.
One of the five fitted parameters lands outside physics, which is a warning about reading a fitted capacity as a measurement. The fitted temperature coefficient runs from −0.0018 to +0.0020 per degree Celsius across the six sites, where a crystalline-silicon module's maximum-power coefficient is negative and datasheets cluster between −0.0045 and −0.0025. A positive value means the fit is using that parameter for something other than the temperature response, and the fitted capacity absorbs whatever derating the temperature coefficient leaves unapplied. The fitted tilts and azimuths are unaffected. For this plan the lesson is that a capacity parameter is only as trustworthy as the terms beside it, which is an argument for the physically-constrained priors candidate B carries.
The effective_capacity table is a single snapshot, not a time series. It holds exactly one row
per time_series_id — 32 rows for 32 series — so code that sorts it by time and takes the last row
silently gets the only row. Nothing errors. The delivery-table
spec describes a time-varying trace, and the
estimator this page plans is what will produce one; until then, anything reading that table is
reading a constant.
A fixed-capacity model's signed error drifts by more than 6 points between neighbouring sites, which bounds how much capacity moved without estimating it. The per-site, per-year drift table is in the write-up. The drift appears on two independently-produced irradiance products and agrees between them to about half a point at the five longer-running sites, so it is in the power rather than in the weather. It remains an upper bound: the residual absorbs degradation, soiling, snow, curtailment, and any site-specific irradiance bias together. Separating them is the estimator contest's job.
A perfect capacity tracker would take about 4% off the error level of a PV forecast on this fleet, which bounds what this estimator can be expected to buy a forecast. An oracle correction subtracts each arm's own mean signed error inside every site-year, reading that mean off the rows being scored, so it beats any real estimator working at annual resolution. It takes the reference arm from 5.333 to 5.266% of P99 output; repeating it inside every site-month reaches 5.124. Both leave every contrast in that experiment where it was, to within 0.0003 points. Two readings follow. A dynamic capacity estimate is worth having for the quantity itself rather than for the forecast accuracy it returns, so the judging criteria are right to score the capacity trace rather than a downstream forecast. And a forecasting experiment on this fleet does not need one, because the error a static denominator leaves is shared across whatever arms are being compared.
Changing the capacity denominator changes the units and nothing else. That experiment normalises every error by each site's 99th percentile of output. Recomputing the headline against the 99.9th percentile, against the highest reading, and against a daylight-only 99th percentile moves the absolute figure between −0.096 and −0.086 percentage points and leaves the relative effect at 1.80% in all four cases. An estimator that reported a higher effective capacity would therefore not be validated by a forecasting score moving, and would not be refuted by one staying still.
The satellite retrieval beats the reanalysis by 4.29 points of mean absolute error on this fleet, which is a local measurement to set beside the literature above. The comparison and its caveats are on the data-sources page. That is a forecasting measurement rather than a capacity one, but it is the same inputs feeding the same physics, and it supports preferring CAMS for the capacity fit on more than a literature argument.
Curtailment needs the export cap, and the export cap needs its go-live date. NGED's setpoint feed reads zero for the six months before the scheme starts enforcing anything, while the generator exports normally. An estimator that took those readings as a constraint would conclude the site was pinned at zero while it ran. Honour the cap only from the first half-hour at which it reaches the connection limit. The evidence is under active network management caps what a generator may export.
A generator's commissioning ramp is signal for this estimator, not noise, and one is already measured. One trial-area solar farm climbed to its settled output through eight months of discrete multi-day plateaus, reaching its settled level on 6 October 2024 — exactly the upward, piecewise-constant movement the metered-capacity prior is built for. The forecasting experiment cuts those rows because a forecast has no ceiling to clamp to; an estimator should track them instead, and the plateau dates and levels make a ready-made test case. The figure and the method are under how a new solar farm reaches full output in stages.
Detecting that ramp needed a reference series rather than a threshold, which is a constraint on how this scales. A part-built array under a clear sky produces what a whole array produces under cloud, so nothing in a single series separates them. Dividing by the median output of neighbouring sites cancels cloud, season and time of day; a clear-sky irradiance model would do the same job where no neighbour is available. At roughly 2,500 series neither is free.
What comes after v0.7
Once the metered assets are accurately tracked, the harder v2 goal — disaggregating the unmetered DERs behind every substation — builds directly on this work. That plan lives on its own canonical page: Net-demand disaggregation.
The fitted plant models can also produce labelled data: run them forward on real weather, and every capacity the estimator must recover is then known by construction. Two conditions keep that v2 idea honest. The simulated capacities must be drawn afresh rather than frozen at the values the estimator itself fitted, or the score measures only that the fit reproduces. And the simulated irradiance must carry a different bias from the fit's, or the aliasing of weather bias into capacity drift becomes invisible by construction. Even then the idea supplements the synthetic fault injection rather than replacing it, because an estimator scored on data generated by its own model family only demonstrates that the parameters are identifiable.