Metrics & leaderboard
How OCF measures the skill of its forecasts and compares forecasting approaches.
Status legend β β Implemented Β· π§ Planned Β· π¬ Research. The
Metricsschema, themetricsDagster asset, the deterministic metrics (MAE, NMAE, RMSE, MBE), and the probabilistic metrics (CRPS, spread-skill ratio, pinball loss, PICP, interval width β see the evaluation-metrics reference) are β implemented. The interactive leaderboard visualisation is π§ planned. The implemented cross-validation protocol has moved out of the roadmap. See the roadmap index for status conventions. The π§ items are tracked under the v0.3 epic #6: baseline forecasters #147 Β· probabilistic evaluation #225 Β· tail & exceedance metrics #254 Β· tricky-days filter #255 Β· fold hygiene #226.
The leaderboard concept π§
Issue: #4
A key deliverable is a leaderboard comparing many forecasting approaches. We plan one
leaderboard per time_series_type, e.g. primary substations, GSPs, BSPs, solar PV sites, wind
farms, BESS, etc.
Each leaderboard will have tens (maybe hundreds) of rows. Each row is one ML experiment: a particular model, trained with a particular set of features, processed a particular way. Entrants must be compared apples-to-apples β same test dataset, same metrics, same assumptions.
Per-experiment configuration, trained weights, and metrics are stored in the project's MLflow database. The leaderboard will be displayed as an interactive table showing multiple metrics at a glance, inspired by the WeirdML leaderboard:

Baseline forecasters π§
Issue: #147
No naive baseline exists anywhere in the codebase (only docstring mentions, e.g.
contracts/power_schemas.py:242). Until the leaderboard carries naive rows, XGBoost's NMAE
numbers aren't interpretable β and, more to the point, we can't answer the question this project
exists to answer: do we beat what NGED does today?
The headline baseline β nged_incumbent
nged_incumbent is a faithful reproduction of
NGED's incumbent forecast β the analogue-ensemble method
they use today, with no weather model and no ML. In brief (full description and the operator's-eye
view are in the background page): for each target half-hour it takes the observed power at the
same weekday & time-of-day from the last 6 weeks and from 49β55 weeks back β 13
analogues β which NGED plot and read by eye (taking the 95th percentile if they need a single
number). Reproducing it matters because it is the bar we have to clear to justify the project β
"XGBoost beats persistence" is table stakes; "XGBoost beats the incumbent" is the deliverable. It is
the first baseline we implement; if we implement only one, it is this one.
It slots into our machinery beautifully, because every one of its 13 members is just a power lag:
- Weekly group (last 6 weeks, same weekday & time):
power_lag_168h, 336h, 504h, 672h, 840h, 1008h - Annual group (49β55 weeks ago, same weekday & time):
power_lag_8232h, 8400h, 8568h, 8736h, 8904h, 9072h, 9240h
So it rides the same audited, no-lookahead pipeline as PersistenceForecaster (below) with zero
new time-series logic. _nullify_leaky_lags already sheds the shortest members as lead time grows
(past 7 days the 168 h member nullifies, past 14 days the 336 h, and so on), leaving the annual
members to carry the full 14-day horizon. Because the shortest member is a week old, the incumbent
has no short-horizon skill from recent power β realistic, since that is exactly what NGED do
today, and a reason to keep the pure PersistenceForecaster as a contrast rather than to sneak a
recent-power member in.
It is also our first probabilistic baseline β and this is the faithful representation, not a
bonus. The plotted spread is the incumbent's output β an operator reads it by eye. We emit the
13 analogues as 13 ensemble_member rows and let the probabilistic
metrics score them for free β scoring
the spread is the closest automatable proxy for the plot a human actually reads. Two consequences
worth stating plainly:
ensemble_memberis overloaded here. For NWP models that column indexes an NWP ensemble member; fornged_incumbentit indexes a historical analogue. Same column, different meaning. We document this on thePowerForecast/AllFeaturesschema so nobody assumesensemble_member β NWP. The incumbent synthesises its ensemble insidepredict()(by unpivoting its analogue-lag columns into member rows) rather than consuming an NWP ensemble; it runs withweather_source: "none".- Deterministic collapse is a property of the metrics layer, not the incumbent. The incumbent
emits its 13 members and nothing else; the metric-matched collapse
decision then scores its MAE
on the members' median (apples-to-apples with every other model's central forecast) and reports
NGED's actual operating point β the 95th percentile β as a labelled secondary number
(
mae/mbeatmetric_param="p95"). Being deliberately conservative, the P95 carries a large positive MBE by design (a peak-safety choice, not a forecasting error), so it belongs beside the central metric, not in place of it. Either way NGED weight the analogues equally ("no further processing at all"), so equiprobable members β and the probabilistic metrics (CRPS etc.) computed over them β are faithful, not an approximation.
A faithful replica and a "cheap upgrades" variant
We implement two closely-related incumbent baselines, and the pair carries a message that is itself a valuable project outcome β most of the benefit may come from a few simple upgrades to what NGED already do, not from heavy ML:
nged_incumbentβ the faithful replica above. No holiday handling; warts and all. Pure lag features.nged_incumbent_holiday_alignedβ the same skeleton, but analogue selection becomes calendar-aware: a bank-holiday target draws from prior bank holidays / the matching day-type (a bank-holiday Monday behaves like a Sunday), and moveable feasts align holiday-to-holiday (EasterβEaster) rather than by fixed week offset. This no longer rides the pure lag machinery β the analogue offset is conditional on the calendar β so it needs a bespoke picker plus a GB bank-holiday calendar (the pure-Pythonholidayspackage), and ships as an immediate follow-up PR, not part of the first one.
Persistence and climatology β diagnostic bookends
The incumbent is really a hybrid β its weekly group is persistence-like recency, its annual group is climatology-like seasonality β so the two pure forms are still worth having: they isolate short-horizon vs long-horizon naive skill (at 0β6 h persistence is famously hard to beat; at day 8β14 seasonal climatology often beats everything). A model could "win" the leaderboard while adding no skill over either, and without these rows nobody would know.
Side benefit: several more BaseForecaster implementations pressure-test the abstraction (the
docs promise the interface is model-agnostic; today only XGBoostForecaster exercises it).
Implementation details β baselines (deleted when they ship)
Five PRs, in order. PRs 1β2 are shared-framework groundwork (no baseline yet); PRs 3β5 add one
baseline each. The nged_incumbent_holiday_aligned variant (described under A faithful replica
and a "cheap upgrades" variant) is a later
sixth PR, out of scope for this arc but given its own tracked issue so it is not lost when #147
closes.
Guiding principle β no special path. New workspace package packages/baseline_forecasters/
mirroring the xgboost_forecaster layout (pyproject.toml, src/baseline_forecasters/,
tests/), added to the root pyproject.toml [tool.uv.sources] and dependencies. Every baseline
subclasses BaseForecaster and rides the identical asset chain
(register_experiment_job β trained_cv_model β cv_power_forecasts β metrics) that
XGBoostForecaster uses β a train() that only records trained_time_series_ids and a
meta.json-only save() cost nothing, and the uniform provenance is exactly what makes re-running
routine. Re-running CV for any algo is then a native Dagster operation: select trained_cv_model++
in the asset graph (two + β metrics is two hops downstream, so a single + would silently stop
at cv_power_forecasts and skip scoring), choose the experiment's partitions, launch a backfill;
the unpartitioned metrics asset materialises once afterwards with its default config (no filter,
leaderboard scope). Verify this drill end-to-end on a smoke-test fold before documenting it β the
Dagster version supports mixed partitioned/unpartitioned backfill selections, but confirm the UI
behaviour rather than assuming it. To re-score a single experiment without touching the rest, use
the metrics asset's PopulationFilter config instead.
PR 1 β deterministic-collapse rework in compute_metrics. Implements the metric-matched
collapse decision. Forecasters
emit members only; every collapse lives in the metrics layer, so there is no per-experiment
collapse config and no designated point-forecast columns on PowerForecast. In
packages/ml_core/src/ml_core/metrics.py:
- In the per-run aggregation, emit both
power_fcst_mean(.mean()) andpower_fcst_median(reuse the already-computedq_p50empirical-quantile column β Polars.quantile(0.5, "linear")is exactly.median()), and keepq_p95. - Score
mae/nmaeon the median error;rmse/mbeon the mean error; the spread-skill denominator stays the mean-error RMSE (its Fortin "1.0 = calibrated" target is defined against the mean). Add extra labelled rows:mae/mbeatmetric_param="p95"(NGED's operating point) andmbeatmetric_param="p50"(bias of the delivered median).METRIC_PARAMSalready contains"p95"and"p50"(both are inDELIVERY_QUANTILES), so noMetricsschema change β and the primary key includesmetric_param, so the new rows do not collide with themetric_param="all"headline rows. - Extend the MLflow allowlist
_MLFLOW_LOGGED_PARAMETRICwith("mbe", "p95")and("mbe", "p50")so the operating-point bias and the median's bias appear on the leaderboard (aggregate keys come out distinct, e.g.mbe_p95__allvsmbe__all). - Update the docstrings that currently say deterministic metrics are "scored on the per-run ensemble
mean" (
METRIC_NAMESincontracts/ml_schemas.py), theMetrics.metric_paramfield description (no longer pinball-only), and the_MLFLOW_LOGGED_PARAMETRICkey-count claims. The promoted evaluation-metrics reference section must state explicitly that MAE/NMAE and RMSE/MBE score different point forecasts and why (otherwise the first person to recompute RMSE from the stored median members files a bug), and note the identitymae β‘ 2 Γ pinball_loss@p50as a deliberate internal consistency check. - Tests: member sets where mean β median β p95, with hand-computed expected values per metric; the single-member ensemble (all collapses coincide; CRPS still reduces to MAE); the pinball-p50 identity. Recompute existing expected values β never relax a test to absorb the shift.
- No standalone re-score needed: the one backfill after PR 2 (below) covers it.
PR 2 β CV predict-path framework: uses_nwp_ensemble, power-lag lookback, ensemble_member
docs. Three changes to the shared rails, none baseline-specific.
uses_nwp_ensemble: ClassVar[bool] = TrueonBaseForecaster.cv_power_forecastspassesensemble_members=[0]to_load_engineering_inputswhen a class sets itFalse. Semantics to document:Trueβ the forecaster consumes the NWP member axis (output members = NWP members);Falseβ the forecaster does not fan out across NWP members β it either passes member 0 through (persistence) or synthesises its own member axis (nged_incumbent: 13 analogues;climatology: quantile-derived members). This is model-family identity, likeMODEL_NAME, so aClassVaris correct β and the class is resolved from the experiment'sforecaster_targetMLflow tag before inputs are loaded, so the flag is available in time. Document alongside it thatweather_source: "none"does not mean "no NWP input": in bulk mode the control-member NWP scan defines the shared(init_time, valid_time)forecast-run grid, which is what keeps every leaderboard row β baseline or ML β scored on the identical grid.- Power-lag lookback at feature-engineering load time. Today
_load_engineering_inputsfilters power to[window_start, window_end], and_apply_power_lagreads lags from that same frame β so any lag longer than the elapsed window is null. This is a real existing gap: XGBoost's 336 h lag is null for the first fortnight of every validation window, and the incumbent's 49β55-week lags would be null for all but roughly the last three weeks of the fold (lag targets only re-enter the window fromval_start + 49 weeks). Fix: apower_lookback: timedeltaparameter that widens only the power scan's lower bound (window_start β power_lookback); callers derive it from the experiment'sselected_featuresviaParsedFeatures(max power-lag hours), so it is automatic per-experiment (~2 weeks for XGBoost, ~55 weeks for the incumbent). Safe because the bulk-mode spine is NWP-centric (_join_nwp_bulk_modeleft-joins power onto NWP rows), so the extra early power rows feed only the lag lookup and add no spine rows and no label rows (labels live on spine rows, bounded by the NWPvalid_timefilter); and_nullify_leaky_lags(which nulls any lag withlead β₯ lag_hours) still guards leakage, so a longer lookback cannot leak. Apply to bothtrained_cv_modelandcv_power_forecasts.LIVE_POWER_HISTORYinproduction_assets.pyis a hard-coded 15-day equivalent for the live path β leave it, but cross-reference it from the new parameter so a future long-lag live model is not silently starved. Add a comment wherecv_power_forecastsre-scans the widened power window perinit_timechunk (cheap next to NWP, but worth flagging so nobody blames the wrong thing when profiling). - Document the
ensemble_memberoverload onPowerForecastandAllFeatures: an NWP-member index for NWP-consuming models, a historical-analogue index fornged_incumbent, a quantile-sample index forclimatology. Nobody may assumeensemble_member β NWP. - Tests: a dummy
uses_nwp_ensemble = Falseforecaster exercising the member-0 path; a feature- engineer test that a 336 h lag at the window edge is non-null with lookback while the spine row count is unchanged; the leak test unchanged; the single-run (live) path unaffected bypower_lookback. - After PRs 1 + 2 land back-to-back, run one
trained_cv_model++backfill over every existing experiment partition β retrain, re-predict, and re-score everything under the fixed lag lookback and the new collapse. This is deliberately the exact "re-run everything after a pipeline fix" drill, and it doubles as the empirical verification of the backfill mechanics before the recipe is written intodocs/ml_experimentation/dagster-workflow.md. Treat both PRs as a single leaderboard epoch event, since each shifts existing numbers.
PR 3 β package skeleton + PersistenceForecaster (seasonal-naive; the cheapest end-to-end
probe). Ships the package and proves the PR-2 framework on the simplest model.
- Shared helper
_meta_io.py: themeta.jsonsave/load round-trip (config dump,trained_time_series_ids, the fully-qualifiedmodel_class). NoStatelessForecasterbase class until a third stateless model exists. MODEL_NAME = "persistence",MODEL_VERSION = 1,uses_nwp_ensemble = False. Config defaultselected_features = {"power_lag_24h", "power_lag_48h", "power_lag_168h", "power_lag_336h"}.train()records the sortedtrained_time_series_ids(the requested ids present in the data, mirroring XGBoost's "no usable rows β not trained" semantics).predict()=pl.coalesceof the lag columns in ascending-lag order β_nullify_leaky_lagsnulls any lag withlead β₯ lag_hours(so the 24 h lag survives all intraday leads), and coalesce then selects the shortest non-leaky lag per row: same-time-yesterday for day-1, last-week for day 2β7, two-weeks-ago beyond. Zero lookahead risk because it rides the audited pipeline. Output keeps the spine'sensemble_member(= 0 only, given member-0 inputs), cast UInt8 β Int8 via an expression cast, followingXGBoostForecaster._build_part(a dict-.cast({...})on the concat-of-group-by frames would hit the Patito cast trap). Rows where all lags are null are dropped, the count logged, and the per-series dropped/coverage counts recorded in asset metadata.conf/model/persistence.yaml:weather_source: "none",training_strategy: "none".- Tests: unit (shortest-non-leaky selection on a hand-built
AllFeaturesfixture; all-null-row dropping;save/loadround-trip freezing ids) plus an integration smoke fold via thetests/test_trained_cv_model.pyfixture pattern. - Register (
smoke_testβfull_cv), materialise the chain, sanity-check: NMAE worse than XGBoost overall but plausibly competitive intraday; single-member CRPS = MAE; spread-skill 0; PICP / interval width degenerate (expected and ignorable for a deterministic baseline). - Coverage caveat: persistence's longest lag is 336 h, so leads in
[336 h, ~360 h]drop out and itsall/extended_rangeaggregates cover a shorter lead population than other models'. Fair CRPS is size-comparable so nothing is wrong, but state the caveat where the sanity numbers are read, backed by the per-series coverage counts in asset metadata.
PR 4 β NGEDIncumbentForecaster (nged_incumbent; the deliverable). The faithful replica,
landing on a now-proven rail.
MODEL_NAME = "nged_incumbent",MODEL_VERSION = 1,uses_nwp_ensemble = False,weather_source: "none". Confign_weekly_analogues = 6andannual_week_span = (49, 55)drive the 13 analogue lags (weekly168h Γ {1..6}; annual168h Γ {49..55}= 8232β¦9240 h β all within the feature parser's 17 520 h cap), withselected_featuresderived from them so variants stay one override cheap.predict()unpivots the 13 analogue-lag columns intoensemble_memberrows (member index = analogue index; Int8 holds 0β12). Members nulled by_nullify_leaky_lags(as lead time grows, the short weekly members shed first) or by insufficient history are dropped; rows where all members are null are dropped with the count logged and per-series surviving-member counts recorded in asset metadata. No point forecast is emitted β PR 1's metrics layer produces the median headline and the p95 / p50 labelled rows.- Depends on PR 2's lookback (the 55-week annual lags are all-null without it).
- Data check before interpreting results:
val_start β 55 weeksβ mid-2024. Confirm which eligible series actually have observations that far back β eligibility requires onlymin_training_monthsof history, so a series can qualify yet have too little for the annual analogues, degrading silently to a weekly-only (β€6-member) ensemble. Nothing crashes and fair CRPS stays size-comparable, but the leaderboard means then average differently-shaped ensembles across series, so surface the per-series member counts rather than discovering it later in a dashboard. - Tests: hand-computed unpivot to the expected members; the median, p95, and p50 collapses match
hand-computed values through
compute_metrics; leaky-lag shedding as lead grows;save/loadround-trip; an integration smoke fold with a synthetic power history spanning the annual lags. Sanity-check: the median is a roughly unbiased central estimate whilembe@p95shows a clear positive bias (the conservative operating point β not a bug); no short-horizon skill (the shortest member is a week old, which is realistic). - Ship-time triage: delete this item's details (summary β PR body); cross-link the NGED's incumbent forecast background page.
PR 5 β ClimatologyForecaster (climatology; the pure probabilistic reference). The calendar-
only skill floor the NWP ensemble must clear at long horizons.
MODEL_NAME = "climatology",MODEL_VERSION = 1,uses_nwp_ensemble = False.train(): from the features frame take(time_series_id, valid_time, power)and dedupe on(time_series_id, valid_time)first β bulk-modeAllFeaturesrepeats each target row once per covering(nwp_init_time, ensemble_member)(~15Γ at a 15-day horizon), so without the dedupe the per-cell samples would be weighted by NWP-run coverage rather than by calendar. Then, per(time_series_id, month, half-hour-of-day, is_weekend)cell, store the empirical quantiles of that cell's power samples. Cell keys derive from local (Europe/London) time computed inside the forecaster fromvalid_time, aligning with the demand rhythm (and matching thelocal_*time features).save()writes the lookup as one parquet +meta.json.- Member emission β deliberate deviation from an earlier draft. Emit members at equiprobable
quantile levels
(i β 0.5)/m, not at the tail-heavyDELIVERY_QUANTILESlevels. Fair CRPS and the per-run empirical delivery quantiles the metrics layer derives from members treat members as an equiprobable sample; feeding 13 members at the delivery levels in with equal weight would put 7.7 % of the mass atq(0.01)andq(0.02), i.e. an ensemble materially wider-tailed than the climatology it represents, corrupting CRPS, PICP, and the derived delivery quantiles. The delivery quantiles are still produced β bycompute_metrics' per-run quantile aggregation, same as every other model. Keepm = 13: over the ~14-month training window a weekend cell holds only ~9β17 samples (weekday ~21β43), so quantile levels below ~0.06 are already min-sample extrapolation and 51 members would be false precision plus ~4Γ morepower_forecastsrows; only raise it with empirical justification. predict(): join the lookup onto the prediction rows by the cell keys (rows in an unseen cell dropped with the count logged), then unpivot the quantile columns intoensemble_memberrows.- Why a distribution, not a mean: a deterministic climatology forecast collapses CRPS to MAE, so
it could only be compared against the ML ensemble on point accuracy β the axis a squared-error
XGBoost is built to win. The claim we need to test is distributional: does the NWP-driven ensemble
know more about day 8β14 power than the plain seasonal/time-of-day distribution, or is it dressing
up climatology in weather-shaped clothing? (See Probabilistic forecasting from NWP
ensembles.)
Only a quantile/ensemble climatology answers that, via
crps__all__extended_rangehead-to-head. Caveat to record where that comparison is read: a deterministic-quantile ensemble slightly out-scores an i.i.d. ensemble of the same size under fair CRPS (a known property of Ferro's correction), so climatology carries a small structural edge in exactly this comparison β read a near-tie as "the NWP ensemble is worth little out here", not as climatology winning outright. - Known limitation, documented not solved: small per-cell samples make the tail quantiles noisy; pooling adjacent months is a possible refinement if the numbers look ragged.
- Tests: hand-computed per-cell quantiles; the dedupe (a duplicated spine must not change the
quantiles);
save/loadround-trip; an integration smoke fold; CRPS flows over the members. - Ship-time triage: unblocks
#354 (the dashboard
climatology reference band). As the last π§ baseline item, delete the whole "Implementation
details β baselines" section (summary β PR body), close #147, and update the status banner plus the
milestone section in
docs/roadmap/index.mdif the arc changed.
Recipe confirmed by NGED (July 2026). No open questions remain. Full write-up in NGED's incumbent forecast; the implementation spec:
- Weekly analogues: the last 6 weeks, same weekday & time-of-day.
- Annual analogues: the seven weeks spanning 49β55 weeks back, same weekday & time.
- Deterministic value: NGED's own operating point is the 95th percentile of all 13 analogue values ("more of a vibe") β reported alongside the metric-matched median headline (PR 1).
- No further processing: no weighting, no holiday handling, no anomaly rejection, no load-growth scaling. (This is precisely why the holiday-aligned variant is a genuine, un-done upgrade β not a reimplementation of something NGED already do.)
Cross-cutting. (1) Issue hygiene: create one tracked sub-issue per PR under epic
#6 / #147 following the
CLAUDE.md issue-creation rules (labels, Type, OCF project fields, sub-issue ordering), including
one for nged_incumbent_holiday_aligned so it survives #147 closing. (2) Re-run recipe: add a
short "Re-running CV for an experiment" subsection to docs/ml_experimentation/dagster-workflow.md
describing the trained_cv_model++ backfill, written only after the drill is verified end-to-end,
and mentioning PopulationFilter for single-experiment scoring. (3) MLflow re-logging on a
re-run is idempotent in effect (latest value wins; history accumulates harmlessly).
Cross-fold validation
The cross-validation protocol is implemented, so it has moved to its permanent home: ML Experimentation β Cross-validation folds. That page covers the expanding-window protocol, the current single fold (and why the available weather data constrains us to it), the target multiple-yearly-fold protocol, and the fold-design alternatives we considered.
Fold hygiene: selection bias and a final-test window π§
Issue: #226
The single leaderboard fold (mid_2025_to_mid_2026 in conf/cv/default.yaml: train 2024-04 β
2025-06, validate 2025-07 β 2026-06) serves as both the model-selection set and the
reported skill number. Every hyperparameter choice, feature ablation, and model comparison is
adjudicated on the same 12 months that the leaderboard reports. With hundreds of planned
experiments (the roadmap mentions LLM-driven auto-experimentation in v0.5), the winner's
reported skill will be optimistically biased β classic leaderboard overfitting. The epoch
mechanism handles data changes but not adaptive selection on a fixed fold.
Until the structural fix lands: leaderboard metrics are selection metrics; differences smaller than fold-level noise should not drive decisions; and the number of experiments per epoch is itself a relevant statistic (visible as the MLflow experiment count).
Implementation details β final-test window (deleted when it ships)
1. Document the caveat (immediately). A short "Selection bias" subsection in
docs/ml_experimentation/cross-validation-folds.md restating the paragraph above.
2. Reserve a final-test window (next leaderboard epoch). Found a new epoch in
conf/cv/default.yaml (the epoch mechanism exists for exactly this):
- Shrink the leaderboard fold's validation window to
2025-07-01 β 2026-03-31. - Add a
final_testfold2026-04-01 β 2026-06-30with a new per-fold flagfinal_test: true(extendCvConfig/ the fold schema inpackages/contracts/src/contracts/hydra_schemas.py; it is neither a leaderboard fold nor a dev fold β_fold_ids_for_run_modeindefs/jobs.py:96must not include it in any run mode, so no experiment trains or scores on it in the normal flow). - Scoring against the final-test window is a deliberate, rare act β only for champion
candidates immediately before promotion β via the
metricsasset withevaluation_scope="ad_hoc"and the window'svalid_timebounds in the existingPopulationFilter. No new asset needed; the discipline is procedural. Note: the model trained for the leaderboard fold is reused as-is (train window unchanged), so final-test scoring needs acv_power_forecastsrun over the reserved window β check whether the existing asset can forecast a window disjoint from the fold'sval_start/val_end, and add a window override to its config if not. - Rule, documented alongside: final-test results are never used to choose between candidates (that re-creates the problem); they exist to report honest skill for the chosen champion and to detect gross overfitting (final-test NMAE β« validation NMAE).
A multi-fold gotcha to handle when folds proliferate. The parent-MLflow-run aggregation in
the metrics asset averages each metric key over only the folds in which that key appears
(exp_metrics.setdefault(key, []).append(value) then sum/len in defs/cv_assets.py). Today
every fold emits the same key set, so this is invisible β but folds with different horizon
coverage (one fold's forecasts stop at 36 h, another's reach day 14) would emit different
per-horizon-slice keys, and a key like rmse__all__extended_range would then silently average
over a different fold subset than rmse__all, with nothing marking the smaller denominator.
Per-time_series_type keys have the same property if fold populations differ. When adding
folds (the multi-fold epoch below, or the final_test fold), either guarantee every
leaderboard fold emits an identical key set, or make the parent-run aggregation record its
per-key denominator.
3. Trade-off (decide at implementation time). This costs 3 of the 12 validation months, on a dataset that is already short. The alternative β accepting documented bias until Dynamical.org backfills enable multiple yearly folds β is defensible; if the backfill is expected within a couple of months, do part 1 now and fold part 2 into the multi-fold epoch instead of spending a separate epoch on it. Decide based on the backfill outlook.
Verification. (1) register_experiment_job in all three run modes never creates a
partition for the final_test fold (extend tests/test_register_experiment_job.py).
(2) Eligibility for the leaderboard fold is unchanged by the shrunk validation window (the
eligibility rule keys off val_start/val_end β re-materialise eligible_time_series and
diff). (3) End-to-end: score one existing experiment against the reserved window via the
ad_hoc metrics path and confirm rows land in forecast_metrics.delta with the window label,
and nothing is logged to the leaderboard MLflow runs.
Evaluation metrics
| Metric | Type | Status | Purpose |
|---|---|---|---|
| Mean absolute error (MAE) | Deterministic | β | Typical error magnitude (MW). |
| Normalised MAE (NMAE) | Deterministic | β | MAE normalised by the series' effective capacity (full-history P99) β comparable across substations of different sizes. |
| Root mean squared error (RMSE) | Deterministic | β | Heavily penalises large misses (one 100 MW error costs more than two 50 MW errors). |
| Mean bias error (MBE) | Deterministic | β | Systematic over/under-prediction. |
| Histogram of errors | Deterministic | π§ | Visual check that errors are ~Normal. |
| Pinball loss (quantile loss) | Quantile | β | Penalises asymmetrically by target quantile, at the 13 NGED delivery quantiles. Averaged across quantiles for a single quantile-skill score. |
| PICP (Prediction Interval Coverage Probability) | Quantile | β | Coverage of six symmetric bands (p1βp99 β¦ p35βp65). Judge against the finite-ensemble calibrated reference (β 0.769 for p10βp90 at 51 members), not the nominal rate. |
| Interval width | Quantile | β | Mean band width (MW) β the sharpness companion that stops PICP being gamed by over-widening. |
| CRPS (Continuous Ranked Probability Score) | Ensemble | β | Probabilistic equivalent of MAE; rewards both accuracy and sharpness. Fair (finite-ensemble-unbiased) form β the one metric comparable across ensemble sizes. |
| Spread-Skill Ratio | Ensemble | β | Fortin-corrected RMS ensemble spread Γ· RMSE of the ensemble mean. 1.0 = well-calibrated; < 1 under-dispersed (overconfident); > 1 over-dispersed (underconfident). |
| Threshold-weighted CRPS (twCRPS) | Ensemble | π§ | CRPS confined to behaviour above a per-series threshold β the ranking-grade metric for skill near the network limit. See Tail & exceedance metrics. |
| Exceedance rate of upper quantiles | Quantile | π§ | "When we said p95, was it exceeded ~5% of the time?" β one-sided calibration check for the delivered tail quantiles. |
| Brier score for threshold exceedance | Ensemble | π§ | Grades the ensemble's "chance load exceeds the threshold" probability β the score for the yes/no warning NGED acts on. |
The
Metricsschema (contracts.ml_schemas.Metrics) stores results as(time_series_id, power_fcst_model_name, fold_id, horizon_slice, metric_name, metric_param, metric_value).metric_paramcarries, e.g., the quantile for Pinball Loss (p10) or the band for PICP (p10_p90). ThemetricsDagster asset computes every β metric above and writes per-series rows toforecast_metricsDelta (partitioned byexperiment_name, fold_id), with per-fold and mean-across-folds aggregates logged to MLflow β see Running an ML experiment end-to-end.
Every metric above is also broken out by named population-filter slices, which appear as extra leaderboard columns. Most are legitimate ranking columns, because the filter is fixed by information known before the outcome: the horizon slice a forecast falls in, the Tricky days calendar, and logged switching events. Two are diagnostic only and must never drive ranking, because they select hours by what actually happened and so reward a model that simply over-forecasts (the forecaster's-dilemma trap): the peak-events slice (the top 5% highest observed demand) and NGED's hand-picked hard examples. Both are described under Tail & exceedance metrics.
Which ensemble collapse defines the deterministic point forecast? π§
Decided: metric-matched collapse, uniform across every model. Implemented as part of the baseline work (PR 1 in the baseline implementation details); the reasoning below is the durable design rationale, promoted to the evaluation-metrics reference when that PR ships.
The problem
The deterministic metrics (MAE, NMAE, RMSE, MBE) score a single point forecast, but every model
on the leaderboard is really an ensemble (51 NWP members for the ML models; 13 historical analogues
for nged_incumbent; a quantile sample for climatology). Something has to collapse each ensemble
to one number. compute_metrics today collapses every ensemble to its mean
(packages/ml_core/src/ml_core/metrics.py), and the risk we were guarding against was that
different models would be scored on different collapses β e.g. the ML models on their mean and
nged_incumbent on the median NGED effectively reads off its analogue spread. Mean and median
diverge for skewed or underdispersed ensembles, so scoring some models on one and some on the other
is not apples-to-apples β a silent trap that quietly mis-ranks models.
Mean versus median β the trade-off
The instinct is to "pick one central statistic and apply it everywhere". But mean and median are not interchangeable, and the reason is the crux of the decision:
- The mean is the point forecast that minimises squared error β so RMSE (and MBE, which is
a mean-of-errors and inherits the mean's clean energy-balance / expectations-aggregate reading)
is consistent with the mean. Our squared-error XGBoost models literally learn a conditional
mean, and the ensemble mean of 51 such members is a coherent estimate of
E[power]. Scoring the mean on RMSE rewards a model for doing the statistically correct thing; the mean also averages out member noise, so it is the more stable statistic on a small ensemble, and it matches standard NWP verification practice. - The median is the point forecast that minimises absolute error β so MAE and NMAE are
consistent with the median. It is robust to the skew that is real in this problem (holiday
weeks, solar clipping) and to the ensemble underdispersion the probabilistic
section documents, it is the faithful reading of NGED's
equally-weighted analogue spread, and it is coherent with the quantile columns already on the
leaderboard: median MAE is exactly
2 Γ pinball_loss@p50, so a median headline makes the deterministic and probabilistic columns tell one story.
The key realisation is that MAE and RMSE elicit different functionals, so no single collapse is fair on both columns. A uniform median is inconsistent for RMSE (it penalises squared-error models for forecasting their honest mean); a uniform mean is inconsistent for MAE (a model could improve its MAE ranking by warping its forecast away from its honest mean β Gneiting 2011, Making and Evaluating Point Forecasts). The apples-to-apples requirement is only that the collapse be uniform across models, not across metrics.
The decision
Score each deterministic metric on the ensemble collapse it is consistent with, the same way for every model:
- MAE / NMAE β ensemble median (NMAE is the headline cross-series metric, so the headline is effectively median-based).
- RMSE / MBE β ensemble mean.
- Spread-skill ratio β ensemble mean, internally, unchanged β its Fortin calibration target
(RMSE of the ensemble mean =
β((m+1)/m) ΓRMS spread, so "1.0 = calibrated") is defined against the mean; switching its internal collapse would silently break that reading.
Everything else is an extra, labelled row, never a headline: NGED's P95 operating point
(mae/mbe at metric_param="p95" β conservative by design, so a large positive MBE that belongs
beside the central number, not in place of it), and the median's own bias (mbe at
metric_param="p50", so the delivered central forecast has an honest bias number distinct from the
mean's energy-balance bias). No model is structurally disadvantaged on any column β which a single
uniform statistic cannot achieve β and the trap is closed because the collapse is uniform across
models. Because the collapse lives entirely downstream of the stored forecasts, switching it
re-scores the whole leaderboard from a single metrics re-materialisation with no retraining or
re-prediction β and doing it now, before the leaderboard adjudicates anything, is the cheapest moment
to shift every existing deterministic number.
Normalising NMAE by effective_capacity
NMAE is MAE divided by a per-series effective capacity, not by the mean or a per-fold P99. A capacity-like denominator is what makes NMAE comparable across asset types: intermittent generators (PV, wind) spend much of their time near zero output, so normalising by the mean would inflate their NMAE relative to a demand substation of similar peak size. Computing the denominator over each series' full history (rather than within the validation window) also keeps it stable across folds β an unusually calm year for a wind farm would otherwise give a low in-window P99 and an inflated NMAE.
Normalisation is also why NMAE is the headline cross-series metric: the aggregate mae__all /
rmse__all values logged to MLflow are unweighted means across series whose scales span roughly
two orders of magnitude, so the GSPs dominate them. They are useful for tracking a single model
over time, not for comparing skill across the population.
The denominator comes from the effective_capacity
Delta table (schema contracts.power_schemas.EffectiveCapacity), consumed by compute_metrics
(ml_core.metrics).
v0.1 representation: one scalar row per series. The effective_capacity asset writes one
row per time_series_id β effective_capacity_mw = P99 of |power| over the whole observation
history, time = the latest observed timestep. compute_metrics joins it onto the per-series
metrics on time_series_id alone and divides.
Why v0.1 is a single row per series, not the value repeated at every half-hour. The
v0.7 upgrade below will store one row per (time_series_id, time) half-hour β but with a
genuinely time-varying value. In v0.1 the value is a single constant per series, so repeating it
across every half-hour would just be a denormalised encoding of one number: at V2 scale (~2,500
series Γ ~4 years Γ 17,520 half-hours/yr β 175M rows) that is hundreds of millions of rows to
express ~2,500 scalars, for zero extra information. It would also not buy forward-compatibility,
because the real v0.1βv0.7 interface change is not the data shape but the join (below). The
EffectiveCapacity schema β (time_series_id, time, effective_capacity_mw) β already accommodates
both the one-row-per-series v0.1 shape and the one-row-per-half-hour v0.7 shape; that is the
forward-compatibility we want. (The v0.7 upgrade does widen the columns β the value becomes a
mean + std pair,
#247 β but the row shape
and the join are unaffected by that.)
v0.7 upgrade: time-varying, and the join changes. The differentiable-physics capacity model produces a value that changes over time (panel degradation, inverter trips, seasonal derating). At that point two things change, and nothing else:
- the
effective_capacityasset body emits one row per(time_series_id, time); and compute_metricschanges its capacity join fromtime_series_id-only to a temporal as-of join on(time_series_id, valid_time)β matching each forecast'svalid_timeto the capacity in effect at that time.
The Metrics schema and the rest of the metrics pipeline are untouched. Note the table is
backward-looking only (it holds no future valid_times): fine for historical CV folds (whose
validation windows lie inside the observed history), but live-forecast scoring
(production monitoring) must
choose which reference time's capacity to apply rather than expecting a row at a future valid_time.
One related distinction to keep straight: the metric denominator may use the full-history smoothed capacity estimate, but any capacity used to normalise model inputs at forecast init time (the two-pass training scheme) must be the causal estimate available at that init time, or backtests gain lookahead β see Capacity estimation.
Tail & exceedance metrics β scoring the question NGED actually asks π§
Issue: #254
NGED's goal is flexibility procurement: paying flexible customers to reduce demand when a substation risks running beyond its capability. The question they ask of a forecast is therefore "will load cross the limit?" β their operator tool plots demand as headroom below a constraint line (see NGED's incumbent forecast), and the requirements name threshold exceedance as the core concept. The existing metrics lean toward the tails (the mean pinball loss is deliberately tail-weighted; PICP covers the p1βp99 band) but they measure upper-quantile skill at every hour equally β including hours when nothing is anywhere near a limit. This section adds the metrics that measure skill near the limit itself.
One design rule governs the whole section: never rank models on a sample selected by what actually happened. Scoring models only on the top-N% highest observed half-hours rewards models that simply bias every forecast upward β the "forecaster's dilemma", explained in plain language in the evaluation-metrics reference. Ranking-grade tail emphasis instead comes from proper scores re-weighted toward the tail and computed over all hours. Concretely, three metrics β definitions, equations, and intuitive explanations in the evaluation-metrics reference:
-
Threshold-weighted CRPS (twCRPS) β CRPS confined to behaviour above a per-series threshold; the headline ranking metric for tail skill. Implementation is nearly free: replace members and observation by
max(value, threshold)and reuse the existing fair-CRPS aggregation unchanged, inheriting its comparability across ensemble sizes. -
Exceedance rate of the upper delivery quantiles (p80βp99) β "when we said p95, was it exceeded ~5% of the time?"; the plain-language calibration check for the tail quantiles NGED reads, one-sided where PICP's bands are symmetric.
-
Brier score for threshold exceedance β grades the ensemble's "chance that load exceeds the threshold" the way one would grade a "70% chance of rain" forecast; the most decision-legible number, scoring exactly the warning NGED acts on.
All three fit the existing Metrics schema β metric_param carries the threshold or quantile
label β so there is no schema change.
Thresholds: static, per-series, quantile-derived. A substation's true limit is not a single number (ratings vary with ambient temperature; transformers tolerate short overloads, so exceedance duration matters; switching changes what a feeder carries β NGED's own limit line is a time-varying "Flex Profile"). We deliberately do not model any of that for scoring. Each series gets two static thresholds β the P90 and P98 of its full observation history, in the series type's constraint-side direction (high load for demand; reverse power flow for generation) β chosen because they guarantee every series a scoreable event rate, mean the same thing across series, and stay stable across CV folds. Physical firm/flex ratings, where NGED supplies them, feed ad-hoc case studies and dashboard overlays instead: a rating that is never breached in the validation window yields zero events, and a warning system cannot be graded on events that never happen. The full rationale, and the explicit "this is a proxy" caveat, live in the threshold-choice section of the reference.
Diagnostic slices β never ranking columns. Two observed-outcome views stay available for understanding a model's errors, clearly labelled as ineligible for ranking (the design rule above): a peak-events slice (the top 5% highest-observed-demand half-hours, via the same population-filter mechanism as the Tricky-days filter) and hand-picked "hard examples" β historically difficult periods NGED supplies, scored as case studies.
These metrics also serve the probabilistic-calibration work: an underdispersed ensemble pushes its exceedance probabilities to 0 or 1 too early, so the Brier score and exceedance rates are the sharpest before/after instruments for Phases C and D.
Implementation details β tail & exceedance metrics (deleted when this ships)
- Thresholds: compute per-series
hist_p90/hist_p98scalars from the full observation history alongside (or within) theeffective_capacityasset β same full-history stability rationale, same join shape (time_series_id-only). Constraint-side direction resolved pertime_series_type; confirm the mapping with NGED for ambiguous types (BESS charges and discharges). - twCRPS: transform members and observation with
pl.max_horizontal(col, threshold)and reuse the existing fair-CRPS expression (sorted-member identity, Float64 accumulation) verbatim, once per threshold rung. - Exceedance rates: compare
yagainst the already-computed empirical quantile columns for p80, p90, p95, p98, p99 β one boolean mean per level. - Brier score: exceedance probability = member fraction above the threshold; outcome
indicator from
y; squared difference, averaged. - MLflow allowlist: extend
_MLFLOW_LOGGED_PARAMETRICwith a small headline subset (e.g.twcrps@hist_p98, exceedance rate at p95,brier@hist_p98); decide the exact set at implementation time and keep it small β everything is in Delta regardless. - Peak-events diagnostic slice: one more named population filter on the shared mechanism (with the Tricky-days filter), flagged in the leaderboard UI as diagnostic-only.
- Verification: hand-computed toy-ensemble values for all three metrics; cross-check
twCRPS against the
scoringrulesreference implementation; Monte-Carlo the finite-ensemble one-sided exceedance references (mirroring the PICP reference-table verification); on a smoke fold, confirmnged_incumbent's P95 operating point scores well on the p95 exceedance rate while paying for its conservatism on Brier/twCRPS sharpness.
Tricky days β a calendar-deterministic metric filter π§
Issue: #255
Alongside the tail & exceedance metrics, we add a "Tricky days" filter: score models separately on the handful of calendar dates whose demand shape departs sharply from the usual weekly rhythm. We scope it deliberately to the calendar-deterministic set β fixed and moveable public holidays (Christmas, Easter, and the rest of the GB bank holidays) plus the two annual daylight-saving transitions β because these are exactly the days our weekday/seasonal analogues are built to mishandle. Days that are hard for data-dependent reasons stay in their own already-planned filters: switching-event days and NGED's hand-picked "hard examples" (above). One shared filter mechanism, several named filters β folding genuinely different failure modes into one bucket would make the number impossible to act on ("bad on tricky days" β is that Christmas, or a switching event?).
Mechanically it is another population filter (the same mechanism the peak-events diagnostic
slice uses): a boolean flag per timestep, derived from valid_time alone. Unlike that
observed-peak slice, this one is a legitimate ranking column: the flag depends only on the
calendar, which every forecaster knew in advance, so it does not fall into the
forecaster's-dilemma
trap. Because it is purely
calendar-driven it shares its calendar module with nged_incumbent_holiday_aligned β the same
GB bank-holiday calendar (the pure-Python holidays package) plus the two DST dates feed both the
holiday-aligned baseline and this metric filter. And the two reinforce each other: nged_incumbent
(no holiday logic) should be visibly worst on tricky days, and nged_incumbent_holiday_aligned
should recover most of the gap β turning "we added holiday alignment" into a measurable number,
exactly the cheap-upgrades story we want to show NGED.
Flag the day and its analogue-relevant neighbours, not just the day itself. The disruption spills onto surrounding timesteps:
- DST transitions: the hard part is not only the 23/25-hour day but that lag and analogue features are misaligned by an hour on the days either side.
- The Christmas run-up: demand in the days before Christmas is already atypical, so the window must cover the run-up, not just the 25th.
So the flag covers a small window around each date rather than a single day; the exact per-event widths are an implementation-time choice.
A subtlety to document now but not model yet. The shape of the Christmas run-up depends not just on the number of days before Christmas but also on which weekday Christmas falls on β the run-up demand pattern shifts year to year with that day-of-week alignment. We record it here as a known effect; the v0.3 tricky-days flag ignores it (it simply marks the window), and we defer any explicit day-of-week-aware modelling of the run-up until there is evidence it moves the leaderboard.
Implementation details β tricky days (deleted when this ships)
- A small calendar module (shared with baseline 2,
nged_incumbent_holiday_aligned) answers, for anyvalid_time, whether it falls inside a tricky-days window. Back it with theholidaysGB calendar plus the two annual DST dates; expose the per-event window widths as config. - Represent the tricky-days slice the same way the peak-events diagnostic slice is represented β
one more named population filter, resolved by the same mechanism, not a new schema axis β
so the leaderboard gains a Tricky days column with no
Metricsschema change. - Verification: unit-test the flag on known dates (a Christmas week, an Easter, both DST
switchovers, and a plain week that must be excluded); on a smoke-test fold, confirm
nged_incumbentscores worse on the tricky-days slice than overall.
Time-slices for performance evaluation
We compute every metric separately per horizon slice, because the driver of model skill changes with lead time:
| Horizon slice | Industry term | Primary driver of model skill |
|---|---|---|
| 0β6 hours | Intraday / Nowcasting | Lagged power & persistence. NWP is often too coarse to beat simple autoregressive features here. |
| 6β36 hours | Day-Ahead | Deterministic NWP. Covers the critical day-ahead market gate; relies on the diurnal cycle + high-res weather. |
| Day 2β7 | Short/Medium Range | Synoptic weather. Skill driven by mapping large weather fronts to power; ensemble spread starts to matter. |
| Day 8β14 | Extended Range | Ensemble probabilities. Deterministic weather is essentially noise; skill comes from processing ensemble uncertainty. |
Measuring performance during switching events π§
We will flag each timestep for whether it contains a switching event, and compute metrics separately
for periods with switching events in the model inputs (or in the forecast's valid_time). This
distinguishes models that perform well only on clean periods from models that handle switching
events in their inputs. In v1 the flags can come straight from NGED's logged switching events
for the trial area; a fleet-scale version would need the discrete detector described in
Switching events & latent demand, whose fate is an open question β see
the decision point.
Delivering the probabilistic metrics π§
Issue: #225
The 51-member ensemble is very likely underdispersed: XGBoost trained with squared error on
the control member (cv_assets.py:373) learns a conditional mean, so pushing 51 members through
it yields spread from weather uncertainty only β no model or observation uncertainty. Such
ensembles are systematically overconfident, worst at short horizons where members haven't
diverged. Flexibility procurement is a tails problem (P90+ peaks), so this hits the use case
directly. Phases A and B below (both shipped) built the measurement machinery β every metric in
the evaluation-metrics reference, per horizon slice; the
remaining phases act on what those numbers show. The planned
tail & exceedance metrics
will sharpen the picture further: an underdispersed ensemble pushes its exceedance
probabilities to 0 or 1 too early, so the Brier score and the quantile exceedance rates are
the clearest before/after instruments for Phases C and D.
The theory behind this diagnosis β the three-term uncertainty decomposition, why a deterministic model driven by an NWP ensemble captures only the weather term, and what the principled fix looks like β is the durable explainer Probabilistic forecasting from NWP ensembles. This section is the plan that applies it.
Fix in four phases, each an independent PR (Phase D is itself several PRs). Phases A and B are pure evaluation (no model changes) and should land before any further MAE-driven experimentation.
Phase A β horizon-sliced metrics β
Shipped. compute_metrics now scores every metric per HORIZON_SLICES band (derived from
valid_time β power_fcst_init_time, with the ensemble collapsed per forecast run), and
build_mlflow_aggregate_metrics logs overall per-slice keys like nmae__all__day_ahead.
Phase B β probabilistic metrics from the existing ensemble β
Shipped. compute_metrics now computes fair CRPS, the Fortin-corrected spread-skill ratio,
pinball loss at the thirteen NGED delivery quantiles (plus their mean), and PICP and interval
width for the six symmetric delivery bands β all member-aware, computed before the
ensemble-mean collapse, per horizon slice. Definitions, equations, and the design decisions
(fair CRPS divisor, RMS spread, delivery quantiles, MLflow allowlist) live in the
evaluation-metrics reference.
Phase C β cheap calibration (after B proves the diagnosis)
Decide based on Phase B's spread-skill numbers β and on the rank (Talagrand) histogram of
the observations among the 51 members, computed ad hoc per horizon slice: a U-shape confirms
plain underdispersion (a single multiplicative inflation can fix it), whereas a sloped or
asymmetric histogram means bias or shape error, which a symmetric inflation cannot repair and
which would push toward a rank-dependent correction instead. Post-hoc spread inflation
(EMOS-lite): per
horizon slice, fit a scalar s on the training window so that inflating members around the
ensemble mean (mean + sΒ·(member β mean)) makes spread match error. Zero schema change, zero
new model β implementable as an optional step in predict or as a wrapper forecaster. Fit on
train, apply on validation (no tuning on the fold being scored).
Spread inflation widens the fan but cannot reshape it (the inflated ensemble is still 51 point forecasts, just pushed apart). It is the stopgap the full fix below must beat to earn its build cost.
Phase D β ensemble of quantile forecasts (Representation 3 β pooled Representation 2)
The full fix from the probabilistic-forecasting explainer: have the model emit a conditional distribution per ensemble member (per-member quantiles β Representation 3 in the delivery tables), then recombine the 51 members with the linear-pool mixture into one set of delivered percentiles (Representation 2). Three steps, one PR each:
- Percentile representations in
PowerForecast(#262) β extend the contract (and the Delta write/read paths) with the Rep 2 and Rep 3 percentile columns, alongside the existing deterministic-ensemble representation. - Quantile XGBoost model family
(#263) β
objective: reg:quantileerrorwith severalquantile_alphas, as a separate experiment/model family emitting Rep 3. Sort each member's quantiles at predict time (monotonic rearrangement fixes quantile crossing). The lead-time feature and training on multiple members (xgboost-improvements β the lead-time feature and ensemble-member training) are the double-counting mitigations discussed in the explainer β land them first or measure without them consciously. - Linear-pool combining step (#264) β pool the per-member quantiles into delivered Rep 2 percentiles (the pseudo-sample recipe in the explainer), with a per-horizon affine recalibration hook (fit on train) applied only if the pooled spread-skill/PICP numbers demand it. Scored with pinball/PICP/CRPS head-to-head against the Phase-C inflated deterministic champion.
Grouping the results
Each ML experiment is tagged with metadata so we can group experiments and compute average performance per group (e.g. "does lagged power always help, regardless of model sophistication?", or "how robust is each model to weather-forecast uncertainty β ERA5 reanalysis vs. operational NWP?"). Example tags:
| Tag | Example values |
|---|---|
time_series_type |
PV, Wind, disaggregated demand (primaries) |
model_family |
nged_incumbent, baseline_persistence, xgboost, pytorch_mlp, pytorch_graph_dp |
weather_source |
none, ecmwf_control, full_ecmwf_ensemble, era5 |
input_features |
datetime, power_lag_24h, power_lag_7d, temperature |
training_strategy |
direct_multistep, horizon_as_feature, end_to_end |
generator_capacity_estimation |
none, simple_p99, convex_envelope, differentiable_physics |
switching_event_detection |
none, simple_statistical |
pre_training |
none, ERA5 |
Estimating cost savings (Β£) attributable to each forecasting approach, per leaderboard row, is a π¬ v2 stretch goal.