Skip to content

NGED Flexpectation

NGED Flexpectation is an NIA-funded project (Network Innovation Allowance project reference NGED_NIA_085) by Open Climate Fix to deliver state-of-the-art, probabilistic power forecasts for National Grid Electricity Distribution (NGED). The forecasts help NGED optimise flexibility procurement and manage electricity network congestion. NGED describe the project on their Flexpectation project page.

Example power forecast

What the forecasts look like

Each forecast is:

  • Probabilistic — expressed as an ensemble of 51 members, one per ECMWF ENS member
  • 14-day horizon, half-hourly temporal resolution
  • Refreshed every 6 hours
  • In MW (active power) or MVA (apparent power) — the unit is given per time_series_id in TimeSeriesMetadata
  • Sign convention depends on substation_type — see Sign convention

Scope

Version 1 (current focus): 32 time series in NGED's trial area — 16 primary substations, 6 solar PV farms, 3 wind farms, 2 grid supply points (GSPs), 2 bulk supply points (BSPs), 1 biofuel generator, 1 battery energy storage system (BESS), and 1 reciprocating gas generator.

Version 2 (future): Scale to approximately 2,500 time series covering all 1,161 of NGED's primary substations, NGED's bulk supply points and grid supply points, and most customer meters.

After the NIA project: the working assumption is that NGED runs the service on its own AWS account. The service is therefore being built to be operable day to day by NGED staff who did not develop the code, working from the runbooks. See Requirements → Operating model & handover and the Handover to NGED design page.

More than a forecast

A large part of this project is building a production forecasting system and researching novel forecasting methods. But NGED's interest goes beyond the forecasts themselves: they also want information — to learn which forecasting approaches actually work well on their data (a major reason we invest in a rigorous leaderboard), and to understand the underlying issues involved in forecasting their electricity network.

This interest in information means a negative result can be just as valuable as a positive result. For example, if we try hard to detect switching events unsupervised and conclude that detection isn't reliably possible from power readings alone, that conclusion is a useful finding in its own right. That finding tells NGED whether extracting switching-event labels from operational systems would help.

The same logic applies to the engineering: our claims about it are written down as falsifiable engineering hypotheses with thresholds attached, and the Design Philosophy section explains why — including which industries we borrow practice from and what we declined.

Documentation

Want to run NGED Flexpectation on your laptop? Start with Getting started — a single walkthrough from a fresh clone to a running Dagster instance that downloads data and trains a model.

  • Design Philosophy — the portable why: the design principles, the falsifiable engineering hypotheses that score them, and the inherent-stability argument in full
  • Background — NGED's electricity network, project requirements, and data quality challenges
  • Techniques — durable explainers of the solution methods: differentiable physics, convex optimisation, encoders, probabilistic forecasting, and evaluation metrics
  • Architecture Overview — what is actually built: technical components and data flow
  • Performance and Scale — the measured performance engineering: storage formats, lazy evaluation, memory bounds, and Polars' row-index ceiling
  • Code Style — code conventions
  • Testing — how the test suite is wired, the house style, and the notable test suites
  • ML Experimentation — methodology for our implemented ML experimentation: cross-validation folds, the leaderboard, and how we evaluate models
  • Live Service — operating the live, 6-hourly production service: promoting a champion model and backfilling missed runs
  • Roadmap — planned future work, plus detailed design docs for the delivery tables, forecast building blocks, metrics & leaderboard, data sources, differentiable physics, switching events, disaggregation evaluation, and encoders

How these docs were written

The ideas, the decisions, and the judgement calls in this documentation are human — they come from the team's own engineering and from reading what other industries do. Much of the prose, though, was drafted and refined with an LLM coding agent (Claude Code) over many hours of back-and-forth. Our experience is that the back-and-forth genuinely improved the writing: an argument that survives being questioned repeatedly tends to end up better evidenced than an argument written in a single pass.

The division of labour matters most for the evidential claims. The performance, size, and cost figures were measured on real data through the real code path rather than estimated. The measure; do not assume principle applies to the documentation as much as to the pipeline. Claims about what the code does are checked against the code. But we will not pretend that every sentence across this many pages has had a human's eye on it next to the source. Where the docs and the code disagree, the code is right, and we would rather hear about the disagreement than have it stand.

New to this repo? See the Documentation Guide for how these sections relate to each other and to GitHub issues — including the rule that roadmap/ holds only not-yet-implemented design, moving out to a permanent home (architecture/ for design rationale, ml_experimentation//live_service/ for step-by-step how-to) once a feature ships.