Dynamical Data API
Dynamical Data Package
Download & process numerical weather predictions from Dynamical.org.
We convert the European Centre for Medium-Range Weather Forecasts' ensemble forecast (ECMWF ENS) from its 0.25 degree latitude/longitude grid to these H3 resolution 5 hexagons. H3 tiles the globe in hexagons at nested resolutions, and resolution 5 is the cell size drawn below:

Note: The generic geospatial logic for mapping latitude/longitude grids to H3 hexagons lives in the
packages/geopackage. This package (dynamical_data) focuses specifically on the ingestion, processing, and storage of time-varying NWP datasets like ECMWF. The H3 grid weights arrive as a Dagster asset from thegeopackage, so no precomputed static file is needed.
Data storage experiments
The storage format itself lives in delta_store.nwp (writer properties, sort order, precision);
this section records the measurements behind it. Full before/after detail is in
PR #271; earlier experiments
(UInt8/Int16 affine quantisation, codec and sort-order sweeps) are in this file's git history. The
same measure-before-you-optimise approach, and the storage numbers for the NWP table and the
power_forecasts table side by side, are summarised in Performance and
Scale.
That page frames both tables' storage choices as an instance of design principle 12, measure; do
not
assume.
Current scheme: physical-unit Float32, every continuous variable rounded to a 13-bit
significand (max relative error 2⁻¹³ ≈ 1.2×10⁻⁴ — measured ≤ 0.004 °C for temperature, ≤ 8 Pa for
mean-sea-level (MSL) pressure), rows sorted init_time → ensemble_member → valid_time → h3_index,
plain ZSTD level 3. init_time is when the weather model was run; valid_time is the moment being
forecast.
Why round at all, when Dynamical.org already rounds? A significand is one bit wider than a
mantissa, because the leading 1 is implicit, so 6–11 mantissa bits are 7–12 significand bits.
Dynamical stores ECMWF ENS with 6–11 mantissa bits per variable (their
binary_rounding.py).
But trailing zeros do not survive arithmetic: our H3 aggregation is a weighted mean over grid
points, and ECMWF publishes wind as an eastward component u and a northward component v, from
which wind speed and direction are derived via sqrt/arctan2, so by the time values reach our
writer their mantissas are full entropy again (measured: 100% of Dynamical-style rounded values have
zeroed low bits; after a weighted mean, 0.07% do). Our 13-significand-bit rounding restores
compressibility while being 1–6 bits finer than the upstream precision, so it discards almost
nothing beyond what Dynamical already dropped.
How much space does Great-Britain-wide ECMWF ENS take? A lead time is one forecast step ahead of
init_time. One daily run (1,671 H3 cells × 51 ensemble members × 85 lead times, up to ~7.24M rows)
averages ~158 MB, so a year is ~58 GB. The full local development table — 899 daily runs (Apr
2024 → Sep 2026, ~6.5 billion rows) — is 142 GB.
Storage — the table below compares the writer configurations against each other, on nine real
partitions spread across every season. The table's absolute figures are older than the NWP table
on disk today. Every row was measured at parquet's default row-group size, against a partition set
averaging 112.9 MB, where a partition now averages ~158 MB. The member-aligned row-group size that
delta_store.nwp writes accounts for 8.6% of that difference, measured across the rewrite of all
899 partitions; the remaining 91.4% predates the member-aligned row-group size and is not accounted
for here:
| Config | avg MB/partition | extrapolated GB/yr |
|---|---|---|
| Previous: Int16 12-bit quantisation, ZSTD-14 | 115.3 | 42.1 |
| Float32 + 13-bit significand, ZSTD-3 | 110.6 | 40.4 |
| Adopted: same + member-early sort | 112.9 | 41.2 |
Same + BYTE_STREAM_SPLIT |
133.8 | 48.8 |
BYTE_STREAM_SPLIT makes this table worse (unlike power_forecasts, where it wins): significand
rounding collapses NWP values into repeats that parquet's default dictionary+RLE encoding captures
directly, and BYTE_STREAM_SPLIT scatters that repetition across four byte planes. Writer
properties are data-dependent — measure per table.
Read path — the member-early sort puts each ensemble member's rows in one contiguous block, and
delta_store.nwp sizes each parquet row group to hold exactly one member. Of the 51 members, member
0 is the unperturbed control run and the other 50 are perturbed. A single-member read (every
training run reads just the control member) therefore matches one row group's min/max range and
skips the other 50 row groups.
The read-path measurements below were taken on their own set of 29 consecutive daily partitions, not
on the nine seasonal partitions the storage table above used. Measured against a valid_time-first
sort of the same 29 daily partitions, reading nine H3 cells and the control member alone: 5.7×
faster and 5.5× less peak memory (170 ms / 2,200 MB → 30 ms / 400 MB), for 3.7% more stored
bytes (4.35 GB → 4.51 GB across the 29 partitions). Both tables were written freshly through
write_nwp, differ only in the sort order, and use the member-aligned row-group size, so the
comparison isolates the row order. Each figure is the warm-cache median of five timed repetitions.
The read decodes 1.96% of each partition censused — one row group in 51 — and that 1.96% holds for
every member. A census of partitions from 2024, 2025, and 2026 found 51 row groups in each. Every
row group spanned a single member, and the 51 together covered members 0 to 50, so no member decodes
extra rows for sitting in the middle of the range. Under the valid_time-first sort the same read
decodes 100%.
The timing and the peak-memory figures were both measured on local disk. On S3 a skipped row group also skips an HTTP range request over the internet, so both figures are a floor: the same read from S3 stands to gain more from row-group skipping.
# Two arms, one partition window, differing only in NWP_SORT_COLS; the second monkeypatches
# delta_store.nwp.NWP_SORT_COLS to ("init_time", "valid_time", "ensemble_member", "h3_index").
# Time each arm in its own process: ru_maxrss is a high-water mark that never falls, so two arms
# in one process report the larger figure twice.
frame = (
Nwp.scan_delta(table)
.filter(
pl.col("init_time").is_between(window_start, window_end),
pl.col("valid_time").is_between(window_start, window_end + timedelta(days=10)),
pl.col("h3_index").is_in(cells), # the 9 cells the 32 V1 series sit in
pl.col("ensemble_member").is_in([0]), # is_in, not ==, as load_engineering_inputs does
)
.collect(engine="streaming")
)
Per-variable keep_bits: considered and rejected (2026-07). Since Dynamical's upstream precision caps the real information at 7–12 significand bits per variable, budgets matched to upstream would compress better than the uniform 13. Measured on four seasonal partitions through the exact production write path (error columns = error added on top of today's stored values; wind speed's power impact is ~3× its relative error because the speed-to-power curve is roughly cubic):
| Config | Size vs today | GB/yr | wind speed rel err | → power (×3) | temp | MSL |
|---|---|---|---|---|---|---|
| Uniform 13 (today) | — | 41.1 | 0 | 0 | 0 | 0 |
| Upstream-matched (temp 8, wind 7, pressure 12, flux 8) | −19.4% | 33.1 | 0.76% | 2.3% | 0.06 °C | 16 Pa |
| Uniform 10 | −15.6% | 34.7 | 0.10% | 0.29% | 0.016 °C | 64 Pa |
| Wind-protected (wind speed stays 13, rest squeezed) | −13.0% | 35.7 | 0 | 0 | 0.06 °C | 16 Pa |
These figures share the Storage table's older ~41 GB/yr basis, so the percentage columns are comparable against each other but the absolute GB/yr sit below today's ~58 GB/yr.
The full squeeze adds ~2.3% power-equivalent wind error — not tolerable. The wind-safe ceiling is −13% ≈ 5.4 GB/yr, and the NWP table covers the whole of Great Britain, so it does not grow with the V2 scale-up to ~2,500 time series: a fixed ~5 GB/yr saving doesn't justify maintaining a dict of per-variable precision budgets. If disk ever becomes a real constraint, the wind-protected config is the one to reach for.
dynamical_data.ecmwf_ens.download
Opening and downloading one ECMWF ENS run from the Dynamical.org catalog.
Attributes
ECMWF_ENS_INSTANTANEOUS_VARS = frozenset(_ECMWF_ENS_VARS_TO_DOWNLOAD) - Nwp.deaccumulated_var_names - Nwp.categorical_var_names
module-attribute
The downloaded variables describing conditions at one instant, under their download names.
None of these is ever legitimately null, anywhere in a run. That is why
dynamical_data.ecmwf_ens.upstream_nulls.assess_upstream_grid_point_nulls counts them separately
from the de-accumulated variables. It is also why the ecmwf_ens asset gates a check on that
count being zero.
Derived from the download list rather than from Nwp's fields, because the two namespaces differ.
We download wind_u_10m/wind_v_10m (and the 100 m pair), and
dynamical_data.ecmwf_ens.convert_to_polars.convert_nwp_xarray_dataset_to_polars_dataframe
derives wind_speed_*/wind_direction_* from them. A set taken from the contract would name
four variables the downloaded dataset does not carry, and indexing it would raise KeyError.
Classes
NwpRunNotYetAvailable
Bases: Exception
Raised when nwp_init_time is not yet in the catalog.
Dynamical.org has not yet published that run.
Source code in packages/dynamical_data/src/dynamical_data/ecmwf_ens/download.py
16 17 18 19 20 | |
Functions:
open_ecmwf_ens_run(nwp_init_time, h3_grid)
Lazily open the ECMWF ENS Icechunk store and slice it to the requested run and H3 grid.
No data is downloaded: the returned dataset is still backed by lazy Dask/Zarr arrays.
Call download_ecmwf_ens_data to actually fetch the data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
nwp_init_time
|
datetime
|
The initialisation time to open. Must be timezone aware. |
required |
h3_grid
|
DataFrame[H3GridWeights]
|
The H3 grid to use for spatial bounds. |
required |
Returns:
| Type | Description |
|---|---|
Dataset
|
The catalog's dataset, sliced to the one |
Dataset
|
bounding box of |
Dataset
|
13 downloaded ECMWF ENS variables. |
Source code in packages/dynamical_data/src/dynamical_data/ecmwf_ens/download.py
57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 | |
download_ecmwf_ens_data(ds_sliced)
Download (compute) a lazily-opened, already-sliced ECMWF ENS dataset.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
ds_sliced
|
Dataset
|
A lazy dataset as returned by |
required |
Returns:
| Type | Description |
|---|---|
Dataset
|
The same variables and coordinates as |
Dataset
|
in-memory |
Dataset
|
downloaded concurrently. |
Source code in packages/dynamical_data/src/dynamical_data/ecmwf_ens/download.py
135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 | |
dynamical_data.ecmwf_ens.convert_to_polars
Converting a downloaded ECMWF ENS xarray dataset into the Nwp Polars contract.
The regular lat/lon grid is aggregated onto H3 cells on the way.
Attributes
Classes
Functions:
convert_nwp_xarray_dataset_to_polars_dataframe(ds, h3_grid)
Convert one downloaded ECMWF ENS run into the validated Nwp frame.
Loops over every (lead_time, ensemble_member) pair in ds. For each pair, aggregates that
chunk's grid points onto h3_grid's H3 cells and derives wind speed and direction from the
downloaded u/v components. Returns the concatenated result, validated against the Nwp
contract.
Guarantees, per cell:
- A numeric variable is the area-weighted mean of the points that supplied a value. The
numerator is
sum(v * p), where theproportionweightspsum to 1 over the whole cell by construction (seegeo.h3.compute_h3_grid_weights). The denominator is the contributing weight rather than that 1.0, so a missing point takes only its own share out of the cell instead of biasing the cell low. - A categorical variable is the category covering most of the cell's area, with an exact tie resolved to the lowest category code. Points that supplied no category are excluded from the ranking rather than competing in it.
- Either kind yields null, never
0.0or a spurious category, when no point contributed.
Each variable is renormalised over its own denominator, which is what keeps one variable's
corruption from nulling the others. The cost a caller must know about: two variables in one
cell can then be averaged over different sub-areas of the hexagon. So if wind_u_* and
wind_v_* ever have different null footprints, the wind vector derived by _calc_wind_speed
and _calc_wind_direction mixes two sub-areas. Upstream corruption has always been co-located
across variables, so that is theoretical today.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
ds
|
Dataset
|
One downloaded ECMWF ENS run, as returned by
|
required |
h3_grid
|
DataFrame[H3GridWeights]
|
The H3 grid weights to aggregate onto. One row per (H3 cell, NWP grid point) pair
whose cell and grid point overlap. |
required |
Returns:
| Type | Description |
|---|---|
DataFrame[Nwp]
|
One row per |
DataFrame[Nwp]
|
The wind components are dropped in favour of the derived speed and direction. Every other |
DataFrame[Nwp]
|
downloaded variable is carried through under its |
Source code in packages/dynamical_data/src/dynamical_data/ecmwf_ens/convert_to_polars.py
32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 | |
dynamical_data.ecmwf_ens.upstream_nulls
Measuring upstream corruption on the raw NWP grid, before the H3 aggregation sees it.
The H3 aggregation renormalises each cell over the grid points that supplied a value, so a corrupt grid point costs only its own share of its cell. That renormalisation is what makes the stored cells robust. That renormalisation is also why counting null cells is a poor proxy for how corrupt the feed was. This module counts the nulls where they arrive, before that renormalisation absorbs most of them. The aggregation mechanics are documented, along with the measurements behind the claim that a corrupt grid point takes only its own share out of its cell, at https://openclimatefix.github.io/nged-substation-forecast/architecture/ecmwf-ens-known-issues/#spatial-aggregation-is-where-a-grid-points-null-is-resolved.
Classes
UpstreamNullRate
dataclass
How much of one ingested NWP run arrived null on the raw grid.
This is the provider channel: the number to quote to Dynamical.org when asking whether their feed is degrading. It counts grid points on the 0.25° lat/lon box we downloaded, before any H3 aggregation.
Read it alongside contracts.weather_schemas.NwpQualityReport, never instead of that report.
NwpQualityReport counts null H3 cells and answers the different question of how much the
power-forecasting model lost. The two are not comparable as rates: different units over
different populations.
See https://openclimatefix.github.io/nged-substation-forecast/architecture/ecmwf-ens-known-issues/.
Source code in packages/dynamical_data/src/dynamical_data/ecmwf_ens/upstream_nulls.py
44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 | |
Attributes
per_variable
instance-attribute
One row per counted variable, with its n_null, n_affected_slices and n_total
grid-point counts, sorted by variable name. Every scalar below is derived from it, so a
breakdown and a total cannot disagree.
n_null_nwp_grid_points
property
Null grid points in the counted variables and steps.
n_total_nwp_grid_points
property
The denominator: counted variables × ensemble members × counted steps × grid points.
n_affected_nwp_slices
property
(variable, ensemble_member, lead_time) slices carrying at least one null grid point.
Separates "one bad slice" from "a hundred" at the same overall rate, which the fraction alone cannot.
affected_nwp_variables
property
The counted variables carrying at least one null grid point.
In per_variable's row order, which is sorted by variable name because
assess_upstream_grid_point_nulls builds it that way. filter preserves row order
rather than imposing a row order.
null_nwp_grid_point_fraction
property
Null grid points as a fraction of those counted; 0.0 when none were counted.
A run with no step left to count has nothing to measure, and a warning path must not raise (rule 7).
is_healthy
property
True when no counted grid point arrived null.
Methods:
__init__(per_variable)
Functions:
assess_upstream_grid_point_nulls(ds, variables, exclude_lead_0)
Count nulls on the raw NWP grid of one downloaded run.
Pure and Dagster-free (unit-testable in isolation). The ecmwf_ens asset calls it twice, once
per null population, and publishes each result on its own WARN check.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
ds
|
Dataset
|
One downloaded ECMWF ENS run, as returned by
|
required |
variables
|
Collection[str]
|
The variables to count over, named as |
required |
exclude_lead_0
|
bool
|
Skip the lead-0 step. True for the de-accumulated variables, which are null there by design, so counting it would report every healthy run as corrupt. False for the instantaneous ones, where lead-0 is an ordinary step and a null in it means what a null in any other step means. |
required |
Returns:
| Type | Description |
|---|---|
UpstreamNullRate
|
An |
UpstreamNullRate
|
by variable name. Each row gives that variable's null grid-point count ( |
UpstreamNullRate
|
of (ensemble_member, lead_time) slices holding at least one null ( |
UpstreamNullRate
|
the total number of grid points counted ( |
Source code in packages/dynamical_data/src/dynamical_data/ecmwf_ens/upstream_nulls.py
113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 | |