Testing
How the test suite is wired up, the house style for writing tests, and the notable test suites that
guard tricky invariants. This repository is a uv workspace monorepo: one root application, plus the
member packages under packages/, all sharing one lockfile and one environment. Its tabular data is
handled by Polars, the dataframe library, and its pipeline is orchestrated by Dagster. One testing
gotcha lives elsewhere because it is not really about tests: Polars row counts wrapping past 2³²
rows, in Performance and Scale.
Where tests and their dependencies live
- Test tooling is declared once, at the workspace root.
pytest,moto, andnumpylive in the rootpyproject.toml[dependency-groups] dev, and every workspace package inherits them. A package that gains atests/directory does not re-declarepytestin its ownpyproject.toml— we run the whole suite from the repo root withuv run pytest, against the root environment. (packages/geodeclares its ownpytest/pytest-cov; treat that as a historical exception, not the pattern to copy.) - Discovery is automatic. The only pytest configuration is the root
[tool.pytest.ini_options]block; there is notestpathssetting, so pytest collects both the top-leveltests/directory and everypackages/*/tests/directory. A brand-newpackages/<pkg>/tests/directory is picked up with no configuration change — provided the package is installed in the root environment, which is automatic only when something already depends on it. A leaf package that nothing depends on (dashboard, a marimo app, andstudies, the machinery the one-off studies call) is not in the default environment, souv run pytestcannot import its tests; add it to the root[dependency-groups] devlist (and give it a[tool.uv.sources]workspace entry) so a plainuv syncinstalls it. -
The
devgroup is also what keeps a research dependency out of the production image. TheDockerfilerunsuv sync --frozen --no-dev, so a package listed there is resolved for developers and for CI and absent from the image.studiesis the case that matters: it pulls inpvlib, which the live-forecast service has no use for. The check is one command, and it belongs in the pull request that adds such a dependency:test "$(uv export --no-dev --format requirements-txt | grep -ciE '^(pvlib|cdsapi)')" -eq 0Write it with
test "$(...)"rather than as a pipeline intogrep -q.grep -cexits 1 when it counts zero matches, so underset -o pipefail— which is what GitHub Actions gives everyrun:step — a pipeline form exits non-zero exactly when the check passes. - Run the whole suite with plainuv run pytest, never--all-packages.uv run pytestexecutes against the root environment, which holds exactly the packages reachable from the root's dependencies and dev group — i.e. every package that has tests, by the rule above.--all-packagesadditionally installs workspace members that have no tests (notebooks) and each member's own dev-groups, so it is heavier for no benefit here. It is the right tool for the pre-committyhook (uv run --all-packages ty check) for a different reason:tytype-checks the source of every workspace member, including leaf packages that are never installed as a dependency of anything, and that source must be present for the check. Type-checking needs the source; running tests needs the package installed — so the two commands legitimately differ. ---import-mode=importlibis set deliberately so that identically-named test modules in different packages (for example, twotest_storage.pyfiles) do not collide during collection. Because of this, test directories do not need__init__.pyfiles. - Test data files go in atests/data/subdirectory and are loaded relative to the test module withPath(__file__).parent / "data" / filename.packages/nged_data/tests/is the canonical example (it also keeps a small script documenting how the fixtures were trimmed down).
Fixtures and mocking
- Define fixtures inline in the test module by default. When a fixture — or a fixture factory —
is shared across more than one test module within a single package, put it in a package-level
tests/conftest.py.packages/dynamical_data/tests/conftest.pyis the example: it builds synthetic Xarray datasets that two test modules share. The only repo-rootconftest.pyholds cross-package pytest plumbing, not fixtures — the network-test gate below, Sentry data source name (DSN) neutralisation, and theOMP_NUM_THREADS/POLARS_MAX_THREADScaps described in Running the suite in parallel. Production code reports its errors to Sentry, and the DSN is the address those reports go to, so the rootconftest.pyblanks it and a test run sends Sentry nothing. - A factory shared across packages goes in the root
tests/directory, not in any one package'stests/. The rootpyproject.tomlsetspythonpath = ["tests"]for the wholeuv run pytestsession. Every module placed at the top level oftests/is therefore importable by bare name from any test suite in the repo —packages/delta_store/tests,packages/ml_core/tests, and the roottests/alike. Two modules already rely on that mechanism.tests/_nwp_test_data.pyrelies on it for its synthetic numerical weather prediction (NWP) writer, which builds frames matching theNwpschema, and for itscast_to_nwp_dtypesdtype helper.tests/_pytest_autoinject.py(see Running the suite in parallel) relies on it to be loadable as a pytest plugin by bare name. Putting a cross-package factory inside one specific package'stests/and importing it from another package's suite would work by accident of that package being installed, but it reads as a dependency of the package under test on another package's test code, which is backwards; the roottests/directory carries no such implication because it is not itself a workspace member. A factory production code needs (not just tests) still belongs incontractsor another library package, never here. - Mock with pytest's
monkeypatchfixture, notunittest.mock. Patch environment variables (monkeypatch.setenv), object attributes, and module-level functions (monkeypatch.setattr(some_module, "open", fake_open)) through the built-in fixture. For S3, drive the in-processmotoserver instead of mocking —tests/test_s3_data_paths.pyis the canonical pattern.motocan run as a real local HTTP server implementing the S3 API, so the code under test makes genuine S3 calls against a fake bucket rather than having its S3 client patched out. - Reset the moto S3 backend per test. The in-process
motoserver keeps its bucket contents in a process-global backend that outlives theThreadedMotoServerobject, so a module-scoped server does not hand each test a clean slate. A test whose write path runs twice against that server — a re-run, or state left behind by an earlier test — reads stale data: an appended Delta table returns double the rows, and anobject_existsprecondition sees a leftover parquet. Keep the server module-scoped for speed, but give each test a function-scoped fixture thatPOSTs to/moto-api/resetand recreates the bucket before the test body runs, so every test starts pristine and independent of execution order. -
Take the
dagster_instancefixture, or enter the instance as a context manager — never leave aDagsterInstance.ephemeral()unowned. The fixture (intests/conftest.py) enters the instance for you, sodispose()runs when the test ends. A module-level test helper that a fixture cannot reach —_run_live_checkintests/test_checks.py— takes the second form instead, and disposes on the way out of the helper.DagsterInstancehas no finaliser, so an undisposed instance defers two cleanups to whenever the garbage collector reaches it. Both then surface as failures owned by no test:- Its run storage and event-log storage each hold one SQLAlchemy connection to an in-memory
SQLite database. At interpreter shutdown the connection-pool finaliser can run after SQLite
has closed the database, printing a bare
Exception during reset or similartraceback after pytest's summary line. TemporaryLocalArtifactStoragedefers itstempfile.TemporaryDirectory()cleanup todispose()as well, leaving it to aweakref.finalizecallback that can fire at any point, not just at shutdown. The callback emitsResourceWarning: Implicitly cleaning up <TemporaryDirectory ...>, which the warnings-are-errors policy raises — and an exception inside a weakref callback is unraisable, so pytest'ssys.unraisablehookturns it into a hardPytestUnraisableExceptionWarningattributed to whatever unrelated test was running when the collector fired.
Neither is rare or environmental: both are what always happens when a used ephemeral instance is collected rather than disposed. Only which test gets blamed varies.
- Its run storage and event-log storage each hold one SQLAlchemy connection to an in-memory
SQLite database. At interpreter shutdown the connection-pool finaliser can run after SQLite
has closed the database, printing a bare
-
In a script, use the context manager directly —
with DagsterInstance.ephemeral() as instance:. Letting the local go out of scope is not enough: once the instance has run a job, Dagster's own caches retain it (aRunDomain, and the partition-loading contexts holding it asdynamic_partitions_store), so it reaches interpreter shutdown with both connections open even on a completely successful run. The context manager also covers the failure path, where an unhandled exception's traceback pins the raising frame. Worked example:scripts/forecasting/run_baseline_experiment.py. Measured with anatexitprobe after one realmaterialize: 2 connections still open with a bare call, 0 with the context manager, on both paths. build_asset_context()(and the otherbuild_*_context()helpers) needs the same treatment when called without an explicitinstance=. It defaults toDagsterInstance.ephemeral(), owned by anExitStackthat closes only on__exit__or__del__. Passed straight into an asset (some_asset(build_asset_context(...))) nothing ever calls__exit__, and inside apytest.raises(...) as exc_info:block the captured traceback keeps the context referenced past the end of that block, so__del__waits on a GC pass. Enter it as a context manager instead:with build_asset_context(...) as context, pytest.raises(...) as exc_info:. Worked example:tests/test_assets.py::test_ecmwf_ens_retries_when_run_not_yet_available— probing right after that test's teardown with no forcedgc.collect()found 2 open connections before the fix, 0 after, every time.build_asset_check_context()is the exception: give itinstance=, because it cannot be entered.DirectAssetCheckExecutionContextdefines no__enter__, so the context-manager form above is aTypeError, and thewithblock has to hold the instance instead. Dagster builds one of these for a directly-invoked check that declares no context parameter too, sopower_data_is_fresh()leaks an instance on a call that mentions Dagster nowhere; pass it a context anyway, which Dagster accepts and uses in place of the one it would build. Worked examples:_run_freshness_checkand_run_live_checkintests/test_checks.py.materialize()andJobDefinition.execute_in_process()do not need this. Both wrap their default instance in their own internalwith ephemeral_instance_if_missing(instance):, entered and exited inside the call, so disposal is already deterministic however the caller uses the return value. The distinguishing question for any Dagster helper that can default-construct an instance is whether the helper itself closes thewithblock, or hands you an object that expects you to.- An autouse fixture fails whichever test leaks one.
_fail_on_an_undisposed_dagster_instanceintests/conftest.pywrapsDagsterInstance.ephemeralandDagsterInstance.disposefor the duration of each test, then asserts at teardown that every instance created was disposed — so the fault is reported by name instead of against an unrelated test. The fixture measures instances created against instances disposed at teardown, which is not the same as the coding mistake that usually causes an undisposed instance. It therefore catches a leak, not an omission. An unreferenced context is freed as soon as the helper that built it returns, so a missinginstance=can still pass green until something pins the frame holding it.
Running the suite in parallel
A plain uv run pytest — no path or node id and no -n of its own — runs under pytest-xdist with
one worker process per physical CPU core (-n auto; pytest-xdist counts physical, not logical,
cores when psutil is installed). The flag is added by the tests/_pytest_autoinject.py plugin,
loaded via -p _pytest_autoinject in addopts, rather than a static addopts entry naming -n
auto directly: the plugin can see the invocation before adding the flag, so a targeted run (uv run
pytest path/to/test_foo.py::test_bar) stays serial instead of paying worker start-up cost for one
test. An invocation that already passes its own -n is left alone; a -k/-m selection is not
treated as a target and still runs in parallel, since a filtered run can still be many tests.
-s/--capture=no is treated as a reason to stay serial, unlike -k/-m: under xdist a
worker's captured-off stdout never reaches the controller, so parallelising a -s run would
silently drop the output -s exists to show — exactly the debug-print step of the "single test"
loop this plugin otherwise protects. The plugin reads pytest's own parsed known_args_namespace
rather than hand-scanning the raw argument list. Every pytest flag that takes a value (-W, -o,
--deselect, …) is therefore handled correctly, not just the ones the plugin happens to name (see
the plugin's own docstring for the one narrow exception).
The auto-injection can't be a pytest_load_initial_conftests hook in the root conftest.py itself
— that hook fires as part of loading the root conftest.py, so a hookimpl defined inside it is
never registered in time to run for that same call. tests/_pytest_autoinject.py is only importable
as a bare-name -p plugin because the root tests/ directory is on pythonpath (see Fixtures and
mocking); removing that ini setting breaks every invocation loudly, with an
ImportError at startup, rather than silently falling back to serial.
The root conftest.py caps OMP_NUM_THREADS and POLARS_MAX_THREADS at 4 for the whole session,
set as plain module-level code rather than inside pytest_configure: Polars reads
POLARS_MAX_THREADS once, the first time it's imported, and tests/conftest.py — loaded as one of
the initial conftests, before pytest_configure runs — imports the Dagster defs module, which
imports Polars transitively. A pytest_configure-time set is already too late and silently caps
nothing in the process running the tests directly (a serial run, or the -n auto controller
process; xdist workers still pick up the cap, inherited via the environment at their own start-up).
XGBoost and Polars each default to a thread pool sized to the machine's core count, so without the
cap, -n auto's one-process-per-core plus each process's own full-width thread pool oversubscribes
the machine by core-count² and every worker's threads fight the others for the same cores — measured
on a 32-thread workstation, removing the caps made the full suite run in 95s with 4755 threads and a
load average of 201, versus 21-24s capped. The cap is 4 rather than 1 because it costs nothing in
wall-clock on that workstation while matching production more closely, where a single job runs with
no thread cap at all.
Warnings are errors
[tool.pytest.ini_options] filterwarnings in pyproject.toml starts with error, so any
warning fails the test that raised it. A deprecation introduced by our own code is therefore caught
by the PR that introduces it, instead of accumulating in the warnings summary.
The exceptions listed after error are third-party deprecations we cannot fix from this repo. Each
one is pinned to the exact warning message and the exact upstream module that raises it — the
module anchor is what stops an entry from ever masking the same deprecation appearing in our own
code — and carries a comment naming the package, the version, and the condition for deleting it.
When a dependency upgrade introduces a new upstream warning, the suite fails loudly. Add a new entry
in that same form rather than widening an existing one; if the warning comes from our code, fix the
code. Later entries take precedence, so error stays first.
A production except BaseException swallows pytest's fail and skip too
pytest.fail() and pytest.skip() raise from BaseException, so a broad except BaseException
in the code under test swallows them along with every real error. defs/checks.py and
defs/assets.py each guard an asset check's body with except BaseException, because a Rust panic
from a compiled dependency does not derive from Exception. Polars and the Delta Lake bindings are
both compiled Rust extensions. See Warn on stale power data with a Dagster asset
check for the full
reasoning. Calling pytest.fail() or pytest.skip() inside such a guarded body, as a "this
branch must not run" sentinel, is caught by the same guard instead of failing the test. Assert after
the call returns instead of relying on either one inside it.
Network-gated tests
Most tests run fully offline, mocking any network call (for example, patching
dynamical_catalog.open to return a synthetic xr.Dataset). A handful of tests are worth running
against a real external service — chiefly to catch the shared-convention blind spot, where a
synthetic fixture and the code under test share the same wrong assumption about the real data's
shape (dimension order, latitude orientation, longitude range, dtypes, units) and both pass.
Mark such a test @pytest.mark.network. The root conftest.py skips every network-marked test
unless the caller passes --run-network, so a plain uv run pytest (local dev and the per-PR CI)
never touches the network. Run them explicitly — the nightly CI job (see Continuous
integration below) or on demand — with:
uv run pytest --run-network # whole suite, network tests included
uv run pytest --run-network -m network # only the network tests
-m network alone names no path or node id, so the nightly CI job's uv run pytest --run-network -m
network (currently one test) still runs under -n auto — pure worker start-up overhead for a
single-test selection (see Running the suite in parallel).
packages/dynamical_data/tests/test_ecmwf_ens_network.py is the canonical example: it drives the
real open → download → convert pipeline against the Dynamical.org ECMWF ENS catalog and asserts
the conventions the offline fixtures merely assume. Dynamical.org publishes the European Centre for
Medium-Range Weather Forecasts' ensemble forecast (ECMWF ENS) as a public cloud-hosted catalog, and
that forecast is this project's weather input.
The gate is a collection hook, not an addopts = "... -m 'not network'". pytest keeps only the
last -m it is given, so any caller-supplied marker expression (e.g. -m "not integration")
would silently replace an addopts -m "not network" and re-include the network tests. A skip
applied during collection cannot be overridden that way — the gate holds whatever -m the caller
passes, and even -m network alone stays skipped until --run-network is added.
Continuous integration
Two GitHub workflows in .github/workflows/ run the checks described on this page:
ci.yml— the per-PR quality gate. Runs on every pull request and every push tomain:ruff check,ruff format --check,ty check, thepymarkdown scancommand from CLAUDE.md,mkdocs build --strict,check_docs_links.py(see below), and the offline test suite (plainuv run pytest— the network gate above keeps CI off the network). The job installs withuv sync --locked --all-packages:--all-packagesbecausetytype-checks the source of every workspace member, including leaf packages that a plain sync would omit, and--lockedso the build fails loudly whenuv.lockis stale. Every subsequent step passesuv run --no-sync, because a bareuv runre-syncs to the root environment and would silently uninstall those extra workspace members. The job also sets dummy values for the three requiredNGED_S3_*Settingsfields, aSettingsobject being what carries the S3 credentials and bucket names that production reads from the environment; NGED is National Grid Electricity Distribution, the network operator whose telemetry this project forecasts. Most tests monkeypatch those fields, but a few constructSettings()directly and locally rely on the developer's.env, which CI doesn't have. Thecijob is a required status check onmain(configured in a GitHub repository ruleset, not in the workflow file).nightly_network_tests.yml— the nightly network job. Runs only the network-gated tests (uv run pytest --run-network -m network) on a daily schedule, plusworkflow_dispatchfor on-demand runs. This is the only CI that touches the real Dynamical.org catalog, and it needs no secrets — the catalog is public. A failure notifies (GitHub emails the workflow author when a scheduled run fails) but deliberately does not block PRs: a red nightly run signals drift in the upstream catalog's conventions, not a defect in whatever PR happens to be open.
Two steps validate links, because neither sees what the other does. mkdocs build --strict
fails on a broken link within docs/. scripts/lint/check_docs_links.py fails on a link into
the published site — the form CLAUDE.md requires from a docstring, a comment or a GitHub issue body.
docs/ is rendered to a public website, and a link from outside docs/ must name that site's URL
rather than a repository path. Renaming a page moves the URL that link points at, and rewriting a
heading kills the anchor it points at.
A script guessing an anchor slug would strip or hyphenate the underscores in a heading, as the
common slug convention does. Python-Markdown's toc extension instead preserves underscores, so a
guessed slug and the real slug differ. check_docs_links.py therefore resolves each anchor by
running the real markdown.Markdown() converter over the target page rather than guessing a slug.
An extension the script cannot load fails the run rather than being skipped: dropping
pymdownx.superfences makes a # comment inside an indented fenced code block parse as a heading,
which would invent anchors the real site does not have and pass links that are broken. The script
also reports a URL that has been reflowed across two lines, because the anchor left stranded on the
second line is no longer part of the link for the reader either. The one gap is a docs/api/ page,
whose anchors mkdocstrings generates at build time, so only the page's existence is checked there.
The check runs in ci.yml and as a pre-commit hook, scanning the whole repo each time because a
link can sit in any text file.
Why a bespoke workflow rather than OCF's template
Open Climate Fix (OCF) is the organisation whose GitHub account holds this repository. OCF's
organisation template (openclimatefix/.github →
workflow-templates/branch_ci.yml)
is a thin caller of the org-wide reusable workflow
branch_ci.yml.
We deliberately don't use it, because it is built for OCF's standard single-package service repos
and fits this repo poorly:
- Single-package assumptions. The reusable workflow expects one
pyproject.tomland one test folder (itstests_folderinput defaults tosrc/tests). This repo is a uv workspace monorepo: the suite must run from the root so pytest collectstests/plus everypackages/*/tests/, and type-checking needs an--all-packagesinstall (see above). - Floating tool versions. It lints and type-checks with
uvx ruff/uvx ty, i.e. whatever version is newest on the day the job runs — so org CI can disagree with the locked dev-group versions used locally and by pre-commit, and a new upstream release can break CI with no change in this repo. We run the locked versions viauv run. - Missing checks, and no offline/network split. It has no equivalent of
ruff format --check(formatting is opt-in and separate) or thepymarkdown scancommand, and no concept of this repo's network-gated nightly job. - Container build baggage. Roughly half the reusable workflow builds and publishes Docker images to ghcr.io; this repo doesn't produce a container image.
- Trigger and pinning mismatch. The template triggers on pushes to non-default branches, whereas
a required status check wants
pull_request(+ push tomain); and it pins the reusable workflow@main, so upstream edits to the org workflow change this repo's CI without any PR here.
We did keep the template's good conventions: a concurrency group with cancel-in-progress, a cached
astral-sh/setup-uv, and locked installs (UV_LOCKED=1 there, uv sync --locked here). If OCF's
reusable workflow ever grows first-class uv-workspace support, revisiting it would shrink ci.yml
to a few lines — but until then the bespoke workflow is smaller than the configuration the template
would need.
Keeping the interpreter and dependencies current
.python-version pins the exact interpreter patch uv resolves, so uv sync gives the same
Python everywhere — local development, CI, and pre-commit — rather than whatever patch happens to
already be installed or however uv's own downloader resolves on the day a job runs. Without a
pin, CI's astral-sh/setup-uv step (see above) can pick up a newer patch on a later run with no
diff in this repo, so a bug that depends on the interpreter's exact patch version can stop
reproducing between two runs of the same commit.
Bump .python-version in the same pull request that runs uv lock --upgrade, not on its own
schedule. Run uv python install <version> then uv python pin <version> for the newest released
patch of the pinned major version (check with uv python list), so the interpreter and the locked
dependencies age together and both get retested at once. A release candidate for the next major
version is not a candidate for the pin: wait until it ships as final and the compiled dependencies
this repo relies on — pyarrow, polars, deltalake — publish wheels for it.
NWP grid → H3 orientation coverage
H3 is a hexagonal spatial index that tiles the globe, and the project aggregates gridded weather
values onto those hexagons. The NWP-grid-to-H3 mapping is the classic place for a silent orientation
bug — a vertically or horizontally flipped weather grid, a transpose (np.meshgrid indexing="ij"
vs "xy"), or a lat/lon swap. Four tests guard it in layers, from cheap-and-synthetic to
real-and-networked. Each was checked by mutation: introducing the bug into the production code and
confirming the test goes red. The table records which mutation each layer catches (✓ = the test
fails when that bug is present; n/a = the test does not exercise the code that mutation changes):
| Mutation in production code | synthetic convert test¹ |
cached real-slice test² | geo landmark test³ | geo snapping test⁴ |
|---|---|---|---|---|
np.meshgrid indexing="ij" → "xy" (transpose) |
✓ | ✓ | n/a | n/a |
| lat/lon swap in the value-join keys | ✓ | ✓ | ✓ | n/a |
| reversed latitude ravel (vertical flip) | ✓ | ✓ | n/a | n/a |
swap cell_to_lat/cell_to_lng when snapping |
n/a | n/a | ✓ | ✓ |
- ¹
dynamical_data/tests/test_convert_to_polars.py::test_convert_maps_each_grid_point_to_its_own_lat_lon— a 2×2 synthetic grid with a distinct value at every corner.convertflattens the latitude/longitude grid into a column of values that must stay aligned with a column of coordinates, and then joins those rows to the H3 weights on coordinate value. This test guards that ravel-alignment step insideconvert. The value-join itself does not depend on row position, so which hexagon owns a given (lat, lon) is delegated to the upstreamh3_grid_weightsasset (the geo test below). - ²
dynamical_data/tests/test_ecmwf_ens_cached.py::test_cached_real_slice_conventions_and_orientation— the same orientation check on a committed real ECMWF ENS slice, so the conventions the synthetic fixture only assumes (descending latitude, °C, dimension order) are exercised on genuine bytes. - ³
geo/tests/test_h3.py::test_grid_weights_preserve_geographic_orientation— provescompute_h3_grid_weightslabels each H3 cell with grid points at the cell's own (lat, lon), using two well-separated landmarks in Great Britain. This is what fixes the hexagon↔(lat, lon) geography theconverttests delegate. - ⁴
geo/tests/test_h3.py::test_grid_weights_snap_to_nearest_grid_centre— builds the expected grid points geographically and compares them against whatcompute_h3_grid_weightsproduced, so it catches thecell_to_lat/cell_to_lngswap as well.
test_ecmwf_ens_network.py (network-gated, above) runs the full open → download → convert pipeline
against the live catalog, but only re-checks orientation and bounds: descending latitude,
longitude in [-180, 180], the slice landing on the requested box, the expected variable names, and a
physical-range sanity check on temperature. It does not re-check the value↔(lat, lon) mapping —
neither the ravel-alignment mutation the synthetic convert test guards nor the hexagon↔(lat, lon)
geography the geo landmark test guards. Both of those bugs misalign indices rather than depending on
the actual numbers, so any values at all reproduce them, and both are therefore already fully proven
offline. What only this test can catch is future upstream drift — a change in Dynamical.org's own
conventions that the committed slice, frozen at capture time, cannot.
Marimo notebooks bind every name their cells reference
Marimo rebuilds a notebook from its with app.setup: block plus its @app.cell functions, and
never runs the module-level statements in between. A name bound at module level is therefore
invisible to every cell: the notebook raises NameError the next time it is opened, while ruff, ty
and pytest all pass, because the file they were handed is valid Python. Two tools produce exactly
that shape from a working notebook — ruff check --fix, which writes an import an autofix needs
into the top-level import block, and marimo check --fix, which deletes such an import and rewrites
the cell that used the name as def _(name), leaving a cell input nothing defines.
scripts/lint/check_marimo_notebooks.py reads each cell's refs and defs and reports any name a
cell references that no cell binds. It runs as a pre-commit hook over changed notebooks, and
tests/test_marimo_notebooks.py runs it over every notebook in packages/notebooks/ and
packages/dashboard/. Three properties are worth knowing:
- It is static. Nothing executes, so the check needs none of the notebooks' runtime dependencies
— only marimo itself, which the root environment has via the
dashboarddev dependency. It cannot catch a notebook that binds every name and still fails inside a Polars or Altair call; executing the notebooks is not an option, because they read real Delta tables and S3. - It rides on private marimo API.
Cell.refsandCell.defsare documented, but loading a notebook without running it is not. So the checker raises rather than reporting "no findings" whenever a file does not parse into at least one cell, and the tests keep a positive control — a deliberately broken notebook, held as a string so ruff never sees it — that fails if a marimo release stops the check detecting a real breakage. - Every
.pyfile directly inside those two directories must be a notebook, and a file that is not one is a finding. The ruff pre-commit hooks share that assumption: they use it to decide which files must never be auto-fixed.
Testing what a notebook's cells actually do is a separate job, and
packages/notebooks/plot_missing_NWP_data.py is the worked example. Its chart-building helper is an
@app.function — marimo's form for a top-level reusable function — so an ordinary test_* function
in the same notebook can exercise it on a synthetic frame. Naming the notebook in python_files is
what makes a plain uv run pytest collect it. The authoring rules for writing one are in the
marimo-notebooks skill.
Assertion style for Patito frames
A tabular contract's shape is declared as a Patito model: a schema class naming the frame's columns and
their dtypes. Bind a frame to its model with set_model, coerce the columns to the model's dtypes
with cast, and validate for the happy path:
df = pt.DataFrame({...}).set_model(MySchema).cast()
df.validate() # happy path: raises nothing
MySchema.validate(existing_df) # or validate a frame produced elsewhere
For the unhappy path, assert that validation raises:
from patito.exceptions import DataFrameValidationError
with pytest.raises(DataFrameValidationError):
bad_df.cast().validate()
packages/contracts/tests/test_geo_schemas.py is the shortest end-to-end example.