ML Experimentation
How we run and evaluate ML forecasting experiments — the durable home for this methodology once it leaves the roadmap, per the Documentation Guide.
Documents
- Running an ML experiment end-to-end — step-by-step recipe for going from
raw data to a trained, MLflow-tracked model using the Dagster pipeline; explains why
trained_cv_modelreads config from MLflow rather than YAML. - Model configuration — how to set hyperparameters and choose features; the full feature vocabulary and the lookahead-bias guardrails.
- Cross-validation folds — the expanding-window CV protocol, the current single fold and why the data constrains us to it, the target multiple-yearly-fold protocol, and how to evaluate a data source whose history is shorter than the folds.
Our approach to MLOps
Modern machine-learning-operations (MLOps) practice, as used in this project, changes throughput and closes the translation gap:
- Throughput. The grunt work of an experiment — assembling features, training, cross-validating, recording results — is automated, so a small team can run hundreds of experiments per month instead of one or two. Humans still decide what to try; the infrastructure runs those experiments. (See Running an ML experiment end-to-end.) This project is betting that the throughput produces a better forecast; the literature has not settled the question. The energy-forecasting review found no study measuring what adopting machine-learning-operations practice delivers, and the case for fast, comparable iteration rests on a structural argument and on practitioner testimony instead.
- No translation gap. The artifact we experimented on is the artifact we deploy. There is no "now rewrite the research code for production" step, because every experiment runs on the exact same code as the production pipeline from the start. The gap is closed by raising research to the production standard, not by lowering production to accept a research notebook: an idea can be explored anywhere, but it only becomes a runnable experiment once it lives in the pipeline's own code. The debt this avoids has a name: Sculley et al. (2015) call the code written to bridge research and production glue code, and the tangle that glue code grows into pipeline jungles. Sculley et al.'s list of debts also constrains how the gap is allowed to close, because dead experimental codepaths is what accrues when experiments run as conditional branches inside production code. In contrast, an ML experiment in a modern MLOps pipeline selects a different configuration and runs the same code path, rather than adding a branch the production pipeline then has to carry.
An analogy
Traditional machine-learning research and development is a chef inventing dishes in their home kitchen: every winning recipe has to be laboriously re-created on the restaurant's equipment before it can go on the menu, and much is lost (or silently changed) in translation. We are building the restaurant where R&D happens on the service line itself: hundreds of tastings a month, every dish judged by the same tasting panel, and the winning dish on the menu the same night — because nothing about it needs translating.
Nothing gets rewritten on the way to production
The model that wins the evaluation is, bit for bit, the model we deploy — not a re-implementation of it. Promotion to production takes minutes, and that speed is a consequence of rigour, not a trade against it. By the time promotion is on the table, the candidate has already been trained, cross-validated (see Cross-validation folds), and evaluated on the same pipeline, under the same standardised protocol, as every model before it. Holding that protocol fixed is what makes two experiments comparable at all. That is the reason Karpathy's autoresearch fixes every training run to the same 5-minute budget: a fixed budget "makes experiments directly comparable regardless of what the agent changes".
Holding the protocol fixed is what makes a one-command promotion safe to press rather than merely quick. In a conventional setup, the artifact measured and the artifact deployed can be two different pieces of code — a risk that does not exist here. The comparison that picked the winner was made against every other candidate on identical folds. And the way back to the previous champion is a single command too. A fast promotion route is worth no more than a slow route if nobody trusts the fast route enough to use it. The speed comes from the protocol rather than from haste, a distinction Karpathy (2019) puts bluntly: "a 'fast and furious' approach to training neural networks does not work and only leads to suffering".