August 4, 2026
How We Accelerate ML Deployment by Empowering Data Scientists (Part I)
Orchestration of ML systems at scale, the easy way
Written by Jordan Melendez
This article builds on Part 1: Building a Reproducible Data Science Environment.
In Part 1, we focused on the foundations of reproducible machine learning: packaging Python code, creating deterministic environments with uv, and standardizing OS-level development with Docker and dev containers.
The choices made in Part 1 exist to make the next stage dramatically simpler. Once structure and environment are in place, the remaining challenge is executing reproducible ML workflows from local experimentation all the way through production.
That’s where orchestration, configuration, and workflow tooling enter the picture.
Running Data Science Code
The tooling we’ve integrated so far solves reproducibility and developer experience. The next question is how those same projects should actually be executed. Once your project has structure and an environment that permits rapid iteration loops, you are ready to begin actually running experiments!
For data scientists and MLEs, Initial iteration may begin in SQL, notebooks, scripts, or BI tools. Eventually you may end up with a series of notebooks or scripts that looks something like
.
├── 00_start.py
├── 01_preprocess.py
├── 01_preprocess_final.py
├── 02_train.py
├── 03a_evaluate.py
├── 03b_evaluate.py
└── 04_cleanup.py
It is then up to the user to ensure that each is run in sequence, with intermediate artifacts propagated from step to step. A colleague, or yourself from the future, may wonder which files are outdated, and whether 01_preprocess.py should be run at all. Good documentation can help, but can go out of date. Furthermore, some of these scripts may have hardcoded values of hyperparameters, data input/output locations, etc. (Hopefully, these scripts have already moved repeated logic to the core package to keep reusable functions clean!)
There are now two concepts to introduce to help clean this up: the specification layer, and the orchestration layer. These layers exist to complement the data science layer that was introduced as the code packaged in core. These are explained in detail in this excellent article. They can be summarized as follows
Data science layer: Describes what is done. It contains the core data-science code specific to the problem we are trying to solve.
Specification layer: Describes how it is done. This includes, for example, the parameters required to run the code (where data lives, what hyperparameters to use, which estimators, etc.). It could also specify where jobs are run, instance sizing, the execution environment, and so on. We’ve already discussed components of the execution environment, such as the
uv.lockandDockerfile. The specification layer only depends on the data-science layer.Orchestration layer: Executes the code defined in the data-science and specification layers. It depends on both the data-science and specification layer.
Importantly, the outer layers only depend on the inner layers, not the other way around. That is, configuration and orchestration logic should not creep into your data-science layer. A preprocessing function may accept a dataframe that was already loaded from a particular location, or an estimator may be trained with particular hyperparameters, but this information comes from the specification layer, it is not hardcoded in your functions.
Let’s discuss these two new layers, and describe what new workflows they enable in the ML lifecycle. I will wrap up by describing how command runners, such as just, help in this ecosystem.
The Specification Layer
I rely primarily on YAML files, combined with pydantic, as the backbone of my specification layer. (Tools like OmegaConf exist for more advanced use cases, which generalize nicely from what is described here.)
The configuration for a training run lives in two files. First, a typed Pydantic model that declares the shape of the config and validates it at load time:
# src/core/config.py
from pydantic import BaseModel
class TrainingConfig(BaseModel):
"""Parameters that govern a training run."""
n_estimators: int
max_depth: int
random_seed: int
Second, a YAML file that holds the actual values:
# config/training.yaml
n_estimators: 100
max_depth: 6
random_seed: 42
Why this combination? A plain dict of hyperparameters works but has no type validation. A new colleague (or yourself from the future) will not get immediate feedback if they use a string in a place that should accept a float. A Pydantic model makes the contract explicit, where typos fail at validation, and every field is typed and self-documenting.
The YAML file is equally important. Because it is plain text, a change from max_depth: 6 to max_depth: 8 is a one-line diff in version control. It appears once in a centrally defined location, not in the middle of a script. Colleagues can easily find what levers are exposed and how to change them, with confidence that the changes will propagate to all downstream code. The YAML files can hold more than hyperparameters, and can be extended to data path locations, flags for where code should run, and even the estimators themselves. Advanced configuration-driven development patterns can follow once the code is split up in this way.
The Orchestration Layer
For the orchestration layer, I have found that Metaflow provides the best developer experience for data science workflows, and the competition is not even close. Key aspects of an effective orchestrator include
The ability to easily run locally or on the cloud without code changes
Enable fast iteration loops for data scientists, including retrying failed jobs, debugging issues, and so on
Integrate with how data scientists naturally work and think, without forcing us into constrained patterns to fit a framework
Metaflow has all of these properties and more. Let’s explore with a simple example.
A Metaflow flow is a Python class that describes a directed graph of steps. An example training flow for this project is shown below:
# flows/training_flow.py
from metaflow import FlowSpec, step, Config
from core import features, models
from core.config import TrainingConfig, pydantic_parser
class TrainingFlow(FlowSpec):
"""End-to-end training pipeline.
Run locally:
python flows/training_flow.py run
Run on AWS Batch:
python flows/training_flow.py run --with batch
"""
config: "TrainingConfig" = Config(
name="config",
default="../config/training.yaml",
parser=pydantic_parser(TrainingConfig),
)
@step
def start(self):
print(f"Starting training run with config: {self.config}")
self.next(self.featurize)
@step
def featurize(self):
import polars as pl
raw_data: pl.DataFrame = pl.DataFrame(...) # load your data here
self.feature_data = features.transform(raw_data)
self.next(self.train)
@step
def train(self):
X = self.feature_data.drop("y")
y = self.feature_data["y"]
params = {
"n_estimators": self.config.n_estimators,
"max_depth": self.config.max_depth,
"random_state": self.config.random_seed,
}
self.model = models.train(X, y, params)
self.next(self.evaluate)
@step
def evaluate(self):
X = self.feature_data.drop("y")
y = self.feature_data["y"]
self.metrics = models.evaluate(self.model, X, y)
self.next(self.end)
@step
def end(self):
print(f"Config: {self.config}")
print(f"Metrics: {self.metrics}")
if __name__ == "__main__":
TrainingFlow()
A few things worth noting here.
The flow is lightweight. It imports from core and accepts configuration as a parameter. Its only job is to wire those pieces together and manage data flow between steps. There is little to no data science logic in the flow itself, these instead live in features.py and models.py. Edits to core immediately impact how the flow runs. Multiple different Metaflow flows can reuse the same logic stored in core without copying code. This respects the separation between the data-science layer and the orchestration layer.
Metaflow natively supports uv environments. By using the same base docker image and uv.lock file in Metaflow as you use for local development, Metaflow will work identically for both local and cloud jobs. This can be enabled via the environment variable METAFLOW_ENVIRONMENT="uv" or by running your jobs with the--environment uv flag, as in uv run flows/training_flow.py --environment=uv run.
There is one nuance here around mixing your core package with Metaflow: you need to symlink your src directory within the same directory that your flows live so that Metaflow can pick it up. You can achieve this via ln -s ../src flows/src. This allows you to retain all the benefits of packaged code but without explicitly installing the package into the image that Metaflow uses. Metaflow recently open-sourced the @package_sources decorator exactly for this use case, and so could help reduce the need for symlinks going forward.
Metaflow gives you data lineage for free. Metaflow artifacts, such as self.feature_data and self.model, are automatically serialized and stored at each step boundary. Reproducing a result from three months ago means finding the run ID and loading its artifacts, rather than tracking down a pickle file on someone’s laptop.
Configuration is separate from orchestration logic. The Config object is Metaflow’s mechanism for loading structured configuration. It reads config/training.yaml, parses it through the Pydantic validator, and makes the typed TrainingConfig available as self.config throughout the run. The config is also automatically persisted as a run artifact, so every historical run has its parameters stored alongside its outputs.
Metaflow can run the same locally as in the cloud. The same flow runs locally with python flows/training_flow.py run and on AWS Batch with python flows/training_flow.py run --with batch. This ships your jobs to the cloud without any changes to your actual code. Outside of AWS Batch, Metaflow also supports running on K8s, Airflow, and Kubeflow. With ephemeral cloud compute, it now becomes trivial to fan out to 1000s of jobs simultaneously training different models, something that would be impossible if limited to your local machine.
Although not Metaflow specific, it’s worth calling out a pattern when running jobs locally and in the cloud: try to keep data paths in config, not in code, and choose I/O libraries that treat local paths and cloud URIs identically. Polars, which this project uses, handles this naturally: pl.read_parquet(“data/local.parquet”) and pl.read_parquet(“s3://my-bucket/data.parquet”) are the same call, just a different string. The same applies to smart-open as a drop-in for Python’s open(), and to any library built on fsspec.
If you’re interested in learning more about Metaflow, I recommend their documentation and, if you’re interested in a deep dive on the design philosophy with concrete examples, see Effective Data Science Infrastructure written by Ville Tuulos, who designed and built Metaflow.
Command runners
Every project accumulates a vocabulary of commands. For example, how to install dependencies, run tests, lint, format, or run a flow. This vocabulary usually lives in a README that goes stale, in contributors’ heads, or nowhere at all. And every team may work slightly differently, making context switching all the more painful. When onboarding a new team member, the first hour is often spent just figuring out which commands to run and in what order.
just solves this. It is a modern command runner and replaces make for most data-science workflows. The justfile at the root of the project is the single source of truth for every workflow command:
sync:
uv sync
test:
uv run pytest
test-one file:
uv run pytest {{file}}
lint:
uv run ruff check .
format:
uv run ruff format .
typecheck:
uv run ty check
check: lint typecheck
train:
cd flows && uv run python training_flow.py --environment uv run
just --list is the entry point for any new contributor. They do not need to read the README, grep through scripts, or ask a colleague. Running just check before opening a PR is the same command CI runs.
CI/CD
The setup outlined above lends itself naturally to CI/CD pipelines. With just in place, the commands I run locally are exactly the commands CI runs. There is no separate CI-specific scripting layer to maintain.
This also keeps CI out of the fast path for day-to-day experimentation. A data scientist iterating on a model should not be waiting for a pipeline. CI runs on pull requests before anything merges, and the deploy workflow runs the flow upon merging to main. It is a quality gate before production, not a bottleneck during research.
Because I have already separated the data science, specification, and orchestration layers, CI can do more than just lint and test. The orchestrated flow that runs locally for experimentation can be run and validated in a staging environment with models and metrics logged to a staging-specific experiment tracker or model registry.
By implementing these practices, moving from the world of “shipping models” towards “shipping model factories” is close to flipping a switch.
Summary
These series of posts have covered three layers of a reproducible ML project: structure, environment, and tooling.
For structure, packaging your data science code with uv and the src/ layout makes it importable anywhere, just as easily as any other package. A separate specification layer, with corresponding Pydantic models, provides a typed, diff-able home in version control instead of scattering magic numbers through scripts.
For the environment, uv pins the Python dependency graph to a lockfile that travels across laptops, CI runners, and cloud batch jobs. Dev containers extend that guarantee to the OS layer, so a new contributor goes from git clone to a running training flow without a setup guide.
For tooling, Metaflow ties it together as the orchestration layer: a thin flow that imports from core, reads config, and runs identically locally or on the cloud with a single flag. The just command runner gives every workflow a single command, and CI runs those same commands before anything ships.
None of these decisions are large in isolation. Together, they make the right way the easy way. This leads to faster onboarding, experimentation, and reproducibility, each of which set you up for scaling to the cloud and to production.
Related Posts

August 04, 2026
How We Accelerate ML Deployment by Empowering Data Scientists (Part I)
Explore how Root empowers data scientists to move machine learning projects from experimentation to production faster with reproducible development environments, standardized tooling, and infrastructure designed to make best practices the easiest path. Read more

September 08, 2021
The Data Science Research-Production Chasm
Explores the challenges that can emerge between data science research and production and practical strategies for building more reliable, scalable machine learning pipelines. Read more

September 11, 2026
How Root Rebuilt Its Telematics Platform for Scale
Learn how Root rebuilt its telematics platform to handle growing scale, improve reliability, and create a more flexible foundation for processing the driving data that powers its insurance products. Read more

December 15, 2021
Six Under-Appreciated Principles of Data Warehousing, Through the Eyes of a Data Scientist
Challenges can emerge between data science research and production. This article explores practical strategies for building more reliable, scalable machine learning pipelines. Read more