{
    "componentChunkName": "component---src-components-pages-blog-post-page-js",
    "path": "/blog/how-we-accelerate-ml-deployment-by-empowering-data-scientists-part-ii/",
    "result": {"data":{"contentfulBlogPostPage":{"id":"56d960e4-853f-5c8c-91aa-55e0afd91e31","publishingMetadata":{"sys":{"type":"Entry","revision":2,"contentType":{"sys":{"type":"Link","linkType":"ContentType","id":"publishingMetadata"}}},"id":"4a92a033-3fb3-5c57-817d-fb46f42b1b37","deIndexPage":null,"includeClarity":null,"includeOptimizely":null,"pageTitle":"2026 // QS Blog / How We Accelerate ML Deployment by Empowering Data Scientists (Part II)","pageDescription":"Explores how empowering data scientists with the right tools, infrastructure, and ownership can streamline machine learning deployment, reduce bottlenecks, and accelerate the transition from research to production.","pageThumbnail":null,"slug":"/how-we-accelerate-ml-deployment-by-empowering-data-scientists-part-ii"},"navigationOptions":null,"blocks":[{"sys":{"type":"Entry","revision":15,"contentType":{"sys":{"type":"Link","linkType":"ContentType","id":"blogHeaderSection"}}},"id":"fc5f8a7d-447a-5d79-9421-2b68664b1030","blogLogo":{"file":{"url":"//images.ctfassets.net/pacigpl3aj13/5PHax8jFbEQIv4fjnCUniG/717b27de99d7b6f53add3fc1b107d09b/Root-blog-Logo.svg"}},"blogLogoAltText":null},{"sys":{"type":"Entry","revision":1,"contentType":{"sys":{"type":"Link","linkType":"ContentType","id":"richTextSection"}}},"id":"4b5a9822-0091-599f-b798-379d92ff98ab","anchor":null,"content":{"json":{"nodeType":"document","data":{},"content":[{"nodeType":"heading-6","data":{},"content":[{"nodeType":"text","value":"August 4, 2026","marks":[],"data":{}}]},{"nodeType":"heading-2","data":{},"content":[{"nodeType":"text","value":"How We Accelerate ML Deployment by Empowering Data Scientists (Part I)","marks":[{"type":"bold"}],"data":{}}]},{"nodeType":"heading-4","data":{},"content":[{"nodeType":"text","value":"Orchestration of ML systems at scale, the easy way","marks":[{"type":"italic"}],"data":{}}]},{"nodeType":"heading-6","data":{},"content":[{"nodeType":"text","value":"Written by Jordan Melendez","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"This article builds on ","marks":[],"data":{}},{"nodeType":"hyperlink","data":{"uri":"https://www.joinroot.com/blog/how-we-accelerate-ml-deployment-by-empowering-data-scientists-part-i/"},"content":[{"nodeType":"text","value":"Part 1: Building a Reproducible Data Science Environment","marks":[],"data":{}}]},{"nodeType":"text","value":".","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"In Part 1, we focused on the foundations of reproducible machine learning: packaging Python code, creating deterministic environments with uv, and standardizing OS-level development with Docker and dev containers.","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"The choices made in Part 1 exist to make the next stage dramatically simpler. Once structure and environment are in place, the remaining challenge is executing reproducible ML workflows from local experimentation all the way through production.","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"That’s where orchestration, configuration, and workflow tooling enter the picture.","marks":[],"data":{}}]},{"nodeType":"heading-5","data":{},"content":[{"nodeType":"text","value":"Running Data Science Code","marks":[{"type":"bold"}],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"The tooling we’ve integrated so far solves reproducibility and developer experience. The next question is how those same projects should actually be executed. Once your project has structure and an environment that permits rapid iteration loops, you are ready to begin actually running experiments!","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"For data scientists and MLEs, Initial iteration may begin in SQL, notebooks, scripts, or BI tools. Eventually you may end up with a series of notebooks or scripts that looks something like","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":".\n├── 00_start.py\n├── 01_preprocess.py\n├── 01_preprocess_final.py\n├── 02_train.py\n├── 03a_evaluate.py\n├── 03b_evaluate.py\n└── 04_cleanup.py","marks":[{"type":"code"}],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"It is then up to the user to ensure that each is run in sequence, with intermediate artifacts propagated from step to step. A colleague, or yourself from the future, may wonder which files are outdated, and whether ","marks":[],"data":{}},{"nodeType":"text","value":"01_preprocess.py","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" should be run at all. Good documentation can help, but can go out of date. Furthermore, some of these scripts may have hardcoded values of hyperparameters, data input/output locations, etc. (Hopefully, these scripts have already moved repeated logic to the core package to keep reusable functions clean!)","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"There are now two concepts to introduce to help clean this up: the ","marks":[],"data":{}},{"nodeType":"text","value":"specification","marks":[{"type":"italic"}],"data":{}},{"nodeType":"text","value":" layer, and the ","marks":[],"data":{}},{"nodeType":"text","value":"orchestration","marks":[{"type":"italic"}],"data":{}},{"nodeType":"text","value":" layer. These layers exist to complement the data science layer that was introduced as the code packaged in core. These are explained in detail in ","marks":[],"data":{}},{"nodeType":"hyperlink","data":{"uri":"https://medium.com/data-science-at-microsoft/a-layered-approach-to-mlops-d935beefca2e"},"content":[{"nodeType":"text","value":"this excellent article","marks":[],"data":{}}]},{"nodeType":"text","value":". They can be summarized as follows","marks":[],"data":{}}]},{"nodeType":"ordered-list","data":{},"content":[{"nodeType":"list-item","data":{},"content":[{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"Data science layer:","marks":[{"type":"bold"}],"data":{}},{"nodeType":"text","value":" Describes ","marks":[],"data":{}},{"nodeType":"text","value":"what","marks":[{"type":"italic"}],"data":{}},{"nodeType":"text","value":" is done. It contains the core data-science code specific to the problem we are trying to solve.","marks":[],"data":{}}]}]},{"nodeType":"list-item","data":{},"content":[{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"Specification layer:","marks":[{"type":"bold"}],"data":{}},{"nodeType":"text","value":" Describes ","marks":[],"data":{}},{"nodeType":"text","value":"how","marks":[{"type":"italic"}],"data":{}},{"nodeType":"text","value":" it is done. This includes, for example, the parameters required to run the code (where data lives, what hyperparameters to use, which estimators, etc.). It could also specify where jobs are run, instance sizing, the execution environment, and so on. We’ve already discussed components of the execution environment, such as the ","marks":[],"data":{}},{"nodeType":"text","value":"uv.lock ","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":"and  ","marks":[],"data":{}},{"nodeType":"text","value":"Dockerfile","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":". The specification layer only depends on the data-science layer.","marks":[],"data":{}}]}]},{"nodeType":"list-item","data":{},"content":[{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"Orchestration layer: ","marks":[{"type":"bold"}],"data":{}},{"nodeType":"text","value":"Executes the code defined in the data-science and specification layers. It depends on both the data-science and specification layer.","marks":[],"data":{}}]}]}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"Importantly, the outer layers only depend on the inner layers, ","marks":[],"data":{}},{"nodeType":"text","value":"not the other way around","marks":[{"type":"italic"}],"data":{}},{"nodeType":"text","value":". That is, configuration and orchestration logic should not creep into your data-science layer. A preprocessing function may accept a dataframe that was already loaded from a particular location, or an estimator may be trained with particular hyperparameters, but this information comes from the specification layer, it is not hardcoded in your functions.","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"Let’s discuss these two new layers, and describe what new workflows they enable in the ML lifecycle. I will wrap up by describing how command runners, such as ","marks":[],"data":{}},{"nodeType":"text","value":"just","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":", help in this ecosystem.","marks":[],"data":{}}]},{"nodeType":"heading-5","data":{},"content":[{"nodeType":"text","value":"The Specification Layer","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"I rely primarily on YAML files, combined with ","marks":[],"data":{}},{"nodeType":"text","value":"pydantic","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":", as the backbone of my specification layer. (Tools like ","marks":[],"data":{}},{"nodeType":"text","value":"OmegaConf","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" exist for more advanced use cases, which generalize nicely from what is described here.)","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"The configuration for a training run lives in two files. First, a typed Pydantic model that declares the shape of the config and validates it at load time:","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"# src/core/config.py\nfrom pydantic import BaseModel\n\nclass TrainingConfig(BaseModel):\n\"\"\"Parameters that govern a training run.\"\"\"\n    n_estimators: int\n    max_depth: int\n    random_seed: int","marks":[{"type":"code"}],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"Second, a YAML file that holds the actual values:","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"# config/training.yaml\nn_estimators: 100\nmax_depth: 6\nrandom_seed: 42","marks":[{"type":"code"}],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"Why this combination? A plain dict of hyperparameters works but has no type validation. A new colleague (or yourself from the future) will not get immediate feedback if they use a string in a place that should accept a float. A Pydantic model makes the contract explicit, where typos fail at validation, and every field is typed and self-documenting.","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"The YAML file is equally important. Because it is plain text, a change from ","marks":[],"data":{}},{"nodeType":"text","value":"max_depth: 6","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" to ","marks":[],"data":{}},{"nodeType":"text","value":"max_depth: 8","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" is a one-line diff in version control. It appears once in a centrally defined location, not in the middle of a script. Colleagues can easily find what levers are exposed and how to change them, with confidence that the changes will propagate to all downstream code. The YAML files can hold more than hyperparameters, and can be extended to data path locations, flags for where code should run, and even the estimators themselves. Advanced ","marks":[],"data":{}},{"nodeType":"hyperlink","data":{"uri":"https://hydra.cc/"},"content":[{"nodeType":"text","value":"configuration-driven development","marks":[],"data":{}}]},{"nodeType":"text","value":" patterns can follow once the code is split up in this way.","marks":[],"data":{}}]},{"nodeType":"heading-5","data":{},"content":[{"nodeType":"text","value":"The Orchestration Layer","marks":[{"type":"bold"}],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"For the orchestration layer, I have found that Metaflow provides the best developer experience for data science workflows, and the competition is not even close. Key aspects of an effective orchestrator include","marks":[],"data":{}}]},{"nodeType":"ordered-list","data":{},"content":[{"nodeType":"list-item","data":{},"content":[{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"The ability to easily run locally or on the cloud without code changes","marks":[],"data":{}}]}]},{"nodeType":"list-item","data":{},"content":[{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"Enable fast iteration loops for data scientists, including retrying failed jobs, debugging issues, and so on","marks":[],"data":{}}]}]},{"nodeType":"list-item","data":{},"content":[{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"Integrate with how data scientists naturally work and think, without forcing us into constrained patterns to fit a framework","marks":[],"data":{}}]}]}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"Metaflow has all of these properties and more. Let’s explore with a simple example.","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"A Metaflow flow is a Python class that describes a directed graph of steps. An example training flow for this project is shown below:","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"# flows/training_flow.py\nfrom metaflow import FlowSpec, step, Config\nfrom core import features, models\nfrom core.config import TrainingConfig, pydantic_parser\n\n\nclass TrainingFlow(FlowSpec):\n    \"\"\"End-to-end training pipeline.\n\n\n    Run locally:\n        python flows/training_flow.py run\n\n\n    Run on AWS Batch:\n        python flows/training_flow.py run --with batch\n    \"\"\"\n\n\n    config: \"TrainingConfig\" = Config(\n        name=\"config\",\n        default=\"../config/training.yaml\",\n        parser=pydantic_parser(TrainingConfig),\n    )\n\n\n    @step\n    def start(self):\n        print(f\"Starting training run with config: {self.config}\")\n        self.next(self.featurize)\n\n\n    @step\n    def featurize(self):\n        import polars as pl\n\n\n        raw_data: pl.DataFrame = pl.DataFrame(...)  # load your data here\n        self.feature_data = features.transform(raw_data)\n        self.next(self.train)\n\n\n    @step\n    def train(self):\n        X = self.feature_data.drop(\"y\")\n        y = self.feature_data[\"y\"]\n        params = {\n            \"n_estimators\": self.config.n_estimators,\n            \"max_depth\": self.config.max_depth,\n            \"random_state\": self.config.random_seed,\n        }\n        self.model = models.train(X, y, params)\n        self.next(self.evaluate)\n\n\n    @step\n    def evaluate(self):\n        X = self.feature_data.drop(\"y\")\n        y = self.feature_data[\"y\"]\n        self.metrics = models.evaluate(self.model, X, y)\n        self.next(self.end)\n\n\n    @step\n    def end(self):\n        print(f\"Config: {self.config}\")\n        print(f\"Metrics: {self.metrics}\")\n\n\n\n\nif __name__ == \"__main__\":\n    TrainingFlow()","marks":[{"type":"code"}],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"A few things worth noting here.","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"The flow is lightweight.","marks":[{"type":"bold"}],"data":{}},{"nodeType":"text","value":" It imports from ","marks":[],"data":{}},{"nodeType":"text","value":"core","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" and accepts configuration as a parameter. Its only job is to wire those pieces together and manage data flow between steps. There is little to no data science logic in the flow itself, these instead live in ","marks":[],"data":{}},{"nodeType":"text","value":"features.py","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" and ","marks":[],"data":{}},{"nodeType":"text","value":"models.py","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":". Edits to ","marks":[],"data":{}},{"nodeType":"text","value":"core","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" immediately impact how the flow runs. Multiple different Metaflow flows can reuse the same logic stored in ","marks":[],"data":{}},{"nodeType":"text","value":"core","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" without copying code. This respects the separation between the ","marks":[],"data":{}},{"nodeType":"text","value":"data-science layer","marks":[{"type":"italic"}],"data":{}},{"nodeType":"text","value":" and the ","marks":[],"data":{}},{"nodeType":"text","value":"orchestration layer","marks":[{"type":"italic"}],"data":{}},{"nodeType":"text","value":".","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"Metaflow natively supports uv environments. By using the same base docker image and ","marks":[],"data":{}},{"nodeType":"text","value":"uv.lock","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" file in Metaflow as you use for local development, Metaflow will work identically for both local and cloud jobs. This can be enabled via the environment variable ","marks":[],"data":{}},{"nodeType":"text","value":"METAFLOW_ENVIRONMENT=\"uv\"","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" or by running your jobs with the","marks":[],"data":{}},{"nodeType":"text","value":"--environment uv","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" flag, as in ","marks":[],"data":{}},{"nodeType":"text","value":"uv run flows/training_flow.py --environment=uv run","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":".","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"There is one nuance here around mixing your core package with Metaflow: you need to symlink your ","marks":[],"data":{}},{"nodeType":"text","value":"src","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" directory within the same directory that your flows live so that Metaflow can pick it up. You can achieve this via ","marks":[],"data":{}},{"nodeType":"text","value":"ln -s ../src flows/src","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":". This allows you to retain all the benefits of packaged code but without explicitly installing the package into the image that Metaflow uses. Metaflow recently open-sourced the ","marks":[],"data":{}},{"nodeType":"text","value":"@package_sources","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" ","marks":[],"data":{}},{"nodeType":"hyperlink","data":{"uri":"https://github.com/Netflix/metaflow/pull/3351"},"content":[{"nodeType":"text","value":"decorator","marks":[],"data":{}}]},{"nodeType":"text","value":" exactly for this use case, and so could help reduce the need for symlinks going forward.","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"Metaflow gives you data lineage for free.","marks":[{"type":"bold"}],"data":{}},{"nodeType":"text","value":" Metaflow artifacts, such as ","marks":[],"data":{}},{"nodeType":"text","value":"self.feature_data","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" and ","marks":[],"data":{}},{"nodeType":"text","value":"self.model","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":", are automatically serialized and stored at each step boundary. Reproducing a result from three months ago means finding the run ID and loading its artifacts, rather than tracking down a pickle file on someone’s laptop.","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"Configuration is separate from orchestration logic.","marks":[{"type":"bold"}],"data":{}},{"nodeType":"text","value":" The ","marks":[],"data":{}},{"nodeType":"text","value":"Config","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" object is Metaflow’s mechanism for loading structured configuration. It reads ","marks":[],"data":{}},{"nodeType":"text","value":"config/training.yaml","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":", parses it through the Pydantic validator, and makes the typed ","marks":[],"data":{}},{"nodeType":"text","value":"TrainingConfig","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" available as ","marks":[],"data":{}},{"nodeType":"text","value":"self.config","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" throughout the run. The config is also automatically persisted as a run artifact, so every historical run has its parameters stored alongside its outputs.","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"Metaflow can run the same locally as in the cloud. The same flow runs locally with ","marks":[],"data":{}},{"nodeType":"text","value":"python flows/training_flow.py run","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" and on AWS Batch with ","marks":[],"data":{}},{"nodeType":"text","value":"python flows/training_flow.py run --with batch","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":". This ships your jobs to the cloud without any changes to your actual code. Outside of AWS Batch, Metaflow also supports running on K8s, Airflow, and Kubeflow. With ephemeral cloud compute, it now becomes trivial to fan out to 1000s of jobs simultaneously training different models, something that would be impossible if limited to your local machine.","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"Although not Metaflow specific, it’s worth calling out a pattern when running jobs locally and in the cloud: try to keep data paths in config, not in code, and choose I/O libraries that treat local paths and cloud URIs identically. Polars, which this project uses, handles this naturally: ","marks":[],"data":{}},{"nodeType":"text","value":"pl.read_parquet(“data/local.parquet”)","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" and ","marks":[],"data":{}},{"nodeType":"text","value":"pl.read_parquet(“s3://my-bucket/data.parquet”)","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" are the same call, just a different string. The same applies to ","marks":[],"data":{}},{"nodeType":"hyperlink","data":{"uri":"https://github.com/piskvorky/smart_open"},"content":[{"nodeType":"text","value":"smart-open","marks":[],"data":{}}]},{"nodeType":"text","value":" as a drop-in for Python’s ","marks":[],"data":{}},{"nodeType":"text","value":"open()","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":", and to any library built on fsspec.","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"If you’re interested in learning more about Metaflow, I recommend their ","marks":[],"data":{}},{"nodeType":"hyperlink","data":{"uri":"https://docs.metaflow.org/"},"content":[{"nodeType":"text","value":"documentation","marks":[],"data":{}}]},{"nodeType":"text","value":" and, if you’re interested in a deep dive on the design philosophy with concrete examples, see ","marks":[],"data":{}},{"nodeType":"hyperlink","data":{"uri":"https://www.oreilly.com/library/view/effective-data-science/9781617299193/"},"content":[{"nodeType":"text","value":"Effective Data Science Infrastructure","marks":[],"data":{}}]},{"nodeType":"text","value":" written by Ville Tuulos, who designed and built Metaflow.","marks":[],"data":{}}]},{"nodeType":"heading-5","data":{},"content":[{"nodeType":"text","value":"Command runners","marks":[{"type":"bold"}],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"Every project accumulates a vocabulary of commands. For example, how to install dependencies, run tests, lint, format, or run a flow. This vocabulary usually lives in a README that goes stale, in contributors’ heads, or nowhere at all. And every team may work ","marks":[],"data":{}},{"nodeType":"text","value":"slightly","marks":[{"type":"italic"}],"data":{}},{"nodeType":"text","value":" differently, making context switching all the more painful. When onboarding a new team member, the first hour is often spent just figuring out which commands to run and in what order.","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"just","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" solves this. It is a modern command runner and replaces ","marks":[],"data":{}},{"nodeType":"text","value":"make","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" for most data-science workflows. The ","marks":[],"data":{}},{"nodeType":"text","value":"justfile","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" at the root of the project is the single source of truth for every workflow command:","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"sync:","marks":[{"type":"bold"},{"type":"code"}],"data":{}},{"nodeType":"text","value":"\n    uv sync\n\n\n","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":"test:","marks":[{"type":"bold"},{"type":"code"}],"data":{}},{"nodeType":"text","value":"\n    uv run pytest\n\n\ntest-one file:\n    uv run pytest {{file}}\n\n\n","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":"lint:","marks":[{"type":"bold"},{"type":"code"}],"data":{}},{"nodeType":"text","value":"\n    uv run ruff check .\n\n\n","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":"format:","marks":[{"type":"bold"},{"type":"code"}],"data":{}},{"nodeType":"text","value":"\n    uv run ruff format .\n\n\n","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":"typecheck:","marks":[{"type":"bold"},{"type":"code"}],"data":{}},{"nodeType":"text","value":"\n    uv run ty check\n\n\n","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":"check: lint typecheck","marks":[{"type":"bold"},{"type":"code"}],"data":{}},{"nodeType":"text","value":"\n\n\n","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":"train:","marks":[{"type":"bold"},{"type":"code"}],"data":{}},{"nodeType":"text","value":"\n    cd flows && uv run python training_flow.py --environment uv run","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":"\n","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"just --list","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" is the entry point for any new contributor. They do not need to read the README, grep through scripts, or ask a colleague. Running ","marks":[],"data":{}},{"nodeType":"text","value":"just check","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" before opening a PR is the same command CI runs.","marks":[],"data":{}}]},{"nodeType":"heading-5","data":{},"content":[{"nodeType":"text","value":"CI/CD","marks":[{"type":"bold"}],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"The setup outlined above lends itself naturally to CI/CD pipelines. With ","marks":[],"data":{}},{"nodeType":"text","value":"just","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" in place, the commands I run locally are exactly the commands CI runs. There is no separate CI-specific scripting layer to maintain.","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"This also keeps CI out of the fast path for day-to-day experimentation. A data scientist iterating on a model should not be waiting for a pipeline. CI runs on pull requests before anything merges, and the deploy workflow runs the flow upon merging to ","marks":[],"data":{}},{"nodeType":"text","value":"main","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":". It is a quality gate before production, not a bottleneck during research.","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"Because I have already separated the data science, specification, and orchestration layers, CI can do more than just lint and test. The orchestrated flow that runs locally for experimentation can be run and validated in a staging environment with models and metrics logged to a staging-specific experiment tracker or model registry.","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"By implementing these practices, moving from the world of “shipping models” towards “shipping model ","marks":[],"data":{}},{"nodeType":"text","value":"factories","marks":[{"type":"italic"}],"data":{}},{"nodeType":"text","value":"” is close to flipping a switch.","marks":[],"data":{}}]},{"nodeType":"heading-5","data":{},"content":[{"nodeType":"text","value":"Summary","marks":[{"type":"bold"}],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"These series of posts have covered three layers of a reproducible ML project: structure, environment, and tooling.","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"For structure, packaging your data science code with ","marks":[],"data":{}},{"nodeType":"text","value":"uv","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" and the ","marks":[],"data":{}},{"nodeType":"text","value":"src/","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" layout makes it importable anywhere, just as easily as any other package. A separate specification layer, with corresponding Pydantic models, provides a typed, ","marks":[],"data":{}},{"nodeType":"text","value":"diff","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":"-able home in version control instead of scattering magic numbers through scripts.","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"For the environment, uv pins the Python dependency graph to a lockfile that travels across laptops, CI runners, and cloud batch jobs. Dev containers extend that guarantee to the OS layer, so a new contributor goes from git clone to a running training flow without a setup guide.","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"For tooling, Metaflow ties it together as the orchestration layer: a thin flow that imports from core, reads config, and runs identically locally or on the cloud with a single flag. The ","marks":[],"data":{}},{"nodeType":"text","value":"just","marks":[{"type":"code"}],"data":{}},{"nodeType":"text","value":" command runner gives every workflow a single command, and CI runs those same commands before anything ships.","marks":[],"data":{}}]},{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"None of these decisions are large in isolation. Together, they make the right way the easy way. This leads to faster onboarding, experimentation, and reproducibility, each of which set you up for scaling to the cloud and to production.","marks":[],"data":{}}]}]}},"textAlignment":"Left","backgroundColor":"White"},{"sys":{"type":"Entry","revision":1,"contentType":{"sys":{"type":"Link","linkType":"ContentType","id":"iconToutSection"}}},"id":"807b202e-9742-5150-9f4a-341186191ac0","anchor":null,"eyebrow":null,"headline":null,"subheadRich":null,"iconSize":"Large","backgroundColor":"White #FFFFFF","contentfulchildren":[{"id":"5bb62ccf-a607-52ab-b25e-e253e28b4a7b","eyebrow":null,"headline":"Want to work with us?","body":{"json":{"nodeType":"document","data":{},"content":[{"nodeType":"paragraph","data":{},"content":[{"nodeType":"text","value":"Join our team of data scientists, actuaries, machine learning engineers, and analysts.","marks":[],"data":{}}]}]}},"asset":{"title":"Root App Badge - [Make]","file":{"url":"//images.ctfassets.net/pacigpl3aj13/6P5qfFeQ6lHMbxBgkstGNd/a077b1e3238f23061cd719073faeb5b5/RootBadge.svg"}},"iconAltText":null,"callToActionButton":{"id":"0f87a3a0-ec2e-56ac-ae83-868aefe080b3","buttonType":"Primary","buttonText":"Explore Careers","buttonLink":"https://inc.joinroot.com/careers/","forwardQueryParams":null,"gTagEventName":null,"modal":null},"form":null}]},{"sys":{"type":"Entry","revision":1,"contentType":{"sys":{"type":"Link","linkType":"ContentType","id":"blogRelatedPostsSection"}}},"id":"bc4ba608-d7b9-5058-ab22-0eacc44826f1","relatedPosts":[{"publishingMetadata":{"slug":"/how-we-accelerate-ml-deployment-by-empowering-data-scientists-part-i"},"blogCategory":[{"categoryTitle":"Data Science","publishingMetadata":{"slug":"/data-science"}}],"customDateCreated":"August 04, 2026","previewTitle":"How We Accelerate ML Deployment by Empowering Data Scientists (Part I)","previewBody":"Explore how Root empowers data scientists to move machine learning projects from experimentation to production faster with reproducible development environments, standardized tooling, and infrastructure designed to make best practices the easiest path.","imageAltText":null,"previewImage":{"gatsbyImageData":{"images":{"sources":[{"srcSet":"//images.ctfassets.net/pacigpl3aj13/M04NhqSa92MHPj1TGZ3Kv/cc8bcbfa2e16bfbe069a29442e171e00/1_n7QvQz3vOVp6ciOI3lvldw.png?w=444&h=222&q=60&fm=webp 444w,\n//images.ctfassets.net/pacigpl3aj13/M04NhqSa92MHPj1TGZ3Kv/cc8bcbfa2e16bfbe069a29442e171e00/1_n7QvQz3vOVp6ciOI3lvldw.png?w=887&h=444&q=60&fm=webp 887w,\n//images.ctfassets.net/pacigpl3aj13/M04NhqSa92MHPj1TGZ3Kv/cc8bcbfa2e16bfbe069a29442e171e00/1_n7QvQz3vOVp6ciOI3lvldw.png?w=1774&h=887&q=60&fm=webp 1774w","sizes":"(min-width: 1774px) 1774px, 100vw","type":"image/webp"}],"fallback":{"src":"//images.ctfassets.net/pacigpl3aj13/M04NhqSa92MHPj1TGZ3Kv/cc8bcbfa2e16bfbe069a29442e171e00/1_n7QvQz3vOVp6ciOI3lvldw.png?w=1774&h=887&q=60&fm=png","srcSet":"//images.ctfassets.net/pacigpl3aj13/M04NhqSa92MHPj1TGZ3Kv/cc8bcbfa2e16bfbe069a29442e171e00/1_n7QvQz3vOVp6ciOI3lvldw.png?w=444&h=222&q=60&fm=png 444w,\n//images.ctfassets.net/pacigpl3aj13/M04NhqSa92MHPj1TGZ3Kv/cc8bcbfa2e16bfbe069a29442e171e00/1_n7QvQz3vOVp6ciOI3lvldw.png?w=887&h=444&q=60&fm=png 887w,\n//images.ctfassets.net/pacigpl3aj13/M04NhqSa92MHPj1TGZ3Kv/cc8bcbfa2e16bfbe069a29442e171e00/1_n7QvQz3vOVp6ciOI3lvldw.png?w=1774&h=887&q=60&fm=png 1774w","sizes":"(min-width: 1774px) 1774px, 100vw"}},"layout":"constrained","width":1774,"height":887,"placeholder":{"fallback":"data:image/jpeg;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAAKCAMAAACDi47UAAACN1BMVEUAAAcAAAAAACoAEUsNGyofAAAuCAAIAAAAFSgDJUEAAAYIAA8AAAsUBQgJAQAAADcALGgyVZRxntJ2pNQvaKYAAB4zEWBVNnM6JU4IACQPHDsAABEEAAAGAAAqCAIZBAwnDAwDAAA6FC5FVpGVve10puNTdraVjpqXflYmAHarctaaeKtaUGwvI1gbNXJVbqFcb58cLF8gAAAwEBMrCxgFAACLSACpaEnRtKm1oJjntDHzuBC+ijxrK47DfN2ZjJuYm5VGQGdIdL6Tt++gu/Fsh8dAQVokHjFdJDddJR4hCB49GxttQwCziTzcoACwfRhGJRg5CTCLTJSDXo9kUm4tEzwnOWtegb58ndZadKtiZG9ma3ZlRVWZMyEzF0wsM18YGVM7N1ePiJKMipFVV2EiJCsiGmBLcKhFca0nNmAhBwhsNkyefYBnU2MNL2ddJkvWbSiNXZRpOYWOZ622p8rZ2N+0tbTS0c9paGsAHmNTjdFtq/A+d9Ypb8wpRHrBalP3ysHFoqNeHlm1YGf70J86Ik9EIWVnTHuNgpfAvsO4t7bIx8VOTE8IIFdQhcF3p+VUeNgAU8NRJ1zZh3ftvbLMpqyqT4bzp6H31rUGG0o+T3wqO2VHRU17eoNsbHMVCkIlK2cHAD4+UHBGXYwmLWNOKC6DREu5hYOegouJPFq1V0eKPRcnBBdLIS0tBhUFAB4AABUAADAOF0kcADQeAEYvR3UfMVk7Qm4AJVQqAAAoAABGGgQBAACNMYzRAAAAPUlEQVQI12NkYASCvwwsjL8Z2H6xM4IBCz+YeifAiAyUGDEBiwjJgnNSkUQLSdGeNhtddA2ICGTYgCTkCgBt+gnEBqsLLwAAAABJRU5ErkJggg=="}}}},{"publishingMetadata":{"slug":"/the-data-science-research-production-chasm"},"blogCategory":[{"categoryTitle":"Data Science","publishingMetadata":{"slug":"/data-science"}}],"customDateCreated":"September 08, 2021","previewTitle":"The Data Science Research-Production Chasm","previewBody":"Explores the challenges that can emerge between data science research and production and practical strategies for building more reliable, scalable machine learning pipelines.","imageAltText":"Image shows chart","previewImage":{"gatsbyImageData":{"images":{"sources":[{"srcSet":"//images.ctfassets.net/pacigpl3aj13/nKJTfBvVPfS2YziAXDxhR/351e40deed4f462b589dbe9ed6941ae1/Screenshot_2026-09-22_at_11.33.05.png?w=610&h=270&q=60&fm=webp 610w,\n//images.ctfassets.net/pacigpl3aj13/nKJTfBvVPfS2YziAXDxhR/351e40deed4f462b589dbe9ed6941ae1/Screenshot_2026-09-22_at_11.33.05.png?w=1220&h=539&q=60&fm=webp 1220w,\n//images.ctfassets.net/pacigpl3aj13/nKJTfBvVPfS2YziAXDxhR/351e40deed4f462b589dbe9ed6941ae1/Screenshot_2026-09-22_at_11.33.05.png?w=2440&h=1078&q=60&fm=webp 2440w","sizes":"(min-width: 2440px) 2440px, 100vw","type":"image/webp"}],"fallback":{"src":"//images.ctfassets.net/pacigpl3aj13/nKJTfBvVPfS2YziAXDxhR/351e40deed4f462b589dbe9ed6941ae1/Screenshot_2026-09-22_at_11.33.05.png?w=2440&h=1078&q=60&fm=png","srcSet":"//images.ctfassets.net/pacigpl3aj13/nKJTfBvVPfS2YziAXDxhR/351e40deed4f462b589dbe9ed6941ae1/Screenshot_2026-09-22_at_11.33.05.png?w=610&h=270&q=60&fm=png 610w,\n//images.ctfassets.net/pacigpl3aj13/nKJTfBvVPfS2YziAXDxhR/351e40deed4f462b589dbe9ed6941ae1/Screenshot_2026-09-22_at_11.33.05.png?w=1220&h=539&q=60&fm=png 1220w,\n//images.ctfassets.net/pacigpl3aj13/nKJTfBvVPfS2YziAXDxhR/351e40deed4f462b589dbe9ed6941ae1/Screenshot_2026-09-22_at_11.33.05.png?w=2440&h=1078&q=60&fm=png 2440w","sizes":"(min-width: 2440px) 2440px, 100vw"}},"layout":"constrained","width":2440,"height":1078,"placeholder":{"fallback":"data:image/jpeg;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAAJCAMAAAAFH/x6AAAABGdBTUEAALGPC/xhBQAACilpQ0NQaWNjAABIiZ2Wd1RT2RaHz703vVCSEIqU0GtoUgJIDb1IkS4qMQkQSsCQACI2RFRwRFGRpggyKOCAo0ORsSKKhQFRsesEGUTUcXAUG5ZJZK0Z37x5782b3x/3fmufvc/dZ+991roAkPyDBcJMWAmADKFYFOHnxYiNi2dgBwEM8AADbADgcLOzQhb4RgKZAnzYjGyZE/gXvboOIPn7KtM/jMEA/5+UuVkiMQBQmIzn8vjZXBkXyTg9V5wlt0/JmLY0Tc4wSs4iWYIyVpNz8ixbfPaZZQ858zKEPBnLc87iZfDk3CfjjTkSvoyRYBkX5wj4uTK+JmODdEmGQMZv5LEZfE42ACiS3C7mc1NkbC1jkigygi3jeQDgSMlf8NIvWMzPE8sPxc7MWi4SJKeIGSZcU4aNkxOL4c/PTeeLxcwwDjeNI+Ix2JkZWRzhcgBmz/xZFHltGbIiO9g4OTgwbS1tvijUf138m5L3dpZehH/uGUQf+MP2V36ZDQCwpmW12fqHbWkVAF3rAVC7/YfNYC8AirK+dQ59cR66fF5SxOIsZyur3NxcSwGfaykv6O/6nw5/Q198z1K+3e/lYXjzkziSdDFDXjduZnqmRMTIzuJw+Qzmn4f4Hwf+dR4WEfwkvogvlEVEy6ZMIEyWtVvIE4gFmUKGQPifmvgPw/6k2bmWidr4EdCWWAKlIRpAfh4AKCoRIAl7ZCvQ730LxkcD+c2L0ZmYnfvPgv59V7hM/sgWJH+OY0dEMrgSUc7smvxaAjQgAEVAA+pAG+gDE8AEtsARuAAP4AMCQSiIBHFgMeCCFJABRCAXFIC1oBiUgq1gJ6gGdaARNIM2cBh0gWPgNDgHLoHLYATcAVIwDp6AKfAKzEAQhIXIEBVSh3QgQ8gcsoVYkBvkAwVDEVAclAglQ0JIAhVA66BSqByqhuqhZuhb6Ch0GroADUO3oFFoEvoVegcjMAmmwVqwEWwFs2BPOAiOhBfByfAyOB8ugrfAlXADfBDuhE/Dl+ARWAo/gacRgBAROqKLMBEWwkZCkXgkCREhq5ASpAJpQNqQHqQfuYpIkafIWxQGRUUxUEyUC8ofFYXiopahVqE2o6pRB1CdqD7UVdQoagr1EU1Ga6LN0c7oAHQsOhmdiy5GV6Cb0B3os+gR9Dj6FQaDoWOMMY4Yf0wcJhWzArMZsxvTjjmFGcaMYaaxWKw61hzrig3FcrBibDG2CnsQexJ7BTuOfYMj4nRwtjhfXDxOiCvEVeBacCdwV3ATuBm8Et4Q74wPxfPwy/Fl+EZ8D34IP46fISgTjAmuhEhCKmEtoZLQRjhLuEt4QSQS9YhOxHCigLiGWEk8RDxPHCW+JVFIZiQ2KYEkIW0h7SedIt0ivSCTyUZkD3I8WUzeQm4mnyHfJ79RoCpYKgQo8BRWK9QodCpcUXimiFc0VPRUXKyYr1iheERxSPGpEl7JSImtxFFapVSjdFTphtK0MlXZRjlUOUN5s3KL8gXlRxQsxYjiQ+FRiij7KGcoY1SEqk9lU7nUddRG6lnqOA1DM6YF0FJppbRvaIO0KRWKip1KtEqeSo3KcRUpHaEb0QPo6fQy+mH6dfo7VS1VT1W+6ibVNtUrqq/V5qh5qPHVStTa1UbU3qkz1H3U09S3qXep39NAaZhphGvkauzROKvxdA5tjssc7pySOYfn3NaENc00IzRXaO7THNCc1tLW8tPK0qrSOqP1VJuu7aGdqr1D+4T2pA5Vx01HoLND56TOY4YKw5ORzqhk9DGmdDV1/XUluvW6g7ozesZ6UXqFeu169/QJ+iz9JP0d+r36UwY6BiEGBQatBrcN8YYswxTDXYb9hq+NjI1ijDYYdRk9MlYzDjDON241vmtCNnE3WWbSYHLNFGPKMk0z3W162Qw2szdLMasxGzKHzR3MBea7zYct0BZOFkKLBosbTBLTk5nDbGWOWtItgy0LLbssn1kZWMVbbbPqt/pobW+dbt1ofceGYhNoU2jTY/OrrZkt17bG9tpc8lzfuavnds99bmdux7fbY3fTnmofYr/Bvtf+g4Ojg8ihzWHS0cAx0bHW8QaLxgpjbWadd0I7eTmtdjrm9NbZwVnsfNj5FxemS5pLi8ujecbz+PMa54256rlyXOtdpW4Mt0S3vW5Sd113jnuD+wMPfQ+eR5PHhKepZ6rnQc9nXtZeIq8Or9dsZ/ZK9ilvxNvPu8R70IfiE+VT7XPfV8832bfVd8rP3m+F3yl/tH+Q/zb/GwFaAdyA5oCpQMfAlYF9QaSgBUHVQQ+CzYJFwT0hcEhgyPaQu/MN5wvnd4WC0IDQ7aH3wozDloV9H44JDwuvCX8YYRNRENG/gLpgyYKWBa8ivSLLIu9EmURJonqjFaMTopujX8d4x5THSGOtYlfGXorTiBPEdcdj46Pjm+KnF/os3LlwPME+oTjh+iLjRXmLLizWWJy++PgSxSWcJUcS0YkxiS2J7zmhnAbO9NKApbVLp7hs7i7uE54Hbwdvku/KL+dPJLkmlSc9SnZN3p48meKeUpHyVMAWVAuep/qn1qW+TgtN25/2KT0mvT0Dl5GYcVRIEaYJ+zK1M/Myh7PMs4qzpMucl+1cNiUKEjVlQ9mLsrvFNNnP1IDERLJeMprjllOT8yY3OvdInnKeMG9gudnyTcsn8n3zv16BWsFd0VugW7C2YHSl58r6VdCqpat6V+uvLlo9vsZvzYG1hLVpa38otC4sL3y5LmZdT5FW0ZqisfV+61uLFYpFxTc2uGyo24jaKNg4uGnupqpNH0t4JRdLrUsrSt9v5m6++JXNV5VffdqStGWwzKFsz1bMVuHW69vctx0oVy7PLx/bHrK9cwdjR8mOlzuX7LxQYVdRt4uwS7JLWhlc2V1lULW16n11SvVIjVdNe61m7aba17t5u6/s8djTVqdVV1r3bq9g7816v/rOBqOGin2YfTn7HjZGN/Z/zfq6uUmjqbTpw37hfumBiAN9zY7NzS2aLWWtcKukdfJgwsHL33h/093GbKtvp7eXHgKHJIcef5v47fXDQYd7j7COtH1n+F1tB7WjpBPqXN451ZXSJe2O6x4+Gni0t8elp+N7y+/3H9M9VnNc5XjZCcKJohOfTuafnD6Vderp6eTTY71Leu+ciT1zrS+8b/Bs0Nnz53zPnen37D953vX8sQvOF45eZF3suuRwqXPAfqDjB/sfOgYdBjuHHIe6Lztd7hmeN3ziivuV01e9r567FnDt0sj8keHrUddv3ki4Ib3Ju/noVvqt57dzbs/cWXMXfbfkntK9ivua9xt+NP2xXeogPT7qPTrwYMGDO2PcsSc/Zf/0frzoIflhxYTORPMj20fHJn0nLz9e+Hj8SdaTmafFPyv/XPvM5Nl3v3j8MjAVOzX+XPT806+bX6i/2P/S7mXvdNj0/VcZr2Zel7xRf3PgLett/7uYdxMzue+x7ys/mH7o+Rj08e6njE+ffgP3hPP78QcZjQAAACBjSFJNAAB6JgAAgIQAAPoAAACA6AAAdTAAAOpgAAA6mAAAF3CculE8AAAAn1BMVEX////x8/Xn7PD19/nk6/HO2+bq7/P///78/f38/f73+vv5+/z+/v7m7PHP2+Ts8PT19vbM2+Xg6fDg6vC5zdzJ2uft8vbK2uXY5Oz6+vrx8fH39/fy9vjn7fL4+PjZ5Ozo7vPo7/TN3ejX5O709/nf6fHn7/T8/Pzr8PTa4+nv8vT8/Pvm5eXu7e3a4+rw9Pf09PT19fXs6+v5+fj29vbwtTprAAAAU0lEQVQI12NgwAEYmZjRBICYhREIfnAyMn7l+QIX5GVEAU8ZWICC/IzPpGEiz6SfQlRqgLmXEWYygQhW1qss75AsAmkHWsFwAdl2kMpvTDdRnQQABnwMkhDDPBAAAAA4dEVYdGljYzpjb3B5cmlnaHQAQ29weXJpZ2h0IChjKSAxOTk4IEhld2xldHQtUGFja2FyZCBDb21wYW55+Vd5NwAAACF0RVh0aWNjOmRlc2NyaXB0aW9uAHNSR0IgSUVDNjE5NjYtMi4xV63aRwAAACZ0RVh0aWNjOm1hbnVmYWN0dXJlcgBJRUMgaHR0cDovL3d3dy5pZWMuY2gcfwBMAAAAN3RFWHRpY2M6bW9kZWwASUVDIDYxOTY2LTIuMSBEZWZhdWx0IFJHQiBjb2xvdXIgc3BhY2UgLSBzUkdCRFNIqQAAAABJRU5ErkJggg=="}}}},{"publishingMetadata":{"slug":"/how-root-rebuilt-its-telematics-platform-for-scale"},"blogCategory":[{"categoryTitle":"Data Science","publishingMetadata":{"slug":"/data-science"}}],"customDateCreated":"September 11, 2026","previewTitle":"How Root Rebuilt Its Telematics Platform for Scale","previewBody":"Learn how Root rebuilt its telematics platform to handle growing scale, improve reliability, and create a more flexible foundation for processing the driving data that powers its insurance products.","imageAltText":"Created by the author using ChatGPT","previewImage":{"gatsbyImageData":{"images":{"sources":[{"srcSet":"//images.ctfassets.net/pacigpl3aj13/6khgGxb7AecT8FWoasAG3a/093b9475639d49b489b670865a6111bc/How_Root_Rebuilt_Its_Telematics_Platform_for_Scale.png?w=350&h=197&q=60&fm=webp 350w,\n//images.ctfassets.net/pacigpl3aj13/6khgGxb7AecT8FWoasAG3a/093b9475639d49b489b670865a6111bc/How_Root_Rebuilt_Its_Telematics_Platform_for_Scale.png?w=700&h=394&q=60&fm=webp 700w,\n//images.ctfassets.net/pacigpl3aj13/6khgGxb7AecT8FWoasAG3a/093b9475639d49b489b670865a6111bc/How_Root_Rebuilt_Its_Telematics_Platform_for_Scale.png?w=1400&h=788&q=60&fm=webp 1400w","sizes":"(min-width: 1400px) 1400px, 100vw","type":"image/webp"}],"fallback":{"src":"//images.ctfassets.net/pacigpl3aj13/6khgGxb7AecT8FWoasAG3a/093b9475639d49b489b670865a6111bc/How_Root_Rebuilt_Its_Telematics_Platform_for_Scale.png?w=1400&h=788&q=60&fm=png","srcSet":"//images.ctfassets.net/pacigpl3aj13/6khgGxb7AecT8FWoasAG3a/093b9475639d49b489b670865a6111bc/How_Root_Rebuilt_Its_Telematics_Platform_for_Scale.png?w=350&h=197&q=60&fm=png 350w,\n//images.ctfassets.net/pacigpl3aj13/6khgGxb7AecT8FWoasAG3a/093b9475639d49b489b670865a6111bc/How_Root_Rebuilt_Its_Telematics_Platform_for_Scale.png?w=700&h=394&q=60&fm=png 700w,\n//images.ctfassets.net/pacigpl3aj13/6khgGxb7AecT8FWoasAG3a/093b9475639d49b489b670865a6111bc/How_Root_Rebuilt_Its_Telematics_Platform_for_Scale.png?w=1400&h=788&q=60&fm=png 1400w","sizes":"(min-width: 1400px) 1400px, 100vw"}},"layout":"constrained","width":1400,"height":788,"placeholder":{"fallback":"data:image/jpeg;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAMAAABI111xAAACi1BMVEWYjoKbkYadlImhmI6jm5GqoZewp5y6saXEuq7KwLTSyLvXzsHVzcPTzMPSzMPVz8bb1cvi3NDp4tXu5tqckoZ6dGxPTkxVVFNdXVxpaGZ4dnOHhIClnpakoZx0fYWBi5SMlqCYoaqrs7m8wcTU087v6Nudk4hoY1wAAAAuKCFPRz5NSEJeWVRcV1ONiIDMwrXUy72HhH8ABh84P0VOVl1lbXSFio+NkpbHx8Xx6d2bkIRqZV5HPTORgnB/bluZiHawopF5bV+HgXnLwLTVy72Kh4KBe3Okm4+qopenoJbIwLSopJzAwb7o4tePhXlhWlJWSz+di3iqmYefk4O6qpiFdWWEfXXFuq3Mw7eGg3+CfXamnZKooZWmn5XOxLi0raWqrazX0saCeW5XUEdzZ1iSfGqej4CYjoC5p5KWgm2NhXzHu67Ow7WGg30zR1VdZ2B+d1+qiF2sfmKdeWyurqzk2cpzamBYU0t3aVehkoGajHulloW3ppOYiHSPhXi1qJnBtKSAe3VFT1VhZl16cl2khWCngWqWenC2tbPo3c9lXFFORj56alefj3yll4WNf2ysk4CqmYizppeXiXuPg3ZdW1iGf3elm4+tpZqwqJy2rqGhm5Wtrave0sN5a1xhU0Z5bmCdi3m9qZOfjXmbjn2cjHq3rZ+xo5Kgk4RkYmB3dXOKiIaQkI6cnZygoqTCv7nq4NCun46jk4KjloehkIG0qJqupJiilYWpmonBsqHIuqrKvKzCtaa/s6a6sKS8sqfCua/NxLra0MTn3M3s4tKxpJWvopWzppi5rZ/AtKfIvK7GuanKvK3Mv6/NwLHOwbLUx7fazLzd0MDh08Li1cTk18fk2Mfl2cnn3MxMnuifAAAAJUlEQVQI12NkYMQELCIw1lsGIbggnCWMTSUj/QQn58FYG4nVDgDmJARofNlbAAAAAABJRU5ErkJggg=="}}}},{"publishingMetadata":{"slug":"/six-under-appreciated-principles-of-data-warehousing-through-the-eyes-of-a-data-scientist"},"blogCategory":[{"categoryTitle":"Data Science","publishingMetadata":{"slug":"/data-science"}}],"customDateCreated":"December 15, 2021","previewTitle":"Six Under-Appreciated Principles of Data Warehousing, Through the Eyes of a Data Scientist","previewBody":"Challenges can emerge between data science research and production. This article explores practical strategies for building more reliable, scalable machine learning pipelines.","imageAltText":"Glasses held up to a computer looking at code","previewImage":{"gatsbyImageData":{"images":{"sources":[{"srcSet":"//images.ctfassets.net/pacigpl3aj13/3kzSUXq3kTTp3jzZAeP77s/b6fe6d4a227a0308259a2d5dbbd25bc0/Glasses___Code.png?w=359&h=245&q=60&fm=webp 359w,\n//images.ctfassets.net/pacigpl3aj13/3kzSUXq3kTTp3jzZAeP77s/b6fe6d4a227a0308259a2d5dbbd25bc0/Glasses___Code.png?w=718&h=490&q=60&fm=webp 718w,\n//images.ctfassets.net/pacigpl3aj13/3kzSUXq3kTTp3jzZAeP77s/b6fe6d4a227a0308259a2d5dbbd25bc0/Glasses___Code.png?w=1436&h=980&q=60&fm=webp 1436w","sizes":"(min-width: 1436px) 1436px, 100vw","type":"image/webp"}],"fallback":{"src":"//images.ctfassets.net/pacigpl3aj13/3kzSUXq3kTTp3jzZAeP77s/b6fe6d4a227a0308259a2d5dbbd25bc0/Glasses___Code.png?w=1436&h=980&q=60&fm=png","srcSet":"//images.ctfassets.net/pacigpl3aj13/3kzSUXq3kTTp3jzZAeP77s/b6fe6d4a227a0308259a2d5dbbd25bc0/Glasses___Code.png?w=359&h=245&q=60&fm=png 359w,\n//images.ctfassets.net/pacigpl3aj13/3kzSUXq3kTTp3jzZAeP77s/b6fe6d4a227a0308259a2d5dbbd25bc0/Glasses___Code.png?w=718&h=490&q=60&fm=png 718w,\n//images.ctfassets.net/pacigpl3aj13/3kzSUXq3kTTp3jzZAeP77s/b6fe6d4a227a0308259a2d5dbbd25bc0/Glasses___Code.png?w=1436&h=980&q=60&fm=png 1436w","sizes":"(min-width: 1436px) 1436px, 100vw"}},"layout":"constrained","width":1436,"height":980,"placeholder":{"fallback":"data:image/jpeg;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAAOCAIAAACgpqunAAACjklEQVQYGQXB224TVxQA0L3PZebMeHwlbSFJG0pBihAPvPQb+mt94Af4lH4EQi2VaASRIlsISGyP7Zk5c657dy18++ZPUvp0CIIlSFlXpjaVVogM/nR0PriQq2mNAp11nSMGYQcnJTAFtbvfqJ9fjMra0RlKvt9tOjddnAkgSY4EpEzJ1jklicqY4nQcvHdSCM6k1g+7NHtxXtR+7MCfDiH2AtYfPyIlJSUiEgMxEAMzMAAR5ZQWq+VytVIoYD6tJoUpy+Zq1cQU7/q043ff159BSgEIAICAwIAMzBKRAZppfXm5VJzzcddmGRWk0N679ttXO7rtNjIigRSAgMDMTAiMAIwAgnRVM6J81uCXv2+WRX56cdE0M4VCKlNz0t5mWcbgwfbaD9F5EUZyo0zusYlSFd3JquXvf3A7Xr9+OXu0HJ07b6pFzPH8yX59e/PPexO659evVKW7/oSIR+uOh8PQb3ef/mqzUaWmqx+mot91GL+fxuF+U8XjfDbRw2H5628ztnVtvJkxlYmhLvyYkqXkmjyfTVXu9z9ePKtqs3X+5sP7h7ZFYCPF49R/VfWMY6Puh2oqmQBIICFQLbE4W+ZJLSiGZjqpJNVKLOZTjaAQTY57KDIKK7R3gWNgJkRmBBSYibKpJs1EKIFZlmrxU3rYDJtPk9DrcUgp9UIzQQTpVTmJtoAsAaTAIjqpFRdlSFll4od9u9/u/r39srVZKwNF0ceImQQiMFpdAorCOcGUEVQhYTIHon07qBDC3X/vgg9a6V+eXvZDdzwNlAEZGQUwZiG8LnJZaoVKYtaqroxO+Xi06ur58uJ6FQiCg5ieDL7/vO4+3IUzNZjutm0zIiMygABAAERM3kdKyfb2fwSKnO6BqjLMAAAAAElFTkSuQmCC"}}}}]}]}},"pageContext":{"id":"56d960e4-853f-5c8c-91aa-55e0afd91e31"}},
    "staticQueryHashes": ["1149588473","253762530","286361453","3663576919","3700698109","633915060","637342082"]}