← All posts

From Notebook to Production: MLOps for Intermediate Developers

Learn how to move your ML models from a single notebook to a scalable, monitored production API using versioning, CI/CD, and observability.

Why MLOps Matters Beyond the Notebook

When you finish a notebook, the model feels finished. But in production the data keeps moving, the model drifts, and the code that trained it is buried in a Jupyter file. Without a systematic approach, reproducibility, monitoring, and collaboration break down, and deployment stalls.

TL;DR

  • Version everything – code, data, experiments, and models – otherwise you’ll lose reproducibility.
  • Automate the pipeline – unit tests, CI, Docker, and IaC keep the model from the notebook to a container in minutes.
  • Expose the model as an API – FastAPI + Docker gives low‑latency inference with minimal overhead.
  • Observe in production – logs, metrics, and alerts turn the model into a reliable service.

Why MLOps Matters Beyond the Notebook

  • Data drift and concept drift become hard to detect without systematic monitoring.
  • Reproducibility gaps arise when experiments are not versioned or logged.
  • Collaboration slows when notebooks are the sole artifact of a model.
  • Deployment latency increases if manual steps are required to move code to production.

Data Drift and Concept Drift

Once a model leaves the notebook, the data it sees can change subtly or dramatically. A single missing feature or a new distribution can reduce accuracy by a few percent, which in a recommendation engine translates to thousands of lost sales. Detecting these changes requires automated data validation and drift alerts—something that notebooks alone cannot provide.

Reproducibility Gaps

A notebook’s “copy‑paste” nature makes it easy to forget the exact library versions, random seeds, or preprocessing steps. When another team member tries to reproduce the result, they may end up training a different model. Version control for code and data, combined with experiment tracking, eliminates this gap.

Collaboration Bottlenecks

Teams that rely on notebooks struggle to share reusable components. A data engineer might need the same preprocessing pipeline that a data scientist wrote in a notebook, but extracting that logic manually is error‑prone. Packaging code into libraries and using a shared registry speeds up collaboration.

Deployment Latency

Without a CI/CD pipeline, moving a model from a notebook to a production API often involves manual copy‑pasting, environment setup, and testing. Each manual step adds risk and delays. Automating the build, test, and deployment cycle reduces friction and the chance of human error.

Key Concepts Every Intermediate Dev Should Own

Concept Why It Matters Typical Tool
Model versioning Keeps track of model weights, hyperparameters, and training code. DVC, MLflow
Experiment tracking Records hyperparameters, metrics, and code hashes for reproducibility. MLflow, Weights & Biases
Infrastructure as Code (IaC) Ensures consistent environments across dev, staging, and prod. Terraform, CloudFormation
Continuous Integration for ML Validates code, tests, and model quality before deployment. GitHub Actions, GitLab CI

Model Versioning

Model artifacts are just files, but without a versioning system you can’t trace which training run produced which weights. DVC stores artifacts in a remote store while keeping a lightweight Git history. MLflow provides a UI to compare runs and roll back to a previous model.

Experiment Tracking

A single notebook run can produce dozens of metrics. Tracking these automatically allows you to compare hyperparameter sweeps and identify the best model without scrolling through notebook cells. The tracking server also stores the exact code commit that produced the run, ensuring you can reproduce it later.

IaC

Spinning up a virtual machine or container with the exact libraries you used in the notebook guarantees that the model behaves the same in staging and prod. Terraform modules or CloudFormation stacks can be versioned alongside your code.

CI for ML

Unit tests for preprocessing, feature engineering, and inference are as important as tests for the training script. A CI pipeline that runs these tests, lints the code, and validates the model against a sanity‑check dataset keeps regressions from slipping into production.

From Notebook to CI Pipeline: A Step‑by‑Step Workflow

  1. Extract notebook logic into reusable Python modules.
    Move preprocessing, feature engineering, and training into separate files under a src/ package. Use __init__ files to expose a clear API.

  2. Write unit tests for data preprocessing and feature engineering.
    Use pytest and hypothesis to generate edge‑case inputs. Store test data in a tests/fixtures/ directory.

  3. Build a Docker image that includes the model and its dependencies.
    A minimal Dockerfile that starts from a slim Python image, installs poetry or pip, copies the source, and runs a lightweight WSGI server.

  4. Configure GitHub Actions to run training, tests, and push the image to a registry.
    The workflow triggers on push to main and pull_request. It runs tests, builds the image, tags it with the commit SHA, and pushes it to Docker Hub or a private registry.

Sample GitHub Actions Workflow

name: CI

on:
  push:
    branches: [main]
  pull_request:

jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v3
      - uses: actions/setup-python@v4
        with:
          python-version: '3.11'
      - run: pip install -r requirements.txt
      - run: pytest tests/

  build:
    needs: test
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v3
      - run: docker build -t ghcr.io/yourorg/model-api:${{ github.sha }} .
      - run: echo ${{ secrets.GHCR_TOKEN }} | docker login ghcr.io -u ${{ github.actor }} --password-stdin
      - run: docker push ghcr.io/yourorg/model-api:${{ github.sha }}

This workflow keeps the model in sync with the code and ensures that every push results in a reproducible, deployable image.

Deploying Models with Minimal Overhead

Deploying a model as a service is more than just running a script. The goal is to expose a contract (an API) that can be called by downstream services with predictable latency.

Wrap the Model in a FastAPI Endpoint

FastAPI is lightweight, async‑friendly, and comes with automatic OpenAPI documentation. It’s a good fit for inference services.

# src/api.py
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel
import joblib
import numpy as np

app = FastAPI(title="Credit Score Predictor")

class CreditRequest(BaseModel):
    age: int
    income: float
    debt: float
    employment_years: int

# Load the pre‑trained model once at startup
model = joblib.load("models/model.pkl")

@app.post("/predict")
def predict(req: CreditRequest):
    try:
        features = np.array([[req.age, req.income, req.debt, req.employment_years]])
        score = model.predict(features)[0]
        return {"credit_score": int(score)}
    except Exception as e:
        raise HTTPException(status_code=500, detail=str(e))

Explanation

  • pydantic validates the incoming JSON and ensures type safety.
  • The model is loaded once when the container starts, avoiding repeated I/O.
  • The endpoint returns a simple JSON with the prediction.

Dockerfile for the API

# Dockerfile
FROM python:3.11-slim

WORKDIR /app

COPY pyproject.toml poetry.lock ./
RUN pip install --no-cache-dir poetry && \
    poetry config virtualenvs.create false && \
    poetry install --no-dev

COPY src/ ./src
COPY models/ ./models

EXPOSE 8000
CMD ["uvicorn", "src.api:app", "--host", "0.0.0.0", "--port", "8000"]

The image is lightweight (~200 MB) and contains only the necessary runtime dependencies.

Docker Compose for Local Development

# docker-compose.yml
version: "3.8"

services:
  api:
    build: .
    ports:
      - "8000:8000"
    volumes:
      - ./logs:/app/logs
    environment:
      - LOG_LEVEL=INFO

This setup allows developers to spin up the service locally and inspect logs.

Blue‑Green Deployment

When updating the model, spin up a new container with the new image, run health checks, and switch traffic. Tools like kubectl rollout or simple Nginx reverse proxies can manage the switch. Blue‑green avoids downtime and allows quick rollback if something goes wrong.

Collect Latency and Error Metrics

FastAPI can be wrapped with middleware that records request latency and error counts. Expose these metrics in a Prometheus‑compatible format (/metrics endpoint). Downstream services can then monitor the API’s performance.

Observability: Seeing the Model in Production

Observability is the ability to understand the internal state of a system from external outputs. For ML, it means tracking predictions, monitoring data quality, and alerting on anomalies.

Log Predictions with Input Metadata

Persist each request and its output to a lightweight database (e.g., SQLite or a cloud table). Include a timestamp, request ID, and input features. This log can be used for auditing, debugging, or retraining.

import sqlite3
from datetime import datetime

conn = sqlite3.connect("/app/logs/predictions.db")
cur = conn.cursor()
cur.execute("""
    CREATE TABLE IF NOT EXISTS predictions (
        id TEXT PRIMARY KEY,
        timestamp TEXT,
        input TEXT,
        output INTEGER
    )
""")
conn.commit()

def log_prediction(request_id, input_json, output):
    cur.execute("""
        INSERT INTO predictions (id, timestamp, input, output)
        VALUES (?, ?, ?, ?)
    """, (
        request_id,
        datetime.utcnow().isoformat(),
        input_json,
        output
    ))
    conn.commit()

Expose Prometheus Metrics

Add a /metrics endpoint that serves counters for total requests, errors, and histograms for latency. Prometheus scrapes this endpoint, and Grafana visualizes it.

Set Alerts on Sudden Drops in Accuracy

If you store predictions, you can compute rolling accuracy or prediction variance. A sudden drop triggers an alert via PagerDuty, Slack, or email. This early warning prevents a silent drift from causing production failures.

Visualize Performance Dashboards

Grafana dashboards can display:

  • Throughput (requests per second)
  • Latency percentiles
  • Error rates
  • Prediction distribution
  • Data drift metrics

These dashboards give teams a real‑time view of the model’s health.

Common Pitfalls and Trade‑offs

Pitfall Why It Happens Mitigation
Over‑engineering pipelines Adding too many tools (Kubernetes, Airflow) before the team is ready. Start with Docker + GitHub Actions; add orchestration only when scaling demands it.
Skipping data versioning Assuming the dataset is static. Use DVC or simple Git LFS to track raw and processed data.
Choosing heavyweight orchestrators early Kubernetes introduces networking, RBAC, and cluster management overhead. Deploy on a managed service (EKS, GKE) only when you need multi‑container workloads.
Balancing flexibility vs reproducibility Branching experiments can lead to divergent code paths. Adopt a clear branching strategy (e.g., main for production, experiment/* for research) and merge back only after validation.

Trade‑offs

  • Speed vs. safety: Rapid iteration may skip tests, but automated tests catch regressions early.
  • Monolithic vs. micro‑service: A single container is easier to manage, but splitting heavy preprocessing into a separate service can improve scalability.
  • Local vs. cloud: Local Docker images are fast to iterate, but cloud registries and CI pipelines provide reproducibility guarantees.

Key Takeaways

  • Version every artifact—code, data, experiments, and models—to eliminate reproducibility gaps.
  • Automate extraction, testing, building, and deployment with a CI pipeline that runs on every push.
  • Package the model as a FastAPI service inside a lightweight Docker image; expose metrics and logs for observability.
  • Use blue‑green deployment to avoid downtime and set up alerts to catch drift or accuracy drops early.
  • Start simple; add complexity only when the team’s scale or reliability needs justify it.