What is MLOps? Machine Learning Operations Explained
Most ML projects never make it to production. The data scientist trains a model, it works in a notebook, and then it gets stuck -- nobody knows how to deploy it, nobody can reproduce the training environment, and nobody trusts that it will keep working once real traffic hits it. MLOps exists to close that gap.
The Production ML Problem
A trained model is not a product. A model is a function that maps inputs to outputs under the specific conditions it was trained on. When those conditions change -- new data patterns, upstream schema changes, feature drift -- the model degrades, often silently. Unlike a crashing application that throws errors, a degrading model just gives worse answers. Users might not notice for weeks.
Getting a model into production and keeping it working requires solving several problems simultaneously:
- Reproducibility: Can you retrain this model exactly? Do you have the data version, the feature logic, and the hyperparameters?
- Deployment: Can you serve the model reliably at scale with predictable latency?
- Monitoring: Can you detect when the model's behaviour diverges from training performance?
- Retraining: Can you trigger a new training run, validate the new model, and promote it to production safely?
MLOps is the set of practices and tooling that makes all of these answerable in production.
The ML Lifecycle
The MLOps lifecycle has more stages than a standard software delivery pipeline:
Data ingestion -> Feature engineering -> Model training
-> Evaluation -> Registry -> Deployment -> Monitoring -> Retraining
Each arrow is a potential failure point. The feature engineering step is particularly tricky: the same feature logic must run identically during training and serving. A mismatch -- called training-serving skew -- is one of the most common sources of production ML bugs and one of the hardest to detect.
Feature stores (like Feast or Tecton) address this by providing a single, versioned source of features that both training pipelines and serving infrastructure read from.
Key MLOps Concepts
Experiment tracking -- During model development, teams run dozens or hundreds of experiments varying data, features, and hyperparameters. Tools like MLflow and Weights and Biases record these runs: what parameters were used, what data was loaded, what metrics resulted. Without this, reproducing a result is guesswork.
Model registry -- A versioned store of trained model artefacts along with their metadata: training data version, performance metrics, who approved them for production. The registry is the handoff point between training and deployment. A model does not go to production without a registry entry.
Pipeline orchestration -- Training is not a one-off script. It is a DAG of steps: data extraction, preprocessing, feature computation, training, evaluation, and registration. Kubeflow Pipelines, Metaflow, and Apache Airflow manage these DAGs, making training pipelines reproducible, schedulable, and auditable.
Model serving -- Deploying a model means wrapping it in a server that accepts requests and returns predictions. Options range from simple FastAPI wrappers to dedicated serving frameworks like Seldon Core or BentoML, which add batching, A/B testing, model versioning, and canary rollouts.
A minimal serving setup with BentoML looks like this:
import bentoml
from bentoml.io import NumpyNdarray
runner = bentoml.sklearn.get("fraud_classifier:latest").to_runner()
svc = bentoml.Service("fraud_classifier", runners=[runner])
@svc.api(input=NumpyNdarray(), output=NumpyNdarray())
def classify(input_data):
return runner.predict.run(input_data)
This gets containerised, pushed to a registry, and deployed to Kubernetes using the same CI/CD pipeline you would use for any other service.
MLOps Maturity
Google's MLOps maturity model defines three levels:
Level 0 -- Manual process. Data scientists train models locally and hand off static model files. No pipeline automation. This is where most teams start.
Level 1 -- ML pipeline automation. Training pipelines run automatically on new data. Models are tracked in a registry. Deployment is still partly manual.
Level 2 -- CI/CD for ML. The full pipeline -- data validation, training, evaluation, deployment -- is triggered automatically and runs without human intervention. New model versions go to production only when automated tests pass.
Most production ML teams are working their way from Level 0 to Level 1. Level 2 is the goal for teams where models are on the critical path and need to retrain frequently.
Why DevOps Skills Matter
MLOps is built on a DevOps foundation. Containerisation with Docker, infrastructure as code, Kubernetes, and CI/CD are not optional extras -- they are the substrate that MLOps tooling runs on. A data scientist who understands how to containerise a training job, write a GitHub Actions workflow, and deploy to Kubernetes is orders of magnitude more effective at getting models into production than one who cannot.
