
MLOps
Production ML pipelines. CI/CD for models, monitoring, feature stores, model serving, and infrastructure automation.

After this course, you'll be able to:
Who this is for
- →Data scientists whose models keep dying somewhere between a notebook and a production endpoint.
- →Backend and platform engineers who have inherited an ML service and discovered it behaves nothing like the rest of the stack.
- →Teams that shipped one model successfully and now cannot explain why the second one is harder.
- →Engineers moving into ML platform or ML infrastructure roles who need the vocabulary and the failure catalog.
- →Not a good fit if you want to learn modeling itself, or if you have no model in production and no near-term plan to put one there. The material is about operations, and operations without a system to operate is abstract.
What you need first
- ·Working Python, including packaging a project rather than only running notebooks.
- ·Git as a daily habit, including branching and pull requests.
- ·Enough understanding of training and evaluation to know what a validation set is and why a metric can be misleading.
- ·Basic Linux, containers and HTTP. Kubernetes is taught from the concepts up, so prior exposure helps but is not assumed.
- ·Access to some environment you can deploy into, even a single virtual machine or a local cluster.
What MLOps is once you take the tooling away
Strip away the vendor diagrams and MLOps is the practice of making a machine learning system reproducible, observable and safely changeable. That sounds like ordinary software engineering, and much of it is. The part that is not comes from a structural difference: a conventional service has one moving part that changes deliberately, which is the code. A machine learning system has three, and only one of them is under version control by default.
Code changes when someone commits. Data changes continuously, without anyone deciding it should, because the world the data describes keeps moving. The model is a function of both, so a model artifact is only meaningful when tied to the exact code and the exact data snapshot that produced it. Break that link and you have an artifact nobody can rebuild, explain or roll forward.
The classic statement of the problem is the paper on hidden technical debt in machine learning systems, published by a group at Google and presented at NIPS, now NeurIPS, in 2015. Its central observation still holds: the model code is a small box in the middle of a much larger diagram, surrounded by configuration, data collection, feature extraction, verification, resource management, serving infrastructure and monitoring. Teams spend their effort on the small box and are then surprised by where the maintenance cost lands. The same paper introduced the entanglement principle, sometimes shortened to "changing anything changes everything": because a model mixes all its inputs into one learned function, you cannot reason about a feature in isolation. Removing an apparently unused input can move the model's behavior in ways no unit test would predict.
Two more properties separate this from normal service operations.
The first is that correctness is statistical, not binary. There is no assertion that fails when a model becomes worse. It returns a well-formed response with a plausible value, and the only signal is an aggregate that moves slowly. By the time it is visible on a weekly chart, the bad predictions have already been acted upon.
The second is that the feedback loop is long and sometimes circular. Ground truth often arrives days or weeks after the prediction, if it arrives at all, so the metric you care about is unavailable when you need it. Worse, systems that influence the world they measure end up training on data shaped by their own previous outputs. A recommendation model that only ever receives feedback on the items it chose to show is learning from a sample it selected. That is not a bug in the training code, and no amount of pipeline hygiene fixes it.
The lifecycle: experiments, registries and promotion gates
The organizing idea of a mature ML platform is that a model should move through defined states, and each transition should be a recorded, reversible event rather than someone copying a file.
The first stage is experiment tracking. Every training run records its parameters, the code revision, the data version, the environment, the evaluation results and the artifact it produced. This feels like bureaucracy for the first month and becomes indispensable the first time someone asks why last quarter's model behaved differently. Without it, the answer is a reconstruction from memory.
The second is data and feature versioning. Code versioning is solved; data versioning is not, and the approaches differ in what they optimize for. Content-addressed storage of dataset snapshots gives exact reproducibility at the cost of storage. Immutable, partitioned tables with time-travel queries give cheaper history at the cost of exactness when upstream corrections land. Recording the transformation logic plus a deterministic extraction query gives you lineage without duplicating the data, but only if the source system never mutates in place, which it usually does.
The third is a model registry. A registry is a small idea with large consequences: models get versions, versions get stages, and stage transitions carry approvals and metadata. It gives you one answer to "what is running in production right now", which is remarkably hard to obtain otherwise, and it makes rollback a lookup instead of an investigation.
The fourth is the promotion gate, which is where continuous integration for models diverges from continuous integration for code. A useful gate checks several things at once: that the training pipeline runs end to end on a small sample, that the evaluation metric clears a threshold on a held-out set, that the metric has not degraded against the currently deployed version, that performance has not dropped on specific slices even if the aggregate improved, that the serving code can load the artifact and produce a prediction, and that inference latency and artifact size are within budget. Slice checks are the ones most often skipped and most often needed, because an aggregate improvement routinely hides a regression on a minority segment.
Underneath all of it sits the same requirement: any run should be re-executable from its recorded inputs. If reproducing an old result requires a person who remembers the details, you do not have a pipeline. You have a habit.
Serving: the architectural decision that is expensive to reverse
Serving looks like an implementation detail and behaves like an architectural commitment. The shape you pick determines what your latency budget is, where features come from and how hard consistency will be.
| Pattern | How predictions are produced | Suits | Principal difficulty |
|---|---|---|---|
| Batch scoring | Scheduled job writes predictions to a table | Predictions that stay valid for hours or days, large populations | Staleness, and no answer for entities that appear between runs |
| Online service | Synchronous request to a model endpoint | Predictions that depend on request-time context | Latency budget, autoscaling, feature availability at low latency |
| Streaming | Predictions computed as events arrive | Continuous decisioning, fraud, monitoring | Operational complexity, replay and exactly-once semantics |
| Embedded | Model runs inside the client or application process | Offline capability, strict latency or privacy requirements | Update distribution, version fragmentation across the fleet |
Cutting across all four is the problem that generates more production incidents than any other in this field: training and serving skew. The features used during training were computed by an analytics job over historical data, in a language and with a set of defaults chosen for that context. The features used at inference are computed by application code, in a different language, under a latency constraint. Any difference between the two, including how nulls are handled, how timestamps are rounded, whether a category encoder saw an unseen value, or which of several similarly named columns was picked, degrades the model in production while every offline evaluation continues to look correct.
Feature stores exist to make this a solved problem rather than a recurring one, by computing a feature once and serving it to both the training path and the inference path from one definition. They also address point-in-time correctness, which is the subtler half of the same problem: when you assemble a training set, every feature value must reflect what was knowable at the moment of the event, not what the table says today. Joining current values onto historical events leaks the future into training and produces offline results that are excellent and meaningless.
Feature stores are not free. They add a system, a query path and a schema discipline. The honest rule of thumb is that they earn their place when several models share features, or when the same feature must be available both in bulk for training and at single-record latency for serving. One model with a handful of features computed in one place does not need one.
For model-serving mechanics, the levers are familiar from ordinary services with one twist. Request batching raises hardware utilization and adds queueing latency, so it is tuned against a latency target rather than maximized. Model optimization, whether quantization, distillation, pruning or graph compilation, buys throughput at some accuracy cost that has to be measured rather than assumed. And loading a large artifact is slow enough that autoscaling policies tuned for stateless web services will scale you up long after the traffic spike has passed.
How ML systems fail quietly
Outages are the easy case. The characteristic ML failure is a system that is fully available, returning valid responses, and steadily getting worse.
Distribution change is the usual cause, and it is worth separating the varieties because they call for different responses. Covariate shift means the inputs have moved while the underlying relationship holds, which retraining on recent data generally fixes. Label shift means the base rate of the outcome has changed, which sometimes needs only a recalibrated threshold. Concept drift means the relationship itself has changed, so yesterday's labels are actively misleading, and retraining on stale history makes things worse rather than better. Then there is the case that is not drift at all: an upstream pipeline changed a unit, a default, a currency or an encoding, and the model is being fed something that is technically valid and semantically different.
That last category is the most common and the least glamorous. Data validation at the pipeline boundary catches most of it: schema and type checks, null-rate bounds, range and cardinality checks, and distributional comparison against a reference window. These are cheap to write and catch the failures that would otherwise be miserable to diagnose from the model's behavior alone.
The monitoring stack for an ML service therefore has layers. Service health tells you the endpoint is up. Data health tells you the inputs still look like the inputs. Prediction health tells you the output distribution has not shifted, which is available immediately and is the best early proxy when labels are delayed. Model quality tells you the accuracy metric against ground truth, which is what you actually care about and usually the last thing you can measure. Business impact closes the loop, and is the only layer that answers whether the model is worth running at all.
Delayed labels are the structural constraint. If ground truth arrives weeks later, quality monitoring is inherently retrospective, so the practical approach is to alert on the leading indicators, use whatever partial or proxy labels exist, and deliberately hold out a small random slice from automated action so you retain an unbiased sample to evaluate against. That last one costs a little performance and is often the only clean signal available.
Finally, there is the accumulating structural debt that no dashboard shows: pipelines that nobody can delete because it is unclear what consumes them, features kept alive for a model retired last year, glue code holding together tools chosen at different times, and configuration spread across enough places that no single person can state the current behavior. This is the cost that eventually dominates, and it is why the operational practices matter more than the choice of framework.
Changing a live model without breaking anything
Deploying a new model is a change to a decision-making system, and it deserves more caution than a code deploy because the effect is probabilistic and often delayed.
Shadow deployment runs the candidate on live traffic while the incumbent continues to serve. Nobody is affected, and you get real-world input distributions, latency behavior under actual load, and a direct comparison of prediction distributions. It catches serving bugs, feature availability problems and skew before any user sees the result. It cannot tell you whether the new model produces better outcomes, because its predictions never influence anything.
Canary release sends a small share of traffic to the candidate and watches operational and quality signals before widening. This is the workhorse. It requires that you decide in advance what would make you stop, and that the stop is automatic, because the failure of manual canaries is that everyone is busy when the numbers start moving.
Online experiments are how you learn whether the model is genuinely better on the outcome you care about. They are also where the most confident mistakes happen. Choosing the metric after seeing the data, stopping the moment significance appears, running many comparisons without accounting for it, ignoring novelty effects in the first days, and randomizing at the wrong unit so that treatment leaks between groups are all routine. So is the harder problem of interference, where the treatment group's behavior changes what the control group experiences, which breaks the independence the whole method assumes. Deciding the metric, the unit of randomization, the duration and the stopping rule before launch is not paperwork. It is the difference between measuring something and confirming a preference.
Automated recovery is the other half. A rollback path that requires a person to find the previous artifact, remember the config and redeploy by hand is not a rollback path. Practical setups keep the previous version loaded or immediately loadable, define health checks that cover data and prediction quality rather than only process liveness, and wire the failure of those checks to an automatic revert. The complement is graceful degradation: a defined fallback, whether a simpler model, a cached prediction or a deterministic rule, so that a failing model produces a known-safe answer rather than an error page or an arbitrary one.
Circuit breakers deserve a mention because they are underused here. If the feature store is slow, an inference service that keeps waiting will exhaust its connection pool and take down everything sharing it. Failing fast to the fallback keeps a degraded model available instead of trading it for an outage.
Cost, control and whether you need this yet
Machine learning infrastructure has an unusual cost profile: expensive hardware, bursty demand and a strong tendency for spending to become invisible. The recurring findings are dull and reliable. Training clusters left running after the job finished. Accelerators reserved for workloads that never needed one. Inference on a large model where a smaller one, distilled or quantized, would have met the requirement. Pipelines re-reading the same data because caching was never wired up. Autoscaling that cannot scale down because loading the artifact takes long enough to make the policy oscillate.
The levers worth learning are right-sizing based on measured utilization rather than requested resources, separating training and serving so their scaling policies do not fight, using interruptible capacity for work that can checkpoint and resume, batching wherever latency permits, and treating model size as a cost decision rather than only an accuracy decision. Cost attribution per model matters too, because a platform that cannot say which model is spending the money cannot make a case for anything.
Security and compliance arrive later than they should. The concrete surface includes access control over training data and feature stores, provenance for anything a regulator might ask about, artifact integrity so the model you validated is the model that loaded, and a defensible answer for how personal data flows through training and inference. Public frameworks help structure that work: the risk management framework published by NIST is a reasonable spine for the governance layer, and sector rules will dictate the rest. The engineering translation is usually less exotic than the policy language suggests, and mostly reduces to lineage, access control, logging and reproducibility, which are the same practices that make the system operable in the first place.
Are you ready for this material
A short diagnostic. Do you have at least one model making decisions that someone would notice if it stopped? Can you currently answer, without asking a colleague, which model version is serving and which data produced it? Has a model already degraded without anyone noticing until a stakeholder mentioned it? If the first is a yes and either of the others is uncomfortable, this is the right time.
If you have no model in production yet, the more useful sequence is to ship one end to end in the simplest form that works, feel where it hurts, and come back. Building a platform before you have a system to run on it is the most reliable way to build the wrong platform.
What compounds with this: data engineering, because the majority of production ML incidents originate upstream of the model; evaluation design, because every rollout decision depends on a metric you trust; and distributed systems fundamentals, because most of the hard parts here are ordinary distributed systems problems wearing a machine learning hat.
Common questions
Is MLOps just DevOps applied to machine learning?
It borrows the practices and adds two things DevOps does not have to handle. Data is a versioned input that changes on its own, so reproducibility means pinning a dataset as well as a commit. And correctness is statistical rather than binary, so testing means evaluating a metric against a baseline and across slices rather than asserting an expected output. Everything else, including CI, containers, infrastructure as code and observability, transfers directly.
Do I need Kubernetes to do MLOps properly?
No. Kubernetes is common because it solves scheduling, scaling and isolation for heterogeneous workloads, and most managed ML platforms are built on it, so understanding it makes those platforms legible. But a single well-managed server with containers, a scheduler and proper monitoring is a perfectly respectable production setup for one or two models, and it is a much better place to learn the principles than a cluster you are also debugging.
What is the difference between MLOps and LLMOps?
The lifecycle is the same and the emphasis moves. With a large language model you often do not train the weights, so the artifacts you version are prompts, retrieval indexes, tool definitions and adapters. Evaluation shifts from a scalar metric to graded rubrics and behavioral test suites, because there is no single accuracy number. Cost moves from training to inference, monitoring gains a safety and abuse dimension, and latency becomes a streaming property rather than a single figure. The registry, promotion gates, canary releases and rollback machinery are unchanged.
How do you monitor model quality when labels arrive weeks late?
You monitor what is available immediately and treat quality as a lagging confirmation. Input distributions, null rates and schema conformance are available at once. The prediction distribution is available at once and is a good early warning, because a model whose output mix has shifted is usually reacting to an input that has shifted. Proxy signals, partial labels and a small randomized holdout that is never acted upon give you an unbiased sample to evaluate against once ground truth lands.
Do I need a feature store?
Only when the problem it solves is one you have. It earns its cost when several models share features, when the same definition must serve both bulk training and low-latency inference, or when point-in-time correctness is difficult to guarantee by hand. For a single model with features computed in one codebase, adding a feature store adds a system to operate and solves nothing you were not already solving.
Which tools should I learn?
Learn the categories rather than the products, because the products churn and the categories do not. You need somewhere to record experiments, somewhere to version data, a model registry, a pipeline orchestrator, a serving runtime, and a monitoring stack that covers data as well as service health. Once you can say what each layer must guarantee, evaluating any specific tool takes an afternoon.
Related reading
Deploying models on Kubernetes
The container orchestration layer in concrete form: manifests, resource requests and node scheduling for a model workload.
Prometheus and Grafana for model serving
Wiring up the metrics layer, which is where any serious monitoring practice starts.
Load balancing inference traffic
Routing across replicas, and why stateless load-balancing assumptions do not always hold for model servers.
Local versus cloud deployment strategies
The hosting decision that sits underneath most cost and compliance discussions.
Self-hosted AI cost calculator
Working through the total cost of running inference yourself instead of paying per request.
Version control at scale
Practical approaches to versioning large data artifacts, the piece Git alone does not cover.
Full syllabus
Model Development Lifecycle
Data Engineering
Model Serving
Container Orchestration
Model Optimization
A/B Testing
Monitoring & Observability
Automated Recovery
Security & Compliance
Cost Optimization
Case Studies
Unlock all 12 chapters
Plus 24 other courses — 549 more chapters included.
Every course, every future course, the Python Lab and eight downloadable kits, nothing to renew. Or subscribe: Pro $8.99/month