MLOps, whose full form is machine learning operations, is the set of practices, tools and team habits that move a machine learning model from a data scientist's notebook into production and keep it accurate once it's live. It borrows version control, CI/CD and monitoring from DevOps, then adds what ML specifically needs: data versioning, experiment tracking, a model registry, drift detection and automated retraining.
Key Highlights of MLOps
- MLOps covers the full model lifecycle: data, training, validation, registry, deployment, monitoring and retraining. Training is only one stage of seven.
- Google Cloud describes three MLOps levels (0, 1, 2) and Microsoft describes five (0 to 4). Most teams starting out sit at Level 0, with manual notebooks and hand-offs.
- A practical 2026 open-source stack is Git plus DVC for data, MLflow for tracking and registry, Docker and Kubernetes (often with KServe) for serving, and Evidently for drift monitoring.
- LLMOps, which Microsoft calls GenAIOps, extends MLOps with prompt versioning, evaluation sets, safety checks and token cost tracking.
- DevOps engineers already know about half the stack, and a focused 90-day plan covers the ML-specific half.
Introduction to MLOps
Anyone who has shipped a model knows the moment. It scores 0.91 AUC in a Jupyter notebook, and then someone asks: "How do we put this behind the app's API, and who notices when it stops working?" MLOps exists to close that gap.
Think of it as treating a model as a living software product. Its code, its training data, its parameters and the model file itself all get versioned. Training runs through a pipeline rather than on someone's laptop, deployment goes through CI/CD, and after release the model is watched for drift the way an SRE watches latency.
Google Cloud's architecture guide on MLOps continuous delivery and automation pipelines makes a point every practitioner eventually learns the hard way: only a small fraction of a real-world ML system is the ML code. The rest is data validation, configuration, serving infrastructure, testing and monitoring. If you'd like a refresher on what machine learning is and how models learn from data, read that first and come back.
Why MLOps Matters: The Problems It Solves?
ML systems fail in ways ordinary software doesn't, and MLOps is the response. Leave a web service's code unchanged and it behaves the same tomorrow as today, whereas a model with unchanged code can quietly get worse because the world feeding it data has moved on.
Model drift and silent accuracy decay
Take a churn model trained on 2025 customer behaviour. After a pricing change in 2026, it may start losing accuracy, and nothing crashes; the predictions simply get less useful. The two usual culprits are data drift, where input distributions shift, and concept drift, where the relationship between the inputs and the outcome shifts. Without monitoring, the business tends to find out from a quarterly report instead of an alert.
The notebook-to-production gap
Production needs a packaged, tested, versioned artifact with a defined input schema. When the hand-off is a pickle file dropped into Slack, every release turns into a negotiation.
Reproducibility
Six months from now, someone will ask how model v7 was built. Which data snapshot, which feature code, which hyperparameters? If those answers live in one person's memory, you can't rebuild it.
Compliance and governance
Lending, insurance and hiring models need lineage that runs from a single prediction back to the model version, its training data and whoever approved the release. A model registry together with pipeline metadata gives you that trail.
MLOps Lifecycle: Stages of a Machine Learning Pipeline
Seven stages make up the MLOps lifecycle: data management, training, validation, registry, deployment, monitoring and retraining. Draw it as a circle with monitoring feeding back into data and training, because a straight line that ends at deployment is exactly how models go stale.
Stage 1: Data ingestion, validation and versioning
Before anything trains on new data, the pipeline checks schema, value ranges and null rates. Each training dataset then gets a version, typically through DVC or a lakehouse table format, so you can point to the exact snapshot behind any model. Our overview of big data tools covers the storage and processing layer underneath.
Stage 2: Feature engineering and the feature store
Features are the transformed inputs a model actually sees, such as "orders in the last 30 days". A feature store like Feast keeps a single definition of each feature for both training and the live API, which prevents training/serving skew.
Stage 3: Model training and experiment tracking
Training runs as a scripted, parameterised job. Every run logs its parameters, metrics, code commit and artifact to a tracker such as MLflow, so comparing 40 runs becomes a table sort.
Stage 4: Model evaluation and validation
A trained model isn't eligible for release until it clears a set of gates: minimum accuracy or AUC on a holdout set, no regression against the current production model, acceptable performance on important slices (new customers versus old, for example), and basic fairness checks where they apply.
Stage 5: Model registry
Models that pass validation get registered. In the MLflow Model Registry each model has numbered versions and mutable aliases, so you might point a champion alias at version 7. Reassigning that alias to version 8 changes what production loads, and nobody has to touch the deployment configuration.
Stage 6: Model deployment and serving
The registered model is then packaged, usually as a container, and deployed as an API endpoint or a batch job. The familiar DevOps rollout patterns apply here, including canary releases, shadow deployments and A/B tests.
Stage 7: Monitoring and retraining
Once the model is in production, you watch two kinds of signal. Operational metrics cover latency and error rate; model metrics cover input drift, prediction distribution and, once labels arrive, accuracy. When drift crosses a threshold or a schedule fires, the pipeline retrains, and the loop starts again.
MLOps vs DevOps vs DataOps vs LLMOps
MLOps applies DevOps principles to machine learning and adds data and model versioning, continuous training and drift monitoring. DataOps is narrower and focuses on reliable data pipelines, while LLMOps adds prompt, evaluation and cost controls for large language model applications.
| Dimension | DevOps | DataOps | MLOps | LLMOps |
| Main artifact | Application code and binaries | Data pipelines and datasets | Trained models plus the data and code behind them | Prompts, retrieval indexes, orchestration code, sometimes fine-tuned models |
| What gets versioned | Code, config, infrastructure | Pipeline code, schemas, data snapshots | Code, data, features, model files, hyperparameters | Prompts, system instructions, chunking and embedding settings, eval sets |
| Typical tests | Unit, integration, end-to-end | Schema and data quality checks | Data validation, model quality gates, slice tests | Groundedness and relevance evals, safety checks, regression evals on fixed prompts |
| Why it breaks in production | Bugs, config errors, capacity | Late or malformed upstream data | Data drift and concept drift | Model provider changes, prompt regressions, stale indexes, hallucination |
| Cost driver | Compute and hosting | Storage and processing | Training compute and GPU serving | Tokens per request and latency |
How does MLOps differ from DevOps in practice?
If you already know the phases of the DevOps lifecycle, three differences stand out. A commit is no longer the only thing that triggers a release, since new data or a drift alert can too. A green build also proves less than you'd hope: a model can pass every unit test and still perform worse than the one it replaces, which is why you need statistical quality gates. And there's an extra artifact to manage. DevOps versions code, while MLOps versions code, data and the model, and all three have to line up.
MLOps Maturity Levels: Google Cloud and Microsoft Models
MLOps maturity levels measure how automated your ML lifecycle is, anywhere from fully manual notebooks to pipelines that retrain and redeploy themselves. The two frameworks people cite most come from Google Cloud and Microsoft.
Google Cloud's MLOps levels 0, 1 and 2
Google's guide defines three levels:
- Level 0, manual process. Data scientists prepare data, train and validate in notebooks, then hand a model to engineers to deploy. Google notes that releases may happen only a few times a year, with no CI/CD and no active performance monitoring.
- Level 1, ML pipeline automation. The training workflow itself is automated as a pipeline, which enables continuous training (CT). This level adds automated data and model validation, metadata tracking, and the same containerised pipeline in development and production. Retraining can be triggered by a schedule, new data, performance degradation or concept drift.
- Level 2, CI/CD pipeline automation. On top of Level 1, changes to the pipeline code are themselves built, tested and deployed through CI/CD, so new ideas reach production quickly and safely.
Microsoft's MLOps maturity model (Levels 0 to 4)
Microsoft's MLOps maturity modelon the Azure Architecture Centre splits the same journey into five levels and assesses people, processes and technology at each one.
| Microsoft level | Name | What it looks like | Roughly maps to Google |
| 0 | No MLOps | Manual builds, manual training and testing, no central performance tracking | Level 0 |
| 1 | DevOps but no MLOps | Application builds and tests are automated, but models are still handed off manually and results are hard to reproduce | Level 0 |
| 2 | Automated training | Training is automated and tracked, models are reproducible, releases are manual but easy | Level 1 |
| 3 | Automated model deployment | Releases go through a CI/CD pipeline, A/B testing is integrated, full traceability from deployment back to data | Level 1 to 2 |
| 4 | Full MLOps automated operations | Drift or regression signals trigger retraining automatically, model promotion is policy-based | Level 2 |
The mapping column is our own reading of the two documents rather than an official crosswalk. For a quick self-check, ask who on your team can name the data snapshot that trained the model running in production today. If nobody can, you're at Level 0, whatever tools you have installed.
MLOps Tools by Stage (2026 Stack)
No single MLOps tool does everything. Teams assemble a stack, roughly one tool per lifecycle stage, and decide stage by stage between open source and a managed cloud service. Treat the table below as a starting map of widely used options; it isn't a ranking. Many of the CI/CD and container tools overlap with the best DevOps tools you may already use.
| Lifecycle stage | Open-source options | Managed cloud options |
| Data versioning | DVC, lakeFS | Cloud object storage with versioning, Databricks Delta tables |
| Orchestration | Apache Airflow, Kubeflow Pipelines, Prefect | Vertex AI Pipelines, SageMaker Pipelines, Azure Machine Learning pipelines |
| Feature store | Feast | Vertex AI Feature Store, SageMaker Feature Store, Azure Machine Learning managed feature store |
| Experiment tracking | MLflow Tracking | Azure Machine Learning (MLflow-compatible), SageMaker Experiments, Vertex AI Experiments |
| Model registry | MLflow Model Registry, Kubeflow Hub | SageMaker Model Registry, Vertex AI Model Registry, Azure Machine Learning registries |
| Serving | KServe, BentoML, Seldon Core | SageMaker endpoints, Vertex AI endpoints, Azure Machine Learning online endpoints |
| Monitoring | Evidently, Prometheus and Grafana | SageMaker Model Monitor, Vertex AI Model Monitoring, Azure Machine Learning model monitoring |
According to the Kubeflow introduction, its subprojects are Pipelines, Katib, Notebooks, Trainer, Hub (a model registry), the SDK, the Spark Operator and Workspaces. Because Kubeflow is Kubernetes-native, it suits teams that already run clusters. KServe, a Kubernetes model-serving platform for predictive and generative models, is now a CNCF incubating project.Evidentlyis Apache 2.0 licensed and handles both tabular drift checks and LLM output evaluation.
MLflow vs Kubeflow: which should you learn first?
Learn MLflow first; the two solve different problems anyway. MLflow is a Python-first library and server for tracking experiments, registering models and packaging them, and it runs happily on a laptop. Kubeflow is a platform of components on top of Kubernetes for running pipelines, training and serving at scale, so add it (or a managed pipeline service) once you genuinely need cluster-scale orchestration. Picking between AWS, Azure, and Google Cloud is a separate decision, and our guide to cloud platforms for production AI walks through it.
MLOps Pipeline Example: From Training to Monitoring
The pipeline below is a compact one for a customer churn model, built from tools you can run for free: GitHub Actions, DVC, MLflow, Docker, Kubernetes with KServe, and Evidently. Commands come from each project's own docs, while the file names and thresholds are illustrative.
Step 1: Version the data with DVC
DVC's getting-started guideputs a dataset under version control alongside Git with four commands, followed by an ordinary Git commit:
| dvc init dvc add data/churn_2026_09.csv dvc remote add -d storage s3://my-bucket/dvc-store dvc push git add data/churn_2026_09.csv.dvc .gitignore && git commit -m "Churn data Sept 2026" |
Git stores only a small .dvc pointer while the CSV itself lives in object storage. Running git checkout and then dvc pull restores any version.
Step 2: Train, log and register with MLflow
In train.py, log parameters and metrics, then register the model by passing registered_model_name to log_model:
| import mlflow, mlflow.sklearn from sklearn.ensemble import GradientBoostingClassifier from sklearn.metrics import roc_auc_score with mlflow.start_run(): model = GradientBoostingClassifier(n_estimators=300, max_depth=3) model.fit(X_train, y_train) auc = roc_auc_score(y_test, model.predict_proba(X_test)[:, 1]) mlflow.log_params({"n_estimators": 300, "max_depth": 3}) mlflow.log_metric("test_auc", auc) mlflow.sklearn.log_model(model, "model", registered_model_name="churn-model") |
Step 3: Gate and automate in CI
A GitHub Actions workflow retrains on a weekly schedule, or whenever the data pointer changes, and fails the job if the new model misses the quality bar:
| name: churn-train on: schedule: - cron: "0 2 * * 1" push: paths: ["data/*.dvc", "src/**"] jobs: train: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - run: pip install -r requirements.txt - run: dvc pull - run: python src/train.py - run: python src/check_gate.py --metric test_auc --min 0.85 --compare-to champion |
You write check_gate.py yourself. It compares the new run's metric with the champion version in MLflow and moves the alias only when the new model passes. Leave it out and the workflow will happily ship regressions on a schedule.
Step 4: Package and serve
The MLflow local deployment docsshow how to test a model as a REST server and then build a container from it:
| mlflow models serve -m "models:/churn-model@champion" -p 5001 mlflow models build-docker -m "models:/churn-model@champion" -n churn-model:v8 |
The local server exposes an /invocations endpoint that accepts JSON such as {"dataframe_split": ...}, along with /ping and /health for probes. On Kubernetes, a KServe InferenceService (API group serving.kserve.io/v1beta1) can serve the same model from object storage with autoscaling; you set a modelFormat and storageUri in the predictor spec. If containers and pods are still unfamiliar, Simpliaxis's Docker and Kubernetes training covers this layer.
Step 5: Monitor drift and close the loop
Every night, a job uses Evidently to compare the last 24 hours of inputs against the training reference set. It writes a report and, if too many features have drifted, calls the retraining workflow through the GitHub API. OurDevOps pipeline tools implementation guide goes deeper on the CI/CD half of this pipeline, including runners, artifacts and environments.
The example deliberately leaves out a feature store, approval workflows and multi-environment promotion. Add them once you have several models or teams.
LLMOps: MLOps for Generative AI Applications
LLMOps extends MLOps to applications built on large language models. Microsoft's guide to GenAIOps for organisations with MLOps investmentsis clear that you don't start over; you extend your existing practice, because the key asset is often the prompt and the orchestration around a model you didn't train. If you're new to the field, start with what generative AI is.
What changes when the model is an LLM
- Prompt versioning. System prompts and templates change behaviour as much as a retrain does. Store them in Git, tag them, and log the prompt version with every request.
- Retrieval data pipelines. For RAG systems, chunking, embedding and indexing jobs join your data pipeline, and someone has to own index freshness.
- Evaluation sets instead of a single metric. Microsoft lists groundedness, relevance, coherence and fluency as RAG evaluation metrics. Keep a fixed set of a few hundred real questions and run it on every prompt or model change.
- Guardrails. Input and output filters for harmful content, prompt injection and leaked personal data sit between the user and the model.
- Cost and latency monitoring. Microsoft recommends tracking latency, token usage and HTTP 429 (throttling) errors. Tokens per request are the LLMOps equivalent of GPU hours.
Fine-tuning a foundation model is the exception. It looks a lot like classic MLOps, with data preparation, training, evaluation, registry and deployment stages you'll already recognise.
MLOps Engineer Role: Skills and Responsibilities
An MLOps engineer builds and runs the platform and pipelines that let data scientists ship models safely, which covers CI/CD for ML, training infrastructure, model serving and monitoring. Smaller companies may call the role "ML engineer" or "ML platform engineer", and the work overlaps heavily.
| Skill area | What you need to know | Tools to practise on |
| Programming | Python for pipelines and services; enough SQL to query feature data; Bash | Python, pandas, FastAPI, pytest |
| ML fundamentals | Train/test splits, common metrics (AUC, precision, recall, RMSE), overfitting, what drift is | scikit-learn, XGBoost |
| Containers and orchestration | Writing small images, resource limits, Deployments, Services, autoscaling | Docker, Kubernetes, Helm |
| CI/CD | Pipelines triggered by code, data or schedule; quality gates; artifact promotion | GitHub Actions, GitLab CI, Jenkins, Argo CD |
| ML lifecycle tooling | Experiment tracking, registry, pipeline orchestration, serving | MLflow, DVC, Kubeflow, KServe, Airflow |
| Cloud | At least one provider's ML platform and IAM basics | SageMaker, Vertex AI or Azure Machine Learning |
| Observability | Metrics, logs, drift reports, alert design | Prometheus, Grafana, Evidently |
MLOps engineer salary figures in India vary widely between aggregators, because they rely on self-reported data. Compare AmbitionBox, Glassdoor and Naukri for your city and experience level, and give more weight to figures backed by a large, recent sample than to any single headline number. For a broader view of adjacent paths, see our guide onhow to become an AI engineer.
MLOps Roadmap: A 90-Day Plan for DevOps and Data Engineers
This MLOps roadmap assumes you work full time and can spare 8 to 10 hours a week. DevOps engineers will move faster through the infrastructure parts, and data engineers through data and orchestration.
Days 1 to 30: ML foundations for engineers
Start with supervised learning, train/validation/test splits, AUC versus accuracy, overfitting, and why a model can look good offline yet fail online. Then train two scikit-learn models end to end on a public tabular dataset (churn or loan default works well) and write down every manual step you take. That list becomes your automation backlog. If Python data work is your weak spot, give this month extra hours rather than rushing ahead.
Days 31 to 60: Tracking, versioning and packaging
Add DVC to the project, then MLflow tracking, and rerun the experiments so you can compare them in the MLflow UI. Register the best model, serve it with mlflow models serve, and build a Docker image you can run locally. Finish the month with a GitHub Actions workflow that retrains and fails on a quality gate. That single workflow is the most convincing thing you can show in an interview.
Days 61 to 90: Production on Kubernetes and monitoring
Deploy the image to a local cluster (kind or minikube), then try KServe for the same model. Add Evidently drift reports on a sample of "production" data that you shift on purpose, and trigger retraining from the drift result. Close with a one-page architecture note on retraining triggers, rollback and what you'd add for a second model, because interviewers ask about exactly these trade-offs. If Kubernetes is new to you, read anintroduction to Kubernetesbefore this phase.
If you're coming from DevOps and want to firm up the cloud CI/CD side of this transition, Simpliaxis's AWS DevOps Engineer certification training pairs well with the plan, since SageMaker pipelines and endpoints sit on the same AWS building blocks.
Common MLOps Mistakes
Most MLOps failures trace back to process gaps rather than missing tools. These six come up again and again:
- Buying a platform before you have a pipeline. Kubeflow on day one for two models adds operational load without fixing the hand-off. Start with scripts, MLflow and CI.
- Versioning code but not data. If you can't reproduce last month's training set, you can't reproduce last month's model.
- No quality gate before promotion. Automated retraining without a comparison against the current champion just automates regressions.
- Training/serving skew. Features computed in pandas for training and in Java for the API will drift apart over time. Share one feature definition.
- Monitoring only infrastructure. A model can return HTTP 200 in 40 ms and still be wrong. Track input drift and, once labels arrive, real accuracy.
- No rollback plan. Keep the previous model version deployable and know the exact alias change that restores it.
Conclusion
MLOps comes down to running machine learning as a product: versioned data and models, pipelines instead of notebooks, quality gates before release, and monitoring that triggers retraining. Start small. One model with DVC, MLflow, a CI quality gate and a drift report will teach you more than any platform demo, and the Google Cloud and Microsoft maturity models show where you stand so you can climb one level at a time.
If the ML half of the job is the newer half for you, Simpliaxis's Introduction to AI and Machine Learning trainingis a sensible place to begin.























