๐ Model Lifecycle and Deployment
Descriptionโ
< What is it? >โ
- The model lifecycle covers defining a task, collecting data, preparing features, training, evaluation, deployment, and ongoing monitoring. It applies to classifiers, regression models, recommendation systems, and other machine-learning applications.
- Deployment makes a trained model usable by an application, through a batch job or an online service. The model is one part of a system that also includes preprocessing, application logic, and monitoring. The Google Machine Learning Crash Course explains this broader production workflow.
Key pointsโ
< The end-to-end journey >โ
The ML model lifecycle spans from understanding the business problem all the way to monitoring the deployed model in production. Here's the end-to-end journey.
1. Scoping & Problem Definition
2. Data Preparation & EDA
3. Experimentation & Model Training
4. Model Validation & Registration
5. Deployment & Serving
6. Monitoring & Maintenance
-
๐ฏ 1. Scoping & Problem Definition
Before any code, you must define what the model needs to do and how you'll know it's working.
Key Tasks: Determine if ML is the right solution. Define prediction targets (classification, regression, etc.) and establish success metrics like accuracy, AUC, or business KPIs.
Output: A design document outlining the ML approach and serving requirements (latency, throughput).
-
๐ 2. Data Preparation & EDA
This phase ensures your data is clean, understood, and ready for modeling.
Exploratory Data Analysis (EDA): Visualize distributions, spot missing values, and identify correlations to inform feature choices.
Data Preprocessing: Clean data (remove duplicates, outliers), handle missing values (imputation), scale features (normalization), and split data into train/validation/test sets to prevent leakage.
Feature Engineering: Create, transform, and select features (e.g., one-hot encoding, PCA) to maximize predictive power. Use a Feature Store to ensure consistency between training and serving.
-
๐งช 3. Experimentation & Model Training
This is the iterative core where you build and tune models.
Model Selection: Start with simpler algorithms (e.g., gradient boosted trees) before moving to deep learning if needed.
Training & Hyperparameter Tuning: Train models, evaluate performance against metrics, and tune hyperparameters (learning rate, depth) to find the optimal architecture.
Experiment Tracking: Use tools like MLflow or SageMaker Experiments to log parameters, metrics, and artifacts for reproducibility.
-
โ 4. Model Validation & Registration
Before deployment, you must rigorously validate the model and register it as a deployable asset.
Validation: Evaluate on a held-out test set. Check for bias using tools like SageMaker Clarify and ensure explainability.
Model Registry: Register the model with its metadata (version, metrics, training data lineage) to create a single source of truth for deployment and rollback.
-
๐ 5. Deployment & Serving
This is where the model moves to production to serve predictions.
Packaging: Package the model (e.g., as a Docker container) with its dependencies.
Serving Infrastructure: Choose based on needs: Kubernetes (full control), Serverless (scale-to-zero for spiky traffic), or Edge (low latency).
Deployment Strategies: Use canary rollouts (small % of traffic) or A/B testing to safely promote the new model without risking the entire user base.
-
๐ก 6. Monitoring & Maintenance
The lifecycle doesn't end at deployment. You must monitor for degradation and maintain the system.
Monitor Key Metrics: Track data drift (input distribution changes), model performance (accuracy, latency), and infrastructure health (error rates).
Retraining Triggers: Set up alerts for sustained drift or performance drops. You can trigger retraining manually or automatically (scheduled or event-driven) to keep the model fresh.
Continuous Improvement: Use the feedback loop to collect new data, retrain, and redeploy, starting the cycle again.
-
๐ก Key Takeaway
The lifecycle is iterative, not linear. If monitoring reveals issues, you circle back to data preparation or experimentation. The goal is to build a closed-loop system where the model continuously learns and adapts
< Batch and online inference >โ
- Batch inference produces predictions for many records on a schedule, such as a nightly demand forecast. Store outputs for downstream applications to read.
- Online inference produces a prediction when a request arrives, such as classifying a submitted document. The service loads the model once and reuses it across requests.
- Hybrid inference precomputes reusable work offline and performs request-dependent work online. Two-tower recommendation retrieval is one example: precompute item embeddings, then compute a user embedding at request time.
Implementationโ
< Package the model and its dependencies >โ
-
Example PyTorch deployment bundle: keep these artifacts with the application code that implements the architecture and preprocessing. Other frameworks use their own model formats.
model-v1/โโโ model.ptโโโ model_config.jsonโโโ preprocessing.jsonโโโ vocabularies.jsonโโโ release_manifest.jsonmodel.ptcontains trained weights. The configuration describes the architecture; preprocessing records transformations; vocabularies preserve category-to-ID mappings when needed. The release manifest identifies compatible artifact versions. Pin application dependencies in the container build. -
Load once at startup: for PyTorch, recreate the architecture, load its
state_dict, callmodel.eval(), and disable gradient tracking during inference, for example withtorch.inference_mode(). See PyTorch's saving and loading guide.
< Serve a PyTorch model with FastAPI and Docker >โ
-
A simple deployment stack: use these tools to serve a small trained MLP. This is one implementation choice; the lifecycle also applies to other model frameworks.
Tool Role PyTorch Loads the trained weights and executes the model FastAPI Exposes an endpoint such as POST /predictand validates request fieldsDocker Packages application code, runtime, and dependencies for a deployment server -
Request flow: validate input, apply saved preprocessing, run the model, perform task-specific postprocessing, and return predictions. Start with CPU inference for a small MLP and benchmark latency and throughput. Docker packages the service; a server or container platform runs it. See FastAPI's container deployment guide.
-
Release and rollback: deploy matching model and preprocessing versions together. Test the packaged release, introduce it gradually where appropriate, and retain the previous complete release for rollback.
< Worked example: two-tower recommendation deployment >โ
-
Training data: for the runnable two-tower example, collect user ID, item ID, impression time, click/no-click labels, and historical profile features. Allow the click-label observation window to complete before labeling an impression unclicked. Use older interactions for training and later periods for validation and testing.
-
Evaluation: measure Recall@K and NDCG@K against a baseline. Record the candidate catalog and negative-sampling policy. For a separate click-probability ranker, also evaluate loss and calibration.
-
Offline and online flow: the item tower can run in a batch job while the user tower runs inside the recommendation service.
Offline:Item profiles โ trained item tower โ item embeddings โ search indexOnline:User request โ user features โ trained user tower โ user embeddingโSearch item indexโFilter / optionally rerankโReturn top item IDs -
Search tools: add FAISS to the PyTorch/FastAPI/Docker stack when the catalog benefits from a vector index. FAISS is a search library, not an HTTP server. For cosine retrieval, normalize both user and item vectors and search by inner product. A small catalog can use direct matrix multiplication; larger catalogs may benefit from approximate search. See FAISS getting started.
-
Additional artifacts: add
item_index.faissanditem_ids.jsonto the deployment bundle; the latter maps index results to catalog IDs. Record their versions alongside the model. Serve recommendations through an endpoint such asPOST /recommend. -
Keep the embedding space consistent: when retraining changes the towers, regenerate item embeddings before switching traffic. Release the matching user tower, item index, and preprocessing together; rollback must restore the matching index too.
-
Keep the catalog current: encode new or changed items with the deployed item tower, update the index, and remove or filter unavailable items. Use a fallback such as eligible popular items when user information is insufficient. Monitor index freshness and recommendation outcomes in addition to service health.
Related ideasโ
- Recommendation System explains retrieval, ranking, and two-tower training.
- Multilayer Perceptron (MLP) describes neural networks that can be served through the example stack.
- Embeddings demonstrates encoding structured user and item profiles.
- Train, Validation, and Test Sets explains data splitting.
- Confidence Calibration covers the reliability of predicted probabilities.
- Model Drift covers changes after deployment.
Referenceโ
- Machine learning Lifecycle (geeksforgeeks.org)
- MLflow Tracking
- MLflow Model Registry
- Amazon SageMaker Experiments
- Amazon SageMaker Clarify: Model explainability
- Feast documentation
- Google Machine Learning Crash Course: Production ML systems
- PyTorch: Saving and loading models
- FastAPI: Containers and Docker
- FAISS: Getting started