0% found this document useful (0 votes)
2 views10 pages

MLOps + ML System Design

The document outlines a comprehensive curriculum for designing, building, deploying, and maintaining production machine learning systems, covering the entire ML lifecycle from problem definition to monitoring and retraining. It includes modules on data engineering, feature engineering, model development, software engineering, serving architecture, containerization, orchestration, CI/CD, monitoring, scaling, optimization, cloud infrastructure, security, and responsible AI. Additionally, it provides guidelines for end-to-end ML system design and an interview checklist for assessing proficiency in these areas.

Uploaded by

Teju aka Tejaswi
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views10 pages

MLOps + ML System Design

The document outlines a comprehensive curriculum for designing, building, deploying, and maintaining production machine learning systems, covering the entire ML lifecycle from problem definition to monitoring and retraining. It includes modules on data engineering, feature engineering, model development, software engineering, serving architecture, containerization, orchestration, CI/CD, monitoring, scaling, optimization, cloud infrastructure, security, and responsible AI. Additionally, it provides guidelines for end-to-end ML system design and an interview checklist for assessing proficiency in these areas.

Uploaded by

Teju aka Tejaswi
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Production Machine Learning Systems (ML System

Design + MLOps)
Goal: Learn how to design, build, deploy, scale, monitor, and maintain production ML systems
from end to end.

Module 1 — Production ML Fundamentals


ML Lifecycle
• Problem Definition
• Data Collection
• Data Labeling
• Data Validation
• Feature Engineering
• Model Training
• Evaluation
• Deployment
• Monitoring
• Retraining

Core Concepts
• MLOps
• ML System Design
• Offline vs Online Inference
• Batch vs Streaming
• Online Learning vs Offline Learning
• Training vs Serving
• Data-Centric AI
• Reproducibility
• Technical Debt in ML

Module 2 — Data Layer


Data Sources
• Databases
• Data Warehouses
• Data Lakes

1
• Streaming Platforms
• APIs

Data Storage
• SQL
• NoSQL
• Object Storage
• Feature Store

Data Engineering
• ETL / ELT
• Data Validation
• Schema Validation
• Data Versioning
• Feature Pipelines

Module 3 — Feature Engineering Infrastructure


Feature Pipelines
• Batch Features
• Real-Time Features
• Feature Reuse
• Feature Freshness
• Point-in-Time Correctness

Feature Stores
• Feast
• Tecton
• Hopsworks

Module 4 — Model Development


Experimentation
• Hyperparameter Tuning
• Cross Validation
• Reproducibility
• Random Seeds

2
Experiment Tracking
• MLflow
• Weights & Biases

Model Registry
• Versioning
• Promotion
• Rollback

Module 5 — Software Engineering


Project Structure
• Modular Design
• Configuration Management
• Dependency Management

Code Quality
• Unit Testing
• Integration Testing
• Logging
• Error Handling
• Documentation

Module 6 — Serving Architecture


Inference Patterns
• Batch Inference
• Real-Time Inference
• Streaming Inference
• Asynchronous Inference

Serving Frameworks
• FastAPI
• BentoML
• Triton
• TorchServe

3
• TensorFlow Serving

API Design
• REST
• gRPC
• Authentication
• Versioning
• Rate Limiting

Module 7 — Containerization
Docker
• Images
• Containers
• Dockerfile
• Multi-stage Builds
• Docker Compose

Module 8 — Orchestration
Kubernetes
• Pods
• Deployments
• ReplicaSets
• Services
• Ingress
• ConfigMaps
• Secrets
• PV / PVC

Deployment Strategies
• Rolling Update
• Canary
• Blue-Green
• Shadow Deployment

4
Module 9 — CI/CD/CT
Continuous Integration
• Testing
• Linting
• Build Automation

Continuous Delivery
• Deployment
• Rollback
• Validation

Continuous Training
• Scheduled Retraining
• Trigger-Based Retraining
• Human Approval
• Automation

Module 10 — Workflow Orchestration


• Airflow
• Prefect
• Dagster
• Kubeflow Pipelines
• Vertex AI Pipelines
• SageMaker Pipelines

Module 11 — Monitoring & Observability


Infrastructure
• CPU
• GPU
• Memory
• Latency
• Throughput
• Availability

5
Model Monitoring
• Prediction Distribution
• Confidence
• Accuracy
• Business Metrics

Drift Detection
• Covariate Drift
• Label Drift
• Concept Drift

Observability
• Logging
• Metrics
• Tracing
• Alerting

Module 12 — Scaling
Scaling Training
• Distributed Training
• Multi-GPU
• Multi-Node

Scaling Inference
• Autoscaling
• Dynamic Batching
• Load Balancing
• GPU Scheduling
• Caching

Module 13 — Optimization
• Quantization
• Pruning
• Knowledge Distillation
• ONNX

6
• TensorRT
• OpenVINO

Module 14 — Cloud Infrastructure


Compute
• Virtual Machines
• Containers
• Serverless

Storage
• Object Storage
• Databases

Managed ML Platforms
• SageMaker
• Vertex AI
• Azure ML

Module 15 — Security & Responsible AI


Security
• IAM
• Secrets
• Authentication
• Authorization
• Encryption

Responsible AI
• Fairness
• Bias
• Explainability
• SHAP
• LIME
• Privacy
• GDPR

7
Module 16 — End-to-End ML System Design
Be able to design the following systems:

Recommendation System

• Candidate Generation
• Ranking
• Feature Store
• Online Serving
• Monitoring

Named Entity Recognition

• Data Pipeline
• Training
• Serving
• Retraining
• Monitoring

OCR Pipeline

• Image Processing
• Detection
• Recognition
• Deployment
• Scaling

Fraud Detection

• Streaming Data
• Feature Engineering
• Low-Latency Inference
• Drift Detection

8
Search Ranking

Document Classification

Image Classification API

Speech Recognition

LLM Inference Service

Chatbot Backend

For every design, discuss:

1. Functional requirements
2. Non-functional requirements
3. Data ingestion
4. Feature engineering
5. Training pipeline
6. Experiment tracking
7. Model registry
8. Deployment
9. API design
10. Scaling
11. Monitoring
12. Drift detection
13. Retraining
14. Rollback strategy
15. Cost optimization
16. Failure recovery

Complete Production ML Pipeline

Raw Data

Data Validation

Feature Engineering

Feature Store

Model Training

Experiment Tracking

9
Model Registry

Docker

CI/CD

Kubernetes

FastAPI / Triton

Load Balancer

Monitoring

Drift Detection

Alerts

Retraining

Model Registry

Production

Interview Checklist
You should be able to confidently answer:

• Design an end-to-end production ML system.


• Deploy an ML model serving millions of requests per day.
• Design a recommendation system.
• Design an OCR pipeline.
• Design a real-time NER service.
• Explain how to monitor production models.
• Detect and respond to data or concept drift.
• Roll back a faulty model deployment.
• Scale inference under heavy traffic.
• Build an automated CI/CD/CT pipeline for ML.
• Optimize latency, throughput, reliability, and cost.

10

You might also like