Machine Learning
Dimensions
• Statisticians
• Data Science Theorists
• Business Applications
• Functional Managers
• Model Users
Coverage
• End to End Machine Learning Problem- Classification
• With clean and structured data
• End to End Machine Learning Problem- Classification
• With missing data and unstructured outliers
• Cross Validation
• Feature Engineering
• Hyper Parameters Tuning
• ROC and AUC
• End to End Machine Learning Problem- Clustering
• Basic Clustering
• Using non numeric data
• Combining Clustering with Classification (Ensemble)
• End to End Machine Learning Problem- Regression
• Machine Learning as a Service
• Azure Machine Learning Studio
End to End Machine Learning
• Models Covered
• Naïve Bayes
• QDA
• Logistic Regression
• Classification- RPART
• Random Forest
• Support Vector Machine
• Gradient Boosting
• XGB
• kNN
• Basic Process
• Use each of this method
• Develop an evaluation matrix
• Compare the models and choose the most appropriate
Consolidated Machine Learning (IBP)
Integrated View of ML
• Training Task
Make Task • Testing Task
• Any of the Cluster, Classif, Regr
Make a Learner • Define Predict Method
Train Learner
for Training Task
Predict on Test
Task
Evaluate
Practical Hands on
• Gestational Diabetes
• Dataset Description
Describing Confusion Matrix
Prior Probability
Good Customers 20 20/30 0.667
Bad Customers 10 10/30 0.333
OPMD- Naïve Bayes Total Customers 30
Define Problem
Likelihood of the new customer given good: 3
20
1
Likelihood of the new customer given bad:
10
3 20
Posterior Probability that customer is good: ∗ =0.1
20 30
Posterior Probability that customer is bad: 1 10
∗ =0.03
10 30
StatQuest with Josh Starmer
OPMD- Discriminant Analysis
OPMD- Logistic Regression
OPMD- Classification
OPMD- Random Forest
OPMD- SVM
Tuning parameters: Type, Kernel and gamma
OPMD- GB and XGB
OPMD- kNN
Real Life Examples of Machine
Learning-1
• Identifying Top 10 Songs
• Dataset Description
• What a Data Scientist is Expected to do?
• General: Build model with Train dataset and predict which songs would come
Top 10 from the Test dataset. Assess the accuracy.
• Specific:
• Which un-tuned algorithm works the best?- (Try out all)
• Tune any one of the algorithm and see the performance.
Real Life Examples of Machine
Learning-1
• Stock Movement Predictions
• Dataset Description
• What a Data Scientist is Expected to do?
• Script is provided for the following activities:
• Basic un-tuned algorithms
• Clustering the stocks with k=3
• Splitting the original dataset into 3 clusters.
• A trial calculations of model improvement with RandomForest and SVM. Since the
improvements are observed, you have to do the following:
• Re-build all the models with Cluster 1 as data. Which models show
improvement and which do not? Why?