100 Machine Learning Algorithm Interview Questions
Linear & Logistic Regression
1. What is the cost function used in linear regression?
2. How does gradient descent optimize linear regression?
3. What are the assumptions of linear regression?
4. How do you detect multicollinearity in regression models?
5. Difference between OLS (Ordinary Least Squares) and Gradient Descent.
6. What is regularization? How do Ridge and Lasso differ?
7. When would you use ElasticNet instead of Ridge or Lasso?
8. How does logistic regression perform classification?
9. What is the logit function?
10. How is decision boundary formed in logistic regression?
Decision Trees
11. How does a decision tree split nodes?
12. Difference between Gini Index and Entropy.
13. What is information gain?
14. What is overfitting in decision trees and how to prevent it?
15. Explain pruning in decision trees.
16. Difference between ID3, C4.5, and CART algorithms.
17. What is the depth parameter in decision trees?
18. How does variance affect decision trees?
19. Explain greedy nature of decision tree splitting.
20. What is the limitation of decision trees?
Ensemble Methods
21. What is bagging? Give an example.
22. What is boosting?
23. Explain random forest algorithm.
24. How does random forest reduce overfitting?
25. Explain out-of-bag (OOB) error in random forests.
26. What is feature importance in random forests?
27. Explain AdaBoost working principle.
28. What is the role of weak learners in AdaBoost?
29. Explain Gradient Boosting algorithm.
30. Difference between Gradient Boosting and AdaBoost.
31. How does XGBoost improve over traditional boosting?
32. What is histogram-based gradient boosting (used in LightGBM)?
33. Difference between XGBoost, LightGBM, and CatBoost.
34. What is stacking in ensemble learning?
35. What are advantages and disadvantages of ensemble methods?
Support Vector Machines
36. What is the objective of SVM?
37. Explain maximum margin classifier.
38. What is the role of kernel functions in SVM?
39. Difference between linear and non-linear SVM.
40. Explain RBF kernel.
41. What is the effect of gamma parameter in SVM?
42. What is the role of C parameter in SVM?
43. How does SVM handle multi-class classification?
44. Difference between hard margin and soft margin SVM.
45. Why is SVM not suitable for very large datasets?
Naive Bayes
46. What is the Bayes theorem formula?
47. Why is Naive Bayes called “naive”?
48. What are the main types of Naive Bayes classifiers?
49. Explain Gaussian Naive Bayes.
50. Explain Multinomial Naive Bayes.
51. Explain Bernoulli Naive Bayes.
52. Why does Naive Bayes work well with text data?
53. How do you handle zero probability in Naive Bayes?
54. What is Laplace smoothing?
55. Compare Naive Bayes with Logistic Regression.
K-Nearest Neighbors (KNN)
56. How does KNN algorithm work?
57. What is the role of K in KNN?
58. How do you choose the optimal K value?
59. Explain the effect of distance metrics in KNN.
60. Difference between Euclidean, Manhattan, and Minkowski distances.
61. Why is KNN considered a lazy learner?
62. How does KNN handle high-dimensional data?
63. What are advantages and disadvantages of KNN?
64. How does KNN handle ties during classification?
65. Difference between weighted KNN and standard KNN.
Clustering Algorithms
66. What is the objective of K-means clustering?
67. Explain the K-means algorithm steps.
68. How do you choose the number of clusters in K-means?
69. What is the K-means++ initialization?
70. What are limitations of K-means clustering?
71. Explain K-medoids clustering.
72. What is hierarchical clustering?
73. Difference between agglomerative and divisive clustering.
74. Explain DBSCAN clustering algorithm.
75. Difference between DBSCAN and K-means.
Dimensionality Reduction Algorithms
76. What is the objective of PCA?
77. How does PCA work?
78. Explain eigenvalues and eigenvectors in PCA.
79. What is explained variance in PCA?
80. Difference between PCA and LDA.
81. What is t-SNE used for?
82. What is the difference between t-SNE and PCA?
83. Explain UMAP algorithm.
84. When would you use dimensionality reduction?
85. Limitations of PCA in high-dimensional data.
Neural Network Algorithms
86. Explain perceptron algorithm.
87. What is backpropagation?
88. Explain gradient descent in neural networks.
89. What is the role of activation functions?
90. Difference between Sigmoid, Tanh, and ReLU.
91. What is dropout in neural networks?
92. Explain CNN architecture.
93. Explain RNN architecture.
94. What are vanishing gradients?
95. Difference between supervised deep learning and unsupervised deep learning
algorithms.
Reinforcement Learning Algorithms
96. What is Q-learning?
97. Explain Bellman equation.
98. What is the difference between SARSA and Q-learning?
99. What is the policy gradient algorithm?
100. Explain Deep Q-Network (DQN).