Contents
1. Useful URL on ML..........................................................................................................................3
2. LLM job descriptions.....................................................................................................................3
3. AI (Artificial intelligence), ML, DL and GENAI and LLM..................................................................5
4. Generative AI................................................................................................................................5
5. Variance , standard deviation, covariance , and correlation..........................................................6
6. standardization vs Normalization..................................................................................................8
7. Understanding Bias, Variance, Overfitting, and Underfitting......................................................11
8. Differences between Data Cleaning and Data Wrangling...........................................................14
9. How to avoid overfitting..............................................................................................................15
10. what are metric or method to identify underfitting or overfitting..........................................16
11. What is Regularization in Machine Learning? Why is Regularization Important?...................20
12. Weights and biases in liner regression....................................................................................22
13. P-Values and Confidence Intervals in Regression....................................................................23
14. Regularization.........................................................................................................................26
15. Loss function and metric table................................................................................................27
16. ***machine learning algorithms along with a brief explanation and an example of each:....30
17. Bayes theorem........................................................................................................................32
18. Basics related to Machine learning model..............................................................................33
19. what is ANOVA , explain with simple example........................................................................35
20. hypothesis testing, regression analysis, ANOVA, experimental design, and probability theory
38
21. Ensemble technique...............................................................................................................41
22. What is AI...............................................................................................................................43
23. Evaluating machine learning models........................................................................45
24. Question-1-> What is Supervised Learning and What is Unsupervised Learning?..................47
1. Question-2-> Can you tell me the difference between A Data Analyst and A Data Scientist?.....49
2. Question-4-> Define p-value and tell me why it is important?....................................................50
3. Question-5-> What is PDF and CDF and why do you think it’s important in Machine Learning? 51
4. Question-6-> Can you explain the Central Limit theorem?.........................................................52
5. Question-7-> What are data structures and what data structures we have in python?..............53
6. Question-8-> What are CNNs, Can you explain any famous CNN architectures?........................54
7. Question-9-> What is Max Pooling and why do we need Max Pooling?.....................................55
8. Question-10-> What is NLP?.......................................................................................................55
9. Question-11-> What are the algorithms to embed a sentence into a vector?............................56
10. Question-12-> What’s your favourite algorithm?...................................................................56
11. Question-13-> What is a hyper parameter?............................................................................57
12. Question-14-> What is Under-fitting.......................................................................................58
13. Question-15->What is Over-Fitting ?......................................................................................58
14. Question-16-> What is model error?......................................................................................59
15. Question-17-> What is One-Hot-Encoding?............................................................................59
16. Question-18-> What is Multi-class classification?...................................................................60
17. Question-19->Can we use all the algorithm if we have Multi-class classification problem?...60
18. Question-20-> What is the performance metric you will use if you have a medical related
problem like “Cancer Prediction”?.......................................................................................................61
19. Question-21-> What is F1 Score?............................................................................................61
20. Question-23-> What is NLTK?.................................................................................................62
21. Question-24-> What is Dropouts?..........................................................................................62
22. Question-25-> What is Batch Normalisation?.........................................................................63
23. Question-26-> What is a Perceptron?.....................................................................................64
24. Question-27-> How logistic Regression is a single neuron model?.........................................65
25. Question-28-> Why the name of logistic regression has regression when it’s a classification
technique.............................................................................................................................................66
26. Question-29-> What is a sigmoid function? Why it is used in logistic regression?..................66
27. Question-30-> What do you mean by a hyperplane and how hyperplane is important with
respect to machine learning?..............................................................................................................67
28. Question-31-> What is Backpropogation?..............................................................................68
29. Question-32-> What is Gradient Descent?..............................................................................68
30. Question-33->What is the problem with gradient descent?...................................................69
31. Question-34-> What is SGD(Stochastic Gradient Descent)?...................................................69
32. Question-35-> Explain the problem of Vanishing Gradient?...................................................70
33. Question-36-> What are the different activation function we used in deep learning?...........71
34. Question-37-> How to deal with Outliers in Machine Learning?............................................72
35. Question-38-> What will you do if you model is overfit?........................................................72
36. Question-39-> What is RFR(Randomisation For Regularisation)?...........................................73
37. Question-40-> Can you explain the difference between Convex and Non Convex function?. 73
38. Question-41-> Find the majority element in a list..................................................................74
39. Question-42-> Rotate the array by d elements.......................................................................75
40. Question-43->Given a string, Reverse the order of strings in each word within a sentence
while preserving white space and initial word order...........................................................................76
41. Question-44-> Given 2 strings, you need to find the common elements from both the strings.
76
42. Question-45-> Write a code in python to print the pattern....................................................77
43. Question-46-> What is the search time complexity of a list and a dictionary?.......................77
44. Question-47->Write a python program to define power function from scratch?...................77
45. Question-48-> Write a python function to find the cubic sum only by recursion?.................78
46. Question-49-> Write a program to find the dot product in python from scratch?..................78
47. Question-50-> Write a python program which gives us euclidian distance from scratch?......79
1. Useful URL on ML
[Link]
2. LLM job descriptions
Job Purpose:
We are seeking an experienced and innovative Senior Machine Learning Engineer to lead the
development of advanced machine learning models and AI-driven solutions. This role involves
working with state-of-the-art large language models (LLMs), multilingual NLP, and reinforcement
learning methods to develop impactful and scalable applications.
Key Responsibilities:
• Develop, implement, and optimize machine learning algorithms and models, including LLMs, for a
variety of applications.
• Preprocess, curate, and manage large datasets, ensuring quality for training and evaluation
purposes.
• Train and fine-tune LLMs (e.g., LLaMA, Gemma) using frameworks such as DeepSpeed and
Accelerate, incorporating techniques like LoRA, QLoRA, and quantization.
• Apply reinforcement learning techniques (e.g., PPO, DPO) to improve model performance.
• Evaluate models using benchmarks such as MMLU, ACVA, MGSM, and ZeroSCROLLS, with an
emphasis on domain-specific datasets and multilingual capabilities (including Arabic).
• Deploy models in production, optimize for latency and throughput using tools like vLLM, Triton
Inference, and TensorRT, and implement monitoring systems to track performance.
• Address Arabic NLP challenges, including tokenization, morphology, syntax, and dialectal variations.
• Stay updated with the latest advancements in AI, machine learning, and NLP research.
• Provide mentorship to junior engineers and data scientists, fostering a collaborative and innovative
team environment.
• Collaborate with cross-functional teams to propose and deliver strategic AI solutions to business
challenges.
Qualifications and Experience:
• 6+ years of experience in machine learning, with at least 3 years in a leadership role.
• Expertise in training and deploying LLMs, with hands-on experience in fine-tuning, quantization,
and model optimization.
• Strong understanding of multilingual models and NLP challenges, particularly with Arabic language
processing (e.g., dialects, tokenization, diacritics).
• Proficiency in Python and frameworks such as PyTorch, Hugging Face Transformers, and FastAPI.
• Familiarity with MLOps tools like ClearML, MLflow, and Weights & Biases for model tracking and
versioning.
• Experience in cloud platforms (AWS, GCP, Azure) and containerization (Docker).
• Solid understanding of reinforcement learning, benchmarking techniques, and deployment
pipelines.
• Native fluency or advanced proficiency in Arabic is highly desirable.
• Strong Python scripting and automation capabilities.
• Knowledge of model optimization techniques (e.g., gradient accumulation, learning rate
schedules).
• Familiarity with academic research and advancements in NLP and LLMs.
• Proficiency with version control systems like Git/GitHub.
Responsabilities :
Proficiency in programming languages like Python
Skills in handling and analysing large datasets, including data cleaning, wrangling, transformation,
and exploratory data analysis (EDA) using tools like pandas.
Knowledge of statistical concepts such as hypothesis testing, regression analysis, ANOVA,
experimental design, and probability theory.
Understanding machine learning algorithms, supervised and unsupervised learning, feature
engineering, model selection and evaluation, and hyperparameter tuning is essential, with libraries
like scikit-learn, TensorFlow, or PyTorch commonly used. Knowledge of YOLO is desirable
Data visualisation skills using libraries like Matplotlib, Seaborn, or ggplot
Familiarity with big data technologies like Apache Hadoop, Apache Spark, or distributed databases
enables the processing and analysis of large-scale datasets.
Proficiency in SQL and database management
Knowledge of data engineering concepts, including data pipelines, integration, warehousing, and
architecture, is beneficial.
Cloud computing platforms (Azure) and services like Azure Machine Learning can be advantageous.
Proficiency in version control systems like Git is crucial for effective collaboration and code
management.
Requirements :
6 to 10 years of experience in projects within the manufacturing industry (aluminum, steel, metals,
mining,...).
Good proficiency in English, both written and spoken."
-------------------------------
3. AI (Artificial intelligence), ML, DL and GENAI and
LLM.
Artificial Intelligence (AI):
The science and engineering of creating computer systems capable of performing tasks that
typically require human intelligence, such as decision-making, language understanding, or
problem-solving.
Machine Learning (ML):
A subset of AI that enables systems to automatically learn from data or the environment,
improving performance or decision-making over time without explicit programming.
Deep Learning (DL):
A specialized subset of ML focused on algorithms modeled after the human brain's neural
networks.
Deep learning involves multiple layers of interconnected nodes (artificial neurons) that can
learn to identify patterns and make decisions.
It mimics the logical structure of human cognition to analyze data and draw insights.
Data Science:
A multidisciplinary field aimed at extracting meaningful insights and knowledge from data
through statistical, computational, and visualization techniques.
Unlike AI and ML, which focus on automation and predictive modeling, data science
emphasizes understanding the data itself.
Many tools and methods overlap with those used in AI and ML
4. Generative AI
Generative AI refers to a class of artificial intelligence systems that have the
capability to generate new content, data, or outputs
based on patterns and information they have learned from existing examples.
Unlike traditional AI systems that are rule-based or rely on explicit
programming, generative AI systems leverage machine learning techniques to
generate content autonomously.
There are different types of generative AI models, and they are often
categorized based on their specific tasks. Two common types of generative AI
models are:
1. Generative Adversarial Networks (GANs):
GANs consist of two neural networks, a generator, and a
discriminator, that are trained simultaneously through
adversarial training. The generator creates new data instances,
and the discriminator evaluates whether these instances are
real or generated. This iterative process helps the generator
improve its ability to create realistic content.
2. Variational Autoencoders (VAEs):
VAEs are another type of generative AI model that learns a
probabilistic mapping between input data and a latent space.
VAEs are commonly used for generating new data samples,
particularly in image generation and other creative applications.
Generative AI is employed in various domains, including:
Image Generation: Creating realistic images, artwork, or photos.
Text Generation: Generating human-like text, such as articles,
stories, or poetry.
Music Composition: Creating new music compositions.
Video Game Content: Generating elements like characters, levels, or
scenarios in video games.
Data Augmentation: Creating synthetic data for training machine
learning models.
Drug Discovery: Generating molecular structures for potential new
drugs.
Generative AI has shown significant advancements in recent years, but it also
raises ethical concerns, particularly when it comes to creating realistic
deepfake content or other potentially malicious uses. Therefore, the
responsible and ethical development and deployment of generative AI
technologies are crucial considerations.
Large Language Model:
In the context of artificial intelligence and natural language
processing, "LLM" might stand for "Large Language Model." Models
like GPT-3 (Generative Pre-trained Transformer 3) fall under the
category of large language models. These models are trained on
extensive datasets to understand and generate human-like
text.
5. Variance , standard deviation, covariance , and
correlation
6. standardization vs Normalization
Both Standardization and Normalization are feature scaling techniques used to transform raw data
into a form that is easier for machine learning models to work with. The choice depends on the
problem and data type. Let’s break it down:
When to Use Which?
1. Standardization:
o Data doesn't have a clear range or has extreme differences in scales (e.g., age in
years vs. income in millions).
o Used in algorithms sensitive to variance or that assume a normal distribution (e.g.,
logistic regression, PCA, SVM).
2. Normalization:
o Data must be scaled to a specific range (e.g., pixel intensities in image data).
o Useful for algorithms that require feature values to be on the same scale (e.g., k-NN,
neural networks).
Summary in Simple Terms
Standardization: Think of it as "centering" your data to look like a bell curve with an average
of 0 and uniform spread.
Normalization: Think of it as "squeezing" your data into a box between 0 and 1.
Analogy
Imagine you’re comparing heights in a room:
Standardization: Makes everyone stand on the same flat surface (mean = 0) and adjusts for
uniformity (variance = 1).
Normalization: Puts everyone’s height into a fixed range (0 = shortest, 1 = tallest), regardless
of their actual values.
Both approaches make comparisons easier for the model to process!
7. Understanding Bias, Variance, Overfitting, and
Underfitting
key concepts in machine learning that affect a model's performance and generalization.
1. Bias
Definition: Bias refers to the error introduced by approximating a real-world problem
(which might be complex) by a simplified model.
High Bias: The model is too simple to capture the underlying patterns, leading to
underfitting.
Low Bias: The model is complex enough to capture the data well.
Example:
A straight line (y=mx+cy = mx + cy=mx+c) trying to fit a curved dataset results in high bias
since it oversimplifies the problem.
2. Variance
Definition: Variance measures how much the model’s predictions change if
trained on different subsets of the data. It reflects the model's sensitivity to small
fluctuations in the training data.
High Variance: The model learns even the noise in the training data, leading to overfitting.
Low Variance: The model is robust to small changes in the training data.
Example:
A highly complex polynomial trying to fit data might exactly pass through all data points but
fail to generalize to new data.
3. Underfitting
Definition: When a model is too simple to capture the underlying patterns in the data. This
happens due to high bias.
Cause: The model lacks the complexity required to learn from the data.
Indicators:
Poor performance on both training and test data.
Example Models:
Linear regression on a dataset with a non-linear relationship.
Decision trees with very low depth.
4. Overfitting
Definition: When a model learns both the underlying patterns and the noise in the training
data. This happens due to high variance.
Cause: The model is too complex relative to the amount of training data.
Indicators:
Excellent performance on training data but poor performance on test data.
Example Models:
Deep neural networks trained for too many epochs without regularization.
Decision trees with no depth restriction.
Relationship Between Bias and Variance
Bias Variance Model Complexity Performance
Poor on both training and test data
High Bias Low Variance Simple Model
(underfitting).
Excellent on training, poor on test data
Low Bias High Variance Complex Model
(overfitting).
Moderate Moderate Optimal Model
Good on both training and test data.
Bias Variance Complexity
Key Tradeoff:
Reducing bias often increases variance and vice versa. The goal is to find the right balance.
How to Identify and Address Each Issue
Problem Symptoms Solutions
Poor training and test
Underfitting - Use a more complex model.
performance.
- Add more features.
Excellent training but poor test
Overfitting - Use regularization (e.g., L1, L2).
performance.
- Reduce model complexity (e.g., pruning decision
trees).
- Use techniques like dropout in neural networks or
early stopping.
Systematic error across training - Use a more flexible/complex model (e.g., polynomial
High Bias
and test data. regression, neural networks).
High Model is too sensitive to
- Increase training data.
Variance training data.
- Use techniques like cross-validation to ensure
stability.
Relevant Models for Bias-Variance Tradeoff
Model Bias Variance
Linear Regression High Bias Low Variance
Polynomial Regression (High Degree) Low Bias High Variance
Model Bias Variance
Decision Trees (Shallow Depth) High Bias Low Variance
Decision Trees (Deep Depth) Low Bias High Variance
Random Forests Moderate Bias Moderate Variance
Neural Networks Low Bias High Variance
Visualizing Bias, Variance, Underfitting, and Overfitting
1. High Bias (Underfitting):
The model is too simple, failing to capture the complexity of the data.
Example: Linear regression on a non-linear dataset.
Visual:
Data points are scattered far from the prediction curve or line.
2. High Variance (Overfitting):
The model is too complex, capturing noise along with the data.
Example: A high-degree polynomial that bends excessively to pass through every point.
Visual:
Prediction curve is overly wiggly, passing exactly through the data points but failing on new data.
3. Balanced Model:
The model generalizes well, capturing the main patterns in the data without overfitting.
Example: A well-tuned random forest.
Visual:
Prediction curve fits the data points well without excessive bending.
Key Takeaways
Underfitting: Fix by increasing model complexity or adding more features.
Overfitting: Fix by reducing model complexity, regularization, or increasing data.
Bias-Variance Tradeoff: Find the sweet spot where bias and variance are balanced for good
generalization.
[Link] between Data Cleaning and Data
Wrangling
Aspect Data Cleaning Data Wrangling
Definition The process of identifying and correcting errors, The process of transforming raw data into a
Aspect Data Cleaning Data Wrangling
inconsistencies, and inaccuracies in the data. structured, usable format for analysis.
Ensures data quality by removing errors,
Prepares data for analysis by transforming and
Purpose handling missing values, and standardizing
integrating it into a usable structure.
formats.
A subset of data wrangling focused on Broader process that includes cleaning,
Scope
correcting and validating data. reshaping, combining, and enriching data.
- Reshaping data (e.g., pivoting, melting)
- Removing duplicates
- Combining datasets (e.g., joins, merges)
- Handling missing or null values
Key Tasks - Creating calculated fields
- Correcting data entry errors
- Handling complex transformations (e.g.,
- Standardizing formats (e.g., dates, units)
parsing text)
Tools & Deduplication, imputing missing values, Data normalization, aggregation, integration,
Techniques applying validation rules. reshaping, and advanced transformations.
Produces analysis-ready data in a structured
Outcome Produces accurate and error-free data.
and meaningful format.
Correcting typos in a dataset where "NYC" is Combining sales data from multiple regions
Example
written as "NYY." into a single dataset with consistent formats.
9. How to avoid overfitting
Simplify the Model:
Use fewer features or variables.
Choose simpler models with fewer parameters (e.g., linear regression over polynomial
regression when appropriate).
Cross-Validation:
Use techniques like k-fold cross-validation to ensure the model generalizes well across
different subsets of data.
from sklearn.model_selection import cross_val_score
scores = cross_val_score(model, X, y, cv=5)
print("Cross-validation scores:", scores)
print("Mean score:", [Link]())
Regularization:
Apply penalties to discourage complex models:
o L1 Regularization (LASSO): Shrinks less important feature coefficients to zero,
promoting sparsity.
o L2 Regularization (Ridge): Penalizes large coefficients, preventing over-complex
models.
from sklearn.linear_model import Lasso, Ridge
lasso = Lasso(alpha=0.1) # Adjust alpha to control regularization strength
ridge = Ridge(alpha=0.1)
Increase Dataset Size:
Collect more data or use data augmentation techniques to expand the dataset artificially.
from [Link] import resample
augmented_data = resample(data, replace=True, n_samples=1000, random_state=42)
Feature Selection:
Remove irrelevant or noisy features to reduce dimensionality and prevent overfitting.
Use feature selection techniques like correlation analysis, Recursive Feature Elimination
(RFE), or LASSO.
Data Augmentation:
Introduce noise or transformations to the dataset to make the model more robust to
variations.
Examples include flipping, cropping, rotating, or adding Gaussian noise (for images).
Early Stopping:
Monitor the validation loss during training and stop when it starts to increase.
from [Link] import EarlyStopping
early_stopping = EarlyStopping(monitor='val_loss', patience=5)
Ensemble Methods:
Combine multiple models to reduce overfitting by averaging their predictions:
o Bagging: Builds multiple models on different subsets of data (e.g., Random Forest).
o Boosting: Builds sequential models that correct the errors of previous models (e.g.,
Gradient Boosting, XGBoost).
from [Link] import RandomForestClassifier
rf = RandomForestClassifier(n_estimators=100, random_state=42)
10. what are metric or method to identify
underfitting or overfitting.
evaluating how a model performs on both training data and test data. The key is to measure the gap
between training and test performance and the absolute performance on both datasets.
Key Metrics
1. Training and Test Accuracy/Error
Purpose: Compare the model's performance on training and test datasets.
Signs of Issues:
o Underfitting:
Low accuracy (or high error) on both training and test sets.
The model is too simple to capture patterns in the data.
o Overfitting:
High accuracy on the training set but low accuracy on the test set.
The model memorizes the training data but fails to generalize.
2. Cross-Validation
Purpose: Evaluate the model's ability to generalize across different subsets of the data.
Signs of Issues:
o Underfitting: Poor performance across all cross-validation folds.
o Overfitting: High variance in performance between folds, with training error much
lower than validation error.
3. Learning Curve
Purpose: Visualize how performance changes as more training data is added.
How it Works:
o Plot training and validation error vs. the amount of training data.
Signs of Issues:
o Underfitting:
Training and validation errors are both high and remain close to each other
as data increases.
o Overfitting:
Training error is low, but validation error is high and does not improve with
more data.
4. Precision, Recall, and F1-Score
Purpose: Evaluate classification performance, especially for imbalanced datasets.
Signs of Issues:
o Underfitting: All metrics (precision, recall, F1-score) are low for both training and
test sets.
o Overfitting: Metrics are high for training data but significantly lower for test data.
5. Residual Analysis (Regression Models)
Purpose: Analyze the difference between predicted and actual values.
Signs of Issues:
o Underfitting: Residuals show clear patterns, indicating the model failed to capture
relationships.
o Overfitting: Residuals are close to zero on training data but erratic on test data.
6. AUC-ROC Curve (Classification Models)
Purpose: Measure the ability of the model to differentiate between classes.
Signs of Issues:
o Underfitting: Low AUC on both training and test datasets.
o Overfitting: High AUC on training data but significantly lower on test data.
7. Regularization Metrics
Use regularization techniques (like L1/L2) to observe performance changes.
Signs of Issues:
o Overfitting is indicated if performance improves significantly with regularization.
Examples of Diagnosis
Scenario 1: Underfitting
Training Accuracy: 65%
Test Accuracy: 60%
Observation: Both training and test accuracy are low, indicating the model is too simple to
learn the data.
Scenario 2: Overfitting
Training Accuracy: 98%
Test Accuracy: 75%
Observation: The model performs well on training data but poorly on unseen data,
suggesting overfitting.
Methods to Identify Overfitting or Underfitting
Method Purpose Underfitting Sign Overfitting Sign
Compare Training vs. Test High training accuracy but
Low accuracy on both.
Accuracy/Error performance. low test accuracy.
Visualize model
Both training and Training error low,
Learning Curve performance as data
validation errors are high. validation error high.
increases.
High variance in
Measure generalization Poor performance across
Cross-Validation performance between
across data splits. folds.
folds.
Residuals good on
Evaluate prediction errors Residuals show patterns or
Residual Analysis training but erratic on test
(regression). trends.
data.
Assess classification Low precision and recall for Metrics high for training
Precision/Recall
performance. both training and test sets. but low for test.
How to Fix
Underfitting Solutions:
1. Increase model complexity:
o Use a more complex algorithm (e.g., polynomial regression instead of linear
regression).
o Add more features.
2. Reduce regularization (if used).
3. Train for more epochs (for neural networks).
Overfitting Solutions:
1. Simplify the model:
o Reduce the number of features or use feature selection.
o Reduce model complexity (e.g., limit tree depth in decision trees).
2. Regularization:
o Apply L1 or L2 regularization.
o Add dropout for neural networks.
3. Increase data:
o Use more training data to help the model generalize.
4. Early Stopping:
o Stop training when validation performance stops improving.
11. What is Regularization in Machine Learning?
Why is Regularization Important?
What is Regularization in Machine Learning?
Regularization is a technique used in machine learning to prevent models from overfitting the
training data. It works by adding a penalty (or constraint) to the model's loss function, discouraging
overly complex models that may not generalize well to unseen data.
Why is Regularization Important?
1. Prevents Overfitting:
Regularization helps the model focus on the underlying patterns in the data instead of
memorizing noise or irrelevant details in the training dataset.
2. Improves Generalization:
By reducing model complexity, regularization ensures better performance on new, unseen
data.
3. Handles Multicollinearity:
Regularization can deal with correlated features by shrinking some feature coefficients closer
to zero, especially in linear models.
4. Simplifies Models:
It enforces sparsity, i.e., some weights are driven to zero, effectively selecting only the most
important features.
Types of Regularization
1. L1 Regularization (Lasso):
o Adds a penalty equal to the absolute value of the coefficients (λ∑∣w∣\lambda \sum |
w|λ∑∣w∣).
o Promotes sparsity by shrinking some coefficients to zero, effectively performing
feature selection.
2. L2 Regularization (Ridge):
o Adds a penalty equal to the square of the coefficients (λ∑w2\lambda \sum
w^2λ∑w2).
o Encourages smaller coefficients but does not necessarily make them zero. Suitable
for multicollinearity.
3. Elastic Net:
o Combines L1 and L2 regularization, balancing sparsity and coefficient shrinkage.
Which Models Need Regularization?
Regularization is especially useful for models that are prone to overfitting or where multicollinearity
is a concern:
1. Linear Models:
o Linear Regression
o Logistic Regression
2. Tree-Based Models:
o Decision Trees (regularized via max depth, minimum samples, etc.)
o Random Forests and Gradient Boosting (regularized via parameters like learning rate,
max depth, and subsampling)
3. Neural Networks:
o Regularized using techniques like L2L_2L2 (weight decay) and dropout to avoid
overfitting in deep networks.
4. Support Vector Machines (SVM):
o The CCC-parameter controls regularization strength.
When is Regularization Not Necessary?
If the dataset is very large, and the model complexity is inherently low, the chances of
overfitting are reduced.
When the model already performs well on both training and test datasets, regularization
might not be required.
In summary, regularization is a crucial component in ML for improving generalization, reducing
overfitting, and creating more interpretable models. Its implementation depends on the type of
model and the problem at hand.
12. Weights and biases in liner regression
13. P-Values and Confidence Intervals in
Regression
Both p-values and confidence intervals are statistical tools used to assess the significance and
reliability of coefficients in a regression model.
2. Confidence Intervals
A confidence interval (CI) provides a range of values within which the true coefficient (β\betaβ) is
likely to fall, with a certain level of confidence (usually 95%).
Key Points:
A 95% CI means that if we repeated the experiment 100 times, 95 of the intervals would
contain the true value of the coefficient.
If the CI does not include zero, the coefficient is statistically significant at the corresponding
level (e.g., 5%).
Example:
Continuing the above example, the 95% confidence intervals might look like this:
Interpreting Together
1. P-Values and CIs Align:
o If p<0.05p < 0.05p<0.05, the 95% CI will not include 0.
o If p≥0.05p \geq 0.05p≥0.05, the 95% CI will include 0.
2. Use Cases:
o P-Values: Binary decision-making (e.g., significant vs. not significant).
o CIs: Provide a range and magnitude of effect, offering more context than p-values
alone.
Practical Example in Python (Using statsmodels):
import [Link] as sm
import pandas as pd
# Example dataset
data = [Link]({
'sales': [2, 4, 5, 7],
'advertising': [1, 2, 3, 4],
'price': [10, 12, 14, 16]
})
# Fit regression model
X = data[['advertising', 'price']]
X = sm.add_constant(X) # Add intercept
y = data['sales']
model = [Link](y, X).fit()
# Summary of results
print([Link]())
Output:
Coefficient Estimate P-Value 95% CI
Intercept 1.5 0.01 [1.0, 2.0]
Advertising 1.3 0.001 [0.9, 1.7]
Price -0.5 0.6 [-1.5, 0.5]
Key Takeaways:
Use p-values to test for statistical significance.
Use confidence intervals to understand the precision and range of the coefficient estimates.
14. Regularization
15. Loss function and metric table
In the context of loss functions, lower values are generally better. This is
because loss functions measure the discrepancy(variation) between the
model's predictions and the actual target values. A lower loss value
indicates that the model's predictions are closer to the true values, which
is desirable.
For example:
In regression tasks, where the goal is to predict continuous values
(e.g., house prices), lower values of loss functions such as Mean
Squared Error (MSE) or Mean Absolute Error (MAE) indicate better
performance. This means that the model's predictions are closer to
the true values of the target variable.
In binary classification tasks, where the goal is to classify instances
into one of two classes (e.g., spam or not spam), lower values of
loss functions like Binary Cross-Entropy Loss (Log Loss) are
preferred. Lower values mean that the model's predicted
probabilities for the correct class are closer to 1 (for correct
predictions) or 0 (for incorrect predictions).
In multi-class classification tasks, lower values of loss functions such
as Categorical Cross-Entropy Loss indicate better performance. This
means that the model's predicted probabilities for the correct class
are higher and closer to 1.
In summary, when evaluating machine learning models based on loss
functions, the aim is to minimize the loss value, as this indicates better
alignment between the model's predictions and the true values of the
target variable.
Metric Parameter Value
Accuracy - 0.85
Precision Positive class 0.78
Negative class 0.89
Recall (Sensitivity) Positive class 0.82
Negative class 0.88
F1 Score - 0.80
ROC AUC - 0.91
Confusion Matrix True Positive (TP) 450
True Negative (TN) 560
False Positive (FP) 90
False Negative (FN) 60
In this table:
Metric: Indicates the evaluation metric used to assess the model's
performance.
Parameter: Provides additional information about the metric, such
as the class for which precision and recall are calculated in binary
classification.
Value: Represents the value obtained for the corresponding metric
after evaluating the model on a dataset.
This table offers a snapshot of various evaluation metrics and their
respective values, providing insights into the model's performance across
different aspects such as accuracy, precision, recall, F1 score, ROC AUC,
and confusion matrix.
Mean Absolute Error (MAE) measures the average absolute
difference between predicted and actual values. Lower MAE values
indicate that, on average, the model's predictions are closer to the
true values.
Mean Squared Error (MSE) measures the average of the squared
differences between predicted and actual values. Since squaring
amplifies larger errors, MSE penalizes larger errors more than MAE.
Lower MSE values indicate that the model's predictions are not only
closer to the true values on average but also exhibit less variability.
----------------------------------------------
choice of loss function depends on the type of problem being solved, such
as classification, regression, or ranking, as well as the specific
requirements of the application. Here are some common types of loss
functions:
1. Regression Loss Functions:
Mean Squared Error (MSE): Calculates the average squared
difference between predicted and actual values.
Mean Absolute Error (MAE): Calculates the average absolute
difference between predicted and actual values.
Huber Loss: A combination of MSE and MAE, providing a
compromise between robustness to outliers and sensitivity to
small errors.
2. Classification Loss Functions:
Binary Cross-Entropy Loss (Log Loss): Used for binary
classification problems, measuring the difference between
predicted probabilities and actual binary outcomes.
Categorical Cross-Entropy Loss: Used for multi-class
classification problems, measuring the difference between
predicted class probabilities and one-hot encoded target
labels.
Hinge Loss: Used in support vector machines (SVMs) for binary
classification, encouraging correct classification with a margin.
3. Ranking Loss Functions:
Pairwise Ranking Loss: Used in ranking problems, penalizing
models for misranking pairs of items.
Listwise Ranking Loss: Considers entire ranked lists of items,
optimizing the order of items in the list directly.
4. Custom Loss Functions:
In some cases, custom loss functions may be defined to
address specific requirements of the problem or to incorporate
domain knowledge.
During the training process, the parameters of the model are adjusted
iteratively to minimize the loss function using optimization algorithms
such as gradient descent or its variants. Evaluating and minimizing the
loss function is fundamental to training accurate and effective machine
learning models.
16. ***machine learning algorithms along with a
brief explanation and an example of each:
Linear Regression: A regression algorithm for predicting a continuous
output based on input features.
Example: Predicting house prices based on features like square footage
and number of bedrooms.
Logistic Regression: A classification algorithm that predicts the probability
of a binary outcome.
Example: Predicting whether an email is spam or not based on its content.
Decision Trees: A tree-like model for making decisions by splitting data
into branches based on feature values.
Example: Predicting whether a passenger survived on the Titanic
based on features like age and class.
Random Forest: An ensemble algorithm that combines multiple decision
trees to improve accuracy and robustness.
Example: Identifying handwritten digits in an image.
Support Vector Machines (SVM): A classification algorithm that finds a
hyperplane to best separate different classes.
Example: Classifying whether an email is about technology or sports
based on keywords.
K-Means Clustering: An unsupervised algorithm that groups similar data
points into clusters.
Example: Grouping customers based on their purchase behavior.
Naive Bayes: A probabilistic classification algorithm based on Bayes'
theorem. Example: Categorizing news articles into topics like politics, sports,
or entertainment.
Neural Networks: A complex model inspired by the human brain, capable of
handling intricate patterns.
Example: Image recognition, like identifying objects or animals in images.
Principal Component Analysis (PCA): A dimensionality reduction
technique to transform data into a lower-dimensional space. Example:
Reducing the dimensions of high-dimensional data like images.
Reinforcement Learning: A learning paradigm where an agent learns by
interacting with an environment and receiving rewards.
Example: Training a robot to navigate a maze to reach a goal.
Gradient Boosting: An ensemble technique that builds multiple
models in a sequence, each correcting the errors of the previous one.
Example: Predicting customer churn based on historical data.
XGBoost: An optimized gradient boosting algorithm known for its
performance and flexibility.
Example: Detecting fraudulent transactions in financial data.
XGBoost is a smart algorithm that builds a series of decision trees, each
one correcting the mistakes of the previous trees. It's good at handling
complex relationships in data, and its flexibility and regularization help
prevent overfitting
Long Short-Term Memory (LSTM): A type of recurrent neural network
designed to model sequences and time series data.
Example: Predicting stock prices based on historical price trends.
Gaussian Mixture Models (GMM): A probabilistic model used for
clustering and density estimation.
Example: Segmenting an image into regions with similar colors.
Recommender Systems: Algorithms that provide personalized
recommendations based on user preferences.
Example: Suggesting movies or products based on a user's past choices.
17. Bayes theorem
Result
The probability that the patient actually has the disease, given they tested positive, is 50%. Despite
the positive test, the rarity of the disease significantly affects the result. This demonstrates how
Bayes' Theorem combines prior knowledge (disease rarity) with new evidence (test result).
18. Basics related to Machine learning model
Statistical Inference
Why: Understanding statistical inference is foundational for analyzing data and deriving
insights.
Topics: Probability, Hypothesis Testing, Descriptive Statistics.
2. Supervised Learning (Classification and Regression)
Why: These are fundamental machine learning tasks. Classification and regression are widely
applied and relatively easier to understand.
Topics:
Classification: kNN, Logistic Regression, Naive Bayes, SVM, Decision Trees.
Regression: Linear Regression, Polynomial Regression, Ridge, Lasso.
3. Unsupervised Learning
Why: After supervised learning, understanding unsupervised methods is essential for
exploring and analyzing data without labeled outcomes.
Topics:
Clustering: k-Means, DBSCAN, Fuzzy C-Means, Mean-Shift.
Pattern Search: Apriori, ECLAT, FP-Growth.
4. Dimensionality Reduction
Why: To deal with high-dimensional datasets, understanding dimensionality reduction
techniques is important for visualization and preprocessing.
Topics: PCA, t-SNE, LDA, SVD, ODA, LLE.
5. Ensemble Models
Why: These models build on classification and regression knowledge and improve predictive
performance.
Topics:
Boosting: AdaBoost, CatBoost, XGBoost, Gradient Boost.
Bagging: Random Forest.
Stacking.
6. Reinforcement Learning
Why: It is more advanced and focuses on decision-making and control, often building upon
the foundations of supervised and unsupervised learning.
Topics:
Algorithms: Q-Learning, SARSA, A3C, Genetic Algorithms, Deep Q-Network (DQN).
19. what is ANOVA , explain with simple example
20. hypothesis testing, regression analysis,
ANOVA, experimental design, and probability
theory
Q5: What is multicollinearity, and how can it affect a regression model?
A: Multicollinearity occurs when two or more independent variables are highly correlated,
leading to difficulties in estimating the coefficients accurately. It can result in:
Unstable coefficient estimates
Increased standard errors, reducing the statistical significance of predictors To address
multicollinearity, you can:
Remove highly correlated variables.
Use techniques like Principal Component Analysis (PCA).
Apply regularization methods like Ridge or Lasso regression.
Q6: Explain the difference between R-squared and Adjusted R-squared.
A:
R-squared: Represents the proportion of variance in the dependent variable explained by
the independent variables. It ranges from 0 to 1.
Adjusted R-squared: Adjusts R-squared for the number of predictors in the model. It
accounts for the possibility of overfitting when more predictors are added.
ANOVA (Analysis of Variance)
Q7: When would you use ANOVA instead of a t-test?
A:
A t-test is used to compare the means of two groups.
ANOVA is used to compare the means of three or more groups. It determines whether there
are statistically significant differences among the group means.
Q8: What are the assumptions of ANOVA?
A:
1. Independence: Samples are independent of each other.
2. Normality: The data in each group should be approximately normally distributed.
3. Homoscedasticity: Variance among the groups should be equal.
Experimental Design
Q10: What is randomization in experimental design, and why is it important?
A: Randomization is the process of assigning subjects to treatment groups randomly. It
reduces bias and ensures that the treatment groups are comparable, making it easier to
attribute differences in outcomes to the treatment effect rather than confounding variables.
Q11: Explain the concept of blocking in experimental design.
A: Blocking is a technique used to reduce variability by grouping experimental units with
similar characteristics into blocks. Within each block, treatments are randomly assigned. This
controls for the effect of the blocking variable.
Q14: What is the difference between discrete and continuous probability distributions?
A:
Discrete distributions: Represent probabilities of discrete outcomes (e.g., Binomial, Poisson).
Continuous distributions: Represent probabilities of continuous outcomes (e.g., Normal,
Exponential).
21. Ensemble technique
"Ensemble technique" in the context of machine learning refers to a methodology where
multiple models are combined to improve the overall predictive power and performance.
Ensemble techniques are commonly used to enhance the accuracy, stability, and
robustness of machine learning models. Instead of relying on a single model's predictions,
ensemble methods leverage the diversity of multiple models to make more accurate
predictions.
Some popular ensemble techniques include:
1. Bagging (Bootstrap Aggregating): Bagging involves training multiple instances
of the same model on different subsets of the training data, created by sampling
with replacement. The final prediction is often an average or majority vote of the
predictions from these individual models.
Random Forest is a well-known algorithm that utilizes bagging.
2. Boosting: Boosting is an iterative technique where multiple weak learners
(models that perform slightly better than random chance) are trained sequentially.
Each subsequent model focuses on correcting the mistakes made by the previous
ones, effectively "boosting" the overall performance. Algorithms like AdaBoost
and Gradient Boosting Machines (GBM) fall under this category.
Example:
In boosting, you train the first model (friend) and see where it makes mistakes.
Now, you bring in a second friend who's good at correcting those specific
mistakes. This second friend focuses on the examples the first friend got wrong
and helps you improve the accuracy.
Then, you bring in a third friend who's even better at correcting the remaining
mistakes. Each new friend (model) you add specializes in fixing the errors made
by the previous friends. This process continues, and with each round, the team of
friends collectively becomes really good at distinguishing between apples and
oranges.
In the end, by combining the opinions of all your friends (models), you're able to
make very accurate predictions about whether a fruit is an apple or an orange,
even though each friend on their own might not have been very reliable. This is
how boosting works – building a strong model by learning from the weaknesses of
multiple weaker models.
3. Stacking: Stacking involves training multiple diverse models and then using a
meta-model to combine their predictions. The predictions of the base models serve
as features for the meta-model. Stacking aims to capture different aspects of the
data through different models, leading to improved overall performance.
4. Voting: Voting methods combine the predictions of multiple models by taking a
majority vote (for classification) or an average (for regression). It works well
when individual models have varying strengths and weaknesses.
5. Ensemble of Experts: In this approach, different models are trained to specialize
in different parts of the data. Each model's prediction is weighted based on its
expertise in a specific region of the feature space.
6. Random Subspace Method: This technique involves training each base model
on a random subset of the input features. It's particularly useful when there are
many features and some of them might be noisy or irrelevant.
Ensemble techniques are generally effective because they mitigate overfitting (when a
model learns the training data too well and performs poorly on new data) and improve
generalization (performing well on new, unseen data). They also provide robustness
against individual model errors or noise. However, ensemble methods can be
computationally more expensive due to training multiple models and combining their
results.
The choice of ensemble technique depends on the problem at hand, the type of data, and
the characteristics of the base models. Each ensemble method has its strengths and
weaknesses, and their performance can vary based on the specific application.
Sure! Imagine you have a bunch of friends who are really good at different things. Some
are good at math, some at art, and others at sports. Now, if you need to solve a problem or
make a prediction, instead of just asking one friend, you ask all your friends and then
decide based on what most of them think.
Ensemble techniques in machine learning are like this. Instead of relying on just one
"friend" (model), you use a group of them to work together and make a better decision.
This way, you get a more accurate answer because each friend brings their own special
knowledge to the table. It's like combining different superpowers to become a superhero
team that's really good at solving problems!
22. What is AI
AI, or Artificial Intelligence, refers to the simulation of human
intelligence in machines, enabling them to perform tasks that
typically require human cognitive functions such as learning,
reasoning, problem-solving, perception, and language
understanding. AI systems aim to replicate or mimic human-
like intelligence to varying degrees, depending on the specific
application and level of advancement.
There are several approaches to AI, including:
1. Narrow or Weak AI: This type of AI is designed for a
specific task and can perform that task very well, often
outperforming humans in that specific domain. Examples
include virtual personal assistants (like Siri or Alexa) and
recommendation systems (like those used by streaming
platforms).
2. General or Strong AI: This refers to AI systems that
possess human-like cognitive abilities and can
understand, learn, and apply knowledge across a wide
range of tasks, similar to human intelligence. General AI is
still largely theoretical and has not been achieved.
3. Machine Learning: A subset of AI, machine learning
involves training algorithms to improve their performance
on a specific task by learning from data. It includes
techniques like neural networks, decision trees, and
clustering algorithms.
4. Deep Learning: This is a subset of machine learning that
focuses on artificial neural networks and their ability to
learn and make decisions. Deep learning has shown
remarkable results in tasks such as image and speech
recognition.
5. Natural Language Processing (NLP): This branch of AI
deals with enabling computers to understand, interpret,
and generate human language. Chatbots, language
translation, and sentiment analysis are examples of NLP
applications.
6. Computer Vision: This field focuses on enabling
machines to interpret and understand visual information
from the world,
similar to how humans perceive and understand images
and videos.
7. Reinforcement Learning: A machine learning paradigm
in which an agent learns to make decisions by interacting
with an environment. It receives feedback in the form of
rewards or penalties, allowing it to learn optimal
strategies.
8. Cognitive Computing: An interdisciplinary field that
combines AI, psychology, neuroscience, and linguistics to
create systems that can simulate human thought
processes.
AI has a wide range of applications, from self-driving cars and
medical diagnostics to financial analysis and entertainment. It
has the potential to transform industries and improve various
aspects of our lives, but it also raises ethical and societal
questions about job displacement, privacy, bias, and more.
23. Confusion Matrix
The confusion matrix is a table that helps evaluate a model's performance in classification. It has 4
terms:
Predicted Positive Predicted Negative
Actual Positive True Positive (TP) False Negative (FN)
Actual Negative False Positive (FP) True Negative (TN)
Memory Trick: Think of Positive and Negative
True Positive (TP): Predicted Positive, and it’s True.
False Positive (FP): Predicted Positive, but it’s False.
True Negative (TN): Predicted Negative, and it’s True.
False Negative (FN): Predicted Negative, but it’s False.
Example for Intuition:
Scenario: Test for a disease.
o TP: People correctly identified as sick.
o FP: Healthy people falsely labeled as sick.
o FN: Sick people missed.
o TN: Healthy people correctly identified as healthy.
Use these questions to remember:
o Precision: "When the test says sick, how often is it right?"
o Recall: "Of all the sick people, how many did the test catch?"
o Specificity: "Of all healthy people, how many were correctly identified?"
24. Evaluating machine learning models
Evaluating machine learning models involves assessing their
performance and generalization ability on unseen data.
Various metrics are used to measure different aspects of a model's
performance. Here are some common types of metrics used for
evaluating machine learning models:
1. Classification Metrics:
Accuracy: The proportion of correctly classified instances
out of the total instances. It is a common metric for
balanced datasets.
Precision: The ratio of true positive predictions to the total
positive predictions. It measures the accuracy of the
positive predictions.
Recall (Sensitivity or True Positive Rate): The ratio of
true positive predictions to the total actual positive
instances. It measures the ability of the model to capture
all positive instances.
F1 Score: The harmonic mean of precision and recall,
providing a balance between the two metrics.
2. Confusion Matrix Metrics:
True Positive (TP): Instances correctly predicted as
positive.
True Negative (TN): Instances correctly predicted as
negative.
False Positive (FP): Instances incorrectly predicted as
positive.
False Negative (FN): Instances incorrectly predicted as
negative.
3. Receiver Operating Characteristic (ROC) Curve and Area
Under the Curve (AUC):
ROC curves plot the true positive rate against the false
positive rate at various threshold settings. AUC measures
the area under the ROC curve, providing a single value to
summarize the model's performance across different
thresholds.
4. Regression Metrics:
Mean Squared Error (MSE): The average of the squared
differences between predicted and actual values. It gives
higher weight to larger errors.
Mean Absolute Error (MAE): The average of the
absolute differences between predicted and actual values.
It provides a more interpretable metric compared to MSE.
R-squared (R2): Measures the proportion of the variance
in the dependent variable that is predictable from the
independent variables. R2 values range from 0 to 1, with
higher values indicating better performance.
5. Clustering Metrics:
Silhouette Score: Measures how similar an object is to its
own cluster compared to other clusters. Values range from
-1 to 1, with higher values indicating better-defined
clusters.
Davies-Bouldin Index: Measures the average "similarity"
between each cluster and its most similar cluster. Lower
values indicate better clustering.
6. Ranking Metrics (for Recommender Systems):
Precision at K: Measures the proportion of recommended
items in the top K that are relevant.
Recall at K: Measures the proportion of relevant items
that are recommended in the top K.
Mean Average Precision (MAP): Averages precision
across different levels of recall.
7. Anomaly Detection Metrics:
Area under the Precision-Recall curve (AUC-PR):
Similar to AUC-ROC but used for imbalanced datasets,
where the positive class is rare.
F1 Score, Precision, Recall: Can be used to evaluate
anomaly detection performance.
8. Fairness and Bias Metrics:
Disparate Impact: Measures the ratio of favorable
outcomes for the protected group compared to the majority
group.
Equalized Odds: Assesses whether the model's
predictions are independent of the protected group.
It's important to choose metrics based on the specific problem, the
nature of the data, and the goals of the machine learning model.
Additionally, cross-validation and hyperparameter tuning are common
practices to ensure robust model evaluation.
25. Question-1-> What is Supervised Learning and
What is Unsupervised Learning?
Ans-1 Let’s say hypothetically I am playing with a small kid
of age 3. I took a tray and put three fruits in-front of him,
First fruit was red apples, Second fruit was pink cherries,
Third fruit was Banana’s.
I am assuming that the kid never saw these fruits, never
ate them, he did not know nothing about these fruits.
Then i told him 100 times, that this red fruits called
“Apples”, this pink fruit is called “Cherry”, this yellow fruit
is called “Banana’s”. I repeat and repeat.
Here I provide my kid the input which is a fruit and the
output which is called “Apple”. The kid is able to learn that
this red looking fruit which is round in shape is called
“Apple”.
That’s what we called Supervised Learning, where you are
giving input and the output both and the machine needs to
learn that how a particular input is mapped to a particular
output.
For example- Gmail Spam/NotSpam Classifier.
Then what is Unsupervised Learning?
Let’s say i got another kid, i just showed him 3 fruits but
did not tell what fruit is called what. I am sure after
looking the tray multiple times, the kid will be able to
differentiate that all the red fruits looks same, all the
yellow fruits looks same and so on. He won’t be able to tell
the name but he will be able to put them in different
categories based on the colours.
That is unsupervised learning where we give machine only
the input, from the input it needs to understand that the
particular input belongs to which group or cluster.
For example- Credit Card Fraud Detection
1. Question-2-> Can you tell me the difference
between A Data Analyst and A Data Scientist?
Ans2- Let’s say i am an employee in Swiggy and my
manager asked me to
“Give the top 5 cities from where we are getting the least
order in last 6 months”. Now let’s say swiggy got 10 crore
orders in last 6 months, to get the least order city
information manually it will take months. The answer to
this question is already there in my data but manually it’s
very difficult to find it so here comes the role of Data
Analyst who uses libraries like Numpy, Pandas, Matplotlib,
even tools like power Bi to find the answer.
But let’s say my manager said to me that, “Give the
estimated sale value of our company by the end of this
year”, now here my job is to predict the sale value, This is
Data Science or Machine Learning.
You go to the past in the data, its data analysis.
You go into the future, its data science.
2. Question-4-> Define p-value and tell me why it is
important?
Ans-4 -When scientists do experiments, they want to know
if their results are actually meaningful or just a
coincidence. The p-value helps them make that
determination.
A p-value is a number that tells us the likelihood that the
results we see in an experiment are due to chance. The
lower the p-value, the less likely it is that the results are
just a coincidence.
For example, let’s say a group of scientists is testing a new
medicine. They give the medicine to one group of people
and a placebo (fake medicine) to another group. Then they
measure how many people in each group get better. If the
group that got the medicine has a much higher percentage
of people who get better, the scientists can calculate the p-
value to see if this difference is likely due to chance or if it
is statistically significant.
In general, scientists consider a p-value of 0.05 or lower to
be statistically significant. This means that there is less
than a 5% chance that the results are due to chance, and
that they are probably meaningful.
3. Question-5-> What is PDF and CDF and why do you
think it’s important in Machine Learning?
Ans-5 PDF stands for Probability Density Function. In
probability theory, a PDF is a function that describes the
relative likelihood of a random variable taking on a
particular value or set of values.
CDF stands for Cumulative Distribution Function. It’s a
function that gives the probability that a random variable
is less than or equal to a certain value.
PDFs and CDFs are important because they allow us to
model the distribution of data. By understanding the
distribution of data, we can make more informed decisions
about how to process and analyze it.
For example, if we have a dataset of images, we might
want to know what the distribution of pixel values looks
like. This information can help us choose appropriate
preprocessing techniques or models that take into account
the structure of the data.
4. Question-6-> Can you explain the Central Limit
theorem?
The Central Limit Theorem is a statistical concept that
helps us understand how the means of random variables
are distributed. It states that if we take a large number of
random samples from a population and calculate the mean
of each sample, the distribution of those means will be
approximately normal, regardless of the shape of the
original population.
For example, imagine we wanted to know the average
height of all students in a school. We could measure the
height of every student, but that would be very time-
consuming. Instead, we could take a random sample of
students and measure their heights. We could then take
the mean height of that sample and repeat the process
many times, each time taking a different random sample of
students. According to the Central Limit Theorem, the
distribution of those means would be approximately
normal, even if the distribution of heights in the original
population was not normal.
5. Question-7-> What are data structures and what
data structures we have in python?
Ans-7 Data Structures means the ways i can store my data,
but the question is why do i need different ways to store
data.
Is storing data in a single structure not enough?
So yes, we can not store data into only one data structure
because every data structures has its own advantages and
has its own disadvantages.
For example- Python has 4 data structures
1. List- Now list is a mutable data structure but it has
a o(n) time complexity which is not good but has
o(1) space time complexity.
2. Tuple- Tuple is immutable and has the same
search time complexity as list.
3. Dictionary-Dictionary is mutable and also have
o(1) search time complexity but o(n) space time
complexity.
4. Set- Set is mutable because we can insert new
elements into list, set does not allow duplicates
also.
6. Question-8-> What are CNNs, Can you explain any
famous CNN architectures?
Ans-8 CNN stands for Convolution Neural Network, where
the convolution means element wise multiplication and
addition. CNNs are generally used in Computer Vision
Problems where the basic problem is to generate vectors
out of images.
How would we gather information from a image to a
mathematical vector so we can use that information for
problems like object detection, face recognition etc.
These are some famous CNN architectures-:
1. LeNET-LeNet (short for LeNet-5) is a convolutional
neural network architecture that was developed by
Yann LeCun, Leon Bottou, Yoshua Bengio, and
Patrick Haffner in 1998. It was one of the first
successful convolutional neural networks for image
recognition and classification tasks.
2. AlexNet- AlexNet is a convolutional neural
network architecture designed by Alex Krizhevsky,
Ilya Sutskever, and Geoffrey Hinton. It won the
ImageNet Large Scale Visual Recognition
Challenge in 2012 and is widely considered as one
of the breakthroughs that sparked the deep
learning revolution.
3. VGG16- VGG16 is a convolutional neural network
architecture that consists of 16 layers, including 13
convolutional layers and 3 fully connected layers. It
was developed by the Visual Geometry Group at
the University of Oxford and achieved state-of-the-
art performance on the ImageNet dataset in 2014.
7. Question-9-> What is Max Pooling and why do we
need Max Pooling?
Ans-9 Max Pooling involves dividing an input image into a
set of non-overlapping rectangular regions and then taking
the maximum value of each region. The result is a down-
sampled version of the input image with reduced spatial
dimensions.
The primary reason for using Max Pooling is to reduce the
size of the feature maps produced by the convolutional
layers in a CNN, while retaining the most important
features. This helps to reduce the computational
complexity of the network, making it easier to train and
faster to process new input images.
8. Question-10-> What is NLP?
NLP stands for Natural Language Processing,it is a branch
of artificial intelligence that focuses on teaching machines
how to understand, interpret, and generate human
language. It involves developing algorithms and models
that can analyze and manipulate text, speech, and other
forms of natural language data.
It enables machines to perform tasks such as sentiment
analysis, language translation, text summarization, chatbot
conversations, and more. NLP is used in a wide range of
applications, from virtual assistants like Siri and Alexa to
spam filters in email, and even in healthcare for medical
diagnosis and treatment planning.
9. Question-11-> What are the algorithms to embed a
sentence into a vector?
Ans-11 Some of the algorithms which are used to embed a
sentence into a vector is:-
1. Bag Of Words
2. TF-IDF
3. Word2Vector
4. BERT(Transfer)
5. GPT Models
10. Question-12-> What’s your favourite
algorithm?
Ans-12 My favourite algorithm is Random Forest.
Random Forest is a popular machine learning algorithm
that uses the concept of ensemble learning to build a
robust and accurate model.
Ensemble learning involves combining multiple models to
make a prediction. In the case of Random Forest, multiple
decision trees are built using randomly selected subsets of
the training data and features. Each decision tree is
trained independently, and the final prediction is made by
taking the average or majority vote of the predictions from
all the decision trees.
The random selection of subsets of data and features helps
to reduce the impact of individual noisy data points or
features on the model’s predictions, and the combination
of multiple decision trees helps to improve the model’s
accuracy and reduce overfitting.
11. Question-13-> What is a hyper parameter?
In machine learning, a hyperparameter is a parameter that
is set before the training process begins and
determines the overall behavior and performance of the
machine learning algorithm.
Hyper parameters are used to control the learning
process and affect how the model is trained, such as the
learning rate, regularization parameter, number of hidden
layers, number of neurons per layer, etc.
The process of choosing the best hyperparameters for a
given problem is known as hyperparameter tuning
12. Question-14-> What is Under-fitting
Ans-14 Underfitting is a common problem in machine
learning where a model is unable to capture the underlying
patterns and relationships in the training data, resulting in
poor performance on both the training and test data.
In simple terms, underfitting occurs when the model is too
simple or not complex enough to represent the data. This
means that the model cannot learn the relevant features
and relationships between the input variables and the
target variable. As a result, the model produces high bias
and low variance, and it performs poorly on both the
training data and new unseen data.
13. Question-15->What is Over-Fitting ?
Ans-15 Overfitting is a common problem in machine
learning where a model is too complex and fits the training
data too closely, resulting in poor generalization
performance on new unseen data.
In simple terms, overfitting occurs when the model is too
complex and learns the noise in the training data along
with the underlying patterns and relationships between the
input variables and the target variable. This means that
the model is not able to generalize well to new data and
produces high variance and low bias, which leads to poor
performance on new data.
14. Question-16-> What is model error?
Ans-16 The model error is defined as-:
Model error= (Bias)²+ Variance
15. Question-17-> What is One-Hot-Encoding?
Ans-17 One hot encoding is a technique used in machine
learning to represent categorical data as numerical data. It
involves creating a binary vector for each category in a
categorical variable, where each vector has a length equal
to the number of unique categories in the variable.
In simple terms, one hot encoding replaces each category
in a categorical variable with a binary vector that has a
value of 1 in the corresponding index and a value of 0 in all
other indices.
For example, suppose we have a categorical variable
“Color” with three unique categories: Red, Green, and
Blue. One hot encoding would create a binary vector of
length three for each color: [1, 0, 0] for Red, [0, 1, 0] for
Green, and [0, 0, 1] for Blue.
16. Question-18-> What is Multi-class
classification?
Ans-18 Multi-class classification is the task of predicting a
target variable that has more than two possible outcomes
or classes. For example, predicting the type of fruit in an
image as either an apple, banana, or orange would be a
multi-class classification problem.
The main difference between binary classification and
multi-class classification is the number of classes or
categories in the target variable. In binary classification,
the target variable has only two classes, while in multi-
class classification, the target variable has more than two
classes.
17. Question-19->Can we use all the algorithm if
we have Multi-class classification problem?
Ans-19 No we can not use all algorithm directly to use
Multi-Class Classification, We have use techniques like
One Vs ALL to make it work.
For example, Logistic Regression is defined for binary
class but we can use One Vs All to make it use for Multi
Class Classification.
18. Question-20-> What is the performance
metric you will use if you have a medical related
problem like “Cancer Prediction”?
Ans-20 One commonly used performance metric in
medical applications is sensitivity, also known as the
true positive rate.
Sensitivity measures the proportion of actual
positive cases that are correctly identified as positive
by the model. In cancer prediction, sensitivity is
important because it reflects the ability of the model to
correctly identify individuals who have cancer and who
may require further medical evaluation.
Another important performance metric for cancer
prediction is specificity, also known as the true negative
rate.
Specificity measures the proportion of actual
negative cases that are correctly identified as
negative by the model. In cancer prediction, specificity is
important because it reflects the ability of the model to
correctly identify individuals who do not have cancer and
who may not require further medical evaluation.
19. Question-21-> What is F1 Score?
Ans-21 F1 score is the harmonic mean of Precision and
Recall.
F1=2PR/P+R
20. Question-23-> What is NLTK?
Ans 23-NLTK is an open-source software library written in
Python that provides tools and resources for natural
language processing (NLP) tasks, such as tokenization,
stemming, tagging, parsing, and semantic reasoning. It
also offers access to a vast collection of language corpora
and lexical resources. NLTK is widely used in academia
and industry for NLP research, development, and
education.
21. Question-24-> What is Dropouts?
Ans 24-Dropout is a regularization technique used in deep
learning neural networks to prevent overfitting. Overfitting
occurs when a model becomes too complex and is trained
to fit the training data too closely, which can cause it to
perform poorly on new, unseen data.
Dropout works by randomly dropping out (i.e.,
setting to zero) some of the neurons in a neural
network during training. This forces the network to
learn more robust features and prevents any single neuron
from becoming too important in making predictions.
During the training process, each neuron in the network is
either dropped out with a specified probability or retained
with probability (1 — dropout rate). The dropout rate is
typically set between 0.1 and 0.5, and the optimal value
can depend on the specific problem and architecture of the
neural network.
22. Question-25-> What is Batch Normalisation?
Ans 25- Batch normalization works by normalizing the
output of each layer before applying the activation
function. During training, the mean and standard deviation
of the output of each layer are computed for each mini-
batch, and these statistics are used to normalize the
output. Specifically, the output is normalized by
subtracting the mean and dividing by the standard
deviation, which is then scaled and shifted by learnable
parameters called gamma and beta.
Batch normalization has several benefits, including:
It reduces the dependence of the model on the
initialization of the weights, making it easier to
train deep networks.
It can improve the generalization performance of
the model by reducing overfitting.
It makes the optimization process more stable,
allowing for the use of higher learning rates.
23. Question-26-> What is a Perceptron?
Ans 26- A perceptron is a type of artificial neural network
that is commonly used in machine learning for binary
classification problems. It is a simple model that takes a
vector of inputs, applies a linear function to the inputs, and
produces a binary output based on a threshold.
The most commonly used activation function in
perceptrons is the step function, which produces an output
of 1 if the weighted sum of the inputs is greater than or
equal to a threshold, and 0 otherwise. Other activation
functions, such as the sigmoid or ReLU functions, can also
be used in more complex models.
24. Question-27-> How logistic Regression is a
single neuron model?
Ans-27 It is a single neuron model in the sense that it is
based on a single output neuron that applies a logistic
function (also known as a sigmoid function) to a linear
combination of the inputs.
In logistic regression, the input features x1, x2, …, xn are
combined linearly using weights w1, w2, …, wn, and an
intercept b to produce a single output z = w1x1 + w2x2 +
… + wn*xn + b. This output is then passed through a
logistic function to produce the predicted probability y_hat
of the positive class:
y_hat = 1 / (1 + exp(-z))
Only one neuron is giving the output, thats we called it
Single Neuron Model.
25. Question-28-> Why the name of logistic
regression has regression when it’s a classification
technique.
Ans 28- The output of the logistic regression is:-
y_hat = 1 / (1 + exp(-z))
y_hat will be always between 0 to 1, so it will any any value
between 0 to 1 which resembles regression. Classification
means 0 or 1, Regression means any value between 0 to 1.
so the nature of the output is regression but then we take a
threshold value to decide whether it belongs to class 0 or
class 1.
For example- If output comes out to be 0.8, and the
threshold is 0.5, then we will say because the value of
output>0.5, it belongs to 0 class.
26. Question-29-> What is a sigmoid function?
Why it is used in logistic regression?
Ans-29 Sigmoid Function is
Sigmoid(x)= 1 / (1 + exp(-x))
Sigmoid function is a differentiable function and has a
probabilistic nature.
It is used in logistic regression because it taper off the
outliers in the data. We use sigmoid function to squash the
outliers from the data.
27. Question-30-> What do you mean by a
hyperplane and how hyperplane is important with
respect to machine learning?
Ans30- Let’s say I have a classification task and i want to
classify two classes, spam and not spam, and I have only 2
features.
Two features means the geometry of the problem lies in
2D, x axis is feature 1 and y axis is feature 2.
Now if you see in the diagram, if i want to classify 2
classes, i would need a line.
If I got three features, to separate these two classes we
need a plane.
If I got more than three features, I would need a
hyperplane.
In most of the machine learning problems, we have more
than 3 features and to solve that i would need a
hyperplane. I would know equation of hyperplane because
hyperplane is actually the decision boundary.
28. Question-31-> What is Backpropogation?
Ans31 Backpropogation is a way where we leverage chain
rule of differentiation and Memoization concept to find the
derivative of the loss function with respect to weights so
that we can update our weights in update equation of
gradient descent.
Backpropogation is the back of deep learning because at
the end we need to find the weights to complete the
training process.
29. Question-32-> What is Gradient Descent?
Ans32- Gradient descent is an optimization algorithm used
to minimize the cost function of a machine learning model.
The cost function is a measure of the difference between
the predicted output of the model and the actual output.
The algorithm works by starting at a random point in the
parameter space of the model and iteratively adjusting the
parameters in the direction of steepest descent of the cost
function, which is determined by the gradient of the cost
function. The gradient is a vector that points in the
direction of the greatest rate of increase of the cost
function, and the algorithm updates the parameters in the
opposite direction of the gradient, in order to move
towards a minimum of the cost function.
30. Question-33->What is the problem with
gradient descent?
Ans33- While gradient descent is a widely used and
effective optimization algorithm for many machine
learning models, there are some potential problems
associated with it:
1. Local minima: Gradient descent can converge to a
local minimum of the cost function, which may not
be the global minimum. This can lead to suboptimal
solutions that do not perform as well as they could.
2. Learning rate: The learning rate determines the
step size taken by the algorithm in the direction of
the gradient. If the learning rate is too small, the
algorithm may converge too slowly, while if the
learning rate is too large, the algorithm may
overshoot the minimum and fail to converge.
31. Question-34-> What is SGD(Stochastic
Gradient Descent)?
Ans34-
1. SGD is a variant of gradient descent used for
training machine learning models.
2. It updates the model parameters after each
individual training example or a small batch of
examples.
3. This makes the algorithm more computationally
efficient and can lead to faster convergence.
4. SGD is commonly used for large datasets and deep
learning models.
32. Question-35-> Explain the problem of
Vanishing Gradient?
Ans35- The problem of vanishing gradients refers to a
situation where the gradients of the cost function with
respect to the model parameters become very small,
particularly in deep neural networks, making it difficult to
train the network effectively.
When the gradients are small, the updates to the model
parameters during training become very small, and the
network may converge to a suboptimal solution or not
converge at all. The problem is particularly acute in deep
neural networks with many layers, where the gradients
can become exponentially small as they propagate through
the layers.
33. Question-36-> What are the different
activation function we used in deep learning?
Ans36-
1. Sigmoid: The sigmoid function maps any input
value to a value between 0 and 1. It is often used in
the output layer of binary classification models.
2. Hyperbolic tangent (tanh): The tanh function
maps any input value to a value between -1 and 1.
It is commonly used as an activation function in
hidden layers.
3. Rectified Linear Unit (ReLU): The ReLU
function outputs the input value if it is positive, and
0 otherwise. It is widely used in deep neural
networks due to its simplicity and effectiveness.
4. Leaky ReLU: The leaky ReLU function is a variant
of the ReLU function that allows for a small
positive gradient when the input is negative. This
can help alleviate the “dying ReLU” problem where
the gradients become zero for negative inputs.
34. Question-37-> How to deal with Outliers in
Machine Learning?
Ans -37- Here are some techniques that can be used to
deal with outliers in machine learning:
1. Detection: Before dealing with outliers, it is
important to first detect them. This can be done
using various statistical techniques such as Z-
score, interquartile range (IQR), and box plots.
2. Removal: One approach to dealing with outliers is
to simply remove them from the dataset. However,
this approach should be used with caution, as
removing too many data points can lead to a loss of
information and biased results.
35. Question-38-> What will you do if you model is
overfit?
Ans 38- Increase the size of the dataset: One way to reduce
overfitting is to increase the size of the training dataset.
This can help the model generalize better to unseen data.
1. Feature selection: Another approach is to
carefully select the most relevant features that are
likely to have a strong impact on the model’s
predictions. This can help reduce noise in the data
and improve the model’s performance.
2. Regularisation: Regularization is a technique that
adds a penalty term to the cost function during
training, which helps to prevent the model from
overfitting the training data. Popular regularisation
techniques include L1 and L2 regularisation,
dropout, and early stopping.
36. Question-39-> What is RFR(Randomisation For
Regularisation)?
Ans39- Randomized Regularization is a technique that
combines the concepts of regularization and randomization
to improve the performance of machine learning models. It
involves adding a random component to the regularization
process, which helps to prevent the model from overfitting
to the training data.
37. Question-40-> Can you explain the difference
between Convex and Non Convex function?
Ans40 -In simple terms, a convex function is a function that
has a bowl-like shape, where any line segment connecting
two points on the function lies above or on the function.
On the other hand, a non-convex function is a function that
has a more complex shape, where there may be multiple
local minimums and maximums.
In optimization problems, convex functions are easier to
optimize because they have a single global minimum,
which can be found efficiently. Non-convex functions, on
the other hand, are more challenging to optimize because
they may have multiple local minimums, and finding the
global minimum requires exploring a larger search space.
CODING ROUND
38. Question-41-> Find the majority element in a
list.
39. Question-42-> Rotate the array by d elements.
40. Question-43->Given a string, Reverse the
order of strings in each word within a sentence
while preserving white space and initial word order.
41. Question-44-> Given 2 strings, you need to
find the common elements from both the strings.
one update in the code, please do the lower case for the
strings and then solve it.
42. Question-45-> Write a code in python to print
the pattern.
43. Question-46-> What is the search time
complexity of a list and a dictionary?
Ans 46- Search Time complexity is O(n) which is not good
when compare to dictionaries which is O(1).
44. Question-47->Write a python program to
define power function from scratch?
45. Question-48-> Write a python function to find
the cubic sum only by recursion?
46. Question-49-> Write a program to find the dot
product in python from scratch?
47. Question-50-> Write a python program which
gives us euclidian distance from scratch?