0% found this document useful (0 votes)
17 views138 pages

FG - Machine Learning Methods Using Python and R-II Edited

This facilitator's guide for the Machine Learning Methods Using Python and R course provides essential materials and tips for effective teaching. It outlines the program's objectives, facilitator roles, and session preparation strategies, including participant engagement and conflict management. The guide also details course content, including topics like time series forecasting, clustering, and neural networks, along with necessary tools and techniques for data analysis.

Uploaded by

pandabiren29
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
17 views138 pages

FG - Machine Learning Methods Using Python and R-II Edited

This facilitator's guide for the Machine Learning Methods Using Python and R course provides essential materials and tips for effective teaching. It outlines the program's objectives, facilitator roles, and session preparation strategies, including participant engagement and conflict management. The guide also details course content, including topics like time series forecasting, clustering, and neural networks, along with necessary tools and techniques for data analysis.

Uploaded by

pandabiren29
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

BSc in Artificial Intelligence and Machine Learning

Undergraduate Diploma - Second Year


SEMESTER-II
Machine Learning Methods Using Python and R - II
FACILITATOR’S GUIDE

Tata Institute of Social Sciences – School of Skill Education


----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Preface and Purpose of the Guide


This guide is designed to help the course facilitator plan and conduct the course.
What will I find in the guide?
This facilitator’s guide is a comprehensive package that contains:
1. Presentation scripts and key points to cover
2. Key points at a glance
3. Facilitation tips
4. Sample of potential questions
5. Checklists of necessary materials and equipment

1. The Session in Perspective

Programme Overview:
This programme is designed to equip participants with critical skills to understand the
importance of programming and get acquainted with various concepts.
Module Learning Goals
Terminal Objectives:
Participants will be able to comprehend the importance of programming and get acquainted
with various concepts.
Enabling objectives:
● Understand and explain the definition of Artificial Intelligence and Machine Learning
● Identify the concept of entities and terms involved in Programming
● Describe the Machine Learning Methods Using R and Python
● Understand and explain the qualities of a High-Level Language
● Familiarize themselves with the relationship between Programming with ML Python and
R.

Programme Preparation
Programme requires a significant amount of preparation. It is crucial that Programme
facilitators familiarize themselves with the material they designed or are expected to deliver
and have adequate time to adapt the content to the specific audience.
The Role of the Facilitator
Who is a facilitator?
A facilitator is someone who is present to assist a group reach its objectives; the group, not the
facilitator, may determine the objectives.

____________________________________________________________________________________
2 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

The facilitator’s role:


When adopting the role of a facilitator, the facilitator needs to:
1. Ensure the more talkative participants do not take over and encourage contributions
particularly from those who may be less confident.
2. Devise non aggressive, friendly ways to deal with difficult participants.
3. Control conflict by stepping in if necessary to help participants learn to deal with conflict
positively.
4. From time to time, get the participants to summarize what has been discussed.
5. Assist ‘weaker’ participants by rephrasing their arguments for them so that these do not get
lost just because they are not forcefully put across.
6. Provide feedback to the group as a whole regarding its performance.
7. Provide the information and resources for the group to function effectively.

____________________________________________________________________________________
3 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Session Preparation

Questions
What
● What is the subject I have been asked to present on/lead/arrange?
Why
● Why I have been asked to do it?
● What is the purpose of the session or the training course?
● To communicate information and knowledge
● To make a proposition
● To test existing knowledge
● To practice skills
● To inspire and motivate
● The first thing to get clear in your mind is the objectives of the entire
course or one session.
When
● At what time of the day will my session (s) take place? After lunch is
known as the graveyard slot; therefore, you should consider making it more
active than, say, a morning session.
How
● How much time have I got?
● How am I going to present my subject?
● Straight talk
● Talk with overheads
● Talk with PowerPoint presentation
● Talk with video
● Give the participants a period in which to discuss aspects of the subject, e.g., by
using a case study
● Combination of any of these methods
● Should I allow questions during the session?
● Always leave time at the end for questions and discussion
Where
● Where is the presentation due to take place?
● How do the windows open/air conditioning work? If using PowerPoint or video,
how do we darken the room?
● What equipment have they got, e.g., video, computer, projector, overhead
projector, etc.?
● Decide on seating arrangements
● Are there likely to be any distractions, e.g., loud air-conditioning, things
happening outside the window, etc.
● Can I be heard at the back of the room?

____________________________________________________________________________________
4 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Who
● Who are the participants? How senior/junior are they?
● How many participants will be present?
● What is the extent of their existing knowledge of the subject I am going to present?
● What will be of interest to them?
● What will their attitudes, preconceptions or expectations be?
● Is there a gender balance within the group?
● Can you foresee or expect any kind of dynamics or potential resistance due to
group composition?
● What can you glean overall from the participants’ list and profile without making
too many assumptions?

____________________________________________________________________________________
5 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Icons Used in this Guide

Icon Description/Guidelines

Facilitator Information

Show a slide < Used to denote the slide to be shown>. Even better,
paste the image of the slide being discussed.

Show a video

<mention video/clip name>

Evaluate – administer assessment

<mention name and guidelines for assessment>

Narrate/Share a story or valid examples

Share insights or ask participants to share insights about the current


topic

Distribute Handouts

Transition from one subject/topic/objective/story etc to another


(also could indicate flow)

____________________________________________________________________________________
6 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Derive objective/key point

Materials required

Ask the following Questions

Group Discussion

<mention guidelines (the number in each group, whether a team


leader is required in each group, etc) and duration>

Play music

< mention file names and duration>

Capture on a flipchart and put up in the class, which can be


referenced at a later point during the class or to summarize the
learning of the session

Summarize the session/day

Activity – describe activity

____________________________________________________________________________________
7 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Table of Contents
Unit Topic/Sub-topic Duration Page No.
I Time Series Forecasting 3 Hrs. 9-24
II Clustering and Classification 6 Hrs. 25-40
III Advanced Data Visualization Using 3 Hrs. 41-57
Matplotlib
IV Data Visualization Using Seaborn, 3 Hrs. 58-63
Bokeh
V Dimension Reduction Methods 3 Hrs. 64-75
VI Model Selection and Evaluation 3 Hrs. 76-96
VII Neural Networks, Text Mining 6 Hrs. 97-123
VIII Reinforcement Learning 3 Hrs. 124-137
References 138

Credits: 2
Hours: 30

____________________________________________________________________________________
8 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Unit I: Time Series Forecasting


● Time Series Analysis and Forecasting Using R and Python
● Time Series – Date and Time-Types of Data and Tools
● Ranges
● Frequencies and Shifting
● Time Zones
● Time Series – Periods and Periodic Arithmetic Sampling
● Resampling
● Frequency Conversion
● Time Series Plotting – Moving Window Functions

Time Series Analysis and Forecasting Using R and Python


Sometimes, data changes over time. This data is called time-dependent data. Using time-
dependent data, quickly analyze the past data to predict the future data. The future prediction
will include time as a variable or attribute, and the output will vary with time. They are using
time-dependent data to find the specific patterns that repeat again and again over time.
A Time Series is a set of observations that collect the data set after the regular time interval.
To plot, the time series always has one of its axes as time.

Figure 1.1: Time Series

Figure 1.2 : Time Arrangement Analysis


____________________________________________________________________________________
9 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Components of Time Series Analysis


The chart below shows the distinctive components of Time Series Analysis:

Figure 1.3: Components of Time Arrangement Analysis


1. Trend: Trend shows the variation of data with time or frequency of data. The use of
trends, how the data can increase or decrease over time, population, stock market, and
production in a company or corporate sector.
2. Seasonality: Seasonality is used to find the variations at regular intervals of time.
Some of the traditional examples are festivals, conventions, seasons, etc. These
variations happen around the same time and affect the data in specific ways that can
be predicted.
3. Irregularity: The time series data do not correspond to the trend. Unforeseeable
circumstances or situations randomly cause these variations or divergences in the time
series.
4. Cyclic: Oscillations in time series that last for more than one year are known as
cyclic.

Time Series Analysis and Forecasting with R and Python

The ARIMA model stands for Auto-Regressive Integrated Moving Normal. It is utilized to
anticipate the long-term values of a time arrangement, using its past values and figure
blunders. Here, the graph shows the components of an ARIMA model:

Figure 1.4: Components of ARIMA


____________________________________________________________________________________
10 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Moving Average: Moving Average could be a strategy that takes the upgraded normal of
values to reduce noise. It takes the average over a particular interval by taking diverse subsets
of the information and finding their specific averages. A bunch of information focuses and
takes their average to discover another normal by expelling the primary esteem of the
information and counting the other values of the series.

Figure 1.5: Stationarity Utilizing Moving Average


Each of the values acts as a parameter for our ARIMA model. The ARIMA shown by these
different operators and models, these are the parameters:

1. p: Past slacked values for each time point. Determined from the Auto-Regressive Model.
2. q: Previous slacked values for the blunder term. It was determined from the Moving
Average.
3. d: The number of times information is differentiated to form it stationary. It is the
number of times it performs integration.

Time Arrangement Investigation in Python: Time Arrangement Investigation in Python


utilizes a dataset that subtle elements the month-to-month cleanser deals over three a long
time.

Figure 1. 6: Bringing in Vital Modules

____________________________________________________________________________________
11 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Here is an example to perform Time Arrangement Investigation in Python:

Output of the code:

For time as the independent variable, Time Series Analysis (TSA) examines the nature of
the response variable. The time variable is the reference point for estimating the target
variable in forecasting or prediction. A range of time-based orders, including years, months,
weeks, days, hours, minutes, and seconds, are represented by TSA. This observation is
derived from the discrete temporal sequence of subsequent intervals. Examples of TSA's
practical uses are weather forecasting models, stock market forecasts, signal processing, and
control systems.
____________________________________________________________________________________
12 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

The following procedures must be followed to complete the time series analysis:
● Collecting the data and cleaning it
● Display of time versus significant features
● To observe the stationarity of the series
● Develop charts to comprehend its nature.
● Model construction: ARMA, ARIMA, MA, and AR
● Concluding forecasts

Time Series – Date and Time

TSA is the foundation of forecasting and prediction analysis, particularly for time-based
problem statements such as the following;
● Examining the trends in the historical dataset
● To understand and match the current situation with patterns derived from the previous
stage. To be aware of the variables that influence particular variables at different times.

We can generate various time-based studies and findings using "Time Series."
● Forecasting: It is the process of estimating any future value.
● Segmentation: Assemble related things into groups.
● Sorting: A group of things into designated classes is called classification.
● To ascertain the contents of a given dataset, use descriptive analysis.
● Analysis of the intervention: Impact of modifying a specific variable on the result.
● Trend: A continuous timeline with no fixed period when there is divergence within the
provided dataset. There would be a neutral, positive, or negative trend.
● Seasonality: A continuous timeline with frequent or defined interval shifts within the
dataset, resembling a saw tooth or bell curve.
● Cyclical: When there is no set.
The time series includes the following constraints, which we must consider when analyzing
the data. Like other models, TSA does not support the missing.
● The relationships between the data points must be linear.
● Data transformations are expensive since they are required.
● Models primarily work on uni-variate data.
There are two significant time series data types: stationary and non-stationary.
Stationary: A dataset should follow the thumb rules below without having the time series'
trend, seasonality, cyclical, and irregularity components.
● Their mean value should be utterly constant in the data during the analysis.
● The variance should be constant to the time frame
● Covariance measures the relationship between two variables.

Non-stationary: The dataset is called non-stationary if the mean-variance or covariance


changes concerning time.

____________________________________________________________________________________
13 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Time Series - Date and Time - Types of Data and Tools

Python and R, handling date and time data is essential, especially in time series analysis.
Timestamps represent a specific time point and include date and time information. Time
Intervals and Periods: Represent continuous periods, such as "1 month" or "1 year". Time
Durations: Represent a length of time independent of any specific start or end point.

Tools for Date and Time in Python:


Pandas is a powerful Python library for data manipulation and analysis. It provides the
Timestamp and DateTime Index data structures for working with date and time data.

Functions like pd.to_datetime() can convert strings or numerical values to DateTime objects,
providing functionalities for date/time arithmetic, slicing, and resampling.

DateTime Module:
Python's built-in datetime module provides classes for manipulating dates and times. Includes
classes like DateTime, date, time, and time delta for representing and performing various
operations on date and time values.

____________________________________________________________________________________
14 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Tools for Date and Time in R:


Lubricate:
The lubridate package in R provides functions to parse, manipulate, and work with date-time
objects. Offers intuitive functions like ymd(), mdy(), hms(), etc., for parsing date-time
strings. Provides functionalities for arithmetic operations, extracting components, and
handling time zones.

Zoo (Z's Ordered Observations):

The Zoo package provides tools for handling ordered observations, including time series data.
It offers functionalities for indexing, merging, and aggregating time series data.
Everyday Operations for Date and Time Data:
Parsing: Converting date and time data from string or numerical formats into datetime
objects.
Formatting: To represent datetime objects in different formats for display or storage.
Arithmetic Operations: Performing operations on datetime objects, such as addition,
subtraction, and comparison.
Indexing and Slicing: Selecting specific date/time ranges or intervals from time series data
for analysis.
Resampling and Aggregation: Using resampling techniques, changing the frequency or
granularity of time series data.

Machine Learning in Python

import pandas as pd

# Creating a DateTime object

date_obj = pd.to_datetime('2024-02-07')

# Performing arithmetic operations

next_day = date_obj + [Link](days=1)

print("Current date:", date_obj)

print("Next day:", next_day)

Machine Learning in R library(lubridate)

# Creating a date-time object

date_time_obj <- ymd_hms("2024-02-07 12:00:00")

# Performing arithmetic operations

____________________________________________________________________________________
15 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

next_day <- date_time_obj + days(1)

print(paste("Current date:", date_time_obj))

print(paste("Next day:", next_day))

Both examples demonstrate creating datetime objects and performing arithmetic operations.

Ranges

In time series analysis, ranges refer to the period the data covers. It includes the start and end
dates or timestamps of the dataset. Understanding the range of the data is crucial for
determining the overall period under consideration.

import pandas as pd

# Assuming 'time_series_data' is a Pandas DataFrame or Series with a range in machine


learning

In machine learning, "range" typically refers to the range of values a feature (or variable) can
take within a dataset. It can also refer to scaling or normalizing features to a specific range to
improve the performance of machine learning algorithms.

Range of Values:
Understanding the range of values for each feature in the dataset is essential for various
reasons:
Data Understanding: Knowing the range of values to understand the distribution and
characteristics of the data.
Feature Selection: It can aid in feature selection by identifying features with a narrow or
wide range of values that may be relevant to the task at hand.
Preprocessing: It guides preprocessing steps such as scaling or normalization, ensuring that
features are appropriately transformed for modeling.
Scaling and Normalization
Scaling and normalization are preprocessing techniques to transform features to a specific
range or distribution. Common methods include the following;
Min-Max Scaling: Rescales features to a fixed range, typically [0, 1] or [-1, 1], preserving
the original shape of the distribution.
Standardization (Z-score normalization): Scales feature for mean value 0 and standard
deviation value of 1 results in a distribution centered around 0.
Robust Scaling: Scales features using robust statistics to mitigate the effect of outliers.

from [Link] import MinMaxScaler

# Assuming 'X' is your feature matrix


____________________________________________________________________________________
16 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

scaler = MinMaxScaler(feature_range=(0, 1))

X_scaled = scaler.fit_transform(X)

R using caret in Machine Learning :

library(caret)

# Assuming 'X' is your feature matrix

X_scaled <- preProcess(X, method = c("range"), range = c(0, 1))

The range of values in features impacts the performance and behavior of machine
learning algorithms:

Gradient Descent: Algorithms like gradient descent may converge faster when features are
within a similar range.
Regularization: Regularization techniques penalize significant coefficients, so scaling
features to a standard range can prevent certain features from dominating the model.
Distance-Based Algorithms: Algorithms that rely on distance metrics, such as k-nearest
neighbors (KNN) or support vector machines (SVM), can be sensitive to the scale of features.
Data Understanding: Knowing the range of values and understanding the distribution and
characteristics of the data.
Feature Selection: It can aid in feature selection by identifying features with a narrow or
wide range of values that may be relevant to the task at hand.
Preprocessing: It guides preprocessing steps such as scaling or normalization, ensuring that
features are appropriately transformed for modeling.
Scaling and Normalization:
Scaling and normalization are preprocessing techniques to transform features to a specific
range or distribution. Common methods include:
Min-Max Scaling: Rescales features to a fixed range, typically [0, 1] or [-1, 1], preserving
the original shape of the distribution.
Standardization (Z-score normalization): Scaling features to a mean of zero and a standard
deviation of one results in a distribution centered around 0.
Robust Scaling: Scales features using robust statistics to mitigate the effect of outliers.

Frequencies and Shifting


Frequencies, shifting, and time zone adjustments are relevant in machine learning when
dealing with time series data or when working with data collected from different time zones.
These concepts are applied in machine learning using Python and R:In machine learning,
frequencies often aggregate or resample time series data to different intervals. It can be
helpful for feature engineering or data preprocessing.
____________________________________________________________________________________
17 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Python (Pandas):

# Resample time series data to a different frequency (e.g., weekly)

weekly_data = time_series_data.resample('W').sum()

# Resample time series data to a different frequency (e.g., weekly)

weekly_data <- aggregate(time_series_data,


[Link]([Link](index(time_series_data))$wday), sum)

Shifting time series data involves moving the data forward or backward in time. IT can help
create lag features or align features with the target variable in supervised learning tasks.

Python (Pandas):

# Create lag features by shifting the time series data

lagged_data = time_series_data.shift(periods=1)

R (xts or zoo):

# Create lag features by shifting the time series data

lagged_data <- lag(time_series_data, k=1)

Time Zones

When working with data collected from different time zones, it is essential to adjust the
timestamps to a standard time zone to ensure consistency in analysis or modeling.

Python (Pandas):

# Assuming 'time_series_data' is in UTC and you want to convert it to US/Eastern time zone

time_series_data = time_series_data.tz_localize('UTC').tz_convert('US/Eastern)

R (xts or zoo):

# Assuming 'time_series_data' is in UTC and you want to convert it to US/Eastern time zone

time_series_data <- [Link](time_series_data, tz = "UTC")

time_series_data <- [Link](format(time_series_data, tz = "US/Eastern"))

____________________________________________________________________________________
18 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Importance in Machine Learning:


Feature Engineering: Frequencies and shifting are essential for creating informative features
from time series data.
Data Preprocessing: Time zone adjustments ensure consistency when merging or comparing
data from different sources.
Model Training: Proper handling of frequencies and time zones can improve the accuracy
and interpretability of machine learning models, especially when dealing with temporal data.

Time Series – Periods and Periodic Arithmetic Sampling


Time series analysis often involves dealing with periodic data, where observations are
collected at regular intervals over time. In this context, "periods" refer to the intervals at
which data is sampled or measured, and "periodic arithmetic sampling" refers to the process
of sampling data at regular intervals.
Both R and Python provide libraries and packages for time series analysis, including methods
for handling periodic data and performing sampling. Here is a brief overview of how you can
handle time series periods and perform periodic arithmetic sampling using R and Python:
Using R:
1. Handling Time Series Data: R provides several packages for time series analysis,
including ts, xts, and zoo. To create time series objects using functions like ts() or xts::xts().
To specify periods, you typically set the frequency parameter in the time series object.
2. Periodic Arithmetic Sampling:
To perform periodic arithmetic sampling using functions like window() or indexing
operations.
The window() function allows you to extract a subset of the time series corresponding to a
specified time window or period.

Using Python:
1. Handling Time Series Data:
Python provides libraries such as pandas and statsmodels for time series analysis.
Create time series objects using the pandas. Series or [Link] class.
Pandas provides the pd.date_range() function for generating date indices with specified
frequencies.
2. Periodic Arithmetic Sampling:
Pandas provides various methods for periodic arithmetic sampling, including slicing
using datetime indices.
Here, use functions like loc[] or slicing with datetime indices to extract data corresponding to
specific periods.

Import pandas as pd

# Create a time series with daily frequency


____________________________________________________________________________________
19 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

dates = pd.date_range(start='2024-01-01', end='2024-01-10', freq='D')

data = [Link](range(len(dates)), index=dates)

# Perform periodic arithmetic sampling (e.g., every two days)

sampled_data = data[::2]

print(sampled_data)

This example creates a time series with daily frequency and then performs periodic arithmetic
sampling to extract data every two days.

Both R and Python offer robust capabilities for handling time series data and performing
periodic arithmetic sampling, with libraries like statsmodels (for R) and pandas (for Python)
being popular choices for such tasks.

Resampling

Resampling in machine learning refers to creating new datasets from an existing dataset
through various sampling techniques. It's commonly used to assess machine learning models'
performance, improve model generalization, and handle imbalanced datasets. Two standard
techniques for resampling are:

Cross-Validation: This technique involves partitioning the dataset into subsets, training the
model on some subsets, and evaluating it on the remaining subset(s). Cross-validation helps
to estimate the model's performance and generalization ability.

Bootstrapping: Bootstrapping is a resampling technique in which multiple datasets are


generated by randomly sampling the original dataset with a replacement. Each bootstrapped
dataset is used to train a model, and the results are aggregated to estimate the model's
performance and uncertainty.

Cross-Validation:
Cross-validation is a widely used resampling technique for estimating the performance of
machine learning models.

The standard form of cross-validation is k-fold, where a dataset is divided into k subsets
(folds). The k times trained model uses k-1 folds as the training data set and the remaining
fold as the validation data.

Python (using sci-kit-learn):

From sklearn.model_selection import cross_val_score, KFold

from sklearn.linear_model import LogisticRegression

____________________________________________________________________________________
20 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

import numpy as np

# Assuming 'X' is your feature matrix and 'y' is your target vector

model = LogisticRegression()

k_fold = KFold(n_splits=5, shuffle=True, random_state=42)

cross_val_scores = cross_val_score(model, X, y, cv=k_fold)

print("Cross-Validation Scores:", cross_val_scores)

print("Mean CV Score:", [Link](cross_val_scores))

R (using caret):

library(caret)

# Assuming 'X' is your feature matrix and 'y' is your target vector

model <- train(X, y, method = "glm", trControl = train control(method = "cv", number = 5,
verboseIter = TRUE))

print(model)

Bootstrapping is a resampling technique that involves repeating the sampling data from the
dataset with replacement to create multiple bootstrap samples. These samples are then used to
train and evaluate multiple models, and the results are aggregated to estimate the model's
performance.

Python (using sci-kit-learn):


From [Link] import resample
from [Link] import accuracy_score
# Assuming 'X_train' and 'y_train' are your training data
# 'base_model' is your base classifier (e.g., DecisionTreeClassifier)
n_iterations = 100
bootstrap_scores = []
for _ in range(n_iterations):

X_boot, y_boot = resample(X_train, y_train, random_state=42)

model = base_model.fit(X_boot, y_boot)

y_pred = [Link](X_test)

____________________________________________________________________________________
21 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

accuracy = accuracy_score(y_test, y_pred)

bootstrap_scores.append(accuracy)

print("Bootstrap Scores:", bootstrap_scores)

print("Mean Bootstrap Score:", [Link](bootstrap_scores))

Using R:
Bootstrapping in R can be implemented using various packages, such as boot or caret.
library(boot)
# Assuming 'X' is your feature matrix and 'y' is your target vector
# 'glm' is your base model
boot_results <- boot(data = [Link](X, y), statistic = function(data, indices) {
model <- glm(y ~ ., data = data[indices, ])
return(model)
}, R = 100)
print(boot_results)
These are basic examples of performing resampling techniques like cross-validation and
bootstrapping in machine learning using Python, and R. Resampling is an essential aspect of
model evaluation and can provide more robust estimates of a model's performance and
generalization ability.

Frequency Conversion
In machine learning, converting frequencies refers to transforming time series data from one
frequency (e.g., daily) to another (e.g., weekly, monthly). This conversion is often used for
feature engineering or data preprocessing purposes. With the given example, we can perform
frequency conversion in Python and R:
Frequency Conversion in Python (using Pandas):
import pandas as pd
# Assuming 'time_series_data' is a Pandas DataFrame or Series with a DateTimeIndex
# Convert daily data to weekly data by summing up the values for each week
weekly_data = time_series_data.resample('W').sum()
# Convert daily data to monthly data by averaging the values for each month
monthly_data = time_series_data.resample('M').mean()
# Convert daily data to quarterly data by taking the maximum value for each quarter
quarterly_data = time_series_data.resample('Q').max()
Frequency Conversion in R (using xts or zoo):

____________________________________________________________________________________
22 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Importance in Machine Learning:


Frequency conversion is essential in machine learning for several reasons:
Feature Engineering: Converting frequencies allows you to create features at different time
resolutions, which can provide additional insights into machine learning models.
Data Preprocessing: Adjusting the frequency of time series data can help handle irregular or
unevenly spaced data, making it more suitable for analysis.
Model Training: Depending on the task and the nature of the data, different time resolutions
may be more suitable for training machine learning models.
By converting frequencies, you can tailor the time series data better to suit the requirements
of your machine learning tasks, potentially improving model performance and insights.

Time Series Plotting - Moving Windows Function


Plotting time series data with moving windows is a common technique used in machine
learning for visualizing trends and patterns over time. Moving windows allow one to
calculate statistics or aggregates over a rolling window of time, which helps in smoothing out
noise and identifying underlying trends.

import pandas as pd#

import [Link] as plt#

# Assuming 'time_series_data' is a Pandas DataFrame or Series with a DateTimeIndex

# Assuming 'window_size' is the size of the moving window


# Calculate rolling mean with a moving window
rolling_mean = time_series_data.rolling(window=window_size).mean()
# Plot original time series data
[Link](figsize=(10, 6))
[Link](time_series_data, label='Original Data')
# Plot rolling mean
[Link](rolling_mean, color='red', label='Rolling Mean (Window
Size={})'.format(window_size))
[Link]('Time Series Data with Moving Window')
[Link]('Date')
[Link]('Value')
[Link]()
[Link]()

Adjust the window_size parameter to control the smoothing effect of the moving window.
Larger window sizes result in smoother trends but may obscure short-term fluctuations, while
____________________________________________________________________________________
23 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

smaller window sizes capture more detail but may be susceptible to noise. To calculate other
statistics or aggregates using different window functions such as rolling_sum(),
rolling_median(), rolling_std(), etc., depending on the specific analysis requirements. In R, to
achieve similar functionality using libraries like zoo or xts)

library(xts)

# Assuming 'time_series_data' is a time series object (e.g., zoo or xts)

# Assuming 'window_size' is the size of the moving window

# Calculate rolling mean with a moving window

rolling_mean <- roll mean(time_series_data, k=window_size, [Link]=TRUE, align='right')

# Plot original time series data

plot(time_series_data, type='l', col='blue', lwd=2, main='Time Series Data with Moving


Window', xlab='Date', ylab='Value')

# Plot rolling mean

lines(rolling_mean, col='red', lwd=2)

legend("top right", legend=c("Original Data", paste("Rolling Mean (Window Size=",


window_size, ")", sep="")), col=c("blue", "red"), lty=1, lwd=2)

This R code calculates the rolling mean using a moving window of size window_size and
plots the original time series data along with the rolling mean.

Adjust the window_size parameter to control the smoothing effect of the moving window, as
explained earlier.

These visualizations can provide insights into the underlying trends and patterns in the time
series data, helping in feature engineering and model selection for machine learning tasks.

____________________________________________________________________________________
24 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Unit II: Clustering and Classification


● Clustering and Reinforcement Learning Using R and Python
● Scikit Learn Library in Python
● K-means Clustering
● Hierarchical Clustering
● Recommendation System
● Apriori Algorithm
● ECLAT Algorithm
● Upper Confidence Bound Learning
● Thompson Sampling

Clustering and Reinforcement Learning Using R and Python

Clustering and Classification is the task of predicting a new observation's category or class
label based on past observations labeled with their corresponding class labels. It is a
supervised learning approach, meaning the algorithm learns from labeled data.
Classification's main aim is to learn a mapping from input features to predefined output
labels. It is used for making predictions on unseen data - for example, spam email detection,
sentiment analysis, disease diagnosis, object recognition, etc.

Reinforcement:
An unsupervised learning technique is used in Clustering to group similar objects or data
points into clusters based on certain features or characteristics. Clustering is utilized in
various fields, such as customer segmentation, image processing, anomaly detection, and
recommendation systems.
Clustering reinforcement are two fundamental techniques in machine learning used for
different types of tasks: Clustering is unsupervised learning, while classification is supervised
learning. Clustering is the grouping of objects so the objects in the same group (called a
cluster) are more similar to each other than those in different groups.
It is an unsupervised learning approach, meaning the algorithm learns the data structure
without any labeled output. The goal of Clustering is to discover inherent groupings in the
data. It helps identify patterns, segment data, and understand the underlying structure—for
example, Customer segmentation, document clustering, image segmentation, anomaly
detection, etc.

#Example using K-means):


from [Link] import KMeans
import numpy as np

____________________________________________________________________________________
25 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

# Generate some random data

[Link](0)
X = [Link](100, 2)
# Perform K-means clustering

means = KMeans(n_clusters=3)
[Link](X)
# Get cluster centers and labels
centers = means.cluster_centers_

labels = means.labels_
print("Cluster Centers:")
print(centers)
print("Cluster Labels:")

print(labels)

Output for the example using google colab:

Reinforcement learning is widely used in machine learning, where the agent learns to take
action in an environment to maximize some of the notion of cumulative reward. It learns
through trial and error by interacting with the environment. Reinforcement learning is applied
in various domains such as robotics, game playing, autonomous vehicles, finance, and
healthcare.

____________________________________________________________________________________
26 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Here is the example with output for Reinforcement Learning:

These examples illustrate the implementations of Clustering using K-means in Python and
reinforcement learning using the CartPole environmerom OpenAI. Depending on the specific
requirements, explore more advanced algorithms and techniques in both clustering and
reinforcement learning.

R Example (K-means):using Clustering andReinforcement Learning

# Load necessary library

library(stats)

# Generate some random data

[Link](123)

data <- matrix(rnorm(100), ncol=2)

# Perform K-means clustering

kmeans_result <- kmeans(data, centers=3)#

# Display cluster centers

print(kmeans_result$centers)

# Display cluster assignments


____________________________________________________________________________________
27 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

print(kmeans_result$cluster)

Python Example (K-means using scikit-learn):

# Load necessary libraries

from [Link] import KMeans

import numpy as np

# Generate some random data

np. [Link](123)#

data = np. [Link](100, 2)###

# Perform K-means clustering

kmeans = KMeans(n_clusters=3)

[Link](data)

# Display cluster centers

print(kmeans.cluster_centers_)

# Display cluster assignments

print(kmeans.labels_)

____________________________________________________________________________________
28 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Scikit Learn Library in Python

Scikit-Learn, commonly called sklearn, is a popular open-source machine-learning library in


Python. It provides the tools for unsupervised and supervised learning and the data set for
preprocessing, model evaluation, and more.

Some key features and components of scikit-learn:


Simple and Efficient Application Programming Interface (API): Scikit-learn offers a
consistent and straightforward API for various machine learning algorithms and tasks,
making it easy to use and learn.
Extensive Algorithm Support: The EAS algorithm implements a wide range of machine
learning algorithms, which include classification, regression, Clustering, dimensionality
reduction, and model selection. Integration with Scientific Libraries: Scikit-learn seamlessly
integrates with other Python libraries such as NumPy, SciPy, and Matplotlib, allowing for
easy manipulation of data arrays and visualization of results.
Model Evaluation and Selection: The library provides tools for model evaluation and
selection, including cross-validation, hyperparameter tuning, and various performance
metrics.
Flexibility and Customization: Scikit-learn allows for customization and fine-tuning of
models through various parameters and options, enabling users to adapt algorithms to specific
tasks and datasets.
Preprocessing and Feature Engineering: It offers many preprocessing techniques for data
cleaning, feature scaling, transformation, and feature selection.
Pipeline Construction: Scikit-learn supports the construction of processing pipelines,
allowing users to chain multiple preprocessing steps and machine learning models into a
single workflow.
Active Development and Community Support: An open-source project, scikit-learn is
actively developed and maintained by a large community of contributors, ensuring regular
updates, bug fixes, and improvements. Scikit-learn is widely used in academia and industry
for various machine-learning tasks due to its simplicity, efficiency, and extensive
functionality. Scikit-learn provides powerful tools for building and deploying machine-
learning models in Python.
K-means Clustering
K-means Clustering is used for unsupervised machine learning algorithms for clustering data
into groups or clusters. It aims to partition data points into the K clusters for each data point
belonging to the cluster with the nearest mean (centroid). It is widely used in various fields,
such as data mining, pattern recognition, and image segmentation. How the algorithm works:
Initialization: Randomly select K data points as the initial centroids, representing the initial
cluster's centers.

____________________________________________________________________________________
29 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Assignment: Each data point is assigned to the nearest centroid based on a distance metric
such as Euclidean distance. Each data point or data set is allocated to the cluster whose
centroid is closest to it.
Update Centroids: After assigning all data points to clusters, recalculate the centroids by
taking the mean of all data points allocated to each cluster.
Repeat: Repeat steps 2 and 3 until convergence. Convergence occurs when the centroids no
longer change significantly or after a certain number of iterations.
Final Clustering: Once the algorithm converges, the final clusters are obtained, and each
data point is associated with a cluster.

The K-means clustering is sensitive to the initial selection of centroids. Depending on the
initial centroids and data distribution, it can converge to a local optimum rather than a global
optimum. The algorithm is often run multiple times with different initializations, and the best
Clustering based on some criterion (e.g., minimizing the total within-cluster variance) is
selected. Additionally, variations and improvements to the basic K-means algorithm, such as
K-means++, provide a better initialization strategy to improve convergence and the
Clustering quality. K-means is efficient and scalable, making it suitable for large datasets.
However, it assumes that clusters are spherical and have similar sizes, which may only
sometimes hold in real-world data. In cases where clusters have complex shapes or varying
densities, other clustering algorithms like DBSCAN or hierarchical Clustering may be more
appropriate. An example of K-means clustering using Python with scikit-learn:

from [Link] import KMeans

import numpy as np
import [Link] as plt

# Generate some random data


[Link](0)
X = [Link](100, 2)
# Initialize the K-means clustering algorithm with 3 clusters

kmeans = KMeans(n_clusters=3)
# Fit the K-means model to the data
[Link](X)
# Get the cluster centers and labels

centers = kmeans.cluster_centers_
labels = kmeans.labels_
# Plot the data points and cluster centers
[Link](X[:, 0], X[:, 1], c=labels, cmap='viridis', alpha=0.5)

____________________________________________________________________________________
30 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

[Link](centers[:, 0], centers[:, 1], c='red', marker='X', s=200)

[Link]('Feature 1')
[Link]('Feature 2')
[Link]('K-means Clustering')

[Link]()

In this above example, we generate random data points in a 2D space, then apply K-means
clustering with 3 clusters. We plot the data points colored by their cluster assignments and the
cluster centers (centroid) in red.

K-means clustering using scikit-learn in Python:


from [Link] import KMeans###
import numpy as np
import matplotlib. pyplot as plt
np. [Link](0)
X = np. [Link](100, 2)
____________________________________________________________________________________
31 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

# Initialize the K-means clustering algorithm with 3 clusters

kmeans = KMeans(n_clusters=3)
[Link](X)
centers = kmeans.cluster_centers_

labels = kmeans.labels_
[Link](X[:, 1], X[:, 0], c=labels, cmap='viridis', alpha=0.5)
[Link](centers[:, 0], centers[:, 1], c='blue', marker='X', s=200)
[Link]('Feature 1')

[Link]('Feature 2')
[Link]('K-means Clustering')
[Link]()

In the above example with output , we generate a random dataset X with 100 data points and
two features. We initialize the K-means clustering algorithm with 3 clusters using
KMeans(n_clusters=3). We fit the K-means model to the data using the fit() method.

____________________________________________________________________________________
32 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Hierarchical Clustering
Hierarchical Clustering is now a popular unsupervised learning technique for grouping
similar objects into clusters based on their pairwise distances. It creates a hierarchy of
clusters where clusters at the same level are more similar to each other than those at higher
levels.
There are two main types of hierarchy:
Agglomerative Hierarchical Clustering: In this approach, each data point initially forms its
cluster, and then pairs of clusters are iteratively merged based on their similarity until all data
points belong to a single cluster. The merging process continues until a stopping criterion is
met, such as a specified number of clusters or a threshold distance.
Divisive Hierarchical Clustering: This approach works opposite to agglomerative
Clustering. It starts with a single cluster containing all data points and recursively splits
clusters into smaller clusters until each data point is in its cluster. It continues until a stopping
criterion is met, similar to agglomerative Clustering.

Hierarchical Clustering is often visualized using dendrograms, which display the merging (or
splitting) process and help interpret the hierarchy of clusters.

Here is an example of agglomerative hierarchical Clustering using scikit-learn:

from [Link] import make_blobs

from [Link] import AgglomerativeClustering


import [Link] as plt

X, _ = make_blobs(n_samples=100, centers=3, cluster_std=1.0, random_state=42)


agg_clustering = AgglomerativeClustering(n_clusters=3)
agg_clustering.fit(X)
[Link](X[:, 0], X[:, 1], c=agg_clustering.labels_, cmap='viridis', alpha=0.5)

[Link]('Feature of 1')
[Link]('Feature of 2')
[Link]('**Agglomerative Hierarchical Clustering**')
[Link]()

____________________________________________________________________________________
33 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

In this example and the output, to generate random data points in 2D space using
make_blobs, then apply agglomerative hierarchical Clustering with 3 clusters using
Agglomerative Clustering. Finally, to visualize the data points colored by their cluster
assignments.

Recommendation System
Recommendation systems are used in various domains to provide personalized
recommendations to users. They can be implemented using clustering and classification
techniques. Clustering algorithms can group similar users or items based on their features or
behavior. Once clusters are formed, recommendations can be made by suggesting popular
items within a cluster or by recommending items liked by similar users in the cluster.
User Clustering: Group users into clusters based on their preferences, behaviour, or
demographic information. For example, K-means clustering can group users into clusters
based on their ratings or purchase history.
Item Clustering: Group items into clusters based on their features or characteristics. For
example, clustering can be applied to group movies or products into clusters based on their
genres, attributes, or tags.
Recommendation Generation: Once clusters are formed, recommendations can be
generated by suggesting popular items within a cluster or by recommending items liked by
similar users in the cluster.
Prediction and Recommendation: Use the trained classifier to predict the probability of a
user liking a new item. Items with high predicted probabilities can be recommended to the
user. Logistic regression or decision trees can be trained on historical user-item interactions
to predict whether a user will like a new item, and recommendations can be made based on
these predictions.

____________________________________________________________________________________
34 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Apriori Algorithm
The Apriori algorithm is a classic algorithm used in association rule mining, a technique in
data mining that discovers exciting relationships hidden in large datasets. While the Apriori
algorithm is not directly related to clustering or classification, it can be combined with
clustering or classification techniques for specific tasks.
The Apriori algorithm primarily identifies frequent item sets in transactional databases. It
works to iteratively find a frequent item set with an increase in the size of the "apriori"
property to state that any subset of a frequent item set must also be frequent.
Frequent Item Set Generation: The Apriori algorithm starts by identifying all individual
items that occur frequently in the dataset (frequent 1-item sets). It then iteratively generates
larger item sets by combining frequent (k-1)-item sets to form candidate sets, which are
subsequently pruned based on the minimum support threshold.
Association Rule Generation: Once frequent item sets are identified, association rules are
generated by forming rules of the form A → B, where A and B are item sets, and calculating
metrics such as support, confidence, and lift to measure the strength of the association
between A and B.
Apriori with Clustering: Apriori algorithm itself is not directly used for clustering, it can be
employed in conjunction with clustering techniques for specific tasks.
As an example Market Basket Analysis: Apriori can be used to discover frequent item sets
in transactional data, such as customer shopping baskets. These item sets can then identify
patterns and associations among products purchased together. Clustering techniques can be
applied to segment customers based on their purchasing behaviour. Apriori can be used
within each cluster to identify specific item sets or associations relevant to that cluster.
Feature Selection: In classification tasks, Apriori can be used for feature selection by
identifying frequent item sets of features that occur together frequently in the dataset. These
frequent item sets can be utilized as input features for classification models to reduce the
dimensionality of the feature space to improve model performance.

ECLAT Algorithm

The ECLAT algorithm is an Equivalence Class Clustering and Bottom-Up Lattice Traversal
algorithm for frequent item set mining in transactional databases, similar to the Apriori
algorithm. While ECLAT is not inherently a clustering or classification algorithm, it can be
used as a preprocessing step or in conjunction with clustering and classification techniques
for specific tasks. ECLAT is designed to mine frequent itemsets from transactional databases
efficiently. Frequent Item Set Generation: ECLAT recursively finds frequent item sets by
exploring the lattice structure of the item sets. It starts by identifying frequent 1-item sets and
then extends them to larger item sets by recursively combining them with other frequent item
sets.
Equivalence Class Clustering: ECLAT uses an equivalence class to reduce the search
space. It clusters transactions that share everyday items into equivalence classes, and the

____________________________________________________________________________________
35 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

search for frequent item sets is performed within these equivalence classes. While ECLAT is
primarily used for frequent item set mining, it can indirectly contribute to clustering tasks:

Feature Extraction: In clustering tasks, ECLAT can be used to identify frequent item sets of
features occurring together in transactional data. These frequent item sets can then be treated
as new features or dimensions in the dataset, potentially capturing meaningful patterns or
associations among features. The clustering algorithms can be applied to the transformed
dataset to group similar instances based on these extracted features.

Feature Selection: ECLAT can be used for feature selection by identifying frequent item
sets of features occurring together frequently in the dataset. These frequent item sets can be
utilized as input features for classification models.

Upper Confidence Bound Learning


The Upper Confidence Bound (UCB) algorithm is typically used in the context of
reinforcement learning and multi-armed bandit problems rather than clustering or
classification directly. However, concepts from UCB, such as balancing exploration and
exploitation, can be adapted and applied in clustering and classification tasks. Let us explore
how UCB concepts can be utilized in these contexts:
While UCB is not directly applied to clustering, balancing exploration (trying new clusters)
and exploitation (using existing clusters) can benefit dynamic clustering scenarios.
Cluster Exploration: When new data points arrive over time, and the number of clusters is
not fixed, UCB-like strategies can decide whether to create a new cluster or assign the data
point to an existing cluster. The uncertainty associated with each cluster's centroid can be
quantified. New data points can be assigned to the cluster with the highest uncertainty (i.e.,
upper confidence bound) to encourage exploration. UCB-like approaches can be used for
feature selection, model selection, or hyper-parameter tuning:
Feature Selection: To deal with a large number of features, UCB can be applied to prioritize
feature selection by estimating the uncertainty or importance of each feature. Features with
high uncertainty or potential significance, as determined by UCB, can be selected for further
analysis or inclusion in the classification model.
Model Selection and Hyperparameter Tuning: UCB strategies can also guide the selection
of different classification models or hyper-parameters. For example, in a grid search for
hyper-parameter tuning, UCB can prioritize hyper-parameter combinations that have shown
promising performance while still allowing for exploration of less-explored regions of the
hyper-parameter space.

Example: UCB-inspired Approach for Model Selection in Classification:

import numpy as np#


import pandas as pd#
from sklearn.model_selection

____________________________________________________________________________________
36 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

import cross_val_score
# UCB-inspired model selection
def ucb_model_selection(estimator, param_grid, X, y, n_iter=50, cv=5):
scores = []
n_params = len(param_grid)
n_iter_per_param = n_iter // n_params
for param in param_grid:
for _ in range(n_iter_per_param):
params = {key: [Link](val) for key, val in param_grid.items()}
clf = estimator.set_params(**params)
cv_scores = cross_val_score(clf, X, y, cv=cv)
[Link]((params, [Link](cv_scores)))
return scores

# Generate some random data


X, y = make_blobs(n_samples=100, centers=3, cluster_std=1.0, random_state=42)

# Example usage
model_scores = ucb_model_selection(RandomForestClassifier(), param_grid, X, y)
best_params = sorted(model_scores, key=lambda x: x[1], reverse=True)[0][0]
print("Best parameters:", best_params)

____________________________________________________________________________________
37 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

In this example, a UCB-inspired approach is used to select the best hyperparameters for a
Random Forest classifier by sampling hyperparameter combinations from the grid and
evaluating their performance using cross-validation. The approach balances exploration
(trying different hyperparameter combinations) and exploitation (selecting hyperparameter
combinations with high estimated performance).

Thompson Sampling

Thompson Sampling is a technique primarily used for decision-making under uncertainty. It


is often applied to sequential decision problems such as the multi-armed bandit problem.
While it's not directly used for clustering or classification, its principles can be adapted to
improve these tasks indirectly. In clustering tasks, Thompson Sampling can be employed to
dynamically select the number of clusters or the appropriate clustering algorithm based on the
observed data:

Dynamic Cluster Selection: Thompson Sampling can be utilized to determine the number of
clusters dynamically or the best clustering algorithm based on uncertainty metrics derived
from the data. As new data points are observed, Thompson Sampling adapts its cluster
assignment strategy to fit the observed patterns best.

Classification in Thompson Sampling


In classification tasks, Thompson Sampling can be applied in various ways, including model
selection, hyperparameter tuning, and active learning:
Model Selection and Hyperparameter Tuning: Thompson Sampling can guide the
selection of different classification models or hyperparameters by iteratively exploring the
space based on uncertainty measures and performance feedback from the data.
Active Learning: Thompson Sampling can be used for active learning in scenarios where
labeled data is limited or costly. By selecting the informative instance for labeling based on
uncertainty estimates, Thompson Sampling can improve the efficiency of the learning
process.
Here is the example for Thompson Sampling-inspired Approach for Model Selection in
Classification:
import numpy as np
from sklearn.model_selection import cross_val_score
from [Link] import RandomForestClassifier
from [Link] import make_classification
# Thompson Sampling-inspired model selection
def thompson_model_selection(models, X, y, n_iter=50, cv=5):
best_score = float('-inf')
best_model = None
____________________________________________________________________________________
38 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

for _ in range(n_iter):
# Sample a model from the available options
model = [Link](models)

# Evaluate the model using cross-validation


cv_scores = cross_val_score(model, X, y, cv=cv)
avg_score = [Link](cv_scores)
# Update the best model if the current one performs better
if avg_score > best_score:
best_score = avg_score
best_model = model
return best_model
X, y = make_classification(n_samples=1000, n_features=20, random_state=42)
# Define a list of models to consider
models = [
RandomForestClassifier(n_estimators=100, max_depth=5),
RandomForestClassifier(n_estimators=200, max_depth=10),
RandomForestClassifier(n_estimators=50, max_depth=3)
]
# Select the best model using Thompson Sampling
best_model = thompson_model_selection(models, X, y)

print("Best model:", best_model)

Output of the above code:

____________________________________________________________________________________
39 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

A Thompson Sampling-inspired approach selects the best classification model from a set of
Random Forest classifiers. The algorithm iteratively selects models based on their
performance on cross-validated data, adapting its selection strategy over time.

____________________________________________________________________________________
40 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Unit III: Advanced Data Visualization Using Matplotlib


● Advanced-Data Visualization in Python
● Data Visualization- Quantile-Quantile Plots
● Statistical characterization
● Parallel Coordinates Plots
● Correlation Plots
● Pareto Charts
● Heatmaps
● Advanced Data Visualization in ggplot
● Using ggrepel
● Legends Movements, etc.

Advanced Data Visualization Using Python

Data Visualization: Data Visualization is an essential part of business activities as


organizations nowadays collect a massive amount of data. Sensors worldwide collect climate
data, user data through clicks, car data for predicting steering wheels, etc. All the data
collected hold key business insights, and visualizations make these insights easy to interpret.

Visualizations are the easiest way to analyze and absorb information. Visuals help to
understand the complex problem quickly. They assist in identifying patterns, relationships,
and outliers in data. It assists in understanding business problems better and quickly. It helps
to build a compelling story based on visuals. Insights gathered from the visuals help create
strategies for businesses. It is also a precursor to many high-level data analyses for
Exploratory Data.

Exploratory Data Analysis (EDA) and Machine Learning (ML). Matplotlib is a 2-D plotting
library that helps in visualizing figures. Matplotlib emulates Matlab-like graphs and
visualizations. Matlab is not accessible, is challenging to scale, and is tedious as a
programming language. So, matplotlib in Python is used as it is a robust, accessible, and easy
library for data visualization. In machine learning, advanced data visualization using
Matplotlib often involves creating visualizations to analyze and understand data, evaluate
model performance, and communicate insights effectively. Matplotlib.

Feature Distribution Visualization: Visualizing the dataset's features' distribution can


provide insights into data characteristics and potential preprocessing requirements.

Correlation Matrix Visualization: Visualizing the correlation matrix helps understand the
relationships between features, which is helpful for feature selection and identifying
multicollinearity.

____________________________________________________________________________________
41 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Model Performance Visualization: Visualizing model performance metrics such as


accuracy, precision, recall, or ROC curve helps evaluate and compare different machine
learning models.
Feature Importance Visualization: Visualizing feature importances helps understand which
features contribute most to the model's predictions, aiding feature selection and model
interpretation.

importances = model.feature_importances_

indices = [Link](importances)[::-1]

# Plot feature importance

[Link]()
____________________________________________________________________________________
42 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

[Link]('Feature Importances')

[Link](range([Link][1]), importance[indices], color='b', align='center')

[Link](range([Link][1]), feature_names[indices], rotation=90)

[Link]([-1, [Link][1]])

[Link]()

Partial Dependence Plots (PDP): PDPs visualize the relationship between a feature and the
target variable while marginalizing the values of other features, providing insights into the
effect of individual features on predictions.

These advanced visualization techniques using Matplotlib can significantly enhance the
understanding of data, model behaviour, and performance in machine learning tasks,
ultimately aiding in model development, evaluation, and interpretation. Matplotlib is a
powerful library in Python that creates static, interactive, and animated visualizations.
Matplotlib offers various features and customization options for advanced data visualization.

Subplots and Grids: Subplots and grids allow creating multiple plots within the exact figure
and comparing different datasets or aspects of the same data.

Here is an example for creating multiple plots and comparing different dataset.

____________________________________________________________________________________
43 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Customizing Styles: Matplotlib allows to customize the appearance of plots extensively,


including colors, line styles, markers, labels, and more.

____________________________________________________________________________________
44 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

3D Plotting: Matplotlib can create 3D plots for visualizing data in three dimensions.

Animations: Matplotlib supports creating animations for visualizing dynamic data over time.

____________________________________________________________________________________
45 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Output for the above code:

Data Visualization - Quantile-Quantile Plots

Quantile-quantile (Q-Q) plots are a type of data visualization commonly used to assess
whether a given dataset follows a particular probability distribution or to compare the
distribution of two datasets. In a Q-Q plot, the quantiles of the observed data are plotted
against the quantiles of a theoretical distribution (e.g., normal distribution) or another dataset.

To create a Q-Q plot in PythonPython using Matplotlib and SciPy:

____________________________________________________________________________________
46 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

In the above example, we generate random data from a normal distribution using numpy.
Random. Usual ().We then use [Link]() to create the Q-Q plot. We pass the
observed data (data) and specify the distribution we want to compare against (dist= "norm"
for a normal distribution). To define the plot to use (plt) to integrate with Matplotlib. Finally,
we add labels and a title to the plot using Matplotlib's xlabel(), ylabel(), and title() functions.
The resulting plot will display points representing the quantiles of the observed data plotted
against the quantiles of the specified theoretical distribution (in this case, the normal
distribution). If the observed data closely follows the restricted distribution, the points will
fall approximately along a diagonal line. Create Q-Q plots to compare two datasets by
passing the two datasets as arguments [Link](). For the given example: #
Generate another dataset from a normal distribution

# Generate another dataset from a normal distribution

data2 = [Link](loc=0, scale=1, size=1000)

# Create Q-Q plots for each dataset

fig, ax = [Link]()

[Link](data, dist="norm", plot=ax)


____________________________________________________________________________________
47 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

[Link](data2, dist="norm", plot=ax)

[Link]('Quantiles of data')

[Link]('Quantiles')

[Link]('Q-Q Plot: Comparison of Datasets')

[Link](['Data 1', 'Data 2'])

[Link]()

Statistical Characterization
Statistical characterization refers to the process of describing and summarizing a dataset
using statistical measures and techniques. It involves analyzing the central tendency,
variability, distribution, and relationships within the data to gain insights into its properties
and behavior. Statistical characterization is essential for understanding the underlying
structure of data, identifying patterns, and making informed decisions. Some common
statistical characteristics include:
Central Tendency:
Mean: The average value of the dataset is calculated by summing all values and dividing by
the number of observations.
Median: The middle value of the dataset, sorted in ascending order, represents the median
value. It represents the central value that divides the dataset into two halves.
Mode: Most frequently occurring value(s) in the dataset.
Variability:
Range: Maximum and minimum difference of the values in the dataset, providing a measure
ofdata spread.
Variance: An average of the squared differences between each data point and the mean. It
quantifies the dispersion of data points around the mean value.
Standard Deviation: The variance measures the average deviation of data points from the
mean and is expressed in the same units as the original data distribution: A histogram is a
graphical representation of the frequency distribution of data, showing the frequency of
values within predefined intervals or bins.
Probability Density Function (PDF): PDF describes the continuous random variable taking
on a particular value. It provides insights into the shape and probability distribution of the
data.

Cumulative Distribution Function (CDF): A function that gives a probability that the
random variable takes a value less than or equal to a given value. It represents the cumulative
probabilities of the data.

____________________________________________________________________________________
48 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Relationships:
Correlation Coefficient: The correlation coefficient is the linear relationship between two
variables. It ranges from -1 to 1 to show a positive correlation. Values too low indicate a
strong negative correlation, and values close to 0 show no linear correlation.
Scatter Plot: A scatter plot graph is a graphical representation of the relations of two
variables, with a plotted on the x-axis for one variable and the other on the y-axis. It helps
visualize patterns and trends in the data. Statistical characterization techniques provide a
comprehensive overview of the dataset's properties and facilitate data exploration, hypothesis
testing, and model development in various fields such as science, engineering, finance, and
social sciences.

Parallel Coordinates Plots

In the Parallel coordinates plots, a clear pattern emerges. Flowers belonging to setosa species
have large seal widths but low seal length, petal width, and length. Flowers belonging to
versicolor species have low Sepal Widths and medium Sepal Lengths, Petal Width, and
Length. Flowers belonging to Virginia species have low to medium Sepal Widths, medium to
large Sepal Lengths, and large Petal Widths and Lengths.

Order: The features can be ordered so that only a few lines intersect, resulting in an
unreadable chart. To make a Parallel Coordinates plot using Python by highlighting one or
more lines? For example use the Olympics dataset for the year of 2020-2021 to illustrate the
use of a parallel coordinates plot. This dataset has details about the teams that have
participated – country, disciplines, athletes who have participated – country, athletes final
medals tally – country, rank, total medals, and the split across gold, silver, bronze medals.

import pandas as pd

# Load data into DataFrames

df_teams = pd.read_excel("data/[Link]")

df_athletes = pd.read_excel("data/[Link]")

df_medals = pd.read_excel("data/[Link]")

# Print information about each DataFrame

print(df_teams.info())

print(df_athletes.info())

Plot using Bar Charts

The bar charts plot each country's athletes, disciplines, ranks, and medals data. For

better readability, use only the top 20 entries, plt. figure(figsize=(20, 5))
____________________________________________________________________________________
49 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

ax = [Link](1,2,1)

ax = df[['Country','Athletes']][:40].[Link](x='Country', xlabel = '', ax=ax)

ax = [Link](1,2,2)

df[['Country','Discipline']][:40].[Link](x='Country', xlabel = '', ax=ax)

[Link](figsize=(20, 5))

ax = [Link](1,2,1)

df[['Country','Rank']][:40].[Link](x='Country', xlabel = '', ax=ax)

ax = [Link](1,2,2)

df[['Country,''Gold Medals,' 'Silver Medals,''Bronze Medals,']][:40].[Link](stacked=True,


x='Country', xlabel = '', ax=ax)

multiple plots

Parallel Coordinates Plots

Parallel Coordinates Plot using Plotly express plotly.graph_objects.Parcoords allow control to


a granular level – range of each axis, tick values, the label of the axis, etc. To define the list
of variables/axes that should be plotted. For each dimension, specify:

Range: start and end values specified as a list or tuple: values where the ticks should be
displayed on this axis tick text: text that should be shown at the ticks label: name of the axis

Values: values that should be plotted on that axis. Then, create a Parcoords, a list of
attributes for the figure.

Parallel Coordinates Plot using Plotly graph objects.

# Adjust the size to fit all the labels

fig.update_layout(width=1200, height=800,margin=dict(l=150, r=60, t=60, b=40))

● show()#Parallel Coordinates Plot using Plotly graph objects

Correlation Plots

Correlation plots, also known as correlation matrices or correlation heatmaps, are a type of
data visualization used to explore the pairwise correlations between variables in a dataset.
They are beneficial for identifying relationships between multiple variables to understand the
strength and direction of those relationships. And the correlation plots work.

____________________________________________________________________________________
50 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Correlation Calculation:
First, the correlation coefficient between each pair of variables in the dataset is calculated.
The most common correlation coefficient used is Pearson's linear relationship between two
continuous variables with measure. Other correlation coefficients, such as Spearman's rank
correlation coefficient, can be used for ordinal or non-linear relationships. Plotting: The
correlation coefficients are then visualized in a matrix format, with each cell representing the
correlation between a pair of variables.
Typically, the matrix is symmetric along the diagonal, as the correlation between two
variables, A and variable B, is the same as between variable B and variable.
● Colour Mapping: To enhance interpretability, the correlation coefficients are often
color-coded, with a gradient ranging from one color (e.g., blue) to another (e.g., red).
This gradient represents the strength and direction of the correlation, with positive
correlations shown in one color and negative correlations in another. Here, the intensity
of the color corresponds to the magnitude of the correlation coefficient.

Interpretation: By examining the correlation matrix, you can quickly identify highly
correlated variables (either positively or negatively) with each other. Strong positive
correlations (close to +1) indicate that the variables tend to increase or decrease together. In
contrast, strong negative correlations (close to -1) suggest that one variable rises as the other
decreases. A correlation coefficient close to 0 indicates little to no linear relationship between
the variables. Correlation plots are commonly used in various fields, such as statistics,
finance, biology, and machine learning, to gain insights into the relations between variables
and to guide further analysis or modeling efforts. They provide a visual summary of the
correlation structure of the data, making it easier to identify patterns and dependencies.

To create a correlation plot (correlation matrix or heatmap) using Matplotlib, follow these
steps:
● Import the necessary libraries.
● Prepare your data and calculate the correlation matrix.
● Create a heatmap to visualize the correlation matrix.
An example to show correlation plots in matplotlib:

____________________________________________________________________________________
51 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

This example generates a random correlation matrix (data) for demonstration purposes.
What each part of the code does:
[Link](data): Calculates the correlation matrix using NumPy's corrcoef function.
● imshow(): Plots the correlation matrix as a heatmap.
cmap='coolwarm': Sets the colormap to coolwarm to better visualize positive and negative
correlations.
● Color bar (): Adds a color bar to the plot to indicate the correlation values.
[Link]() and [Link](): Sets the labels for the x and y axes with appropriate tick
positions.

Adjust the figure size, colormap, and other parameters according to the preferences.
Additionally, if the column names for the dataset are different, we can replace the tick labels
with the actual variable names for better readability.

____________________________________________________________________________________
52 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Pareto Charts

using ggplot2 library ggplot2 in R for machine learning# Sample data

data <- [Link](Defect_Type = c('A,' 'B,' 'C,' 'D,' 'E'),Frequency = c(50, 30, 20, 15, 10))

# Sort the data by frequency in descending order

data_sorted <- data[order(-data$Frequency), ]

# Calculate the cumulative percentage


____________________________________________________________________________________
53 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

data_sorted$Cumulative_Percentage<-cumsum(data_sorted$Frequency)
/sum(data_sorted$Frequency) *100

# Plot Pareto chart

ggplot(data_sorted, aes(x=Defect_Type)) +geom_bar(aes(y=Frequency), stat="identity",


fill="blue") +geom_line(aes(y=Cumulative_Percentage * max(data_sorted$Frequency) /
100), color="red") +geom_point(aes(y=Cumulative_Percentage *
max(data_sorted$Frequency) / 100), color="red") +scale_y_continuous([Link] =
sec_axis(~./max(data_sorted$Frequency)*100, name = "CumulativePercentage"))
+xlab("Defect Type") +ylab("Frequency") +ggtitle("Pareto Chart")

Using Python and R, generate Pareto charts based on the provided data. In Python, using
Matplotlib for plotting, while in R, utilize the library ggplot2. The key concept is to sort the
data in descending order of frequency and then calculate the cumulative percentage.

Heatmaps

Heatmaps in Python and R are graphical representations of data where values in a matrix are
represented as colors. They are commonly used in machine learning for visualizing various
types of data, including correlation matrices, confusion matrices, and feature importance
matrices, among others. Python, Seaborn is often used for creating heatmaps due to its
simplicity and aesthetics, while in R, ggplot2 provides powerful tools for visualization,
including heatmaps.

A heatmap with some sample data:

____________________________________________________________________________________
54 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

The above example, np. [Link](10, 10) generates a 10x10 array of random numbers or
variables between 0 and 1, representing the data. [Link]() displays the
[Link]='hot' specifies the colormap;
● Color bar () adds a color bar to the plot for reference.
● Title (), plot. Label (), and plot. Label () adds a title and labels to the plot.

Advanced Data Visualization in ggplot

Advanced data visualization in ggplot refers to the use of advanced techniques and
functionalities within the ggplot2 package in R to create highly customized and intricate
visualizations. The library ggplot2 is a flexible package for creating static, interactive, and
multi-layered plots based on the grammar of graphics concepts. Some of the advanced data
visualization techniques in ggplot include:
Faceting: Faceting allows you to create small multiples of plots based on one or more
categorical variables.
Layering: ggplot2 allows you to add multiple layers to a plot, allowing for complex
visualizations. You can add layers for points, lines, bars, text, etc., and customize each layer
independently.
Geometric objects: ggplot2 provides a wide range of geometric objects (geoms) such as
points, lines, bars, polygons, and more. You can customize the appearance of these geoms
using aesthetics (aes) mappings.

Themes: Themes in ggplot2 allow you to customize the appearance of the plot elements such
as background, grid lines, text, and overall style. Create the custom themes or use pre-defined
themes provided by ggplot2.

____________________________________________________________________________________
55 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Scale customization: ggplot2 allows you to customize the scale of axes, colors, sizes, etc.
You can change the scale type (e.g., linear, logarithmic), labels, breaks, and limits.
Statistical transformations: ggplot2 provides various statistical transformations (stats) that
can be applied to the data before plotting. It includes summarizing data (e.g., mean, median),
smoothing data (e.g., loess, smoothing splines), and more.
Annotations: You can add annotations such as text, labels, arrows, and shapes to highlight
specific aspects of the plot. Using the functions like geom_text(), geom_label(), and
annotate().
Interactive visualizations: While ggplot2 primarily generates static plots, you can integrate
it with other packages like ggplotly to create interactive plots that can be explored
dynamically.

Using ggrepel Legends Movements

In Matplotlib, ggrepel-like functionality for legends, where the labels adjust their positions
automatically to avoid overlap, is not directly available. However, there is similar
functionality by changing the legend manually using the bbox_to_anchor parameter of the
plt. legend() function.

Here is an example of how you can create a scatter plot with legends that avoid overlap:

____________________________________________________________________________________
56 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

In this example, Line2D is used to create custom legend elements with markers that match
the colors of the scatter plots,bbox_to_anchor=(1, 1) places the legend outside the plot area in
the upper right corner. By adjusting the bbox_to_anchor values, the legend's position is
controlled to avoid overlapping with data points.

____________________________________________________________________________________
57 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Unit IV: Data Visualization Using Seaborn, Bokeh


● Using the Seaborn library to reconstruct all the charts done using matplotlib
● Pros and Cons of the Two Visualization Libraries
● Introduction to Data Plotting Using Bokeh
● Creating Interactive Charts with Plot
● Data Visualization Using Seaborn

Using the Seaborn library to reconstruct all the charts done using matplotlib

The Seaborn library is a Python data visualization library and provides a high-level interface
for creating attractive and informative statistical graphics using matplotlib. It complements
Matplotlib's functionalities and simplifies the process of creating complex visualizations.
Seaborn offers a variety of built-in themes and color palettes and provides functions for
creating complex statistical plots with minimal code. Seaborn provides similar functions for
various types of plots, and it's often preferred for its aesthetic appeal and ease of use in
creating complex statistical visualizations. Matplotlib package is a powerful tool for low-
level customization and flexibility in plotting the statistical graphics and visualizations in
Python.

1. High-Level interface: The Seaborn library provides the APIs to create statistical visuals.
2. Attractive Defaults: Seaborn comes with attractive default styles and color palettes,
making it easy to create visually appealing plots without much customization.
3. Statistical Plotting: Seaborn offers specialized functions for visualizing statistical
relationships, such as scatter plots, line plots, bar plots, box plots, violin plots, pair plots,
and more. These functions are designed to handle complex data structures like Pandas
DataFrames.
4. Faceted Plotting: Seaborn supports faceted plotting, allowing to easily create grid-based
layouts of plots based on one or more categorical variables.
5. With Pandas: Seaborn seamlessly integrates with Pandas DataFrames, enabling easy
manipulation and visualization of data.
6. Customization: While Seaborn's default styles are attractive, it also provides options for
customizing plots, including color palettes, plot styles, axis labels, titles, and more.
Seaborn is a powerful tool for data visualization in Python, mainly for exploratory data
analysis and for creating publication-quality graphics for presentations, reports, and
publications. It simplifies the process of making seaborn is a popular Python data
visualization library based on Matplotlib.

____________________________________________________________________________________
58 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Seaborn provides a high-level interface for drawing statistical graphics:


Here is an example for Seaborn statistical graph:

____________________________________________________________________________________
59 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

# Example 2: Line plot

____________________________________________________________________________________
60 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

# Example 3: Histogram

# Example 4: Box plot

[Link](x="species", y="petal_length", data=iris)

[Link]("Petal Length Distribution by Species")

[Link]()

# Example 5: Heatmap (correlation matrix)

correlation_matrix = [Link]()

[Link](correlation_matrix, annot=True, cmap="cool warm")

[Link]("Correlation Matrix of Iris Dataset")

[Link]()

In each of the charts using Seaborn functions. Seaborn provides tasks that are specifically
tailored for making these types of visualizations, often with improved aesthetics and default
settings.
____________________________________________________________________________________
61 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Pros and Cons of the Two Visualization Libraries

Both Seaborn and Matplotlib are powerful data visualization libraries in Python, each with its
own set of pros and cons.

Pros are:

Wide Adoption: Matplotlib is one of the oldest and most widely used plotting libraries in
Python. It's been around for a long time and has a large user base.

Customizability: Matplotlib offers a high degree of customization, allowing users to have


fine-grained control over almost every aspect of a plot.

Low-level Interface: Matplotlib provides a low-level interface, which means you can create
almost any type of plot from scratch if needed.

Integration: It integrates well with other libraries and tools in the Python ecosystem.

Cons are:
Verbose Syntax: Matplotlib's syntax can be verbose and requires more lines of code to
create complex plots, especially compared to higher-level libraries like Seaborn.

Aesthetics: While Matplotlib can produce publication-quality plots, achieving aesthetically


pleasing plots requires additional effort due to its extensive customization options and low-
level interface. Matplotlib has a steeper learning curve for beginners.

High-level Interface: Seaborn provides a high-level interface for creating attractive


statistical graphics. It simplifies the process of creating complex plots by abstracting away
many of the tedious details.

Attractive Defaults: Seaborn has attractive default styles with color palettes and makes it
easy to create visually appealing plots with minimal customization.

Statistical Plotting: Seaborn offers specialized functions for visualizing statistical


relationships, making it particularly useful for exploratory data analysis.

Integration with Pandas: Seaborn seamlessly integrates with Pandas DataFrames, allowing
for easy manipulation and visualization of data.

Limited Customization: While Seaborn's default styles are attractive, they can be less
customizable compared to Matplotlib, especially for highly specialized or unconventional
plots.

Dependency on Matplotlib: Seaborn is built on top of Matplotlib, so understanding


Matplotlib concepts can be beneficial for using Seaborn effectively.

Less Flexibility: In some cases, the high-level nature of Seaborn's interface may limit
flexibility for users who require highly customized plots.
____________________________________________________________________________________
62 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Introduction to Data Plotting Using Bokeh

Bokeh is a powerful Python library for creating interactive and visually appealing data
visualizations for the web. It allows creating interactive plots, dashboards, and applications
directly in Python, Bokeh's key features include:

Interactive Visualization: Bokeh provides tools for creating interactive plots with zooming,
panning, and hovering capabilities, enabling users to explore data interactively.

Flexible and Expressive: Bokeh offers a high-level and flexible API for creating a wide
range of plots, including line plots, scatter plots, bar charts, heat maps, and more. It also
supports complex visualizations like linked plots and interactive dashboards.

Web-Based: Bokeh generates plots as HTML and JavaScript, allowing them to be easily
embedded into web applications or shared as standalone HTML files without the need for a
server backend.

Integration with Jupyter Notebooks: Bokeh can be seamlessly integrated into Jupyter
Notebooks, enabling interactive data exploration and visualization within the notebook
environment.

High-Performance Rendering: Bokeh leverages modern web browser technologies to


render plots efficiently, making it suitable for handling large datasets and streaming data.

Here is a basic introduction to plotting data using Bokeh:


Installation:
To install Bokeh using pip command
Pip install Bokeh

Creating Interactive Charts with Plotly

Plotly is a versatile Python library for creating interactive and visually appealing data
visualizations and a wide range of chart types, including line plots, scatter plots, bar charts,
heat maps, and more. Plotly's key features include:
Interactive Visualization: Plotly generates interactive plots with zooming, panning, and
hovering capabilities, enabling users to explore data interactively.
Web-Based: Plotly generates plots as HTML and JavaScript, making them easily
embeddable into web applications or standalone HTML documents without the need for a
server backend.
High-Quality Graphics: Plotly produces high-quality, publication-ready graphics suitable
for presentations, reports, and publications.
Ease of Use: Plotly offers a simple and intuitive API for creating a wide range of plots with
minimal code.
Integration: Plotly integrates seamlessly with other Python libraries, such as Pandas,
NumPy, and Matplotlib, allowing for easy data manipulation and visualization.

____________________________________________________________________________________
63 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Unit V: Dimension Reduction Methods


● Dimensionality Reduction Techniques
● Principal Component Analysis
● Linear Discriminant Analysis
● Multiple Discriminant Analysis
● Markov Chain Monte Carlo (MCMC) Methods
● Hidden Markov Models
● Laplace Approximation and BIC

Dimensionality Reduction Techniques

Dimensionality reduction techniques are methods used in data analysis and machine learning
to reduce the number of input variables, and "user" typically refers to an individual
interacting with a system or a platform. It is meant to Various Dimensionality Reduction
Techniques. Dimensionality decrease is a method utilized in AI and measurements to lessen
the number of information factors in a dataset while saving its fundamental elements. It is
beneficial when dealing with high-dimensional data, as it can help improve computational
efficiency, reduce noise, and prevent overfitting. Here are some standard dimensionality
reduction methods:
1. Principal Component Analysis (PCA): PCA is one of the most widely used
dimensionality reduction techniques. It transforms the actual features into a new set of
uncorrelated variables called principal components. These components are ordered by
the amount of variance.

2. t-Distributed Stochastic Neighbor Embedding (t-SNE): t-SNE is a non-straight


dimensionality decrease method that is wildly successful for envisioning high-layered
information in a few aspects. It focuses on preserving local relationships between data
points, making it helpful in exploring clusters or groups in the data.

3. Uniform Manifold Approximation and Projection (UMAP): UMAP is another


non-linear dimensionality reduction method similar to t-SNE but is often faster and
has some advantages in preserving both local and global structures in the data. It has
gained popularity for visualizing and exploring high-dimensional datasets.

4. Linear Discriminant Analysis (LDA): LDA is a supervised dimensionality


reduction technique that maximizes the separation between classes in the data. It is
commonly used in classification problems where the goal is to find a projection that
maximizes the distance between class means while minimizing the spread within each
class.

____________________________________________________________________________________
64 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

5. Autoencoders: They are a neural network engineering utilized for unsupervised


learning. They comprise an encoder and a decoder, and the center layer (inactive
space) fills in as a packed portrayal of the information.

6. Factor Analysis: Factor analysis models the observed variables as linear


combinations of underlying factors and error terms. It helps identify latent variables
that explain the observed correlations in the data. Factor analysis can be considered a
probabilistic version of PCA.

7. Random Projection: Random projection is an easy and computationally efficient


technique that uses random matrices to project high-dimensional data onto a lower-
dimensional subspace. It may not capture complex relationships in the data. The
choice of dimensionality reduction method depends on the characteristics of data.

8. Principal Component Analysis (PCA) using Python and R: PCA is a


dimensionality scaling down technique widely used in data analysis and machine
learning. It transforms high-dimensional data into a new coordinate system, capturing
the most significant information in the first few principal components. These Parts are
linear combinations of the actual features and are chosen to maximize the variance in
the data. By selecting a subset of these components, PCA enables a reduction in the
dimensionality of the dataset, simplifying computation and aiding in visualization.

9. The first principal module corresponds to the direction of maximum variance,


followed by subsequent components in decreasing order. PCA is beneficial for noise
reduction, feature extraction, and revealing underlying patterns in complex datasets.
Implementation in python, often using libraries like scikit-learn, involves fitting the
PCA model to the data and transforming it into a lower-dimensional representation.
Principal Component Analysis is a widely used dimensionality reduction technique in
R. It transforms the actual features of a dataset into a new set of uncorrelated
variables, called principal components. The goal is to capture the maximum variance
in the data, providing a lower-dimensional representation while retaining essential
information.

Here is an example of how to perform PCA in R:


# Load a sample dataset (using the iris dataset as an example)
data(iris)
# Extract the features
X <- iris[, 1:4]
# Perform PCA
pca_result <- prcomp(X, scale. = TRUE)
# Summary of PCA results
____________________________________________________________________________________
65 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

summary(pca_result)
# Accessing principal components
pcs <- pca_result$x

The prcomp function is then used to perform PCA. The scale. = TRUE argument scales the
variables to have unit variance before [Link] the plot function, visualize the variance
explained and the loadings of each variable on each principal component. PCA captures the
main patterns of the data while reducing it to a lower-dimensional space. The central head
part addresses the course of the most significant fluctuation, and the second head part is
symmetrical to the first and addresses the following most elevated change.
It is important to note that, in practice, PCA is often used as a preprocessing step before other
analyses or modeling tasks, such as clustering or regression, to reduce the data's
dimensionality while preserving its essential characteristics.

Linear Discriminant Analysis


LDA (Linear Discriminant Analysis) is a supervised dimensionality reduction technique
primarily used for classification tasks. It seeks to identify a linear combination of features
that maximizes the dataset's separation of multiple classes. By projecting the data onto this
discriminant subspace, LDA maximizes the distance between class means while minimizing
each class's spread (variance). This results in a lower-dimensional representation of the data
that retains the most discriminative information for classification. LDA assumes that the
features are typically distributed and that the classes have a standard covariance matrix. It is
particularly effective when the classes are well-separated, and the assumptions are met. In
practical terms, LDA is employed in scenarios where identifying the most relevant features
for class discrimination is crucial, making it a valuable tool in pattern recognition and
machine learning applications.

Example:
LDA (Linear Discriminant Analysis) is a supervised dimensionality reduction and
classification technique. It finds the linear combinations of features that best separate
multiple classes in a dataset. LDA aims to maximize the distance between the means of
different classes while minimizing the spread (variance) within each class. It is commonly
used for pattern recognition and classification tasks.
Here is a brief explanation along with a simple example in python using the `scikit-learn`
library:

____________________________________________________________________________________
66 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Linear Discriminant Analysis Example:

In the example, we use the famous Iris dataset. For visualization purposes, LDA is applied to
reduce the data to two dimensions (`n_components=2`). The resulting plot shows how LDA
transforms the original data into a lower-dimensional space while maximizing the separation
between the three classes of iris flowers. LDA is used for visualization and as a preprocessing
step in classification tasks, where the reduced feature space often leads to improved
classification performance.
____________________________________________________________________________________
67 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Multiple Discriminant Analysis


Multiple Discriminant Analysis (MDA) is a dimensionality reduction and classification
statistical technique. It is closely connected to linear discriminant analysis (LDA) but is
extended to handle multiple classes. In Python and R, we can perform Multiple Discriminant
Analyses using various libraries. One popular library is scikit-learn, which provides a
LinearDiscriminantAnalysis class that can handle multiple classes.
Here is a simple example of how to use Multiple Discriminant Analysis in Python with scikit-
learn:

In the example, the LinearDiscriminantAnalysis class is used to fit the model on the training
data and make predictions on the test data. The model's accuracy is then calculated using the
accuracy_score function from scikit-learn.
Another Example of R:

# Load necessary library (if not already installed)


# [Link]("MASS")
library(MASS)
# Load the Iris dataset as an example
data(iris)
# Split the dataset into training and testing sets
[Link](123)
indices <- sample(1:nrow(iris), nrow(iris)*0.8)
train_data <- iris[indices, ]

____________________________________________________________________________________
68 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

test_data <- iris[-indices, ]


# Fit LDA model
lda_model <- lda(Species ~ ., data = train_data)
# Make predictions on the test data
lda_predictions <- predict(lda_model, test_data)
# Confusion matrix
conf_matrix <- table(lda_predictions$class, test_data$Species)
print(conf_matrix)
# Calculate accuracy
accuracy <- sum(diag(conf_matrix)) / sum(conf_matrix)
print(paste("Accuracy:", accuracy))

In the example, the lda function is used to fit the model to the training data and make
predictions on the test data. The confusion matrix is then calculated, and accuracy is
computed based on its diagonal elements, the Python example; you might need to perform
additional steps such as data preprocessing, parameter tuning, and cross-validation for a more
comprehensive analysis in a real-world scenario.

Markov Chain Monte Carlo (MCMC) Methods

Markov Chain Monte Carlo (MCMC) methods are a class of algorithms used for sampling
from complex probability distributions, especially when direct sampling is difficult. MCMC
is widely used in Bayesian statistics and other fields. In python, the most commonly used
library for MCMC is PyMC3. PyMC3 is a probabilistic programming library that allows
users to define probabilistic models using a high-level syntax and then perform Bayesian
inference using MCMC methods.
Here is a basic example of using PyMC3 to perform MCMC in Python:

import pymc3 as pm
import 83atpl as np
import [Link] as plt

# Generate synthetic data


[Link](42)
true_slope = 2
true_intercept = 1

____________________________________________________________________________________
69 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

data_x = [Link](0, 10, 100)


data_y = true_slope * data_x + true_intercept + [Link](0, 1, 100)

# Define a probabilistic model


with [Link]() as linear_model:
# Priors
slope = [Link](‘slope’, mu=0, sd=10)
intercept = [Link](‘intercept’, mu=0, sd=10)
# Likelihood
likelihood = [Link](‘y’, mu=slope * data_x + intercept, sd=1, observed=data_y)
# Use MCMC to sample from the posterior distribution
trace = [Link](1000, tune=1000)
# Plot the posterior distribution
pm.plot_posterior(trace, var_names=[‘slope’, ‘intercept’])
[Link]()

In this example, define a simple linear regression model using PyMC3 then specify the slope
and intercept priors and use a normal distribution as the likelihood. The [Link] function
runs the MCMC algorithm to generate samples from the posterior distribution.

Here is a basic example using rstan to perform MCMC in R:


Example:
# Install and load necessary library (if not already installed)
# [Link]("rstan")
library(rstan)
# Generate synthetic data
[Link](123)
true_slope <- 2
true_intercept <- 1
data_x <- runif(100, 0, 10)
data_y <- true_slope * data_x + true_intercept + rnorm(100)
# Define a probabilistic model using Stan

____________________________________________________________________________________
70 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

stan_code <- "data {


int<lower=0> N; // number of data points
real x[N]; // input data
real y[N]; // output data
}
parameters {
real slope; // slope parameter
real intercept; // intercept parameter
real<lower=0> sigma; // standard deviation of the noise
}
model {
y ~ normal(slope * x + intercept, sigma); // likelihood
}"
# Prepare data for Stan
stan_data <- list(N = length(data_x), x = data_x, y = data_y)
# Compile the Stan model
stan_model <- stan_model(model_code = stan_code)
# Run MCMC sampling
stan_fit <- sampling(stan_model, data = stan_data, iter = 1000, chains = 4)
# Print summary of the MCMC results
print(stan_fit)

In this example, a simple linear regression model using the Stan modeling language within
the rstan package. The sampling algorithm runs the MCMC function to obtain posterior
distribution samples. The stan_fit object contains the MCMC samples, and you can use
various functions to analyze and visualize the results of the MCMC parameters, perform
model checking, and ensure convergence. PyMC3 provides tools for diagnosing convergence
and assessing the quality of the samples generated by the MCMC algorithm.

Hidden Markov Models


Hidden Markov Models (HMMs) are statistical models that describe systems that evolve and
are characterized by unobservable (hidden) states. These models assume that the system's
states are not directly observable but produce observable outcomes or emissions. HMMs have

____________________________________________________________________________________
71 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

applications in many fields, including speech recognition, bioinformatics, finance, and natural
language processing. In python, the hmmlearn library is commonly used for working with
Hidden Markov Models. Here is a simple example of how to use hmmlearn to create and
train an HMM:
Here is a basic example in Python:

from hmmlearn import hmm


import numpy as np
# Generate synthetic data
[Link](42)
observed_states = [Link]([0, 1, 2, 1, 0, 2, 1, 0, 0, 1]).reshape(-1, 1)
# Create a Hidden Markov Model
model = [Link](n_components=3, n_iter=100)
# Fit the model to the data
[Link](observed_states)
# Predict the hidden states for the observed data
hidden_states = [Link](observed_states)
# Print the model parameters
print("Transition matrix:")
print(model.transmat_)
print("\nEmission probabilities:")
print(model.emissionprob_)

In this example, create a simple HMM using the MultinomialHMM class from hmmlearn.
We generate synthetic data, which represents the observed states over time. The HMM is then
fitted to the data, and the predict method is used to predict the hidden states based on the
observed [Link] hmmlearn library also provides functionalities for generating samples
from an HMM, evaluating likelihoods, and dealing with continuous observations.
To install: pip install hmmlearn
Additionally, explore other Python libraries for hidden Markov models such as pyhsmm or
pomegranate based on your specific needs.

____________________________________________________________________________________
72 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Laplace Approximation and Bayesian Information Criterion (BIC)

The Laplace approximation is a method used for approximating complex probability


distributions with simpler ones, particularly in Bayesian statistics. It provides a Gaussian
(normal) approximation to the posterior distribution, which can be easier to work with than
the actual distribution, especially when the posterior is unimodal and has a peak. This
technique is often used when an analytical solution to a Bayesian problem is challenging or
impossible.
In python, the [Link].multivariate_normal class and other numerical libraries can be used
to implement the Laplace approximation.
Here is a simplified example:

import numpy as np from [Link] import multivariate_normal


# Assume we have some mean and covariance matrix from Bayesian analysis
mean_vector = [Link]([1, 2])
covariance_matrix = [Link]([[2, 0.5], [0.5, 1]])
# Create a multivariate normal distribution using the Laplace approximation
laplace_approximation = multivariate_normal(mean=mean_vector, cov=covariance_matrix)
# Now you can use the laplace_approximation object for further calculations or sampling

The Laplace approximation assumes that a Gaussian near its peak can approximate the
posterior distribution well. In R, the laplace() function in the stats package can be used to
perform the Laplace approximation. However, in more complex scenarios, you may need to
implement the approximation manually using numerical optimization techniques.

Here is a simplified example of using the laplace() function:


# Load necessary library (if not already installed)
# [Link]("stats")
library(stats)
# Define the log-likelihood function
log_likelihood <- function(x) {
# Define your log-likelihood function here
# Example: -0.5 * (x - mu)^2 / sigma^2 - log(sigma) - 0.5 * log(2 * pi)
}
# Perform Laplace approximation
laplace_result <- laplace(log_likelihood, x0 = initial_guess)

____________________________________________________________________________________
73 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

# Print the result


print(laplace_result)

Replace log_likelihood() with your specific log-likelihood function. The laplace() function
performs the Laplace approximation based on the provided log-likelihood function and an
initial [Link] Bayesian Information Criterion (BIC) is a model selection criterion that
balances the fit of a model to the data with the complexity of the model. When choosing
between different models in statistical modeling and machine learning, it is often used. The
BIC is calculated using the formula:
B I C = −2 log(L) + klog(n)
Where:
● L is the maximized likelihood of the model,
● k is the number of parameters in the model, and
● n is the number of data points.
The BIC can be implemented in Python using libraries like scikit-learn or stats models. Here
is a simplified example:

Example:
from [Link] import GaussianMixture
from [Link] import make_blobs
# Generate synthetic data
X, _ = make_blobs(n_samples=100, centers=3, random_state=42)
# Fit a Gaussian Mixture Model with different numbers of components
for n_components in range(1, 6):
gmm = GaussianMixture(n_components=n_components)
[Link](X)
# Calculate BIC
bic = [Link](X)
print(f"Number of components: {n_components}, BIC: {bic}")

A Gaussian Mixture Model (GMM) is fitted to the synthetic data, and the BIC is calculated
for models with different numbers of components. The idea is to choose the model with the
lowest BIC value, indicating a good balance between model fit and complexity.
In R, you can calculate the BIC using functions from various packages such as stats, MASS,
or specific modeling packages like glm for generalized linear models.

____________________________________________________________________________________
74 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Here is a simplified example of using the BIC with the MASS package using R:
Example
# Load necessary library (if not already installed)
# [Link]("MASS")
library(MASS)
# Fit a model (e.g., linear regression)
model <- lm(y ~ x, data = my_data)
# Calculate BIC
bic <- BIC(model)
# Print the BIC
print(bic)

Replace lm(y ~ x, data = my_data) with the appropriate model fitting function and dataset.
The BIC() function calculates the BIC value for the specified model, which includes the
maximized likelihood and the number of parameters in the model. Lower BIC values indicate
better model fit while penalizing for model complexity.

Unit VI: Model Selection and Evaluation


● Model Evaluation

____________________________________________________________________________________
75 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

● Model Selection
● Model Boosting - Ensemble Method
● Gradient Boosting
● Xgboost
● Adaptive Boosting
● Model Parameters
● Hyperparameter Tuning

Model Evaluation

Model evaluation is the process of evaluating the performance of a prepared model on unseen
data to determine how well it generalizes to new, unseen samples. Model evaluation is
essential for understanding the model's effectiveness and gaining insights into its behavior.
Model evaluation typically involves several key steps and techniques:
1. Holdout Method: The dataset is divided into two subsets- training and testing sets. The
model is prepared on the training set and then evaluated on the separate testing set to
assess its performance on unseen data.
2. Cross-Validation: Instead of a single train-test split, the dataset is divided into multiple
subsets (folds). The model is prepared on k-1 folds and evaluated on the remaining fold;
the process of the test set is repeated k times (every time using a different fold as the test
set). Cross-validation gives a more reliable estimate of the model's performance than the
holdout method, especially with smaller datasets.
3. Performance Metrics: Various metrics are used to quantify the model's performance,
depending on the type of task (classification, regression, clustering, etc.). For
classification tasks, standard metrics include accuracy, precision, recall, F1-score, ROC
AUC (Receiver Operating Characteristic Area under the Curve), etc. For regression
tasks, metrics include Mean Squared Error (MSE), mean absolute error (MAE), R-
squared, etc.
4. Confusion Matrix: For classification tasks, a confusion matrix is often used to visualize
the model's performance in terms of true +ve, true -ve, false +ve, and false -ve. It
provides insights into the model's ability to classify different classes correctly.
5. ROC Curve and Precision-Recall Curve: ROC (Receiver Operating Characteristic)
curves and Precision-Recall curves are graphical representations that are used to assess
binary classification models. They visualize the trade-off between actual positive and
false favorable rates (ROC curves) or precision and recall (Precision-Recall curves) at
different threshold values.
6. Learning Curves: Learning curves plot the model's performance (e.g., training and
validation error) as a function of the training dataset size. They help diagnose issues like
overfitting (high variance) or underfitting (high bias) and provide insights into whether
collecting more data would be beneficial.
In Python, libraries such as scikit-learn (sklearn), TensorFlow, and PyTorch, and in R
libraries such as caret, MLmetrics, and pROC provide built-in functions and utilities for
____________________________________________________________________________________
76 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

model evaluation. These libraries offer a wide range of performance metrics, evaluation
techniques, and visualization tools to assess the quality and efficacy of machine learning
models. Python's rich ecosystem of scientific computing libraries allows custom evaluation
techniques and metrics to be implemented. Additionally, custom evaluation techniques and
metrics can be implemented using R's rich ecosystem of statistical and machine-learning
packages.

Model Selection

Model selection in Python refers to choosing the best-performing model among a set of
candidate models or algorithms for a given task. The model selection goal is to identify the
model that achieves the best balance between model complexity and predictive performance.
Here's an example of model selection using Python and scikit-learn:

from [Link] import load_iris


from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from [Link] import DecisionTreeClassifier
#from [Link]#
#import RandomForestClassifier#
#from [Link]#
#import accuracyscore
## Load the iris dataset here
iris = load_iris()#
x, y = [Link], [Link]
## split the dataset in the testing and training dataset
#xtrain,xtest, ytrain, ytest = traintestsplit(x, y, test_size=0.3, random_state=43)#
# Define a list of candidate models
models = [LogisticRegression(),
DecisionTreeClassifier(),
RandomForestClassifier()]
# Train and evaluate each model
best_model = None
best_accuracy = 0
For a model in models:
____________________________________________________________________________________
77 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

[Link](X_train, y_train)
y_pred = [Link](X_test)
accuracy = accuracy_score(y_test, y_pred)
print(f"{model.__class__.__name__} Accuracy: {accuracy}")
# Update the best model if the current model performs better
if accuracy > best_accuracy:
best_accuracy = accuracy
best_model = model
# Print the best model
print("\nBest Model:")
print(best_model)

In this example:
1. We load the Iris dataset and divide it into training and testing sets.
2. The list of candidate models includes the Logistic Regression, Decision Tree, and
Random Forest classifiers.
3. To train and evaluate each model on the training and testing data, we print the
accuracy of each model.
4. Finally, we select the model with the highest accuracy as the best model.
Here is an example of model selection using R:
# Load necessary libraries (if not already installed)
# [Link]("caret")
# [Link]("e1071")
library(caret)
library(e1071)
# Load the Iris dataset
data(iris)
####[Link](123) # for reproducibility
trainIndex <- createDataPartition(iris$Species, p = 0.8,
list = FALSE,
times = 1)
data_train <- iris[trainIndex, ]

____________________________________________________________________________________
78 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

data_test <- iris[-trainIndex, ]


# Define a list of candidate models

models <- list(


"Logistic Regression" = train(Species ~ ., data = data_train, method = "glm"),
"Decision Tree" = train(Species ~ ., data = data_train, method = "rpart"),
"Random Forest" = train(Species ~ ., data = data_train, method = "rf")
)# Evaluate each model
results <- lapply(models, function(model) {
predictions <- predict(model, newdata = data_test)
accuracy <- confusionMatrix(predictions, data_test$Species)$overall["Accuracy"]
return(accuracy)
})
# Print the results
print(results)
# Identify the best model
best_model <- names([Link](unlist(results)))
print(paste("Best Model:", best_model))

1. Here is an example for loading the iris dataset

2. —split the dataset into training and testing sets.

3. We define a list of candidate models, including Logistic Regression, Decision Tree,


and Random Forest.

4. Utilizing the train function from the caret package, we train each model using the
training data.

5. We evaluate each model on the testing data and compute the accuracy of each model.

6. Finally, we identify the best-performing model based on the highest accuracy.

This example demonstrates a fundamental approach to model selection in Python and R.


Depending on the specific task and dataset, may need to consider additional factors such as
hyperparameter tuning, cross-validation, and more sophisticated model selection techniques.

____________________________________________________________________________________
79 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Model Boosting – Ensemble Method

In machine learning, model boosting refers to a meta-algorithm combining multiple weak


learners to create a strong learner. The boosting algorithm iteratively trains a sequence of
weak models, with each subsequent model focusing on the errors made by the previous ones,
and by continuously adjusting the weights of misclassified data points, boosting aims to
improve the overall performance of the ensemble same data.
One of the most popular boosting algorithms is AdaBoost (Adaptive Boosting), which
assigns higher weights to incorrectly classified data points and lower weights to correctly
classified ones. AdaBoost sequentially fits a series of weak classifiers to the data, and each
subsequent weak learner gives more weight to the previously misclassified observations, thus
focusing on the problematic instances. Python's scikit-learn library implements AdaBoost
through the AdaBoostClassifier and AdaBoostRegressor classes for classification and
regression tasks.
Here is an example of using AdaBoost for classification in Python:

from [Link] import load_iris


from sklearn.model_selection import train_test_split
from [Link] import AdaBoostClassifier
from [Link] import accuracy_score#
# Load the Iris dataset#
iris = load_iris()#
X, y = [Link], [Link]#
# Split the dataset into training and testing sets#
X_train, X_test, y_train, y_test = train_test_split(X, y,test_size=0.2, random_state=42)
# Create an AdaBoost classifier
clf = AdaBoostClassifier(n_estimators=50, random_state=42)
# Train the classifier
[Link](X_train, y_train)
# Make predictions on the test set
y_pred = [Link](X_test)
# Calculate accuracy
accuracy = accuracy_score(y_test, y_pred)
print("Accuracy:", accuracy)

____________________________________________________________________________________
80 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

In this example,
● Load the iris dataset and divide it into training and testing sets.
● To create an AdaBoost classifier with 50 weak estimators.
● To train the classifier on the training data.
● To make predictions on the test set and calculate the model's accuracy.
In R, the adabag package implements AdaBoost for classification tasks. The gbm
(Generalized Boosted Regression Models) package also offers boosting algorithms for
regression tasks.
Here's a simple example of using AdaBoost for classification in R:

# Install and load necessary packages (if not already installed)


# [Link]("adabag")
library(adabag)
# Load the Iris dataset
data(iris)
####[Link](42)
train_index <- sample(1:nrow(iris), 0.8*nrow(iris))
train_data <- iris[train_index, ]
test_data <- iris[-train_index, ]
# Train an AdaBoost classifier
model <- boosting(Species ~ ., data = train_data, boos = TRUE, final = 50)
# Make predictions on the test set
predictions <- predict(model, newdata = test_data)
# Calculate accuracy
accuracy <- mean(predictions$class == test_data$Species)
print(paste("Accuracy:", accuracy))

In this example,
● To load the Iris dataset and split it into training and testing sets.

____________________________________________________________________________________
81 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

● To train an AdaBoost classifier using the boosting function from the adabag
package.
● To make predictions on the test set using the trained model.
● To calculate the model's accuracy by comparing the predicted classes to the true
classes in the test set.
Boosting algorithms like AdaBoost are powerful techniques for improving the performance
of weak learners and are widely used in practice for various classification and regression
tasks.
Machine learning, an ensemble method, combines the predictions of multiple individual
models (often called base learners or weak learners) to produce a final prediction. The idea
behind ensemble methods is to leverage the individual models' diversity to improve the
prediction's overall performance and robustness.
There are several types of ensemble methods, including:
1. Voting: Each base learner independently predicts this method; final prediction for the
majority of the votes (for classification of the tasks) or averaging (for regression
tasks) of the individual predictions.
2. Bagging (Bootstrap Aggregating): Bagging involves training multiple base learners
independently on different bootstrap samples (random subsets with replacement) of
the training data. The final prediction is typically obtained by averaging the
predictions of all base learners.
3. Boosting: Boosting algorithms sequentially train a series of base learners, with each
subsequent learner focusing on the errors made by the previous ones. The final
prediction is usually a weighted combination of the projections of all base learners.
4. Stacking: Stacking, also known as stacked generalization, combines the predictions
of multiple base learners using a meta-learner (often a simple linear model). The base
learners' predictions serve as features for training the meta-learner.

Here is an example of using the VotingClassifier ensemble method from scikit-learn in


Python:

from [Link] import load_iris


from sklearn.model_selection import train_test_split
from [Link] import VotingClassifier
from [Link] import DecisionTreeClassifier
from [Link] import KNeighborsClassifier
from sklearn.linear_model import LogisticRegression
from [Link] import accuracy_score#
# Load the iris dataset#
iris = load_iris()#
____________________________________________________________________________________
82 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

x, y = [Link], [Link]
## split the datasets training and testing sets here #
x_train, x_test, y_train, y_test = train_test_split(x, y, test_size=0.2, random_state=42)#
# Define base classifiers
tree_clf = DecisionTreeClassifier()
knn_clf = KNeighborsClassifier()
log_reg_clf = LogisticRegression()
# Define a voting classifier combining the base classifiers
voting_clf = VotingClassifier(
estimators=[('tree', tree_clf), ('know, knn_clf), ('log_reg', log_reg_clf)],
voting='hard' # Use majority voting
)# Train the voting classifier
voting_clf.fit(X_train, y_train)
# Make predictions on the test set
y_pred = voting_clf.predict(X_test)
# Calculate accuracy
accuracy = accuracy_score(y_test, y_pred)
print("Accuracy:", accuracy)

In this example,
● To load the Iris dataset and split it into training and testing sets.
● To define three base classifiers: a decision tree classifier, a k-nearest neighbors classifier,
and a logistic regression classifier.
● To create a VotingClassifier ensemble by specifying the base classifiers and the voting
strategy (in this case, majority voting).
● To train the ensemble or same classifier on the training data for making the predictions
on the test set.
● To calculate the accuracy of the ensemble classifier on the test set.

Here is an example of using the caret package in R to implement ensemble methods:

# Install and load necessary packages (if not already installed)


# [Link]("caret")
library(caret)

____________________________________________________________________________________
83 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

# Load the Iris dataset


data(iris)
# Define training control with 5-fold cross-validation
train_control <- trainControl(method="cv", number=5, verboseIter=FALSE)
# Define base models
base_models <- list(
"glm" = glm(Species ~ ., data = iris),
"rpart" = rpart::rpart(Species ~ ., data = iris),
"svmRadial" = train(Species ~ ., data = iris, method = "svmRadial")
)
# Create an ensemble model using stacking
ensemble_model <- caretEnsemble::caretStack(models = base_models, metric =
"Accuracy")
# Train the ensemble model
ensemble_fit <- train(Species ~ ., data = iris, method = ensemble_model, trControl =
train_control)
# Summarize the ensemble model
print(ensemble_fit)

In this example,
 To load the Iris dataset.
 To define three base models: logistic regression (glm), decision tree (rpart), and support
vector machine with radial kernel (svmRadial).
 To use the caretStack function from the caretEnsemble package to create an ensemble
model using stacking.
 To train the ensemble model on the training data using 5-fold cross-validation.
 Finally, print the summary of the ensemble model, which includes information about the
base models and their performance.

Gradient Boosting

Gradient Boosting is a popular ensemble learning technique that builds a robust predictive
model by sequentially adding weak learners (usually decision trees) and fitting them to
the residuals of the previous models. This sequential approach allows gradient Boosting
to continuously increase the model's performance by focusing on the mistakes made by
the last weak learners. In Python, the most widely used implementation of Gradient
Boosting is provided by the GradientBoostingClassifier and
GradientBoostingRegressor classes in the scikit-learn library.
____________________________________________________________________________________
84 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Here is an example of using Gradient Boosting for classification in Python:

from [Link] import load_iris


from sklearn.model_selection import train_test_split
from [Link] import GradientBoostingClassifier
from [Link] import accuracy_score
# Load the Iris dataset
iris = load_iris()
X, y = [Link], [Link]
# Split the dataset into training and testing sets
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
# Create a Gradient Boosting classifier
clf = GradientBoostingClassifier(n_estimators=100, learning_rate=0.1, random_state=42)
# Train the classifier
[Link](X_train, y_train)
# Make predictions on the test set
y_pred = [Link](X_test)
# Calculate accuracy
accuracy = accuracy_score(y_test, y_pred)
print("Accuracy:", accuracy)

In this example,
● To load the Iris dataset and split it into training and testing sets.
● To create a GradientBoostingClassifier with 100 weak learners (decision trees) and
a learning rate of 0.1.
● To train the Gradient Boosting classifier on the training data.
● To make predictions on the test set using the trained classifier.
● To calculate the accuracy of the model on the test set.
Similarly, one can use GradientBoostingRegressor for regression tasks by importing it
from [Link] and applying it to regression datasets.

However, it is essential to be mindful of hyperparameters such as the number of


estimators, learning rate, and depth of trees, as these can significantly impact the model's
performance and computational efficiency.

____________________________________________________________________________________
85 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Here is an example of using gradient boosting for classification in R:

# Install and load necessary packages (if not already installed)


# [Link]("gbm")
library(gbm)
library(caret)
# Load the iris dataset
data(iris)###
[Link](42) # for reproducibility
train_index <- createDataPartition(iris$Species, p = 0.8, list = FALSE)
train_data <- iris[train_index, ]
test_data <- iris[-train_index, ]
# Train a gradient boosting classifier
gbm_model <- gbm(Species ~ ., data = train_data, [Link] = 100, [Link] = 3,
shrinkage = 0.1, distribution = "multinomial")
# Make predictions on the test set
predictions <- predict(gbm_model, newdata = test_data, [Link] = 100, type = "response")
# Convert predictions to class labels
predicted_classes <- apply(predictions, 1, [Link])
# Calculate accuracy
accuracy <- mean(predicted_classes == test_data$Species)
print(paste("Accuracy:", accuracy))

In this example:
- To load the datasets for iris.
- To split the training and testing datasets with the `createDataPartition` function from the
`caret` package.
- To train a gradient boosting classifier using the `gbm` function from the `gbm` package.
We specify the number of trees (`[Link]`), the most depth of each tree
(`[Link]`), the shrinkage parameter (`shrinkage`), and the distribution for
classification (`distribution = "multino mial"`). - To predict the test set using the trained
gradient boosting model.
- To convert the predicted probabilities to class labels.
- Finally, calculate the model's accuracy on the test set.

____________________________________________________________________________________
86 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Gradient boosting is a powerful technique for classification and regression tasks, and the
gbm package in R provides a flexible and efficient implementation of this algorithm.

XGBboost

XGBoost, short for eXtreme Gradient Boosting, implements the gradient boosting algorithm.
It is widely used for supervised learning tasks, including classification, regression, and
ranking problems. XGBoost is known for its speed and performance and has won numerous
machine-learning competitions.
Key features of XGBoost include:
1. Regularization: XGBoost includes built-in support for L1 (Lasso) and L2 (Ridge)
regularization to prevent overfitting.
2. Parallelization: XGBoost is highly optimized for parallel computation, enabling
efficient training on large datasets.
3. Tree Pruning: XGBoost uses tree pruning to remove splits that provide little or no
additional information, leading to more efficient and accurate trees.
4. Customizable Objective Functions: XGBoost allows users to define custom loss
functions and evaluation criteria, making it flexible for various tasks.
5. Handling Missing Values: XGBoost can automatically handle the missing values
dataset during the training and prediction data set.

In Python, XGBoost is implemented through the `xgboost` library, which gives an easy-to-
use interface for building and training gradient-boosting models.
To install the library using pip:
pip install xgboost
Here is a simple example of using XGBoost for classification in Python:

import xgboost as xgb


from [Link] import load_iris
from sklearn.model_selection import train_test_split
from [Link] import accuracy_score
# Load the Iris dataset
iris = load_iris()
x, y = [Link], [Link]
# Split the dataset into training and testing sets
x_train, x_test, y_train, y_test = train_test_split(x, y, test_size=0.2, random_state=42)
# Convert the data into DMatrix format
dtrain = [Link](x_train, label=y_train)
____________________________________________________________________________________
87 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

dtest = [Link](x_test, label=y_test)


# Define hyperparameters
params = {
'objective': 'multi:softmax,' # Multiclass classification
'num_class': 3, # Number of classes
'max_depth': 3, # Maximum depth of each tree
'eta': 0.3, # Learning rate
'eval_metric': 'merror' } # Evaluation metric}
# Train the XGBoost model
num_rounds = 10
model = [Link](params, dtrain, num_rounds)
# Make predictions on the test set
y_pred = [Link](dtest)
# Calculate accuracy
accuracy = accuracy_score(y_test, y_pred)
print("Accuracy:", accuracy)

In R, XGBoost is implemented through the `xgboost` package, which provides functions


for building, training, and evaluating XGBoost models. Firstly, install the package from
CRAN using:
Firstly, install the package.
[Link]("xgboost")
Here is a simple example of using XGBoost for classification in R:

library(xgboost)
library(caret)
# Load the Iris dataset
data(iris)
[Link](42) # for reproducibility
train_index <- createDataPartition(iris$Species, p = 0.8, list = FALSE)
train_data <- iris[train_index, ]
test_data <- iris[-train_index, ]

____________________________________________________________________________________
88 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

# Convert the data into DMatrix format


dtrain <- [Link](data = [Link](train_data[, -5]), label = train_data$Species)

dtest <- [Link](data = [Link](test_data[, -5]), label = test_data$Species)


# Define hyperparameters
params <- list(
objective = "multi:softmax", # Multiclass classification
num_class = 3, # Number of classes
max_depth = 3, # Maximum depth of each tree
eta = 0.3 # Learning rate
)
# Train the XGBoost model
model <- xgboost(params = params, data = dtrain, nrounds = 10)
# Make predictions on the test set
y_pred <- predict(model, test)
# Calculate accuracy
accuracy <- mean(y_pred == test_data$Species)
print(paste("Accuracy:", accuracy))

In this example:
- To load the iris datasets using the `createDataPartition` function from the `caret`
package.
- To convert the data into XGBoost's `DMatrix` format using the `[Link]`
function.
- To define hyperparameters such as the objective function, higher tree depth,
learning rate, and number of rounds.
- To train the XGBoost model using the `xgboost` function.
- To make predictions on the test set using the trained model.
- To calculate the accuracy of the model on the test set.

XGBoost is a powerful and versatile library for gradient boosting. It is widely used in
research and industry for various machine-learning tasks.

____________________________________________________________________________________
89 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Adaptive Boosting

Adaptive Boosting, commonly known as AdaBoost, is an ensemble learning method that


builds a robust classifier by sequentially combining multiple weak classifiers. It assigns
weights to training instances, with higher weights given to incorrectly classified
instances, allowing subsequent weak learners to focus more on complex cases. In Python,
AdaBoost is implemented in the AdaBoostClassifier class provided by scikit-learn.
Here is an example of using AdaBoost for classification in Python:

from [Link] import load_iris


from sklearn.model_selection import train_test_split
from [Link] import AdaBoostClassifier
from [Link] import DecisionTreeClassifier
from [Link] import accuracy_score
# Load the Iris dataset
iris = load_iris()
x, y = [Link], [Link]
# Split the dataset into training and testing sets
x_train, x_test, y_train, y_test = train_test_split(x, y, test_size=0.2, random_state=42)
# Define base classifier
base_classifier = DecisionTreeClassifier(max_depth=1)
# Create AdaBoost classifier with base classifier
ada_boost = AdaBoostClassifier(base_classifier, n_estimators=50, learning_rate=1.0,
random_state=42)
# Train the AdaBoost classifier
ada_boost.fit(X_train, y_train)
# Make predictions on the test set
y_pred = ada_boost.predict(X_test)
# Calculate accuracy
accuracy = accuracy_score(y_test, y_pred)
print("Accuracy:", accuracy)

- Decision trees are often used as beginning learners in AdaBoost.


- To create an AdaBoost classifier with the base classifier, specifying the number of
estimators (weak learners) and the learning rate.
____________________________________________________________________________________
90 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

- To train the AdaBoost classifier on the training data.


- To make predictions on the test set using the trained AdaBoost classifier.
- To calculate the accuracy of the model on the test set using scikit-learn's
`accuracy_score` function.

In R, AdaBoost is implemented through the `adabag` package. Here's an example of using


AdaBoost for classification in R:

# Install and load necessary packages (if not already installed)


# [Link]("adabag")
library(adabag)
# Load the Iris dataset
data(iris)
[Link](42) # for reproducibility
train_index <- sample(1:nrow(iris), 0.8*nrow(iris))
train_data <- iris[train_index, ]
test_data <- iris[-train_index, ]
# Train an AdaBoost classifier
ada_model <- boosting(Species ~ ., data = train_data, boos = TRUE, final = 50)
# Make predictions on the test set
predictions <- predict(ada_model, newdata = test_data)
# Calculate accuracy
accuracy <- mean(predictions$class == test_data$Species)
print(paste("Accuracy:", accuracy))

- To train an AdaBoost classifier using the `boosting` function from the `adabag`
package.
- To specify the formula for the model (`Species ~ .`), the data, and set `boos =
TRUE` to indicate that AdaBoost should be used.
- To make predictions on the test set using the trained AdaBoost model.
- To calculate the model's accuracy by comparing the predicted classes to the actual
classes in the test set.

AdaBoost is a robust algorithm for classification tasks and is ultimately used in practice
due to its simplicity and effectiveness, especially when combined with weak learners.

____________________________________________________________________________________
91 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

AdaBoost is a robust algorithm for classification tasks. It is particularly effective when


combined with weak learners. Due to its simplicity and effectiveness, it is commonly
used in practice.

Model Parameters

In the machine learning models, parameters are the configuration settings or variables that the
model learns from the training data. These parameters define the structure and behavior of the
model and are typically adjusted during the training process to optimize the model's
performance.
The Parameters can be classified into two types:
1. Hyperparameters: These parameters are set before the training begins and are not
learned from the data. Examples of hyperparameters include the learning rate,
regularization strength, and the number of layers in a neural network.
Hyperparameters are tuned manually or through automated techniques such as grid or
random search to find the optimal values that yield the best model performance.

2. Model Parameters: These are the parameters that the model learns from the training
data. They represent the relationships and patterns in the data the model captures
during training. For example, the model parameters are the coefficients assigned to
each feature in a linear regression model. In a neural network, the model parameters
are weighted on the neurons in the network.

Here is an example in Python using a linear regression model from scikit-learn:

from [Link] import load_diabetes


from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from [Link] import mean_squared_error
# Load the diabetes dataset
diabetes = load_diabetes()
x, y = [Link], [Link]
# Split the dataset into training and testing sets
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
# Create a linear regression model
model = LinearRegression()
# Train the model on the training data
[Link](X_train, y_train)

____________________________________________________________________________________
92 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

# Get the learned model parameters (coefficients)


model_parameters = model. coef_
intercept = model. intercept_
print("Model Parameters (Coefficients):", model_parameters)
print("Model Intercept:", intercept)

In this example,
● To load the diabetes dataset and split it into training and testing sets.
● To create a linear regression model using the LinearRegression class from scikit-
learn.
● To train the model on the training data using the fit method.
● To obtain the learned model parameters (coefficients) using the coef_ attribute and the
intercept using the intercept_ attribute.
These learned model parameters represent the linear relationship and the features of the
target variable in the training data, as captured by the linear regression model.
Here's an example in R using a linear regression model:

# Load the mtcars dataset


data(mtcars)
# Fit a linear regression model
lm_model <- lm(mpg ~ cyl + disp + hp + wt, data = mtcars)
# Get the learned model parameters (coefficients)
model_parameters <- coef(lm_model)
print(model_parameters)

In this example,
- To load the mtcars dataset.
- A linear regression model uses the lm function, where mpg is the response
variable, and cyl, disp, hp, and wt are the predictor variables.
- To extract the learned model parameters (coefficients) using the coef function.
The model_parameters object contains the coefficients of the linear regression model,
which represent the relationships between the predictor variables and response variables.
Model parameters can vary depending on the type of model being used. For example, in a
neural network, model parameters would include the weights and biases of the neurons.
In contrast, model parameters would consist of the splitting thresholds and leaf values in a
decision tree.

____________________________________________________________________________________
93 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Hyperparameter Tuning

Hyperparameters are configuration settings set before the training process begins and are not
learned from the data. Hyperparameters include learning rate, regularization strength, and the
number of layers in a neural network. Tuning these hyperparameters is crucial for achieving
the best performance of the model. Several techniques for hyperparameter tuning exist,
including grid search, random search, and Bayesian optimization. In Python using grid search
for hyperparameter tuning with a support vector machine (SVM) classifier:

from [Link] import load_iris


from sklearn.model_selection import train_test_split, GridSearchCV
from [Link] import SVC
from [Link] import accuracy_score
# Load the Iris dataset
iris = load_iris()
X, y = [Link], [Link]
# Split the dataset into training and testing sets
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
# Define the parameter grid to search
param_grid = {
'C': [0.1, 1, 10, 100], # Regularization parameter
'gamma': [0.001, 0.01, 0.1, 1], # Kernel coefficient
'kernel': ['rbf'] # Kernel type
}
# Create an SVM classifier
svm = SVC()
# Perform grid search with 5-fold cross-validation
grid_search = GridSearchCV(estimator=svm, param_grid=param_grid, cv=5)
# Fit the grid search to the training data
grid_search.fit(X_train, y_train)
# Get the best hyperparameters found by grid search
best_params = grid_search.best_params_
# Get the best model
best_model = grid_search.best_estimator_
____________________________________________________________________________________
94 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

# Make predictions on the test set using the best model


y_pred = best_model.predict(X_test)
# Calculate accuracy
accuracy = accuracy_score(y_test, y_pred)
print("Best hyperparameters:", best_params)
print("Accuracy:", accuracy)
In this example,
● To load the Iris dataset and split it into training and testing sets.
● To define a grid of hyperparameters to search through. In this case, we're tuning the
regularization parameter C, the kernel coefficient gamma, and the kernel type kernel.
● To create an SVM classifier.
● To perform a grid search using 5-fold cross-validation (cv=5) to find the best
hyperparameters from the grid and fit the grid search to the training data.
● To extract the best hyperparameters found by the grid search.
● To get the best model obtained after tuning.
● To make predictions on the test set using the best model.
● To calculate the accuracy of the best model on the test set.
In R, hyperparameter tuning is performed using random search, grid search, and Bayesian
optimization techniques, a specified grid of hyperparameter values. Here is an example in R
using grid search for hyperparameter tuning with a support vector machine (SVM) classifier:

# Load the necessary packages


library(e1071)
library(caret)
# Load the Iris dataset
data(iris)
[Link](42)
train_index <- createDataPartition(iris$Species, p = 0.8, list = FALSE)
train_data <- iris[train_index, ]
test_data <- iris[-train_index, ]
# Define the parameter grid to search
param_grid <- [Link](
C = c(0.1, 1, 10, 100), # Regularization parameter
gamma = c(0.001, 0.01, 0.1, 1), # Kernel coefficient
kernel = "radial" # Kernel type)

____________________________________________________________________________________
95 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

# Create a training control with 5-fold cross-validation


train_control <- trainControl(method = "cv", number = 5)
# Perform grid search with SVM using 5-fold cross-validation
svm_grid <- train(
Species ~ .,
data = train_data,
method = "svmRadial",
trControl = train_control,
tuneGrid = param_grid)
# Get the best hyperparameters found by grid search
best_params <- svm_grid$bestTune
# Get the best model
best_model <- svm_grid$finalModel
# Make predictions on the test set using the best model
predictions <- predict(best_model, newdata = test_data)
# Calculate accuracy
accuracy <- mean(predictions == test_data$Species)
print(paste("Best hyperparameters:", best_params))
print(paste("Accuracy:", accuracy))

In this example,
● To define a grid of hyperparameters to search through. In this case, we're tuning the
regularization parameter C, the kernel coefficient gamma, and the kernel type kernel.
● To create a training control object with 5-fold cross-validation.
● To perform a grid search using the train function from the caret package to find the
best hyperparameters from the grid.
● To extract the best hyperparameters found by the grid search.
● To get the best model obtained after tuning.
● To make predictions on the test set using the best model.
● To calculate the accuracy of the best model on the test set.
Grid search is a simple yet effective method for hyperparameter tuning, but it can be
computationally expensive, especially for large hyperparameter grids. Other techniques like
random search or Bayesian optimization may be more efficient in some cases.

____________________________________________________________________________________
96 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Unit VII: Neural Networks, Text Mining


● Introduction to Artificial Intelligence
● Introduction to Neural Networks
● Artificial Neural Networks (ANNs)
● Convolutional Neural Networks (CNNs)
● RNN Recurrent Neural Networks (RNNs)
● Basics of Backpropagation
● Implementation of Simple Neural Networks
● Introduction to Text Mining and Sentiment Analysis, Word Cloud

Introduction to Artificial Intelligence


AI can perform tasks that typically require human intelligence. These tasks include learning,
reasoning, problem-solving, perception, understanding natural language, and interacting with
the environment. AI aims to replicate human-like intelligence in machines, enabling them to
understand, adapt, and make decisions in complex and uncertain environments. The concept
of AI dates back to ancient times. Still, significant advancements have been made in recent
decades, driven by developments in computer science, mathematics, cognitive psychology,
neuroscience, and other fields. AI techniques and approaches continue to evolve rapidly,
leading to various applications across different domains, including healthcare, finance,
transportation, education, entertainment, and more. AI can be categorized into two main parts:
Narrow AI (also known as Weak AI) and General AI (also known as Strong AI).

1. Narrow AI refers to AI systems designed and trained for specific tasks or domains. These
systems excel at performing particular tasks but cannot generalize their knowledge
beyond their predefined scope. Examples of narrow AI include virtual assistants (e.g.,
Siri, Alexa), image recognition systems, spam filters, and recommendation systems.
2. General AIrefers to AI systems with human-like intelligence capable of understanding,
learning, and reasoning across diverse tasks and domains. General AI can transfer the
knowledge and skills from one to another domain, adapt to new situations, and exhibit
creativity and consciousness. However, achieving true general AI remains an aspirational
goal and is currently a subject of research and speculation.
Essential techniques and approaches in AI include the following:
Machine learning: A subset of AI that creates algorithms and models that allow the
computers to learn from the data and make predictions and decisions without being explicitly
programmed. Machine learning includes supervised learning, unsupervised learning, and
reinforcement learning.
Deep learning: It is the part of machine learning that uses multilayer artificial neural
networks (deep neural networks) that can learn complex patterns and representations from
large amounts of data. Deep learning has been incredibly successful in tasks such as image
recognition, natural language processing, and speech recognition.

____________________________________________________________________________________
97 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Natural Language Processing (NLP): It is a field of computer science that focuses on the
interactions between computers and humans, enabling computers to understand and generate
human language. NLP techniques are used in machine translation, sentiment analysis,
chatbots, and text summarization applications.
Computer Vision: The field of computer vision allows machines to interpret and
comprehend visual information from the real world. Computer vision techniques are used in
object detection, image classification, facial recognition, and medical image analysis.
Robotics: It is the intersection of AI, engineering, and robotics aimed at designing and
developing intelligent robots capable of performing tasks autonomously or collaboratively
with humans. Robotics applications range from industrial automation and autonomous
vehicles to assistive robots in healthcare and home environments.
AI has the potential to revolutionize the various aspects of society, driving innovation,
efficiency, and progress. However, it also raises ethical, societal, and philosophical questions
regarding privacy, bias, job displacement, autonomy, and the future of humanity. As AI
advances, it is essential to ensure responsible development and deployment practices that
prioritize ethical considerations, transparency, fairness, and accountability.

Introduction to Neural Networks


Neural networks are a fundamental concept in artificial intelligence and machine learning,
inspired by the structure and function of the human brain. They are powerful computational
models composed of interconnected nodes (neurons) organized into layers. Each neuron
processes input data, performs computations, and passes the result to neurons in the next
layer.

Neural networks can learn the complex patterns and relationships in data by adjusting their
weights and biases through training. This process typically involves feeding input data
through a network, comparing the predicted output with the actual production, and updating
the network's parameters using optimization algorithms such as gradient descent. In Python,
neural networks are commonly implemented using libraries such as TensorFlow, Keras, and
PyTorch. Here is a simple example of building a neural network for image classification
using Keras:
Here is an Example:
import numpy as np
from [Link] import mnist
from [Link] import Sequential
from [Link] import Dense, Flatten
from [Link] import to_categorical
# Load the MNIST dataset

____________________________________________________________________________________
98 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

(x_train, y_train), (x_test, y_test) = mnist.load_data()


# Preprocess the data
x_train = x_train / 255.0
x_test = x_test / 255.0
# Flatten the input images
x_train = x_train.reshape((-1, 28 * 28))
x_test = x_test.reshape((-1, 28 * 28))
# Convert class vectors to binary class matrices
y_train = to_categorical(y_train, num_classes=10)
y_test = to_categorical(y_test, num_classes=10)
# Define the neural network architecture
model = Sequential([
Flatten(input_shape=(28, 28)),
Dense(128, activation='relu'),
Dense(64, activation='relu'),
Dense(10, activation='softmax')])
# Compile the model
[Link](optimizer='adam', loss='categorical_crossentropy', metrics=['accuracy'])
# Train the model
[Link](x_train, y_train, epochs=5, batch_size=32, validation_split=0.2)
# Evaluate the model
test_loss, test_acc = [Link](x_test, y_test)
print('Test accuracy:', test_acc)

In this example,
● To load the MNIST dataset, consisting of grayscale handwritten digit images (0-9).
● To preprocess the data by scaling pixel values to the range [0, 1].
● To define a neural network architecture using Keras' Sequential API. The network
consists of an input layer (Flatten), two hidden layers (Dense), and an output layer
(Dense) with softmax activation.
● To accumulate the model with the Adam enhancer and categorical cross-entropy loss
function.
● To train the model on the training data for five epochs using a batch size of 32.
● Finally, assess the trained model on the test data and print the test accuracy.

____________________________________________________________________________________
99 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

In R, neural networks can be implemented using packages such as ‘keras’, ‘nnet’, or ‘neuralnet’.
Here's a simple example of building a neural network for binary classification using the ‘keras’
package:

# Install and load the necessary packages


# [Link]("keras")
library(keras)
# Define the neural network architecture
model <- keras_model_sequential() %>%
layer_dense(units = 64, activation = 'relu', input_shape = c(20)) %>%
layer_dense(units = 32, activation = 'relu') %>%
layer_dense(units = 1, activation = 'sigmoid')
# Compile the model
model %>% compile(
optimizer = 'adam',
loss = 'binary_crossentropy',
metrics = c('accuracy'))
# Generate sample data
[Link](42)
x_train <- matrix(rnorm(2000), ncol = 20)
y_train <- rbinom(100, 1, 0.5)
x_test <- matrix(rnorm(400), ncol = 20)
y_test <- rbinom(100, 1, 0.5)
# Train the model
history <- model %>% fit(
x_train, y_train,
epochs = 20, batch_size = 32,
validation_split = 0.2)
# Evaluate the model
model %>% evaluate(x_test, y_test)

____________________________________________________________________________________
100 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

In this example,

- To explain a simple neural network architecture using the ‘keras_model_sequential()


‘function from the keras package. The network consists of an input layer (`layer_dense`), two
hidden layers, and an output layer with a single neuron and sigmoid activation function for
binary classification.
- To accumulate the model with the Adam enhancer and binary cross-entropy loss function.
- To generate sample data for training and testing.
- To train the model on the training data for 20 epochs with a batch size 32.
- To calculate the trained model on the test data.

Neural networks are versatile models capable of learning complex patterns in various data
types, making them suitable for multiple applications, including image classification, natural
language processing, and reinforcement learning.

Artificial Neural Networks (ANNs)


Artificial Neural Networks (ANNs) are computational models that draw inspiration from the
organization and operation of the human brain. They consist of interconnected nodes
(neurons) organized into layers. Each neuron processes input data, performs computations
using weights and biases, and passes the result to neurons in the next layer. ANNs can learn
complex patterns and relationships in data by adjusting their weights and biases through
training, typically using optimization algorithms such as gradient descent.
ANNs can be implemented in Python using libraries such as TensorFlow, Keras, PyTorch,
and scikit-learn. Here is a simple example of building a feedforward neural network for
binary classification using Keras:

import numpy as np
from [Link] import Sequential
from [Link] import Dense
from [Link] import make_classification
from sklearn.model_selection import train_test_split
from [Link] import StandardScaler
from [Link] import accuracy_score
# Generate synthetic data for binary classification
X, y = make_classification(n_samples=1000, n_features=20, n_classes=2, random_state=42)
# Split the data into training and testing sets
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
# Scale the features
____________________________________________________________________________________
101 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = [Link](X_test)
# Build the neural network architecture
model = Sequential([
Dense(128, activation='relu', input_shape=(20,)),
Dense(64, activation='relu'),
Dense(1, activation='sigmoid')
])
# Compile the model
[Link](optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])
# Train the model
[Link](X_train_scaled, y_train, epochs=10, batch_size=32, validation_split=0.2)
# Evaluate the model
y_pred = model.predict_classes(X_test_scaled).flatten()
accuracy = accuracy_score(y_test, y_pred)
print("Test Accuracy:", accuracy)

In this example,
● To generate synthetic data for binary classification using the make_classification
function from scikit-learn.
● To split the data into training and testing sets.
● To normalize the features to have a mean of 0 and a unit variance.
● To define a feedforward neural network with three dense layers using the Sequential
API from Keras. The network has two hidden layers with ReLU activation and a put
layer with sigmoid activation for binary classification.
● To accumulate the model with the Adam enhancer and a binary cross-entropy loss
function.
● This model is trained on the training dataset for 10 epochs, utilizing a batch size of 32
for each epoch.
● Then, assess the trained model on the test data and print the test accuracy.

ANNs can be implemented using keras, nnet, or neuralnet packages in R. Here is a simple
example of building a feedforward neural network for binary classification using the keras
package:

____________________________________________________________________________________
102 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

# Install and load the necessary packages


# [Link]("keras")
library(keras)
# Generate synthetic data for binary classification
[Link](42)
n <- 1000
p <- 20
X <- matrix(rnorm(n * p), nrow = n, ncol = p)
y <- rbinom(n, 1, 0.5)
# Split the data into training and testing sets
split_index <- sample(1:n, 0.8*n)
X_train <- X[split_index, ]
y_train <- y[split_index]
X_test <- X[-split_index, ]
y_test <- y[-split_index]
# Scale the features
scaler <- keras::layer_standardize()
# Build the neural network architecture
model <- keras_model_sequential() %>%
scaler %>%
layer_dense(units = 128, activation = 'relu', input_shape = c(p)) %>%
layer_dense(units = 64, activation = 'relu') %>%
layer_dense(units = 1, activation = 'sigmoid')
# Compile the model
model %>% compile(
optimizer = 'adam',
loss = 'binary_crossentropy',
metrics = c('accuracy'))
# Train the model
history <- model %>% fit(
X_train, y_train,
epochs = 10, batch_size = 32,

____________________________________________________________________________________
103 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

validation_split = 0.2)
# Evaluate the model
model %>% evaluate(X_test, y_test)
In this example:
● To generate synthetic data for binary classification.
● To split the data into training and testing sets.
● To use the Keras package to define a feedforward neural network with three dense
layers. The network has two hidden layers with ReLU activation and a put layer with
sigmoid activation for binary classification.
● To utilize the Adam optimizer and binary cross-entropy loss function to compile the
model.
● The model undergoes training on the training data for 10 epochs, utilizing a batch size
of 32.
● Finally, assess the trained model on the test data.

Artificial Neural Networks are powerful models capable of learning complex patterns in
various data types, making them suitable for multiple applications, including classification,
regression, image recognition, and natural language processing.

Convolutional Neural Networks (CNNs)


Convolutional Neural Networks (CNNs) are commonly used for analyzing visual imagery.
They have proven highly effective in image classification, object detection, and image
segmentation tasks.
Here is an example of building a simple CNN for image classification using Python with
TensorFlow and Keras:

import numpy as np
import [Link] as plt
from [Link] import mnist
from [Link] import Sequential
from [Link] import Conv2D, MaxPooling2D, Flatten, Dense
from [Link] import to_categorical

# Load the MNIST dataset (handwritten digits)


(x_train, y_train), (x_test, y_test) = mnist.load_data()

# Preprocess the data


x_train = np.expand_dims(x_train, axis=-1) / 255.0
x_test = np.expand_dims(x_test, axis=-1) / 255.0

____________________________________________________________________________________
104 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

y_train = to_categorical(y_train, num_classes=10)


y_test = to_categorical(y_test, num_classes=10)

# Define the CNN architecture


model = Sequential([
Conv2D(32, (3, 3), activation='relu', input_shape=(28, 28, 1)),
MaxPooling2D((2, 2)),
Conv2D(64, (3, 3), activation='relu'),
MaxPooling2D((2, 2)),
Flatten(),
Dense(64, activation='relu'),
Dense(10, activation='softmax')
])

# Compile the model


[Link](optimizer='adam', loss='categorical_crossentropy', metrics=['accuracy'])

# Train the model


history = [Link](x_train, y_train, epochs=5, batch_size=32, validation_split=0.2)

# Evaluate the model


test_loss, test_acc = [Link](x_test, y_test)
print('Test accuracy:', test_acc)

# Plot training and validation accuracy


[Link]([Link]['accuracy'], label='Training Accuracy')
[Link]([Link]['val_accuracy'], label='Validation Accuracy')
[Link]('Epoch')
[Link]('Accuracy')
[Link]()
[Link]()

____________________________________________________________________________________
105 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

In this example,
- To load the MNIST dataset, which consists of 28x28 pixel grayscale images of handwritten
digits (0-9).
-To ingest the MNIST dataset, comprising 28x28 pixel grayscale images representing
handwritten digits ranging from 0 to 9.
- To preprocess the data by scaling pixel values to the range [0, 1].
- To define a CNN architecture using Keras's `Sequential API. The architecture of CNN
includes two convolutional layers with ReLU activation followed by max-pooling and fully
connected layers.
- To build the model, we utilize the Adam optimizer and the categorical cross-entropy loss
function.
- This model is trained on the training dataset for 10 epochs, utilizing a batch size of 32 for
each epoch. Finally, we assess the trained model on the test data and print the test accuracy.
We also plot the training and validation accuracy over epochs.

In R, Convolutional Neural Networks (CNNs) can be implemented using the keras package,
which provides an interface to TensorFlow and allows for building and training deep learning
models. Here's an example of building a simple CNN for image classification using the keras
package in R:

Install and load necessary packages


# [Link]("keras")
library(keras)
# Load the MNIST dataset (handwritten digits)
mnist <- dataset_mnist()
x_train <- mnist$train$x
y_train <- mnist$train$y
x_test <- mnist$test$x
y_test <- mnist$test$y
# Preprocess the data
x_train <- array_reshape(x_train, c(nrow(x_train), 28, 28, 1)) / 255
x_test <- array_reshape(x_test, c(nrow(x_test), 28, 28, 1)) / 255
y_train <- to_categorical(y_train, 10)
y_test <- to_categorical(y_test, 10)
# Define the CNN architecture

____________________________________________________________________________________
106 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

model <- keras_model_sequential() %>%


layer_conv_2d(filters = 32, kernel_size = c(3, 3), activation = 'relu', input_shape = c(28, 28,
1)) %>%

layer_max_pooling_2d(pool_size = c(2, 2)) %>%


layer_conv_2d(filters = 64, kernel_size = c(3, 3), activation = 'relu') %>%
layer_max_pooling_2d(pool_size = c(2, 2)) %>%
layer_flatten() %>%
layer_dense(units = 64, activation = 'relu') %>%
layer_dense(units = 10, activation = 'softmax')
# Compile the model
model %>% compile(
loss = 'categorical_crossentropy',
optimizer = optimizer_adam(),
metrics = c('accuracy'))
# Train the model
history <- model %>% fit(
x_train, y_train,
epochs = 5, batch_size = 32,
validation_split = 0.2)
# Evaluate the model
model %>% evaluate(x_test, y_test)

In this example,
● To load the MNIST dataset, which consists of 28x28 pixel grayscale images of
handwritten digits (0-9).
● To preprocess the data by reshaping the images and scaling pixel values to the range
[0, 1].
● To define a CNN architecture using the keras_model_sequential() function from the
keras package. The architecture of CNN comprises two convolutional layers with
ReLU activation, followed by max-pooling and fully connected layers.
● To compile a model with Adam optimizer and categorical cross-entropy loss function.
● This trains the model on the training data for 5 epochs with a batch size of 32.
● To evaluate the trained model on the test data and print the test accuracy.

____________________________________________________________________________________
107 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

CNN are powerful models that are widely used for image classification in computer vision
applications. The keras package in R provides a convenient interface for building and training
CNNs, leveraging the capabilities of the TensorFlow backend. CNNs have revolutionized
computer vision and are used in image recognition, object detection, and image segmentation.
They are particularly effective in extracting hierarchical features from images, making them
suitable for understanding complex visual patterns.

Recurrent Neural Networks (RNNs)


Recurrent Neural Networks (RNNs) are a class of neural networks designed to model
sequential data by maintaining an internal state or memory. Tasks such as time series
prediction, natural language processing, and speech recognition are areas where they excel.
Here is an example of building a simple RNN for time series prediction using Python with
TensorFlow and Keras:

‘’’python
import numpy as np
import [Link] as plt
from [Link] import Sequential
from [Link] import SimpleRNN, Dense
from sklearn.model_selection import train_test_split
# Generate synthetic time series data
def generate_time_series_data(n_samples, n_steps):
freq1, freq2, offset1, offset2 = [Link](4, n_samples, 1)
time = [Link](0, 1, n_steps)
series = 0.5 * [Link]((time - offset1) * (freq1 * 10 + 10)) # wave 1
series += 0.2 * [Link]((time - offset2) * (freq2 * 20 + 20)) # + wave 2
return series[..., [Link]].astype(np.float32)
# Generate synthetic time series data
n_samples = 10000
n_steps = 50
series = generate_time_series_data(n_samples, n_steps)
# Split the data into training and testing sets
X_train, X_test, y_train, y_test = train_test_split(series[:, :n_steps-1, :], series[:, 1:, :],
test_size=0.2, random_state=42)
____________________________________________________________________________________
108 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

# Define the RNN architecture


model = Sequential([
SimpleRNN(20, return_sequences=True, input_shape=[None, 1]),
SimpleRNN(20, return_sequences=True),
Dense(1)])
# Compile the model
[Link](optimizer='adam', loss='mse')
# Train the model
history = [Link](X_train, y_train, epochs=20, validation_split=0.2)
# Plot training and validation loss
[Link]([Link]['loss'], label='Training Loss')
[Link]([Link]['val_loss'], label='Validation Loss')
[Link]('Epoch')
[Link]('Loss')
[Link]()
[Link]()
# Predict time series values
y_pred = [Link](X_test)
# Plot example predictions
[Link](X_test[0, :, 0], label='Input')
[Link](y_test[0, :, 0], label='Target')
[Link](y_pred[0, :, 0], label='Prediction')
[Link]()
[Link]()

In this example,
- To generate synthetic time series data with two sine waves of different frequencies.
- To split the data into training and testing sets.
- To define an RNN architecture using Keras's `Sequential API. The RNN consists of two
Simple RNN layers with 20 units each and a dense output layer.
-To use the Adam optimizer and mean squared error (MSE) loss function to compile the
model.
____________________________________________________________________________________
109 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

- To train the model on the training data for 20 epochs.


- Finally, plot the training and validation loss curves and visualize example predictions made
by the model.

Recurrent Neural Networks (RNNs) can be implemented using the `keras` package, which
provides an interface to TensorFlow and allows for building and training deep learning
models. Here is an example of building a simple RNN for time series prediction using the
`keras` package in R:
```R
# Install and load necessary packages
# [Link]("keras")
library(keras)
library(tidyverse)
# Generate synthetic time series data
generate_time_series_data <- function(n_samples, n_steps) {
freq1 <- runif(n_samples)
freq2 <- runif(n_samples)
offset1 <- runif(n_samples)
offset2 <- runif(n_samples)
time <- seq(0, 1, [Link] = n_steps)
series <- 0.5 * sin((time - offset1) * (freq1 * 10 + 10)) +
0.2 * sin((time - offset2) * (freq2 * 20 + 20))
return([Link](series))
}
# Generate synthetic time series data
n_samples <- 10000
n_steps <- 50
series <- generate_time_series_data(n_samples, n_steps)

# Prepare input and target data


X <- series[, -n_steps]
y <- series[, -1]
____________________________________________________________________________________
110 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

# Reshape data for RNN input


X <- array_reshape(X, c(n_samples, n_steps - 1, 1))
y <- array_reshape(y, c(n_samples, n_steps - 1, 1))

# Split the data into training and testing sets


split_index <- sample(1:n_samples, 0.8 * n_samples)
X_train <- X[split_index, , ]
y_train <- y[split_index, , ]
X_test <- X[-split_index, , ]
y_test <- y[-split_index, , ]

# Define the RNN architecture


model <- keras_model_sequential() %>%
layer_simple_rnn(units = 20, return_sequences = TRUE, input_shape = c(n_steps - 1, 1))
%>%
layer_simple_rnn(units = 20, return_sequences = TRUE) %>%
layer_dense(units = 1)

# Compile the model


model %>% compile(
optimizer = optimizer_adam(),
loss = "mse"
)

# Train the model


history <- model %>% fit(
X_train, y_train,
epochs = 20,
validation_split = 0.2
)

____________________________________________________________________________________
111 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

# Plot training and validation loss


plot(history)

# Predict time series values


y_pred <- model %>% predict(X_test)

# Plot example predictions


plot(y_test[1, , 1], type = "l", col = "blue", ylim = c(-1, 1), xlab = "Time Step", ylab =
"Value", main = "Example Predictions")
lines(y_pred[1, , 1], col = "red")
legend("topright", legend = c("Target", "Prediction"), col = c("blue", "red"), lty = 1)
```

In this example,
- To generate synthetic time series data with two sine waves of different frequencies.
- To split the data into training and testing sets.
- To define an RNN architecture using the `keras_model_sequential()` function from the
`keras` package. The RNN consists of two `SimpleRNN` layers with 20 units each and a
dense output layer.
- To compile the model with the Adam optimizer and mean squared error (MSE) loss
function.
- To train the model on the training data for 20 epochs.
- Finally, plot the training and validation loss curves and visualize example predictions made
by the model.

RNNs are powerful models for sequential data modeling and prediction tasks. Temporal
dependencies in data can be captured by them, making them suitable for a wide range of
applications.

Basics of Backpropagation
Backpropagation is a fundamental algorithm used for training artificial neural networks,
including deep learning models. Efficiently compute gradient of loss function network
weights, essential for updating during training.

____________________________________________________________________________________
112 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Here is a basic explanation of backpropagation with an example implemented in python:


### Basics of Backpropagation:
1. Forward Pass: In the forward pass, input data is passed through the neural network, and
the output is computed layer by layer until the final output is obtained.
2. Compute Loss: After obtaining the output, the loss function is calculated to measure the
difference between the predicted and actual output
3. Backpropagation involves calculating the gradients of the loss function in relation to each
parameter in the network, including weights and biases. This involves the application of the
chain rule in calculus, which allows the error to be propagated backwards through the
network.
4. Update Weights: After computing the gradients, the weights of the network are updated
using an optimization algorithm (e.g., stochastic gradient descent) to minimize the loss
function.

### Example of Backpropagation in Python:


Let us consider a simple example of training a neural network to perform binary classification
using backpropagation:
```python
import numpy as np
# Define the sigmoid activation function
def sigmoid(x):
return 1 / (1 + [Link](-x))
# Define the derivative of the sigmoid function
def sigmoid_derivative(x):
return x * (1 - x)
# Define input data
X = [Link]([[0, 0],
[0, 1],
[1, 0],
[1, 1]])
# Define target data
y = [Link]([[0],
[1],
____________________________________________________________________________________
113 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

[1],
[0]])
# Initialize random weights and biases
[Link](42)
weights_input_hidden = [Link](2, 3)
weights_hidden_output = [Link](3, 1)
biases_hidden = [Link](1, 3)
bias_output = [Link](1, 1)
learning_rate = 0.1
# Training loop
for epoch in range(10000):
# Forward pass
hidden_layer_input = [Link](X, weights_input_hidden) + biases_hidden
hidden_layer_output = sigmoid(hidden_layer_input)
output_layer_input = [Link](hidden_layer_output, weights_hidden_output) + bias_output
predicted_output = sigmoid(output_layer_input)
# Compute loss
loss = [Link]((y - predicted_output) ** 2)
# Backward pass (backpropagation)
output_error = (y - predicted_output) * sigmoid_derivative(predicted_output)
hidden_error = [Link](output_error, weights_hidden_output.T) *
sigmoid_derivative(hidden_layer_output)
# Update weights and biases
weights_hidden_output += [Link](hidden_layer_output.T, output_error) * learning_rate
bias_output += [Link](output_error, axis=0, keepdims=True) * learning_rate
weights_input_hidden += [Link](X.T, hidden_error) * learning_rate
biases_hidden += [Link](hidden_error, axis=0, keepdims=True) * learning_rate
# Print loss every 1000 epochs
if epoch % 1000 == 0:
print(f'Epoch: {epoch}, Loss: {loss}')
____________________________________________________________________________________
114 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

# Print final predictions


print("Final Predictions:")

print(predicted_output)

In this example,
- To define a simple neural network with one hidden layer and one output layer.
- To initialize random weights and biases.
- To perform the forward pass to compute the predicted output.
- To compute the loss using mean squared error.
- To perform the backward pass (backpropagation) to compute gradients.
- To update weights and biases using the gradients and a learning rate.
- To repeat this process for multiple epochs until convergence.

#Let us consider a simple example of training a neural network to perform binary


classification using backpropagation:
# Install and load necessary packages
# [Link]("keras")
library(keras)
# Define neural network architecture
model <- keras_model_sequential() %>%
layer_dense(units = 3, activation = "sigmoid", input_shape = c(2)) %>%
layer_dense(units = 1, activation = "sigmoid")
# Compile the model
model %>% compile(
optimizer = optimizer_sgd(lr = 0.1),
loss = "binary_crossentropy",
metrics = "accuracy")
# Define input data
X <- matrix(c(0, 0, 0, 1, 1, 0, 1, 1), ncol = 2)
# Define target data
y <- matrix(c(0, 1, 1, 0), ncol = 1)
____________________________________________________________________________________
115 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

# Train the model


history <- model %>% fit(X, y, epochs = 1000, batch_size = 1)
# Print summary of the trained model
summary(model)
# Print training history
plot(history)
# Predictions
predictions <- model %>% predict(X)
print(predictions)

In this example,
- To define a simple neural network with one hidden layer and one output layer using the
`keras` package.
- To compile the model with stochastic gradient descent (SGD) optimizer and binary cross-
entropy loss function.
- To define input and target data for binary classification.
- To train the model on the input and target data for 1000 epochs with a batch size of 1.
- To visualize the training history and print the summary of the trained model, to have
completed the training process, can utilize the model to forecast outcomes.

This is an example of backpropagation in python and R for training a neural network.


Backpropagation is a foundational concept in deep learning and is used in more complex
models for tasks such as image classification, natural language processing, and reinforcement
learning.

Implementation of Simple Neural Networks


Implement a simple neural network using Python with TensorFlow and Keras. To create a
primary neural network for binary classification using synthetic data.

```python
import numpy as np
from [Link] import Sequential
from [Link] import Dense
from [Link] import make_classification
from sklearn.model_selection import train_test_split

____________________________________________________________________________________
116 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

from [Link] import StandardScaler


from [Link] import accuracy_score
# Generate synthetic data for binary classification
X, y = make_classification(n_samples=1000, n_features=20, n_classes=2, random_state=42)
# Split the data into training and testing sets
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
# Scale the features
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = [Link](X_test)
# Build the neural network architecture
model = Sequential([
Dense(128, activation='relu', input_shape=(20,)),
Dense(64, activation='relu'),
Dense(1, activation='sigmoid')
])
# Compile the model
[Link](optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])
# Train the model
[Link](X_train_scaled, y_train, epochs=10, batch_size=32, validation_split=0.2)
# Evaluate the model
test_loss, test_acc = [Link](X_test_scaled, y_test)
print('Test Accuracy:', test_acc)
In this example,
- To generate synthetic data for binary classification using `make_classification` from scikit-
learn.
- To split the data into training and testing sets.
- To normalize the features, ensure a mean of 0 and standard deviation of 1.
- To define a feedforward neural network with three dense layers using the `Sequential` API
from Keras. The network has two hidden layers with the ReLU activation and output layer
with sigmoid activation for binary classification.
-To build our model, we utilize the Adam optimizer and binary cross-entropy loss function.
- This model is trained on the training dataset for 10 epochs, utilizing a batch size of 32 for
each epoch.
- Finally, calculate the trained model on the test data and print the test accuracy.
____________________________________________________________________________________
117 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

```R
# Install and load necessary packages
# [Link]("keras")
library(keras)
library(tidymodels)
# Generate synthetic data for binary classification
[Link](42)
data <- synthetic_classification(n = 1000, noise = 0.1) %>%
[Link]()
# Split the data into training and testing sets
split <- initial_split(data, prop = 0.8, strata = Class)
train_data <- training(split)
test_data <- testing(split)
# Preprocess the data
preprocess <- recipe(Class ~ ., data = train_data) %>%
step_scale(all_predictors()) %>%
step_center(all_predictors()) %>%
prep()
train_data_prep <- bake(preprocess, train_data)
test_data_prep <- bake(preprocess, test_data)
# Define the neural network architecture
model <- keras_model_sequential() %>%
layer_dense(units = 128, activation = 'relu', input_shape = ncol(train_data_prep) - 1) %>%
layer_dense(units = 64, activation = 'relu') %>%
layer_dense(units = 1, activation = 'sigmoid')
# Compile the model
model %>% compile(
loss = 'binary_crossentropy',
optimizer = optimizer_adam(),
metrics = c('accuracy'))
# Train the model
history <- model %>% fit(
x = [Link](select(train_data_prep, -Class)),
y = [Link](train_data_prep$Class) - 1, # Convert to numeric and zero-indexed
____________________________________________________________________________________
118 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

epochs = 10,
batch_size = 32,
validation_split = 0.2)
# Evaluate the model
model %>% evaluate(
x = [Link](select(test_data_prep, -Class)),
y = [Link](test_data_prep$Class) - 1 # Convert to numeric and zero-indexed)

In this example,
- To generate synthetic data for binary classification using the `synthetic_classification`
function from the `tidymodels` package.
- To use the 'initial_split' function from the 'tidymodels' package to divide data into training
and testing sets.
- To preprocess the data by scaling and centering the predictors using the `recipe` and `bake`
functions from the `tidymodels` package.
- To use the `keras_model_sequential` function from the `keras` package to define a neural
network architecture that consists of two hidden layers and an output layer.
- To compile the model with the binary cross-entropy loss function, Adam optimizer, and
accuracy metric.
- This model is trained on the training dataset for 10 epochs, utilizing a batch size of 32
for each epoch.
-To calculate the trained model on the test data using the `evaluate` method from the `keras`
package.
This is a simple example of implementing a neural network for binary classification using
Python with TensorFlow and Keras. The architecture, hyperparameters, and other aspects of
the model to suit your specific problem.

Introduction to Text Mining

Text mining involves extracting meaningful insights from unstructured text data. Sentiment
analysis involves identifying subjective information from text, such as opinions and
emotions, and determining their polarity as positive, negative, or neutral.

### Text Mining and Sentiment Analysis Example:```python


import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.naive_bayes import MultinomialNB
____________________________________________________________________________________
119 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

from [Link] import accuracy_score, classification_report


# Load the dataset
data = pd.read_csv('amazon_reviews.csv')
# Explore the dataset
print([Link]())
# Split the dataset into training and testing sets
X_train, X_test, y_train, y_test = train_test_split(data['Review'], data['Label'], test_size=0.2,
random_state=42)
# Convert text data into numerical features using Bag-of-Words
vectorizer = CountVectorizer()
X_train_vec = vectorizer.fit_transform(X_train)
X_test_vec = [Link](X_test)
# Train a classifier (e.g., Naive Bayes) for sentiment analysis
clf = MultinomialNB()
[Link](X_train_vec, y_train)
# Make predictions on the testing set
y_pred = [Link](X_test_vec)
# Evaluate the classifier
accuracy = accuracy_score(y_test, y_pred)
print("Accuracy:", accuracy)
print("\nClassification Report:")
print(classification_report(y_test, y_pred))

In this example,
- To load a dataset containing Amazon product reviews and their corresponding labels
(positive or negative sentiment).
- To split the dataset into testing and training sets.
- To use the Bag-of-Words (BoW) approach to convert text data into numerical features.
- To train a classifier (Multinomial Naive Bayes) using the training data.
- To make predictions on the testing set and evaluate the classifier's performance using
accuracy and a classification report.

### Text Mining and Sentiment Analysis Example:


```R
# Install and load necessary packages

____________________________________________________________________________________
120 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

# [Link]("tidyverse")
# [Link]("textdata")
# [Link]("tm")
# [Link]("wordcloud")
library(tidyverse)
library(textdata)
library(tm)
library(wordcloud)
# Load the dataset
data <- [Link]("amazon_reviews.csv")
# Explore the dataset
head(data)
# Preprocess the text data
corpus <- Corpus(VectorSource(data$Review))
corpus <- tm_map(corpus, content_transformer(tolower))
corpus <- tm_map(corpus, removeNumbers)
corpus <- tm_map(corpus, removePunctuation)
corpus <- tm_map(corpus, removeWords, stopwords("en"))
corpus <- tm_map(corpus, stripWhitespace)
# Convert the corpus to a document term matrix
dtm <- DocumentTermMatrix(corpus)
# Build a word cloud to visualize most frequent words
word_freq <- colSums([Link](dtm))
wordcloud(names(word_freq), word_freq, [Link] = 50, [Link] = FALSE)
# Perform sentiment analysis
sentiment <- textdata_sentiment(data$Review)
print(sentiment)

In this example,
- To load a dataset containing Amazon product reviews and their corresponding text.
- To preprocess the text data, including converting text to lowercase and removing numbers,
punctuation, stopwords, and extra whitespaces.
- To create a document term matrix (DTM) to represent the frequency of words in the corpus.
- To visualize the most frequent words using a word cloud.

____________________________________________________________________________________
121 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

- To perform sentiment analysis using the `textdata_sentiment` function from the `textdata`
package, which provides sentiment scores (positive, negative, neutral) for each review.

A basic approach to sentiment analysis using text mining techniques in python. There are
more advanced methods and techniques available, including deep learning models (e.g.,
recurrent neural networks, transformers) and pretrained language models. Depending on the
specific task and requirements, to explore and experiment with different approaches to
achieve better performance.

Word Cloud

A word cloud is a visual representation of text data where the size of each word indicates its
frequency or importance.
```python
from wordcloud import WordCloud
import [Link] as plt
# Sample text data
text = "Python is an amazing programming language. It is used for data analysis, machine
learning, and web development. Python has a large community and extensive libraries."
# Generate a word cloud
wordcloud = WordCloud(width=800, height=400, background_color='white').generate(text)
# Plot word cloud
[Link](figsize=(10, 5))
[Link](wordcloud, interpolation='bilinear')
[Link]('off')
[Link]()
In this example,
- First import the `WordCloud` class from the `wordcloud` library and `[Link]` for
plotting.
- Define some sample text data.
- Generate the word cloud using the `WordCloud` class, specifying the width, height, and
background color.
- Finally, plot the word cloud using `[Link]()` and configure the plot with
`[Link]('off')` to remove axes and `[Link]()` to display the plot.
In R, create a word cloud using the `wordcloud` package, which provides a simple interface
for generating word clouds. Here is an example of creating a word cloud in R:
```R
# Install and load necessary packages
____________________________________________________________________________________
122 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

# [Link]("wordcloud")
library(wordcloud)
# Sample text data
text <- "Python is an amazing programming language. It is used for data analysis, machine
learning, and web development. Python has a large community and extensive libraries."
# Generate word cloud
wordcloud(words = strsplit(text, " ")[[1]], freq = NULL, scale = c(3, 0.5), [Link] = 1,
[Link] = 100, [Link] = TRUE, [Link] = 0.35, colors = [Link](8,
"Dark2"))

In this example,
- First load the `wordcloud` package.
- Define some sample text data.
- Generate the word cloud using the `wordcloud` function. To provide the words as a vector
(extracted from the text using `strsplit`) and specify various parameters such as scale,
[Link], [Link], [Link], [Link], and colors.
- The `[Link]` function is used to specify colors for the word cloud.
This will generate a word cloud visualization where the size of each word represents its
frequency in the text. The size of a word in a word cloud is proportional to its frequency in
the text.

____________________________________________________________________________________
123 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Unit VIII: Reinforcement Learning


● Reinforcement Learning
● Value Function
● Bellman's Equation
● Application of Neural Networks in Current Industry

Reinforcement Learning
Value function, Bellman's equation, Application of neural networks in current industry.
Reinforcement Learning (RL): Interacting with an environment, an agent learns to make
decisions using machine learning. The agent learns through trial and error, receiving feedback
in the form of rewards or penalties, and adjusts its actions to maximize long-term benefits.
Reinforcement Learning implemented in python using the OpenAI Gym environment:
Here is an example for RL:

import numpy as np
import gym
# Create the environment
env = [Link]('Taxi-v3')
# Initialize Q-table with zeros
Q = [Link]([env.observation_space.n, env.action_space.n])
# Set hyperparameters
alpha = 0.1 # learning rate
gamma = 0.6 # discount factor
epsilon = 0.1 # exploration-exploitation trade-off
# Number of episodes
num_episodes = 10000
# Q-learning algorithm
for episode in range(num_episodes):
state = [Link]()
done = False
while not done:
# Choose action using epsilon-greedy policy
if [Link](0, 1) < epsilon:
____________________________________________________________________________________
124 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

action = env.action_space.sample() # exploration


else:
action = [Link](Q[state, :]) # exploitation
# Take action and observe the next state and reward
next_state, reward, done, _ = [Link](action)
# Update Q-table using Q-learning equation
Q[state, action] = Q[state, action] + alpha * (reward + gamma * [Link](Q[next_state, :])
- Q[state, action])
state = next_state
# Evaluate the trained agent
total_rewards = 0
num_episodes_eval = 100
for _ in range(num_episodes_eval):
state = [Link]()
done = False
while not done:
action = [Link](Q[state, :])
state, reward, done, _ = [Link](action)
total_rewards += reward
# Calculate the average reward
average_reward = total_rewards / num_episodes_eval
print("Average reward:", average_reward)
# Close the environment
[Link]()

In this example,
- Import the necessary libraries, including OpenAI Gym, which provides RL environments.
- Create an instance of the Taxi-v3 environment from OpenAI Gym.
- Initialize a Q-table with zeros. Each row corresponds to a state and each column
corresponds to an action.

____________________________________________________________________________________
125 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

- Define hyperparameters such as learning rate (alpha), discount factor (gamma), and
exploration-exploitation trade-off (epsilon).
The Q-learning algorithm for a fixed number of episodes, during which the agent explores the
environment, updates the Q-table based on received rewards, and gradually learns an optimal
policy.
- To evaluate the trained agent's performance by running it for a fixed number of episodes
and calculating the average reward obtained.

# Install and load necessary packages


# [Link]("RLearn")
library(RLearn)
# Create environment (a simple grid world)
env <- makeEnvironment(
states = c("A", "B", "C", "D", "E"),
actions = c("left", "right"),
transition = function(state, action) {
next_state <- state
if (state == "A" && action == "right") {
next_state <- "B"
} else if (state == "B" && action == "left") {
next_state <- "A"
} else if (state == "B" && action == "right") {
next_state <- "C"
} else if (state == "C" && action == "left") {
next_state <- "B"
} else if (state == "C" && action == "right") {
next_state <- "D"
} else if (state == "D" && action == "left") {
next_state <- "C"
} else if (state == "D" && action == "right") {
next_state <- "E"

____________________________________________________________________________________
126 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

} else if (state == "E" && action == "left") {


next_state <- "D"
}
return(list(state = next_state))
}
)
# Initialize Q-table with zeros
Q <- matrix(0, nrow = length(env$states), ncol = length(env$actions), dimnames =
list(env$states, env$actions))
# Set hyperparameters
alpha <- 0.1 # learning rate
gamma <- 0.9 # discount factor
epsilon <- 0.1 # exploration-exploitation trade-off
# Q-learning algorithm
num_episodes <- 1000
for (episode in 1:num_episodes) {
state <- sample(env$states, 1)
while (state != "E") {
# Choose action using epsilon-greedy policy
if (runif(1) < epsilon) {
action <- sample(env$actions, 1) # exploration
} else {
action <- names([Link](Q[state, ])) # exploitation
}
# Take action and observe the next state
next_state <- env$transition(state, action)$state
# Update Q-table using Q-learning equation
Q[state, action] <- Q[state, action] + alpha * (0 + gamma * max(Q[next_state, ]) - Q[state,
action])
state <- next_state

____________________________________________________________________________________
127 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

}
}
# Print Q-table
print("Q-table:")
print(Q)
# Choose optimal policy
optimal_policy <- apply(Q, 1, function(row) names([Link](row)))
print("Optimal policy:")
print(optimal_policy)

In this example,
- To use the `RLearn` package to create a simple grid world environment.
- Initialize a Q-table with zeros. Each row represents a state, and each column represents an
action.
- Define hyperparameters such as learning rate (alpha), discount factor (gamma), and
exploration-exploitation trade-off (epsilon).
- To run the Q-learning algorithm for a fixed number of episodes, during which the agent
explores the environment, updates the Q-table based on received rewards, and gradually
learns an optimal policy.
- Finally, print the learned Q-table and the optimal policy chosen by the agent.

Value Function

In reinforcement learning, a value function estimates the expected return (total future
rewards) that an agent can receive from a particular state or state-action pair. There are two
types of value functions: state value function (V(s)) and action value function (Q(s, a)).

Here is an example of implementing a state value function in Python using the Gridworld
environment:
import numpy as np
# Define the Gridworld environment
class Gridworld:
def __init__(self, size):
[Link] = size
____________________________________________________________________________________
128 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

[Link] = [Link]((size, size))


self.start_state = (0, 0)
self.goal_state = (size-1, size-1)
def get_actions(self):
return [(0, 1), (0, -1), (1, 0), (-1, 0)]
def is_valid_state(self, state):
x, y = state
return 0 <= x < [Link] and 0 <= y < [Link]
def is_terminal_state(self, state):
return state == self.goal_state
def step(self, state, action):
x, y = state
dx, dy = action
new_state = (x + dx, y + dy)
if self.is_valid_state(new_state):
return new_state
else:
return state
# Define a function to calculate the state value function using value iteration
def value_iteration(env, gamma=0.9, theta=1e-5):
V = [Link](([Link], [Link]))
while True:
delta = 0
for i in range([Link]):
for j in range([Link]):
state = (i, j)
if env.is_terminal_state(state):
continue
v = V[i, j]

____________________________________________________________________________________
129 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

max_value = -[Link]
for action in env.get_actions():
next_state = [Link](state, action)
x, y = next_state
reward = 0 if env.is_terminal_state(next_state) else -1
value = reward + gamma * V[x, y]
max_value = max(max_value, value)
V[i, j] = max_value
delta = max(delta, abs(v - V[i, j]))
if delta < theta:
break
return V
# Create the Gridworld environment
env = Gridworld(size=5)
# Calculate the state value function using value iteration
V = value_iteration(env)
# Print the state value function
print("State Value Function:")
print(V)

In this example,
- Define a simple Gridworld environment where the agent can move in four directions: up,
down, left, and right.

- To implement the value iteration algorithm to calculate the state value function (V(s)) for
each state in the Gridworld.
- The state value function represents the expected return (total future rewards) that an agent
can receive from each state.
- To print the calculated state value function for the Gridworld environment.

____________________________________________________________________________________
130 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

Here is an example of implementing a state value function in R using the Gridworld


environment:

# Define the Gridworld environment


Gridworld <- function(size) {
grid <- matrix(0, nrow = size, ncol = size)
start_state <- c(1, 1)
goal_state <- c(size, size)
# Define a function to check if a state is valid
is_valid_state <- function(state) {
x <- state[1]
y <- state[2]
return (x >= 1 && x <= size && y >= 1 && y <= size)
}
# Define a function to check if a state is terminal
is_terminal_state <- function(state) {
return (all(state == goal_state))
}
# Define a function to get valid actions
get_actions <- function() {
return (list(c(0, 1), c(0, -1), c(1, 0), c(-1, 0)))
}
# Define the function to take a step in the environment
step <- function(state, action) {
new_state <- state + action
if (is_valid_state(new_state)) {
return (new_state)
} else {
return (state)
}
}
____________________________________________________________________________________
131 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

return (list(grid = grid, start_state = start_state, goal_state = goal_state,


is_valid_state = is_valid_state, is_terminal_state = is_terminal_state,
get_actions = get_actions, step = step))
}
# Define a function to calculate the state value function using value iteration
value_iteration <- function(env, gamma = 0.9, theta = 1e-5) {
V <- matrix(0, nrow = nrow(env$grid), ncol = ncol(env$grid))
while (TRUE) {
delta <- 0
for (i in 1:nrow(env$grid)) {
for (j in 1:ncol(env$grid)) {
state <- c(i, j)
if (env$is_terminal_state(state)) {
next
}
v <- V[i, j]
max_value <- -Inf
for (action in env$get_actions()) {
next_state <- env$step(state, action)
x <- next_state[1]
y <- next_state[2]
reward <- ifelse(env$is_terminal_state(next_state), 0, -1)
value <- reward + gamma * V[x, y]
max_value <- max(max_value, value)
}
V[i, j] <- max_value
delta <- max(delta, abs(v - V[i, j]))} }
if (delta < theta) {
break

____________________________________________________________________________________
132 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

}
}
return (V)
}
# Create the Gridworld environment
env <- Gridworld(size = 5)
# Calculate the state value function using value iteration
V <- value_iteration(env)
# Print the state value function
print("State Value Function:")
print(V)

In this example,
- To define a simple Gridworld environment where the agent can move in four directions: up,
down, left, and right.
- To implement the value iteration algorithm to calculate the state value function (V(s)) for
each state in the Gridworld.
- The state value function represents the expected return (total future rewards) that an agent
can receive from each state.
- To print the calculated state value function for the Gridworld environment, to implement
and calculate the state value function in Python for a simple reinforcement learning problem.

Bellman's Equation
Bellman's equation is a fundamental concept in dynamic programming and reinforcement
learning. It expresses the relationship between the value of a state or state-action pair and the
value of its successor states or state-action pairs. Bellman's equation has two forms: the
Bellman Expectation Equation and the Bellman Optimality Equation.
The Bellman Expectation Equation for the state value function (V(s)) is given by:
\[ V(s) = \sum_{a} \pi(a|s) \sum_{s'} P(s'|s, a) [R(s,a,s') + \gamma V(s')] \]
Where:
- \( V(s) \) is the value of state \( s \).
- \( \pi(a|s) \) is the policy, representing the probability of taking action \( a \) in state \( s \).

____________________________________________________________________________________
133 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

- \( P(s'|s, a) \) is the transition probability, representing the probability of transitioning to


state \( s' \) given that action \( a \) is taken in state \( s \).
- \( R(s,a,s') \) is the reward obtained after transitioning from state \( s \) to state \( s' \) by
taking action \( a \).
- \( \gamma \) is the discount factor, which discounts future rewards.

Here is an example of implementing the Bellman Expectation Equation for the state value
function in python:

#python
import numpy as np
# Define transition probabilities
P = [Link]([
[[0.8, 0.2], [0.1, 0.9]], # Transition probabilities for action 0 (left)
[[0.7, 0.3], [0.2, 0.8]] # Transition probabilities for action 1 (right)
])# Define rewards
R = [Link]([
[[1, 0], [-1, 0]], # Rewards for action 0 (left)
[[0, -1], [0, 1]] # Rewards for action 1 (right)])
# Define policy (uniform random policy)
pi = [Link]([[0.5, 0.5], [0.5, 0.5]])
# Define discount factor
gamma = 0.9
# Initialize state value function
V = [Link](2)
# Bellman Expectation Equation
for s in range(2):
V_s = 0
for a in range(2):
for s_prime in range(2):
V_s += pi[s, a] * P[a, s, s_prime] * (R[a, s, s_prime] + gamma * V[s_prime])
V[s] = V_s
____________________________________________________________________________________
134 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

print("State Value Function:")


print(V)

Here is an example,
- To define transition probabilities (\( P \)) and rewards (\( R \)) for a simple Markov Decision
Process (MDP) with two states and two actions.
- To define a uniform random policy (\( \pi \)).
- To define a discount factor (\( \gamma \)).
- To initialize the state value function (\( V \)).
- To iterate over each state and apply the Bellman Expectation Equation to update the value
of each state based on the expected future rewards to implement the Bellman Expectation
Equation for the state value function in python.
In R, Bellman's equation can be implemented for various reinforcement learning tasks, such
as solving Markov decision processes (MDPs) or training agents in environments using
algorithms like Q-learning or value iteration.
Here is a conceptual example of implementing Bellman's equation for Q-learning in R:

# Define the Q-learning function using Bellman's equation


q_learning <- function(env, num_episodes, learning_rate, discount_factor, epsilon) {
# Initialize Q-table with zeros
q_table <- matrix(0, nrow = nrow(env$states), ncol = ncol(env$actions))
# Iterate over episodes
for (episode in 1:num_episodes) {
# Reset environment to initial state
state <- env$reset()
# Iterate until the episode terminates
while (!env$is_done()) {
# Choose action using epsilon-greedy policy
if (runif(1) < epsilon) {
action <- sample(env$actions, 1)
} else {
action <- [Link](q_table[state, ])
}
____________________________________________________________________________________
135 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

# Take action and observe the next state and reward


next_state <- env$step(action)
reward <- env$get_reward()
# Update Q-value using Bellman's equation
q_table[state, action] <- q_table[state, action] +
learning_rate * (reward + discount_factor * max(q_table[next_state, ]) - q_table[state,
action])
# Move to the next state
state <- next_state
}
}
# Return learned Q-table
return(q_table)
}
# Define environment (e.g., grid world)
environment <- list(
states = c(1, 2, 3, 4),
actions = c("up", "down", "left", "right"),
reset = function() {
return(sample(environment$states, 1))
},step = function(action) {
# Implement transition function based on current state and action
# Return next state
},is_done = function() {
# Implement termination condition
}, get_reward = function() {
# Implement reward function based on current state and action
# Return reward})
# Set hyperparameters
num_episodes <- 1000
____________________________________________________________________________________
136 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

learning_rate <- 0.1


discount_factor <- 0.9

epsilon <- 0.1


# Run Q-learning
q_table <- q_learning(environment, num_episodes, learning_rate, discount_factor, epsilon)

In this example,
- To define a Q-learning function `q_learning` that implements the Q-learning algorithm
using Bellman's equation. Inside the function, iterate over episodes, take actions, observe
rewards, and update Q-values based on Bellman's equation.
- To define an example environment (e.g., grid world) with states, actions, a transition
function, a termination condition, and a reward function.
- To set hyperparameters such as the number of episodes, learning rate, discount factor, and
exploration rate (epsilon). Run the Q-learning algorithm on the defined environment to learn
the optimal Q-values.
Bellman's equation can be used in R to implement reinforcement learning algorithms like Q-
learning to solve optimization problems in dynamic environments. Neural networks have
found widespread application across various industries due to their ability to learn complex
patterns and relationships from data.

____________________________________________________________________________________
137 | Machine Learning Methods Using Python and R- II | FG
----------------------------------------------------------- TISS SSE ----------------------------------------------------------

References
1. Bishop, C. M. (2006). Pattern Recognition and Machine Learning. Cambridge: Springer.
Retrieved from [Link]
2. Geron, A. (2019). Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFLow.
United State of America: O’Reilly Media, Inc., 1005 Gravenstein Highway North,
Sebastopol, CA 95472.
3. Grolemund, H. W. (2017). R for Data Science. United State of America: O'Reilly,
Sebastopol, cop.
4. Guido, A. C. (2016). Introduction to Machine Learning with Python: A Guide for Data
Scientists. United State of America: Sebastopol, CA: O'Reilly Media, Inc.
5. Ian H. Witten, E. F. (2011). Data Mining: Practical Machine Learning Tools and Techniques.
Burlington, Massachusetts: Morgan Kaufmann. doi:[Link]
19715-5
6. Murphy, K. P. (2012). Machine Learning: A Probabilistic Perspective. Massachusetts
London: The MIT Press Cambridge, England. Retrieved from
[Link]
5nh9osgl8qq0
7. Raschka, S. (2017). Python Machine Learning. Birmingham- Mumbai: Packt Publishing.

____________________________________________________________________________________
138 | Machine Learning Methods Using Python and R- II | FG

You might also like