ML Lab
ML Lab
Course Objectives:
This course will enable students to learn and understand different Data sets in implementing the
machinelearning algorithms.
Experiment-2:
For a given set of training data examples stored in a .CSV file, implement and demonstrate the
Candidate-Elimination algorithm to output a description of the set of all hypotheses consistent
with the training examples.
Experiment-3:
Write a program to demonstrate the working of the decision tree based ID3 algorithm. Use an
appropriatedata set for building the decision tree and apply this knowledge to classify a new
sample.
Experiment-4:
Exercises to solve the real-world problems using the following machine learning methods: a)
LinearRegression b) Logistic Regression c) Binary Classifier
Experiment-5: Develop a program for Bias, Variance, Remove duplicates , Cross Validation
Experiment-6: Write a program to implement Categorical Encoding, One-hot Encoding
Experiment-7:
Build an Artificial Neural Network by implementing the Back propagation algorithm and test
the sameusing appropriate data sets.
Experiment-8:
Write a program to implement k-Nearest Neighbor algorithm to classify the iris data set. Print
bothcorrect and wrong predictions.
2| P a g e
Machine Learning with Python Lab
Experiment-9: Implement the non-parametric Locally Weighted Regression algorithm in order
to fit data points. Select appropriate data set for your experiment and draw graphs.
Experiment-10:
Assuming a set of documents that need to be classified, use the naïve Bayesian Classifier model
to perform this task. Built-in Java classes/API can be used to write the program. Calculate the
accuracy, precision, and recall for your data set.
Experiment-11: Apply EM algorithm to cluster a Heart Disease Data Set. Use the same data set
for clustering using k-Means algorithm. Compare the results of these two algorithms and
comment on the quality of clustering. You can add Java/Python ML library classes/API in the
program.
Experiment-13:
Write a Python program to construct a Bayesian network considering medical data. Use this
model todemonstrate the diagnosis of heart patients using standard Heart Disease Data Set
Experiment-14:
Write a program to Implement Support Vector Machines and Principle Component Analysis
Experiment-15:
Write a program to Implement Principle Component Analysis
Course Outcomes (Cos): At the end of the course, student will be able to
• Implement procedures for the machine learning algorithms
• Design and Develop Python programs for various Learning algorithms
• Apply appropriate data sets to the Machine Learning algorithms
• Develop Machine Learning algorithms to solve real world problems
3| P a g e
Machine Learning with Python Lab
LIST OF EXPERIMENTS
MACHINE LEARNING WITH PYTHON LAB
Experiment-1:
Implement and demonstrate the FIND-S algorithm for finding the most specific hypothesis based on a
given set of training data samples. Read the training data from a .CSV file.
Experiment-2:
For a given set of training data examples stored in a .CSV file, implement and demonstrate the
Candidate-Elimination algorithm to output a description of the set of all hypotheses consistent with the
training examples.
Experiment-3:
Write a program to demonstrate the working of the decision tree based ID3 algorithm. Use an
appropriate data set for building the decision tree and apply this knowledge to classify a new sample.
Experiment-4:
Exercises to solve the real-world problems using the following machine learning methods: a) Linear
Regression b) Logistic Regression c) Binary Classifier
Experiment-5:
Experiment-6:
Experiment-7:
Build an Artificial Neural Network by implementing the Back propagation algorithm and test the same
using appropriate data sets.
Experiment-8:
Write a program to implement k-Nearest Neighbor algorithm to classify the iris data set. Print both
correct and wrong predictions.
4| P a g e
Machine Learning with Python Lab
Experiment-9:
Implement the non-parametric Locally Weighted Regression algorithm in order to fit data points. Select
appropriate data set for your experiment and draw graphs.
Experiment-10:
Assuming a set of documents that need to be classified, use the naïve Bayesian Classifier model to
perform this task. Built-in Java classes/API can be used to write the program. Calculate the accuracy,
precision, and recall for your data set.
Experiment-11:
Apply EM algorithm to cluster a Heart Disease Data Set. Use the same data set for clustering using k-
Means algorithm. Compare the results of these two algorithms and comment on the quality of
clustering.
Experiment-12:
Experiment-13:
Write a Python program to construct a Bayesian network considering medical data. Use this model to
demonstrate the diagnosis of heart patients using standard Heart Disease Data Set.
Experiment-14:
Write a program to Implement Support Vector Machines and Principle Component Analysis.
Experiment-15:
5| P a g e
Machine Learning with Python Lab
Experiment-1
1. AIM: Implement and demonstrate the FIND-S algorithm for finding the most specific
hypothesis based on a given set of training data samples. Read the training datafrom a
.CSV file.
Source Code:
import csv
for i in your_list:
print(i)
if i[-1] == "True":
j=0
for x in i:
if x != "True":
if x != h[0][j] and h[0][j] == '0':
h[0][j] = x
elif x != h[0][j] and h[0][j] != '0':
h[0][j] = '?'
else:
pass
j=j+1
print("Most specific hypothesis is")
print(h)
Output
6| P a g e
Machine Learning with Python Lab
Experiment-2
2. AIM: For a given set of training data examples stored in a .CSV file, implement and
demonstrate the Candidate-Elimination algorithm to output a description of the set of all
hypotheses consistent with the training examples.
Source Code:
class Holder:
factors= {} #Initialize an empty dictionary
attributes = () #declaration of dictionaries parameters with an arbitrary length
'''
Constructor of class Holder holding two parameters,
self refers to the instance of the class
'''
def init (self, attr): #
self. Attributes =
attrfor i in attr:
self. Factors[i]=[]
class CandidateElimination:
Positive= {} #Initialize positive empty dictionary
Negative={} #Initialize negative empty dictionary
def run_algorithm(self):
'''
Initialize the specific and general boundaries, and loop the dataset against the
algorithm
'''
G = [Link]()
S = [Link]()
'''
Programmatically populate list in the iterating variable trial_set
'''
count=0
for trial_set in [Link]:
if self.is_positive(trial_set): #if trial set/example consists of positive examples
7| P a g e
Machine Learning with Python Lab
G = self.remove_inconsistent_G(G,trial_set[0]) #remove inconsitent data from
the general boundary
print (S)
print (G)
def initializeS(self):
''' Initialize the specific boundary '''
S = tuple(['-' for factor in range(self.num_factors)]) #6 constraints in the vector
return [S]
def initializeG(self):
''' Initialize the general boundary '''
G = tuple(['?' for factor in range(self.num_factors)]) # 6 constraints in the vector
return [G]
def is_positive(self,trial_set):
''' Check if a given training trial_set is positive '''
if trial_set[1] == 'Y':
8| P a g e
Machine Learning with Python Lab
return True
elif trial_set[1] == 'N':
return False
else:
raise TypeError("invalid target value")
def match_factor(self,value1,value2):
''' Check for the factors values match,
necessary while checking the consistency of
training trial_set with the hypothesis '''
if value1 == '?' or value2 == '?':
return True
elif value1 == value2 :
return True
return False
def consistent(self,hypothesis,instance):
''' Check whether the instance is part of the hypothesis '''
for i,factor in enumerate(hypothesis):
if not self.match_factor(factor,instance[i]):
return False
return True
def remove_inconsistent_G(self,hypotheses,instance):
''' For a positive trial_set, the hypotheses in G
inconsistent with it should be removed '''
G_new = hypotheses[:]
for g in hypotheses:
if not [Link](g,instance):
G_new.remove(g)
return G_new
def remove_inconsistent_S(self,hypotheses,instance):
''' For a negative trial_set, the hypotheses in S
inconsistent with it should be removed '''
S_new = hypotheses[:]
for s in hypotheses:
if [Link](s,instance):
S_new.remove(s)
return S_new
def remove_more_general(self,hypotheses):
''' After generalizing S for a positive trial_set, the hypothesis in S
general than others in S should be removed '''
S_new = hypotheses[:]
for old in hypotheses:
9| P a g e
Machine Learning with Python Lab
for new in S_new:
if old!=new and self.more_general(new,old):
S_new.remove[new]
return S_new
def remove_more_specific(self,hypotheses):
''' After specializing G for a negative trial_set, the hypothesis in G
specific than others in G should be removed '''
G_new = hypotheses[:]
for old in hypotheses:
for new in G_new:
if old!=new and self.more_specific(new,old):
G_new.remove[new]
return G_new
def generalize_inconsistent_S(self,hypothesis,instance):
''' When a inconsistent hypothesis for positive trial_set is seen in the specific
boundary S,
it should be generalized to be consistent with the trial_set ... we will get one
hypothesis'''
hypo = list(hypothesis) # convert tuple to list for mutability
for i,factor in enumerate(hypo):
if factor == '-':
hypo[i] = instance[i]
elif not self.match_factor(factor,instance[i]):
hypo[i] = '?'
generalization = tuple(hypo) # convert list back to tuple for immutability
return generalization
def specialize_inconsistent_G(self,hypothesis,instance):
''' When a inconsistent hypothesis for negative trial_set is seen in the general
boundary G
should be specialized to be consistent with the trial_set.. we will get a set of
hypotheses '''
specializations = []
hypo = list(hypothesis) # convert tuple to list for mutability
for i,factor in enumerate(hypo):
if factor == '?':
values = [Link][[Link][i]]
for j in values:
if instance[i] != j:
hyp=hypo[:]
hyp[i]=j
hyp=tuple(hyp) # convert list back to tuple for immutability
[Link](hyp)
return specializations
10| P a g e
Machine Learning with Python Lab
def get_general(self,generalization,G):
''' Checks if there is more general hypothesis in G
for a generalization of inconsistent hypothesis in S
in case of positive trial_set and returns valid generalization '''
for g in G:
if self.more_general(g,generalization):
return generalization
return None
def get_specific(self,specializations,S):
''' Checks if there is more specific hypothesis in S
for each of hypothesis in specializations of an
inconsistent hypothesis in G in case of negative trial_set
and return the valid specializations'''
valid_specializations = []
for hypo in specializations:
for s in S:
if self.more_specific(s,hypo) or s==[Link]()[0]:
valid_specializations.append(hypo)
return valid_specializations
def exists_general(self,hypothesis,G):
'''Used to check if there exists a more general hypothesis in
general boundary for version space'''
for g in G:
if self.more_general(g,hypothesis):
return True
return False
def exists_specific(self,hypothesis,S):
'''Used to check if there exists a more specific hypothesis in
general boundary for version space'''
for s in S:
if self.more_specific(s,hypothesis):
return True
return False
def more_general(self,hyp1,hyp2):
''' Check whether hyp1 is more general than hyp2 '''
hyp = zip(hyp1,hyp2)
for i,j in hyp:
if i == '?':
continue
11| P a g e
Machine Learning with Python Lab
elif j == '?':
if i != '?':
return False
elif i != j:
return False
else:
continue
return True
def more_specific(self,hyp1,hyp2):
''' hyp1 more specific than hyp2 is
equivalent to hyp2 being more general than hyp1 '''
return self.more_general(hyp2,hyp1)
dataset=[(('sunny','warm','normal','strong','warm','same'),'Y'),(('sunny','warm','high','stron
g','warm','same'),'Y'),(('rainy','cold','high','strong','warm','change'),'N'),(('sunny','warm','hi
gh','strong','cool','change'),'Y')]
attributes =('Sky','Temp','Humidity','Wind','Water','Forecast')
f = Holder(attributes)
f.add_values('Sky',('sunny','rainy','cloudy')) #sky can be sunny rainy or cloudy
f.add_values('Temp',('cold','warm')) #Temp can be sunny cold or warm
f.add_values('Humidity',('normal','high')) #Humidity can be normal or high
f.add_values('Wind',('weak','strong')) #wind can be weak or strong
f.add_values('Water',('warm','cold')) #water can be warm or cold
f.add_values('Forecast',('same','change')) #Forecast can be same or change
a = CandidateElimination(dataset,f) #pass the dataset to the algorithm class and call the
run algoritm method
a.run_algorithm()
Output
12| P a g e
Machine Learning with Python Lab
Experiment-3
3. AIM: Write a program to demonstrate the working of the decision tree based ID3
[Link] an appropriate data set for building the decision tree and apply this
knowledge to classify a new sample.
Source Code:
import numpy as np
import math
from data_loader import read_data
class Node:
def init (self, attribute):
[Link] = attribute
[Link] = []
[Link] = ""
for x in range([Link][0]):
for y in range([Link][0]):
if data[y, col] == items[x]:
count[x] += 1
for x in range([Link][0]):
dict[items[x]] = [Link]((int(count[x]), [Link][1]), dtype="|S32")
pos = 0
for y in range([Link][0]):
if data[y, col] == items[x]:
dict[items[x]][pos] = data[y]
pos += 1
if delete:
dict[items[x]] = [Link](dict[items[x]], col, 1)
def entropy(S):
items = [Link](S)
if [Link] == 1:
13| P a g e
Machine Learning with Python Lab
return 0
for x in range([Link][0]):
total_size = [Link][0]
entropies = [Link](([Link][0], 1))
intrinsic = [Link](([Link][0], 1))
for x in range([Link][0]):
ratio = dict[items[x]].shape[0]/(total_size * 1.0)
entropies[x] = ratio * entropy(dict[items[x]][:, -1])
intrinsic[x] = ratio * [Link](ratio, 2)
for x in range([Link][0]):
total_entropy -= entropies[x]
return total_entropy / iv
split = [Link](gains)
node = Node(metadata[split])
14| P a g e
Machine Learning with Python Lab
metadata = [Link](metadata, split, 0)
items, dict = subtables(data, split, delete=True)
for x in range([Link][0]):
child = create_node(dict[items[x]], metadata)
[Link]((items[x], child))
return node
def empty(size):
s = ""
for x in range(size):
s += " "
return s
print(empty(level), [Link])
Data_loader.py
import csv
def read_data(filename):
with open(filename, 'r') as csvfile:
datareader = [Link](csvfile, delimiter=',')
headers = next(datareader)
metadata = []
traindata = []
for name in headers:
[Link](name)
for row in datareader:
[Link](row)
return (metadata, traindata)
15| P a g e
Machine Learning with Python Lab
[Link]
outlook,temperature,humidity,wind,
answer sunny,hot,high,weak,no
sunny,hot,high,strong,no
overcast,hot,high,weak,yes
rain,mild,high,weak,yes
rain,cool,normal,weak,yes
rain,cool,normal,strong,no
overcast,cool,normal,strong,yes
sunny,mild,high,weak,no
sunny,cool,normal,weak,yes
rain,mild,normal,weak,yes
sunny,mild,normal,strong,yes
overcast,mild,high,strong,yes
overcast,hot,normal,weak,yes
rain,mild,high,strong,no
Output
outlook
overcast
b'yes'
rain
wind
b'strong'
b'no'
b'weak'
b'yes'
sunny
humidity
b'high'
b'no'
b'normal'
b'yes
16| P a g e
Machine Learning with Python Lab
Experiment-4
4. AIM: Exercises to solve the real-world problems using the following machine learning methods:
a) Linear Regression b) Logistic Regression c) Binary Classifier.
a) Linear Regression
Source Code:
import numpy as np
X = [Link](([2, 9], [1, 5], [3, 6]), dtype=float)
y = [Link](([92], [86], [89]), dtype=float)
X = X/[Link](X,axis=0) # maximum of X array longitudinally
y = y/100
#Sigmoid Function
def sigmoid (x):
return 1/(1 + [Link](-x))
#Variable initialization
epoch=7000 #Setting training iterations
lr=0.1 #Setting learning rate
inputlayer_neurons = 2 #number of features in data set
hiddenlayer_neurons = 3 #number of hidden layers neurons
output_neurons = 1 #number of neurons at output layer
#weight and bias initialization
wh=[Link](size=(inputlayer_neurons,hiddenlayer_neurons))
bh=[Link](size=(1,hiddenlayer_neurons))
wout=[Link](size=(hiddenlayer_neurons,output_neurons))
bout=[Link](size=(1,output_neurons))
#draws a random range of numbers uniformly of dim x*y
for i in range(epoch):
#Forward Propogation
hinp1=[Link](X,wh)
hinp=hinp1 + bh
hlayer_act = sigmoid(hinp)
outinp1=[Link](hlayer_act,wout)
outinp= outinp1+ bout
output = sigmoid(outinp)
#Backpropagation
EO = y-output
outgrad = derivatives_sigmoid(output)
d_output = EO* outgrad
EH = d_output.dot(wout.T)
hiddengrad = derivatives_sigmoid(hlayer_act)#how much hidden layer wts
17| P a g e
Machine Learning with Python Lab
contributed to error
d_hiddenlayer = EH * hiddengrad
wout += hlayer_act.[Link](d_output) *lr# dotproduct of nextlayererror and
currentlayerop
# bout += [Link](d_output, axis=0,keepdims=True) *lr
wh += [Link](d_hiddenlayer) *lr
#bh += [Link](d_hiddenlayer, axis=0,keepdims=True) *lr
print("Input: \n" + str(X))
print("Actual Output: \n" + str(y))
print("Predicted Output: \n" ,output)
output
Input:
[[ 0.66666667 1. ]
[ 0.33333333 0.55555556]
[ 1. 0.66666667]]
Actual Output:
[[ 0.92]
[ 0.86]
[ 0.89]]
Predicted Output:
[[ 0.89559591]
[ 0.88142069]
[ 0.8928407 ]]
18| P a g e
Machine Learning with Python Lab
b) Logistic Regression
Source Code:
x = [5,7,8,7,2,17,2,9,4,11,12,9,6]
y = [99,86,87,88,111,86,103,87,94,78,77,85,86]
def myfunc(x):
return slope * x + intercept
[Link](x, y)
[Link](x, mymodel)
[Link]()
Output:
c) Binary Classifier
Source Code:
import sklearn as sk
import pandas as pd
import pandas as pd
import os
rom sklearn.linear_model import LogisticRegression
from sklearn import svm
from [Link] import RandomForestClassifier
from sklearn.neural_network import MLPClassifier
[Link]('/Users/stevenhurwitt/Documents/Blog/Classification')
heart = pd.read_csv('[Link]', sep=',', header=0)
[Link]()
y = [Link][:,9]
19| P a g e
Machine Learning with Python Lab
X = [Link][:,:9]
vowel_train = pd.read_csv('[Link]', sep=',', header=0)
vowel_test = pd.read_csv('[Link]', sep=',', header=0)
vowel_train.head()
y_tr = vowel_train.iloc[:,0]
X_tr = vowel_train.iloc[:,1:]
y_test = vowel_test.iloc[:,0]
X_test = vowel_test.iloc[:,1:]
OutPut:
20| P a g e
Machine Learning with Python Lab
Experiment-5
5. AIM: Develop a program for Bias, Variance, remove duplicates, Cross Validation
Source Code:
Output:
MSE: 22.487
Bias: 20.726
Variance: 1.761
21| P a g e
Machine Learning with Python Lab
Experiment-6
Source Code:
import numpy as np
import pandas as pd
print(data['Gender'].unique())
print(data['Remarks'].unique())
Output:
# importing libraries
import pandas as pd
import numpy as np
from [Link] import OneHotEncoder
# Retrieving data
data = pd.read_csv('Employee_data.csv')
22| P a g e
Machine Learning with Python Lab
data[['Gen_new', 'Rem_new']]).toarray())
# Merge with main
New_df = [Link](enc_data)
print(New_df)
Output:
23| P a g e
Machine Learning with Python Lab
Experiment-7
7. AIM: Build an Artificial Neural Network by implementing the Back propagation algorithm
and test the same using appropriate datasets.
Source Code:
Import numpy as np
X =[Link](([2,9],[1,5],[3,6]),dtype=float)
y=[Link](([92],[86],[89]),dtype=float)
X=X/[Link](X,axis=0) #maximumofXarraylongitudinallyy= y/100
#Sigmoid Functiondefsigmoid(x):
return1/(1+[Link](-x))
#DerivativeofSigmoidFunctiondefderivatives_sigmoid(x):
returnx* (1-x)
#Variableinitialization
epoch=7000#Settingtrainingiterationslr=0.1#Settinglearning rate
inputlayer_neurons = 2 #number of features in data
sethiddenlayer_neurons=3#numberofhiddenlayersneuronsoutput_neurons = 1 #number of neurons at
output layer#weightand biasinitialization
wh=[Link](size=(inputlayer_neurons,hiddenlayer_neurons))
bh=[Link](size=(1,hiddenlayer_neurons))wout=[Link](size=(hiddenlayer_ne
urons,output_neurons))bout=[Link](size=(1,output_neurons))
#drawsarandomrangeofnumbersuniformlyofdimx*yforiin range(epoch):
#BackpropagationEO=y-output
outgrad=derivatives_sigmoid(output)d_output= EO* outgrad
EH=d_output.dot(wout.T)
hiddengrad=derivatives_sigmoid(hlayer_act)#how muchhiddenlayerwtscontributedto error
d_hiddenlayer=EH*hiddengrad
wout+=hlayer_act.[Link](d_output)*lr#dotproductofnextlayererrorandcurrentlayerop
# bout += [Link](d_output, axis=0,keepdims=True) *lrwh+=[Link](d_hiddenlayer) *lr
#bh+=[Link](d_hiddenlayer,axis=0,keepdims=True)*lrprint("Input:\n"+ str(X))
print("Actual Output: \n" + str(y))print("PredictedOutput:\n",output)
24| P a g e
Machine Learning with Python Lab
Output:
Input:
[[ 0.666666671. ]
[0.333333330.55555556]
[1. 0.66666667]]
Actual Output:[[ 0.92]
[0.86]
[0.89]]
Predicted Output: [[0.89559591]
[0.88142069]
[0.8928407]]
25| P a g e
Machine Learning with Python Lab
Experiment-8
8. AIM: Write a program to implement k-Nearest Neighbor algorithm to classify the iris data set.
Print both correct and wrong predictions.
Source Code:
import numpy as np
import pandas as pd
from [Link] import KNeighborsClassifier
from sklearn.model_selection import train_test_split
from sklearn import metrics
Output:
26| P a g e
Machine Learning with Python Lab
27| P a g e
Machine Learning with Python Lab
Experiment-9
9. AIM: Implement the non-parametric Locally Weighted Regression algorithm in order to fit
data points. Select appropriate data set for your experiment and draw graphs.
Source Code:
29| P a g e
Machine Learning with Python Lab
Experiment-10
10. AIM: Assuming a set of documents that need to be classified, use the naïve Bayesian Classifier
model to perform this task. Built-in Java classes/API can be used to write the program. Calculate
the accuracy, precision, and recall for your dataset.
Source Code:
import pandas as pd
msg=pd.read_csv('[Link]',names=['message','label'])
print('The dimensions of the dataset',[Link])msg['labelnum']=[Link]({'pos':1,'neg':0})
X=[Link]=[Link](X)
print(y)
#splittingthedataset intotrainandtestdata
fromsklearn.model_selectionimporttrain_test_splitxtrain,xtest,ytrain,ytest=train_test_split(X,y)
print([Link])
print([Link])print([Link])print([Link])
#outputofcountvectoriserisasparsematrix
fromsklearn.feature_extraction.textimportCountVectorizercount_vect= CountVectorizer()
xtrain_dtm=count_vect.fit_transform(xtrain)xtest_dtm=count_vect.transform(xtest)print(count_vect.get
_feature_names())
df=[Link](xtrain_dtm.toarray(), columns=count_vect.get_feature_names())
print(df)#tabularrepresentation
print(xtrain_dtm) #sparsematrixrepresentation
#TrainingNaiveBayes(NB)[Link].naive_bayes importMultinomialNB
clf=MultinomialNB().fit(xtrain_dtm,ytrain)predicted=[Link](xtest_dtm)
'''docs_new=['Ilikethisplace','Mybossisnotmysaviour']
X_new_counts=count_vect.transform(docs_new)
predictednew =[Link](X_new_counts)
fordoc,categoryinzip(docs_new,predictednew):
print('%s->%s'%(doc, [Link][category]))'''
Output:
['about','am','amazing','an','and','awesome','beers','best','boss','can','deal',
'do','enemy','feel','fun','good','have','horrible','house','is','like','love','my',
'not','of','place','restaurant','sandwich','sick','stuff','these','this','tired','to',
'today','tomorrow','very','view','we', 'went','what','will','with','work']
0 10 0 0 0 01 0 0 0 ... 0
1 00 0 0 0 00 1 0 0 ... 0
2 00 1 1 0 0 0 0 0 0 ... 0
3 00 0 0 0 00 0 0 0 ... 1
4 00 0 0 0 00 0 0 0 ... 0
5 01 0 01 0 0 0 0 0 ... 0
6 00 0 0 0 00 0 0 1 ... 0
7 00 0 0 0 00 0 0 0 ... 0
8 01 0 0 0 00 0 0 0 ... 0
9 00 0 1 0 10 0 0 0 ... 0
10 0 0 0 0 0 0 0 0 0 0 ... 0
11 0 0 0 0 0 0 0 0 1 0 ... 0
12 0 0 0 1 0 1 0 0 0 0 ... 0
1 0 0 0 0 00 0 0 1
2 0 0 0 0 00 0 0 0
3 0 0 0 0 10 0 0 0
4 0 0 0 0 00 0 0 0
5 0 0 0 0 00 0 0 0
6 0 0 0 0 00 0 1 0
7 1 0 0 1 00 1 0 0
8 0 0 0 0 00 0 0 0
31| P a g e
Machine Learning with Python Lab
Experiment-11
11. AIM: Apply EM algorithm to cluster a set of data stored in a. CSV file. Use the same data set
for clustering using k-Means algorithm. Compare the results of the set two algorithms and
comment on the quality of clustering. You can add Java/Python ML library classes/API in the
program.
Source Code:
import numpy as np
import [Link] as plt
[Link].samples_generatorimportmake_blobsX, y_true = make_blobs(n_samples=100,
centers =4,Cluster_std=0.60,random_state=0)
X =X[:,::-1]
#flipaxesforbetterplotting
from [Link]
def draw_ellipse(position, covariance, ax=None,
**kwargs);“””Drawanellipsewithagivenpositionandcovariance”””
Ax=[Link]()
#Convertcovariancetoprincipalaxes
[Link] ==(2,2):
U,s,Vt=[Link](covariance)
Angle=[Link](np.arctan2(U[1,0],U[0,0]))Width,height= 2 * [Link](s)
else:
angle=0
width,height=2*[Link](covariance)
#DrawtheEllipse
fornsiginrange(1,4):
ax.add_patch(Ellipse(position,nsig*width,nsig*height,angle,**kwargs))
defplot_gmm(gmm,X,label=True,ax=None):ax= ax or [Link]()
labels=[Link](X).predict(X)iflabel:
[Link](X[:,0],x[:,1],c=labels,s=40,cmap=‟viridis‟,zorder=2)
else:
[Link](X[:,0],x[:,1],s=40,zorder=2)[Link](„equal‟)
w_factor=0.2/gmm.weights_.max()
32| P a g e
Machine Learning with Python Lab
forpos,covar,winzip(gmm.means_,gmm.covariances_,gmm.weights_):
draw_ellipse(pos,covar, alpha=w*w_factor)
gmm=GaussianMixture(n_components=4,random_state=42)plot_gmm(gmm,X)
gmm=GaussianMixture(n_components=4,covariance_type=‟full‟,random_state=42)
plot_gmm (gmm, X)
Output :
[[1,0, 0, 0]
[0,0,1,0]
[1,0,0,0]
[1,0,0,0]
[1,0, 0, 0]]
33| P a g e
Machine Learning with Python Lab
34| P a g e
Machine Learning with Python Lab
Experiment-12
12. AIM: Exploratory Data Analysis for Classification using Pandas or Matplotlib
Source code:
Import pandas as pd
Import matplotlib. pyplot as plt
DF =pd. read_csv("[Link]
Df =pd.read_csv("[Link] / Rdatasets / csv / car / [Link]")
[Link]()
/ fivethirtyeight / data / master / airline-safety / [Link]")
y =list([Link])
plt. Boxplot(y)
plt. show()
DF["education"].value_counts()
[Link](['education', 'vote']).mean()
From [Link] importf_oneway
# Perform ANOVA
f_statistic, p_value =f_oneway(group1, group2, group3)
35| P a g e
Machine Learning with Python Lab
Output:
36| P a g e
Machine Learning with Python Lab
Experiment-13
13. AIM: Write a program to construct a Bayesian network considering medical data. Use this
model to demonstrate the diagnosis of heart patients using standard Heart Disease Data Set. You
can use Java/Python ML library classes/API
Theory:
A Bayesian network is a directed a cyclic graph in which each edge corresponds to a conditional
dependency,
and each node corresponds to a unique random variable.
Bayesiannetworkconsistsoftwomajorparts: adirectedacyclicgraphandasetofconditionalprobability
distributions
• The directed acyclic graph isa setoff random variables represented by nodes.
• Theconditionalprobabilitydistributionofanode(randomvariable)isdefinedforeverypossible
outcomeoftheprecedingcausalnode(s).
Fig: Directedacyclicgraphrepresentingtwoindependentpossiblecausesofacomputerfailure.
Thegoalistocalculatetheposteriorconditionalprobabilitydistributionofeachofthepossibleunobserve
dcausesgiventhe observed evidence, i.e. [Cause|Evidence].
DataSet:
37| P a g e
Machine Learning with Python Lab
Database: 0 1 2 3 4 Total
Cleveland: 164 55 36 35 13 303
Attribute Information:
1. age: ageinyears
2. sex: sex (1 =male;0=female)
3. cp: chest pain type
• Value1: typicalangina
• Value2:atypicalangina
• Value3:non-anginalpain
• Value4:asymptomatic
4. trestbps: restingbloodpressure(inmmHgonadmissiontothehospital)
5. chol:serumcholestoralinmg/dl
6. fbs:(fastingblood sugar >120 mg/dl)(1=true; 0 =false)
7. restecg: restingelectrocardiographicresults
• Value0: normal
• Value1: havingST-Twaveabnormality
(Twaveinversionsand/orSTelevationordepressionof>0.05mV)
• Value2: showingprobableordefiniteleftventricularhypertrophybyEstes’criteria
8. thalach: maximumheartrateachieved
9. exang: exercise induced angina (1 =yes;0=no)
10. oldpeak=STdepressioninduced byexerciserelative torest
11. slope: the slope of thepeakexercise STsegment
• Value 1: upsloping
• Value 2: flat
• Value3: down sloping
12. ca=number of majorvessels(0-3) colored byflourosopy
13. thal:3=normal;6=fixed defect; 7=reversable defect
Heart disease:
Itisintegervaluedfrom0(nopresence)[Link](angiographicdiseasestatus)
38| P a g e
Machine Learning with Python Lab
ag se cp trestbps cho fb restec thalac exan oldpea slop c thal Heartdisea
e x l s g h g k e a se
63 1 1 145 233 1 2 150 0 2.3 3 0 6 0
67 1 4 160 286 0 2 108 1 1.5 2 3 3 2
67 1 4 120 229 0 2 129 1 2.6 2 2 7 1
41 0 2 130 204 0 2 172 0 1.4 1 0 3 0
62 0 4 140 268 0 2 160 0 3.6 3 2 3 3
60 1 4 130 206 0 2 132 1 2.4 2 2 7 4
Program:
import numpy as np
import csv
import pandas aspd
frompgmpy. modelsimportBayesianModel
from pgmpy.
[Link]
#readClevelandHeartDiseasedata heartDisease=pd.read_csv('[Link]')
heartDisease=heartDisease. Replace (‘?’, [Link])
#displaythedata
print('Few examples from the dataset are given below')
print([Link]())
#LearningCPDsusingMaximumLikelihoodEstimators
print ('\n Learning CPD using Maximum likelihood estimators')[Link](heartDisease,
estimator=MaximumLikelihoodEstimator)
#InferencingwithBayesianNetwork
print('\n Inferencing with Bayesian Network:')
HeartDisease_infer=VariableElimination(model)
#computingtheProbabilityofHeartDiseasegivenAge
print('\n 1. Probability of HeartDisease given
Age=30')q=HeartDisease_infer.query(variables=['heartdisease'],evidence
={'age':28})
print(q['heartdisease'])
Output:
Fewexamplesfromthedatasetaregivenbelow
agesexcptrestbps ...slopecathalheartdisease0 63 1 1 145
... 3 0 6 0
1 67 1 4 160 ...2 3 3 2
2 67 1 4 120 ...2 2 7 1
3 37 1 3 130 ... 3 0 3 0
4 41 0 2 130 ... 1 0 3 0
[5rowsx14columns]
EstimatorsInferencingwithBayesianNetwork:
40| P a g e
Machine Learning with Python Lab
2. Probability of Heart Disease given cholesterol=100
╒════════════════╤═════════════════════╕
│heartdisease │ phi(heartdisease)│
╞════════════════╪═════════════════════╡
│heartdisease_0│ 0.5400│
├────────────────┼─────────────────────┤
│heartdisease_1│ 0.1533│
├────────────────┼─────────────────────┤
│heartdisease_2│ 0.1303│
├────────────────┼─────────────────────┤
│heartdisease_3│ 0.1259│
├────────────────┼─────────────────────┤
│heartdisease_4│ 0.0506│
╘════════════════╧═════════════════════╛
41| P a g e
Machine Learning with Python Lab
Experiment-14
14. AIM: Write a program to Implement Support Vector Machines and Principle Component
Analysis.
Source code:
#Data Pre-processing Step
# importing libraries
import numpy as np
import matplotlib. pyplot as mtp
import pandas as pd
#importing datasets
data_set= pd.read_csv('user_data.csv')
43| P a g e
Machine Learning with Python Lab
Experiment-15
Source code:
44| P a g e
Machine Learning with Python Lab
45| P a g e
Machine Learning with Python Lab