0% found this document useful (0 votes)
2 views6 pages

Mod 5 Python

The document provides notes on Python programming for Semester III, focusing on data handling techniques such as reading and writing various file formats, pickling, data preparation, transformation, and aggregation. It includes examples of discretization, binning, permutation, and random sampling using the pandas library. Additionally, it discusses the importance of detecting and filtering outliers in data analysis.

Uploaded by

inshreya22
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views6 pages

Mod 5 Python

The document provides notes on Python programming for Semester III, focusing on data handling techniques such as reading and writing various file formats, pickling, data preparation, transformation, and aggregation. It includes examples of discretization, binning, permutation, and random sampling using the pandas library. Additionally, it discusses the importance of detecting and filtering outliers in data analysis.

Uploaded by

inshreya22
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

lOMoARcPSD|62311163

Python SEM III BDS306B Mod - 5 notes @vtunetwork

Political science (Ram Krishna Dharmarth Foundation University)

Scan to open on Studocu

Studocu is not sponsored or endorsed by any college or university


Downloaded by Shreya Thakur (inshreya22@[Link])
lOMoARcPSD|62311163

Semester : III
Subject : Python Subject Code : BDS306B

Module 5

Contents
Reference - Textbook2 – Chapter 5 and Chapter 6

 Reading and Writing data - CSV and textual files, HTML files, XML files,
Microsoft excel files, JSON data.
 Pickle python object serialization.
 Data preparation.
 Data transformation - discretization binning, permutation, string manipulation
 Data aggregation group iteration.

Abbreviation
HTML – Hyper Text Markup Language
CSV – Comma separated values
JSON – Java script object notation
XML – Extended Markup Language

Pickling of Objects in Python

Serialization is the process of converting complex data or an object into byte stream. This process is
called pickling in python. Complex data or object can be recreated back by deserializing. This
process is called unpickling in python. Library pickle or _pickle is used to pickle and unpickle in
python.

Library pandas can also be used to pickle and unpickle

Example:
# pickling using pickle library
d1 = {"USN01":"xyz", "USN02":"abc"}
file1 = open("student","wb")

Downloaded by Shreya Thakur (inshreya22@[Link])


lOMoARcPSD|62311163

file1 = [Link](file1, d1)


[Link]()
# unpickling using pickle library
file2 = open("[Link]", "rb")
d1 = [Link](file2)
print(d1)
[Link]()

# using pandas library


d1 = {"USN01":["xyz"], "USN03":["abc"]}
df1 = [Link](d1)
df1.to_pickle("[Link]") --- pickling
df2 = pd.read_pickle("[Link]") ----- unpickling
print(df2)

Discretization and Binning

Discretization is the process of converting continuous variable to categorical variable. Categorical


variable is one which stores discrete and finite values. Example - result (pass, fail), color
(RED,BLUE, ...), days (Sunday, Monday, ....), Month (Jan, Feb, ......). Continuous variable stores
continuous numeric values like percentage, height, weight, etc.

Pandas provides two functions cut() and qcut() to perform discretization.

Example -

perc = [56.23,67.23,44.56, 89.99,76.99, 99.9,72.65, 45.34,82.34]


bins = [40,50,60,70,80,90,100]
bin_names = ["F", "E","D","C","B","A"]
grade = [Link](perc,bins, labels=bin_names)
print(grade)
output :
['E', 'D', 'F', 'B', 'C', 'A', 'C', 'F', 'B']
Categories (6, object): ['F' < 'E' < 'D' < 'C' < 'B' < 'A']

# if bins are not specified, and no. of bins are specified


perc = [56.23,67.23,44.56, 89.99,76.99, 99.9,72.65, 45.34,82.34]
bin_names = ["E","D","C","B","A"]
cat = [Link](perc, 5, labels = bin_names) # value_counts are not equal
print(cat)
output :
['D', 'C', 'E', 'A', 'C', 'A', 'C', 'E', 'B']
Categories (5, object): ['E' < 'D' < 'C' < 'B' < 'A']

Downloaded by Shreya Thakur (inshreya22@[Link])


lOMoARcPSD|62311163

# using qcut() ---- value_counts are equal but edges vary


perc = [56.23,67.23,44.56, 89.99,76.99, 99.9,72.65, 45.34,82.34]
bin_names = ["E","D","C","B","A"]
cat = [Link](perc, 5, labels=bin_names) # value_counts are equal
print(cat)
output :
['D', 'D', 'E', 'A', 'B', 'A', 'C', 'E', 'B']
Categories (5, object): ['E' < 'D' < 'C' < 'B' < 'A']

Downloaded by Shreya Thakur (inshreya22@[Link])


lOMoARcPSD|62311163

Permutation

Random reordering of Series or rows of a DataFrame is called Permutation.

Example :
df = [Link]([Link](30).reshape(5,6))
print(df)
new_order = [Link](5)
print([Link](new_order))
output :
0 1 2 3 4 5
0 0 1 2 3 4 5
1 6 7 8 9 10 11
2 12 13 14 15 16 17
3 18 19 20 21 22 23
4 24 25 26 27 28 29

Random subet of a dataframe can also be created.

Example :
df = [Link]([Link](30).reshape(5,6))
print(df)
new_order = [2,3,0]
print([Link](new_order))
output:
0 1 2 3 4 5
2 12 13 14 15 16 17
3 18 19 20 21 22 23
0 0 1 2 3 4 5

Random Sampling

Extract a subset of DataFrame randomly using randomint() function in numpy is Random


Sampling.

Example :
df = [Link]([Link](30).reshape(6,5))
print(df)
sample = [Link](0,len(df), size= 3)
print([Link](sample))

Downloaded by Shreya Thakur (inshreya22@[Link])


lOMoARcPSD|62311163

output :
0 1 2 3 4
1 5 6 7 8 9
5 25 26 27 28 29
3 15 16 17 18 19

Detecting and Filtering outlier

Outlier is an unusual value which is very high or very low. Outliers can be considered as those
values which is greater 3 times standard deviation. It is important in data analysis to detect and
remove outliers from dataframe before model building as its presence affects accuracy. Any()
method can be used to detect outliers in a dataframe.

Downloaded by Shreya Thakur (inshreya22@[Link])

You might also like