0% found this document useful (0 votes)
41 views74 pages

Nptel - Data Analytics With Python

The document outlines a comprehensive course on data analytics using Python, covering various topics over 12 weeks, including statistics, probability, hypothesis testing, regression analysis, and clustering. It emphasizes the importance of understanding data and analytics to make informed decisions in business contexts. The course aims to equip students with practical skills in data analysis and interpretation, highlighting the distinction between data analysis and data analytics.

Uploaded by

anujgupta7208
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
41 views74 pages

Nptel - Data Analytics With Python

The document outlines a comprehensive course on data analytics using Python, covering various topics over 12 weeks, including statistics, probability, hypothesis testing, regression analysis, and clustering. It emphasizes the importance of understanding data and analytics to make informed decisions in business contexts. The course aims to equip students with practical skills in data analysis and interpretation, highlighting the distinction between data analysis and data analytics.

Uploaded by

anujgupta7208
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

INDEX

S. No Topic Page No Week 1


1 Introduction to data analytics 1 2 Python Fundamentals - I 33 3 Python Fundamentals -
II 54 4 Central Tendency and Dispersion - I 83 5 Central Tendency and Dispersion - II
108 Week 2
6 Introduction to Probability- I 127 7 Introduction to Probability- II 155 8 Probability
Distributions - I 177 9 Probability Distributions - II 198 10 Probability Distributions - III
225 Week 3
11 Python Demo for Distributions 246 12 Sampling and Sampling Distribution 256 13
Distribution of Sample Means, population, and variance 287 14 Confidence interval
estimation: Single population - I 304 15 Confidence interval estimation: Single population
- II 324 Week 4
16 Hypothesis Testing- I 342 17 Hypothesis Testing- II 364 18 Hypothesis Testing- III 380
19 Errors in Hypothesis Testing 394 20 Hypothesis Testing: Two sample test- I 422
Week 5
21 Hypothesis Testing: Two sample test- II 442 22 Hypothesis Testing: Two sample test-
III 464 23 ANOVA - I 480 24 ANOVA - II 494 25 Post Hoc Analysis(Tukey’s test)
513 Week 6
26 Randomize block design (RBD) 542 27 Two Way ANOVA 563 28 Linear Regression -
I 583 29 Linear Regression - II 601 30 Linear Regression - III 614 Week 7
31 Estimation, Prediction of Regression Model Residual Analysis 634 32 Estimation,
Prediction of Regression Model Residual Analysis - II 652 33 Multiple Regression Model
- I 674 34 Multiple Regression Model-II 695 35 Categorical variable regression 714
Week 8
36 Maximum Likelihood Estimation- I 744 37 Maximum Likelihood Estimation-II 761 38
Logistic Regression- I 785 39 Logistic Regression-II 802 40 Linear Regression Model Vs
Logistic Regression Model 818 Week 9
41 Confusion matrix and ROC- I 838 42 Confusion Matrix and ROC-II 860 43
Performance of Logistic Model-III 883 44 Regression Analysis Model Building - I 895 45
Regression Analysis Model Building (Interaction)- II 910 Week 10
46 Chi - Square Test of Independence - I 928 47 Chi-Square Test of Independence - II 949
48 Chi-Square Goodness of Fit Test 971 49 Cluster analysis: Introduction- I 990 50
Clustering analysis: part II 1009

Week 11
51 Clustering analysis: Part III 1026 52 Cluster analysis: Part IV 1046 53 Cluster analysis:
Part V 1068

54 K- Means Clustering 1083 55 Hierarchical method of clustering -I 1109 Week 12


56 Hierarchical method of clustering- II 1134 57 Classification and Regression Trees
(CART : I) 1162 58 Measures of attribute selection 1187 59 Attribute selection Measures
in CART : II 1206 60 Classification and Regression Trees (CART) - III 1224
Data Analytics with Python
Prof. Ramesh Anbanandam
Department of management studies
Indian Institute of Technology, Roorkee

Lecture No 1
Introduction to Data Analytics

Welcome students this course on data analytics with the Python today is the introduction
class. This lecture is on introduction to data analytics.
(Refer Slide Time: 00.34)

The objective of this course is to introduce the conceptual understanding using simple and
practical examples rather than repetitive and point clique mentality, here most of the students
generally they are, how they are using the software for doing data analytics. Just they want to
just click it, they want to get the result, they do not want bother about exactly what is
happening inside the software. This course should make you comfortable using analytics in
your career and your life.

You will know how to work with a real data and you might have learnt the many different
methodologies, but choosing the right methodology is important. This course will focus you
will help you how to choose the right data analytical tools.
(Refer Slide Time: 01:17)

1
Objective of the course is, when you look at this picture. How this person is using this tool,
there is a ladder. He was not knowing correctly how to use this ladder for the, for the purpose
it is intended. So the danger in using quantitative method does not generally lie in the
inability to perform the calculation, because of the computer development in computer
technology.

There are many packages are available for doing data analytics. But, the real threat is lack of
fundamental understanding of why to use particular technique or procedures and how to use it
correctly and, how to correctly interpret the result. This course will focus on how to choose
the right technique and how to use it correctly and how to interpret the result.
(Refer Slide Time: 02:01)

2
So what was the learning objective of this class that is; after completing this lecture what you
will learn one is you can define what is data and its importance. You can define what are data
analytics and types. You can explain why analytics is in today's business environment is so
important. Then we can see how statistics, analytics and data science are interrelated, there
seems to be some overlap in this we will clarify that what is the difference, how these are
overlapped how these are interrelated.
In this course we are going to use a package called Python. I will explain how and why it is
important to use the Python in this course, at the end of this session we will explain the four
important levels of data that is nominal, ordinal, interval and ratio. Now we will go to the
content;
(Refer Slide Time: 02:54)

We will define data


and its
importance. There are
three term one is variable,
measurement and data.
Next we will see what is
generating so much data.
Next we will see how data add value to the business, and then we will say why data is
important.
(Refer Slide Time: 03:11)

3
See the variable, measurement and data these are the terms which we are going to use
frequently in this course. So what is a variable? Variable is a characteristic of any entity being
studied that is capable of taking on different values. Say for example, X is the variable it can
take any values it may be 1 it may be 2 or it may be 0 and so on. The measurement is, when
you standard processes used to assign numbers to your particular attributes or characteristics
of variable are called a measurement.

For that X, you want to substitute some values. For that value, you have to measure the
characteristics of the variable, that is nothing but your measurement. So then, what is the
data? Data are recorded measurement. So there is a variable you measure the phenomena,
after measuring the phenomena you are substituting some value for the variable so the
variable will take a particular value that value is nothing but your data.

So X is the variable for example number 5 is the data. How you are measuring that 5, that is
called measurement. Then what is generating so much of data.
(Refer Slide Time: 04:33)

Data can be generated different way humans, machines, and human - machines combines.
The humans, machines and human - machines combines in the sense, now seen everybody is
having the various Facebook account, we have LinkedIn account, we are in various social
network sites. Now the availability of the data is not the problem. It can be generated
anywhere where the information is generated and stored and structured or unstructured
format.
(Refer Slide Time: 05:06)

So how the data add value to the business? So the data after getting from various sources
assume that it is a store in the form of data warehouse. So from the data warehouse the data
can be used for development of a data product. Here we are using the word data product and
in the coming slides, I will explain exactly what is the data product with some examples. So

5
the same data, if we look at the right hand side that can be used to get more insights from the
data.

Okay, what do you mean the data product? For example, algorithm solutions in production,
marketing and sales, example of some data product. For example, recommendation engine
one of the example for data product. Suppose, if you go for Flipkart or Amazon for buying a
particular product in that package, that software itself, we will recommend to you what is the
next product, possible product that you can buy. That is nothing but the recommendation
engine.

Even if you watch some YouTube videos on particular topic, that YouTube itself will suggest
to you what are the relevant videos are available. So that is a recommendation engine. That is
one of the examples of your data product; with the help of data so that will help you to
forming a data product or you can get an insight from the data. That will add your business
value to you.
(Refer Slide Time: 06:27)
See this is an example of your data products, this is the driverless car. Google car, so the
whole concept of Google car is with the help of data. It is detecting all other requirements for
driving the cars. The next example is for recommendation engine, as I told you previously
when you buy any product they will suggest you that along with this product, the other
product also can be purchased.

6
Another very common example for a data product is Google. The Google is lot of
applications, there one of the application of example for data product is Google mapping. So
the Google mapping is helping you to find out what is the right route, which road there is a
traffic, in which road there is a toll booth, so this kind of information we can get it from the
Google map. So this Google map is the one of an example of your data production.
(Refer Slide Time: 07:20)

Now why data is important? The data helps in making better decisions, data helps in solve
problem by finding the reason for underperformance. Suppose some company it is not
performing properly by collecting the data we can identify what was the reason for this under
performance. The data helps one to evaluate the performance. So what is the current
performance, the data also can be used for benchmarking the performance of your business
organization.

And after benchmarking data helps one improving the performance also, so data also can help
one understand the consumers and the markets, especially the marketing context. You can
understand who are the right consumers and what kind of preferences they are having in the
market.
(Refer Slide Time: 08:16)

Next we will define what is a data analytics and its types? So in this coming two, three slides
we are going to discuss, we will define what is data analytics? Then you say why analytics is
important? Then we will see that data analysis? Then we will see how data analytics is
different from data analysis? At the end will we see types of data analytics?
(Refer Slide Time: 08:40)
We will define data analytics is the scientific process of transforming data into insights for
making better decisions. See it is a scientific process for transforming the data into for
making better decisions, even without the data also even without doing analytics also you can
make the decision but you cannot make the better decision without analytics. By the virtue of
your experience on intuitions you can take the decisions that also sometimes may be correct.

8
But about the help of data if you are making the decision then that will enable you to make
the better decisions. Another professor James Evans, he has defined the data analytics in this
way. it is the use of the data information technology, statistical analysis, quantitative methods
and mathematical or computer-based models to help managers gain improved insight about
their business operations and make better, fact-based decisions.

You see that there are many terms which are appearing here, one is IT, next one is a statistical
analysis, and next one is the quantitative methods, then mathematical knowledge and
computer-based models. So when we will see how these are interrelated in coming slides.
Generally, among the students, there is a confusion whether the analysis and analytics is same
or different?
(Refer Slide Time: 10:13)
Why analytics is important. The opportunity abounds for the use of analytics and big data
such as: for determining the credit risk, for developing new medicines, especially in
healthcare. The healthcare analytics is an emerging, that is helping you to identify what is the
correct medicines. Finding more efficient ways to deliver product and services. For example:
in the banking context data analytics is used for preventing the fraud, and it is uncovering the
cyber threats.

With the help of data analytics you can find out the possible cyber crimes and we can detect it
we can prevent it. And data analytics are also important for retaining the most valuable
customers. We can identify who is your valuable customer or non valuable customers. So we
can focus on more on our valuable customers. Okay,

9
(Refer Slide Time: 11:08)

Now what is the data analysis? Is the process of examining, transforming and arranging raw
data in a specific way to generate useful information from it. So data analysis allows for the
evaluation of data through analytical and logical reasoning to lead to some sort of outcome or
conclusion in some context. Data analysis is a multi-faceted process that involves a number
of steps approaches and diverse techniques. That we will see in coming lecture.
(Refer Slide Time: 11:41)

So now we will see what is the analysis is data analysis and data analytics. When you say
analysis when you say data analysis it is something about what has happened in the past. So
we will explain why that has happened? We will explain how it has happened? We can
explain why it has happened? For example, when we say data analysis that is nothing about

10
studying about what has happened it is like kind of a post-mortem analysis. What has
happened in the past?
(Refer Slide Time: 12:13)

Okay, in the contrary the analytics is studying about what will happen in future and with the
help of analytics. We can predict explore possible potential future events. (Refer Slide
Time: 12:25)
So the analytics is maybe qualitative or quantitative. For example in analytics if we say
qualitative analytics. So it is the decision mostly based on the intuition. But if you say in
quantitative where with the help of formulas with the help of algorithms will make the
decisions.
(Refer Slide Time: 12:44)

11

So in the analysis data analysis also we can go for qualitative. We can explain how and why
a story ends in that way it did? When we say in quantitative we can say, how the sales
decreased the last summer. When I say as I am repeating, when you say analysis is something
studying about what has happened in the past.
(Refer Slide Time 13:12)
Okay, so it is not exactly analysis equal to analytics. Similarly when you say data analysis is
different data analytics is different.

Similarly business analytics is different business analytics. When you say analytics is nothing
but studying about the future events with the help of the past data.
(Refer Slide Time 13:34)

12

Next we will go for classification of data analytics, based on the phase of workflow and the
kind of analysis required, there are four major types of data analytics. One is descriptive
analytics, diagnostic analytics, predictive analytics and prescriptive analytics. We will see
these four types of analytics in detail in coming classes:
(Refer Slide Time 13:57)
If we look at the difficulty and the kind of value which we can get from different types of
analytics; this picture shows for example: when you see the descriptive analytics that will
answer what happened? Diagnostic analytics, will help you to answer why did it happen?
Predictive analytics will help you what will happen? Prescriptive analytics will help you to
answer how can we make it happen? There is one context when you look at the level of
difficulty you see that the descriptive analytics is the level of difficulty is very less.

13
And the contrary when you look at the prescriptive analytics the difficulty level is more and
the value also, value in the sense business value which adds to you also more. so when there
is a more difficulty there is a more value. Okay,
(Refer Slide Time 14:54)

Then we listen what is the descriptive analytics? Descriptive analytics is the conventional
form of business intelligence or data analysis. It seeks to provide the depiction or summary
view of facts and figures in an understandable format. These either inform or prepare data for
further analysis. so descriptive analysis or we can say another way in statistics can summarize
raw data and convert it into your form that can be easily understood by humans. They can
describe in detail about an even that has occurred in the past. Okay,
(Refer Slide Time 15:40)

14
Some of the examples of descriptive analytics is a common example of descriptive analytics
are company reports that simply provide the historic review like: data queries, reports,
descriptive statistics, data visualization and data dashboard. Okay,
(Refer Slide Time 16:00)

The next one will go to the diagnostic analytics. Diagnostic analytics is a form of advanced
analytics which examines data or content to answer the question why did it happen? So we
are diagnosing, suppose we are meeting a doctor for consulting, so he will try to understand
why this has happened? Okay so that kind of analytics nothing but diagnostic analytics. So
the diagnostic analytical tools aid and analyst to dig deeper into an issue.

So that, they can arrive at the source of the problem. So doctor also will identify you
somebody has got some disease what was the sources of the problem. Similarly the
diagnostic analytics also if something has happened for example the company's not
performing well that diagnostic abilities will help you to identify what was the core reason
for that. In a structured business environment tools for both descriptive and diagnostic
analytics go parallel.

So when you look at the whether it is a prescriptive or diagnostic analytics, the tools,
analytical tools which are using can be same only the purpose may be different. (Refer Slide
Time 17:09)

15

For example: data discovery, data mining, and correlations. These tools can be used for your
prescriptive analytics also. Okay,
(Refer Slide Time 17:20)

Now we will go for predictive analytics, predictive analytics helps to forecast trends based on
the current events. When you say predicting obviously it say, that it is discussing about what
will happen in future? Predicting the probability of an event happening in future are
estimating accurate time it will happen can all be determined with the help of predictive
analytical models. Many different but co-dependent variables are analysed to predict a trend
in this type of analysis.

So in the predictive analytics one of the tool is the regression analysis. There may be some
independent variables, some dependent variables, sometimes more dependent variable, more

16
than one dependent variable and how these variables are inter-related. So that kind of study is
nothing but your predictive analytics.
(Refer Slide Time 18:11)

When you look at this picture, you see that with the help of historical data by using different
algorithm, predictive algorithms you can come with a model. Once the model is developed a
new data can be fit into this model so we can get some predictions about the past events.
(Refer Slide Time 18:35)

Example is linear regression, time series analysis and forecasting and data mining. These are
the techniques for predictive analytics.
(Refer Slide Time 18:46)
17

The last one is the prescriptive analytics. A set of techniques to indicate the best course of
action. It tells what decision to make to optimize the outcome. The goal of prescriptive
analytics is to enable: quality improvements, service enhancements, cost reductions and
increasing productivity. Okay,
(Refer Slide Time 19:13)

In the prescriptive analytics, some of the tools which we can use is optimization models,
simulation model, and decision analysis. These are the tools under prescriptive analytics.
(Refer Slide Time 19:27)
18

Next is we are going to see, why the analytics so important? In this section we will see what
is happening the demand for data analytics and we look at the different elements of data
analytics.
(Refer Slide Time 19:44)

This picture shows, Google Trends, this was up to 2017. See for example, the blue represents
the data scientist; this orange represents the statistician operation researchers. You see the
trend is it is increasing that means people are searching in the Google search engine the word
data scientist more number of times. See the search count is increasing. That means there is a
demand for that particular say job.
(Refer Slide Time 20:19)

19

You see, if you look at this is the newspaper clipping from Times of India. There are so many
news are coming about data scientists and the future requirement of data scientists. You see
the data scientist earning more than CA’s and engineers. You can look at this link for further.
(Refer Slide Time 20:37)

And you see the demand for data analytics. This also newspaper clipping with companies
across industries striving to bring their research and analysis department up to speed, the
demand for qualified data scientist is rising. So there is an emerging field. so many
companies are looking for the qualified data scientist. So if you take this course and end up
the course that you may be qualified for getting into these companies.
(Refer Slide Time 21:07)
20

Many times you see what is data analytics, Statistics, data mining, optimizations. These are
students having different understanding on that. So when we say data analytics, there are
different element one is statistics, next one is the business intelligence information systems,
then modelling and optimizations, then simulation and risk. We can say if you are able to do
what if analysis? That is nothing but sensitivity analysis, visualization, data mining. These are
the components of data analytics and how these different domains are interrelated?
(Refer Slide Time 21:47)

Next we will see, what kind of skill set is required to become a data analyst? then we will see
the small difference between data analyst and data scientist?
(Refer Slide Time 21:59)
21

See to become a data analyst is the basic fundamental knowledge is you need to have
knowledge of mathematics. Next you need to have the knowledge of technology is nothing
but hacking skill. Hacking skills in the sense, if the data is given hacking is done and looked
at the positive way. How to use the data to get more information? The next skill is business
and strategy acumen; you should have the knowledge of the domain and knowledge of the
business and you knew to the strategy equipment.

So these three skills are required for a good data scientist. It is very difficult to have a one
person will have all these three skills that's why availability of good data analyst is becoming
very difficult. Because somebody may be very good at mathematics but they may not have
very good knowledge and business, some people may be very well at technology, technology
in the sense information technology, they may not have good knowledge on the business
knowledge.

So we need to have the combination of all these three skills otherwise the group of people
some people from mathematics department or mathematics area, some people from computer
science, some people from the domain knowledge. They were to work together to form a
good data scientist team, so these forms data science.
(Refer Slide Time 23:31)
22

Now what is the difference between data analysts and data scientists and the difference is
what kind of role they are doing? For example; the role of your data analyst is, see in your
business context, he may have the knowledge of business domain. For example; if he is good
at doing analytics in the area of marketing, he can be called as a marketing analyst. If the
person is from finance area, he can be called it as a finance analyst. So he is the analyst, data
analyst.

But the role of data scientist is little bigger, because the data scientist need to have the
knowledge of advanced algorithms and machine learning and able to come out with a data
product. Which I told you in the previously, so the data scientist can come out with a data
product. Okay,
(Refer Slide Time 24:30)
23
In this course we are going to use Python. In this in the next lecture, I will tell you the basic
introduction about the Python. Here we will see why we are going to use the Python?.
Because python is very simple and easy to learn. Most importantly it is a free software and
open source. It uses interpreted, it is not the compiler. Suppose what do you my compiler and
interpreter is you need a compiler to solve the whole program but interpreter need not be in
that way.

It can solve, even you can interpret one sentence also, one line in the programming line also.
it is dynamically typed, dynamically type in the sense in some other programs every time you
have to declare the variable. What is the nature of the variable? Whether it is integer?
Whether it is a float? But here you need not do. It is dynamically takes the value. it is
extensible, extensible in the sense if you make a code in some other language that can be
extended with the help of Python.

And can be embedded, embedded in the sense you have made some program in Python it can
be embedded with the some other platforms and it has extensive library. (Refer Slide Time
25:45)

The usability of Python is it is a desktop and web applications, it can be used for data
applications, it can be used for networking applications, most importantly it can be used for
data analyst, data science can be used for machine learning, it can be used for IoT Internet of
Things and artificial intelligence applications and can be used for games.
(Refer Slide Time 26:05)

24
Another reason for choosing Python is most of the companies, they use Python is a language
in their company. Like for example; Google, Facebook, NASA, Yahoo and eBay. They use
Python is a programming language.
(Refer Slide Time 26:23)

In this Python also we are going to use Jupyter notebook. In the next class I will explain you
because it is the client-server application is edit code on web browser. It is easy in
documentation, easy in demonstration and user friendly interface.
(Refer Slide Time 26:39)

25
This was the last session of this lecture; we will explain four different levels of the data. What
is the type of variables? Levels of data measurement? Compare for different level of data:
will say nominal, ordinal, interval and ratio. We will see that why and what is the usage of
knowing this different level of data?
(Refer Slide Time 27:03)

The one way for classifying the data is the categorical data, one is a numerical data. In
categorical data; you see marital status, political party, and eye color. These are categorical
data. Numerical data; it can be discrete or continuous. Discrete data may be a number of
children and defects per hour. So this is the discrete data. In the continuous data may be
weight and voltage. These are the example of continuous data.

26
So what is the difference between discrete and continuous is, you say a number of children
you may say two children or three children 2.5 children was not possible but in continuous, if
you look at between 0 & 1 the numbers are continuing there are infinite number of values that
are there between 0 & 1. So it is a continuous variable.
(Refer Slide Time 27:56)

Next will you see the different level of data measurement? Easily we have seen the
classification of data. We classified as the categorical data and numerical data. There is
another way of classification is, classifying into nominal data, ordinal data, interval data and
ratio data.
(Refer Slide Time 28:14)

We will look at, what is a nominal data? Nominal scale classifies data into distinct categories
in which no ranking is implied. The example of nominal data is gender, marital status. For

27
example; gender suppose you are conducting a questionnaire. Suppose you captured the
gender male 0, female 1. This 0 1 represents just the gender. You cannot do any arithmetic
operations with the help of the 0 & 1.

For example, you cannot find out the average, software will give you some number but there
is no meaning for that. Similarly, marital status, whether it is married or unmarried. This is
the example of nominal data.
(Refer Slide Time 29:01)
The next level of data is the ordinal scale. It classifies data into distinct categories in which
the ranking is implied. Here the numbers are the ranked. For example; you may ask the
customer to give a ranking about their level of satisfaction. For example, satisfied, neutral,
unsatisfied. The faculty ranking, for example; professor, associate professor, assistant
professor.

You see that their rank is followed for example 1 professor, 2 associate professor, 3 three
professor. Student grades, A, B, C, D, E, F. These are ranking, because the numbers 1, 2, 3
represents the rank.
(Refer Slide Time 29:45)

28

The next level of data is interval scale. The interval scale is ordered scale, in which the
difference between measurements is a meaningful quantity but the measurement do not have
to zero point. The example of interval scale is, for example year. Say now, this here is 2019,
you can add and subtract something. You can add another five years, its 2024 or you can
subtract another nine years, its 2010.

But you cannot multiply, if you multiply that number for example 2019 and 2020 you will
end up with the big number there is no meaning for that. Because, there is no meaning for
zero. Another example of interval scale is your Fahrenheit temperature. For example, in the
Fahrenheit scale, the zero represents freezing point but it is not the absence of the seat but
absence of the temperature but at the same time in the Kelvin for example minus 273 it is
absence of heat. So Kelvin will be the some other scale. That you will see the next one,
(Refer Slide Time 30:52)

29

The ratio scale is the ordered scale in which the difference between the measurements is a
meaningful quantity and the measurements have the true zero point. Weight, age, salary and
the Kelvin temperature comes under ratio scale. Because 0 Kelvin that means the absence of
the heat. So in the ratio scale, he can do all kinds of arithmetic operation. For example the
nominal, you cannot do any arithmetic operation. In ordinal you cannot do in arithmetic
operation.

In the interval you can add and subtract but you cannot multiply. But in the ratio data, you
can do all kinds of arithmetic operations. You can add. You can subtract, you can multiply,
and you can divide.
(Refer Slide Time 31:35)

30
You see the usage potential various level of data. For example the usage potential of nominal
data is not that much. The next one is ordinal; next one is interval, next one ratio. So the ratio
data is having the highest to use its potential. The nominal data is having the least usage
potential.
(Refer Slide Time 31:56)

This is more important, why we have to still know the different types of data. Because this
types of data is helping to choose the right analytical tools for doing analysis. For example; if
the data is the nominal data. You can do only nonparametric tests. For example the data is
ordinal, here also you can do only nonparametric test. But if the data is interval, you can do
parametric test. You see that interval; you can do all above plus addition and subtraction.
In the ratio, if you can do all of the above plus multiplication and division and statistical
methods. You can go for parametric methods. So the purpose of classifying the data into
nominal, ordinal, interval, ratio is to choose the right analytical tools with it whether it is a
parametric or non parametric. The other reason is sometime for if we want to do a non
parametric analysis that is used only for nominal data.

Sometime the students they will, the data may be nominal but they may go for a parametric
test that, should not be done. That is the purpose of knowing what kind of, what is the nature
of this data. So in this class we have seen the introduction for data analytics. We have seen
the importance of data analytics. We have seen the classification of data analytics. Then we
can we have seen what is the analytics and analyst and we have seen different types of data.

31
The next class we will learn about what is Python? How to install the Python and what kind
of descriptive analysis we can do with the help of Python? So the next class will meet you
with another lecture. Thank you very much.
32
Data Analytics with Python
Prof. Ramesh Anbanandam
Department of management studies
Indian Institute of Technology, Roorkee

Lecture No 2
Python fundamentals -1

Good morning students, in the last class, that was the introduction class, we have seen the
importance of data analytics and we have seen certain classification of data analytics. This is
my second lecture that is Python fundamentals because we are going to use this Python. In
this lecture I have 3 objectives.
(Refer Slide Time: 00:50)

One is I will tell you how to


install Python second one I will see some fundamentals of the Python, third one some data
visualization. In the data visualization I am going to give only theory in this class. The next
class we are going to use Python and we are going to take some sample data and we have to
visualize the data using Python software.
(Refer Slide Time: 01:13)

33

As I told you the 1st one is how


to install this Python. There are 5 steps is there. Step 1, we are going to see in detail in
coming slides. In step 1 we are going to visit this website [Link] at the address
bar of the web browser 2nd one we are going to click on download button 3rd one will
download python 3.7 version for Windows operating system. Then we will double click that
is a 4th step we will double click on file to run the application.

The 5th one will follow the instruction until the completion of installation process. What I
have done I have taken the screenshot of all these, all the 5 steps while installing the laptop. I
am going to show each steps in the form of screenshot.
(Refer Slide Time: 02:03)

The 1st one is type


this [Link] at the address bar of the web browser. (Refer Slide
Time: 02:13)

34

2nd one is once you typed it you can see this screen, here you see this location you can see
here. This location there is a download option. When you click that you see that the left side
also I have rounded there is a download option you download it.
(Refer Slide Time: 02:36)
In the 3rd step is there are two
versions of python, python 3.7 and python 2.7. In these courses we are going to use the latest
version that is the Python 3.7.
(Refer Slide Time: 02:46)

35

In the 4th step double click on


file to run the application it will get downloaded. when you double click for example; I have
stored this anaconda in F folder.
(Refer Slide Time: 02:59)
Step 5 is
just to keep on click Next
(Refer Slide Time: 03:02)

36

You have to
agree for their agreements, terms and conditions. (Refer Slide Time:
03:06)
Next you
select, just me recommended click Next. (Refer Slide Time: 03:10)

37

Then it is installed in C drive click Next.


(Refer Slide Time: 03:15)
Then install,
Installation process is started, then installation is completed. (Refer Slide
Time: 03:25)

38

Again click
Next
(Refer Slide Time: 03:29)
Then click
and finish Ok
(Refer Slide Time: 03:34)

39

Now we have installed the


anaconda. So I will explain you how to open Jupyter notebook. I will switch the screen
(Refer Slide Time: 03:46)
Yeah, this is the screen. The
initially there are you see, I will see what is this some box it is showing in blue color
sometimes it will show in green color that I will show you later. So this is the Jupyter
notebook look like.
(Refer Slide Time: 04:04)

40

The next one is why there are


some more interfaces there for using Python. There is a spider is there Jupyter is it but we
prefer Jupyter for some reasons because it is edit code on web browser, it is easy in
documentation, it is easy in demonstration and it is user- friendly interface. That was the
reason we are using Jupyter it is not necessary if you already you are comfortable in some
other interface you can continue with that.
(Refer Slide Time: 04:34)
See that in anaconda it consists
of two software one is python that is on the left hand side the another side the right hand side
the Jupyter applications these are combined together and kept in the Anaconda software
package.
(Refer Slide Time: 04:49)

41

When you
from the start when you type Jupyter you will get this screen (Refer Slide
Time: 04:58)
Then when you click launch
you will get this one. So now from the start I am going to explain how to start this Python
jupyter notebook.
(Video Starts: 05:08)
You have to type Jupiter and Jupiter notebook .When you click it you will get this one.
Suppose if you want to type in a new go new Python 3. Yeah? Here there is a Jupyter there is
a it is coming untitled 2 there you can change the name. You give the name as introduction to
Python, introduction Okay?
(Video Ends: 05:40)

(Refer Slide Time: 05:40)

42

You see there is a box is


appearing this is called cell I have made it in the red color, it is a cell it can be the cell can be
accessed using Enter key.
(Refer Slide Time: 05:51)
You see sometime that box will
look like a green color, Green color indicates it is in edit mode sometime the box will look
like in blue color.
(Refer Slide Time: 06:01)

43

The blue color indicates it is a


command mode. See when you go to below the help there is a file name is called mark down
.There if you type something then you select mark down that is used for making
documentation. So it contains documentation, here text not executed as a code it is only for
our understanding purpose.
(Refer Slide Time: 06:22)
Okay? Now about the Jupyter
Notebook Command mode allowed editing notebook as a whole. To close edit mode press
Escape key. Execution can be done in three ways you can simultaneously we can press
Ctrl+Enter. So what will happen when you press Ctrl+Enter output field cannot be modified,
another way is to press Shift+Enter output field is modified. Then there is a third way is there
is a run button on the jupyter interface.

44
That you can directly you can click that. Then your code will get executed command line is
written proceeding with # tag symbol. So when you want to make some understanding on
your program you can use the # symbol, so that will not be executed.
(Refer Slide Time: 07:17)

That you only for your


understanding purpose but there are about the Jupyter notebook important shortcut keys.
When you press A that is used to create a cell above when you press B that is to create a cell
below when you press D+ D for deleting cell. When you press M that will made a say mark
down cell, when you press Y that is for coding cell. (Refer Slide Time: 07:46)
For example;
when I am entering B
(Refer Slide Time: 07:54)

45

We will go to the next one fundamentals of Python and you see loading here .What we are
going to see in coming slides. Loading a simple delimited data file counting how many rows
and columns were loaded and determining which type of data was loaded. Then looking at
different parts of data by subsetting rows and columns because these activities are more
important because once we loaded a data that may have n number of cells n number of rows,
column and rows.

Sometime we need to do some operation using only few rows are few cells .You should know
how from a big data file how to use only a particular row or how to use a particular column.
Sometimes we can have a collection of rows also, collection of columns also for doing our
specific operations.
(Refer Slide Time: 08:49)

46
This was the reference book which I am following for this course and the book name is
Pandas for everyone especially for this lecture. It is the professor Daniel Y. Chen he is the
author of this book.
(Refer Slide Time: 09:04)

Now we are going to learn how


to load a simple delimited data file. This is the fundamental because before doing data
analysis the first step is how to load the data into the Python. For that purpose we are going to
import some basic libraries one is pandas numpy another is [Link] as plt. So, first
we are going to import these three basic library .Then we are going to load the data. The data
sources it is taken from [Link] /gennybc/gapminder.

So I have downloaded this data set already I am going to tell you how to load the data set in
to the python. Before that I am going to open that excel file, I am going to show what is the
column? What is the row open the excel file?
(Video Starts 10:07)
When you look at this I am reading the column see that there is a country, year, population,
continent, life expectancy that is given as the short name life exp then gdp per capita. So in
rows there are, how many rows is there I will tell you how many rows is there I am coming
down and this is a this is csv file format. How many rows are there? There are 1705.
The last row is Zimbabwe Right? Please look at the data Zimbabwe year 2007. I think it is a
population, continent, life expectancy this is a per capita income. Okay? Now this data, this
csv file I am going to import into the Python. You see that I am going to call this data set df.

47
df= pd because pd is the short-form of pandas, Pandas nothing but the panel data
Pandas.read_csv.

Why I am using csv because the csv file is I am going to read it. The location of the file given
the path of that file you can directly copy that path but one thing you have to note it down.
See, C: this will be this should be \ because when you copy that path directly. Generally you
will get here / but you have to change it. So I changed it back C: / users / ET cell / desktop /
gapminder-five year [Link].

Look at their it should be in the code. Now I am going to read the df, Yes? once I read it you
see that, the row is starting from 0 . That is a very important. It is a 0 indexing 0, 1, 2, 3, 4, 5,
6, 7, 8 I am able to see whatever I have seen in the csv file just a few minutes before. You see
I am able to see the country, year, population, continent, life expectancy and gdp per capita.
Okay? What I how I have read it pd. read_csv

Suppose I have installed, I have loaded that data I want to say what are the headings of that
file. Heading means what are the columns. For that there is, in Python there is a two type
print and open the parentheses [Link] when you execute this one you will get 1st 5 rows that
means 0, 1, 2, 3, 4. So that means when you execute this one you can see 1st 5 rows from the
data set Yes? You are able to see that, Okay?

I will go to the next command, suppose I want to know the size of that file that is I want to
know how many rows and how many column is there. For that there is a command called the
shape. So print [Link], df is they were finally because we outer loading that csv file we
have named in the variable called df. Okay? So when you type print [Link] then we will
come to know how many rows are there. How many columns are there?

So, I am typing print df shape. One more thing you should not type this parenthesis because it
is the shape is without parentheses. So I am going to remove this parenthesis again I got to
run it. Yes, it is showing how many rows? How many columns? Okay? We will go to the next
one now I want to know how many column names? What are the column names? So if I type
print [Link] Right?
48
Here, please note that here also there is no parenthesis if I type print [Link]. This was the
output which I copied see what is output disappearing country, year, population, continent,
life expectancy, gdp per capita, data type is object; I will show you how this comment is
running. Type print df., yes you are able to get this way. So what the students what you have
to do while looking at the video you have to open your laptop you have to type this
command.

Then you have to see you can verify the answer. Okay? The next command is to get the data
type of each column; you have to type this command print [Link]. That will give you the
summary of the all data set and what is the nature of the data. We will see that how it is
appearing. So, I am going to type print [Link]. Now you we will see the data type of each
column. For that you have to use this command dtypes.

So print [Link], this is the output which you will get it. I will show you in the Python, first
we will look at what is the subject output you see countries object, year is an integer,
population pop it is a variable that is in the float. Float means there is a decimal continent it is
an object that means a character, life expectancy it is a float that means you are going to get
that value in decimals.

Similarly, gdp per cap that also going to get in the decimals then data type is object now we
will go to the Jupiter. We run this command so you see that you see line number 8. Print df.
dtypes you are getting whatever it was there in this or whatever I have shown in the slide is
there. Say country object, year integer, population float type of data and so on.
( Video Ends 17:10)
(Refer Slide Time: 17:11)

49
This is a classification of types
of data in the perspective of pandas, in the perspective of Python. See when they say string it
is a most common data type it is a character. When I say ‘Int’ it is a whole number integers. It
is a float number with the decimals. Date, time is that is to represent the data it is not loaded
by default that need to be imported. Whenever it is a requirement is there that we will see
(Refer Slide Time: 17:38)

That one more command is to get more information about the data. So you type [Link] you
will get the full details about each columns. We will do that one.
(Video starts: 17:53)
Look at this when it print df. info so I am getting data columns there were 6 columns country
there are 1704 rows is there Non null object that means all the data is there is no missing
values. Similarly year 1704 rows is there, Non-null that is an integer, Non null means, that all

50
the values are filled. There is no missing cell so population, float, continent object, life exp
float, gdp per capita float, memory usage is this much.

Suppose, there is a big data file is there we want to see the specific rows are specific columns.
How to do that? Now get the country column and save it to its own variable. So country if
you look at the data which I will show initially countries one of the column. So I want to pick
up only that country column I am going to save it. I am going to give the name for that a
country_df= df you see that you have to open the square bracket, Square bracket within
quote.

Suppose, in the country column I want to see 1st 5 rows Okay? You type print, open
parenthesis country_df.head that shows 1st 5 rows and see that now from the full data we
have fetched only the country column. That we have seen there are 1st 5 rows, that is 0th row
is a, 0, 1, 2, 3, 4, 5 to 5 rows we are able to see when you from the big file. Suppose there
may be requirement you need to see what are the last five observations for that purpose.

You type print country_df.tail, then you can see from the bottom we can see last 5 rows. You
will see how it is appearing, Yes? So what is it we are able to see last 5 rows from the
country, country_df file.
(Video ends: 21:15)

(Refer Slide Time: 21:15)

51
There may be requirement you need to see more than one column at a time. So I am going to
save in the form of another file name that is called a subset, Subset=df. You see there is a
double square bracket so I want to switch the country columns, continent columns and year
columns. Then I going to see what are the heading that means I want to see what does the 5
rows of these subsets so we will go to they go to Python.
(Video starts: 17:46)
I am going to call it a subset continent. Suppose I want to see the 1st 5 rows of this file called
a subset. Data set called subset. You see that I am able to fetch 3 columns at a time that is on
the country, continent and year. The same way we from the subsets file I want to see last 5
rows so print [Link]. Let us see what we are getting we will get this output Yes? You see
that there are 3 columns.
There were the last 5 rows from the button. So far we looking at different columns now we
want to subset rows by index label there is one command called loc. So first we look at the
file initial file that is a print [Link]. Next you see that I want to locate the 0th row so for that
purpose, print [Link] see it is a square bracket you type 0 because if you suppose we want to
know the 1st row i out to enter because I would enter 0 because Python counts from 0, so
print [Link] 0 that will show the 1st row.

You see 0th row access country Afghanistan year 1952 population is this much continent is
Asia. Suppose I want to access this 1st row that means 0th row, Yeah? 0th row you can verify
0th row is the country Afghanistan, year 1952 this is a way to access a particular row. Dear
students whatever comments which I am typing that I will be given to you when you take this
course you can practice yourself.

You need not bother about in case we are not getting at this stage this all the commands all
the codes will be given to you .You can practice on yourself. Suppose I want to get the 100th
row how to access from the file df? I want to look at 100th row so you type print [Link] 99.
You can exactly access in 100th row what is the element is there? Suppose I want to access
100th row df, 100th row is the country Bangladesh, year 1967, population this is.

This is the way to access different rows for our calculation purpose. So far we have seen how
to load csv file into the Python, we have seen some basic commands.
(Video Ends: 25:21)

52
We have seen how to know the size of the file then we have seen how to access a particular
row and also we have seen how to subset from the given big file? How to subset different
small data file? So that can be used for our further analysis. So the next class we will see how
to access different columns that will continue in the next lecture. Thank you.
53
Data Analytics with Python
Prof. Ramesh Anbanandam
Department of management studies
Indian Institute of Technology, Roorkee

Lecture No 3
Python fundamentals

Okay? We will continue our lecture. How to access different rows and columns, because, it
has very important applications.
(Refer Slide Time 00:33)
When the data file is very big
sometimes you need to access only some rows or some columns for your calculation purpose.
That we will learn how to access a particular rows or particular columns there is looking at
columns, rows and cells. When you look at this see print [Link] when I use this command
and getting there are different. For example; the first column says 0, 1, 2, 3, 4, country, year,
population, continent, life expectations, gdppercapita. (Refer Slide Time: 01:04)

54

Suppose I want to get the first


row as we know that the Python counts from 0. If you want to know the first row you type a
print [Link], it is a location in square bracket 0. Will do that you will get the details which are
there in the first row.
(Refer Slide Time: 01:26)
So first if I want to know
hundredth row so printed [Link] 99. We knew that python count from zero. If I want to know
100th row you have to type 99. So it should be in Square bracket you can see the details in
the 100th row.
(Refer Slide Time: 01:42)

55

Suppose we want to know the last row in the data set. So print [Link] n equal to 1. If you type
n equal to – 1, it will not work, that we will see why if you want to know the last row simply
type to [Link] n equal to 1, you will get to know that what is the last two, So we will see that.
(Video Starts: 02:01)
Now we are going to use this command to see the last row, that is a detail about the last rows.
Now we can subset a multiple rows at a time. For example; there will be requirement we have
to select 100th row, 1st row 100th rows and 1000th rows. For that purpose you type this
command print [Link]. You see there are two square brackets 0, 99, 999, you will see what
output where getting. So type print [Link]. Yes, so we are able to see the 1st row, 100th row,
1000th row.

There is another way we can subset rows by row number by using this command iloc.
Previously loc, now we are going to use iloc. Suppose for type I want to get the 2nd row, if I
type print [Link] 1. I will get the details about the 2nd row. Okay? Yeah, this is a detail about
the 2nd row. Suppose I want to know 100th row by using iloc command so go there. Yes?
That is the details about the 100th row. You see that if I want to access the last row by using
iloc command.

So you can directly type print [Link] in squared bracket - 1. So that will be the details of the
last row. So what you can do we can open our Excel file you can verify what was the title, the
last row and soon.
(Video Ends: 04:27)

56
(Refer Slide Time: 04:27)

See then important note here


with iloc command. We can pass in the - 1 to get the last row, but same thing that we could
not do with loc. That is the difference between loc command and iloc command.
(Refer Slide Time: 04:42)
Suppose we want to get the first
100th and 1000th rows, using iloc command. So we are going to type this print [Link] 0, 99,
999. Let us see what answer we are getting.
(Video Starts: 04:58)
Yes? See, we have getting 1st, 100th and 1000th row.
(Video Ends: 05:21)

57
(Refer Slide Time: 05:21)

So far we are seeing subsetting


rows. Now we will see subsetting columns, the Python slicing syntax used a colon, colon
represents all the rows. If you have just a colon that attribute refers to everything. So if you
just want to get the first columns using a loc, or iloc syntax. We can write something like
[Link][ : , which column we need to refer, to subset the columns. (Refer Slide Time: 05:49)
The next slide I need to show
that we are going to subset the columns with loc, not the position of the colon. It is used to
select all rows.
(Refer Slide Time: 05:58)

58

You see that, subset equal


[Link]: , I want to see only two columns that is year and population. So when you type this
way you will get all the rows only two columns details that is year and population. You will
type this so when you type print [Link]. You can get the first 5 rows. So you will see
how it appearing.
(Video Starts: 06:26)
Subsets equal to Subset is object because from the df is the initial object which has all the
details. Now I am going to fetch only few columns from the df object that I am going to
saved in the name subset, subset is the object. So all the rows but I need only year column
and population column so I am going to type I want to see the first 5 rows, see that I am able
to see 1 st 5 rows, only for 2 cells. That is year and population. This is the way to get only 2
cells from the 2 columns from the Big Files.
(Video Ends: 07:21)
(Refer Slide Time: 07:21)
59

There is another example subset column with iloc, iloc will allow us to use integers - 1 will
select the last column. The same thing whatever we have seen in the previously so subset
equal to [Link]:, represents all the rows. Then [2 , 4 , - 1, then we can see by using this
command print [Link] 1st 5 rows.
(Video Starts: 07:46)
See that we are able to see the last column and the population column, life expectancy
column. You can open our Excel sheet you can verify whether we are getting the right answer
or not. (Video Ends: 08:26)
(Refer Slide Time: 08:26)
60
Sometime there is another way for subsetting columns by using the command called range.
First will make range of numbers we are going to save that range of a number in object called
small _ range, so small _ range equal to list range 5. Print small range will get 0, 1, 2, 3, and
4. Now this small _ range, object can be used to access the corresponding columns.
(Refer Slide Time: 08:57)

So if I type a subset equal to


[Link]:, small _ range I can get.
(Refer Slide Time: 09:04)
1st column, 2nd column, 3rd column, 4th column and 5th column, so we will try this. (Video
Starts: 09:09)

61
small_range is an object, we are going to create a range. Suppose we want to see what
small_range is. So it is up to 0 to 4, that means 1 to 5. Now we are going to subset using that
object called small _ range using ilocation command. [Link]:small_ range we see that
here we are able to see 5 column that is a country, year, population, continent and life
expectancy.
(Video Ends: 10:21)
(Refer Slide Time: 10:22)

So far we have seen subsetting only rows and columns. Now we are going to subset rows and
columns simultaneously. For example; using loc command so if you type print [Link] 42
countries. We can check in the 42 label in country columns. What is the cell name, there cell
name is Angola. Will try this.
(Video Starts: 10:47)
Going to see in that file in 42nd label in country column what value is there so that is an
Angola, Yes?
(Video Ends: 11:09)
(Refer Slide Time: 11:09)
62

Yes, we can see what is in the


using the same ilocation we can see in 42nd label in 0th column. Now we can represent
column also with 0 columns, what value it is, you will see that. You can verify you have to
get to the answer. You can open the Excel file. You can verify we are correctly accessing the
cell or not.
(Video Starts: 11:29)
Print [Link] in 42nd label 0th column what is the value it is Angola.
(Video Ends: 11:46)
(Refer Slide Time: 11:46)

Next we can subset multiple


rows and columns. For example; get the 1st, 100th and 100th rows from the 1st, 4th and 6th
column. So now we are going to, simultaneously we are going to fetch

63
rows and columns and corresponding cells. So print to [Link] 0, 99, 999. Similarly column
labels is 0, 3, 5. Let us see what answer.
(Video Starts: 12:13)
This accessing rows and columns are very important functions because nowadays data file
comes with a lot of rows and lot of columns. We need not use all the columns, all the rows for
further analysis. Sometimes we need only specific rows or specific columns. So these basic
commands will help you, how to access a particular rows and columns, that will be very
useful when we do further analysis using Python. Yeah? This is the value so that means 1st
row, 100th row 1000th row, 1st column and soon.
(Video Starts: 13:08)
(Refer Slide Time: 13:08)

And there is another way if you


use the column names directly it makes the code a bit easier to read. In terms of number and
so you see number column. If you use for representing column, if you use column name we
can see what is there, so simply type the column name. So we use this command, print [Link]
0, 99, 999. Then directly will type the column name country, life expectancy, gdpPercap you
see there is a square bracket here.
(Video Starts: 13:36)
That you have to do as the same that Life capital Exp, Yes? This is because country, life
expectation this is the easy way to because we cannot remember column name. (Video Ends
14:48)

64
(Refer Slide Time 14:49)
This was not only that instead
of see suppose if you put a 10 column 13 that corresponding rows will be displayed. So print
[Link] 10 to 13, the 10th row 11th, row 12th, row 13th, row will be shown and in columns
country and life expectancy and gdpPercap so we will try this command. (Video Starts:
15:11)
That means we can see the range of rows at a time. You are able to see the 10th row, 11th,
12th and 13th.
(Video Ends: 16:17)
(Refer Slide Time: 16:17)

Okay? Next see print df. head


st
we can see we can able to see 1 , 10 rows.

65
(Refer Slide Time: 16:23)
The 10th row some time for each
year in our data what was the average life expectancy. To answer this question we need to
split our data into parts per year and then we can get the life expectancy column and calculate
the mean.
(Refer Slide Time: 16:38)

So what is happening there is a


command which I go to use called groupby, we look at the data it is not grouped. So when
you use this command print [Link] year,and life expectancy and corresponding mean.
The mean of the on the in the year 1952, the mean of the life expectancy variable is 49.05. In
57, 51.09. We look at the data; it is not in this order. So the groupby by year

66
this command is grouping all the values, with respect to year. So we will see what is the
answer for this, we will verify this.
(Video Starts: 17:15)
When you open that Excel file you will see that the Excel file will be in some other form it is
not grouped by year, different years are appearing at different places. So this command that is
a group by will help you to group the data in year wise. Yes, you see that you are able to get
1952 the life expectancy was 49 years you see that when you look at this data. When year
increases the life expectancy year also increases due to advancement of medical facility
available and the standard of life is also increasing.
(Video Ends: 18:42)
(Refer Slide Time: 18:42)

Now, we can form a stacked


table. Stacked table is using the group by command. So you type this multi _ group _ variable
= df . \ . See the \ represents to breaking the command we can use \. Otherwise you can write
straightaway also no problem. [Link] by year, continent, life expectancy,gdp per capita,
then we can find the mean. Then we will get this output for that means in 1952, in Africa, the
life expectancies 39, in America 53, in Asia 46 in Europe 64 will try this command.
(Video Starts: 19:28)
When we takes these command you will get an output, that is a stacked table. That is very
useful for interpreting the whole dataset, is kind of a way of summarizing the data in the form
of table.

67
Multi_group. You see that now year wise. It is very, very useful command it is year by 1952,
some country Africa. What was the average year 1957 Africa. We see that if you look at only
the Africa data. 52 to 39 in 57 41, in 62 43, in 67 45, see that we can interpret this way, by
looking at the, this table. Suppose you have to flatten this.
(Video Ends 21:24)
(Refer Slide Time: 21:24)

If, you need to flatten the data


frame. You can use this reset underscore index method, just to type flat = multi _ group _ var
. reset _ index. Then you see now the data is again. Now it is flattened. The same data set,
which was it in the table form now it in the simple learned form. So we will try this comment.
(Video Starts 21:48)
This is what you are doing the data manipulation, because from the big data file, we have to
learn this kind of fundamental data manipulation methods that will be very useful, in coming
classes. So able to use reset _ index command to flatten the, that stacked table. See that now
we can see first 15 rows. Now it is data is flattened into the normal form.
(Video Ends 22:41)
(Refer Slide Time: 22:41)

68

The next one is grouped


frequency counts. By using nunique command, we can get a count of unique values on the
panda series. So when you type print df. groupby continent, country. nunique, you can get
unique values that means frequency. Okay, will try this command. (Video Starts: 23:04)
Print, See Africa 52, America is 25, Asia 33. When you look at the data, again, you go to
excel,Excel data you can interpret what is the 52 means, what is the America 25 and soon.
(Video Ends: 23:49)
(Refer Slide Time: 23:49)
Now, some basic plot a way to
construct two things one is year and life expectancy. So we are going to create a new object
that is called Global _ yearly_ life _expectancy. By grouping year

69
and life expectancy, with respect to its mean. Then we are going to print it. So you are going
to get two values one is year. Next one is life expectancy. That is a mean life expectancy, you
will see this.
(Video Starts: 24:17)
There is a new object. The object name is called Global _ yearly _ life expectancy. Yes, see
that year, and supposed we want to plot it. We will see we are going to plot this data, how we
are going to plot it.
(Video Ends: 25:28)
(Refer Slide Time: 25:28)

Simply, just that object name.


plot. That automatically takes this was output, which I got is in x axis in a year, in y axis,
average life expectancy. We will run this.
(Video Starts 25:40)
So, what this data says that, when the year 1950 - 2000 you see when the year increases, the
life expectancy also increases.
(Video Ends: 26:07)
(Refer Slide Time 26:07)

70

Just we have seen only the


simple plot, in coming classes, we will see some of the visual representation of the data. We
are going to see a histogram, frequency polygon, ogive curves,pie chart, stem and leaf plot
and pareto chart and scatter plot .
(Refer Slide Time: 26:21)

Suppose, this is the data, see what is there in East, west, north. In column first quarter, second
quarter, third quarter, fourth quarter.
(Refer Slide Time: 26:30)
71

Suppose the very easiest way is


the graph. By using this is called bar graph, bar chart. Bar chart is different regions are
labeled as different colors. This is a method of visual representation of the data. If you look at
this, the eastern side in third quarter, there are more sales. Okay. (Refer Slide Time: 26:53)

The another way to represent


visually, the data is pie chat, is the first quarter, third quarter. You look at this, third quarter,
which is in blue in color. There are more sales. And most importantly the pie chart, we can
get pie chart only for categorical variable. The variable is continuous, you cannot use bar
chart, you cannot use pie chart. So the pie chart is used only for categorical variable. That is
for only count data.
(Refer Slide Time: 27:31)

72
The another one is the Multiple
bar chart. This is another way to represent the data visually. (Refer Slide Time: 27:39)

Another one is a simple


pictogram.
(Refer Slide Time: 27:43)

73
See, this is the frequency table.
(Refer Slide Time: 27:25)

See, next one is frequency polygon. This figure is drawn from the previous table, which was
shown in the previous slide. So below 20 around 13,14. This represents frequency polygon.
When you connect the midpoint, you see that this is the. This is called frequency polygon.
Then the, this one is the cumulative frequency. It is not always, you cannot connect the
midpoint, you have to be very careful with the data is continuous, then only you can connect
one this bar. The data is not continuous, you cannot connect it.
(Refer Slide Time: 28:24)

74
Next one is a histogram .The histogram was constructed from the given table. You see.
(Refer Slide Time: 28:30)

The lower limit of the table values is going to in x axis. The frequency is shown in the y axis.
You see that this is data in continuous data. Okay, that was histogram. The purpose of
histogram is, the histogram will give you a rough idea what is the nature of the data whether,
what kind of distribution it follows. Whether it is following bell shaped curve, whether the
data is skewed right or skewed left.
(Refer Slide Time: 29:03)

75
Next one is the frequency
polygon which I have shown you. If, the midpoint of histogram are connected then there is
called frequency polygon. Because, the frequency polygon is used to know the trend.
(Refer Slide Time: 29:20)

Trend of the data. The next one


is ogive curve. This is cumulative frequency curve .So what is happening in the, for example
20- under 30, the upper limit of the interval is taken the x axis, the cumulative frequency is
taken in the y axis. For example, the first interval.20 - 30. So 30 the upper interval is 6. For
40, upper interval is to 24, that is marked.

76

You might also like