0% found this document useful (0 votes)
0 views18 pages

DATA SCIENCE Module 2

The document provides a comprehensive overview of statistical analysis, covering its definition, key steps, and the importance of both statistical and non-statistical methods in data analysis. It outlines major categories of statistics, including descriptive and inferential statistics, and explains measures of central tendency and dispersion for both populations and samples. Additionally, it highlights the significance of statistical analysis across various fields such as healthcare, business, and social sciences.

Uploaded by

Ashique Iqbal
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
0 views18 pages

DATA SCIENCE Module 2

The document provides a comprehensive overview of statistical analysis, covering its definition, key steps, and the importance of both statistical and non-statistical methods in data analysis. It outlines major categories of statistics, including descriptive and inferential statistics, and explains measures of central tendency and dispersion for both populations and samples. Additionally, it highlights the significance of statistical analysis across various fields such as healthcare, business, and social sciences.

Uploaded by

Ashique Iqbal
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

MITM, Jamshedpur STATISTICAL ANALYSIS Module: 02

STATISTICAL ANALYSIS:-

Introduction to statistics, Statistical and Non-Statistical Analysis:-

Introduction to Statistics:

Statistics is a branch of mathematics and a fundamental tool for data analysis. It involves collecting,
organizing, interpreting, presenting, and summarizing data to make meaningful inferences and decisions.
Statistics plays a crucial role in various fields, including science, business, economics, social sciences,
healthcare, and more. Its primary purpose is to help us understand and describe the world by using data-
driven methods.

Statistical Analysis: Statistical analysis is the process of using statistical techniques to examine and
analyze data. It aims to uncover patterns, relationships, and insights within datasets. Statistical analysis
involves
several key steps:
1. Data Collection: The first step in statistical analysis is to gather relevant data. This can be done through
surveys, experiments, observations, or by acquiring existing datasets.
2. Data Cleaning and Preparation: Raw data often contain errors, outliers, or missing values. Data
cleaning involves correcting these issues to ensure the dataset is accurate and complete. Data
preparation includes transforming, aggregating, and organizing the data for analysis.
3. Descriptive Statistics: Descriptive statistics are used to summarize and describe the main features of a
dataset. Common measures include mean (average), median (middle value), mode (most frequent
value), and measures of variability like standard deviation and range.
4. Inferential Statistics: Inferential statistics involve making predictions or inferences about a population
based on a sample of data. This includes hypothesis testing, confidence intervals, and regression
analysis.
5. Data Visualization: Data visualization techniques, such as graphs and charts, help in presenting data in
a visually understandable format. This aids in communicating findings and patterns effectively.
6. Interpretation: After conducting statistical analysis, it's essential to interpret the results and draw
meaningful conclusions. This often involves assessing the statistical significance of findings and
considering their practical implications.

Non-Statistical Analysis: Non-statistical analysis refers to data analysis methods that do not rely on
formal statistical techniques. While statistics is a powerful tool, there are situations where non-statistical
approaches are more appropriate. Here are some examples of non-statistical analysis methods:

1. Qualitative Analysis: Qualitative research methods focus on exploring and understanding the
underlying reasons, motivations, and contexts behind phenomena. This can involve techniques like
interviews, content analysis, and thematic analysis.
2. Expert Opinion: In some cases, experts in a particular field may provide insights and analysis based on
their knowledge and experience, without relying on statistical data.

Mr. Koushik Dey (Assistant professor)


Department of Computer Science & Engineering Page No. 1
MITM, Jamshedpur STATISTICAL ANALYSIS Module: 02
3. Case Studies: Case studies involve in-depth examination of a single or a few cases to gain insights into
specific situations. This approach is often used in fields like psychology, sociology, and
business.
4. Historical Analysis: Historical analysis examines past events, trends, and patterns without necessarily
using statistical methods. It relies on historical records, documents, and narratives.
5. Observational Analysis: Observational studies involve direct observation of phenomena without
manipulation or formal statistical analysis. This is common in fields like ecology and ethnography.
6. Content Analysis: Content analysis is a method of examining textual, visual, or audio content to
identify patterns, themes, or sentiments. It is commonly used in fields like media studies and
communication research.
7. Market Research: Market research often uses non-statistical methods like surveys, focus groups, and
customer feedback to gather insights into consumer behaviour and preferences.
Non-statistical analysis methods are particularly useful when the research question is exploratory, when the
data are primarily qualitative or subjective in nature, or when the goal is to gain a deep understanding of a
specific phenomenon. These methods can complement statistical analysis and provide a more
comprehensive view of complex issues.
In summary, statistical analysis is a rigorous and systematic approach to analysing data, while non-statistical
analysis encompasses a range of methods that may not involve formal statistics but can still provide
valuable insights and information. The choice between statistical and non-statistical analysis depends on
the research question, data availability, and the specific goals of the analysis.

Importance of Statistical Analysis:-


Statistical analysis is of paramount importance in various fields and disciplines due to its numerous
advantages and contributions to decision-making, research, and problem-solving. Here are some key
reasons why statistical analysis is important:
1. Data Interpretation: Statistical analysis helps in interpreting large and complex datasets bysummarizing
the information into meaningful insights. It allows researchers and analysts to understand patterns, trends,
and relationships within the data.
2. Inference and Generalization: Statistical analysis enables researchers to make inferences about
populations based on samples. By conducting hypothesis tests and calculating confidence intervals, one can
draw conclusions and generalize findings to larger populations.
3. Objectivity: Statistical methods provide an objective and systematic way to analyze data. This reduces
the potential for bias and subjectivity in decision-making and research.
4. Risk Assessment: In fields such as finance and insurance, statistical analysis is crucial for risk assessment.
It helps in predicting and managing financial risks, determining insurance premiums, and making
investment decisions.

Mr. Koushik Dey (Assistant professor)


Department of Computer Science & Engineering Page No. 2
MITM, Jamshedpur STATISTICAL ANALYSIS Module: 02
5. Quality Improvement: Statistical tools like Six Sigma and quality control charts are used in manufacturing
and process industries to monitor and improve product quality, leading to cost savings and customer
satisfaction.
6. Evidence-Based Decision-Making: In healthcare, education, and policy-making, statistical analysis is vital
for evaluating the effectiveness of interventions, treatments, and policies. It ensures that decisions are
based on empirical evidence rather than intuition or anecdotal evidence.
7. Predictive Modelling: Statistical modelling, including regression analysis and machine learning, is used
for making predictions and forecasts. This is valuable in fields like marketing (predicting consumer
behaviours), weather forecasting, and stock market analysis.
8. Scientific Research: In scientific research, statistical analysis is essential for testing hypotheses, drawing
conclusions, and publishing research findings. It ensures that research is rigorous and replicable.
9. Market Research: Businesses use statistical analysis to understand consumer preferences, market trends,
and customer behavior. This information guides product development, marketing strategies, and pricing
decisions.
10. Social Sciences: Statistical analysis is widely used in sociology, psychology, political science, and other
social sciences to study human behavior, attitudes, and social phenomena. It helps in drawing meaningful
conclusions from survey data and experiments.
11. Environmental Studies: Environmental scientists use statistical analysis to assess environmental impact,
analyse pollution data, and model ecological systems. This is critical for conservation efforts and
environmental policy-making.
12. Public Health: Epidemiologists use statistical analysis to track disease outbreaks, study the effectiveness
of public health interventions, and inform healthcare policy decisions.
13. Education: Educational researchers use statistical analysis to evaluate teaching methods, assess student
performance, and identify factors that influence learning outcomes.

In summary, statistical analysis is a versatile and powerful tool that plays a pivotal role in research, decision-
making, and problem-solving across a wide range of disciplines. It provides a systematic, objective, and data-
driven approach to understanding, analyzing, and making informed choices based on empirical evidence.
Its importance extends to both academic and practical applications, making it an indispensable skill in
today's data-driven world.

Example of Statistical Analysis:-

Look at the standard deviation sample calculation given below to understand more about statistical
analysis.

The weights of 5 pizza bases in cms are as follows:

Mr. Koushik Dey (Assistant professor)


Department of Computer Science & Engineering Page No. 3
MITM, Jamshedpur STATISTICAL ANALYSIS Module: 02

Particulars (Weight in cms) Mean Deviation Square of Mean Deviation

9 9-6.4 = 2.6 (2.6)2 = 6.76

2 2-6.4 = - 4.4 (-4.4)2 = 19.36

5 5-6.4 = - 1.4 (-1.4)2 = 1.96

4 4-6.4 = - 2.4 (-2.4)2 = 5.76

12 12-6.4 = 5.6 (5.6)2 = 31.36

Calculation of Mean = (9+2+5+4+12)/5 = 32/5 = 6.4

Calculation of mean of squared mean deviation = (6.76+19.36+1.96+5.76+31.36)/5 = 13.04

Sample Variance = 13.04

Standard deviation = √13.04 = 3.611

Major Categories of Statistics:-


Statistics is a broad field that can be categorized into several major categories based on its various
applications and methodologies. Here are some of the major categories of statistics:

1. Descriptive Statistics: Descriptive statistics involve summarizing and presenting data in a meaningful
way. Common techniques in this category include measures of central tendency (mean, median, mode),
measures of dispersion (range, variance, standard deviation), and graphical representations like histograms,
bar charts, and scatterplots.
2. Inferential Statistics: Inferential statistics is concerned with making inferences or drawing conclusions
about a population based on a sample of data. It includes techniques such as hypothesis testing,
confidence intervals, and regression analysis.
3. Probability Theory: Probability theory deals with the study of randomness and uncertainty. It provides
the foundation for statistical inference and includes concepts such as probability distributions, random
variables, and probability laws.
4. Bayesian Statistics: Bayesian statistics is a branch of statistics that uses Bayesian probability theory to
update beliefs or make predictions based on prior knowledge and observed data. It's particularly useful in
situations with limited data.

Mr. Koushik Dey (Assistant professor)


Department of Computer Science & Engineering Page No. 4
MITM, Jamshedpur STATISTICAL ANALYSIS Module: 02
5. Non-parametric Statistics: Non-parametric statistics are used when the assumptions of traditional
parametric statistics are not met. These methods do not rely on specific distributional assumptions and
include techniques like the Mann-Whitney U test and the Wilcoxon signed-rank test.
6. Multivariate Statistics: Multivariate statistics involve the analysis of data with more than one variable.
Techniques in this category include multivariate regression, principal component analysis, and factor
analysis.
7. Time Series Analysis: Time series analysis deals with data that is collected over time, such as stock
prices or weather data. It includes methods for modeling and forecasting time-dependent patterns and
trends.
8. Survival Analysis: Survival analysis is used to analyze time-to-event data, such as the time until a patient
recovers from a disease or the time until a machine fails. It uses techniques like Kaplan-Meier survival
curves and Cox proportional hazards models.
9. Experimental Design: Experimental design is about planning and conducting experiments to investigate
the effects of one or more variables on an outcome. It includes techniques like randomization and control
groups to minimize bias and draw valid conclusions.
10. Statistical Software and Tools: This category focuses on the practical application of statistics using
software and tools like R, Python, SPSS, Excel, and specialized statistical software packages.
11. Biostatistics: Biostatistics is the application of statistical methods to biological and medical data. It plays
a crucial role in clinical trials, epidemiology, and healthcare research.
12. Business and Economic Statistics: Business and economic statistics are used to analyze economic data,
market trends, and financial performance. It includes methods for forecasting, risk assessment, and
economic modeling.
These major categories of statistics cover a wide range of techniques and applications, making statistics a
versatile and essential field in various disciplines and industries.

Mr. Koushik Dey (Assistant professor)


Department of Computer Science & Engineering Page No. 5
MITM, Jamshedpur STATISTICAL ANALYSIS Module: 02

Population And Sample Measure Of Central Tendency And Dispersion:-

Population and sample statistics are used to describe and summarize data, and they include measures of
central tendency and measures of dispersion. These measures help us understand the characteristics and
variability of a dataset, whether we are working with an entire population or a sample from that
population.

Population Measures:
1. Population Mean (μ): The population mean is the average of all values in the entire population. It is
calculated by adding up all the values and dividing by the total number of values.

2. Population Median (Med): The population median is the middle value when all data points are
arranged in ascending or descending order. If there is an even number of data points, it's the average of
the two middle values.

3. Population Mode (Mo): The population mode is the value that appears most frequently in the
population dataset.
4. Population Variance (σ^2):The population mode is the value that appears most frequently in the The
population variance measures how data points deviate from the population mean. It is calculated as the
average of the squared differences between each data point and the population mean.

5. Population Standard Deviation (σ): The population standard deviation is the square root of the variance
and provides a measure of how spread out the data is.

Sample Measures:
When working with a sample from a population, we use slightly different formulas to estimate population
parameters:
1. Sample Mean (x¯ ): The sample mean is the average of all values in the sample. It is calculated in the
same way as the population mean.

Mr. Koushik Dey (Assistant professor)


Department of Computer Science & Engineering Page No. 6
MITM, Jamshedpur STATISTICAL ANALYSIS Module: 02

2. Sample Median (Med): The sample median is calculated the same way as the population median, using
the data points in the sample.
3. Sample Mode (Mo): The sample mode is the value that appears most frequently in the sample dataset.
4. Sample Variance (s^2): The sample variance measures how data points deviate from the sample mean. It
is calculated as the average of the squared differences between each data point and the sample mean, with
a correction factor (n-1) to account for the degrees of freedom.

5. Sample Standard Deviation (s): The sample standard deviation is the square root of the sample variance
and provides a measure of how spread out the sample data is.

It's important to note that when dealing with a sample, the use of (n-1) in the variance formula is known as
Bessel's correction, and it accounts for the fact that we are estimating population parameters based on a
sample rather than the entire population.
These measures of central tendency (mean, median, mode) and measures of dispersion (variance, standard
deviation) are fundamental tools in statistical analysis and are used to summarize and understand datasets,
whether they represent populations or samples from those populations.

Question: Consider the following data set representing the ages of 10 individuals in a population:
38, 42, 36, 45, 49, 32, 50, 40, 44, 41
1. Calculate the mean, median, and mode of the population.
2. Calculate the population variance and standard deviation.
Solution:
1. Calculate the Mean, Median, and Mode of the Population:
Mean: Mean (μ) = (Sum of all values) / (Number of values) μ
= (38 + 42 + 36 + 45 + 49 + 32 + 50 + 40 + 44 + 41) / 10
Calculate the mean.
Median: To calculate the median, first arrange the data in ascending order and then find the middle value.
Since there are 10 values, the median will be the average of the 5th and 6th values when arranged in
ascending order.
Mode: The mode is the value that appears most frequently in the data set.

Mr. Koushik Dey (Assistant professor)


Department of Computer Science & Engineering Page No. 7
MITM, Jamshedpur STATISTICAL ANALYSIS Module: 02
2. Calculate the Population Variance and Standard Deviation:
Population Variance (σ²): Variance measures the spread or dispersion of data points around the mean. The
population variance is calculated as the average of the squared differences between each data point and
the mean.
σ² = Σ(xi - μ)² / N Where xi is each data point, μ is the mean, and N is the number of data points.
Population Standard Deviation (σ): Standard deviation is the square root of the variance and is a measure
of how spread out the data is.
σ = √σ²
Calculate the population variance and standard deviation.

Moments: -
In statistics and probability theory, moments are mathematical quantities that provide information about
the shape and characteristics of a probability distribution or a dataset. Moments are used to describe
various aspects of a distribution, including its center, spread, skewness, and kurtosis. There are several
types of moments, with each type providing different insights into the data. The most commonly used
moments are:
1. Zeroth Moment (Moment about the Origin):
The zeroth moment is simply the constant 1. It doesn't provide much information about the data but is
used in some calculations, especially in generating moment-generating functions.
2. First Moment (Mean or Expectation):
The first moment is the average or expected value of a dataset. It is the center of the distribution and is
often denoted by E(X) or μ for a random variable X. For a discrete dataset, it can be calculated as:

For a continuous random variable, the mean is calculated using integration:

3. Second Moment (Variance):


The second moment measures the spread or dispersion of a dataset. It is often denoted by Var(X) or σ^2
for a random variable X. The variance is calculated as:

Mr. Koushik Dey (Assistant professor)


Department of Computer Science & Engineering Page No. 8
MITM, Jamshedpur STATISTICAL ANALYSIS Module: 02

4. Third Moment (Skewness):


The third moment measures the asymmetry or skewness of a distribution. A positively skewed distribution
has a long tail on the right, while a negatively skewed distribution has a long tail on the left. The skewness is
often denoted by Skew(X) and can be calculated as:

5. Fourth Moment (Kurtosis):


The fourth moment measures the "tailedness" or kurtosis of a distribution. It tells us about the heaviness of
the tails compared to a normal distribution. Kurtosis is often denoted by Kurt(X) and can be calculated as:

Kurtosis is often compared to the kurtosis of a normal distribution (which has a kurtosis of 3), leading to
discussions of "excess kurtosis" (Kurt(X) - 3).
Higher-order moments beyond the fourth are less commonly used and provide more detailed information
about the shape of a distribution.
Moments are fundamental in probability theory and statistics because they help quantify and describe the
key characteristics of data distributions, making it easier to compare and analyze datasets. They are often
used in applications such as hypothesis testing, risk assessment, and data modeling.
Example Data: Consider the following dataset of exam scores for a class of 8 students:

{75,80,85,90,95,100,105,110} {75,80,85,90,95,100,105,110}

We'll calculate the first four moments:

1. Zeroth Moment (0th moment or Moment about the Origin):

The zeroth moment is simply the count of data points.


In this case, it's n=8 because there are 8 data points.

2. First Moment (1st moment or Mean):

The first moment is the mean (average) of the data.


Calculate the mean as follows-

Mr. Koushik Dey (Assistant professor)


Department of Computer Science & Engineering Page No. 9
MITM, Jamshedpur STATISTICAL ANALYSIS Module: 02
3. Second Moment (2nd moment or Variance):

The second moment is the variance, which measures the spread or dispersion of the data.
First, calculate the squared differences from the mean for each data point:

Then, calculate the variance as the average of these squared differences.


After performing these calculations, you should find the variance. In this case, it's approximately 170.

4. Third Moment (3rd moment or Skewness):

The third moment measures the skewness or asymmetry of the data distribution.
Calculate the skewness using the formula:

Here, the standard deviation can be calculated as the square root of the variance from the second
moment.
After performing these calculations, you should find the skewness.

5. Fourth Moment (4th moment or Kurtosis):

The fourth moment measures the kurtosis or tailedness of the data distribution.
Calculate the kurtosis using the formula:

The standard deviation should still be calculated as the square root of the variance from the second
moment.
After performing these calculations, you should find the kurtosis.

By applying these formulas to the given dataset, you can calculate the values of the first four moments.
They describe the dataset's central tendency (mean), spread (variance), skewness (third moment), and
tailedness (kurtosis).

Skewness And Kurtosis :-


Skewness And Kurtosis:- Skewness and kurtosis are two statistical measures used to describe the shape of a
probability distribution or the distribution of data in a dataset. They provide insights into the departure of a
dataset from a normal distribution (bell-shaped curve) and can help statisticians and data analysts.
Skewness is a measure of the asymmetry of a probability distribution or dataset. It helps you understand
the direction and degree of skew (departure from symmetry) in the data. Here are some key points about
skewness:

Mr. Koushik Dey (Assistant professor)


Department of Computer Science & Engineering Page No. 10
MITM, Jamshedpur STATISTICAL ANALYSIS Module: 02
1. Direction of Skewness:

• Negative Skewness (Left-Skewed): When the distribution has a long left tail and most of the data
points are concentrated on the right side. The mean is typically less than the median in this case
because the extreme values on the left side pull the mean in that direction.
• Positive Skewness (Right-Skewed): When the distribution has a long right tail, and most data points
are concentrated on the left side. The mean is usually greater than the median in this case because
the extreme values on the right side pull the mean in that direction.
• Zero Skewness (Symmetric): In a perfectly symmetric distribution, the mean and median are equal,
and the skewness is zero.
2. Calcultion of Skewness:

• One common formula for skewness is based on the third standardized moment, as mentioned in the
previous response. However, there are other methods and formulas for calculating skewness,
including Pearson's First Coefficient of Skewness.
3. Interpretation

• Skewness provides insights into the shape and distribution of data.


• Understanding skewness helps in making decisions about data transformations and selecting
appropriate statistical tests.
• For financial data, skewness can indicate whether returns are more likely to be positive (right-
skewed) or negative (left-skewed)
Mathematically, skewness can be calculated using various formulas, with one common formula being based
on the third standardized moment:

Kurtosis:
• Kurtosis measures the "tailedness" of the probability distribution or dataset. It indicates how heavy
the tails of the distribution are compared to a normal distribution.
• There are two common measures of kurtosis: excess kurtosis and sample kurtosis.
• Excess kurtosis: It subtracts 3 from the sample kurtosis to make it zero for a normal distribution.
Positive excess kurtosis indicates heavier tails than a normal distribution, while negative excess
kurtosis indicates lighter tails.
• Sample kurtosis: This is the fourth standardized moment of the data.
High kurtosis (positive) means that the dataset has heavy tails and is more peaked around the mean, while
low kurtosis (negative) means lighter tails and a flatter peak. Mathematically, sample kurtosis is calculated
as:

Mr. Koushik Dey (Assistant professor)


Department of Computer Science & Engineering Page No. 11
MITM, Jamshedpur STATISTICAL ANALYSIS Module: 02
Kurtosis measures the "tailedness" of the probability distribution or dataset. It quantifies how data in the
tails of the distribution behaves compared to a normal distribution. Here are some key points about
kurtosis:
1 . Types of Kurtosis:
• Mesokurtic (Normal Kurtosis): When the kurtosis value is 3. This indicates that the distribution has the
same tail behavior as a normal distribution.
• Leptokurtic (Excess Kurtosis >3): When the distribution has fatter tails and a more peaked central region
compared to a normal distribution. It implies that extreme values are more likely to occur.
• Playkurtic (Excess Kurtosis <3): When the distribution has thinner tails and a flatter central region
compared to a normal distribution. It suggests that extreme values are less likely.
2. Calculation of Kurtosis:
• Kurtosis is typically calculated using the fourth standardized moment, as mentioned in the previous
response. The excess kurtosis subtracts 3 from this value to make it zero for a normal distribution.
3. Interpretation:
• High kurtosis (leptokurtic) indicates that the data has heavier tails and is more peaked around the mean.
This suggests that extreme values are more likely, and the distribution has higher volatility.
• Low kurtosis (platykurtic) indicates that the data has lighter tails and a flatter central region. Extreme
values are less likely, and the distribution has lower volatility.
Practical Applications:
1. Finance: Skewness and kurtosis are used in finance to understand the distribution of returns, assess
investment risk, and develop trading strategies.
2. Statistics: They help in choosing appropriate statistical tests, especially when data does not follow a
normal distribution.
3. Data Preprocessing: Skewness and kurtosis can guide data transformations (e.g., logarithmic
transformation) to make data more suitable for analysis.
4. Risk Assessment: In risk management, skewness and kurtosis are used to assess the probability of
extreme events in various fields like insurance and climate science.
5. Quality Control: They are used in quality control to evaluate the consistency and reliability of
manufacturing processes.
In summary, skewness and kurtosis are fundamental statistical measures that provide insights into the
shape and characteristics of data distributions. Understanding these measures is valuable in various fields
for data analysis, modeling, and decision-making.

Mr. Koushik Dey (Assistant professor)


Department of Computer Science & Engineering Page No. 12
MITM, Jamshedpur STATISTICAL ANALYSIS Module: 02
In summary, skewness and kurtosis are important statistical measures for understanding the shape of data
distributions. They can help in making decisions about which statistical techniques are appropriate for
analyzing a dataset and in identifying potential outliers or departures from normality. However, they should
be used in conjunction with other exploratory data analysis techniques for a more comprehensive
understanding of the data.

Example Data: Consider the following dataset of exam scores for a class of 20 students:
{75,80,85,90,95,100,105,110,115,120,125,130,135,140,145,150,155,160,165,170}
{75,80,85,90,95,100,105,110,115,120,125,130,135,140,145,150,155,160,165,170}
We'll calculate both skewness and kurtosis for this dataset.
Skewness Calculation:
To calculate skewness, we'll use the formula mentioned earlier:

1. Calculate the mean (average):

2. Calculate the median (middle value):


• Since we have 20 data points, the median is the average of the 10th and 11th values, which are 125 and
130.

3. Calculate the standard deviation:


• First, calculate the squared differences from the mean for each data point:

• Then, calculate the variance as the average of these squared differences.


• Finally, take the square root of the variance to get the standard deviation.
• After performing these calculations, we find that the standard deviation is approximately 30.35.
Now, we can calculate skewness:

Since the skewness is positive (greater than 0), this dataset is right-skewed.

Mr. Koushik Dey (Assistant professor)


Department of Computer Science & Engineering Page No. 13
MITM, Jamshedpur STATISTICAL ANALYSIS Module: 02
Kurtosis Calculation:
To calculate kurtosis, we'll use the formula for excess kurtosis:

1. Calculate the fourth power of the differences from the mean for each data point:

2. Calculate the sum of these fourth powers.


3. Plug these values into the kurtosis formula:

After performing these calculations, you will find the kurtosis value. A positive value indicates leptokurtic
(heavy-tailed) behavior, while a negative value indicates platykurtic (light-tailed) behavior. In this case, you
should find that the kurtosis is positive, indicating heavier tails compared to a normal distribution.

Correlation And Regression:-


Correlation and Regression are two fundamental concepts in statistics and data analysis. They are used to
explore and quantify the relationships between variables, predict outcomes, and make informed decisions.
Let's dive into each of these concepts with detailed explanations and numerical examples.
Correlation: Its measures the degree and direction of the linear relationship between two continuous
variables. It indicates whether and how much one variable changes when the other changes. Correlation is
often represented by the correlation coefficient, denoted as r.
Correlation Coefficient (r):
• r ranges from -1 to 1.
• r = 1 indicates a perfect positive correlation, meaning that as one variable increases, the other also
increases in a linear fashion.
• r = -1 indicates a perfect negative correlation, meaning that as one variable increases, the other decreases
in a linear fashion.
• r = 0 indicates no linear correlation between the variables.
• The sign of r indicates the direction of the correlation (positive or negative), while the absolute value ∣r∣
measures the strength of the correlation.

Mr. Koushik Dey (Assistant professor)


Department of Computer Science & Engineering Page No. 14
MITM, Jamshedpur STATISTICAL ANALYSIS Module: 02

Where:
• n is the number of data points.
• represents summation.
• x and y are the two variables.

Theoretical Distributions- Binomial, Poisson, Normal:-


Theoretical probability distributions, such as the Binomial, Poisson, and Normal distributions, are
fundamental concepts in statistics and probability theory. These distributions are used to model and
describe various real-world phenomena and provide insights into the behavior of random variables. Let's
explore each of these distributions in detail:

1. Binomial Distribution: The Binomial distribution models the number of successes (typically denoted as
"x") in a fixed number of independent Bernoulli trials, where each trial has only two possible outcomes
(success or failure) and the probability of success (denoted as "p") remains constant across all trials.
• Parameters:
• n: The number of trials.
• p: The probability of success on each trial.
• Probability Mass Function (PMF):

• Where represents the binomial coefficient, which is the number of ways to choose x
successes out of n trials.
• Mean (Expected Value): E(X)=np
• Variance: Var(X)=np(1−p)
• Use Cases: The Binomial distribution is used in scenarios where you have a fixed number of trials, each
with a binary outcome, and you want to calculate the probability of a specific number of successes.
Examples include coin flips, pass/fail rates, and the number of defective items in a sample.
2. Poisson Distribution:

Mr. Koushik Dey (Assistant professor)


Department of Computer Science & Engineering Page No. 15
MITM, Jamshedpur STATISTICAL ANALYSIS Module: 02
• The Poisson distribution models the number of events that occur in a fixed interval of time or space
when the events are rare and random, and the average rate of occurrence (denoted as "λ") is known.
• Parameter:
• λ: The average rate of events per interval.
• Probability Mass Function (PMF): !P(X=x)=x!e−λ⋅λx
• Mean (Expected Value): E(X)=λ
• Variance: Var(X)=λ
• Use Cases: The Poisson distribution is commonly used to model rare events such as accidents,
customer arrivals at a service center, or the number of emails received per hour.
3. Normal Distribution (Gaussian Distribution):
• The Normal distribution is a continuous probability distribution that is characterized by its bell-
shaped curve. It is widely used due to its symmetry and the Central Limit Theorem, which states that the
distribution of the sum (or average) of a large number of independent, identically distributed random
variables approaches a Normal distribution.
• Parameters:
• μ (Mu): The mean (center) of the distribution.
• σ (Sigma): The standard deviation (spread) of the distribution.
• Probability Density Function (PDF): f(x)=2πσ21⋅e−2σ2(x−μ)2
• Mean (Expected Value): E(X)=μ
• Variance: 2Var(X)=σ2
• Use Cases: The Normal distribution is frequently used in statistical inference and hypothesis testing.
Many real-world measurements, such as height, weight, IQ scores, and errors in scientific experiments,
tend to follow a Normal distribution.

Each of these theoretical distributions plays a critical role in statistical analysis, hypothesis testing, and
modeling real-world data. Understanding these distributions is fundamental for making statistical
inferences and predictions.

Mr. Koushik Dey (Assistant professor)


Department of Computer Science & Engineering Page No. 16
MITM, Jamshedpur STATISTICAL ANALYSIS Module: 02

Theoretical Distributions- Binomial, Poisson, Normal:-


Theoretical probability distributions, such as the Binomial, Poisson, and Normal distributions, are
fundamental concepts in statistics and probability theory. These distributions are used to model and
describe various real-world phenomena and provide insights into the behavior of random variables. Let's
explore each of these distributions in detail:
1. Binomial Distribution: The Binomial distribution models the number of successes (typically denoted as
"x") in a fixed number of independent Bernoulli trials, where each trial has only two possible outcomes
(success or failure) and the probability of success (denoted as "p") remains constant across all trials.

Parameters:

• n: The number of trials.


• p: The probability of success on each trial.
• Probability Mass Function (PMF):

• Where represents the binomial coefficient, which is the number of ways to choose x successes out of n
trials.
• Mean (Expected Value): E(X)=np
• Variance: Var(X)=np(1−p)
• Use Cases: The Binomial distribution is used in scenarios where you have a fixed number of trials, each
with a binary outcome, and you want to calculate the probability of a specific number of successes.
Examples include coin flips, pass/fail rates, and the number of defective items in a sample.
2. Poisson Distribution:
• The Poisson distribution models the number of events that occur in a fixed interval of time or space
when the events are rare and random, and the average rate of occurrence (denoted as "λ") is known.
Parameter:
• λ: The average rate of events per interval.
• Probability Mass Function (PMF): !P(X=x)=x!e−λ⋅λx
• Mean (Expected Value): E(X)=λ
• Variance: Var(X)=λ
• Use Cases: The Poisson distribution is commonly used to model rare events such as accidents, customer
arrivals at a service center, or the number of emails received per hour.

Mr. Koushik Dey (Assistant professor)


Department of Computer Science & Engineering Page No. 17
MITM, Jamshedpur STATISTICAL ANALYSIS Module: 02
3. Normal Distribution (Gaussian Distribution):
The Normal distribution is a continuous probability distribution that is characterized by its bell-shaped
curve. It is widely used due to its symmetry and the Central Limit Theorem, which states that the
distribution of the sum (or average) of a large number of independent, identically distributed random
variables approaches a Normal distribution.
Parameters:
• μ (Mu): The mean (center) of the distribution.
• σ (Sigma): The standard deviation (spread) of the distribution.
• Probability Density Function (PDF): f(x)=2πσ21⋅e−2σ2(x−μ)2
• Mean (Expected Value): E(X)=μ
• Variance: 2Var(X)=σ2
• Use Cases: The Normal distribution is frequently used in statistical inference and hypothesis testing. Many
real-world measurements, such as height, weight, IQ scores, and errors in scientific experiments, tend to
follow a Normal distribution.
Each of these theoretical distributions plays a critical role in statistical analysis, hypothesis testing, and
modeling real-world data. Understanding these distributions is fundamental for making statistical
inferences and predictions.

Mr. Koushik Dey (Assistant professor)


Department of Computer Science & Engineering Page No. 18

You might also like