0% found this document useful (0 votes)
87 views22 pages

Primary vs. Secondary Data Collection

Uploaded by

baiz.p.s77
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
87 views22 pages

Primary vs. Secondary Data Collection

Uploaded by

baiz.p.s77
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

MODULE I

Chapter 1
Concept of primary and secondary data, Methods of collection
Primary and Secondary Data
Primary Data
Data collected firsthand for a specific [Link], raw, and collected directly from the [Link]
reliable but can be time-consuming and costly.
Examples:Surveys,Interviews,Experiments,Observations.

Advantages of Primary Data:

High accuracy and reliability.​


Collected specifically for the intended research purpose.​
Up-to-date and relevant.

Disadvantages of Primary Data:

Expensive and time-consuming.​


Requires proper planning and execution.​
May involve bias if not collected properly.

Secondary Data
Data that has been previously collected by someone [Link] for analysis without collecting new
[Link] expensive but may be less accurate or outdated.

Examples:Government reportsBooks, journals, articles,Company records.

Advantages of Secondary Data:

Saves time and cost.​


Easily accessible.​
Useful for historical comparisons and trends.

Disadvantages of Secondary Data:

May be outdated or irrelevant.​


Accuracy depends on the original source.​
Limited control over how the data was collected.

Methods of Data Collection

A. Methods for Primary Data Collection


1.​ Surveys & Questionnaires – Structured set of questions for data collection.
2.​ Interviews – Face-to-face, phone, or online conversations for deeper insights.
3.​ Observations – Directly watching and recording behaviors or events.
4.​ Experiments – Conducting scientific or controlled tests to gather data.
5.​ Focus Groups – Discussions with a small group to gain qualitative insights.
B. Methods for Secondary Data Collection
1.​ Published Sources – Books, research papers, government reports.
2.​ Online Databases – Digital libraries, statistical reports, business records.
3.​ Internal Records – Company sales reports, financial statements.
4.​ Census & Government Reports – Population statistics, economic reports.

Methods of Data Collection


A. Methods of Primary Data Collection
There are a number of methods of collecting primary data, Some of the common methods are as
follows:
1. Interviews: Collect data through direct, one-on-one conversations with individuals. The
investigator asks questions either directly from the source or from its indirect links.

1.​ Direct Personal Investigation: The method of direct personal investigation involves
collecting data personally from the source of origin. In simple words, the investigator makes direct
contact with the person from whom he/she wants to obtain information. For example, direct contact
with the household women to obtain information about their daily routine and schedule.
2.​ Indirect Oral Investigation: In the indirect oral investigation method of collecting primary
data, the investigator does not make direct contact with the person from whom he/she needs
information, instead they collect the data orally from some other person who has the necessary
information. For example, collecting data of employees from their superiors or managers.
●​ Advantage: Provides real-time, natural data; no reliance on self-reported information.
●​ Disadvantage: Observer bias; limited to what can be seen; may influence subjects’ behavior.
●​ Suitable Use Case: Behavioral studies, user experience research.
2. Questionnaires: Collect data by asking people a set of questions, either online, on paper, or
face-to-face. In this method the investigator prepares a questionnaire to collect Information
through Questionnaires and Schedules , while keeping in mind the motive of the study . The
investigator can collect data through the questionnaire in two ways:
Mailing Method: This method involves mailing the questionnaires to the informants for the
collection of data. The investigator attaches a letter with the questionnaire in the mail to define the
purpose of the study or research.
Enumerator’s Method: This method involves the preparation of a questionnaire according to the
purpose of the study or research. However, in this case, the enumerator reaches out to the informants
himself with the prepared questionnaire.
●​ Advantage: Can reach a large audience quickly and cost-effectively.
●​ Disadvantage: Responses may be biased or inaccurate; low response rates.
●​ Suitable Use Case: Customer satisfaction surveys, market research.
3. Observations: The observation method involves collecting data by watching and recording
behaviors, events, or conditions as they naturally occur. The observer systematically watches and
notes specific aspects of a subject’s behavior or the environment, either covertly or overtly.
●​ Advantage: Provides real-time, authentic data without reliance on self-reported information.
●​ Disadvantage: Observer bias can influence the results, and the presence of an observer
might alter subjects’ behavior.
●​ Suitable Use Case: Studying user interactions with a product in a natural setting, monitoring
wildlife behavior, or assessing classroom dynamics.
4. Experiments: The experiment method involves manipulating one or more variables to determine
their effect on another variable, within a controlled environment. Researchers create two groups
(control and experimental), apply the treatment or variable to the experimental group, and compare
the outcomes between the groups.

●​ Advantage: Allows for the establishment of cause-and-effect relationships with high


precision.
●​ Disadvantage: Experiments can be artificial, limiting the ability to generalize findings to
real-world settings, and they can be resource-intensive.
●​ Suitable Use Case: Testing the efficacy of a new drug, assessing the impact of a new
teaching method, or evaluating the effect of a marketing campaign.
5. Focus Group: The focus group method involves gathering a small group of people to discuss a
specific topic or product, facilitated by a moderator. A group of 6-12 participants engages in a
guided discussion led by a moderator who asks open-ended questions to elicit opinions, attitudes,
and perceptions.

●​ Advantage: Provides in-depth insights and diverse perspectives through interactive


discussions, revealing the reasoning behind participants’ thoughts and feelings.
●​ Disadvantage: Results can be influenced by dominant participants or groupthink, and the
findings are not easily generalizable due to the small, non-representative sample size.
●​ Suitable Use Case: Exploring customer attitudes towards a new product, gathering feedback
on a marketing campaign, or understanding public opinion on social issues.
6. Information from Local Sources or Correspondents: In this method, for the collection of data,
the investigator appoints correspondents or local persons at various places, which are then furnished
by them to the investigator. With the help of correspondents and local persons, the investigators can
cover a wide area.
Methods of Secondary Data Collection

1.​ Published Sources

Books, research papers, newspapers, and magazines.

Example: A student using a textbook for research.

2.​ Online Databases


Digital sources like Google Scholar, research journals, and statistical reports.

Example: A researcher using WHO’s website for global health data.

3.​ Internal Company Records

Businesses use their own records for sales, financial, and employee analysis.

Example: A company analyzing past sales reports to predict future demand.

4.​ Census & Government Reports


Data collected by government agencies, such as population census or employment statistics.

Example: A policymaker using census data to design public services

Differences between primary and secondary data


The term primary data refers to the data originated by the researcher for the first time. Secondary data is the
already existing data, collected by the investigator agencies and organisations earlier.

1.​ Primary data is real-time data whereas secondary data is one which relates to the past.

2.​ Primary data is collected for addressing the problem at hand while secondary data is collected for
purposes other than the problem at hand.

3.​ Primary data collection is a very involved process. On the other hand, the secondary data collection
process is rapid and easy.

4.​ Primary data collection sources include surveys, observations, experiments, questionnaires, personal
interviews, etc. On the contrary, secondary data collection sources are government publications, websites,
books, journal articles, internal records etc.

5.​ Primary data collection requires a large amount of resources like time, cost and manpower. Conversely,
secondary data is relatively inexpensive and quickly available.

6.​ Primary data is always specific to the researcher’s needs, and he controls the quality of research. In
contrast, secondary data is neither specific to the researcher’s need, nor he has control over the data quality.

7.​ Primary data is available in the raw form whereas secondary data is the refined form of primary data. It
can also be said that secondary data is obtained when statistical methods are applied to the primary data.

8.​ Data collected through primary sources are more reliable and accurate as compared to the secondary
sources.

Chapter 2
Measures of Central Tendency (Measures of Averages)
Definition or features
An average is a single figure which represents the whole [Link] lies between minimum and maximum values
in the given [Link] gives a general idea about whole data
Importance or uses or functions of averages
1.​ Averages give a general idea about the whole group
2.​ Averages can be used for summarising the data
3.​ Averages help comparison
4.​ Averages help in decision making
5.​ Averages constitute the basis of statistical analysis
6.​ Averages represent the universe
Essential properties or characteristics of a good average
1.​ It Should be clearly defined
2.​ Based on all observation of the data
3.​ Easy to calculate and simple to follow
4.​ Not be influenced by sampling fluctuation
5.​ Need to further algebraic treatment
6.​ Not affected by extreme items

Types of averages
There are five important averages
1.​ Arithmetic Mean (AM)
2.​ Median
3.​ Mode
4.​ Geometric mean (GM)
5.​ Harmonic mean (HM)
Arithmetic Mean
It is a simple mathematical average. It is the ratio between sum of numbers and their numbersIt can be divided
into two
1.​ Simple arithmetic mean
2.​ Weighted arithmetic mean
Simple arithmetic mean is mean of items which are given equal importance
Weighted arithmetic mean is mean of items which are given different weights in accordance with there
relative importance
Merits or Properties of AM

►​ Simple to understand

►​ Easily calculated

►​ Determined in most of the cases

►​ Based on all observations

►​ Capable of more algebraic treatment

►​ Stable
Demerits of AM

►​ Affected by extreme values

►​ Not suitable for averaging ratios and percentages

►​ Cannot be calculated for qualitative data

►​ Gives absurd results in some cases


Why is Arithmetic mean considered to be best average

►​ Simple to understand
►​ Easily calculated

►​ Determined in most of the cases

►​ Based on all observations

►​ Capable of more algebraic treatment

►​ Stable

►​ It is well defined

►​ It is an average of everyday life


Median
It is middle most item when items are arranged in ascending or descending [Link] divides the data into two
equal [Link] part has values less than [Link] part has values greater than [Link] is a positional
average
Merits of median

►​ Very simple measure

►​ Sometimes it can be located even by inspection

►​ Not affected by extreme items

►​ Suitable for even such data which are not capable of numerical expression

►​ Determined graphically

►​ Suitable average in the case of open end class series


Demerits of median

►​ Not based on all the observation

►​ Not capable of algebraic treatment

►​ Requires arraying (ascending or descending order)

►​ In the case of continues series interpolation formula is to be used the value thus obtained may give
only an approximate value
Mode
Most frequent item.
Merits of Mode

►​ Easy To Calculate & Simple To Understand

►​ It Is The Best Representative Of The Data

►​ Not Affected By The Value Of Extreme Items

►​ No Need Of Complete Data

►​ Useful For Both Quantitative & Qualitative Data

►​ It Can Be Determined Graphically With The Help Of Histogram.

Demerits of mode

►​ Uncertain and vague measure of central tendency .it is I’ll defined in some cases

►​ Not capable of further algebraic treatment

►​ Some items grouping becomes necessary to identify the modal value

►​ Not based on all the values of the series

►​ When the extreme value appears most frequently it becomes mode which gives an absurd result .(for
example 0)
Uses of mode

►​ In trading sector

►​ Average expenditure Average income

►​ In meteorological department:-Average rainfallAverage temperature


Positional averages and mathematical averages
Positional averages
Basis of the [Link] not change along with changes in other values
Mathematical averages
Obtained by applying mathematical [Link],GM,HM are mathematical [Link] along with
changes with value of the series
Compare and contrast Mean, Median and Mode
►​ Mean
Mathematical averages
Based on all observation
Capable of mathematical treatment
Cannot located graphically
Obtained by calculation
Affected by extreme items
Mean are defined in all cases

►​ Median and Mode


Median is middle item
. Mode is most frequently item
. Values found in the series
Graphically located
Obtained by inspection
Mode is I’ll defined in some cases
Median is well defined
Geometric mean
Nth root of the product of those n values
Uses of geometric mean

►​ To find the average% increases in sales production etc

►​ To find the index number (as they show relative changes)

►​ When large weights are given to small items and small weights to large items (when there are extreme
value)
Merits

►​ Geometric mean in based on all observation

►​ It is rigidly defined

►​ It is useful in averaging ratios and percentages It is not affected very much by extreme values

►​ It is capable of algebraic treatment


Demerits
It is difficult to understand
It can’t be computed when there are both negative and positive values or when one or more of the value is zero
Harmonic mean
It is the reciprocal of the mean of the reciprocals of the values
Merits

►​ Possess most of the important properties of a good average

►​ Based on all the observation


►​ Amenable to more algebraic treatment

►​ Not affected much by sampling fluctuation

►​ Rigidly defined

►​ Used to average relative value


Demerits

►​ Difficult to calculate

►​ Not easy to understand

►​ Gives greater weight to large items

►​ Can’t be computed when there are both negative and positive items or when an item is zero

►​ Not a popular average


Uses of Harmonic mean

►​ For computing the average rate of increase in profit

►​ Average rate of speed

►​ Average price

Positional Averages
A positional average is a measure of central tendency that is determined by the position of values in a data set.
The most common positional averages are the median and the mode.
Relationship between AM,GM,and HM
2
The relationship between AM, GM and HM is given by:AM x HM = 𝐺𝑀

Chapter 3
Measures of dispersion (Measures of Variability)

Definition
Refers to the variability in the size of items
Speaks about the scatter or spread of the values in a series
Indicates lack of uniformity in the size of items
Purpose of measuring variation
To test the reliability of an average
To compare the variability of two or more series
To exercise control over variability
Desirable properties of a good measure of dispersion
Simple to understand
Easy to calculate
Clearly defined
Based on all items of the distribution
Necessary to further algebraic treatment
Have sampling stability
Not affected by extreme items
Types of Measures of dispersion
Absolute measures
Relative measures
Absolute measure of dispersion
Measure of dispersion
Same units in which data are collected
Measure variability in a series

Types of absolute Measures of dispersion


Range
Quartile Deviation (QD)
Mean Deviation (MD)
Standard Deviation (SD)
Relative measures of dispersion (Coefficient of Dispersion)
To make comparison between two or more distributions
Expressed in different units
Ratio between relative measure and appropriate average
Greater relative measure ,greater variability

Types of relative measures of dispersion

Coefficient of range
Coefficient of Quartile Deviations
Coefficient of Mean Deviation
Coefficient of Variation
Comparison between relative measure and absolute measure of dispersion

Absolute measure Relative Measure


Computed directly from the data Computed from relative measure
Expressed in the same units. Free from units
Measure variability in a series. Measure variability in two series
Range
Difference between highest and lowest values in a series
Measure maximum variation in the values of a series
Uses of Range
Measure temperature fluctuation of patients
In quality control
In shares and interest rates
In weather forecasts
Merits of Range
Simplest measure of dispersion
Easily calculated
Understood even by an ordinary man
Demerits of Range
Not based on all items
Highly affected by sampling fluctuations
Cannot be computed in the open end distribution
Quartiles
In statistics, a quartile divides a data set into four equal parts. Each part represents a quarter of the data. There
are three main quartiles:

1.​ First Quartile (Q1): Also known as the lower quartile, it marks the 25th percentile of the data. This
means 25% of the data points are below this value.
2.​ Second Quartile (Q2): This is the median of the dataset, marking the 50th percentile. Half of the data
points lie below this value.
3.​ Third Quartile (Q3): Also known as the upper quartile, it marks the 75th percentile, meaning 75% of
the data points are below this value

Quartiles are useful for understanding the spread and central tendency of a dataset.
Quartile Deviation (QD/Semi interQuartile Range)
Half the distance between third and first quartile
Merits of QD
Easy to understand
Easy to calculate
Not affected by extreme values
Demerits the QD
Ignores extreme items
Not capable of more algebraic treatment
Mean Deviation (MD)
AM of deviations of all the values in a series from their average
All deviations are positive

Merits of MD
Very simple
Easily understood
Based on all items
Less affected by extreme values

Demerits of MD
Inaccuracy (Ignores +/-)
Not capable of further algebraic treatment
Not reliable measure
Uses of MD
In economic and social phenomena
In measure of wealth and income
Standard Deviation (SD)
Square root of mean of squares of deviations of all values of a
series from their AM
First used by Karl Pearson
Known as root mean square deviation
Measure extent of spread of values q
Have minimum value zero
Non negative

Merits of SD
Based on all values of series
Clear and definite measure
Not affected by sampling fluctuations
Need for further algebraic treatment

Demerits of SD
Difficult to calculate
Gives more weight to extreme values

Features of SD
Deviations are measured from AM
Signs e deviations are not ignored

SD considered to be the best measure of dispersion


Rigidly defined
Based on all observations
Need to more algebraic treatment
Possesses many mathematical properties
Not affected by sampling fluctuations
Does not ignore signs of Deviations
Find out for two or more groups
CV is based on SD
To find out statistical measure

Coefficient of Variation(CV)
Percentages ratio between SD and mean
Relative measure of dispersion
To determine variability of two or more series
Used for comparing variability of two series when their means are
different
CV is less that series is more stable

MODULE II
Chapter 2
STATISTICAL INFERENCE AND CORRELATION ANALYSIS
CORRELATION ANALYSIS
Definition
Two or more variables are said to be correlated if the change in one variable results in a corresponding change
in the other variable.
Correlation Coefficient
Correlation analysis is actually an attempt to find a numerical value to express the extent of relationship that
exists between two or more variables. The numerical measurement showing the degree of correlation
coefficient. Correlation coefficient ranges between-1 and +1

Classification of Correlation
Correlation can be classified in different [Link] following are the most important classifications.
[Link] and Negative correlation
2 Simple, partial and multiple correlation
3. Linear and Non-linear correlation
Positive and Negative Correlation
Positive Correlation
When the variables are varying in the same direction, it is called positive correlation. In other words, if an
increase in the value of one variable is accompanied by an increase in the value of another variable or if a
decrease in the value of one variable is accompanied by a decrease in the value of another variable, it is called
positive correlation.

Negative Correlation
When the variables are moving in opposite directions, it is called negative correlation. In other words, if an
increase in the value of one variable is accompanied by a decrease in the value of another variable or if a
decrease in the value of one variable is accompanied by an increase in the value of another variable, it is called
negative correlation.
Simple, Partial and Multiple Correlation Simple Correlation
In a correlation analysis, if only two variables are studied it is called simple correlation. Eg. The study of the
relationship between price & demand of a product or price and supply of a product is a problem of simple
correlation.
Multiple correlation
In a correlation analysis, if three or more variables are studied simultaneously, it is called multiple correlation.
For example, when we study the relationship between the yield of rice with both rainfall and fertilizer together,
it is a problem of multiple correlation.
Partial correlation
In a correlation analysis, we recognize more than two variables, but consider one dependent variable and one
independent variable and keep the other Independent variables as constant. For example, the yield of rice is
influenced by the amount of rainfall and the amount of fertilizer used. But if we study the correlation between
yield of rice and the amount of rainfall by keeping the amount of fertilizers used as constant, it is a problem of
partial correlation.
Linear and Non-linear correlation Linear Correlation
In a correlation analysis, if the ratio of change between the two sets of variables is same, then it is called linear
correlation.
For example, when 10% increase in one variable is accompanied by 10% increase in the other variable, it is
the problem of linear correlation
Here the ratio of change between X and Y is the same. When we plot the data in graph paper, all the plotted
points would fall on a straight line.
Non-linear correlation
In a correlation analysis if the amount of change in one variable does not bring the same ratio of change in the
other variable, it is called non linear correlation.
Here the change in the value of X does not being the same proportionate change in the value of Y.
This is the problem of non-linear correlation, when we plot the data on a graph paper, the plotted points would
not fall on a straight line.
Degrees of Correlation
Correlation exists in various degrees
1. Perfect positive correlation
If an increase in the value of one variable is followed by the same proportion of increase in other related
variables or if a decrease in the value of one variable is followed by the same proportion of decrease in other
related variables, it is a perfect positive correlation.
2. Perfect Negative correlation
If an increase in the value of one variable is followed by the same proportion of decrease in other related
variables or if a decrease in the value of one variable is followed by the same proportion of increase in other
related variables it is Perfect Negative Correlation. For example if 10% rise in price results in 10% fall in its
demand the correlation is perfectly negative. Similarly if 5% fall in price results in 5% increase in demand, the
correlation is perfectly negative.

3. Limited Degree of Positive Correlation


When an increase in the value of one variable is followed by a non-proportional increase in another related
variable, or when a decrease in the value of one variable is followed by a non-proportional decrease in another
related variable, it is called a limited degree of positive correlation.
For example, if a 10% rise in price of a commodity results in a 5% rise in its supply, it is a limited degree of
positive correlation. Similarly if 10% fall in price of a commodity results in 5% fall in its supply, there is a
limited degree of positive correlation.
4. Limited degree of Negative correlation
When an increase in the value of one variable is followed by a non-proportional decrease in another related
variable, or when a decrease in the value of one variable is followed by a non-proportional increase in another
related variable, it is called a limited degree of negative correlation.
For example, if a 10% rise in price results in a 5% fall in its demand, there is a limited degree of negative
correlation. Similarly, if 5% fall in price results in 10% increase in demand, it is a limited degree of negative
correlation.
5. Zero Correlation (Zero Degree Correlation)
If there is no correlation between variables it is called zero correlation. In other words, if the values of one
variable cannot be associated with the values of the other variable, it is zero correlation.
Methods of measuring correlation
Correlation between 2 variables can be measured by graphic methods and algebraic methods.
I Graphic Methods
1. Scatter Diagram
2. Correlation graph
II Algebraic methods (Mathematical Methods or statistical methods or Coefficient of correlation
methods)
1. Karl Pearson's Coefficient of correlation
2. Spearman's Rank correlation method
3. Concurrent deviation method
Graphic Methods
Scatter Diagram
This is the simplest method for ascertaining the correlation between variables. Under this method all the values
of the two variables are plotted in a chart in the form of dots. Therefore, it is also known as a dot chart. By
observing the scatter of the various dots, we can form an idea of whether the variables are related or not.

Merits of Scatter Diagram method
1. It is a simple method of studying correlation between variables.
2. It is a non-mathematical method of studying correlation between the variables. It does not require any
mathematical calculations.
3.​ It is very easy to understand. It gives an idea about the correlation between variables even to a layman.
4. It is not influenced by the size of extreme items.
5. Making a scatter diagram is, usually, the first step in investigating the relationship between two variables.
Demerits of Scatter diagram method
[Link] gives only a rough idea about the correlation between variables.
2. The numerical measurement of correlation coefficient cannot be calculated under this method.
3. It is not possible to establish the exact degree of relationship between the variables.

Correlation graph Method


Under correlation graph method the individual values of the two variables are plotted on a graph paper. Then
dots relating to these variables are joined separately so as to get two curves. By examining the direction and
closeness of the two curves, we can infer whether the variables are related or not. If both the curves are
moving in the same direction (either upward or downward) correlation is said to be positive. If the curves are
moving in the opposite directions, correlation is said to be negative.
Merits of Correlation Graph Method
[Link] is a simple method of studying the relationship between the variables.
2. This does not require mathematical calculations.
3. This method is very easy to understand
Demerits of correlation graph method
1.A numerical value of correlation cannot be calculated.
2. It is only a pictorial presentation of the relationship between variables.
3. It is not possible to establish the exact degree of relationship between the variables.
Algebraic methods
Karl Pearson's Coefficient of Correlation
It is also called product moment correlation coefficient.
Interpretation of Coefficient of Correlation
Pearson's Coefficient of correlation always lies between +1 and -1. The following general rules will help to
interpret the Coefficient of correlation:
1. When​ r-+1,​ it​ means​ there​ is​ perfect positive relationship between variables.
2. When r = -1, it means there is a perfect negative relationship between variables.
3. When r = 0, it means there is no relationship between the variables.
4. When 'r' is closer to +1, it means there is a high degree of positive correlation between variables.
[Link] 'r' is closer to -1, it means there is a high degree of negative correlation between variables.
6. When 't' is closer to 'O', it means there is less relationship between variables.
Properties of Pearson's Coefficient of Correlation
1. If there is correlation between variables, the Co- efficient of correlation lies between +1 and -1.
[Link] there is no correlation, the coefficient of correlation is denoted by zero (ie r=0)
3. It measures the degree and direction of change
4. It simply measures the correlation and does not help to predict cansation.

5. It is the geometric mean of two regression coefficients.


Probable Error and Coefficient of Correlation
Probable error (PE) of the Coefficient of correlation is a statistical device which measures the reliability and
dependability of the value of coefficient of correlation.
If the value of coefficient of correlation (r) is less than the PE, then there is no evidence of correlation.
If the value of 'r' is more than 6 times of PE, the correlation is certain and significant.
By adding and submitting PE from the coefficient of correlation, we can find out the upper and lower limits
within which the population coefficient of correlation may be expected to lie.
Uses of PE
1. PE is used to determine the limits within which the population coefficient of correlation may be expected to
lie.
2. It can be used to test whether the value of correlation coefficient of a sample is significant with that of the
population.
Coefficient of Determination
One very convenient and useful way of interpreting the value of coefficient of correlation is the use of the
square of coefficient of correlation. The square of coefficient of correlation is called coefficient of
determination.
Coefficient of determination = r2
Merits of Pearson's Coefficient of Correlation
[Link] is the most widely used algebraic method to measure coefficients of correlation.
2. It gives a numerical value to express the relationship between variables.
[Link] gives both direction and degree of relationship between variables.
4. It can be used for further algebraic treatment such as coefficient of determination, coefficient of non-
determination etc.
5. It gives a single figure to explain the accurate degree of correlation between two variables
Demerits of Pearson's Coefficient of correlation
[Link] is very difficult to compute the value of the coefficient of correlation.
2. It is very difficult to understand
[Link] requires complicated mathematical calculations
4. It takes more time
5. It is unduly affected by extreme items
[Link] assumes a linear relationship between the variables. But in real life situations, it may not be so.
Spearman's Rank Correlation Method
The correlation coefficient obtained from ranks of the variables instead of their quantitative measurement is
called rank correlation.
Merits of Rank Correlation method
[Link] correlation coefficient is only an approximate measure as the actual values are not used for
calculations.
[Link] is very simple to understand the method.
[Link] can be applied to any type of data, i.e. quantitative and qualitative.
4. It is the only way of studying correlation between qualitative data such as honesty, beauty etc.
5. As the sum of rank differences of two qualitative data is always equal to zero, this method facilitates a cross
check on the calculation.
Demerits of Rank Correlation method
[Link] correlation coefficient is only an approximate measure as the actual values are not used for
calculations.
2. It is not convenient when number of pairs (ie. N) is large
3. Further algebraic treatment is not possible.
[Link] correlation coefficient of different series cannot be obtained as in the case of mean and
standard deviation. In case of mean and standard deviation, it is possible to compute combined arithmetic
mean and combined standard deviation.
Regression Analysis
Definition
´ It means the estimation or the prediction of the unknown value if one variable from the known value of the
other variable in terms of original units of data
´ It is a statistical device
´ Used to study the relationship between two or more variables that are related

Dependent variable
Value is to be predicted is called dependent variable
Independent variable
´ Value used for prediction is independent variable
Types of regression equation
Simple ,Multiple, Linear ,Nonlinear regressions
Simple regression
´ There are only two variables
Multiple regression
´ There are more than two variables and try to find out the effect of two or more independent variable on one
dependent variable
Nonlinear regression
´ The points so obtained on the scatter diagram will more or less concentrate around a curve
Linear regression
´ The curve is a straight line
´ Its form is y= a+ bx​
Line of best fit(Fitting of straight line)
If the points of the scatter diagram concentrate around a straight line,that line is called the line of best fit. That
is the line of best fit is closer to the points of the scatter [Link] line is called regression line
Curve Fitting(method of least squares)
It is a method of drawing a regression line by applying the principle of least squares.​
Principle of least squares(Least Square Method of Computing Regression Equation)
It states that the line of best fit should be drawn in such a manner that the sum of squares of difference between
the known values of the dependent variable and the corresponding values of it obtained from the line of best fit
should be the least
Uses of study of regression
´ To obtain most probable values of one series for given values of other related series
´ Used in physical sciences where the data are generally in functional relationship
´ To describe the relationship between the two variables and to show the rates of change in one factor in terms
of another
Difference between Correlation and regression
Correlation Regression
1. We study the degree of relationship 1. We study the nature of relationship
between the variables.
[Link] choice of dependent and independent [Link] has to decide which variable shall be
variable is purely personal choice and is of taken as dependent and which as independent
no practical significance
[Link] is not for the purpose of prediction 3. It basically used for prediction purpose
Regression Lines
Regression line is a graphic technique to show the functional relationship between the two variables X and Y.
It is a line which shows the average relationship between two variables X and Y.
Properties and Regression lines
1. The two regression lines cut each other at the point of average of X and average of Y (i.e X and Y)
[Link] r = 1, the two regression lines coincide with each other and give online.
3. When r = 0, the two regression lines are mutually perpendicular.

Properties of Regression Coefficient


1. There are two regression coefficients. They are byxy and byx
2. Both the regression coefficients must have the same signs. If one is +ve, the other will also be a +ve value.
3. The geometric mean of regression coefficients will be r
4. If x and y are the same,then the regression coefficient and correlation coefficient will be the same.
Difference between linear and nonlinear regression

Linear regression Non linear regression


[Link] curve is a straight line 1. Regression curve is not a straight line
[Link] change in the dependent variable is 2. The dependent variable does not
proportionate to the change in the change by a constant amount of
independent variable. change in the independent variable.
MODULE III
Chapter 1
THEORY OF PROBABILITY
Definition of Probability
The probability of a given event may be defined as the numerical value given to the likely hood of the
occurrence of that event. It is a number lying between '0' and '1' '0' denotes the event which cannot occur, and
'1' denotes the event which is certain to occur. For example, when we toss on a coin, we can enumerate all the
possible outcomes (head and tail), but we cannot say which one will happen. Hence, the probability of getting
a head is neither 0 or 1 but between 0 and 1. It is 50% or ½.
Terms used in Probability
A random experiment is an experiment that has two or more outcomes which vary in an unpredictable manner
from trial to trial when conducted under uniform conditions.
In a random experiment, all the possible outcomes are known in advance but none of the outcomes can be
predicted with certainty. For example, tossing a coin is a random experiment because it has two outcomes
(head and tail), but we cannot predict any of them with certainty.
Random experiments
An experiment that has two or more outcomes which vary in an unpredictable manner from trial to trial when
conducted under uniform conditions is called a random experiment
Eg: the tossing of a coin it has two specified outcomes Head and Tail
Sample point
Every indecomposable outcome of a random experiment is called a sample point. It is also called a simple
event or elementary outcome.
Eg:When a die is thrown, getting '3' is a sample point.
Sample Space
Sample space of a random experiment is the set containing all the sample points of that random experiment.
Eg: When a coin is tossed, the sample space is (Head, Tail)
Event
An event is the result of a random experiment. It is a subset of the sample space of a random experiment.

Sure Event (Certain Event)


An event whose occurrence is inevitable is called a sure event.
Eg:Getting w white ball from a box containing all white balls.
Impossible Events
An event whose occurrence is impossible, is called an impossible event. Eg:- Getting a white ball from a box
containing all red balls.
Uncertain Events
An event whose occurrence is neither sure nor impossible is called an uncertain event.
Eg: Getting a white ball from a box containing white balls and black balls.
Equally likely Events
Two events are said to be equally likely if any of them cannot be expected to occur in preference to others.
For example, getting a herd and getting tail when a coin is tossed are equally likely events.
Mutually exclusive events
A set of events are said to be mutually exclusive if the occurrence of one of them excludes the possibility of
the occurrence of the others.
Eg: Getting an Ace and a King when a card is drawn from a pack of cards are mutually Exclusive
Exhaustive Events:
A group of events is said to be exhaustive when it includes all possible outcomes of the random experiment
under consideration.
Eg:When a die is thrown ,outcomes 1,2,3,4,5 and 6 togethetr will form exhaustive events because one of tehm
will occur when a die is thrown
Dependent Events:
Two or more events are said to be dependent if the happening of one of them affects the happening of the
other.
Eg:From a pack of 52 cards if one card is drwn the 51 cards are [Link] another card is drawn without replacing
the first chance of the second draw is affected by the first draw.

Independent Events
Two or more events are said to be independent if the occurrence of one of them in no way affects the
occurrence of the other or the others
Eg:In the tossing of a coin twice ,the result of the second tossing is not affected by the result of the first toss.
Operations on Events
Intersection of Two Events

The intersection of two events A and B denoted by A∩B is the set of sample
points common to both A and B

Eg: A = getting a multiple of 5 and B= getting a multiple of 3

A∩B = 15

Union of Two Events


The union of two events A and B denoted by AUB is the set of sample points in A or in B or in both.
Eg: A = getting a multiple of 5 and B= getting a multiple of 3
AUB = getting a multiple of 5 or 3
Complement of event A
The event A and the event not A are called complementary events.
𝐶
𝐴 = U-A
𝐶
Eg: U = {1,2,.......,10},A = {2,4,6} 𝐴 ={ 1,3,5,7,8,9,10}
MODULE IV
ADVANCED PROBABILITY DISTRIBUTION
Binomial Distribution
Situations where It can be Applied
●​ The random experiment has two outcomes,which can be called’ success ‘and ‘failure’
●​ Probability for success in a single trial remains constant from trial to trial of the experiment
●​ The experiment is repeated finite number of times
●​ Trials are independent
Characteristics or Properties of Binomial Distribution
●​ It is a discrete probability distribution
●​ The shape and location changes as p changes for a given n
●​ It has one or two model values
●​ Mean increases as n increases with p remaining constant
●​ If n is large and if neither p nor q is too close to zero,it may approximated to normal distribution
●​ Mean = np , S.D = 𝑛𝑝𝑞
●​ If two independent random variable follow Binomial distribution ,their sum also follows Binomial
distribution
Poisson Distribution
Uses/situations or Importance of Poisson Distribution
●​ To count number of telephone calls arriving at the telephone switchboard in unit time
●​ To count number of customers arriving at the supermarket
●​ To count number of defects per unit of a manufactured product
●​ To count number of radioactive disintegrations of a radioactive element per unit time
●​ to count number of bacterias per unit
●​ To count number of defective materials
●​ To count the number of casualties due to a rarte disease such as heart attack in a year
●​ To count the number of accidents taking place in a day on a busy road
Characteristics or properties
●​ It is discrete probability distribution
●​ If x follows a Poisson distribution ,the x takes values 0,1,2,..... to infinity
●​ It has a single parameter [Link] m is known all the terms can be found out.
●​ Mean = Variance = m
●​ It is positively skewed distribution
Normal Distribution
Uses/situations or Importance of Normal Distribution
●​ It is a continuous curve
●​ The ordinate at mean divides the whole area into two equal parts
●​ Coefficient of skewness = 0
●​ It has only one mode
●​ The point of inflexion occur at ㅆ+σ
●​ Q1 and Q3 are equidistant from median
●​ M D = (⅘)σ Q D = (⅔)σ
●​ It is bell shaped
●​ It is symmetric about the mean
●​ Mean = Median = Mode
●​ The height of the normal curve is at its maximum at the mean
Characteristics or properties
●​ Most of the discrete probability distribution tend to normal distribution as n becomes large
●​ Almost all sampling distribution -t,F,Z-Distributions etc conform to the normal distribution for large
values of n
●​ The various tests of significance like t-test F-test etc based on the assumption that the parent population
from which the samples have been drawn follows normal distribution
●​ It is extensively used in large sample theory to find the estimates of parameters from statistics
,confidence limits ,etc
●​ In theoretical statistics as well as applied works many problems can be solved only under the
assumption of a normal population
●​ The normal curve is reasonably close to many distributions
●​ It finds applications in statistical quality control and industrial experiments.

Common questions

Powered by AI

Pearson's Coefficient of Correlation may not be ideal for financial time series due to its sensitivity to outliers and inability to distinguish between causation and correlation. The requirement for linear relationships can be a limitation as financial markets often exhibit non-linear patterns .

The Harmonic Mean (HM) meets several characteristics of a good average as it is based on all observations, amenable to algebraic treatment, and is not significantly affected by sampling fluctuation; however, it is difficult to calculate and understand, and it gives greater weight to smaller values, which may affect its practical applicability .

The Median is more appropriate in scenarios involving skewed data or when dealing with outliers, since it is not affected by extreme values. It provides a central point that divides the data into two equal parts, making it useful for ordinal datasets or datasets with open-ended classes .

The Arithmetic Mean (AM) is affected by extreme values, making it less suitable for datasets with significant outliers. On the other hand, the Geometric Mean (GM) is not significantly affected by extreme values, which makes it more appropriate for datasets where such outliers might skew the results .

Using both graphic (e.g., scatter diagrams) and algebraic (e.g., Pearson's Coefficient) methods provides a comprehensive analysis. Graphic methods offer a visual representation and qualitative insight, while algebraic methods provide quantitative measurement, ensuring a thorough understanding of correlation between variables .

The Coefficient of Determination can be misleading if used inappropriately, such as interpreting high values as indicators of causality instead of correlation strength or in models with non-linear relationships where it might overestimate the explanatory power .

The Mode is considered the most representative measure in datasets where the most frequent item is of interest, such as in categorical data or in cases where the most common occurrence is more meaningful than the average, for instance in fashion or consumer preference data .

The Range, being the simplest measure of dispersion, is relevant in quality control processes because it provides a quick assessment of the variability within a product batch. It is used to determine maximum variation, which can be crucial for identifying defects and ensuring consistent product quality .

In financial analysis, these means are used to analyze different aspects: the Arithmetic Mean is used for overall average calculations; the Geometric Mean effectively calculates average growth rates over time; the Harmonic Mean is useful for average ratios like price-to-earnings. The relationship between them, expressed as AM x HM = GM², reflects their dependence on each other's values .

Probable Error provides an indication of the reliability of the correlation coefficient value. A correlation value less than the Probable Error suggests no significant correlation, whereas a value more than six times the PE indicates the correlation is significant, helping analysts assess the robustness of the relationship .

You might also like