SRI BALAJI UNIVERSITY
SY MCA Semester - III
RESEARCH METHODOLOGY AND ETHICS
Dr. Savita V. Mohurle
Data Collection
• The primary data are those which are collected afresh and for the first
time, and thus happen to be original in character.
• The secondary data, on the other hand, are those which have already been
collected by someone else and which have already been passed through
the statistical process.
Data Collection
• Primary data can be collected either through experiment or through survey.
• If the researcher conducts an experiment, he observes some quantitative
measurements, or the data, with the help of which he examines the truth
contained in his hypothesis.
• But in the case of a survey, data can be collected by any one or more of the
following ways:
➢By observation – Refer Reference Book
➢Through personal interview
➢Through telephone interviews
➢By mailing of questionnaires
➢Through schedules
The researcher should select one of these methods of collecting the
data taking into consideration the nature of investigation, objective
and scope of the inquiry, financial resources, available time and the
desired degree of accuracy.
In this context Dr A.L. Bowley very aptly remarks that in collection of
statistical data commonsense is the chief requisite and experience
the chief teacher.
Data Collection
• The secondary data, on the other hand, are those which have already been
collected by someone else and which have already been passed through the
statistical process.
• The researcher would have to decide which sort of data he would be using
(thus collecting) for his study and accordingly he will have to select one or
the other method of data collection.
Other methods
Warranty cards: Warranty cards are usually postal sized cards which are used
by dealers of consumer durables to collect information regarding their
product.
Distributor or store audits: Distributor or store audits are performed by
distributors as well as manufactures through their salesmen at regular
intervals the retail stores audit is performed through salesmen and use such
information to estimate market size, market share, seasonal purchasing
pattern and so o
Pantry audits: Pantry audit technique is used to estimate consumption of the
basket of goods at the consumer level. The investigator collects an inventory
of types, quantities and prices of commodities consumed. The usual objective
in a pantry audit is to find out what types of consumers buy certain products
and certain brands.
COLLECTION OF SECONDARY DATA
Secondary data may either be published data or unpublished data. Usually published data
are available in:
• various publications of the central, state are local governments;
• various publications of foreign governments or of international bodies and their
subsidiary organisations;
• technical and trade journals;
• books, magazines and newspapers;
• reports and publications of various associations connected with business and industry,
banks, stock exchanges, etc.;
• reports prepared by research scholars, universities, economists, etc. in different fields;
and
• public records and statistics, historical documents, and other sources of published
information.
• The sources of unpublished data are many; they may be found in diaries, letters,
unpublished biographies and autobiographies and also may be available with scholars and
research workers, trade associations, labour bureaus and other public/private individuals
and organisations.
Other methods
Consumer panels: An extension of the pantry audit approach on a regular
basis is known as ‘consumer panel’, where a set of consumers are arranged
to come to an understanding to maintain detailed daily records of their
consumption and the same is made available to investigator on demands.
Content-analysis: Content-analysis consists of analysing the contents of
documentary materials such as books, magazines, newspapers and the
contents of all other verbal materials which can be either spoken or printed.
Depth interviews: Depth interviews are those interviews that are designed to
discover underlying motives and desires and are often used in motivational
research. Such interviews are held to explore needs, desires and feelings of
respondents.
Characteristics of Secondary data
By way of caution, the researcher, before using secondary data, must see that they
possess following characteristics:
1. Reliability of data: The reliability can be tested by finding out such things about
the said data:
• Who collected the data?
• What were the sources of data?
• Were they collected by using proper methods
• At what time were they collected?(e) Was there any bias of the compiler?
• What level of accuracy was desired? Was it achieved ?
2. Suitability of data: The data that are suitable for one enquiry may not necessarily
be found suitable in another enquiry. Hence, if the available data are found to be
unsuitable, they should not be used by the researcher.
3. Adequacy of data: If the level of accuracy achieved in data is found inadequate for
the purpose of the present enquiry, they will be considered as inadequate and
should not be used by the researcher. The data will also be considered inadequate,
if they are related to an area which may be either narrower or wider than the area
of the present enquiry.
SELECTION OF APPROPRIATE METHOD FOR DATA COLLECTION
[Link], scope and object of enquiry: This constitutes the most important factor
affecting the choice of a particular method. The method selected should be such
that it suits the type of enquiry that is to be conducted by the researcher. This factor
is also important in deciding whether the data already available (secondary data)
are to be used or the data not yet available (primary data) are to be collected.
[Link] of funds: Availability of funds for the research project determines to a
large extent the method to be used for the collection of data. Finance, in fact, is a
big constraint in practice and the researcher has to act within this limitation.
[Link] factor: Availability of time has also to be taken into account in deciding a
particular method of data collection. The time at the disposal of the researcher,
thus, affects the selection of the method by which the data are to be collected.
[Link] required: Precision required is yet another important factor to be
considered at the time of selecting the method of collection of data.
Assignment
Explain Case Study Method in detail.
Sample Design
• All items in any field of inquiry constitute a ‘Universe’ or ‘Population.’
• A complete enumeration of all items in the ‘population’ is known as a census
inquiry.
• It can be presumed that in such an inquiry, when all items are covered, no
element of chance is left and highest accuracy is obtained
• A sample design is a definite plan for obtaining a sample from a given population.
• It refers to the technique or the procedure the researcher would adopt in
selecting items for the sample.
• Sample design may as well lay down the number of items to be included in the
sample i.e., the size of the sample.
• Sample design is determined before data are collected.
Sample Design
• The selected respondents constitute what is technically called a ‘sample’
and the selection process is called ‘sampling technique.’
• The survey so conducted is known as ‘sample survey’.
• Algebraically, let the population size be N and if a part of size n (which is <
N) of this population is selected according to some rule for studying some
characteristic of the population, the group consisting of these n units is
known as ‘sample’
Sample Design
In research methodology, sample design refers to the specific plan or
procedure used to select a subset of individuals or items (the sample)
from a larger group (the population) for study.
The goal is to choose a sample that accurately represents the
population, allowing researchers to draw reliable conclusions and make
generalizations about the entire population based on the sample data.
[Link]
STEPS IN SAMPLE DESIGN
• Type of universe: The first step in developing any sample design is to clearly define the set
of objects, technically called the Universe, to be studied. The universe can be finite or
infinite.
• Sampling unit: A decision has to be taken concerning a sampling unit before selecting
sample. Sampling unit may be a geographical one such as state, district, village, etc., or a
construction unit such as house, flat, etc., or it may be a social unit
• Source list: It is also known as ‘sampling frame’ from which sample is to be drawn. It
contains the names of all items of a universe (in case of finite universe only).such as family,
club, school, etc., or it may be an individual.
• Size of sample: This refers to the number of items to be selected from the universe to
constitute a sample.
• Parameters of interest: In determining the sample design, one must consider the question
of the specific population parameters which are of interest.
• Budgetary constraint: Cost considerations, from practical point of view, have a major
impact upon decisions relating to not only the size of the sample but also to the type of
Sample.
• Sampling procedure: Finally, the researcher must decide the type of sample he will use i.e.,
he must decide about the technique to be used in selecting the items for the sample. In
fact, this technique or procedure stands for the sample design itself
CRITERIA OF SELECTING A SAMPLING PROCEDURE
• Researcher must keep in view the two causes of incorrect inferences
viz., systematic bias and sampling error.
• A systematic bias results from errors in the sampling procedures, and it
cannot be reduced or eliminated by increasing the sample size.
• At best the causes responsible for these errors can be detected and
corrected.
CRITERIA OF SELECTING A SAMPLING PROCEDURE
• Inappropriate sampling frame: If the sampling frame is inappropriate i.e., a
biased representation of the universe, it will result in a systematic bias.
• Defective measuring device: If the measuring device is constantly in error, it
will res
• Non-respondents: If we are unable to sample all the individuals initially
included in the sample, there may arise a systematic bias.
• Indeterminacy principle: Sometimes we find that individuals act differently
when kept under observation than what they do when kept in non-
observed situations.
• Natural bias in the reporting of data: Natural bias of respondents in the
reporting of data is often the cause of a systematic bias in many inquiries
Sample Design
Confidence Interval
Definition: A confidence interval provides a range of values within which a
population parameter is likely to fall.
Example: If a survey reports that 60% of respondents prefer a certain
product, with a margin of error of ±3%, the confidence interval would be
57% to 63%. This means you are confident that the true percentage of
people who prefer the product in the entire population falls within that
range.
Key Components: A confidence interval is defined by a lower and upper
limit.
Confidence Level: The confidence level, usually expressed as a percentage
(e.g., 95%), indicates the probability that the true population parameter
falls within the calculated interval.
Sample Design
Margin of Error
Definition: The margin of error is a statistic that expresses the amount of
random sampling error in survey results.
Calculation:
• It is often calculated as half the width of the confidence interval.
• Relationship with Confidence Interval:
• A wider confidence interval indicates a larger margin of error, while a
narrower confidence interval indicates a smaller margin of error.
Interpretation: A larger margin of error suggests that the results are less
precise, while a smaller margin of error suggests greater precision.
Sample Design
CHARACTERISTICS OF A GOOD SAMPLE DESIGN
(a) Sample design must result in a truly representative sample.
(b) Sample design must be such which results in a small sampling error.
(c) Sample design must be viable in the context of funds available for the
research study.
(d) Sample design must be such so that systematic bias can be controlled in a
better way.
(e) Sample should be such that the results of the sample study can be
applied, in general, for the universe with a reasonable level of confidence
[Link]
Sampling Techniques
Non Probability sampling
• In this type of sampling, items for the sample are selected deliberately
by the researcher; his choice concerning the items remains supreme.
• In other words, under non-probability sampling the organisers of the
inquiry purposively choose the particular units of the universe for
constituting a sample on the basis that the small mass that they so select
out of a huge one will be typical or representative of the whole.
Sampling Techniques
Quota sampling
• Quota sampling is also an example of non-probability sampling.
• Under quota sampling the interviewers are simply given quotas to be
filled from the different strata, with some restrictions on how they are to
be filled. In other words, the actual selection of the items for the sample
is left to the interviewer’s discretion.
• This type of sampling is very convenient and is relatively inexpensive. But
the samples so selected certainly do not possess the characteristic of
random samples.
• Quota samples are essentially judgement samples and inferences drawn
on their basis are not amenable to statistical treatment in a formal way.
Sampling Techniques
Judgment sampling
• Judgmental sampling, also called purposive sampling or authoritative
sampling, is a non-probability sampling technique in which the sample
members are chosen only on the basis of the researcher’s knowledge and
judgment.
• As the researcher’s knowledge is instrumental in creating a sample in this
sampling technique, there are chances that the results obtained will be
highly accurate with a minimum margin of errors.
• Judgmental sampling is most effective in situations where there are only a
restricted number of people in a population who own qualities that a
researcher expects from the target population.
Sampling Techniques
Judgment sampling - Advantages
• Consumes minimum time for execution: In this sampling bias approach,
researcher expertise is important and there are no other barriers
involved due to which selecting a sample becomes extremely
convenient.
• Allows researchers to approach their target market directly: There are
no criteria involved in selecting a sample except for the researcher’s
preferences. Due to this, he/she can communicate directly with
the target audience of their choice and produce desired results.
• Almost real-time results: A quick poll or survey can be conducted with
the sample using judgmental sampling since the members of the sample
will possess appropriate knowledge and understanding of the subject.
Snowball sampling
• Snowball sampling is a non-probability sampling technique where existing
study participants recruit future subjects from among their
acquaintances.
• This method is particularly useful when the population is hard-to-reach or
when traditional sampling methods are not feasible. It's also known as
chain-referral or network sampling.
• When is it used?
➢When studying sensitive or stigmatized topics.
➢When the population is hidden or hard to locate.
➢When traditional sampling methods are not feasible.
➢When studying social networks and relationships.
Sampling Techniques
Probability sampling:
• Probability sampling is also known as ‘random sampling’ or ‘chance
sampling’.
• Under this sampling design, every item of the universe has an equal
chance of inclusion in the sample.
• It is, so to say, a lottery method in which individual units are picked up
from the whole group not deliberately but by some mechanical process.
• Here it is blind chance alone that determines whether one item or the
other is selected.
Sampling Techniques
In brief, the implications of random sampling (or simple random sampling) are:
(a) It gives each element in the population an equal probability of getting into
the sample; and all choices are independent of one another.
(b) It gives each possible sample combination an equal probability of being
chosen.
Sampling Techniques
Systematic sampling
• The most practical way of sampling is to select every ith item on a list.
• Sampling of this type is known as systematic sampling.
• An element of randomness is introduced into this kind of sampling by using random
numbers to pick up the unit with which to start.
• For instance, if a 4 per cent sample is desired, the first item would be selected
randomly from the first twenty-five and thereafter every 25th item would
automatically be included in the sample.
• Thus, in systematic sampling only the first unit is selected randomly and the
remaining units of the sample are selected at fixed intervals.
• Although a systematic sample is not a random sample in the strict sense of the
term, but it is often considered reasonable to treat systematic sample as if it were a
random sample.
Sampling Techniques
Stratified sampling:
• If a population from which a sample is to be drawn does not constitute a
homogeneous group, stratified sampling technique is generally applied in order to
obtain a representative sample.
• Under stratified sampling the population is divided into several sub-populations
that are individually more homogeneous than the total population (the different
sub-populations are called ‘strata’) and then we select items from each stratum to
constitute a sample.
• Since each stratum is more homogeneous than the total population, we are able to
get more precise estimates for each stratum and by estimating more accurately
each of the component parts, we get a better estimate of the whole.
Cluster sampling and area sampling
• Cluster sampling and area sampling: Cluster sampling involves grouping the
population
• and then selecting the groups or the clusters rather than individual elements
for inclusion in the sample.
• Suppose some departmental store wishes to sample its credit card holders. It
has issued its cards to 15,000 customers. The sample size is to be kept say
450. For cluster sampling this list of 15,000 card holders could be formed into
100 clusters of 150 card holders each.
• Three clusters might then be selected for the sample randomly.
• The sample size must often be larger than the simple random sample to
ensure the same level of accuracy because is cluster sampling procedural
potential for order bias and other sources of error is usually accentuated.
• The clustering approach can, however, make the sampling procedure
relatively easier and increase the efficiency of field work, specially in the case
of personal interviews.
Data Analysis
• The term analysis refers to the computation of certain measures along with
searching for patterns of relationship that exist among data-groups.
• Thus, “in the process of analysis, relationships or differences supporting or
conflicting with original or new hypotheses should be subjected to
statistical tests of significance to determine with what validity data can be
said to indicate any conclusions”.
• Data analysis in the research process is a systematic approach to
inspecting, cleansing, transforming, and modeling data with the goal of
discovering useful information, drawing conclusions, and supporting
decision-making.
• It involves various techniques to interpret data from different sources and
formats, both structured and unstructured, to uncover insights, identify
patterns, and understand trends.
Data Analysis
Key aspects of data analysis in research:
• Informed Decision-Making: Data analysis provides valuable insights that
support informed decision-making by enabling researchers to make data-
driven choices.
• Identifying Patterns and Relationships: Data analysis helps to uncover
hidden patterns, relationships, and trends within the data, leading to a
deeper understanding of the research topic.
• Supporting Hypotheses: Data analysis can be used to test hypotheses and
determine whether they are supported by the evidence.
• Drawing Conclusions: Researchers use data analysis to draw meaningful
conclusions about their research questions or hypotheses.
• Validating Results: Data analysis can be used to validate the results of
previous studies or to provide new insights into existing research.
Types
The data analysis can be categorized into the following six main
methods (Taherdoost, 2021):
● Descriptive
● Exploratory
● Inferential
● Predictive
● Explanatory or Causal
● Mechanistic
Nominal Symbols or names (or numbers) of Occupation (teacher, dentist, programmer, farmer, and
things so on)
Symmetric - if both of its states are equally valuable
and carry the same weight;
A nominal attribute with only two Example: 0 or 1, true or false, male or female, yes or
Binary categories or states no
Qualitative
Asymmetric if the outcomes of the states are not
equally important, such as the positive and negative
outcomes of a medical test for HIV.
An attribute with possible values
that have a meaningful order or
Ordinal ranking among them, but the Example: Small, medium, large
magnitude between successive
values is not known.
Interval scaled - Are measured on a scale of equal-size
units
It is a measurable quantity, Example: A temperature attribute is interval-scaled.
Numeric represented in integer or real Ratio scaled - Is a numeric attribute with an inherent Quantitative
values. zero-point.
Example: height, money, age, weight, speed, and time
periods.
Statistical Description of Data
• Science of collecting, organizing, presenting, interpreting and analyzing.
• Data scientist is a person who is better at statistics than any programmer and
better at programming than any statistician.
Descriptive statistics summarizes or describes the characteristics of a data set.
➢Descriptive statistics consists of three basic categories of measures: measures
of central tendency, measures of variability (or spread), and frequency
distribution.
❑Measures of central tendency describe the center of the data set (mean,
median, mode).
❑Measures of variability describe the dispersion of the data set (variance,
standard deviation).
❑Measures of frequency distribution describe the occurrence of data within
the data set (count, frequency table, histogram, bar chart).
Statistical Description of Data
Descriptive Analysis
Elements Std Dev. Mean Median Mode Q1 (25%) Q2 (50%) Q3 (75%) Minimum Maximum
Moisture 7.936389 19.12252 19.0652 23.1 14.18 19.065 22.56 0 56.5
N 2.828273 1.153824 0.78 0.78 0.5 0.78 1.12 0 68
P205 1.147476 1.625204 1.43 0.95 0.8975 1.43 1.9 0.099 9.25
K20 2.440252 2.624122 1.59 0.98 0.97 1.59 3.88 0 19.8
Cd 2.120466 4.282424 4.12 5.1 2.99 4.12 5.34 0 16.2
C 8.628163 20.37754 17.4 16.34 15.7 17.4 22.1 0 67.89
C:N 8.658163 20.32102 19.66 20.1 16.873 19.66 20.873 2.3 56.2
Zn 144.6234 993.5245 1000 999 965 1000 1020.5 189.6 1500
Pb 14.16045 99.29923 99.89 100.1 98.665 99.89 101 0 267
Cu 64.09558 278.225 299.1 0 281.26 299.1 302 0 401.1
pH 1.718799 7.217256 7.32 7.23 6.5 7.32 8.1 0.5 24
Density 2.226921 1.76008 0.7675 0.45 0.42 0.7675 2.11 0.001 15.6
Conductivity 2.7050982 5.5016157 5.45 6.45 3.435 5.45 7.4574 0 19.2
As 2.026374 7.672502 7.45 6.55 7.45 8.5 0.5 0 24
Ni 16.34339 22.74485 18.685 4.5 9.2415 18.685 33.223 1.2 91.34
Hg 0.37413 0.184395 0.023 0.01 0.01 0.023 0.128 0 3.2
Cr 15.62833 20.15139 14.7 4.5 7.89 14.7 31.32 0 66.4
Statistical Description of Data
• Inferential statistics helps study a sample of data and make conclusions
about its population.
• A sample is a smaller data set drawn from a larger data set called the
population.
• If the sample does not represent the population, one cannot make accurate
estimations related to the latter.
• The purpose of studying inferential statistics is to infer the behavior of a
population.
• The accuracy of inferential statistics depends largely on the accuracy of
sample data and how it represents the larger population.
• This can be effectively done by obtaining a random sample. Results that are
based on non-random samples are usually discarded.
Descriptive Statistics Inferential Statistics
Make inferences and draw conclusions about a population
Purpose Describe and summarize data
based on sample data
Analyzes and interprets the characteristics of a Uses sample data to make generalizations or predictions
Data Analysis
dataset about a larger population
Focuses on a subset of the population (sample) to draw
Population Vs Sample Focuses on the entire population or dataset
conclusions about the entire population
Provides measures of central tendency and Estimates parameters, tests hypotheses, and determines the
Measurements
dispersion level of confidence or significance in the results
Mean, median, mode, standard deviation, Hypothesis testing, confidence intervals, regression analysis,
Examples
range, frequency tables ANOVA (analysis of variance), chi-square tests, t-tests, etc.
Generalize findings to a larger population, make predictions,
Goal Summarize, organize, and present data test hypotheses, evaluate relationships, and support
decision-making
Estimated using sample statistics (e.g., sample mean as an
Population Parameters Not typically estimated
estimate of population mean)
Crucial; the sample should be representative of the
Sample Representativeness Not required
population to ensure accurate inferences
Processing operation
• Editing: Editing of data is a process of examining the collected raw data
(specially in surveys) to detect errors and omissions and to correct these
when possible.
➢Field editing
➢Central editing
• Coding: Coding refers to the process of assigning numerals or other
symbols to answers so that responses can be put into a limited number of
categories or classes. Such classes should be appropriate to the research
problem under consideration.
Processing operation
• Classification: Most research studies result in a large volume of raw data
which must be reduced into homogeneous groups if we are to get
meaningful relationships.
➢Classification according to attributes
➢Classification according to class-intervals
• Tabulation: When a mass of data has been assembled, it becomes
necessary for the researcher to arrange the same in some kind of concise
and logical order. This procedure is referred to as tabulation
❑It conserves space and reduces explanatory and descriptive statement to a
minimum.
❑It facilitates the process of comparison.
❑It facilitates the summation of items and the detection of errors and omissions.
❑It provides a basis for various statistical computations.
STATISTICS IN RESEARCH
• Descriptive statistics concern the development of certain indices from the
raw data
• Inferential statistics concern with the process of generalisation.
• Inferential statistics are also known as sampling statistics and are mainly
concerned with two major type of problems:
(i) the estimation of population parameters,
(ii) the testing of statistical hypotheses.
Important statistical measures
(1) measures of central tendency or statistical averages
(2) measures of dispersion
(a) range
(b) mean deviation
(c) standard deviation
(3) measures of asymmetry (skewness)
(4) measures of relationship
(5) other measures.
The three most important ones are the arithmetic
➢Average or mean
➢Median
➢Mode
➢Geometric mean
➢Harmonic mean
Important statistical measures
• The measures of dispersion, variance, and its square root—the standard
deviation are the most often used measures.
• In respect of the measures of skewness and kurtosis, we mostly use the
first measure of
• Other measures:
➢Skewness, based on quartiles or on the methods of moments, are also
used sometimes. skewness based on mean and mode or on mean and
median.
➢ Kurtosis is also used to measure the peakedness of the curve of the
frequency distribution.
Hypothesis testing
• Hypothesis is usually considered as the principal instrument in research.
• Its main function is to suggest new experiments and observations.
“Students who receive counselling will show a greater increase in creativity
than students not receiving counselling”
Or
“The automobile A is performing as well as automobile B.”
These are hypotheses capable of being objectively verified and tested
Hypothesis testing
• In hypothesis testing, the null hypothesis (H0) assumes no effect or
relationship, while the alternative hypothesis (Ha) proposes that
there is a significant effect or relationship.
• If testing a new drug's effectiveness, the null hypothesis might be that
the drug has no effect, and the alternative would be that it does have
an effect on the patient.
Hypothesis testing
Null Hypothesis (H0)
• It's a statement of no effect, no difference, or no relationship.
• It's the default position that we try to disprove.
• Examples:
• "There is no difference in test scores between students using method A and
method B".
• "The new drug has no effect on blood pressure".
• "There is no relationship between the amount of text highlighted and exam
scores".
• "The average height of students is the same as the national average".
• "The average salary of data scientists is $113,000".
Hypothesis testing
Alternative Hypothesis (Ha)
• It contradicts the null hypothesis and suggests a specific effect, difference,
or relationship.
• It's what the researcher is trying to find evidence for.
• Examples:
• "Students using method A will have significantly different test scores than those
using method B".
• "The new drug will significantly lower blood pressure".
• "There is a relationship between the amount of text highlighted and exam scores".
• "The average height of students is different from the national average".
• "The average salary of data scientists is not $113,000".
Hypothesis testing
Key Points
• The null and alternative hypotheses are mutually exclusive; they
cannot both be true at the same time.
• The goal of hypothesis testing is to determine whether there is
enough evidence to reject the null hypothesis in favor of the
alternative hypothesis.
• There are different types of alternative hypotheses (one-tailed and
two-tailed), depending on the direction of the expected effect.
• In essence, the null hypothesis is a baseline assumption, while the
alternative hypothesis proposes a change or effect that the
researcher hopes to demonstrate with their data.
Hypothesis testing
• Mr. Mohan of the Civil Engineering Department wants to test the load
bearing capacity of an old bridge which must be more than 10 tons, in that
case he can state his hypotheses as under:
➢Null hypothesis H0 : m = 10 tons
➢Alternative Hypothesis Ha: m > 10 tons
• The average score in an aptitude test administered at the national level is
80. To evaluate a state’s education system, the average score of 100 of the
state’s students selected on random basis was 75. The state wants to know
if there is a significant difference between the local scores and the national
scores. In such a situation the hypotheses may be stated as under:
➢Null hypothesis H0: m = 80
➢Alternative Hypothesis Ha: m(not equal to) 80
Hypothesis testing
• Hypothesis testing is used to assess the plausibility of a hypothesis by
using sample data.
• The test provides evidence concerning the plausibility of the hypothesis,
given the data.
• Statistical analysts test a hypothesis by measuring and examining a
random sample of the population being analyzed.
• The four steps of hypothesis testing include stating the hypotheses,
formulating an analysis plan, analyzing the sample data, and analyzing
the result.
Hypothesis testing
• If, for example, a person wants to test that a penny has exactly a 50%
chance of landing on heads, the null hypothesis would be that 50% is
correct, and the alternative hypothesis would be that 50% is not correct.
• Mathematically, the null hypothesis would be represented as Ho: P = 0.5.
The alternative hypothesis would be denoted as "Ha" and be identical to
the null hypothesis, except with the equal sign struck-through, meaning
that it does not equal 50%.
Hypothesis testing
• The level of significance is the measurement of the statistical significance.
It defines whether the null hypothesis is assumed to be accepted or
rejected.
• The standard deviation of a random variable, sample, statistical population,
data set, or probability distribution is the square root of its variance.
• A z-test is used in hypothesis testing to evaluate whether a finding or
association is statistically significant or not. In particular, it tests whether
two means are the same (the null hypothesis). A z-test can only be used if
the population standard deviation is known and the sample size is 30 data
points or larger.
• One-tailed tests allow for the possibility of an effect in one direction. Two-
tailed tests test for the possibility of an effect in two directions—positive
and negative.
Hypothesis testing
Left-tailed test:
• The alternative hypothesis includes a "less than" (<) symbol.
• You're testing if the parameter is significantly smaller than a
hypothesized value.
• The rejection region is in the left tail of the distribution.
Right-tailed test:
• The alternative hypothesis includes a "greater than" (>) symbol.
• You're testing if the parameter is significantly larger than a
hypothesized value.
• The rejection region is in the right tail of the distribution.
Hypothesis testing
Example: Let's say you're testing if a new
medication lowers blood pressure.
•Null Hypothesis (H0): The medication has no
effect on blood pressure (mean blood pressure
is the same).
•Alternative Hypothesis (Ha): The medication
lowers blood pressure (mean blood pressure is
less than before).
•Because the alternative hypothesis uses "less
than", you would use a left-tailed test.
Z-test formula
The Z-test formula is used to determine whether two population
means are different when the population standard deviation is
known.
The formula is: z = (x̄ - μ) / (σ / √n)
Here's a breakdown of the formula:
•z: The z-statistic, which is the test statistic for the Z-test.
•x̄ (x-bar): The sample mean.
•μ (mu): The population mean.
•σ (sigma): The population standard deviation.
•n: The sample size.
[Link]
Z-test formula
• In statistics, a critical value is a threshold used in hypothesis testing to
determine whether to reject the null hypothesis.
• It's a value that defines the boundary of a region (the "critical region"
or "rejection region") where, if the calculated test statistic falls within,
the null hypothesis is rejected in favor of the alternative hypothesis.
When to use what test
• If your alternative hypothesis uses < (less than), use a left-tailed test.
• If your alternative hypothesis uses > (greater than), use a right-tailed test.
• If your alternative hypothesis uses ≠ (not equal to), use a two-tailed test.
❖One-tailed tests
(either left or right) are used when the alternative hypothesis specifies a direction.
➢Left-tailed test: Used when you hypothesize that the true value is less than the value
stated in the null hypothesis (e.g., a new drug is less effective than the old one).
➢Right-tailed test: Used when you hypothesize that the true value is greater than the
value stated in the null hypothesis (e.g., a new drug is more effective than the old one).
❖Two-tailed test:
Used when the alternative hypothesis states that the true value is different from the value
stated in the null hypothesis, without specifying a direction (e.g., a new drug's
effectiveness is different from the old one, without specifying whether it's better or
worse).
One sample Z test
• Suppose a company claims that their new smartphone has an average
battery life of 12 hours. A consumer group tests 100 phones and finds
an average battery life of 11.8 hours with a known population
standard deviation of 0.5 hours.
Two sample Z test
• Example: There are two groups of students preparing for a competition: Group
A and Group B. Group A has studied offline classes, while Group B has studied
online classes. After the examination the score of each student comes. Now
we want to determine whether the online or offline classes are better.
➢Group A: Sample size = 50, Sample mean = 75, Sample standard deviation = 10
➢Group B: Sample size = 60, Sample mean = 80, Sample standard deviation = 12
➢Assuming a 5% significance level perform a two-sample z-test to determine if
there is a significant difference between the online and offline classes.
How to Calculate Z Test Statistic?
• The most important step in calculating the z test statistic is to interpret
the problem correctly. It is necessary to determine which tailed test
needs to be conducted and what type of test does the z statistic belong
to. Suppose a teacher claims that his section's students will score
higher than his colleague's section. The mean score is 22.1 for 60
students belonging to his section with a standard deviation of 4.8. For
his colleague's section, the mean score is 18.8 for 40 students and the
standard deviation is 8.1. Test his claim at α = 0.05. The steps to
calculate the z test statistic are as follows:
The Z-test is used to compare a sample mean to a population mean or to
compare two sample means. It assumes that the data is normally
distributed and that the population standard deviation is known.
How to Calculate Z Test Statistic?
Example
• A teacher claims that the mean score of students in his class is
greater than 82 with a standard deviation of 20. If a sample of 81
students was selected with a mean score of 90 then check if there is
enough evidence to support this claim at a 0.05 significance level.
• A light bulb manufacturer claims that its' energy saving light bulbs last an
average of 60 days.
Set up a hypothesis test to check this claim and comment on what sort of
test we need to use.
Solution: So, we have
• H0: The mean lifetime of an energy-saving light bulb is 60 days.
• H1: The mean lifetime of an energy-saving light bulb is not 60 days.
Because of the “is not” in the alternative hypothesis, we have to consider
both the possibility that the lifetime of the energy-saving light bulb is
greater than 60 and that it is less than 60. This means we have to use a
two-tailed test.
• The manufacturer now decides that it is only interested whether the mean
lifetime of an energy-saving light bulb is less than 60 days. What changes
would you make from Example 1?
Solution: So, we have
• H0: The mean lifetime of an energy-saving light bulb is 60 days.
• H1: The mean lifetime of an energy-saving light bulb is less than 60 days.
• Now we have a “less than” in the alternative hypothesis. This means that
instead of performing a two-tailed test, we will perform a left-sided one-
tailed test.
• A z-test is a statistical test to determine whether two population
means are different when the variances are known and the sample size
is large.
• A z-test is a hypothesis test in which the z-statistic follows a normal
distribution.
• A z-statistic, or z-score, is a number representing the result from the z-
test.
• Z-tests are closely related to t-tests, but t-tests are best performed
when an experiment has a small sample size.
• Z-tests assume the standard deviation is known, while t-tests assume it
is unknown.
Thank You