0% found this document useful (0 votes)
10 views47 pages

Biometrical Analysis Techniques Overview

The document provides an overview of biometrical techniques used in research, focusing on univariate, bivariate, and multivariate analyses to assess variability and relationships among treatments. It outlines the objectives of a course on biometrical analysis, including understanding descriptive statistics, experimental designs, and hypothesis testing. Additionally, it defines key concepts such as population, sample, parameter, and various sampling techniques, emphasizing the importance of representative samples in statistical research.

Uploaded by

mihiretu muluneh
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
10 views47 pages

Biometrical Analysis Techniques Overview

The document provides an overview of biometrical techniques used in research, focusing on univariate, bivariate, and multivariate analyses to assess variability and relationships among treatments. It outlines the objectives of a course on biometrical analysis, including understanding descriptive statistics, experimental designs, and hypothesis testing. Additionally, it defines key concepts such as population, sample, parameter, and various sampling techniques, emphasizing the importance of representative samples in statistical research.

Uploaded by

mihiretu muluneh
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

CHAPTER I.

GLOSARY AND DESCRIPTIVE


STATISTICS

1. INTRODUCTION
Researchers use biometrical techniques to assess variability/diversity among
and within treatments, to study interrelationship between/among characters,
and GxE interaction and varietal performance stability. The common types of
statistical analysis include:

Univariate analysis: how do the treatments vary for a single trait at a time?
• E.g. descriptive statistics, 2-test, ANOVA, Stability, GxE
interaction, heritability, genetic advance, GCA, SCA, etc.

Bivariate analysis: How do the treatments co-vary for two traits at a time.
Bivariate analysis examines how two variables are related to each other. The
most common bivariate statistic is the bivariate correlation (often, simply
called “correlation”), which is a number between -1 and +1 denoting the
strength of the relationship between two variables. Let’s say that we wish to
study how age is related to self-esteem in a sample of 20 respondents, i.e., as
age increases, does self-esteem increase, decrease, or remains unchanged.
Other examples of bivariate analysis include regression, ANCOVA, co-
heritability, correlated response, etc.
Multivariate analysis: Multivariate statistical methods, or simply
multivariate methods, are statistical methods for the simultaneous analysis
of data on several variables. As the name indicates, multivariate analysis
consists of a collection of methods that can be used when several
measurements are made on each individual or object in one or more samples.
• E.g. diversity analysis, cluster analysis, population structure,
multiple correlation, multiple regression, PCA, etc.

The major objectives of the course include acquainting students with concepts
and principles of biometrical analysis of data and interpretation of the results.
At the end of the course students will at least be able to:
 Understand the basic concepts of descriptive statistics including types of
variables, measures of central tendency and dispersion
 Organize and present data set graphically and numerically with a
meaningful editorial and scientific standards
 Grasp the core principles of experimental designs and apply them in
various experiments
 Acquainted with the use of testing hypothesis for different parameters,
models of ANOVA/ANCOVA and analyze data using variances

1
 Use statistical inferences for comparing two and more than two means
 Explain the differences and special use of each experimental designs
 Investigate the relationships between parameters and develop knowledge
on using correlation and regression analysis

2. GLOSARY AND IMPORTANT DEFINITIONS


Biostatistics/biostatistics: A combination of biology and statistics - is the
application of statistics to a wide range of topics in biology. The science of
biostatistics encompasses the design of biological experiments, especially in
medicine and agriculture; the collection, summarization, and analysis of data
from those experiments; and the interpretation of, and inference from, the
results.

Research versus Experiment


Research: A structured enquiry that utilizes acceptable scientific methodology to
solve problems and create new knowledge that is generally applicable. Research
is also defined as the systematic process of collecting and analyzing
information to increase our understanding of the phenomenon under study.
The general aims of research are to observe and describe, to determine causes
and explain. Research employs a scientific method including systematic
observation, classification and interpretation of data to arrive at valid results.

Experiment: An experiment is any process or study which results in the


collection of data, the outcome of which is unknown. In statistics, the term is
usually restricted to situations in which the researcher has control over some
of the conditions under which the experiment takes place. For example, before
introducing a new drug treatment to reduce high blood pressure, the
manufacturer carries out an experiment to compare the effectiveness of the
new drug with that of one currently prescribed. Newly diagnosed subjects are
recruited from a group of local general practices. Half of them are chosen at
random to receive the new drug, the remainder receiving the present one. So,
the researcher has control over the type of subject recruited and the way in
which they are allocated to treatment.

Type of research: Research can be classified based on the application of


research results as: basic/pure research and applied research. Pure
research involves developing and testing theories and hypotheses that are
intellectually challenging to the researcher but may or may not have practical
application at the present time or in the future. The knowledge produced
through pure research is sought in order to add to the existing body of research
methods. Applied research is done to solve specific, practical questions; for
policy formulation, administration and understanding of a phenomenon. It can
be exploratory, but is usually descriptive. It is almost always done on the basis
of basic research. Applied research can be carried out by academic or

2
industrial institutions. Often, an academic institution such as a university will
have a specific applied research program funded by an industrial partner
interested in that program. Note that it is difficult to draw a clear boundary
between the two types of research; i.e. they should be regarded as being
mutually exclusive. Basic research lays the foundations for applied research
that follows. Applied research uses methodology that is not as rigorous as that
of basic/pure research, and its findings are evaluated in terms of local
applicability and not in terms of universal validity.

The research process: Generally, research involves what research questions


the researcher wants to answer and what method the researcher uses to find
answers to his research questions. There are several steps in the research
process and at each step; the researcher needs to choose the appropriate
method that enables him to conduct the research successfully.
1. Formulating the Research Problem
2. Extensive Literature Review
3. Developing the objectives
4. Preparing the Research Design including Sample Design
5. Collecting the Data
6. Analysis of Data
7. Generalization and Interpretation
8. Preparation of the Report or Presentation of Results-Formal write ups
of conclusions reached

3
Population versus Sample
Population: A population is any entire collection of people, animals, plants or
things from which we may collect data. It is the entire group we are interested
in, which we wish to describe or draw conclusions about. In order to make any
generalizations about a population, a sample, that is meant to be
representative of the population, is often studied. For each population there are
many possible samples. A sample statistic gives information about a
corresponding population parameter. For example, the sample mean for a set
of data would give information about the overall population mean. It is
important that the investigator carefully and completely defines the population
before collecting the sample, including a description of the members to be
included.

Target Population: The target population is the entire group a researcher is


interested in; the group about which the researcher wishes to draw
conclusions. For example, suppose we take a group of men aged 35-40 who
have suffered an initial heart attack. The purpose of this study could be to
compare the effectiveness of two drug regimes for delaying or preventing
further attacks. The target population here would be all men meeting the same
general conditions as those actually included in the study.

Sample: A sample is a group of units selected from a larger group (the


population). By studying the sample it is hoped to draw valid conclusions
about the larger group. A sample is generally selected for study because the
population is too large to study in its entirety. The sample should be
representative of the general population. This is often best achieved by random
sampling. Also, before collecting the sample, it is important that the researcher
carefully and completely defines the population, including a description of the
members to be included.

Matched Samples: Two samples in which the members are clearly paired, or
are matched explicitly by the researcher are called matched samples. For
example, matched samples can arise when IQ measurements are taken on
pairs of identical twins. Similarly, matched samples can also arise when the
same attribute, or variable, is measured twice on each subject, under different
circumstances. For example when the milk yields of cows is recorded before
and after being fed a particular diet. Sometimes, the difference in the value of
the measurement of interest for each matched pair is calculated, i.e. the
difference between before and after measurements for an appropriate statistical
analysis. We will re-examine this aspect in a more detail when we will be
discussing about paired-t test.

4
Parameter versus Statistic
Parameter and statistic: A parameter is a value, usually unknown (and which
therefore has to be estimated), used to represent a certain population
characteristic. For example, the population mean is a parameter that is often
used to indicate the average value of a quantity. Within a population, a
parameter is a fixed value which does not vary. Each sample drawn from the
population has its own value of statistic that is used to estimate this
parameter. A statistic is a quantity that is calculated from a sample of data. It
is used to give information about unknown values in the corresponding
population. For example, the mean of the data in a sample is used to give
information about the overall mean in the population from which that sample
was drawn. It is possible to draw more than one sample from the same
population and the value of a statistic will in general vary from sample to
sample. For example, the average value in a sample is a statistic. The average
values in more than one sample, drawn from the same population, will not
necessarily be equal. Statistics are often assigned Roman letters (e.g. m and s),
whereas the equivalent unknown values in the population (parameters) are
assigned Greek letters (e.g. µ and ).

Precision and Bias: Precision is a measure of how close an estimator is


expected to be to the true value of a parameter. Precision is usually expressed
in terms of imprecision and related to the standard error of the estimator.
Less precision is reflected by a larger standard error. Bias is a term which
refers to how far the average statistic lies from the parameter it is estimating,
that is, the error which arises when estimating a quantity. Errors from chance
will cancel each other out in the long run, those from bias will not.

SAMPLING IN STATISTICS
Sampling design: All items in any field of inquiry constitute a ‘Universe’ or
‘Population.’ A complete enumeration of all items in the ‘population’ is known
as a census inquiry. It can be presumed that in such an inquiry, when all
items are covered, highest accuracy is obtained. Example: the government
adopts this in very rare cases such as population census. But in practice this
type of inquiry involves a great deal of time, money and energy. Sometimes it is
possible to obtain sufficiently accurate results by studying only a part of total
population, a representative sample. The first stage is defining the target
population. A population can be defined as all people or items (unit of analysis)
with the characteristics that one wishes to study. Sometimes the population is
obvious. At other times, the target population may be a little harder to
understand. Note that samples may not entirely be representative of the
population at large, and if so, inferences derived by such a sample may not be
generalizable to the population.

5
CHARACTERISTICS OF A GOOD SAMPLE DESIGN
(a) Sample design must result in a truly representative sample
(b) Sample design must be such which results in a small sampling error
(c) Sample design must be viable in the context of funds available for the
research study
(d) Sample design must be such so that systematic bias can be controlled in
a better way, and
(e) Sample should be such that the results of the sample study can be
applied, in general, for the universe with a reasonable level of confidence.

Probability: A probability provides a quantitative description of the likely


occurrence of a particular event. Probability is conventionally expressed on a
scale from 0 to 1; a rare event has a probability close to 0, a very common
event has a probability close to 1. The probability of an event has been defined
as its long-run relative frequency. It has also been thought of as a personal
degree of belief that a particular event will occur (subjective probability). In
some experiments, all outcomes are equally likely. For example, when tossing a
coin, we assume that the results 'heads' or 'tails' each have equal probabilities
of 0.5. This is the equally-likely outcomes model and is defined to be:

number of outcomes corresponding to event E


P(E) =
total number of outcomes

Subjective Probability: A subjective probability describes an individual's


personal judgment about how likely a particular event is to occur. It is not
based on any precise computation but is often a reasonable assessment by a
knowledgeable person. Like all probabilities, a subjective probability is
conventionally expressed on a scale from 0 to 1; a rare event has a subjective
probability close to 0, a very common event has a subjective probability close to
1. A person's subjective probability of an event describes his/her degree of
belief in the event.

Conditional Probability: In many situations, once more information becomes


available, we are able to revise our estimates for the probability of further
outcomes or events happening. For example, suppose you go out for lunch at
the same place and time every Friday and you are served lunch within 15
minutes with probability 0.9. However, given that you notice that the
restaurant is exceptionally busy, the probability of being served lunch within
15 minutes may reduce to 0.7. This is the conditional probability of being
served lunch within 15 minutes given that the restaurant is exceptionally busy.

Independent Events: Two events are independent if the occurrence of one of


the events gives us no information about whether or not the other event will
occur; that is, the events have no influence on each other.

6
Mutually Exclusive Events: Two events are mutually exclusive (or disjoint) if
it is impossible for them to occur together. Formally, two events A and B are
mutually exclusive if and only if . For example a subject in a study
cannot be both male and female, nor can they be aged 20 and 30. A subject
could however be both male and 20, or both female and 30.

Probability Sampling: Probability sampling is a technique in which every unit


in the population has a chance (non-zero probability) of being selected in the
sample, and this chance can be accurately determined. Sample statistics thus
produced, such as sample mean or standard deviation, are unbiased estimates
of population parameters, as long as the sampled units are weighted according
to their probability of selection. All probability sampling have two attributes in
common:
(1) Every unit in the population has a known non-zero probability of
being sampled, and
(2) The sampling procedure involves random selection at some point.

TECHNIQUES OF PROBABILITY SMAPLING


Simple random sampling: In this technique, all possible subsets of a
population (more accurately, of a sampling frame) are given an equal
probability of being selected and, hence, sample statistics are unbiased
estimates of population parameters, without any weighting and the inferences
are most generalizable amongst all probability sampling techniques.

Systematic sampling: In this technique, the sampling frame is ordered


according to some criteria and elements are selected at regular intervals
through that ordered list. Systematic sampling involves a random start and
then proceeds with the selection of every element from that point onwards. This
process will ensure that there is no overrepresentation of any order in your
sample, but rather that firms of all orders are generally uniformly represented,
as it is in your sampling frame. In other words, the sample is representative of
the population, at least on the basis of the sorting criterion.

Stratified sampling: If a population from which a sample is to be drawn does


not constitute a homogeneous group, stratified sampling technique is generally
applied in order to obtain a representative sample. Under stratified sampling
the population is divided into several sub-populations that are individually
more homogeneous than the total population (the different sub-populations
are called ‘strata’) and then we select items from each stratum to constitute a
sample. Since each stratum is more homogeneous than the total population,
we are able to get more precise estimates for each stratum and by estimating
more accurately each of the component parts, we get a better estimate of the
whole. In brief, stratified sampling results in more reliable and detailed
information. The strata should be formed on the basis of common
characteristic(s) of the items to be put in each stratum. This means that
various strata be formed in such a way as to ensure elements being most

7
homogeneous within each stratum and most heterogeneous between the
different strata.

The following three questions are highly relevant in the context of stratified
sampling:
(1) How to form strata?
(2) How should items be selected from each stratum?
(3) How many items be selected from each stratum or how to allocate the
sample size of each stratum?

In respect of the second question, we can say that the usual method, for
selection of items for the sample from each stratum, resorted to is that of
simple random sampling.

Systematic sampling can also be used if it is considered more appropriate in


certain situations.

Regarding the third question, we usually follow the method of proportional


allocation under which the sizes of the samples from the different strata are
kept proportional to the sizes of the strata. That is, if Pi represents the
proportion of population included in stratum i, and n represents the total
sample size, the number of elements selected from stratum i is n*Pi.

To illustrate the last point, let us suppose that we want a sample of size n = 30
to be drawn from a population of size N = 8000 which is divided into three
strata of size N1 = 4000, N2 = 2400 and N3 = 1600. Adopting proportional
allocation, we shall get the sample sizes as under for the different strata: For
strata with N1 = 4000, we have P1 = 4000/8000 and hence n1 = n*P1 = 30
(4000/8000) = 15. Similarly, for strata with N2 = 2400, we have n2 = n*P2 = 30
(2400/8000) = 9, and for strata with N3 = 1600, we have n3 = n*P3 = 30
(1600/8000) = 6. This is called ‘optimum allocation’ in the context of
disproportionate sampling.

Cluster sampling: If the total area of interest happens to be a big one, a


convenient way in which a sample can be taken is to divide the area into a
number of smaller non-overlapping areas and then to randomly select a
number of these smaller areas (usually called clusters), with the ultimate
sample consisting of all (or samples of) units in these small areas or clusters.
Thus in cluster sampling the total population is divided into a number of
relatively small subdivisions which are themselves clusters of still smaller units
and then some of these clusters are randomly selected for inclusion in the
overall sample. Cluster sampling, no doubt, reduces cost by concentrating in
selected clusters but certainly it is less precise than random sampling.

Area sampling: If clusters happen to be some geographic subdivisions, in that


case cluster sampling is better known as area sampling. In other words, cluster

8
designs, where the primary sampling unit represents a cluster of units based
on geographic area, are distinguished as area sampling.

Multi-stage sampling: Multi-stage sampling is a further development of the


principle of cluster sampling. Suppose we want to make germplasm collection
in a country and we want to take a sample of few hotspots of genetic diversity
for this purpose. The first stage is to select large primary sampling unit such as
states/regions in a country. Then we may select certain districts and exploit
diversity hotspots the chosen districts. This would represent a two-stage
sampling design with the ultimate sampling units being clusters of districts.

Ordinarily multi-stage sampling is applied in big inquires extending to a


considerable large geographical area, say, the entire country. There are two
advantages of this sampling design namely:

(1) It is easier to administer than most single stage designs mainly


because of the fact that sampling frame under multi-stage sampling is
developed in partial units.
(2) A large number of units can be sampled for a given cost under
multistage sampling because of sequential clustering, whereas this is
not possible in most of the simple designs.

Non-probability sampling: A sampling technique in which some units of the


population have zero chance of selection or where the probability of selection
cannot be accurately determined. Typically, units are selected based on certain
non-random criteria, such as quota or convenience. Because selection is non-
random, non-probability sampling does not allow the estimation of sampling
errors, and may be subjected to a sampling bias. Therefore, information
from a sample cannot be generalized back to the population.

TECHNIQUES OF NON-PROBABILITY SAMAPLING

There are many types of non-probability sampling techniques which include:

 Convenience sampling. Also called, accidental or opportunity


sampling, is a technique in which a sample is drawn from that part of
the population that is close to hand, readily available, or convenient. For
instance, if you stand outside a shopping center and hand out
questionnaire surveys to people or interview them as they walk in, the
sample of respondents you will obtain will be a convenience sample. This
is a non-probability sample because you are systematically excluding all
people who shop at other shopping centers. This type of sampling is
most useful for pilot testing, where the goal is instrument testing or
measurement validation rather than obtaining generalizable inferences.

 Quota sampling. In this technique, the population is segmented into


mutually exclusive subgroups (just as in stratified sampling), and then a
9
non-random set of observations is chosen from each subgroup to meet a
predefined quota.

 Expert sampling. This is a technique where respondents are chosen in a


non-random manner based on their expertise on the phenomenon being
studied. The advantage of this approach is that since experts tend to be
more familiar with the subject matter than non-experts, opinions from a
sample of experts are more credible than a sample that includes both
experts and non-experts, although the findings are still not
generalizable to the overall population at large.

 Snowball sampling. In snowball sampling, you start by identifying a few


respondents that match the criteria for inclusion in your study, and then
ask them to recommend others they know who also meet your selection
criteria. For instance, if you wish to survey computer network
administrators and you know of only one or two such people, you can
start with them and ask them to recommend others who also do network
administration. Although this method hardly leads to representative
samples, it may sometimes be the only way to reach hard-to reach
populations or when no sampling frame is available.

 Blinding: In a medical experiment, the comparison of treatments may be


distorted if the patient, the person administering the treatment and those
evaluating it know which treatment is being allocated. It is therefore
necessary to ensure that the patient and/or the person administering the
treatment and/or the trial evaluators are 'blind to' (don't know) which
treatment is allocated to whom. Sometimes the experimental set-up of a
clinical trial is referred to as double-blind, that is, neither the patient nor
those treating and evaluating their condition are aware (they are 'blind'
as to) which treatment a particular patient is allocated. A double-blind
study is the most scientifically acceptable option. Sometimes however, a
double-blind study is impossible, for example in surgery. It might still be
important though to have a single-blind trial in which the patient only is
unaware of the treatment received, or in other instances, it may be
important to have blinded evaluation.

STATISTICS OF SAMAPLING
Imagine that you took three different random samples from a given population,
and for each sample, you derived sample statistics such as sample mean and
standard deviation. If each random sample was truly representative of the
population, then your three sample means from the three random samples will
be identical (and equal to the population parameter), and the variability in
sample means will be zero. But this is extremely unlikely, given that each
random sample will likely constitute a different subset of the population, and
hence, their means may be slightly different from each other.

10
You can take these three sample means and plot a frequency histogram of
sample means. If the number of such samples increases from three to 10 to
100, the frequency histogram becomes a sampling distribution. Hence, a
sampling distribution is a frequency distribution of a sample statistic (like
sample mean) from a set of samples, while the commonly referenced frequency
distribution is the distribution of a response (observation) from a single sample.
Just like a frequency distribution, the sampling distribution will also tend to
have more sample statistics clustered around the mean (which presumably is
an estimate of a population parameter), with fewer values scattered around the
mean. With an infinitely large number of samples, this distribution will
approach a normal distribution.

A response is a measurement value provided by a sampled unit. Responses of


different treatments to the same item or observation can be graphed into a
frequency distribution based on their frequency of occurrences. For a large
number of responses in a sample, this frequency distribution tends to resemble
a bell-shaped curve called a normal distribution, which can be used to
estimate overall characteristics of the entire sample, such as sample mean
(average of all observations in a sample) or standard deviation (variability or
spread of observations in a sample).

These sample estimates are called sample statistics (a “statistic” is a value


that is estimated from observed data). Populations also have means and
standard deviations that could be obtained if we could sample the entire
population. However, since the entire population can never be sampled,
population characteristics are always unknown, and are called population
parameters (and not “statistic” because they are not statistically estimated
from data). Sample statistics may differ from population parameters if the
sample is not perfectly representative of the population; the difference
between the two is called sampling error.

Theoretically, if we could gradually increase the sample size so that the sample
approaches closer and closer to the population, then sampling error will
decrease and a sample statistic will increasingly approximate the
corresponding population parameter. If a sample is truly representative of the
population, then the estimated sample statistics should be identical to
corresponding theoretical population parameters. How do we know if the
sample statistics are at least reasonably close to the population parameters?
Here, we need to understand the concept of sampling distribution.

The variability or spread of a sample statistic in a sampling distribution (i.e.,


the standard deviation of a sampling statistic) is called its standard error. In
contrast, the term standard deviation is reserved for variability of an observed
response from a single sample. The mean value of a sample statistic in a
sampling distribution is presumed to be an estimate of the unknown
population parameter. Based on the spread of this sampling distribution (i.e.,
based on standard error), it is also possible to estimate confidence intervals for
11
that prediction population parameter. Confidence interval is the estimated
probability that a population parameter lies within a specific interval of
sample statistic values.

Confidence Level: The confidence level is the probability value


associated with a confidence interval. It is often expressed as a percentage. For
example, say , then the confidence level is equal to (1-0.05) =
0.95, i.e. a 95% confidence level. Confidence limits are the lower and upper
boundaries/values of a confidence interval, that is, the values which define the
range of a confidence interval.

Confidence Interval for the Difference between Two Means: A confidence


interval for the difference between two means specifies a range of values within
which the difference between the means of the two populations may lie. These
intervals may be calculated. The confidence interval for the difference between
two means contains all the values of µ1 - µ2 (the difference between the two
population means) which would not be rejected in the two-sided hypothesis
test of: H0: µ1 = µ2 against H1: µ1 not equal to µ2 i.e. H0: µ1 - µ2 = 0 against
H1: µ1 - µ2 not equal to 0. If the confidence interval includes 0 we can say that
there is no significant difference between the means of the two populations, at
a given level of confidence. The width of the confidence interval gives us some
idea about how uncertain we are about the difference in the means. A very wide
interval may indicate that more data should be collected before anything
definite can be said. We calculate these intervals for different confidence levels,
depending on how precise we want to be. We interpret an interval calculated at
a 95% level as we are 95% confident that the interval contains the true
difference between the two population means. We could also say that 95% of all
confidence intervals formed in this manner (from different samples of the
population) will include the true difference.

A General Rule of Probability Sampling: All normal distributions tend to


follow a 68-95-99 percent rule (see Figure below), which says that over 68% of
the cases in the distribution lie within one standard deviation of the mean
value (μ + 1σ), over 95% of the cases in the distribution lie within two standard
deviations of the mean (μ +2σ), and over 99% of the cases in the distribution lie
within three standard deviations of the mean value (μ + 3σ). Since a sampling
distribution with an infinite number of samples will approach a normal
distribution, the same 68-95-99 rule applies, and it can be said that: Sample
statistic + one standard error represents a 68% confidence interval for the
population parameter. Similarly sample statistic + two standard errors
represents a 95% confidence interval for the population parameter and sample
statistic + three standard errors represents a 99% confidence interval for the
population parameter.

A sample is “biased” (i.e., not representative of the population) if its sampling


distribution cannot be estimated or if the sampling distribution violates the 68-

12
95-99 percent rule. As an aside, note that in most regression analysis where we
examine the significance of regression coefficients with p<0.05, we are
attempting to see if the sampling statistic predicts the corresponding
population parameter (true effect size) with a 95% confidence interval.

13
DISTRIBUTION IN STATISTICS

Normal Distribution: The normal frequency distribution of a continuous


random variable tends to resemble a bell-shaped curve as stated earlier. For
example, height at a given age for a given gender in a given racial group is
adequately described by a normal random variable even though heights must
be positive. Many distributions arising in practice can be approximated by a
normal distribution and other random variables may be transformed to
normality. The simplest case of the normal distribution, known as the
Standard Normal Distribution, has expected value zero and variance one. This
is written as N(0,1).

A continuous random variable X, taking all real values in the range is


said to follow a Normal distribution with parameters µ and :

Poisson Distribution: Poisson distributions of discrete random variables a


count of the number of events that occur in a certain time interval or spatial
area. For example, the number of cars passing a fixed point in a 5 minute
interval, or the number of calls received by a switchboard during a given period
of time.

Symmetry and Skewness: Symmetry is implied when data values are


distributed in the same way above and below the middle of the sample.
Symmetrical data sets: (1) are easily interpreted; (2) allow a balanced attitude
to outliers, that is, those above and below the middle value ( median) can be
considered by the same criteria; and (3) allow comparisons of spread or
dispersion with similar data sets. Many standard statistical techniques are
appropriate only for a symmetric distributional form. For this reason, attempts

14
are often made to transform skewed data so that they become roughly
symmetric.

Skewness is defined as asymmetry in the distribution of the sample data


values. Values on one side of the distribution tend to be further from the
'middle' than values on the other side. For skewed data, the usual measures of
location will give different values, for example, mode<median<mean would
indicate positive (or right) skewness. Positive (or right) skewness is more
common than negative (or left) skewness. If there is evidence of skewness in the
data, we can apply transformations, for example, taking logarithms of positive
skew data.

A discrete random variable X is said to follow a Poisson distribution with


parameter m, written X ~ Po(m):
x = 0, 1, 2, ..., n and m > 0.

Expected Value: The expected value (or population mean) of a random variable
indicates its average or central value. It is a useful summary value (a number)
of the variable's distribution. Stating the expected value gives a general
impression of the behavior of some random variable without giving full details
of its probability distribution. Two random variables with the same expected
value can have very different distributions. There are other useful descriptive
measures which affect the shape of the distribution, for example variance. The
expected value of a random variable X is symbolized by E(X) or µ.

15
Estimation and Statistical Inference: Estimation is the process by which
sample data are used to indicate the value of an unknown quantity in a
population. Results of estimation can be expressed as a single value, known as
a point estimate, or a range of values, known as a confidence interval.
Statistical Inference makes use of information from a sample to draw
conclusions (inferences) about the population from which the sample was
taken.

Estimator: An estimator is any quantity calculated from the sample data


which is used to give information about an unknown quantity in the
population. For example, the sample mean is an estimator of the population
mean. Estimators of population parameters are sometimes distinguished from
the true value by using the symbol 'hat'. For example, = true population
standard deviation and = estimated (from a sample) population standard
deviation. The usual estimator of the population mean is

where n is the size of the sample and X1, X2, X3, ......., Xn are the values of the
sample. If the value of the estimator in a particular sample is found to be 5,
then 5 is the estimate of the population mean µ.

Estimate: An estimate is an indication of the value of an unknown quantity


based on observed data. More formally, an estimate is the particular value of
an estimator that is obtained from a particular sample of data and used to
indicate the value of a parameter.

Central Limit Theorem: The Central Limit Theorem states that whenever a
random sample of size n is taken from any normal distributed population with
mean µ and variance , then the sample mean will be approximately
normally distributed with mean µ and variance /n. The larger the value of
the sample size n, the better the approximation to the normal. This is very
useful when it comes to inference as the tests use the sample mean , which
the Central Limit Theorem tells us will be approximately normally distributed.

TYPES OF DATA
Discrete Data: A set of data is said to be discrete if the values /observations
belonging to it are distinct and separate, i.e. they can be counted (1,2,3,....).
Examples might include the number of kittens in a litter; the number of
patients in a doctors surgery; gender (male, female); blood group (O, A, B, AB).

Categorical Data: A set of data is said to be categorical if the values or


observations belonging to it can be sorted according to category. Each value is

16
chosen from a set of non-overlapping categories. For example, shoes in a
cupboard can be sorted according to colour: the characteristic 'colour' can have
non-overlapping categories 'black', 'brown', 'red' and 'other'. People have the
characteristic of 'gender' with categories 'male' and 'female'. Categories should
be chosen carefully since a bad choice can prejudice the outcome of an
investigation. Every value should belong to one and only one category, and
there should be no doubt as to which one.

Nominal Data: A set of data is said to be nominal if the values / observations


belonging to it can be assigned a code in the form of a number where the
numbers are simply labels. You can count but not order or measure nominal
data. For example, in a data set males could be coded as 0, females as 1;
marital status of an individual could be coded as Y if married, N if single.

Ordinal Data: A set of data is said to be ordinal if the values/observations


belonging to it can be ranked (put in order) or have a rating scale attached. You
can count and order, but not measure, ordinal data. The categories for an
ordinal set of data have a natural order. Suppose a group of people were asked
to taste varieties of biscuit and classify each biscuit on a rating scale of 1 to 5,
representing strongly dislike, dislike, neutral, like, strongly like. A rating of 5
indicates more enjoyment than a rating of 4, for example, so such data are
ordinal. However, the distinction between neighboring points on the scale is not
necessarily always the same. For instance, the difference in enjoyment
expressed by giving a rating of 2 rather than 1 might be much less than the
difference in enjoyment expressed by giving a rating of 4 rather than 3.

Interval Scale: An interval scale is a scale of measurement where the distance


between any two adjacent units of measurement (or 'intervals') is the same but
the zero point is arbitrary. Scores on an interval scale can be added and
subtracted but cannot be meaningfully multiplied or divided. For example, the
time interval between the starts of years 1981 and 1982 is the same as that
between 1983 and 1984, namely 365 days. The zero point, year 1 AD, is
arbitrary; time did not begin then. Other examples of interval scales include the
heights of tides, and the measurement of longitude.

Continuous Data: A set of data is said to be continuous if the


values/observations belonging to it may take on any value within a finite or
infinite interval. You can count, order and measure continuous data. For
example height, weight, temperature and the amount of sugar in an orange.

17
TEST STATISTICS/TEST OF SIGNIFICANCE
Statistical testing is always probabilistic, because we are never sure if our
inferences, based on sample data, apply to the population, since our sample
never equals the population. The probability that a statistical inference is
caused by pure chance is called the p-value. The p-value is compared with the
significance level (α), which represents the maximum level of risk that we are
willing to take that our inference is incorrect. For most statistical analysis, α is
set to 0.05. A p-value less than α=0.05 indicates that we have enough
statistical evidence to reject the null hypothesis, and thereby, indirectly accept
the alternative hypothesis. If p>0.05, then we do not have adequate statistical
evidence to reject the null hypothesis or accept the alternative hypothesis.

Hypothesis Test:

The hypotheses are often statements about population parameters like


expected value and variance; for example Ho might be that the expected value
of the height of ten year old boys in a population is not different from that of
ten year old girls. A hypothesis might also be a statement about the
distributional form of a characteristic of interest, for example that the height of
ten year old boys is normally distributed within the a population. The outcome
of a hypothesis test is "Reject Ho in favor of H1" or "Do not reject Ho”.

Null Hypothesis
The null hypothesis, H0, represents a theory that has been put forward, either
because it is believed to be true or because it is to be used as a basis for
argument, but has not been proved. For example, in a clinical trial of a new
drug, the null hypothesis might be that the new drug is no better, on average,
than the current drug. We would write:
H0: there is no difference between the two drugs on average.

We give special consideration to the null hypothesis. This is due to the fact that
the null hypothesis relates to the statement being tested, whereas the
alternative hypothesis relates to the statement to be accepted if/when the null
is rejected. The final conclusion once the test has been carried out is always
given in terms of the null hypothesis. We either "Reject H0 in favor of H1" or
"Do not reject H0"; we never conclude "Reject H1", or even "Accept H1".

If we conclude "Do not reject H0", this does not necessarily mean that the null
hypothesis is true; it only suggests that there is not sufficient evidence against
H0 in favor of H1. Rejecting the null hypothesis then, suggests that the
alternative hypothesis may be true.

Alternative Hypothesis
The alternative hypothesis, H1, is a statement of what a statistical hypothesis
test is set up to establish. For example, in a clinical trial of a new drug, the
alternative hypothesis might be that the new drug has a different effect, on
average, compared to that of the current drug. We would write:
18
H1: the two drugs have different effects, on average.
The alternative hypothesis might also be that the new drug is better, on
average, than the current drug. In this case we would write:
H1: the new drug is better than the current drug, on average.

THE 2 TEST

19
THE Z TEST

20
21
22
23
24
25
26
27
28
29
30
31
32
SOME COMMON METHODS OF SUMMARIZING DATA (Descriptive)
Frequency Table: A frequency table is a way of summarizing a set of data. It is
a record of how often each value (or set of values) of the variable in question
occurs. It may be enhanced by the addition of percentages that fall into each
category. A frequency table is used to summarize categorical, nominal, and
ordinal data. It may also be used to summarize continuous data once the data
set has been divided up into sensible groups. When we have more than one
categorical variable in our data set, a frequency table is sometimes called a
contingency table because the figures found in the rows are contingent upon
(dependent upon) those found in the columns.

Example, the frequencies of the different diseases scores of a given crop can be
summarized as:

Score Frequency Frequency (%)

0 4 13%

1 3 10%

2 5 17%

3 5 17%

4 6 20%

5 7 23%

Pie Chart: A pie chart is a way of summarizing a set of categorical data. It is a


circle which is divided into segments. Each segment represents a particular
category. The area of each segment is proportional to the number of cases in
that category. Suppose that the area covered by three major cereals in a given
year in Ethiopia was 6 million hectares; 3 million to teff, 2 million to maize and
1 million to sorghum. The area allotment can be summarized using a pie chart:

33
50%

33.3
33.3 %
%

Bar Chart: A bar chart is a way of summarizing a set of categorical data. It is


often used in exploratory data analysis to illustrate the major features of the
distribution of the data in a convenient form. It displays the data using a
number of rectangles, of the same width, each of which represents a particular
category. The length (and hence area) of each rectangle is proportional to the
number of cases in the category it represents, for example, age group. Bar
charts are used to summarize nominal or ordinal data. Bar charts can be
displayed horizontally or vertically and they are usually drawn with a gap
between the bars (rectangles), whereas the bars of a histogram are drawn
immediately next to each other.

Dot Plot: A dot plot is a way of summarizing data, often used in exploratory
data analysis to illustrate the major features of the distribution of the data in a
convenient form. For nominal or ordinal data, a dot plot is similar to a bar
chart, with the bars replaced by a series of dots. Each dot represents a fixed
number of individuals. For continuous data, the dot plot is similar to a

34
histogram, with the rectangles replaced by dots. A dot plot can also help detect
any unusual observations (outliers), or any gaps in the data set. The Figures
presented below shows the revenues of 60 companies in dot plot and a bar
chart for your comparison. Note that the dot plot is less cluttered, less
redundant, and uses less ink as compared to the bar graph.

The power of the dot plot becomes evident if we wish to combine the
information from two or more information into a single chart. For example,
both the revenues and the profits of the companies stated above can be easily
presented in a single dot plot as given below. The presentation of such
information would be much more cluttered and more difficult to interpret with
a bar chart. Another advantage of dot plot is that it does not depend on color so
that it can be used in black and white publications with no loss of clarity. The
two groups can be distinguished by using different symbols.

35
Histogram: A histogram is a way of summarizing data that are measured on
an interval scale (either discrete or continuous). It is often used in exploratory
data analysis to illustrate the major features of the distribution of the data in a
convenient form. It divides up the range of possible values in a data set into
classes or groups. For each group, a rectangle is constructed with a base
length equal to the range of values in that specific group, and an area
proportional to the number of observations falling into that group. This means
that the rectangles might be drawn of non-uniform height. The histogram is
only appropriate for variables whose values are numerical and measured on an
interval scale. It is generally used when dealing with large data sets (>100
observations), when stem and leaf plots become tedious to construct. A
histogram can also help detect any unusual observations (outliers), or any gaps
in the data set. Compare it with the bar chart.

Box and Whisker Plot (Box plot): A box and whisker plot is a way of
summarizing a set of data measured on an interval scale. It is often used in
exploratory data analysis. It is a type of graph which is used to show the shape
of the distribution, its central value, and variability. The picture produced
consists of the most extreme values in the data set (maximum and minimum
values), the lower and upper quartiles, and the median. A box plot (as it is
often called) is especially helpful for indicating whether a distribution is skewed
and whether there are any unusual observations (outliers) in the data set (see
Figure below). Box and whisker plots are very useful when large numbers of
observations are involved and when two or more data sets are being compared.

36
Scatter Plot: A scatter plot is a useful summary of a set of bivariate data (two
variables), usually drawn before working out a linear correlation coefficient or
fitting a regression line. It gives a good visual picture of the relationship
between the two variables, and aids the interpretation of the correlation
coefficient or regression model. Each unit contributes one point to the scatter
plot, on which points are plotted but not joined. The resulting pattern indicates
the type and strength of the relationship between the two variables, i.e. the
more the points tend to cluster around a straight line, the stronger the linear
relationship between the two variables (the higher the correlation). If the line
around which the points tends to cluster runs from lower left to upper right,
the relationship between the two variables is positive (direct). Similarly, if the
line around which the points tends to cluster runs from upper left to lower
right, the relationship between the two variables is negative (inverse). However,
if there exists a random scatter of points, there is no relationship between the
two variables (very low or zero correlation). Very low or zero correlation could
result from a non-linear relationship between the variables. If the relationship
is in fact non-linear (points clustering around a curve, not a straight line), the
correlation coefficient will not be a good measure of the strength. A scatterplot
will also show up a non-linear relationship between the two variables and
whether or not there exist any outliers in the data. Note that we can also use a
three-dimensional graph when we are dealing with three variable (but not two).

37
[Link] STATISTICS

What is statistics?
Statistics is a branch of mathematics that deals with the collection,
organization, and analysis of numerical data and with such problems as
experimental design and decision making. In conducting agriculture research,
huge data is collected in various experiments of breeding, agronomy, crop
protection, and etc. Hence, knowledge of statistics is essential for the research
for data collection, organization, summarizing and analysis and proper
interpretation of the results. In every statistical analysis some widely used
statistical estimates include mean, range, standard deviation, standard
error, variances and coefficient of variation. Some of them are defined
below:

Samples and Populations


In agricultural research it is important to make inferences (draw conclusions)
about a population, which is defined as the collection of all possible
observations of interest. In studying statistics, it is not always possible to deal
with populations because of their enormous size. For example, if we want to
know the average height of the human population, it is impossible to measure
each and every human being. If we want to know the average yield of barley in
a particular country, we cannot weigh all what harvested in that country. In
this case, we have to deal with a sample where proper statistical analysis can
be done. A sample is a representative group taken at random from a population
and the number of observations in the sample is called the sample size. When
a sample is properly taken, the statistics from that sample can be applied to
the population. Measured characteristics of the sample are called statistics
(e.g. sample mean) and characteristics of the population are called parameters
(e.g. population mean).

Descriptive statistics
In examining large collections of numbers, such as census data, it is helpful to
be able to present a number that provides a summary of the data. Such
numbers are often called descriptive statistics. The arithmetic mean is probably
the best-known descriptive statistic. The mean is often called the average, but
it is actually only one of several kinds of averages, such as the median and the
mode.

38
Measures of central tendency

Mean

The sample mean is an estimator available for estimating the population mean
. It is a measure of location, commonly called the average. Its value depends
equally on all of the data which may include outliers. It may not appear
representative of the central region for skewed data sets. It is especially useful
as being representative of the whole sample for use in subsequent calculations.

Given the set of observations, symbolized by xi: 3, 6, 2, 5, 4, 3


We see that, for this set, n6. The mean x , of the set is
6
x = ( xi) / 6  (x1 x2 x3 x4 x5 x6)/6
i 1

 (362543)/6  23/6  3.83

The median and the mode

The median and the mode are two other measures of central tendency a set of
discrete data. Mode, i.e. the number in a given set of numbers that appears
most frequently. Let the x's be arranged in numerical order; if n is odd, the
median is the middle x; if n is even, the median is the average of the two middle
x's. The mode is the x that occurs most frequently. If two or more distinct x's
occur with equal frequencies, but none with greater frequency, the set of x's
may be said not to have a mode or to be bimodal, with modes at the two most
frequent x's, or trimodal, with modes at the three most frequent x's. In the set
{3, 4, 6, 7, 10, 10, 13}, for example, the mode of the set is 10. If two or more
numbers are tied for most frequent appearances the set has multiple modes.
The modes of the set {1, 1, 2, 2, 3, 4, 4, 5}, for example, are 1, 2, and 4. Other
sets, such as {5, 7, 9, 11}, have no modes because all the numbers occur with
equal frequency.

The mode provides a way to summarize a set of numbers without examining all
of the numbers in the set. Knowing that the mode of a class’s scores on an
algebra test was 100 percent, for example, suggests that the test may have
been too easy, since a perfect score was the most common result. Like the

39
mean (or average) and mode of a set of numbers, the median can be used to get
an idea of the distribution or spread of values within a set when examining
every value individually would be overwhelming or tedious. The median is the
value halfway through the ordered data set, below and above which there lies
an equal number of data values. The median of the set {1, 3, 7, 8, 9}, for
example, is 7, because 7 is the member of the set that has an equal number of
members on each side of it when the members are arranged from lowest to
highest. If a set contains an even number of values, there is no single middle
member. In such cases the median is the mean of the two values closest to the
middle. The median of the set {1, 3, 9, 10}, for example, is (3 + 9)/2 = 6.

The mean is a more precise measure than the median, but can be greatly
affected by a few numbers that are very different from the other members of a
set. For example, the mean of the set {2, 4, 5, 7, 8, 934}—calculated by adding
the members of the set together and dividing the sum by the total number of
members—is 160, which is much higher than all but one of the values in the
set. In cases such as this the median, 6, is used to give a better overall
impression of the typical values of the numbers because it ignores outlying
values.

Weighted Mean
When we compute a simple arithmetic mean of a set of data, we assume that
all the observed values are of equal importance and we give them equal weight
in our calculation. In situations where the numbers are not equally important
we can assign to each a weight which is proportional to its relative importance
and calculate the weighted mean. Let V1, V2, …, Vk be a set of k values, and
let w1, w2, …, wk be the weights assigned to them. The weighted mean is found
by dividing the sum of the products of the values and their weights by the sum
of the weights; that is:

V = w1V1 + w2V2 + w3V3 + …. + wkVk


w1 + w2 + w3 + …. + wk
In summation notation, this is
V =  wiVi

 wi
Every student is familiar with the concept of the weighted mean; the grade-
point is such a measure. It is the mean of the numerical values of the letter
grades weighed by the number of credit hours in which the various grades are
earned. Suppose these numerical values are A=4, B=3, C=2, D=1. If a student
takes a 3-credit course and makes an A it (w1 = 3, V1 = 4) a 5-credit course
and makes a B (w2 = 5, V2 = 3), another 3-redit course makes a C (W3 = 3, V3
= 2) and a 2-credit course and makes A (w4 = 2, V4 = 4),, then the grade-point
average for the term is:

40
V = (3 x 4) + (5 x 3) + (3 x 2) + (2 x 4)
3+5+3+2
The weighting procedure is also used to find the mean when several sets of
data are combined. Suppose we have three sets of data consisting of n1, n2,
and n3 observed values and having the means X1, X2, X3, respectively. Then
the mean for the combined data is the weighted average of the individual
means, the respective weights being the sample sizes n1, n2, and n3.
X = n1X1 + n2X2 + n3X3
N1 + n2 + n3
Failure to weight the means when combining data is not an uncommon error.
Imagine the male student body of a school split into two groups. In the first
group the mean height is 75 inches, and in the second group the mean height
is 69 inches. The average (75 + 69)/2 = 72 is not the mean height of male
students in the school if the first group consists of the 15 members of the
basketball team and the second group consists of the remaining 568 male
students in the school. The true mean then is:
(15 x 75) + (568 x 69) = 69.15 inches
583
in accordance with the equation above for two (rather than three means).

Measures of dispersion (variability)

The data values in a sample are not all the same because of sampling
variability. Sampling variability refers to the different values which a given
function of the data takes when it is computed for two or more samples drawn
from the same population. This variation between values is called dispersion.
When the dispersion is large, the values are widely scattered; when it is small
they are tightly clustered. The width of diagrams such as dot plots and box
plots is greater for samples with more dispersion and vice versa. There are
several measures of dispersion, the most common being the standard
deviation. These measures indicate to what degree the individual observations
of a data set are dispersed or 'spread out' around their mean. In measurement,
high precision is associated with low dispersion.

Range
The investigator frequently is concerned with the variability of the distribution,
that is, whether the measurements are clustered tightly around the mean or
spread over the range.

41
Variance
The variance of a set of n observation, x1 x2, … xn, is the sum of squares of the
difference between the observations and their mean divided by one less than
the number of observations. The variance is usually symbolized by S 2 (s-
squared) using summation notation, the defining formula for the variance is:
n
S2   ( xi  x )2/(n-1)
i1
Note that it can be shown that
n n n

 ( xi  x )2 
i1
 x i2 – (  xi )2/n
i1 i1
We should examine this expression in detail:
n

 xi 2  x12  x22  . . .  xn2 . This term is called the raw or uncorrected, sum
i1
of square.
n
(  xi )2/n  (x1 x2 . . .  xn)2/n. This term is called the correction term,
i1
symbolized c.t. The raw sum of squares minus the correction term is called
sum of squares, symbolized as SS. The divisor of S2 n-1 is called the degree
of freedom symbolized as d.f. From this we see that the following relationships
hold:
n n n
S2   ( xi  x )2/(n-1)  [  xi 2 - (  xi )2/n] / (n-1)
i1 i1 i1

 [Raw SS – c.t.] / (n-1)  SS/d.f.

Suppose we compute the variance, S2, for the set of numbers


3, 6, 2, 5, 4, 3, for which we computed the mean:
6
1. (  xi )  (362543)  23
i1
6
2. C.T.  (  xi )2/n  (23)2/6  526/6  88.17
i1
n
3.  xi 2  3262 22524232  99.00  Raw SS.
i1

4. SS  Raw SS – C.T.  99.00-88.17  10.83


5. S2  SS/d.f.  SS/(n-1)  10.83/(6-1)  2.17

Standard deviation
The standard deviation is a measure of variability that is more convenient in
analysis of statistical data. The square, σ2, of the standard deviation is called
the variance. If the standard deviation is small, the measurements are tightly
clustered around the mean; if it is large, they are widely scattered. The
standard deviation is defined as the root-mean-square (RMS) deviation of the
values from their mean, or as the square root of the variance. It is the most

42
important measure of statistical dispersion, measuring how widely spread the
values in the data set are. If many data points are close to the mean, the
standard deviation is small; if many data points are far from the mean, then
the standard deviation is large. If all data values are equal, then the standard
deviation is zero. A useful property of standard deviation is that, unlike
variance, it is expressed in the same units as the data.

Standard Errors
The standard error (SE) is the measure of the difference between sample mean
( x ) and the population mean (). Thus it is the measure of uncontrolled
variation present in a sample. The standard error of the mean of n
observations, x1 x2, … xn, is the square root of the variance of the observations
divided by the number of observations. The standard error of the mean is
usually symbolizes S x . By definition we have
Sx  S2 /n
For the set of data we have been examining we have S 2  2.17 and n  6. The
standard error is S x  S2 /n  2.17 / 6  0.36  0.60

43
44
45
46
REFERENCES

47

You might also like