Basic Statistics
Definition of statistics:
Statistics is a branch of mathematics that deals with collecting, organizing, and analyzing numerical data to
solve real-world problems. It is widely regarded as a distinct scientific discipline due to its vast applications
across various fields.
It encompasses both descriptive methods for summarizing data and inferential methods for drawing
conclusions about a larger population based on a sample.
Here's a more detailed breakdown:
1. Collection: This involves gathering data through various methods like surveys, experiments, or
observations.
2. Organization: Collected data is structured and arranged in a meaningful way, often using tables,
charts, or graphs.
3. Analysis: Statistical techniques are applied to examine the data, identify patterns, and extract
meaningful insights.
4. Interpretation: The results of the analysis are explained and understood in the context of the problem
or question being investigated.
5. Presentation: The findings are communicated effectively through visual aids or reports.
Types of Statistics
There are 2 types of statistics:
1. Descriptive Statistics
2. Inferential Statistics
Descriptive Statistics
Descriptive statistics uses data that describes the population either through numerical calculated graphs or
tables. It provides a graphical summary of data.
It is simply used for summarizing objects, etc. There are two categories in this as follows.
Measure of Central Tendency
Measure of Variability
Inferential Statistics
Inferential Statistics makes inferences and predictions about the population based on a sample of data taken
from the population. It generalizes a large dataset and applies probabilities to conclude.
It is simply used for explaining the meaning of descriptive stats. It is simply used to analyze, interpret
results, and draw conclusions. Inferential Statistics is mainly related to and associated with hypothesis
testing, whose main target is to reject the null hypothesis.
Types of Inferential Statistics
One sample test of difference/One sample hypothesis test
Confidence Interval
Contingency Tables and Chi-Square Statistic
T-test or Anova
Pearson Correlation
Bivariate Regression
Multi-variate Regression
Data:
Data refers to a collection of facts, figures, or information gathered through observations, measurements, or
analysis. It can be numerical, such as measurements or counts, or non-numerical, like descriptions or
categories.
Data serves as the raw material for statistical analysis, providing the basis for understanding patterns, trends,
and relationships within a specific context.
Geographical data:
Geographical datasets mostly consist of numerical measurements. Each number is a measurement of a
specified property of one particular entity. In statistical terminology, geographical objects or samples form
the entity and the measured attributes on these are called variables. Thus, a variable is a characteristic or
attribute that can assume different values. It may be independent or dependent.
The independent variable is also called the explanatory variable. In an experimental study, this is
the one that is being manipulated by the researcher. It is always plotted on the x-axis.
The resultant variable is called the dependent variable or the outcome variable and is essentially
plotted on the y-axis.
Variables can be classified as either qualitative or quantitative.
Qualitative variables are those that can be placed into distinct categories (gender: male/female).
Quantitative variables, on the other hand, are numerical and can be ordered or ranked (age: people
can be ranked).
Again, these may be further divided into two categories, discrete or continuous.
Discrete variables are values that can be counted (e.g., number of children in a family, number of
students in a class, frequency, count data, population density, etc.).
Continuous variables, on the contrary, can assume an infinite number of values between any two
specific values, obtained by measuring, often fractions and decimals (rainfall, elevation, weight, area,
length, etc).
Type of Geographical Data
Terminology of statistics
1. Population: A population is any complete group with at least one characteristic in common.
Populations are not just people. Populations may consist of, but are not limited to, people, animals,
businesses, buildings, motor vehicles, farms, objects or events.
2. Sample: A sample is a subset of the population that is selected for analysis. It's impractical or
impossible to collect data from an entire population in many cases, so researchers often work with a
sample to make inferences about the population. For instance, a researcher might select a group of
100 students from the school to represent the entire student population.
3. Variable: A variable is a characteristic of the individuals or objects being studied that can vary or
take on different values. For example, in the study of student heights, the height of each student
would be a variable.
4. Parameter: A parameter is a numerical characteristic of a population. It's a fixed value that describes
the entire population, but it is usually unknown and estimated using sample statistics. In the student
height example, the average height of all students in the school would be a population parameter.
Sampling: Concepts, Types and Significance.
Sampling is a statistical technique for efficiently analyzing large datasets by selecting a representative
subset. Rather than analyzing an entire dataset, sampling analyzes a small portion so researchers can make
conclusions about a larger population. This allows for informed decision-making without exhaustive data
collection.
Sampling is the process of selecting units (e.g., people, organizations) from a population of interest so that
by studying the sample we may fairly generalize our results back to the population from which they were
chosen. External validity is related to generalizing, which is the degree to which the conclusions in a
study would hold for other persons in other places and at other times.
Population:
A population is a particular group of individuals or items. It can be any size or even infinite.
The population is the entire group of subjects the researcher wants information on.
Represents the complete set of individuals or items about which information is desired.
Can be large and potentially infinite.
Sample:
A sample is a subset of the population that represents the entire group.
A sample is a selection (subset) of data from a larger group of data, (called the population.) A sample should
be representative of the population, this means the sample and the population should have similar properties.
A sample is a smaller set of data that a researcher chooses or selects from a larger population using a pre-
defined selection bias method. These elements are known as sample points, sampling units, or observations.
A smaller, manageable subset of the population.
Selected to represent the characteristics of the entire population.
Used when studying the entire population is impractical, costly, or time-consuming.
Population vs Sample
Population Sample
The population includes all members of a A sample is a subset of the population.
specified group.
Collecting data from an entire population can Samples offer a more feasible approach to studying
be time-consuming, expensive, and populations, allowing researchers to draw
sometimes impractical or impossible. conclusions based on smaller, manageable datasets
Includes all residents in the city. Consists of 1000 households, a subset of the entire
population.
In statistics, a population is the entire group of individuals or items that are of interest in a study,
while a sample is a smaller, representative subset of that population.
Sampling Frame:
The sampling frame (also known as the “sample frame” or “survey frame”) is indeed the actual collection of
units. A sample has now been taken from this. A basic random sample gives all units in it an equal
probability of being drawn and appearing in the sample. In the ideal scenario, the sample frame should
match the sample of people.
Importance of sampling
Sampling is used in practice for a variety of reasons such as:
1. Sampling can save time and money. A sample study is usually less expensive than a census study and
produces results at a relatively faster speed.
2. Sampling may enable more accurate measurements for a sample study is generally conducted by
trained and experienced investigators.
3. Sampling remains the only way when population contains infinitely many members.
4. Sampling remains the only choice when a test involves the destruction of the item under study.
5. Sampling usually enables to estimate the sampling errors and, thus, assists in obtaining information
concerning some characteristic of the population.
Applications of sampling:
1. Sampling is widely used in various fields, including:
2. Market Research: Understanding consumer preferences, behaviors, and trends by surveying a
representative sample of the target audience.
3. Social Sciences: Studying social phenomena, opinions, and attitudes by sampling relevant
populations.
4. Medical Research: Conducting clinical trials and studies on a subset of patients to assess the
effectiveness and safety of treatments.
5. Quality Control: Inspecting a sample of products to ensure they meet quality standards.
6. Polling and Elections: Predicting election outcomes or gauging public opinion by surveying a sample
of voters.