0% found this document useful (0 votes)
7 views8 pages

Introduction to Statistics Concepts

Uploaded by

adam.ibr1234
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views8 pages

Introduction to Statistics Concepts

Uploaded by

adam.ibr1234
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Chapter 1 - Introduction

This is a course in statistics, so let us start by finding out what is statistics?

Statistics is the science of how to collect, organize, present, analyse, and interpret data. (where
data is the information that is collected)

‘Statistics may be defined as “a body of methods for making wise decisions in the face of
uncertainty.”’ by W.A. Wallis.
In this chapter we are going to first introduce some basic terminology used in statistics and
then discuss ways of collecting data, mainly through sampling

Population Versus Sample:


A population consists of all elements – individuals, items, or objects – whose characteristics are
be studied. The population that is being studied is also called the target population.
Example of population.
All students registered at AUB for Summer Session 2020

What defines a population is the research problem or question we are working with Population
in statistics are not necessarily defined by geographic boundaries
A survey that includes every member of the population is called a census.

A Sample is a portion (or subset) of the larger population.
Example of Sample:
Five hundred students picked from the list of the AUB students
Every 5th student who enters to AUB through the Main Gate

The aim is to study that portion (the sample) to gain information about the population. The
technique of collecting information from a portion of the population is called a sample survey.
One of the main concerns in the field of statistics is how well the sample represents the
population. A sample that represents the characteristics of the population as closely as possible
is called a representative sample. The reason for that is to be able to use the sample data and
generalize to the population
Why Sample ?

 Reduced cost
 Faster results
 Increase precision

What is the problem with taking the students registered In this class as a sample of all AUB
students ?

 Identify each of the following data sets as either a population or a sample:


a. The grade point averages (GPAs) of all students at a college.

b. The GPAs of a randomly selected group of students on a college campus.

c. The ages of the 128 members of the Lebanese Parliament on January1,2023.

d. The gender of every second customer who enters a movie theater.

Parameters vs Statistics
Parameters are fixed values describing the population based on data collected for all units in
the population
Example: In a study of all cloned sheep it is found that their mean age is 2.7 years.
2.7 years in this example is a parameter since all cloned sheep were studied

Statistics are values or measures computed from sample data. They are a value describing the
sample that is picked from the population
Example: In studying a random sample of 12 Canada geese it was found that their mean
number of offspring is 3.8.
3.8 in this example is a statistic since it is the average of 12 geese that were studied and not the
whole population of Canada geese.

As a memory help:
Parameters start with “P” population starts with “P”
Statistics start with “S”, sample starts with “S”
Types of Statistics
To ensure that our sample statistic is a good estimate for the true population parameter, we
need to make sure that we obtain a representative sample – a sample in which the
characteristics of the individuals closely match the characteristics of the overall population.

Distinguish between a parameter and a statistic:


Which of the following is a parameter which is a sample?

a. 43% of teachers in a high school are females


b. Of 50 birds that were selected the mean width of their wing is 20.3 cm.
c. The average age of students at a university is 20.8 years

Types of Statistics
 Descriptive Statistics: consists of methods for organizing, displaying, and describing data
by using tables, graphs, and summary measures, to look for patterns in a data set, to
summarize the information and to present the data in a clear and meaningful way.
 The attempts to count phenomena such as people, animals, crops, and land and make
decisions based on these numbers existed since ancient civilization.

 There are 3 common types of descriptive statistics:
 Frequency distributions (tables) – they help us understand how the values are
distributed
 Graphical Representations – helps us visualize the data
 Numerical summary measures- these are numerical values calculated based on the
data which gives us an idea of the location and spread of the data ( for example the
average )

We will spend the next few lectures learning the above criteria

Inferential Statistics: consists of methods that generalize from samples to populations.


Use sample results to help make decisions or predictions about a population.
(Estimation, Testing Hypothesis, Predictions and relations between variables)
We will look at inferential statistics in the second half of the course
The whole idea behind inferential statistics is that we want to obtain some values about
the population but it is not feasible to study the whole population so we choose a
sample, collect data on the sample and use this data to draw conclusions about the
general population

Exercise: Identify whether the statement describes inferential statistics or descriptive


statistics:
a. The average age of the students in STAT 201 is 19 years

b. Based on data from a previous study of 1000 students at AUB, it is predicted that
82% of the students will vote during the student council election.

c. A research report issued by the Centers for Disease Control and Prevention
states that American adults are nearly 25 pounds heavier now than they were in
1960. Source: [Link]

d. In a survey of all postsecondary institutions, 66% of institutions offered distance


education classes in 2006.

Basic Terms

An element or member of a sample or population is a specific subject or object (for example, a


person, firm, item, state, or country) about which the information is collected
The value of a variable for an element is called an observation or measurement.

A data set is a collection of observations on one or more variables.
A variable is a characteristic under study that assumes different values for different elements.
In contrast to a variable, the value of a constant is fixed.

Exercise 1.5
The following table gives the number of dog bites reported to the police in six cities last year

City Number of Bites


Center City 52
Elm Grove 32
Franklin 43
Bay City 44
Oakdale 12
Sand Point 2
Briefly show which is a member, a variable, a measurement, and a data set with reference to
the table above.
With reference to this table, we have the following definitions:
Member: Each city included in the table
Variable: Number of dog bites reported last year
Measurement: Number of dog bites in a specific city (for example 52 in Center City )
Data set: Collection of dog bite numbers for the six cities listed in the table.

Types Of variables :

A variable is the recorded data about the unit we are studying. It assumes different values for
different units.

The first thing to do when you start learning statistics is get acquainted with the data types that
are used, such as numerical and categorical variables.
Different types of variables require different types of statistical analysis and visualization
approaches. Therefore, it is important that you understand how to classify the data you are
working with.

Qualitative or Categorical Variables


Are variables that describe data that fits into categories: characteristics or attributes.
Categories are non- measurable. Qualitative data are also often called categorical data. The
values are usually letters or words. For example
Hair color( brown, black , blond , Red, grey )
blood type( A , B, AB, O )
the car a person drives, and the street a person lives on.

Exception: Sometimes a variable that is reported as a number is considered as a Qualitative


variable for Example
Student number although is a number (for example 2023xxxx) but it is an
identification of the student
Car License Plates,
Telephone numbers
Quantitative Variables:
are always numbers. the result of counting or measuring attributes of a population

Discrete Variables
 A variable whose values are countable is called a discrete variable. In other words,
a discrete variable can assume only certain values with no intermediate values.
Example : the number of eggs that hens lay

Continuous Variables
A variable that can assume any numerical value over a certain interval or intervals
is called a continuous variable.
Example : the amount of milk that a cow yields in a week

Identify the following measures as either quantitative or qualitative:

a. The gender of the first 40 newborns in a hospital one year.

b. The hobbies of 150 randomly selected teenagers

c. The ages of 20 randomly selected fashion models.

d. The distance in kilometers travelled per 20 liters of fuel of 30 new cars purchased last month.

e. The last four digits of social security numbers of all students in a class
f. The number of credits taken by a random sample of 125 students during this semester

g. The numbers on the jerseys of 15 football players on a team.

A researcher wishes to estimate the average weight of newborns in Lebanon in the last five years.
He takes a random sample of 235 newborns and obtains an average of 3.27 kilograms

 the population of interest:


All newborn babies in Lebanon in the last five years.
 The sample :
The 235 babies that were selected
 the variable of interest:
average weight of newborns in the last five years
 Is the average obtained a parameter or a statistic?
Statistic since it was collected form the sample of 235 babies
 Type of the Variable:
Quantitative continuous
Overview of statistical procedure

Population
Collection of all
Units of interest

Analyze the data ,


draw conclusions Select a random
and generalize to Sample
the population Representative
Inferential sample
Statistics

Organize and
summarize the
data Collect Data on
the sample
Descriptive
Statistics

Summary of the statistical procedure


The aim of statistics is to draw conclusions and make predictions on the population
based on the data collected from the sample. In order to do that the sample must be
representative of the population from which it was drawn and that is why we use probability
samples.

Common questions

Powered by AI

Using a statistic, which is derived from a sample, instead of a parameter, which encompasses the entire population, implies that conclusions are based on estimations and are subject to sampling error . This approach is fundamental when it is impractical or impossible to collect data from the entire population. The choice between using a statistic or a parameter is determined by the feasibility of studying the whole population; if a census is possible, parameters are used, otherwise statistics are relied upon .

A representative sample is crucial in ensuring that the characteristics of the sample closely match the characteristics of the entire population, which allows for accurate and reliable statistical inferences . This representation is vital for minimizing bias and error in conclusions drawn about the population. Factors that might jeopardize this representation include sampling bias, such as when the sample is not randomly selected, or if some groups within the population are underrepresented . Failure to accurately represent the population can lead to incorrect generalizations.

Data collection via sampling involves selecting a subset of the population to analyze and is generally more cost-effective, faster, and can provide increased precision in data collection when done properly . Meanwhile, a census involves collecting data from every member of the population, providing complete data with no sampling error but at a much higher cost and longer duration. The drawbacks of sampling include potential sampling bias and error, whereas the ultimate benefit of a census is obtaining definitive insights without inferential estimations .

Descriptive statistics involves methods for organizing, displaying, and describing data using tables, graphs, and summary measures like averages or medians to summarize information effectively . For example, describing the average grades of a class involves descriptive statistics. Inferential statistics, on the other hand, allows for generalizations from a sample to a population, using techniques like hypothesis testing and estimation to make predictions about a population based on sample data . Predicting overall student performance at a university based on a sample of student grades is an example of inferential statistics.

Discrete variables assume countable numerical values, such as the number of students in a class, while continuous variables can assume any value within a range, such as height or weight . The correct classification affects the choice of statistical tests and analysis methods. If a continuous variable is mistaken for discrete, it might lead to using inappropriate statistical methods, thus affecting the validity of research outcomes by potentially misinterpreting the variable's behavior and relationships .

Data visualization plays a crucial role in descriptive statistics as it transforms complex data sets into graphical representations, making it easier to identify trends, patterns, and outliers . Graphs and charts help communicate findings succinctly and effectively, facilitating understanding and facilitating better decision-making. This step is critical because it provides an intuitive overview of data distribution and relationships, often revealing insights not immediately obvious through numerical summaries alone .

Frequency distributions summarize data by indicating how often each different value in a set of data occurs, which helps in visualizing the data's dispersion and identifying patterns or trends . For example, in analyzing student performance, a frequency distribution might be used to display the number of students achieving each letter grade within a class, allowing educators to quickly identify common performance levels and better understand the overall class achievement.

This definition highlights the comprehensive role of statistics in managing data's lifecycle, emphasizing its importance in ensuring data-driven decision-making in academic research . Collecting data efficiently requires understanding sampling methods; organizing and presenting data involve using descriptive statistics; and analyzing and interpreting data rely on inferential statistics to draw meaningful conclusions from research. These stages are interdependent, and their effective execution is essential for producing credible and generalizable research findings .

Categorical variables describe data fitting into categories, typically expressed in words or labels, such as 'gender' or 'blood type.' They do not measure a quantity but rather represent a quality . Quantitative variables, however, involve numerical data, where numbers express a count or measure, such as 'age' or 'distance.' Misclassification might occur if numbers are treated as identifiers instead of measurements, such as student ID numbers, which are numerical but serve as labels (qualitative) and do not measure a characteristic .

Ensuring a sample is representative requires careful sampling design, such as using random sampling techniques, to include all relevant subgroups of the population proportionally . In educational research, considering diversity in demographics, including socio-economic status, gender, and academic performance bands, is crucial to prevent bias and enhance the applicability and reliability of study outcomes. Additionally, the sample size should be large enough to capture the heterogeneity of the population .

You might also like