0% found this document useful (0 votes)
2 views13 pages

Week # 2

The document provides an overview of key statistical concepts including data, elements, datasets, variables, and types of measurement scales. It explains the distinction between qualitative and quantitative variables, as well as independent and dependent variables, and discusses sampling methods like random and systematic sampling. Additionally, it emphasizes the importance of understanding levels of measurement for appropriate statistical analysis and the implications of sampling methods on research results.

Uploaded by

deuce R7
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views13 pages

Week # 2

The document provides an overview of key statistical concepts including data, elements, datasets, variables, and types of measurement scales. It explains the distinction between qualitative and quantitative variables, as well as independent and dependent variables, and discusses sampling methods like random and systematic sampling. Additionally, it emphasizes the importance of understanding levels of measurement for appropriate statistical analysis and the implications of sampling methods on research results.

Uploaded by

deuce R7
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Week # 2

1. Data:

Data is a collection of facts, numbers, words, measurements, or observations that we collect for
analysis.

Example: If we survey five students about their ages, the data we collect might be:​
Ages: 18, 20, 19, 21, 22

2. Element:

An element is a single object or individual from which we collect data.

Example: If we collect data from five students, each student is an element in our dataset.

●​ Student 1 (Age: 18)


●​ Student 2 (Age: 20)
●​ Student 3 (Age: 19)
●​ Student 4 (Age: 21)
●​ Student 5 (Age: 22)

Each student is an element because we collect data from them.

3. Dataset:

A dataset is a collection of data organized in a structured way, usually in tables, where each row
represents an element and each column represents a characteristic.

Example (Dataset of Students' Age and Scores):

Student Age Exam Score

1 18 85

2 20 90

3 19 78

4 21 88

5 22 92

●​ The dataset contains 5 elements (students)


●​ The dataset has 2 types of data (Age & Exam Score)
Summary:

●​ Data: Individual values (like 18, 20, 85, 90)


●​ Element: A single object from which data is collected (like each student)
●​ Dataset: A structured collection of data (like a table of students' ages and scores)

A collection of data values forms a data set. Each value in the data set is called a data value or
a datum.

Variable:

A variable is something that can change or have different values in a study or experiment.

Examples of Variables:

✔ Age of people (It changes from person to person)​


✔ Temperature of a city (It changes every day)​
✔ Number of students in a class (Different classes have different numbers)

🔹 A variable can be represented using letters like x, y, z etc.​


🔹 If something does not change, it is called a constant (like π = 3.1416).

a)​ population b) sample c) parameter d) statistics e) variable f) data


Dependent and Independent Variables:

In statistics and experiments, variables are classified into two main types:

1️ Independent Variable (Cause)

●​ The variable that we change or control to see its effect.


●​ It is the cause in a cause-and-effect relationship.

Example: In a study on how studying time affects exam scores:

Studying time is the independent variable (because we control it).

2️ Dependent Variable (Effect)

●​ The variable that we measure to see the effect of the independent variable.
●​ It is the effect in a cause-and-effect relationship.

Example: In the same study on exam scores:

Exam score is the dependent variable (because it depends on studying time).

Example #1: Can blueberries slow down aging? A study indicates that antioxidants found in
blueberries may slow down the process of aging. In this study, 19-month-old rats (equivalent to
60-year-old humans) were fed either their standard diet or a diet supplemented by either
blueberry, strawberry, or spinach powder. After eight weeks, the rats were given memory and
motor skills tests. Although all supplemented rats showed improvement, those supplemented
with blueberry powder showed the most notable improvement.

1. What is the independent variable?

2. What are the dependent variables?

Example #2: Does beta-carotene protect against cancer? Beta-carotene supplements have
been thought to protect against cancer. However, a study published in the Journal of the
National Cancer Institute suggests this is false. The study was conducted with 39,000 women
aged 45 and up. These women were randomly assigned to receive a beta-carotene supplement
or a placebo, and their health was studied over their lifetime. Cancer rates for women taking the
beta-carotene supplement did not differ systematically from the cancer rates of those women
taking the placebo.

1. What is the independent variable?

2. What is the dependent variable?


Qualitative and Quantitative Variables:

An important distinction between variables is between qualitative variables and quantitative


variables.

Qualitative variables are those that express a qualitative attribute such as hair color, eye color,
religion, favorite movie, gender, and so on. The values of a qualitative variable do not imply a
numerical ordering. Values of the variable “religion” differ qualitatively; no ordering of religions is
implied. Qualitative variables are sometimes referred to as categorical variables.

Quantitative variables (numerical) are those variables that are measured in terms of numbers.
Some examples of quantitative variables are height, weight, and shoe size.

Example: In the study on the effect of diet discussed previously, the independent variable was
the type of supplement: none, strawberry, blueberry, and spinach. The variable “type of
supplement” is a qualitative variable; there is nothing quantitative about it. In contrast, the
dependent variable “memory test” is a quantitative variable since memory performance was
measured on a quantitative scale (number correct)

Quantitative variables can be further classified into two groups: discrete and continuous.
Discrete variables can be assigned values such as 0, 1, 2, and 3 and are said to be countable.
Examples of discrete variables are the number of children in a family, the number of students in
a classroom, and the number of calls received by a call center each day for a month.

Discrete variables assume values that can be counted.

Continuous variables can assume an infinite number of values between any two specific
values. They are obtained by measuring. They often include fractions and decimals.
EXAMPLE 1–2 Discrete or Continuous Data Classify each variable as discrete or continuous.

a. The number of hours per day that children 6 to 12 years old reported that they played video
games

b. The number of runs a Major League player made each year of his career.

c. The amount of money drivers spend on gasoline each week

d. The weights of the players on a hockey team

Class Boundaries

●​ In statistics, continuous data is measured rather than counted.


●​ Because of measurement limitations, values are often rounded to the nearest unit.
●​ To represent the true range of a recorded value, we use class boundaries.

Why Do We Round Measurements?

●​ Measuring devices have limits in accuracy.


●​ Example:
○​ A thermometer may measure to the nearest 0.1°C.
○​ A weighing scale may measure to the nearest pound.
○​ A height measurement may be rounded to the nearest inch.

✔️ If rounding to the nearest 0.1 unit → Boundaries are ± 0.05.​


✔️ If rounding to the nearest 1 unit → Boundaries are ± 0.5.​
✔️ If rounding to the nearest 10 units → Boundaries are ± 5.
Example 1: Height Measurement

●​ Suppose a person’s recorded height is 73 inches.


●​ This means their actual height could be anywhere between 72.5 inches and 73.5 inches.
●​ Boundary notation: 72.5 – 73.5 inches.

Defining Class Boundaries

●​ The boundary of a number is the actual range of values before rounding.


●​ Boundaries are always written one decimal place more than the measured value.
●​ The boundaries always end in .5 to show the full range.

Example 2: Weight Measurement

●​ Suppose a person’s recorded weight is 86 pounds.


●​ The actual weight could be anywhere between 85.5 and 86.5 pounds.
●​ Boundary notation: 85.5 – 86.5 pounds.

Example 3: Temperature Measurement

●​ A thermometer records a temperature of 37°C.


●​ The actual temperature could be between 36.5°C and 37.5°C.
●​ Boundary notation: 36.5 – 37.5°C.

Example 4: Exam Scores Rounded to the Nearest 10

●​ If a student scores 90 marks, the actual score is between 85 and 95.


●​ Boundary notation: 85 – 95 marks.

🔹 Question 1: If a recorded value is 55 cm (rounded to the nearest cm), what are the class
🔹 Question 2: A student scored 75 marks (rounded to the nearest 5 marks). What are the
boundaries?​

🔹 Question 3: If a digital scale records weight as 120.0 kg (rounded to the nearest 0.1 kg),
boundaries?​

find the boundaries.


Level of Measurement/ Types of Scale:

Before analyzing data, we need to measure it properly. However, the way we measure
something depends on the type of data we have. Different types of information require different
ways of measurement.

For example:

●​ To measure reaction time, we use a stopwatch.


●​ To measure people’s opinions, we use a rating scale (e.g., "very favorable" to "not
favorable").
●​ To record a favorite color, we simply note the color name (e.g., "red" or "blue").

Since different kinds of data need different ways of measurement, we classify them into four
main types of scales:

1. Nominal Scale (Naming/Labeling)

Data that is measured using a nominal scale is qualitative(categorical). Categories, colors,


names, labels, and favorite foods along with yes or no responses are examples of nominal-level
data. Nominal scale data are not ordered.

For example, trying to classify people according to their favorite food does not make any sense.
Putting pizza first and sushi second is not meaningful.

Smartphone companies are another example of nominal scale data. The data are the names of
the companies that make smartphones, but there is no agreed-upon order of these brands,
even though people may have personal preferences. Nominal scale data cannot be used in
calculations.

2. Ordinal Scale (Ranking)

Data that is measured using an ordinal scale is similar to nominal scale data but there is a big
difference. The ordinal scale data can be ordered.

An example of ordinal scale data is a list of the top five national parks in the United States. The
top five national parks in the United States can be ranked from one to five but we cannot
measure differences between the data.

Another example of using the ordinal scale is a cruise survey where the responses to questions
about the cruise are “excellent,” “good,” “satisfactory,” and “unsatisfactory.” These responses
are ordered from the most desired response to the least desired. However, the differences
between the two pieces of data cannot be measured. Like the nominal scale data, ordinal scale
data cannot be used in calculations.
3. Interval Scale (No True Zero)

Data that is measured using the interval scale is similar to ordinal level data because it has a
definite ordering but there is a difference between data. The differences between interval scale
data can be measured though the data does not have a starting point.

Temperature scales like Celsius (C) and Fahrenheit (F) are measured by using the interval
scale. In both temperature measurements, 40° is equal to 100° minus 60°. Differences make
sense. But 0 degrees does not because, in both scales, 0 is not the absolute lowest
temperature. Temperatures like -10° F and -15° C exist and are colder than 0.

Interval-level data can be used in calculations, but one type of comparison cannot be done. 80°
C is not four times as hot as 20° C (nor is 80° F four times as hot as 20° F). There is no
meaning to the ratio of 80 to 20 (or four to one).

4. Ratio Scale (True Zero)

Data measured using the ratio scale solves the ratio problem and provides the most information.
Ratio scale data is like interval scale data, but it has a 0 point, and ratios can be calculated. For
example, four multiple-choice statistics final exam scores are 80, 68, 20, and 92 (out of a
possible 100 points).

The exams are machine-graded. The data can be put in order from lowest to highest: 20, 68,
80, 92.

The differences between the data have meaning. The score of 92 is more than the score of 68
by 24 points. Ratios can be calculated. The smallest score is 0. So 80 is four times 20. The
score of 80 is four times better than the score of 20.

We use levels of measurement because they help determine the appropriate statistical
techniques and graphical methods for analyzing data. Even though we might not directly use the
level names while calculating mean, variance, or creating graphs, understanding them is
essential for making correct decisions about data analysis.

Here’s why levels of measurement matter:

1.​ Choice of Statistical Measures​

○​ Nominal & Ordinal: Mean and variance don’t make sense. Instead, we use
mode or median.
○​ Interval & Ratio: Mean, variance, and standard deviation are meaningful.

2.​ Type of Graphs​

○​ Nominal & Ordinal: Bar charts or pie charts are suitable.


○​ Interval & Ratio: Histograms, line graphs, and scatter plots are more
appropriate.

3.​ Statistical Tests​

○​ Nominal: Chi-square test, mode, proportion comparisons.


○​ Ordinal: Median tests, Mann-Whitney U test.
○​ Interval & Ratio: T-tests, ANOVA, correlation, regression.

Without understanding levels of measurement, we might apply the wrong statistical method,
leading to incorrect conclusions.
Identify the following as nominal level, ordinal level, interval level, or ratio level data.

1. Flavors of frozen yogurt ________________

2. Amount of money in savings accounts________________

3. Students classified by their reading ability: Above average, Below average, Normal
________________

4. Letter grades on an English essay ________________

5. Religions ________________

6. Commuting times to work ____________

7. Ages (in years) of art students ________________

8. Ice cream flavor preference ________________

9. Years of important historical events ________________

10. Instructors classified as: Easy, Difficult or Impossible ________________

11. Social Security numbers ________________

12. Telephone numbers ________________

13. Years ending in a double zero ________________

14. Amount of money spent on pet care per year ________________

15. Debts of college students ________________

16. Ratings of high schools based on teachers’ salaries ________________

17. Number of CT scans an imaging center completes ________________

18. Horsepower of automobile engines ________________


Sampling Method

When researchers study a large group of people (called a population), they cannot ask every
single person because it takes too much time and effort. Instead, they select a sample, which is
a smaller group that represents the whole population. However, this sample should be chosen
carefully to avoid bias (wrong or misleading results).

For example, if you want to know people's opinions about a new school rule but only ask
students in one classroom, your results might not represent the whole school. To get fair results,
we use proper sampling methods.

There are various methods of sampling, but today we will focus on two of the most common:
Random Sampling and Systematic Sampling.

1. Random Sampling

Random sampling ensures that each member of the population has an equal chance of being
selected. This method helps eliminate bias and ensures that the sample accurately represents
the population.

Example

Imagine Lisa wants to form a study group with three other students from her pre-calculus class,
which has 31 members (excluding herself).

She can select her sample in two ways:

1.​ Traditional Method:​

○​ Write all 31 names on slips of paper.


○​ Put them in a hat, mix them, and pick three names randomly.
○​ This ensures each student has an equal chance of being selected.
2.​ Using Random Numbers (Technology-Based Method)​

○​ Lisa assigns each student a two-digit number (from 00 to 30).


○​ She uses a random number generator on a calculator to generate random
numbers.
○​ She reads the numbers in two-digit groups and selects three valid student IDs.
Suppose she generates the following numbers:​
0.94360; 0.99832; 0.14669; 0.51470; 0.40581; 0.73381; 0.04399

●​ The valid two-digit numbers she extracts are 14, 05, and 04.
●​ Referring to the class roster, these correspond to:
○​ 14 → Macierz
○​ 05 → Cuningham
○​ 04 → Cuarismo

Thus, Lisa’s study group will include herself, Macierz, Cuningham, and Cuarismo.

Advantages of Simple Random Sampling

✔ Fair and unbiased (everyone has an equal chance of selection).​


✔ Easy to understand and implement.

2. Systematic Sampling

In systematic sampling, the researcher selects every nth individual from the population after
choosing a random starting point.

Example of Systematic Sampling

Suppose a research team needs to conduct a phone survey with 400 participants from a city
phone directory containing 20,000 names.
Steps to conduct a systematic sample:

1.​ Assign each resident a number from 1 to 20,000.


2.​ Use random sampling to select a starting point (e.g., number 12).
3.​ Select every 50th name after the starting point:
○​ 12, 62, 112, 162, 212, ..., until reaching 400 names.
4.​ If the list ends before reaching 400 participants, loop back to the beginning.

Advantages of Systematic Sampling

✔ Faster and more convenient than random sampling.​


✔ Ensures even coverage across the population.​
✔ Easy to implement with an ordered list.

Summary

Both random sampling and systematic sampling are effective methods for obtaining
representative samples.

●​ Random Sampling is best when the population is small and diverse.


●​ Systematic Sampling is more efficient when dealing with large populations.

Using these methods correctly helps researchers gather reliable data while minimizing bias and
ensuring accurate conclusions.

Example: Use the following information to answer the next four exercises: A study was done to
determine the age, number of times per week, and duration (amount of time) of residents using
a local park in San Antonio, Texas. The first house in the neighborhood around the park was
selected randomly, and then the resident of every eighth house in the neighborhood around the
park was interviewed.

1. The sampling method was

a. simple random; b. systematic; c. stratified; d. cluster

2. “Duration (amount of time)” is what type of data?

a. qualitative(categorical); b. quantitative discrete; c. quantitative continuous

3. The colors of the houses around the park are what kind of data?

a. qualitative(categorical); b. quantitative discrete; c. quantitative continuous

4. The population is ______________________

You might also like