0% found this document useful (0 votes)
8 views14 pages

Sample vs. Trimmed Mean Explained

The document discusses various statistical measures including sample mean, trimmed mean, median, and mode, explaining their definitions, uses, and effects of outliers. It provides examples of data analysis from two companies and demonstrates how curing temperature affects strength in materials. Additionally, it covers types of data such as quantitative, qualitative, organized, unorganized, and big data, along with their characteristics and applications.

Uploaded by

fatimaazka352
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
8 views14 pages

Sample vs. Trimmed Mean Explained

The document discusses various statistical measures including sample mean, trimmed mean, median, and mode, explaining their definitions, uses, and effects of outliers. It provides examples of data analysis from two companies and demonstrates how curing temperature affects strength in materials. Additionally, it covers types of data such as quantitative, qualitative, organized, unorganized, and big data, along with their characteristics and applications.

Uploaded by

fatimaazka352
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

EXERCISES

Book: Probability & Statistics


1.1
Part (f) Question:

Is the sample mean more or less descriptive as a center of location than the trimmed mean?

✅ First, understand the two terms:

1. Sample Mean:
o This is the normal average (add all values and divide by how many values).
o It uses all values, even the very big or very small ones (called outliers).
2. Trimmed Mean (20%):
o This is also an average but after removing some smallest and largest values (in
this case, 3 from each side).
o It ignores extreme values, so it gives a more stable or balanced center.

What happened in your data:

• Your mean = 3.79


• Your trimmed mean = 3.68
• The mean is a little higher because of big numbers like 5.6 and 5.2
• These big values pull the mean up

✅ Final Answer in Easy Words:

The trimmed mean is more descriptive because it removes extreme values and gives a better
idea of the "typical" drying time.
So, in this case:

✔ The trimmed mean is better than the sample mean as a center of location
1.2
(d) Evidence of Outliers Using Mean, Median, and Trimmed Mean

• Mean = 20.77, Median = 20.61, Trimmed Mean = 20.74


• All three values are very close to each other.
• ✅ Conclusion: No strong evidence of outliers because there’s no big difference
between the three.

1.3
(d) Medians

• No Aging (median) = 221.5


• Aging (median) = 210.0

✅ Both mean and median show a similar drop.


✔ This supports the conclusion that aging reduces strength.

1.4
Data:

Company A:
9.3, 8.8, 6.8, 8.7, 8.5, 6.7, 8.0, 6.5, 9.2, 7.0

Company B:
11.0, 9.8, 9.9, 10.2, 10.1, 9.7, 11.0, 11.1, 10.2, 9.6

🔹 (a) Calculate Mean and Median


Company A:

Step 1: Add all values

9.3+8.8+6.8+8.7+8.5+6.7+8.0+6.5+9.2+7.0=79.59.3 + 8.8 + 6.8 + 8.7 + 8.5 + 6.7 + 8.0 + 6.5 + 9.2 + 7.0 =
79.59.3+8.8+6.8+8.7+8.5+6.7+8.0+6.5+9.2+7.0=79.5

Step 2: Divide by 10

Mean (A)=79.510=7.95\text{Mean (A)} = \frac{79.5}{10} = 7.95Mean (A)=1079.5=7.95

Step 3: Sort values to find Median

Sorted:
6.5, 6.7, 6.8, 7.0, 8.0, 8.5, 8.7, 8.8, 9.2, 9.3

Middle two values (5th and 6th) = 8.0 and 8.5

Median (A)=8.0+8.52=8.25\text{Median (A)} = \frac{8.0 + 8.5}{2} = 8.25Median (A)=28.0+8.5=8.25

Company B:

Step 1: Add all values

11.0+9.8+9.9+10.2+10.1+9.7+11.0+11.1+10.2+9.6=102.611.0 + 9.8 + 9.9 + 10.2 + 10.1 + 9.7 + 11.0 + 11.1


+ 10.2 + 9.6 = 102.611.0+9.8+9.9+10.2+10.1+9.7+11.0+11.1+10.2+9.6=102.6

Step 2: Divide by 10

Mean (B)=102.610=10.26\text{Mean (B)} = \frac{102.6}{10} = 10.26Mean (B)=10102.6=10.26

Step 3: Sort values to find Median

Sorted:
9.6, 9.7, 9.8, 9.9, 10.1, 10.2, 10.2, 11.0, 11.0, 11.1

Middle two values (5th and 6th) = 10.1 and 10.2

Median (B)=10.1+10.22=10.15\text{Median (B)} = \frac{10.1 + 10.2}{2} = 10.15Median (B)=210.1+10.2


=10.15
✅ Final Answers:
Company Mean Median

A 7.95 8.25

B 10.26 10.15

🔹 (b) Plot and Impression

If you draw a dot plot or line chart:

• Company A's values are lower and more spread out.


• Company B's values are higher and more tightly grouped around 10–11.

Impression:

✔ Company B’s steel rods are more flexible on average.


✔ Company A has more variation and lower flexibility.

1.6
🔹 (c) Does curing temperature affect strength?

✅ Yes!
When the temperature increases from 20°C to 45°C, the average tensile strength increases
from 2.11 to 2.32 MPa.

So, higher temperature = stronger rubber.

🔹 (d) Anything else influenced?

Yes — the variation/spread of the values is also wider at 45°C.

You can see:

• 20°C values are mostly around 2.0 to 2.2


• 45°C values range from 1.99 to 2.52, so they are more spread out
This means:
✔ Tensile strength increases
❗ But also becomes a bit more variable (less consistent)

✅ 1. Mean (Average)

What is it?

The mean is the total of all values divided by how many values there are.

Interpretation:

• It tells us the overall average of the data.


• Affected by extreme values (outliers).
• If one number is too big or too small, the mean can move up or down.

✅ Use it when:

• You want a summary of the whole group.


• Data is evenly spread (no extreme outliers).

✅ 2. Median (Middle Value)

What is it?

The middle value when the numbers are arranged in order.


(If even numbers, take average of two middle values.)

Interpretation:

• It shows the center of the data.


• Not affected by outliers (like one very large or small number).
• Half of the data is below it, and half is above it.

✅ Use it when:

• There are extreme values or skewed data.


• You want to find the true center.
✅ 3. Mode (Most Frequent Value)

What is it?

The number that appears the most in the dataset.

Interpretation:

• It tells what is common or typical in the group.


• There can be no mode, one mode, or multiple modes.

✅ Use it when:

• You're interested in what happens most often (e.g., most sold size, most repeated score).
• Useful for categorical or discrete data.

🎯 Summary Table
Measure What it tells you Affected by Outliers? Best Use

Mean The overall average ✅ Yes Balanced data

Median The middle value ❌ No Skewed or uneven data

Mode The most frequent value ❌ No Most common item or repeated value
1.22
(c) Bell-Shaped?

• Yes, the data is fairly symmetrical


• Most values are near the mean (6.77), and frequency drops gradually on both sides.
• So, yes, there’s an indication that this is approximately bell-shaped.

1.23
(c) Interpretation

• 1980 emissions are higher and more spread out.


• 1990 emissions are lower and more consistent.
• So yes, emissions clearly dropped, and variability also reduced, suggesting cars got
cleaner and more regulated.

1.25
Interpretation:

• Mean (34.41) is higher than median (27.4) due to some very large values (like 89.2).
• Trimmed mean (30.65) is closer to median, showing how trimming reduces influence
of extreme values.

THEORY
✅ 1. Quantitative Data

Definition: Data that can be measured and written in numbers.

Examples:

• Age = 25 years
• Temperature = 37°C
• Height = 5.6 ft
• Marks = 80

It answers “How much?”, “How many?”, or “What number?”

✅ 2. Qualitative Data

Definition: Data that describes qualities, categories, or characteristics; usually non-


numerical.
Examples:

• Gender (Male/Female)
• Eye Color (Brown, Blue)
• Religion (Islam, Christianity)
• Nationality (Pakistani, Indian)

It answers “What kind?”, “Which type?”, or “What category?”

✅ 3. Organized Data

Definition: Data that is arranged clearly in a table, graph, or chart, so it's easy to read and
analyze.

Examples:

• A marksheet arranged by student names


• Sales data in Excel
• Data shown in pie charts or bar graphs

✅ 4. Unorganized Data

Definition: Data that is scattered, raw, or jumbled, without any structure.

Examples:

• Randomly listed survey responses


• A list of numbers without order
• Handwritten responses without sorting

✅ 5. Big Data

Definition: Very large and complex data that is difficult to process using traditional
methods.

Examples:

• Facebook user data


• Google search trends
• Online shopping history of millions of users
Used in AI, machine learning, cloud computing, etc.

✅ 6. Numerical Data

Definition: A type of quantitative data that specifically deals with numbers.

Examples:

• Number of students = 50
• Income = Rs. 30,000
• Speed = 90 km/h

Note: All numerical data is quantitative, but not all quantitative data is just numbers (it may
involve units, ratios, etc.).

Summary Table:

Type Numerical? Describes... Example

Quantitative ✅ Yes Amount or number Height = 160 cm

Qualitative ❌ No Category or label Eye Color = Blue

Organized Can be Properly arranged Marks in a table

Unorganized Can be Messy or raw Random responses

Big Data ✅ Mostly Massive, complex data Data from 1M users

Numerical ✅ Yes Numbers only Age = 20, Salary = 50,000

Same but more


✅ 1. Quantitative Data (Also called Numerical Data)

Definition:
Quantitative data refers to data that can be measured or counted and expressed in numbers. It
deals with quantities, meaning how much, how many, or how often. This data can be
mathematically analyzed, averaged, or used for statistical operations.

Types of Quantitative Data:

• Discrete: Countable (e.g. number of students)


• Continuous: Measurable (e.g. height, weight, temperature)

Examples:

• A student scored 92 marks in Math.


• The temperature today is 38°C.
• A bag weighs 5.5 kg.
• A shop sells 27 pens in a day.

Use:

Useful for creating graphs like bar charts, histograms, and calculating mean, median, standard
deviation, etc.

✅ 2. Qualitative Data (Also called Categorical Data)

Definition:

Qualitative data describes the qualities, categories, or attributes of a thing. It is non-numeric,


and cannot be measured in numbers. Instead, it classifies or labels data based on characteristics.

Types of Qualitative Data:

• Nominal: No natural order (e.g. gender, religion)


• Ordinal: With an order (e.g. grades: A, B, C)

Examples:

• Hair color: Black, Brown, Blonde


• Gender: Male, Female
• Blood group: A+, B-, O+
• Feedback: Excellent, Good, Poor

Use:
Used in surveys, interviews, and opinions. Cannot be averaged but can be used to calculate mode
or shown in pie charts.

✅ 3. Organized Data

Definition:

Organized data is data that has been arranged in a meaningful structure for easy reading and
interpretation. It is presented in a systematic form like tables, charts, frequency distribution,
or graphs.

Examples:

• A table showing student names with their marks.


• A pie chart showing sales distribution.
• A bar graph of population in different cities.

Use:

Organized data helps in quick analysis, decision-making, and reporting. It is suitable for
statistical and graphical representation.

✅ 4. Unorganized Data

Definition:

Unorganized data (also called raw data) is data in its original, unarranged, and unprocessed
form. It lacks a proper structure, and it's hard to read or analyze directly.

Examples:

• Survey responses written randomly on paper.


• A list of numbers like: 45, 87, 34, 29, 99, 16.
• Data collected from different sources but not cleaned.

Use:

Before using this type of data, it must be cleaned, sorted, and organized. Data scientists and
researchers often start with unorganized data.
✅ 5. Big Data

Definition:

Big Data refers to extremely large and complex datasets that cannot be easily managed,
processed, or analyzed using traditional tools like Excel. It includes data generated from social
media, sensors, websites, mobile devices, and more.

Characteristics of Big Data (5 Vs):

1. Volume – Huge amount of data


2. Velocity – Data comes in fast (live streams)
3. Variety – Different formats (text, images, videos)
4. Veracity – Data accuracy and trustworthiness
5. Value – Useful insights that can be taken from the data

Examples:

• Facebook data (likes, comments, shares)


• Google search data from billions of users
• YouTube viewing history

Use:

Big Data is used in Artificial Intelligence, Machine Learning, Business Analytics, and
forecasting.

✅ 6. Numerical Data (Subset of Quantitative Data)

Definition:

Numerical data is data that consists of numbers and can be operated on using arithmetic
operations. It is mostly the same as quantitative data, but more focused on its mathematical
nature.

Examples:

• Income: Rs. 50,000


• Number of books: 12
• Speed: 60 km/h
• Temperature: 23.5°C

Use:

Ideal for all statistical calculations like mean, median, mode, variance, standard deviation,
correlation, regression, etc.

📘 Summary Table:
Type of Data Description Example

Countable or measurable values in


Quantitative Age = 20, Marks = 90
numbers

Qualitative Descriptive or categorical info Gender = Male, Color = Blue

Organized Structured and arranged clearly Tables, Graphs, Frequency chart

Unorganized Raw, messy, unprocessed Random survey responses

Big Data Huge, fast, and complex data Twitter data, Website logs

Number-based data (subset of


Numerical Height = 5.6ft, Income = 30K
quantitative)

You might also like