0% found this document useful (0 votes)
48 views4 pages

PGDM Hybrid 2023: Basic Statistics Assignment

This document contains 6 assignments related to analyzing data sets using descriptive statistics such as mean, median, mode, standard deviation, and frequency tables. The assignments include identifying the level of measurement for different types of survey data, creating a frequency table and bar graph to represent house room counts, drawing and analyzing box plots of temperature data from different cities, and calculating measures of central tendency and variability for waiting time, road accident, exam score, and factory production data. The final question involves comparing the standard deviations of production from two factory lines to determine which has more consistent output and proposing a strategy to improve stability.

Uploaded by

Leena Choudhary
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
48 views4 pages

PGDM Hybrid 2023: Basic Statistics Assignment

This document contains 6 assignments related to analyzing data sets using descriptive statistics such as mean, median, mode, standard deviation, and frequency tables. The assignments include identifying the level of measurement for different types of survey data, creating a frequency table and bar graph to represent house room counts, drawing and analyzing box plots of temperature data from different cities, and calculating measures of central tendency and variability for waiting time, road accident, exam score, and factory production data. The final question involves comparing the standard deviations of production from two factory lines to determine which has more consistent output and proposing a strategy to improve stability.

Uploaded by

Leena Choudhary
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

PGDM Hybrid 2023

Basic Statistics

Assignment-1

1) A marketing research team is conducting a survey to gather data from customers about their
preferences and experiences with a new product. The collected data falls into various categories.
Identify the type of data for each of the following scenarios: nominal, ordinal, interval, or ratio.

a) Customers are asked to rate their satisfaction with the product on a scale of 1 to 5, with
1 being "Very Dissatisfied" and 5 being "Very Satisfied."

b) The survey collects data on the number of products purchased by each customer in the
last month.

c) Customers are asked to indicate their age groups: 18-24, 25-34, 35-44, 45-54, 55+.

d) Customers are asked to rank three product features in order of importance: price,
durability, and aesthetics.

e) The survey asks customers to choose their preferred payment method: credit card, cash,
online payment, or mobile payment.

For each scenario, explain your reasoning behind categorizing the data as nominal, ordinal, interval,
or ratio.

2) The given list provides the count of rooms in 50 houses located in Goa:
2643344754
5375544562
6344586553
3375445416
5448623364
a) Create a frequency table and a bar graph to visually represent this data.
b) If a new real estate developer plans to construct an apartment complex with units
having an equal number of rooms, which count of rooms should they choose
according to this data? Provide an explanation.
3) The following four data sets give the daytime temperature in Mumbai, Bangalore, Hyderabad,
and Chennai for all the 28 days in February 2023.

Mumbai

31 30 30 30 30 29 31

30 31 29 29 30 31 29

29 30 28 29 29 29 29

28 29 27 29 28 29 29

Bangalore

20 29 22 25 20 19 28

24 22 23 25 26 21 22

20 16 23 26 22 18 21

20 19 24 24 22 18 20

Hyderabad

26 35 23 27 26 25 33

26 25 29 30 28 24 28

23 21 24 32 27 23 24

24 23 31 32 35 23 23

Chennai

28 28 29 29 30 27 30

25 24 25 24 29 26 28

29 31 23 26 29 31 26

27 26 27 25 29 37 25

Draw a box-plot for each data set and comment on the differences in the shape, spread, and
location of these box-plots.

4) Ten patients at a doctor’s surgery wait for varying lengths of time to see their doctor. The
waiting times, in minutes, are as follows:
5 mins, 17 mins, 8 mins, 2 mins, 55 mins, 9 mins, 22 mins, 11 mins, 16 mins, 5 mins.

a) Calculate the mean, median, and mode for the given waiting times. For each
calculation, show your steps clearly.
b) Considering the nature of the data and its distribution, discuss which measure of central
tendency (mean, median, or mode) would be most appropriate to represent the typical
waiting time in this scenario. Explain your reasoning.

5) For each dataset, calculate the first quartile (Q1), median (Q2), third quartile (Q3), and the
interquartile range (IQR). Provide clear steps for your calculations and show your final answers.

a) The data represents the time in minutes that twelve employees took to commute to
work on a particular day:
18, 34, 68, 22, 10, 92, 46, 52, 38, 29, 45, 37, 10, 50, 30, 70, 90.

b) The data provides the number of people killed in road traffic accidents in Delhi from
2013 to 2021:
1820, 1671, 1622, 1591, 1584, 1690, 1463, 1196.

c) The following dataset presents the final marks of 40 students for the Basic Statistics
course:
61, 77, 51, 85, 55, 77, 70, 56, 41, 61, 28, 87, 23, 22, 86, 63, 99, 94, 38, 25,
90, 59, 87, 53, 29, 86, 33, 87, 75, 50, 59, 77, 77, 71, 99, 78, 70, 93, 78, 93.

6) The management team of a manufacturing company is examining the production output of two
assembly lines, A and B, over the past 7 days. The number of units produced per day is recorded
for each assembly line. The team aims to understand the average production and variability to
make informed decisions.

Assembly Line A:

Day 1: 250 units

Day 2: 270 units

Day 3: 260 units

Day 4: 220 units

Day 5: 255 units

Day 6: 245 units

Day 7: 250 units

Assembly Line B:

Day 1: 230 units


Day 2: 240 units

Day 3: 255 units

Day 4: 245 units

Day 5: 235 units

Day 6: 250 units

Day 7: 240 units

a) Calculate the mean production and the standard deviation for both Assembly Line A and
Assembly Line B over the 7-day period.

b) Considering the company's focus on maintaining consistent production, analyze the


calculated standard deviations for both assembly lines. Propose a strategy or action that the
management team could consider to enhance production stability.

Common questions

Powered by AI

Examining IQR is crucial as it measures the variability within the middle 50% of a dataset, offering insights into the data's spread and central tendency unaffected by potential outliers. This is particularly helpful for skewed distributions where the mean is not representative .

Yes, by analyzing the mean and standard deviation, management can identify consistent performance and variability. Lower variability in production suggests stable processes. For instance, if Assembly Line A has a lower standard deviation, it indicates more consistent output, prompting management to investigate and replicate these conditions in Assembly Line B to stabilize its production .

Box-plot analysis shows differences in climate by revealing variations in the temperature distributions. For example, Mumbai and Chennai may have similar medians, indicating similar solar intensity patterns, but their interquartile ranges (spread) could differ, illustrating variance in daily temperature consistency. Differences in outliers between cities reveal extreme weather days .

Customer satisfaction ratings on a scale of 1 to 5 are considered ordinal data because the ratings signify a ranking or order (from 'Very Dissatisfied' to 'Very Satisfied') but the intervals between the numbers are not necessarily equal .

Age groups fall under ordinal data classification because they represent ordered categories (e.g., 18-24 to 55+) that signify ranks but not reflect precise age differences among them. Each group is distinct, and while you can compare them, mathematical operations between groups aren't meaningful .

The developer should choose 4 rooms for the new apartments as it is the most frequently appearing number in the dataset, making it the mode. Since the most common characteristic in the existing homes is 4 rooms, offering this number may align with customer preferences .

The mode is advantageous because it identifies the most common category, offering direct insights into the most frequent occurrence. However, its limitations include being less informative in datasets with uniform distribution or multiple modes and not providing an understanding of distribution spread or variation .

Frequency tables summarize data distribution facilitating pattern detection, while bar graphs visually highlight trends such as the most common number of rooms in houses. Together, they offer a clear, intuitive comprehension of data structure and differences, aiding in data-driven decision-making .

Strategies may include adopting uniform procedures across shifts, enhancing equipment maintenance schedules, and adjusting input material quality to reduce variability. Monitoring the standard deviation regularly helps assess the impact of these changes, aiming for a lower deviation is indicative of consistency .

The median is the most appropriate measure of central tendency for the waiting times since it is less influenced by outliers (such as the 55-minute wait) and provides a central value that better represents the typical waiting experience compared to the mean .

You might also like