0% found this document useful (0 votes)
23 views3 pages

Boston College Employee Data Analysis

This document contains instructions for a managerial statistics problem set analyzing employee data from Dunnen Dockery's Boston office. It asks students to [1] access an SPSS data file containing information on 474 Boston employees, [2] analyze variables like job category, salary, experience to understand the employee population, and [3] consider questions about using the sample to make inferences about all Dunnen Dockery employees nationwide. It also contains a second data set on crime rates in US states for additional analysis questions.

Uploaded by

prem
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
23 views3 pages

Boston College Employee Data Analysis

This document contains instructions for a managerial statistics problem set analyzing employee data from Dunnen Dockery's Boston office. It asks students to [1] access an SPSS data file containing information on 474 Boston employees, [2] analyze variables like job category, salary, experience to understand the employee population, and [3] consider questions about using the sample to make inferences about all Dunnen Dockery employees nationwide. It also contains a second data set on crime rates in US states for additional analysis questions.

Uploaded by

prem
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Managerial Statistics Problem Set

Populations and Distributions


Boston College

Problems

Part I.
Dunnen Dockery is a building maintenance service company that opened a Boston office over a
decade ago. They have hired you to analyze the company employee files for the Boston office
so they understand their current Boston office employees better. They have provided you with a
data file that includes information on every employee at the Boston office. The file measures
each of the following for each employee:
 Employee ID (1 for the first employee hired in the Boston office, 2 for the second, etc.)
 Gender
 Birthdate
 Education (years of school completed)
 Job Category (type of job held at the office)
 Salary (current annual salary)
 Beginning salary (salary upon date of hire)
 Job time (months since date of hire)
 Previous experience (total months of full-time work experience before hire)
 Minority status (1 if employee identifies as a minority, 0 otherwise)

Access the data file named “Employee data” with SPSS. Use it to answer the questions that
follow.

1. How many members are there in the population under consideration?


474

2. Do you consider the people included in the file to be representative of the population? Why?
No

3. Suppose you wanted to gather further data on the employees as part of a study of employee
satisfaction. But you cannot get the needed data from every employee. So instead you put the
names of all the managers in a hat and put out 30 of them at random. Would this give you a
representative sample?

4. Suppose you wanted a number that indicates how familiar each person was with the working
world upon entering the company. What data item might be a good measure of this?
5. In the previous example, what do we call something like “how familiar each person was with
the workaday world upon entering the company”, in other words, some property of a member of
the population that we are interested in?

6. In the previous example, what do we call the data item we use to measure the property?
Why do we call it that?

7. Ignore “id” for this question. Of the 9 random variables in the file, which are nominal, which
are ordinal, which are scale, and which are binary? (Note: in the data file, the “Measure” of
each variable in the file has been switched to “Nominal” to disguise the variable type for each
variable.

8. Of the 9 random variables in the file, which are discrete and which are continuous?

9. Create a histogram for jobcat (employment category). Is this a distribution of the variable?
If not, could it be made into one?

10. Create a histogram for salary. Is this a distribution of the variable? If not, could it be made
into one?

11. Create a histogram for job time (Months since hire) with exactly 6 bins. Based on the
results, does this variable appear to have roughly any particular distribution? Which one?

Part II.

Dunnen Dockery now wants you to use the data from the Boston office employees to reach
conclusions about all of its employees across the country.

12. Would this qualify as the use of a representative sample?

13. Might this sample be biased in some way? If so, how?


Part III.

The FBI has complied statistics on the 50 U.S. states and their crime rates. Variables
included are:
 City ID
 Crime Rate
 Age (Males aged 14-24 per 1000)
 Per Capita Expenditure on Police
 Labor (Labor Force per 1000 males)
 Males (Males per 1000 females)
 Population (State pop in 100,000s)
 Unemployment (Unemployed males per 1000)
 Income (Families per 1000 below average income)

Download the data file named “Crime Rate data” and open it in SPSS. Complete the following
tasks and fill in your responses below. Once you’ve finished, you may want to save your output
page to look back in the future.

14. Create a histogram for Population (State Population in thousands). Based on the results,
does this variable appear to have any particular distribution? Which one?
Sample distribution

15. Create a histogram for Males (males per 1000 females). Based on the results, does this
variable appear to have any particular distribution? Which one?
Normal Distribution

16. Create a histogram for the random variable Unemployment (Unempl. of males per 1000 of
population). Assume the country can be divided into ten “regions”, each with 5 states in it. Now
imagine that you could calculate the average of Unemployment in each of the ten regions and
plot the 10 resulting data points on a histogram. How would this look compared with the
histogram you just created, with the 50 unemployment levels, one for each state?

You might also like