0% found this document useful (0 votes)
3 views4 pages

SPSS Data Transformation Exercises

The document outlines a series of exercises related to data transformation and analysis using SPSS, including calculating probabilities for student demographics, evaluating a screening test for Down syndrome, and constructing various data visualizations. It also includes instructions for recoding data into different formats and categories. Each exercise is assigned a point value and requires the application of statistical methods and interpretation of results.

Uploaded by

herdsnerds
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views4 pages

SPSS Data Transformation Exercises

The document outlines a series of exercises related to data transformation and analysis using SPSS, including calculating probabilities for student demographics, evaluating a screening test for Down syndrome, and constructing various data visualizations. It also includes instructions for recoding data into different formats and categories. Each exercise is assigned a point value and requires the application of statistical methods and interpretation of results.

Uploaded by

herdsnerds
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

‭Assignment 2‬

‭SPSS-TRANSFORMING DATA‬

‭ xercise 1:‬‭(2 points)‬


E
‭The number of students in the MPH program by county is‬
‭ PH program‬
M
‭given below. If a student is selected at random from this‬
‭dataset, find the probability that:‬ ‭Number of‬
‭a.‬ ‭The student is from Escambia county‬‭0.212‬ ‭ ounty‬
C ‭students‬ ‭Percentage‬
‭b.‬ ‭The student is from Santa Rosa county‬‭0.30‬ ‭Escambia‬ ‭106‬ ‭21.2‬
‭c.‬ ‭The student is from either Escambia or Santa Rosa‬ ‭Walton‬ ‭151‬ ‭30.2‬
‭county‬‭0.512‬ ‭Santa Rosa‬ ‭150‬ ‭30‬
‭d.‬ ‭The student is not from Santa Rosa county‬‭0.70‬ ‭Okaloosa‬ ‭93‬ ‭18.6‬
‭Total‬ ‭500‬ ‭100‬

‭ xercise 2:‬‭(2 points)‬


E
‭A crosstabulation of the number of students in the MPH program by county and whether the students owned‬
‭a computer when they started the program is given below. If the student is selected at random from this data‬
‭set, determine the probability of selecting a student who is from Santa Rosa County‬‭or‬‭a student who had‬‭a‬
‭computer at the start of the program. (Show your work)‬‭0.774‬
‭ xercise 3:‬‭(2 points)‬
E
‭Investigators‬‭conducted‬‭a‬‭study‬‭to‬‭evaluate‬‭the‬‭use‬‭of‬‭a‬‭screening‬‭test‬‭of‬‭specific‬‭hormones‬‭in‬‭the‬‭blood‬‭to‬
‭assess‬ ‭whether‬ ‭or‬ ‭not‬ ‭the‬ ‭fetus‬ ‭of‬ ‭a‬ ‭pregnant‬ ‭women‬ ‭is‬ ‭likely‬ ‭to‬ ‭have‬ ‭Down‬ ‭syndrome. ‬ ‭10000‬‭pregnant‬
‭women‬ ‭underwent‬ ‭the‬ ‭screening‬ ‭test‬ ‭and‬ ‭scored‬ ‭either‬ ‭positive‬ ‭or‬ ‭negative‬ ‭depending‬ ‭on‬ ‭the‬ ‭levels‬ ‭of‬
‭hormones‬‭in‬‭the‬‭blood.‬‭There‬‭were‬‭1255‬‭women‬‭with‬‭positive‬‭test;‬‭from‬‭them,‬‭285‬‭women‬‭had‬‭an‬‭affected‬
‭fetus.‬ ‭There‬ ‭were‬ ‭9700‬ ‭unaffected‬ ‭fetus;‬ ‭8730‬‭of‬‭these‬‭had‬‭negative‬‭test‬‭results.‬‭What‬‭was‬‭the‬‭sensitivity,‬
‭specificity,‬‭positive‬‭predictive‬‭value‬‭and‬‭negative‬‭predictive‬‭value‬‭of‬‭the‬‭test?‬ ‭Interpret‬‭the‬‭results‬‭(Show‬
‭your work including developing a 2 X 2 table).‬

‭ ffected‬
a f‭ etus NOT‬ ‭total‬
‭fetus‬ ‭affected‬

‭ ositive‬
p ‭285‬ ‭970‬ ‭1,255‬
‭screening‬

‭ egative‬
n ‭15‬ ‭8,730‬ ‭8,745‬
‭screening‬

‭total‬ ‭300‬ ‭9,700‬ ‭10,000‬

‭Exercise 4:‬‭(2 points)‬

‭ ‬ I‭ QR = Q3-Q1 = 39.00-2.25 =‬‭36.75‬



‭●‬ ‭Fence‬‭Lower‬‭=‬ ‭-52.875‬
‭Q1 - (1.5)(IQR) = 2.25 - (1.5)(36,75) = 2.25 - 55.125 = -52.875‬
‭●‬ ‭are there outside values below the lower fence?‬‭yes‬
‭●‬ ‭Fence‬‭Upper‬‭=‬ ‭94.125‬
‭Q3 + (1.5)(IQR) = 39.00 + (1.5)(36.75) = 39.00 + 55.125 =‬
‭94.125‬
‭●‬ ‭are there outside values above the upper fence?‬‭yes‬
‭lower fence is the lowest number in the dataset‬
‭upper fence is the highest number in the dataset‬
‭.‬

‭Exercise 5:‬‭(2 points)‬


‭●‬ C
‭ onstruct a steam-and-leaf plot of the data‬
‭Seizures following bacterial meningitis.‬
‭●‬ ‭Construct a boxplot of the data‬‭Seizures following bacterial meningitis.‬

‭‬ A
● ‭ re there any outside values in this data set?‬‭yes‬
‭●‬ ‭Does the boxplot show evidence of asymmetry?‬‭no, the data set is skewed‬

‭ xercise 6 (optional):‬
E
‭(1 point extra for each output – 1, 2, 3)‬
‭This week we are going to learn how to recode data into the same variable and how to record data into a‬
‭different variables. This is a very helpful tool in SPSS.‬

‭1.‬ R
‭ ecoding data into single values‬
‭The first data set represents runs scored by 5 baseball players in a national tournament. We want to‬
‭recode this data so that the players are rank ordered by their number of runs, with the player with the‬
‭highest runs given a code of “1” and the player with the lowest score given a 5.‬
‭a.‬ ‭Enter the following data in SPSS: # of runs by players‬
‭b.‬ ‭Recode the data so that the players are rank ordered by their number of runs, with the player‬
‭with the highest runs given a code of "1" and the batsman with the lowest runs given a "5".‬
‭c.‬ ‭Run frequencies of the new created variable.‬
‭2.‬ R
‭ ecoding data into a given range of values‬
‭The second data set represents the scores of 10 MPH students in their final biostatistics exam. We‬
‭want to recode the data giving a code of 1 to scores between 75 - 100, code 2 to scores between 61 -‬
‭75, code 3 to scores between 41 - 60 and code 4 to scores between 0 - 40.‬
‭a.‬ ‭Enter the following data in SPSS‬
‭b.‬ ‭Recode the data giving code "1" to scores between 75 - 100, code 2 to scores between 61 -‬
‭75, code 3 to scores between 41 - 60 and code 4 to scores between 0 – 40‬
‭c.‬ ‭Run frequencies of the new created variable.‬

‭3.‬ R
‭ ecoding data into two categories‬
‭The‬ ‭third‬ ‭dataset‬ ‭represents‬ ‭12‬ ‭patient‬ ‭satisfaction‬ ‭scores‬ ‭for‬ ‭a‬ ‭dental‬ ‭provider.‬ ‭The‬ ‭satisfaction‬
‭scores‬ ‭goes‬ ‭from‬ ‭1‬ ‭to‬ ‭10‬ ‭with‬ ‭10‬ ‭being‬ ‭extremely‬ ‭satisfied‬ ‭and‬ ‭1‬ ‭extremely‬ ‭dissatisfied.‬ ‭The‬
‭provider‬‭wants‬‭to‬‭code‬‭all‬‭those‬‭who‬‭responded‬‭by‬‭giving‬‭ratings‬‭above‬‭5‬‭a‬‭"Satisfactory"‬‭code‬‭and‬
‭those below 5 a "Dissatisfactory" code.‬
‭a.‬ ‭Enter the data in SPSS‬
‭b.‬ ‭Recode‬ ‭the‬ ‭data‬ ‭so‬ ‭that‬ ‭all‬ ‭those‬ ‭who‬ ‭responded‬ ‭by‬ ‭giving‬ ‭ratings‬ ‭above‬ ‭5‬ ‭have‬ ‭a‬
‭"Satisfactory" code and those below 5 have a "Dissatisfactory" code‬
‭c.‬ ‭Run frequencies of the new created variable‬

Common questions

Powered by AI

Sensitivity is calculated as the proportion of true positives identified by the test, which is 285/300 = 0.95. Specificity is the proportion of true negatives correctly identified, which is 8730/9700 = 0.90. The positive predictive value is the probability that subjects with a positive screening test truly have the condition, calculated as 285/1255 = 0.227. The negative predictive value is the probability that subjects with a negative test truly don’t have the condition, calculated as 8730/8745 = 0.998 .

Recoding data into a range of values simplifies the interpretation and analysis of data by categorizing continuous numerical scores into ordinal data. In the case of MPH students' scores, recoding is done by assigning codes (1 to 4) based on predefined score brackets. This changes the nature of data from interval to ordinal, which can influence statistical tests suitable for analysis. It can improve clarity when analyzing trends or differences between groups but may reduce informational content inherent in original, more granular data .

To calculate this probability, you use the formula for the union of two probabilities: P(A or B) = P(A) + P(B) - P(A and B). From the data provided: P(A) (student from Santa Rosa County) = 0.30, P(B) (student had a computer) is not directly provided but can be deduced, and P(A and B) is also deduced based on given overlap probabilities. The result of these calculations yields a probability of 0.774 .

The positive predictive value (PPV) indicates the probability that subjects with a positive test truly have Down syndrome; a PPV of 0.227 means that 22.7% of positive results are true positives. The negative predictive value (NPV) is the probability that subjects with a negative test truly don’t have Down syndrome; an NPV of 0.998 suggests that 99.8% of negative results are true negatives. These values are crucial for evaluating test reliability: high NPV and low PPV may indicate that the test is better at confirming the absence rather than the presence of the condition .

The interquartile range (IQR) is calculated by subtracting the first quartile (Q1) from the third quartile (Q3). In the example, IQR = Q3 - Q1, giving 39.00 - 2.25 = 36.75. The IQR represents the spread of the middle 50% of the dataset, indicating variability and helping identify potential outliers. A large IQR may indicate data spread over a wide range, while a small IQR suggests that data points are close together .

A stem-and-leaf plot organizes data values into stems (representing ranges of values) and leaves (representing individual data points) to display their distribution. For seizures following bacterial meningitis, each y-value (seizure count) is split into a stem (all digits except the final one) and a leaf (final digit). This plot allows easy visualization of data distribution, showing shape, centrality, and variability while retaining actual data points, facilitating a quick yet detailed overview .

Boxplot analysis reveals the distribution of data through quartiles and highlights potential outliers. If central lines (median) aren’t centered between quartiles or if whiskers are unequal, asymmetry might exist. In this context, no evidence of asymmetry implies a symmetrical distribution, but the presence of outside values denotes outliers. This can suggest variability or exceptional cases within the dataset .

Recoding patient satisfaction scores simplifies data interpretation by transforming continuous data into categorical data. For example, coding scores above 5 as 'Satisfactory' and scores below 5 as 'Dissatisfactory' consolidates responses into manageable groups, aiding comparative analysis, and trend identification without losing clarity. This method helps decision-making by providing clear metrics for service improvements and efficacy, while allowing targeted strategies for enhancing patient satisfaction .

Data visualizations like stem-and-leaf plots and boxplots provide detailed insights into data distribution, central tendency, variance, and potential outliers. Stem-and-leaf plots preserve raw data and enable quick visual assessments of frequency and distribution. Boxplots offer a succinct view of data quartiles, skewness, and extreme values. Together, these tools allow for both detailed and summary views, facilitating comprehensive data analysis and understanding of complex datasets .

A specificity of 0.90 indicates that 90% of truly unaffected fetuses are correctly identified by the test as negative. This high specificity suggests the test is reliable in minimizing false positives, reducing unnecessary anxiety and further testing for many women. However, the remaining 10% may still yield false positives, indicating a trade-off between detecting all true negatives and limiting false positive results .

You might also like