1/12/24, 12:15 PM Assignment_College.
R - Colaboratory
keyboard_arrow_down Assignment
keyboard_arrow_down [Link]
Use the [Link]() function to read the data into R. Call the loaded data college. Make sure that you have the
set to the correct location for the data.
df <- [Link]('/content/[Link]')
rownames(df) <- df[, 1]
df <- df[, -1]
str(df)
'[Link]': 777 obs. of 18 variables:
$ Private : chr "Yes" "Yes" "Yes" "Yes" ...
$ Apps : int 1660 2186 1428 417 193 587 353 1899 1038 582 ...
$ Accept : int 1232 1924 1097 349 146 479 340 1720 839 498 ...
$ Enroll : int 721 512 336 137 55 158 103 489 227 172 ...
$ Top10perc : int 23 16 22 60 16 38 17 37 30 21 ...
$ Top25perc : int 52 29 50 89 44 62 45 68 63 44 ...
$ [Link]: int 2885 2683 1036 510 249 678 416 1594 973 799 ...
$ [Link]: int 537 1227 99 63 869 41 230 32 306 78 ...
$ Outstate : int 7440 12280 11250 12960 7560 13500 13290 13868 15595 10468 ...
$ [Link] : int 3300 6450 3750 5450 4120 3335 5720 4826 4400 3380 ...
$ Books : int 450 750 400 450 800 500 500 450 300 660 ...
$ Personal : int 2200 1500 1165 875 1500 675 1500 850 500 1800 ...
$ PhD : int 70 29 53 92 76 67 90 89 79 40 ...
$ Terminal : int 78 30 66 97 72 73 93 100 84 41 ...
$ [Link] : num 18.1 12.2 12.9 7.7 11.9 9.4 11.5 13.7 11.3 11.5 ...
$ [Link]: int 12 16 30 37 2 11 26 37 23 15 ...
$ Expend : int 7041 10527 8735 19016 10922 9727 8861 11487 11644 8991 ...
$ [Link] : int 60 56 54 59 15 55 63 73 80 52 ...
keyboard_arrow_down Rename columns because of alphanumeric characters
# Rename column names
colnames(df)[colnames(df) %in% c("[Link]", "[Link]", "[Link]", "[Link]", "[Link]", "[Link]" )] <- c("S_F_Ratio"
# display to check the column names
head(df)
A [Link]: 6 × 18
Private Apps Accept Enroll Top10perc Top25perc [Link] [Link] Outstate [Link] Books Personal P
<chr> <int> <int> <int> <int> <int> <int> <int> <int> <int> <int> <int> <in
Abilene
Christian Yes 1660 1232 721 23 52 2885 537 7440 3300 450 2200
University
Adelphi
Yes 2186 1924 512 16 29 2683 1227 12280 6450 750 1500
University
Adrian
Yes 1428 1097 336 22 50 1036 99 11250 3750 400 1165
College
Agnes
Scott Yes 417 349 137 60 89 510 63 12960 5450 450 875
College
Alaska
Pacific Yes 193 146 55 16 44 249 869 7560 4120 800 1500
University
Albertson
Yes 587 479 158 38 62 678 41 13500 3335 500 675
College
# Visualising distribution private vs non-private colleges using bar plot
barplot(table(df$Private),
ylab = "Frequency",
xlab = "Private", col = c("green", "orange"))
[Link] 1/6
1/12/24, 12:15 PM Assignment_College.R - Colaboratory
keyboard_arrow_down [Link]
View the datasets and summarize the data-frame using summary() function. Provide interpretations from the
of the function.
summary(df)
Private Apps Accept Enroll
Length:777 Min. : 81 Min. : 72 Min. : 35
Class :character 1st Qu.: 776 1st Qu.: 604 1st Qu.: 242
Mode :character Median : 1558 Median : 1110 Median : 434
Mean : 3002 Mean : 2019 Mean : 780
3rd Qu.: 3624 3rd Qu.: 2424 3rd Qu.: 902
Max. :48094 Max. :26330 Max. :6392
Top10perc Top25perc S_F_Ratio Grad_Rate
Min. : 1.00 Min. : 9.0 Min. : 139 Min. : 1.0
1st Qu.:15.00 1st Qu.: 41.0 1st Qu.: 992 1st Qu.: 95.0
Median :23.00 Median : 54.0 Median : 1707 Median : 353.0
Mean :27.56 Mean : 55.8 Mean : 3700 Mean : 855.3
3rd Qu.:35.00 3rd Qu.: 69.0 3rd Qu.: 4005 3rd Qu.: 967.0
Max. :96.00 Max. :100.0 Max. :31643 Max. :21836.0
Outstate perc_alumni Books Personal
Min. : 2340 Min. :1780 Min. : 96.0 Min. : 250
1st Qu.: 7320 1st Qu.:3597 1st Qu.: 470.0 1st Qu.: 850
Median : 9990 Median :4200 Median : 500.0 Median :1200
Mean :10441 Mean :4358 Mean : 549.4 Mean :1341
3rd Qu.:12925 3rd Qu.:5050 3rd Qu.: 600.0 3rd Qu.:1700
Max. :21700 Max. :8124 Max. :2340.0 Max. :6800
PhD Terminal Room_Board F_Undergrad
Min. : 8.00 Min. : 24.0 Min. : 2.50 Min. : 0.00
1st Qu.: 62.00 1st Qu.: 71.0 1st Qu.:11.50 1st Qu.:13.00
Median : 75.00 Median : 82.0 Median :13.60 Median :21.00
Mean : 72.66 Mean : 79.7 Mean :14.09 Mean :22.74
3rd Qu.: 85.00 3rd Qu.: 92.0 3rd Qu.:16.50 3rd Qu.:31.00
Max. :103.00 Max. :100.0 Max. :39.80 Max. :64.00
Expend P_Undergrad
Min. : 3186 Min. : 10.00
1st Qu.: 6751 1st Qu.: 53.00
Median : 8377 Median : 65.00
Mean : 9660 Mean : 65.46
3rd Qu.:10830 3rd Qu.: 78.00
Max. :56233 Max. :118.00
Observations/interpretations
The range of total applications is from 81(min) to 48094(max)
The highest number of enrollments are from for Texas A&M Univ i.e 6392
It looks like there aren't any missing or null values, so we are good to continue
Private column can be converted to factor variable to check the count of private and non-private schools
I think the maxvalue of graduation rate 118 is an outlier, since it is impossible to have a graduation rate that is above 100%
The mean and median values are quite similar for columns like P_undergrad,F_undergrad, Room_Board, Terminal etc, they might form a
normal distribution.
Aside from some outlier schools with very high costs, there isn't a wide gap for the median non-tution costs between private schools and
non-private schools.
[Link] 2/6
1/12/24, 12:15 PM Assignment_College.R - Colaboratory
keyboard_arrow_down [Link]
Produce a scatterplot matrix of the first ten columns or variables of the data. Provide interpretations from the
of those plots.
# change Private to a factor variable
df$Private <- [Link](df$Private)
pairs(df[,1:10], panel=[Link], main="scatter plots of all pairs of variables")
Observations
As for enrollment and number of applications accepted we can see that there is a positive relationship between the two. We can also see
that - Positive correlation between Accept and Enroll: As the number of applicants accepted by a college increases, the number of enrolled
students also tends to increase.
We can observe that as the outstate tuition increases, the number of students enrolling in the university is pretty low.
There is a positive relationship between Private schools and perc_alumini as there are a higher percent of alumni who donate
We observe a negative relation between SF ratio and Graduation rate. In other words, as the ratio increases, graduation rates tend to
decrease.
keyboard_arrow_down 4. Use the plot() function to produce side-by-side boxplots of ‘Outstate’ versus ‘Private’.
plot(df$Private, df$Outstate, xlab = "Private University", ylab = "Tuition in $", col = c("red","green"))
Observations
The above boxplot for Private universities provides insights into the distribution of out-of-state students
[Link] 3/6
1/12/24, 12:15 PM Assignment_College.R - Colaboratory
The boxplot indicates that the median for Private universities is higher and the interquartile range is larger, it suggests that, on average,
Private universities have more out-of-state students compared to Non-private universities.
We can also observe the outliers for non-private universities in the plot
Boxplots of Outstate versus Private: Private universities have more out of state students as compared to non-private universities
keyboard_arrow_down 5:quantitative
Use the hist() function to produce some histograms with differing numbers of bins for a few of the
variables.
par(mfrow = c(3, 3))
#Applications, Applications accepted, students enrolled
hist(df$Apps, xlab = "Number of applicants", main = "Histogram for all colleges", col="red", breaks = 20)
hist(df$Accept, xlab = "Number of applicants", main="Histogram for applications accepted",col="red",breaks=20)
hist(df$Enroll, xlab = "Number of students", main="Histogram for new students enrolled",col="red")
# Applications by Private or non private
hist(df$Apps[df$Private == "Yes"], xlab = "Number of applicants", main = "Histogram for applicants in private schools", col="green", bre
hist(df$Apps[df$Private == "No"], xlab = "Number of applicants", main = "Histogram for applicants in non-private schools", col="green")
keyboard_arrow_down Observations
The histograms above for total applications, Applications accepted, students enrolled, applications for private and non-private are right
skewed/ positive skewed distributions
# Expend
par(mfrow = c(2, 2))
hist(df$Expend, xlab = "Instructional expenditure per student (dollars)", main = "Histogram for expenditure in all colleges", col="red",
hist(df$Expend[df$Private == "Yes"], xlab = "Instructional expenditure per student (dollars)", main = "Histogram for expenditure in priv
hist(df$Expend[df$Private == "No"], xlab = "Instructional expenditure per student (dollars)", main = "Histogram for expenditure in non-p
[Link] 4/6
1/12/24, 12:15 PM Assignment_College.R - Colaboratory
keyboard_arrow_down Observations
Kept the same scle on x-axis for all the histograms above to be able to compare well.
The histograms above for expenditure for private and non-private are right skewed/ positive skewed distributions.
# PHD, Books, Persona, Outstate and Alumini donate
par(mfrow = c(3, 3))
hist(df$PhD,main="Histogram for Percent of faculties with PhD",col="red",breaks=10)
hist(df$perc_alumni,main="Histogram for Percent of alumini donate",col="green")
hist(df$Books,main="Histogram for College Books",col="blue",breaks=100)
hist(df$Personal,main="Histogram for Personal",col="orange")
hist(df$S_F_Ratio, col=3, breaks=100, xlab = "Student/faculty ratio", main = "Histogram for Student/faculty ratio")
hist(df$Outstate, main = "Histogram for Outstate")
Observations
Kept the same scale on x-axis for all the histograms above to be able to compare well.
Except for PHD i.e left skewed, rest of the histograms above for are right skewed/ positive skewed distributions.
6: Find out the name of the university with the most students in the top 10% of class. Tips: you should create a
keyboard_arrow_down new variable that divides universities into two groups based on whether or not the proportion of students
coming from the top 10% of their high school classes exceeds 50%.
[Link] 5/6
1/12/24, 12:15 PM Assignment_College.R - Colaboratory
# Create a new variable to divide into two groups, taking above 50% as 'YES' while others as 'NO'
Super <- rep("No", nrow(df))
Super[df$Top10perc > 50] <- "Yes"
Super <- [Link](Super)
df <- [Link](df, Super)
# Picking
# Checkingthe
theuniversity
summary with most students requires them to fall in the Super category(where top 10perc is above>50% and should have hig
summary(df$Super)
df[df$Super == 'Yes' & df$Top10perc > 95,]
No: 699 Yes: 78
Private Apps Accept Enroll Top10perc Top25perc S_F_Ratio Grad_R
<fct> <int> <int> <int> <int> <int> <int> <i
Massachusetts
Institute of Yes 6411 2140 1078 96 99 4481
Technology
Observation:
Massachusetts Institute of Technology is the university with most students in top 10 % class
keyboard_arrow_down 7. Find out the name of the university that has the smallest acceptance rate.
# Calculate acceptance rate by dividing Applications accepted by total applications per college
acceptance_rate <- df$Accept / df$Apps
# University that has the smallest acceptance rate
[Link](df)[[Link](acceptance_rate)]
'Princeton University'
Observation:
'Princeton University' has the smallest acceptance rate
keyboard_arrow_down 8. Find out the name of the university that has the highest acceptance rate.
# University that has the highest acceptance rate
[Link](df)[[Link](acceptance_rate)]
'Emporia State University'
[Link] 6/6