0% found this document useful (0 votes)
12 views27 pages

Midwest Colleges with High Default Rates

This document is a notebook focused on using R for data exploration to analyze student debt in colleges using the US Department of Education's College Scorecard Database. It provides a step-by-step introduction to basic R commands, data manipulation techniques, and how to filter and analyze college data. The goal is to determine which colleges help their graduates succeed financially and which do not.

Uploaded by

papanamaburner
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
12 views27 pages

Midwest Colleges with High Default Rates

This document is a notebook focused on using R for data exploration to analyze student debt in colleges using the US Department of Education's College Scorecard Database. It provides a step-by-step introduction to basic R commands, data manipulation techniques, and how to filter and analyze college data. The goal is to determine which colleges help their graduates succeed financially and which do not.

Uploaded by

papanamaburner
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

Data Science Project: Use data to determine the best


and worst colleges for conquering student debt.
Notebook 1: Basic R Commands & Data Exploration
Does college pay off? We'll use some of the latest data from the US Department of
Education's College Scorecard Database to answer that question. In this first notebook,
you'll get a gentle introduction to R - a coding language used by data scientists to
analyze large datasets. Then, you'll begin diving into the college scorecard data yourself.
By the end of this notebook, you'll get a general sense of which colleges set up their
graduates for success and which colleges ... don't.
In [9]: ## Run this code but do not edit it. Hit Ctrl+Enter to run the code
# This command downloads a useful package of R commands
library(coursekata)

── CourseKata packages ──────────────────────────────────── coursekata 0.18.


0 ──
✔ dslabs 0.7.6 ✔ Metrics 0.1.4
✔ Lock5withR 1.2.2 ✔ lsr 0.5.2
✔ fivethirtyeightdata 0.1.0 ✔ mosaic [Link]
✔ fivethirtyeight 0.6.2 ✔ supernova 2.5.7

1.0 - Exploring the dataset


To begin, let's download our data. Our full dataset is included in a file named
[Link] , which we're retrieving from the [Link] website. The
command below downloads the data from the file and stores it into an R dataframe
object called dat .
In [10]: ## Run this code but do not edit it. Hit Ctrl+Enter to run the code
# This command downloads data and stores it in the object `dat`
dat <- [Link]('[Link]

[Link] 1_ Basic R %… 1/27


5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

The <- operator is used to store values. For example, x<-10 stores the value of 10 in
x , meaning the value 10 is saved in the object x .

To get a quick view of the dataframe ( dat ), we can use the head command to print
out its first several rows.
In [11]: ## Run this code but do not edit it. Hit Ctrl+Enter to run the code
# This command prints out the first several rows of the dataset
head(dat)

OPEID name city state region median_debt default_rate highes


<int> <chr> <chr> <chr> <chr> <dbl> <dbl>
Alabama A
1 100200 &M Normal AL South 15.250 12.1
University
University
2 105200 of Alabama Birmingham AL South 15.085 4.8
at
Birmingham
3 2503400 Amridge Montgomery AL South 10.984 12.9
University
University
4 105500 of Alabama Huntsville AL South 14.000 4.7
in
Huntsville
Alabama
5 100500 State Montgomery AL South 17.500 12.8
University
The
6 105100 University Tuscaloosa AL South 17.671 4.0
of Alabama
The vertical columns of the dataframe are called variables , and their elements are
called values . For example, the variable city has values Normal , Birmingham ,
Montgomery , Huntsville , etc.

The horizontal rows of the dataframe are called observations . For example, the first
observation is Alabama A & M University , which is located in AL (Alabama), in
the city of Normal , and has a median student debt of $15,250 . For this dataframe,
each observation describes a specific college.
1.1 - Of the variables displayed, identify one that is quantitative, one that is
categorical, and one that is a unique identifier.

[Link] 1_ Basic R %… 2/27


5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

Double-click to type a response: The median debt would be a quantitative variable.


The region would be a categorical variable. The OPEID would be the unique identifier
variable.
The head command only displays several rows of the dataframe. To see the full
dimensions of the dataframe, we can use the dim command.
1.2 - Use the dim command on dat to display the dimensions of the dataframe.
In [12]: # Your code goes here
dim(dat)

4435 · 26
Check yourself: Your code should have printed out two numbers: 4435 and 26.
The first number outputted by dim is the number of horizontal rows in the dataframe.
This represents the number of observations (number of colleges). The second number is
the number of vertical columns in the dataframe. This represents the number of
variables. What are all these variables? See the description of the dataset below, along
with links to descriptions of all the variables.

The Dataset
General description - The US Department of Education's College Scorecard
Database shows various metrics of cost, enrollment, size, student debt, student
demographics, and alumni success. It describes almost every University, college,
community college, trade school, and certificate program in the United States. The
data is current as of the 2020-2021 school year.
Description of all variables: See here
Detailed data file description: See here
With such a large dataset, to make your life easier, you may want to work with only a few
variables at a time. In the following code, we use the select command to select only
the variables name , median_debt , ownership , admit_rate , and hbcu and
save them in a new dataframe called example_dat .
In [13]: ## Run this code but do not edit it
# Select certain columns from dat, store into example_dat
example_dat <- select(dat, name, median_debt, ownership, admit_rate, hbcu)

[Link] 1_ Basic R %… 3/27


5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

# Display head of example_dat


head(example_dat)

A [Link]: 6 × 5
name median_debt ownership admit_rate hbcu
<chr> <dbl> <chr> <dbl> <chr>
1 Alabama A & M University 15.250 Public 89.65 Yes
2 University of Alabama at 15.085 Public 80.60 No
Birmingham
3 Amridge University 10.984 Private NA No
nonprofit
4 University of Alabama in Huntsville 14.000 Public 77.11 No
5 Alabama State University 17.500 Public 98.88 Yes
6 The University of Alabama 17.671 Public 80.39 No

1.3 - Use the command to select the variables name , region ,


select
default_rate , ownership , and pct_PELL from dat . Store your new
dataframe in an object called my_dat and display its head.
In [14]: # Your code goes here

my_dat <- select(dat, name, region, default_rate, ownership, pct_PELL)


head(my_dat)

A [Link]: 6 × 5
name region default_rate ownership pct_PELL
<chr> <chr> <dbl> <chr> <dbl>
1 Alabama A & M University South 12.1 Public 70.95
2 University of Alabama at South 4.8 Public 33.97
Birmingham
3 Amridge University South 12.9 Private 74.52
nonprofit
4 University of Alabama in Huntsville South 4.7 Public 24.03
5 Alabama State University South 12.8 Public 73.68
6 The University of Alabama South 4.0 Public 17.18
In addition to filtering out columns (variables), we can also filter out rows (observations).
For example, if I only wanted to analyze colleges that are HBCUs and that have an
admissions rate below than 40%, I can use the subset command on example_dat
like this:

[Link] 1_ Basic R %… 4/27


5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

In [15]: ## Run this code but do not edit it


# Subset example_dat to only HBCUs with admissions rates lower than 40%
subset(example_dat, hbcu == "Yes" & admit_rate < 40)

A [Link]: 7 × 5
name median_debt ownership admit_rate hbcu
<chr> <dbl> <chr> <dbl> <chr>
461 Delaware State University 18.264 Public 39.34 Yes
473 Howard University 19.500 Private 38.64 Yes
nonprofit
491 Florida Agricultural and 18.750 Public 32.98 Yes
Mechanical University
503 Florida Memorial University 17.155 Private 38.41 Yes
nonprofit
1376 Alcorn State University 16.895 Public 37.72 Yes
1401 Rust College 11.226 Private 29.47 Yes
nonprofit
2747 Hampton University 18.500 Private 36.00 Yes
nonprofit
A total of 7 colleges fit these conditions.
Note that R has different conventions for comparative statements. For example...
== means equals exactly
!= means does not equal
< means less than
> means greater than
<= means less than or equal to
>= means greater than or equal to

Here are some other common conditional symbols


| means or
& means and

1.4 - Use the subset command to find the colleges in my_dat that are located in
the Midwest region of the United States and have more than a third of their
students (greater than 33%) default on their loans.
In [16]: # Your code goes here
subset(my_dat, region=="Midwest" & default_rate>33)

[Link] 1_ Basic R %… 5/27


5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

A [Link]: 2 × 5
name region default_rate ownership pct_PELL
<chr> <chr> <dbl> <chr> <dbl>
815 West Michigan College of Midwest 34.4 Private for- 42.47
Barbering and Beauty profit
4382 Kenny's Academy of Barbering Midwest 44.4 Private for- 72.33
profit

Check yourself: You should find that 2 schools match your selection criteria.

1.5 - What do you notice about the observations that fit your selection criteria? What
do you wonder?
Double-click to type a response:
Suppose you're interested in a particular college, such as Howard University. We can use
the subset command to filter the example_dat dataframe and focus solely on the
information pertaining to that college.
In [17]: ## Run this code but do not edit it
# Subset example_dat to only show Howard University
subset(example_dat, name == "Howard University")

A [Link]: 1 × 5
name median_debt ownership admit_rate hbcu
<chr> <dbl> <chr> <dbl> <chr>
473 Howard University 19.5 Private nonprofit 38.64 Yes

1.6 - Select a college that interests you. Then use the subset command to locate
and extract information about the college from my_dat . Note: The exact spelling of
the names of all the colleges in the dataset can be found here.
In [18]: # Your code goes here
subset(my_dat, name=="Amridge University")

A [Link]: 1 × 5
name region default_rate ownership pct_PELL
<chr> <chr> <dbl> <chr> <dbl>
3 Amridge University South 12.9 Private nonprofit 74.52

[Link] 1_ Basic R %… 6/27


5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

One further way to explore a dataset is to reorder its observations. For example, we can
use the arrange command to order the colleges in example_dat by their admission
rate:
In [14]: ## Run this code but do not edit it
# Arrange data in order of their admission rates
arrange(example_dat, admit_rate)

[Link] 1_ Basic R %… 7/27


5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

A [Link]: 4435 × 5
name median_debt ownership admit_rate hbcu
<chr> <dbl> <chr> <dbl> <chr>
Curtis Institute of Music 16.250 Private 2.44 No
nonprofit
Harvard University 12.072 Private 5.01 No
nonprofit
Stanford University 11.000 Private 5.19 No
nonprofit
Princeton University 10.355 Private 5.63 No
nonprofit
Yale University 12.000 Private 6.53 No
nonprofit
Columbia University in the City of New 19.250 Private 6.66 No
York nonprofit
California Institute of Technology 9.867 Private 6.69 No
nonprofit
Massachusetts Institute of Technology 12.000 Private 7.26 No
nonprofit
University of Chicago 13.000 Private 7.31 No
nonprofit
The Juilliard School 25.000 Private 7.64 No
nonprofit
Brown University 12.000 Private 7.67 No
nonprofit
Duke University 12.500 Private 7.74 No
nonprofit
Pomona College 10.000 Private 8.62 No
nonprofit
University of Pennsylvania 14.000 Private 8.98 No
nonprofit
Swarthmore College 14.000 Private 9.06 No
nonprofit
Bowdoin College 14.000 Private 9.16 No
nonprofit
Dartmouth College 14.500 Private 9.22 No
nonprofit
Northwestern University 14.000 Private 9.31 No
nonprofit
Colby College 17.500 Private 10.27 No
nonprofit

[Link] 1_ Basic R %… 8/27


5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

name median_debt ownership admit_rate hbcu


<chr> <dbl> <chr> <dbl> <chr>
Cornell University 13.108 Private 10.71 No
nonprofit
Rice University 10.500 Private 10.89 No
nonprofit
Johns Hopkins University 11.750 Private 11.06 No
nonprofit
Tulane University of Louisiana 19.000 Private 11.11 No
nonprofit
Vanderbilt University 12.420 Private 11.62 No
nonprofit
Amherst College 12.000 Private 11.83 No
nonprofit
Circle in the Square Theatre School 16.000 Private 12.12 No
nonprofit
Claremont McKenna College 12.070 Private 13.34 No
nonprofit
Colorado College 15.045 Private 13.60 No
nonprofit
Barnard College 16.250 Private 13.60 No
nonprofit
Bates College 12.610 Private 14.10 No
nonprofit
⋮ ⋮ ⋮ ⋮ ⋮

National Personal Training Institute- 6.333 Private for- NA No


Tampa profit
Mobile Technical Training 3.800 Private for- NA No
profit
California Institute of Arts & Technology 9.500 Private for- NA No
profit
Elite Cosmetology Barber & Spa 6.054 Private for- NA No
Academy profit
Gwinnett Institute 9.500 Private for- NA No
profit
Manuel and Theresa's School of Hair 6.494 Private for- NA No
Design profit
Peloton College 9.500 Private for- NA No
profit
Ross Medical Education Center - 8.089 Private for- NA No
Kalamazoo profit
[Link] 1_ Basic R %… 9/27
5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

name median_debt ownership admit_rate hbcu


<chr> <dbl> <chr> <dbl> <chr>
Ross College-Canton 8.347 Private for- NA No
profit
Ross College-Grand Rapids North 7.125 Private for- NA No
profit
American Institute-Somerset 9.176 Private for- NA No
profit
Bull City Durham Beauty and Barber 9.833 Private for- NA No
College profit
Fortis College-Cutler Bay 12.667 Private for- NA No
profit
Unitech Training Academy-Baton Rouge 6.991 Private for- NA No
profit
Empire Beauty School-Tampa 7.917 Private for- NA No
profit
Empire Beauty School-Lakeland 7.667 Private for- NA No
profit
Galen College of Nursing-ARH 16.500 Private for- NA No
profit
Tricoci University of Beauty Culture- 8.468 Private for- NA No
Janesville profit
Lynnes Welding Training-Bismarck 3.385 Private for- NA No
profit
No Grease Barber School 9.833 Private for- NA No
profit
Pima Medical Institute-San Marcos 8.910 Private for- NA No
profit
The College of Health Care Professions- 8.852 Private for- NA No
South San Antonio profit
Drury University-College of Continuing 12.124 Private NA No
Professional Studies nonprofit
Salon Success Academy-West Covina 7.089 Private for- NA No
profit
Indiana Institute of Technology-College 13.781 Private NA No
of Professional Studies nonprofit
Toni & Guy Hairdressing Academy-Rio 9.500 Private for- NA No
Rancho profit
The Salon Professional Academy- 7.030 Private for- NA No
Washington DC profit
Fortis Institute-Cookeville 10.107 Private for- NA No
profit
[Link] 1_ Basic R … 10/27
5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

name median_debt ownership admit_rate hbcu


<chr> <dbl> <chr> <dbl> <chr>
Fortis College-Landover 9.253 Private for- NA No
profit
Stautzenberger College-Rockford 13.300 Private for- NA No
Career College profit
As we can see, the most selective schools now top the list. You'll see some NA values
from admit_rate at the bottom of the arranged dataset. These are missing values,
which we'll discuss later.
To arrange the data in descending order of admission rates (highest admission rates on
top), we can use the desc argument within our arrange command:
In [19]: ## Run this code but do not edit it
# Arrange data in descending order of their admission rates
arrange(example_dat, desc(admit_rate))

[Link] 1_ Basic R … 11/27


5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

A [Link]: 4435 × 5
name median_debt ownership admit_rate hbcu
<chr> <dbl> <chr> <dbl> <chr>
University of Arkansas Community 6.250 Public 100 No
College-Morrilton
Design Institute of San Diego 31.000 Privateprofit
for- 100 No
Naropa University Private
16.390 nonprofit 100 No
VanderCook College of Music Private
27.000 nonprofit 100 No
Saint Elizabeth School of Nursing Private
20.291 nonprofit 100 No
Maharishi International University Private
13.085 nonprofit 100 No
Grace Christian University Private
9.708 nonprofit 100 No
Sacred Heart Major Seminary Private
7.343 nonprofit 100 No
JFK Muhlenberg Harold B. and Dorothy Private
15.750 nonprofit 100 No
A. Snyder Schools
Arnot Ogden Medical Center Private
11.744 nonprofit 100 No
Neighborhood Playhouse School of the Private
12.000 nonprofit 100 No
Theater
Samaritan Hospital School of Nursing Private
14.250 nonprofit 100 No
Trinity Bible College and Graduate Private
12.835 nonprofit 100 No
School
Trinity Health System School of Nursing Private
13.625 nonprofit 100 No
Warner Pacific University Private
24.382 nonprofit 100 No
New Castle School of Trades 8.729 Privateprofit
for- 100 No
Saint Charles Borromeo Seminary- Private
16.500 nonprofit 100 No
Overbrook
Universidad Adventista de las Antillas Private
11.850 nonprofit 100 No
Greene County Career and Technology 16.325 Public 100 No
Center

[Link] 1_ Basic R … 12/27


5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

name median_debt ownership admit_rate hbcu


<chr> <dbl> <chr> <dbl> <chr>
Western Area Career & Technology 16.500 Public 100 No
Center
Hussian College-Daymar College 9.500 Privateprofit
for- 100 No
Clarksville
Eastern Center for Arts and Technology 8.554 Public 100 No
Greater Lowell Technical School 5.500 Public 100 No
Cass Career Center 9.500 Public 100 No
Orange Ulster BOCES-Practical Nursing 11.880 Public 100 No
Program
Washington Saratoga Warren Hamilton 12.800 Public 100 No
Essex BOCES-Practical Nursing Program
Mifflin County Academy of Science and 11.756 Public 100 No
Technology
Living Arts College 10.003 Privateprofit
for- 100 No
Cayuga Onondaga BOCES-Practical 7.709 Public 100 No
Nursing Program
Delaware County Technical School- 16.500 Public 100 No
Practical Nursing Program
⋮ ⋮ ⋮ ⋮ ⋮

National Personal Training Institute- 6.333 Privateprofit


for- NA No
Tampa
Mobile Technical Training 3.800 Privateprofit
for- NA No
California Institute of Arts & Technology 9.500 Privateprofit
for- NA No
Elite Cosmetology Barber & Spa 6.054 Privateprofit
for- NA No
Academy
Gwinnett Institute 9.500 Privateprofit
for- NA No
Manuel and Theresa's School of Hair 6.494 Privateprofit
for- NA No
Design
Peloton College 9.500 Privateprofit
for- NA No
Ross Medical Education Center - 8.089 Privateprofit
for- NA No
Kalamazoo
Ross College-Canton 8.347 Privateprofit
for- NA No

[Link] 1_ Basic R … 13/27


5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

name median_debt ownership admit_rate hbcu


<chr> <dbl> <chr> <dbl> <chr>
Ross College-Grand Rapids North 7.125 Privateprofit
for- NA No
American Institute-Somerset 9.176 Privateprofit
for- NA No
Bull City Durham Beauty and Barber 9.833 Privateprofit
for- NA No
College
Fortis College-Cutler Bay 12.667 Privateprofit
for- NA No
Unitech Training Academy-Baton Rouge 6.991 Privateprofit
for- NA No
Empire Beauty School-Tampa 7.917 Privateprofit
for- NA No
Empire Beauty School-Lakeland 7.667 Privateprofit
for- NA No
Galen College of Nursing-ARH 16.500 Privateprofit
for- NA No
Tricoci University of Beauty Culture- 8.468 Privateprofit
for- NA No
Janesville
Lynnes Welding Training-Bismarck 3.385 Privateprofit
for- NA No
No Grease Barber School 9.833 Privateprofit
for- NA No
Pima Medical Institute-San Marcos 8.910 Privateprofit
for- NA No
The College of Health Care Professions- 8.852 Privateprofit
for- NA No
South San Antonio
Drury University-College of Continuing Private
12.124 nonprofit NA No
Professional Studies
Salon Success Academy-West Covina 7.089 Privateprofit
for- NA No
Indiana Institute of Technology-College Private
13.781 nonprofit NA No
of Professional Studies
Toni & Guy Hairdressing Academy-Rio 9.500 Privateprofit
for- NA No
Rancho
The Salon Professional Academy- 7.030 Privateprofit
for- NA No
Washington DC
Fortis Institute-Cookeville 10.107 Privateprofit
for- NA No
Fortis College-Landover 9.253 Privateprofit
for- NA No
[Link] 1_ Basic R … 14/27
5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

name median_debt ownership admit_rate hbcu


<chr> <dbl> <chr> <dbl> <chr>
Stautzenberger College-Rockford Career 13.300 Privateprofit
for- NA No
College

1.7 - Use the arrange command to organize the colleges in my_dat such that the
colleges with the highest student loan default rates are at the top.
In [20]: # Your code goes here
arrange(my_dat, desc(default_rate))

[Link] 1_ Basic R … 15/27


5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

A [Link]: 4435 × 5
name region default_rate ownership pct_PELL
<chr> <chr> <dbl> <chr> <dbl>
Tomorrow's Image Barber And South 57.1 Private for- 88.89
Beauty Academy of Virginia profit
Bull City Durham Beauty and South 57.1 Private for- 53.33
Barber College profit
No Grease Barber School South 57.1 Private for- 94.12
profit
Barber Institute of Texas Rockies & 56.4 Private for- 100.00
Southwest profit
Natural Images Beauty College Rockies & 53.1 Private for- 56.79
Southwest profit
Nuvani Institute Rockies & 51.6 Private for- 96.55
Southwest profit
B-Unique Beauty and Barber South 48.4 Private for- 62.50
Academy profit
Kenny's Academy of Barbering Midwest 44.4 Private for- 72.33
profit
Louisiana Academy of Beauty South 43.7 Private for- 15.00
profit
Denmark Technical College South 43.1 Public 47.82
Vibe Barber College South 41.1 Private for- 78.31
profit
Champ's Barber School Northeast 40.0 Private for- 92.11
profit
Bennett Career Institute Northeast 39.7 Private for- 45.75
profit
Bos-Man's Barber College South 38.2 Private for- 100.00
profit
Virginia University of Lynchburg South 35.7 Private
nonprofit 89.08
Lane College South 34.5 Private 87.13
nonprofit
Jacksonville College-Main Campus Rockies & 34.5 Private 34.87
Southwest nonprofit
West Michigan College of Midwest 34.4 Private for- 42.47
Barbering and Beauty profit
Barber Tech Academy South 34.1 Private for- 60.71
profit

[Link] 1_ Basic R … 16/27


5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

name region default_rate ownership pct_PELL


<chr> <chr> <dbl> <chr> <dbl>
Sebring Career Schools-Huntsville Rockies & 33.8 Privateprofit
for- 60.98
Southwest
Sebring Career Schools-Houston Rockies & 33.8 Privateprofit
for- 41.78
Southwest
Ponca City Beauty College Rockies & 33.3 Privateprofit
for- 65.52
Southwest
University Academy of Hair Design South 33.3 Privateprofit
for- 52.63
More Tech Institute South 33.3 Privateprofit
for- 98.33

United Tribes Technical College Midwest Private


31.1 nonprofit 76.04
Livingstone College South Private 84.04
30.6 nonprofit
Southwestern Christian College Rockies & Private 94.34
30.1 nonprofit
Southwest
Buckner Barber School Rockies & 30.0 Privateprofit
for- 67.92
Southwest
Southwest School of Business and Rockies & 29.8 Privateprofit
for- 91.94
Technical Careers-San Antonio Southwest
P&A Scholars Beauty School Midwest 29.6 Privateprofit
for- 53.49
⋮ ⋮ ⋮ ⋮ ⋮

Appalachian Bible College South 0 Private 41.33


nonprofit
West Virginia Junior College- South 0 Private for- 82.68
Charleston profit
West Virginia Junior College- South 0 Private for- 75.04
Morgantown profit
Bellin College Midwest 0 Private 13.57
nonprofit
Franklin County Career and Northeast 0 Public 50.00
Technology Center
West Virginia Junior College- Northeast 0 Private for- 80.77
United Career Institute profit
Eastern Center for Arts and Northeast 0 Public 40.64
Technology
School of Automotive Machinists & Rockies & 0 Private for- 47.32
Technology Southwest profit
[Link] 1_ Basic R … 17/27
5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

name region default_rate ownership pct_PELL


<chr> <chr> <dbl> <chr> <dbl>
Soka University of America Far West 0 Private 20.69
nonprofit
Ohio State School of Midwest 0 Private for- 57.28
Cosmetology-Heath profit
Pinnacle Institute of Cosmetology South 0 Private for- 54.17
profit
Franklin W Olin College of Northeast 0 Private 12.44
Engineering nonprofit
West Virginia Junior College- South 0 Private for- 72.73
Bridgeport profit
ATA College Far West 0 Private for- 94.12
profit
CES College Far West 0 Private
nonprofit 54.49
Career Development Institute Inc Far West 0 Private for- 66.89
profit
The University of Aesthetics & Midwest 0 Private for- 43.08
Cosmetology profit
Salon & Spa Institute Rockies & 0 Private for- 67.42
Southwest profit
Aveda Institute-Boise Rockies & 0 Private for- 57.69
Southwest profit
Medical Allied Career Center Far West 0 Private for- 65.56
profit
Medspa Academies Rockies & 0 Private for- 29.64
Southwest profit
School of Missionary Aviation Midwest 0 Private 17.86
Technology nonprofit
Southern Texas Careers Academy Rockies & 0 Private for- 86.54
Southwest profit
Academy of Interactive Far West 0 Private 32.08
Entertainment nonprofit
Lionel University Far West 0 Private for- 32.20
profit
Center for Ultrasound Research & Northeast 0 Private for- 35.32
Education profit
Saint Michael College of Allied Northeast 0 Private for- 81.69
Health profit
Image Maker Beauty Institute South 0 Private for- 45.68
profit
[Link] 1_ Basic R … 18/27
5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

name region default_rate ownership pct_PELL


<chr> <chr> <dbl> <chr> <dbl>
CAAN Academy of Nursing Midwest 0 Private 76.92
nonprofit
Hoss Lee Academy Far West 0 Privateprofit
for- 55.48

1.8 - What patterns do you notice among the programs that have the highest student
loan default rates? What do you wonder?
Double-click to type a response:
Reference Guide for R (student resource) - Now that you've seen a number of
different commands in R, check out our reference guide for a full listing of useful R
commands for this project.

2.0 - Finding summary statistics


When analyzing variables of interest, it's often helpful to calculate summary statistics.
For quantitative variables, we can use the summary command to find the five-number
summary (minimum, Q1, median, Q3, maximum) and the average (mean) of the values.
The code block shows how we find these summary statistics for the admit_rate
variable.
Note: The $ sign in R is used to isolate a single variable ( admit_rate ) from a full
dataframe ( dat ).
In [21… ## Run this code but do not edit it
# Find summary statistics for admit_rate
summary(dat$admit_rate)

Min. 1st Qu. Median Mean 3rd Qu. Max. NA's


2.44 59.79 74.68 70.81 86.11 100.00 2731

A few interesting facts about admit_rate that are revealed by this summary:
As expected, no schools have a 0% admissions rate (the minimum admissions rate
is 2.4%).
The maximum admissions rate was 100%. So, there's at least one school that
admits every applicant.
The first quartile (Q1) is a 59.79% admissions rate. This means only 25% of
schools have admissions rates lower than 59.79%.

[Link] 1_ Basic R … 19/27


5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

For 2,731 schools, we have missing data. R uses the sybmol NA to represent
missing values. If we use admit_rate in future analyses, we should pay
attention to which schools have missing data and, ideally, investigate why their
data is missing.
2.1 - Use the summary command to get summary statistics for the
default_rate variable in the dat dataframe.
In [22… # Your code goes here
summary(dat$default_rate)

Min. 1st Qu. Median Mean 3rd Qu. Max.


0.00 4.40 8.20 9.06 12.30 57.10

Check yourself: The median should be 8.20

2.2 - Comment on what these summary statistics reveal about the


default_rate values in our dataset.

Double-click to type a response:


For categorical data, it doesn't make sense to find means and medians. Instead, it's
helpful to look at value counts and proportions. We can use the table command to
find the counts of the different values for highest_degree :
In [23… ## Run this code but do not edit it
# Find counts of values for highest_degree, store in object 'degree_counts
degree_counts <- table(dat$highest_degree)

# Print table stored in 'degree_counts'


degree_counts

Associates Bachelors Certificate Graduate


1096 501 1374 1464

1464 of the institutions in our dataset are Universities that offer graduate degrees. On
the other end of the spectrum, 1374 of the institutions aren't Universities at all. Rather,
they are career-oriented programs that offer trade certificates.
To get a better sense of scale, we can turn these raw counts into proportions by
dividing them by the total:
In [24… ## Run this code but do not edit it
# Sum all counts in table, store in object 'total'
total <- sum(degree_counts)

[Link] 1_ Basic R … 20/27


5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

# Print the value stored in 'total'


total

4435
In [25… ## Run this code but do not edit it
# Divide the table by the total to get proportions
degree_counts / total

Associates Bachelors Certificate Graduate


0.2471251 0.1129651 0.3098083 0.3301015

As you can see, you can use R just like a calculator. Addition, subtraction,
multiplication, division ... it's all there. Universities offering graduate degrees make up
about 33% of the institutions in our dataset. These are about three times more
prevalent than 4-year colleges (Bachelors) that don't offer graduate degrees.
2.3 - Use the table command to get the value counts for the ownership
variable.
In [26… # Your code goes here
own <- table(dat$ownership)
own

Private for-profit Private nonprofit Public


1684 1212 1539

Check yourself: There are 1539 public schools in the dataset

2.4 - Find the proportion of all institutions that are public, private nonprofit, and
private for-profit.
In [28… # Your code goes here
total1 <- sum(own)
own/total1

Private for-profit Private nonprofit Public


0.3797069 0.2732807 0.3470124

Check yourself: About 34.7% of the schools in the dataset are public schools

3.0 - Visualizing data (histograms, barplots, and boxplots)


In addition to summary statistics, a great way to get an overall impression of our data is
to visualize it. In this section, we'll walk through different types of visualizations we can
create in R. Note: We're saving scatterplots for the next notebook in our series.

[Link] 1_ Basic R … 21/27


5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

One of the most useful visualizations for displaying a quantitative variable is a


histogram. Here, we use the gf_histogram command to display the histogram for
admit_rate .

In [29… ## Run this code but do not edit it


# Create histogram for admit_rate
gf_histogram(~admit_rate, data = dat)

Warning message:
“Removed 2731 rows containing non-finite outside the scale range (`stat_bin
()`).”

Note: A warning message was displayed about removing rows. This is R telling us that
it's choosing not to visualize the missing data values ( NA ) that we discovered for
admit_rate earlier in the notebook.

As we suspected from the summary statistics, it appears that most programs have
admissions rates well above 50%, and only a small subset of programs have highly
selective admissions rates. In statistics, we call this distribution left skew, since
there's a tail on the left side. So, institutions with low values (low admissions rates) are
relatively unusual compared to most of the other institutions in our dataset.
3.1 - Create a histogram to visualize all the default_rate values in the dat
dataframe.
In [30… # Your code goes here

[Link] 1_ Basic R … 22/27


5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

gf_histogram(~admit_rate, data=dat)

Warning message:
“Removed 2731 rows containing non-finite outside the scale range (`stat_bin
()`).”

3.2 - Describe the distribution and note any features of interest.


Double-click to type a response: The distribution is skewed to the left with a center
around 78.
To visualize categorical variables, we can use the gf_bar command to make bar
plots. Here we create a bar plot for highest_degree :
In [ ]: ## Run this code but do not edit it
# Create bar plot for highest_degree
gf_bar(~highest_degree, data = dat)

As shown here, most of the institutions in our dataset are Universities that graduate
degrees or trade programs that offer professional certificates. There are about 500
colleges that only offer bachelors degrees (without offering graduate degrees).
3.3 - Create a bar plot to visualize the ownership values from the dat
dataframe.

[Link] 1_ Basic R … 23/27


5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

In [31… # Your code goes here


gf_bar(~ownership, data=dat)

3.4 - Describe the distribution and note any features of interest.


Double-click to type a response: The barplots are fairly similar, with the count for
private for-profit being the greatest at 1700 while the lowest is private nonprofit is
1200.
Sometimes, we may want to explore the relationship between two variables by
visualizing them both at once. When we want to visualize the relationship between a
categorical variable and quantitative variable, we can use boxplots. Here, we show how
to use gf_boxplot to visualize the relationship between highest_degree
(categorical) and admit_rate (quantitative).
In [32… ## Run this code but do not edit it
# Create boxplots for admit_rates of institutions with different highest_de
gf_boxplot(admit_rate ~ highest_degree, data = dat)

Warning message:
“Removed 2731 rows containing non-finite outside the scale range
(`stat_boxplot()`).”

[Link] 1_ Basic R … 24/27


5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

In this case, we're using highest_degree as the predictor variable and


admit_rate as the outcome variable. In other words, we can use the degree level
of an institution (certificate, associates, bachelors, etc.) to help predict its admission
rate. That's because certain levels of institutions typically have lower admissions rates
than others. So, knowing the level of an institution can help us better predict its
admissions rate.
Note: This predictor-outcome relationship is coded in R through the syntax outcome
~ predictor , as in gf_boxplot(admit_rate ~ highest_degree,...) .

We see that admission rates tend to be lower (lower medians) for colleges /
Universities that grant bachelors and graduate degrees. However, it's worth noting that
for every institution-type, the first quartile is higher than a 50% admissions rate. So,
most programs admit more than half their applicants, regardless of insitution-type.
Indeed, we see that the most prestigious Universities with admissions rates lower than
25% are outliers (visualized as dots on the boxplot) among other Universities that offer
graduate degrees.
3.5 - Create boxplots to visualize the relationship between ownership and
default_rate from the dat dataframe.

In [34… # Your code goes here


gf_boxplot(ownership ~ default_rate, data=dat)

[Link] 1_ Basic R … 25/27


5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

3.6 - Using your boxplot visualization, describe the relationship between institution
ownership and studen loan default rates.
Double-click to type a response: Shows that private for-profit colleges typically have
the highest student default loan rates, followed by public institutions. Meanwhile,
private non-profit have the lowest rate. This implies students at for profit schools are
more likely to default on their loans.

Summer Opportunity: Do you want to learn more about Data


Science & AI?
Join our Data Science & AI Summer Bootcamp, where you'll take your learning from
this project to the next level. No prior coding or statistics experience required!
Designed by Harvard grads, the bootcamp allows students from all experience
levels to dive deeper into data science concepts, from the basics (e.g. linear
regression) to the advanced (e.g. AI neural networks). Students learn in a
supportive and collaborative environment, and they walk away with their own real-
world project that can be shared on college and internship applications.
📢 Scholarships are available! We’re committed to making this opportunity
accessible to all students.
📝 Applications are considered on a rolling basis. Final application deadline: May
30, 2025
[Link] 1_ Basic R … 26/27
5/21/25, 12:25 AM Notebook 1_ Basic R & Data Exploration

🔗 Learn more and apply here: [Link]

[Link] 1_ Basic R … 27/27

You might also like