Unit-1 2BCA R Programming
Unit-1 2BCA R Programming
[Link]
UNIT-I
[Link] 2
Introduction of the Language
Introduction
R is a programming language and free software environment designed for
statistical computing and graphics. It is widely used by statisticians, data
scientists, and researchers for data analysis, visualization, and reporting.
Developed in the early 1990s by Ross Ihaka and Robert Gentleman, R was
built to simplify complex data manipulation and create clear, customizable
visualizations. Over time, it has gained popularity among statisticians, data
scientists and researchers because of its capabilities and the vast array of
packages available.
[Link] 3
Why Choose R Programming?
R is a unique language that offers a wide range of features for data analysis,
making it an essential tool for professionals in various fields. Here’s why R is
preferred:
Free and Open-Source: R is open to everyone, meaning users can modify,
share and distribute their work freely.
Designed for Data: R is built for data analysis, offering a comprehensive
set of tools for statistical computing and graphics.
Large Package Repository: The Comprehensive R Archive Network
(CRAN) offers thousands of add-on packages for specialized tasks.
Cross-Platform Compatibility: R can work on Windows, Mac and Linux
operating systems.
Great for Visualization: With packages like ggplot2, R makes it easy to
create informative, interactive charts and plots.
Key Features of R
Cross-Platform Support: R works on multiple operating systems, making
it versatile for different environments.
Interactive Development: R allows users to interactively experiment with
data and see the results immediately.
Data Wrangling: Tools like dplyr and tidyr help simplify data cleaning and
transformation.
Statistical Modeling: R has built-in support for various statistical models
like regression, time-series analysis and clustering.
Reproducible Research: With R Markdown, users can combine code,
output and narrative in one document, ensuring their analysis is
reproducible.
[Link] 4
Applications of R Programming in the Real World.
[Link] 5
4. Data Visualization & Reporting
R excels in creating high-quality, customizable data visualizations, from
basic charts to complex interactive dashboards, the powerful packages are
used like ggplot2 and plotly. The Shiny package allows for building
interactive web applications directly from R code.
5. Finance:
Data Science is most widely used in the financial industry. The finance
industry relies on R for portfolio optimization, risk analysis, and asset
pricing.
Libraries in R simplify moving averages, auto-regression, time-series
analysis, stock-market modeling, financial data mining, and downside risk
assessment. This helps in better decision-making and produces good results.
We can display the results of the analysis using high-quality candlestick
charts, density, and drawdown plots.
Companies like American Express, Bajaj Allianz Insurance, JP Morgan,
Standard Chartered, etc., use R.
6. Banking:
Banking industries also use R for credit risk modeling, fraud detection, and
other risk analytics. Packages, such as caret and randomForest, help develop
credit scoring models and also help to identify fraud patterns.
7. E-Commerce:
E-commerce goes beyond in its usage of data science. R helps by providing
valuable insights into customer behavior, sales forecasting, and personalized
marketing. E-commerce platforms use data analysis and representation tools
to understand customer preferences and different markets. It also helps
optimize pricing strategies.
R is used to improve cross-product selling. When a customer buys a product,
the site suggests extra products that complement their original order. These
suggestions also work for products purchased by the customer in the past.
[Link] 6
8. IT Sector:
IT companies use R for analyzing data, machine learning, and software
development. Agencies use R for business intelligence. They also offer these
services to other businesses of different sizes.
IT companies use R's machine learning abilities to detect network breaches,
filter spam, and make recommendations.
Big IT companies like Accenture, IBM, Infosys, Paytm, TCS, and Wipro use
R.
9. Social Media:
Today, a lot of data is generated and regulated on social media. Therefore,
data science is widely used in the social media industry. R is a tool used for
social analytics. It helps companies get useful results from their social media
data. Companies can learn about customers' feelings by studying data,
checking how people see their brand, and spotting new trends.
R is also used to analyze traffic, user sessions, and content. Organizations
use it to improve users' suggestions based on their history, mood, and
recent posts and content views.
10. Data Exploration and Cleaning:
Data exploration and cleaning are foundational steps in any data analysis
process. R's capabilities make it an ideal choice for handling missing values,
outliers, and ensuring overall data quality before diving into in-depth
analysis.
11. Environmental Science and Climate Research:
R contributes significantly to environmental science by analyzing climate
data, predicting environmental trends, and assessing the impact of human
activities on ecosystems. Its applications in climate research are vital for
understanding and addressing environmental challenges. With increasing
concerns around the sustainability of our planet, data scientists are relying
[Link] 7
on R for environmental research and climate modeling, showcasing the
practical applications of R in addressing global issues.
Advantages of R Programming.
[Link] 8
functions and tools to implement complex algorithms in time time-
efficient manner.
Academia and Research: R is best suited for academic and research use;
it allows researchers to perform advanced data analysis to publish it.
RMarkdown and knitr Support: It allows reproducible research by
combining data, code, and analysis results in a single document.
Customs Function and Tailor Analysis: R’s extensibility and flexible
nature let users customize functions as per their specific needs. Additionally,
users can also connect R with different programming languages such as C,
Python, and Java.
Provides Libraries for Data Visualization: It provides libraries
like ggplot2 that give the flexibility to create various impressive data
visualizations such as publication-quality plots, charts, and graphs.
Therefore, it helps businesses and organizations to communicate insights
through this data effectively.
Cost-Effective Solutions: R is free to use and redistributable, which makes
it an excellent choice for organizations and individuals to leverage data
analysis capabilities at a low cost.
Strong Foundation for Statistics: R is a powerful tool to conduct linear
and nonlinear modeling, time series analysis, and hypothesis testing
for Statistical analysis. Users can use its built-in functions and libraries to
perform complex statistical tasks.
[Link] 9
Disadvantages of R.
No Robust OOP Support: R does not provide robust support for Object-
Oriented Programming which may cause restrictions on many software
designs and programming structures.
[Link] 10
Package Fragmentation: R offers an array of packages, which can lead to
overlapping functionalities. Sometimes it is not well maintained and
potentially causes compatible issues.
[Link] 11
Advantages of R over Other Programming Languages
R, compared to other programming languages, offers distinct advantages,
particularly in the realm of statistical computing and data analysis.
1. Designed for Statistical Computing and Data Analysis:
R was specifically created for statistical analysis, providing a comprehensive
suite of tools for statistical modeling, hypothesis testing, time-series
analysis, classification, and clustering. This specialization provides a more
integrated and optimized environment for statistical tasks compared to
general-purpose languages.
2. Rich Ecosystem of Packages:
The Comprehensive R Archive Network (CRAN) hosts thousands of user-
contributed packages that significantly extend R's capabilities. These
packages provide specialized functions for almost any data-related task,
from advanced statistical methods to machine learning algorithms and data
visualization.
3. Advanced Data Visualization:
R is renowned for its powerful and flexible data visualization capabilities.
Packages like ggplot2 allow for the creation of highly customizable,
publication-quality plots and charts, enabling users to effectively
communicate insights from data.
4. Open Source and Free:
R is an open-source language, R is free to use, distribute, and modify. This
eliminates licensing costs and fosters a large, active community that
contributes to its development and provides extensive support resources.
5. Reproducible Research:
Tools like R Markdown facilitate reproducible research by allowing users to
combine code, output, and narrative in a single document. This ensures that
analyses can be easily shared, understood, and replicated.
[Link] 12
6. Cross-Platform Compatibility:
R is platform-independent, meaning it can run seamlessly on various
operating systems, including Windows, macOS, and Linux, without requiring
significant modifications.
7. Active Community Support:
R benefits from a large and vibrant global community of users and
developers. This provides a wealth of online resources, forums, tutorials, and
workshops for learning and problem-solving.
R Software -Installations
Click the below link to download R Software.
[Link]
After installing R Software, window will look like this.
[Link] 13
To Install R Studio, click the below link.
[Link]
[Link] 14
R Script File
R Studio is an integrated development environment (IDE) for R. IDE is a
GUI, where you can write your quotes, see the results and also see the
variables that are generated during the course of programming. R is
available as an Open Source software for Client as well as Server Versions.
1. Creating an R file:
There are two ways to create an R file in R studio:
You can click on the File tab, from there when you click it will give a drop-
down menu, where you can select the new file and then R script, so that,
you will get a new file open.
[Link] 15
Once you open an R script file, this is how an R Studio with the
script file open looks like.
[Link] 16
4. Execution of an R file:
There are several ways in which the execution of the commands that are
available in the R file is done.
[Link] 17
Handling Packages in R:
Handling packages in R involves several key steps: installation, loading, and
managing them. Packages are collections of R functions, data, and compiled
code that extend the functionality of base R.
Write the Package name in Package, and click on install. Your package is
installed.
[Link] 18
Keywords:
Keywords in R are reserved words that have special meaning within the
language and cannot be used as identifiers (e.g., variable names, function
names). Here is a list of keywords in R:
if, else, repeat, while, function, for, in, next, break, TRUE, FALSE, NULL, Inf,
NaN, NA, NA_integer_, NA_real_, NA_complex_, NA_character_
Variable
In computer programming, a variable is a named memory location where
data is stored.
Rules to be followed while naming a variable.
➢ A variable name in R can be created using letters, digits, periods, and
underscores.
➢ You can start a variable name with a letter or a period, but not with
digits.
➢ If a variable name starts with a dot, you can't follow it with digits.
➢ R is case sensitive. This means that name, Name and NAME are
treated as different variables.
➢ Reserved words cannot be used as variable names.
[Link] 19
Variable Assignment
The variables can be assigned values using leftward, rightward and equal to
operator.
[Link] 20
Syntax of Variables
Creating variables in R and giving them values requires the assignment
operator, which can be either <- or -> or =.
The following is the standard syntax for generating variables in R:
variable_name <- value
or
variable_name -> value
or
variable_name = value
433 -> b
The value 433 is assigned to variable b.
x = 114
The value 114 is assigned to x.
[Link] 21
Data Types in R
Data types in R define the kind of values that variables can hold. Choosing
the right data type helps optimize memory usage and computation.
[Link] 22
1. Numeric Data Type
Decimal values are called numeric in R. It is the default R data type for
numbers in R.
Real numbers with a decimal point are represented using this data type in R.
It uses a format for double-precision floating-point numbers to represent
numerical values.
[Link] 23
Even if an integer value 46 is assigned to a variable x, it is still saved as a
numeric value.
[Link] 24
2. Integer Data Type
The integer data type specifies real values without decimal points. We use
the suffix L to specify integer data.
[Link] 25
[Link] 26
4. Logical Data Type
The logical data type in R is also known as Boolean data type. It can only
have two values: TRUE and FALSE.
[Link] 27
5. Complex Data Type
The complex data type is used to specify purely imaginary values in R. We
use the suffix i to specify the imaginary part.
[Link] 28
6. Raw Data Type
A raw data type specifies values as raw bytes. You can use the following
methods to convert character data types to a raw data type and vice-versa:
charToRaw() - converts character data to raw data
rawToChar() - converts raw data to character data
[Link] 29
In this program,
We have first used the charToRaw() function to convert the string "Kalyan"
to raw bytes.
This is why we get "raw" as output when we print the class of raw_variable.
Then, we have used the rawToChar() function to convert the data in
raw_variable back to character form.
This is why we get "character" as output when we print the class of
char_variable.
Operators in R Programming
Operators are the symbols directing the compiler to perform various kinds of
operations between the operands
R language is rich in built-in operators and provides following types of
operators.
Types of Operators
We have the following types of operators in R programming −
➢ Arithmetic Operators
➢ Relational Operators
➢ Logical Operators
➢ Assignment Operators
➢ Miscellaneous Operators
[Link] 30
Arithmetic Operators
Arithmetic operators in R programming perform a variety of mathematical
operations, such as addition, subtraction, multiplication, division, and
modulo. You can perform specific operations using the operators between
two or more operands. These operands can be vector, scalar or complex
values.
[Link] 31
Relational Operators
The relational operators in R programming perform a variety of comparison
operations between the operands. It returns a Boolean value of TRUE if the
relation between the first and second operands is satisfied.
[Link] 32
Logical Operators
Logical operators in R programming perform a variety of decision operations
and gives output as TRUE or FALSE. You can consider a non-zero number as
True.
[Link] 33
Assignment Operators
Assignment Operators in R are used to assigning values to various data
objects in R. The objects may be integers, vectors, or functions. These
values are then stored by the assigned variable names. There are two kinds
of assignment operators: Left and Right. There are two types of assignment
operators in R Programming, Right and Left.
[Link] 34
Miscellaneous Operators in R
These operators are used for specific purposes and are not general
mathematical or logical computers. These operators include the colon
operator, %in% operator, and %*% operator.
a) Colon Operator (:)
Creates a sequence of numbers.
x <- 1:5 #Creates a vector with elements 1, 2, 3, 4, 5
print(x)
Output
[1] 1 2 3 4 5
[Link] 35
Data Structure in R
Data structures are used to store and organize values. R provides several
fundamental data structures to organize and store data, each suited for
different types of data and analytical tasks. These can be broadly classified
by their dimensionality and whether they are homogeneous (elements of the
same data type) or heterogeneous (elements of different data types).
[Link] 36
Vectors:
Vector is one of the basic data structures in R. It is homogenous, which
means that it only contains elements of the same data type. Data types can
be numeric, integer, character, complex, or logical.
Vectors are created by using the c() function. The typeof() function is used
to check the data type of the vector, and the class() function is used to
check the class of the vector.
Syntax:
vector_name <- c(value1, value2, ...)
Example:
# Creating a numeric vector
numeric_vector <- c(1, 2, 3.5, -4, 0)
print(numeric_vector)
Output:
1.0 2.0 3.5 -4.0 0.0
[Link] 37
Lists:
A list is a non-homogeneous data structure, which implies that it can contain
elements of different data types. It accepts numbers, characters, lists, and
even matrices and functions inside it. It is created by using the list()
function.
Syntax:
list_name <- list(element1, element2, ...)
Example-1:
# Creating a list of various data types
my_list <- list("Nisarga", 28, “BCA”,TRUE, c(1, 2, 3))
print(my_list)
Output:
[[1]]
[1] "Nisarga"
[[2]]
[1] 28
[[3]]
[1] "BCA"
[[4]]
[1] TRUE
[[5]]
[1] 1 2 3
[Link] 38
Example-2:
empId = c(101, 102, 103, 104)
empName = c("MadanRaj", "Lokesh", "Aruna", "Roja")
Weight = c(62.5,70,60.4,45.8)
empList = list(empId, empName, Weight)
print(empList)
Output:
[[1]]
[1] 101 102 103 104
[[2]]
[1] "MadanRaj" "Lokesh" "Aruna" "Roja"
[[3]]
[1] 62.5 70.0 60.4 45.8
Matrices:
A matrix is a two-dimensional, homogeneous data structure where all
elements are of the same data type, arranged in rows and columns.
Syntax:
matrix_name <- matrix(data, nrow = num_rows, ncol = num_cols)
Example:
# Creating a matrix
data_matrix <- matrix(1:12, nrow = 3, ncol = 4)
print(data_matrix)
[Link] 39
Output:
Data Frames:
A data frame is a two-dimensional, heterogeneous data structure that is
used to store data in tabular form. These are lists of vectors of equal
lengths. To create a data frame we use the [Link]() function.
Data frames have the following constraints placed upon them:
➢ A data-frame must have column names and every row should have a
unique name.
➢ Each column must have the identical number of items.
➢ Each item in a single column must be of the same data type.
➢ Different columns may have different data types.
Syntax:
dataframe1 <- [Link](
first_col = c(val1, val2, ...),
second_col = c(val1, val2, ...),
...
)
Example:
Emp_dataframe <- [Link] (
EmpId = c("E101", "E102", "E103"),
EmpName = c("Arjun", "Afifa","Anna"),
DeptName= c("Accounts","Computer","Finance")
)
print(Emp_dataframe)
[Link] 40
Output:
Example-2:
Name = c("Afifa", "Roopa", "MadanRaj")
Language = c("R", "Python", "Java")
Age = c(22, 25, 45)
df = [Link](Name, Language, Age)
print(df)
Output:
Arrays:
Arrays are multi-dimensional data structures that can store elements of the
same data type. They are used for more complex data arrangements, such
as three-dimensional data.
For example, if we create an array of dimensions (2, 3, 3) then it creates 3
rectangular matrices each with 2 rows and 3 columns. They are
homogeneous data structures.
[Link] 41
Syntax:
array_name <- array(data, dim = c(num_rows, num_cols,
num_dimensions))
Example:
A = array(c(11, 22, 33, 44, 55, 66, 77, 88), dim = c(2, 2, 3))
print(A)
Output:
[Link] 42
Factors:
Factors are data structure used to categorize and store data on multiple
levels . The main advantage is that it can store both integer and character
types of data.
In R programming, levels are always stored in alphabetical order.
Factor can be ordered or unordered and are essential for statistical analysis
and plotting.
Syntax:
factor_name <- factor(vector_of_categories)
To create a factor, use the factor() function and add a vector as argument:
Example:
Output:
You can see from the example above that that the factor has four levels
(categories): Classic, Jazz, Pop and Rock.
[Link] 43
To only print the levels, use the levels() function:
Output:
Factor Length:
Use the length() function to find out how many items there are in the factor:
Output:
Access Factors:
To access the items in a factor, refer to the index number, using [] brackets:
Output:
[Link] 44
Change Item Value:
To change the value of a specific item, refer to the index number:
Change the value of the third item:
Output:
[Link] 45
Matrix Operations:
Addition, subtraction, multiplication and division can also be performed on
matrices in R.
#Create two 2x2 matrices.
matrix1 <- matrix(c(1:4), nrow = 2)
matrix2 <- matrix(c(5:8), nrow = 2)
# Add the matrices
result1 <- matrix1 + matrix2
# Subtract the matrices
result2 <- matrix1 - matrix2
# Multiply the matrices
result3 <- matrix1 * matrix2
# Divide the matrices
result4 <- matrix1 / matrix2
# Print
print("The first matrix is:")
print(matrix1)
print("The second matrix is:")
print(matrix2)
# Print results
print("The result of Addition of is:")
print(result1)
print("The result of Subtraction of is:")
print(result2)
print("The result of Multiplication of is:")
print(result3)
print("The result of Division of is:")
print(result4)
[Link] 46
Output:
#Create two 2x3 matrices.
matrix1 <- matrix(c(1:4), nrow = 2)
matrix2 <- matrix(c(5:8), nrow = 2)
# Add the matrices
result1 <- matrix1 + matrix2
# Subtract the matrices
result2 <- matrix1 - matrix2
# Multiply the matrices
result3 <- matrix1 * matrix2
# Divide the matrices
result4 <- matrix1 / matrix2
[1] "The first matrix is:"
[,1] [,2]
[1,] 1 3
[2,] 2 4
[1] "The second matrix is:"
[,1] [,2]
[1,] 5 7
[2,] 6 8
[Link] 47
[2,] -4 -4
[1] "The result of Multiplication of is:"
[,1] [,2]
[1,] 5 21
[2,] 12 32
[1] "The result of Division of is:"
[,1] [,2]
[1,] 0.2000000 0.4285714
[2,] 0.3333333 0.5000000
[Link] 48
Special Values:
[Link] 49
NULL – Absence of a Value:
NULL signifies an empty or undefined object, often returned by functions
expecting no result. It is different from NA because NULL means the object
does not exist, while NA means a value is missing. The key properties are:
➢ NULL is a zero-length object, while NA has a placeholder.
➢ Cannot be part of a vector.
➢ Functions return NULL if they operate on a NULL object.
➢ Use [Link]() to check for NULL.
Example:
y <- NULL
[Link](y)
Output:
[1] TRUE
Inf and -Inf – Infinity:
Inf and -Inf represent positive and negative infinity in R. These values occur
when numbers exceed the largest finite representable value. Inf arises from
operations like division by zero or overflow. The key properties are:
➢ Often results from division by zero.
➢ Can be used in comparisons (Inf > 1000 returns TRUE).
Example:
1/0
log(0)
[Link](1/0)
Output:
[1] Inf
[1] -Inf
[1] TRUE
Note that Infinite values can be checked with [Link](x). Inf and -
Inf results in NaN.
[Link] 50
NaN – Not a Number:
NaN results from undefined mathematical operations, like 0/0. One can
check NaN values by using [Link]() function. Let us see how to check for
NaN using R example:
Note that NULL is different from NA and NaN; it means no value exists. It is
commonly used for empty lists, missing function arguments, or when an
object is undefined.
Example:
0/0
[Link](0 / 0)
[Link](NaN)
Output:
[1] NaN
[1] TRUE
[1] TRUE
[Link] 51
Classes in R Programming
Classes and Objects are core concepts in Object-Oriented Programming
(OOP), modeled after real-world entities. In R, everything is treated as an
object. An object is a data structure with defined attributes and methods. A
class is a blueprint that defines a set of properties and methods shared by all
objects of that type.
R has a unique three-class system: S3, S4, and Reference Classes. Each
of these class systems has distinct characteristics and is used to define and
manage objects and their methods effectively.
1. S3 Class
S3 is the most widely used OOP system in R, but it lacks a formal definition
and structure. An object of this type can be created simply by adding an
attribute to it.
First we create a list with various components then we create a class using
the class() function. For example,
Example:
# create a list with required components
student1 <- list(name = "Nuthan Gowda", age = 21, GPA = 8.5)
# name the class appropriately
class(student1) <- "Student_Info"
# create and call an object
student1
In the above example, we have created a list named student1 with three
components. Notice the creation of class,
class(student1) <- "Student_Info"
[Link] 52
Here, Student_Info is the name of the class. And to create an object of this
class, we have passed the student1 list inside class().
Finally, we have created an object of the Student_Info class and called the
object student1.
S4 Class in R
S4 class is an improvement over the S3 class.
In R, we use the setClass() function to define a class. For example,
setClass("Student_Info", slots=list(name="character", age="numeric",
GPA="numeric"))
Here, we have created a class named Student_Info with three slots (member
variables): name, age, and GPA.
Now to create an object, we use the new() function. For example,
student1 <- new("Student_Info", name = "John", age = 21, GPA = 3.5)
We have successfully created the object named student1.
Example:
# create a class "Student_Info" with three member variables
setClass("Student_Info", slots=list(name="character", age="numeric",
GPA="numeric"))
# create an object of class
student1 <- new("Student_Info", name = "MadanRaj", age = 21, GPA
= 8.5)
# call student1 object
student1
[Link] 53
Output:
An object of class "Student_Info"
Slot "name":
[1] "MadanRaj"
Slot "age":
[1] 21
Slot "GPA":
[1] 8.5
Reference Class in R
Reference classes were introduced later, compared to the other two.
Defining a reference class is similar to defining a S4 class. Instead of
setClass() we use the setRefClass() function.
Example:
student <- setRefClass("student",
fields = list(name = "character", age = "numeric", GPA = "numeric"))
s <- student(name = "MadanRaj", age = 21, GPA = 8.5)
s
Output:
Reference class object of class "student"
Field "name":
[1] "MadanRaj"
Field "age":
[1] 21
Field "GPA":
[1] 8.5
In the above example, we have created a reference class named Student
using the setRefClass() function. we have used our generator function
Student() to create a new object s.
[Link] 54
Input/Output Functions in R
With R, we can read inputs from the user or a file using simple and easy-to-
use functions. Similarly, we can display the complex output or store it to a
file using the same.
Output Function in R
Here is the list of the functions that we use to display the output of any
program in R.
➢ print()
➢ cat()
➢ paste()
➢ sprintf()
➢ message()
a) print() Function
In R, we use the print() function to print values and variables. For example,
Example:
# print string
print("234")
print("Enter the value for N")
# print variables
y=12
print(y)
Output:
[1] "234"
[1] "Enter the value for N"
[1] 12
[Link] 55
Printing a number:
Let us try to print the value of pi and limit of digits to 3.
Example:
print(pi)
print(pi, digits =3)
Output:
[1] 3.141593
[1] 3.14
The first print(pi) displays the pi value.
The second print(pi, digit=3) displays the pi value with limited total 3 digits
only.
Printing a dataset:
Let us try to print a whole dataset using the print function.
Example:
Employee_Data <- [Link](Name = c("Ashwini", "Ramya", "Priya"),
Age = c(25, 30, 35))
print(Employee_Data)
Output:
[Link] 56
b) cat() Function:
The cat() function in R can also be used to display the results of a program.
It can concatenate different variables and display the output. It does not add
line breaks by default.
Example-1:
cat("Hello", "I am","from","Mysore",987654321,TRUE)
Output:
Example-2:
a=35007
cat("The value of a = ",a)
Output:
The value of a = 35007
c) paste() Function:
You can also print a string and variable together using the print() function.
For this, you have to use the paste() function inside print().
Example-1:
company <- "GKMV Kalyan"
# print string and variable together
print(paste("Welcome to", company))
Output:
[1] "Welcome to GKMV Kalyan"
Notice the use of the paste() function inside print(). The paste() function
takes two arguments: string - "Welcome to" and variable - company
By default, you can see there is a space between string Welcome to and the
value GKMV Kalyan.
[Link] 57
d) sprintf() Function
The sprintf() function is similar to the printf function in C Programming. It is
used to print formatted output using format specifiers (e.g., %s for string,
%d for integer, %f for float). It returns a formatted character vector.
Example-1:
value <- 412.3867
formatted_string <- sprintf("The value is %.2f", value)
print(formatted_string)
Output:
[1] "The value is 412.39"
Example-2:
# sprintf() with integer variable
myInteger <- 123
sprintf("Integer Value: %d", myInteger)
Output:
[1] "Integer Value: 123"
[1] "Float Value: 12.340000"
[Link] 58
e) message() Function
In R programming, the message() function is used to generate diagnostic
messages. These messages are distinct from warnings and errors, as they
are intended to provide informative output without necessarily indicating a
problem that needs fixing or stopping code execution.
The message() function takes one or more arguments, which are typically
character strings or objects that can be coerced to character strings.
Output:
This is an informational message.
The value of x is: 10
When to use message():
✓ To provide non-critical information to the user during function
execution.
✓ To indicate default values being used in a function.
✓ To offer guidance on interpreting function results.
✓ In package development, to provide informative output that can be
easily controlled by the user.
[Link] 59
Input Function in R
Reading input from the console enables the user to provide data or
parameters directly to a program during runtime. R provides several
methods to accomplish this task, including readline() and scan() functions.
These functions offer flexibility and convenience for capturing user input in
different scenarios.
In R, there are three methods to take user input.
➢ readline() method
➢ scan() method
➢ [Link]() method (To read CSV files)
readline() function
The readline() method is useful whenever we want to take a single line input
from the user. The readline simply means read a line from the terminal. This
method prompts the user to enter input from the console. Once entered, it
returns the input value as a string.
Syntax
readline (prompt = " ")
prompt: It is an optional parameter. It is used to specify the user what type
of input the program is expecting.
Example-1:
input_read <- readline()
print(input_read)
Output:
[Link] 60
Example-2:
#Taking user input
var <- readline()
#Printing type of variable
print(paste("Datatype: ",typeof(var)))
#Printing variable
print(paste("User Input:" ,var))
Output:
We can convert character type values to some other data type using
various methods in R.
[Link] 61
Let us look at some examples using these methods.
#Taking user input
val <- readline(prompt = "Enter the number: ")
#Printing type of variable
print(paste("Old datatype: ",typeof(val)))
#Converting into integer type
val <- [Link](val)
#Printing the type of variable
print(paste("New datatype: ",typeof(val)))
#printing the variable
print(val)
Output:
[Link] 62
scanf() function
The scan() method is another way using which we can take inputs in R. This
method is used to read inputs from the console or even a file. This is a
flexible method that allows you to specify the format or data type of an
input. This method is very useful when a user wants to take inputs that are
separated by spaces or new lines.
Note: This method takes input in the form of a vector or a list continuously.
Syntax:
scan( file = “ “, what = “ “, nmax = -1, …)
file: It is an optional parameter that is used to specify the name of the file
that a user wants to take input from. If it is not provided, then the input is
taken from the console.
what: It is an optional parameter that is used to specify the data type or the
format of the input.
nmax: It is also an optional parameter that is used to specify the maximum
number of inputs to read from the console/file.
Example 1:
#Taking input using scan()
print("Enter the input:")
val <- scan()
#Printing variable
print(val)
Note: This method will continuously take input from the user in the console.
In order to terminate the process, press the Enter key 2 times which will
take the control to the next line and stop the process.
[Link] 63
Example-2:
#Taking first input
val1 <- scan(what= " ", nmax = 1)
#Printing variable
print(val1)
#Taking second input
val2 <- scan(what = " ", nmax = 2)
#Printing variable
print(val2)
Output:
In the above example, the scan() method is used to take user input, and the what = “
“ argument is used to specify the string input type. In the first input, the nmax value is
set to 1; thus, the compiler only reads a single string value from our input, While in the
second input, the nmax value is 2, thus the compiler reads both string values.
[Link] 64
Scan():to read from file
We can read a text file using the scan() method. In order to read a text file,
we need to specify the location of the file and the data type of the file to the
scan() method.
Syntax:
scan("[Link]", what = "character")
[Link]: Text file to be scanned.
Returns: Scanned output.
Example-3:
# Taking file input using scan()
data <- scan("[Link]", what = "character")
# Printing the text file
print(data)
Output:
[Link] 65
Example 4: Scan Excel CSV File
[Link] 66
Graph Plotting in R Programming
Data Visualization becomes the most desirable way, it is always better to
visualize that data through charts and graphs, to gain meaningful insights.
a) Box Plotting
A box plot generates a rectangle that covers the area spanned by the
column of the dataset. It can be produced as follows:
boxplot(mtcars$mpg, col="green")
Note that the thick line in the rectangle depicts the median of the mpg
column, i.e. 19.20 as seen in the Five Point Summary. The col=”green”
simply colours the plot green.
b) Histograms
In R programming, histograms are a fundamental tool for visualizing the
distribution of a continuous numerical variable. They display the frequency
or count of data points falling into predefined intervals, known as "bins."
The ‘breaks’ argument essentially alters the width of the histogram bars. It
is seen that as we increase the value of the break, the bars grow thinner.
[Link] 67
temperatures <- c(67 ,72 ,62 ,76 ,66 ,65 ,59 ,61, 79)
# histogram of temperatures vector
result <- hist(temperatures)
hist(temperatures,
main = "Histogram of Temperature",
xlab = "Temperature in degrees Fahrenheit",
ylab = "Frequency",
col = "lightblue",
breaks = 5)
c) Bar Plotting
Bar Plots is one of the most efficient ways of representing data’s. It can be
used to summarize large data in visual form.
Bar graphs have the ability to represent data that shows changes over time,
which helps us to visualize trends.
we use the barplot() function to create bar plots. For example,
temperatures <- c(22, 27, 26, 24, 23, 26)
result <- barplot(temperatures,
main = "Maximum Temperatures in a Week",
xlab = "Degree Celsius",
ylab = "Day",
col = "blue")
[Link] 68
d) Scatter Plot:
A scatter plot in R programming is a graphical representation used to
visualize the relationship between two continuous numerical variables. Each
point on the plot represents a single observation, with its position
determined by the values of the two variables on the x and y axes.
[Link] 69
e) Line
A line graph is a type of chart that helps us visualize data through a series of
points connected by straight lines. It is commonly used to show changes
over time, making it easier to track trends and patterns. In a line graph, we
plot data points on the X and Y axes and connect them with lines, which
helps us understand how values move over time or across categories.
[Link] 70
Pie Chart
A pie chart is a circular statistical graphic, which is divided into slices to
illustrate numerical proportion.
Pie charts represents data visually as a fractional part of a whole, which can
be an effective communication tool.
[Link] 71
****
[Link] 72