DATA SCIENCE AND VISUALIZATION
21CS644
Module-1: Introduction to Data Science
2
What Is Data Science?
1. Data science is the study of data to extract meaningful insights for business
2. It is a multidisciplinary approach that combines principles and practices from
the fields of mathematics, statistics, artificial intelligence, and computer
engineering to analyze large amounts of data
3. This analysis helps data scientists to ask and answer questions like what
happened, why it happened, what will happen, and what can be done with
the results. 3
Why is data science important?
1. Data science is important because it combines tools, methods, and
technology to generate meaning from data
2. Modern organizations are inundated with data; there is a proliferation of
devices that can automatically collect and store information
3. Online systems and payment portals capture more data in the fields of e-
commerce, medicine, finance, and every other aspect of human life. We
have text, audio, video, and image data available in vast quantities 4
History of data science
1. While the term data science is not new, the meanings and connotations have
changed over time
2. The word first appeared in the ’60s as an alternative name for statistics
3. In the late ’90s, computer science professionals formalized the term
4. A proposed definition for data science saw it as a separate field with three
aspects: data design, collection, and analysis. It still took another decade
5
for the term to be used outside of academia
Future of data science
1. Artificial intelligence and machine learning innovations have made data
processing faster and more efficient.
2. Industry demand has created an ecosystem of courses, degrees, and job
positions within the field of data science.
3. Because of the cross-functional skillset and expertise required, data science
shows strong projected growth over the coming decades.
6
What is data science used for?
Data science is used to study data in four main ways:
[Link] analysis
a. Descriptive analysis examines data to gain insights into what
happened or what is happening in the data environment.
b. It is characterized by data visualizations such as pie charts, bar
charts, line graphs, tables, or generated narratives.
c. For example, a flight booking service may record data like the
number of tickets booked each day.
d. Descriptive analysis will reveal booking spikes, booking slumps, and
7
high-performing months for this service
What is data science used for?
Data science is used to study data in four main ways:
[Link] analysis
a. Diagnostic analysis is a deep-dive or detailed data examination to understand why
something happened.
b. It is characterized by techniques such as drill-down, data discovery, data mining, and
correlations.
c. Multiple data operations and transformations may be performed on a given data set to
discover unique patterns in each of these techniques.
d. For example, the flight service might drill down on a particularly high-performing month
to better understand the booking spike. This may lead to the discovery that many
customers visit a particular city to attend a monthly sporting event
8
What is data science used for?
Data science is used to study data in four main ways:
[Link] analysis
a. Predictive analysis uses historical data to make accurate forecasts about data patterns
that may occur in the future.
b. It is characterized by techniques such as machine learning, forecasting, pattern
matching, and predictive modeling.
c. In each of these techniques, computers are trained to reverse engineer causality
connections in the data.
d. For example, the flight service team might use data science to predict flight booking
patterns for the coming year at the start of each year.
e. The computer program or algorithm may look at past data and predict booking spikes for certain
destinations in May. Having anticipated their customer’s future travel requirements, the company could
9
start targeted advertising for those cities from February
What is data science used for?
Data science is used to study data in four main ways:
[Link] analysis
a. Prescriptive analytics takes predictive data to the next level.
b. It not only predicts what is likely to happen but also suggests an optimum
response to that outcome.
c. It can analyze the potential implications of different choices and recommend
the best course of action.
d. It uses graph analysis, simulation, complex event processing, neural
networks, and recommendation engines from machine learning
10
What is data science used for?
Data science is used to study data in four main ways:
[Link] analysis
e. Back to the flight booking example, prescriptive analysis could look at
historical marketing campaigns to maximize the advantage of the upcoming
booking spike.
f. A data scientist could project booking outcomes for different levels of
marketing spend on various marketing channels.
g. These data forecasts would give the flight booking company greater
confidence in their marketing decisions.
11
Introduction to Big Data
12
What is Big Data?
1. The definition of big data is data that contains greater variety, arriving in
increasing volumes and with more velocity.
2. This is also known as the three Vs.
3. Put simply, big data is larger, more complex data sets, especially from new
data sources.
4. These data sets are so voluminous that traditional data processing software
just can’t manage them.
5. But these massive volumes of data can be used to address business
problems you wouldn’t have been able to tackle before. 13
History of Big Data?
1. Around 2005, people began to realize just how much data users generated
through Facebook, YouTube, and other online services.
2. Hadoop (an open-source framework created specifically to store and analyze
big data sets) was developed that same year.
3. NoSQL also began to gain popularity during this time.
4. The development of open-source frameworks, such as Hadoop (and more
recently, Spark) was essential for the growth of big data because they make
big data easier to work with and cheaper to store.
5. In the years since then, the volume of big data has skyrocketed. 14
History of Big Data?
6. With the advent of the Internet of Things (IoT), more objects and devices are
connected to the internet, gathering data on customer usage patterns and product
performance.
7. The emergence of machine learning has produced still more data.
8. While big data has come far, its usefulness is only just beginning.
9. Cloud computing has expanded big data possibilities even further.
10. The cloud offers truly elastic scalability, where developers can simply spin up ad hoc
clusters to test a subset of data.
11. And graph databases are becoming increasingly important as well, with their ability to
display massive amounts of data in a way that makes analytics fast and
15
Big data use cases
1. Product development: Companies like Netflix and Procter & Gamble use big data to
anticipate customer demand.
2. Predictive maintenance: Factors that can predict mechanical failures may be deeply
buried in structured data, such as the year, make, and model of equipment, as well as
in unstructured data that covers millions of log entries, sensor data, error messages,
and engine temperature.
3. Customer experience: The race for customers is on. A clearer view of customer
experience is more possible now than ever before.
4. Machine learning: We are now able to teach machines instead of program them. The
availability of big data to train machine learning models makes that possible. 16
Big data challenges
1. It’s not enough to just store the data.
2. Data must be used to be valuable and that depends on curation.
3. Clean data, or data that’s relevant to the client and organized in a way that enables
meaningful analysis, requires a lot of work.
4. Data scientists spend 50 to 80 percent of their time curating and preparing data
before it can actually be used. Finally, big data technology is changing at a rapid
pace.
5. A few years ago, Apache Hadoop was the popular technology used to handle big
data. Then Apache Spark was introduced in 2014. Today, a combination of the two
frameworks appears to be the best approach. Keeping up with big data technology is
17
Datafication
18
Datafication
1. Datafication is a process of “taking all aspects of life and turning them into data.” As
examples, they mention that Google’s augmented-reality glasses datafy the gaze.
2. Twitter datafies stray thoughts.
3. LinkedIn datafies professional networks.
4. Its importance is with respect to people’s intentions about sharing their own data.
5. We are being datafied, or rather our actions are, and when we “like” someone or
something online, we are intending to be datafied, or at least we should expect to be.
6. But when we merely browse the Web, we are unintentionally, or at least passively,
being datafied through cookies that we might or might not be aware of.
19
Datafication
1. And when we walk around in a store, or even on the street, we are being datafied in a
completely unintentional way, via sensors, cameras, or Google glasses.
2. This spectrum of intentionality ranges from us gleefully taking part in a social media
experiment we are proud of, to all-out surveillance and stalking. But it’s all
datafication.
3. Once we datafy things, we can transform their purpose and turn the information into
new forms of value.
20
The Current Landscape
1. Data science, as it’s practiced, is a blend of Red-Bull-fueled hacking and espresso-
inspired statistics.
2. But data science is not merely hacking—because when hackers finish debugging
their Bash one-liners and Pig scripts, few of them care about non-Euclidean distance
metrics. In general, the phrase underscores the distinction between hacking and data
science, emphasizing that while they may share some technical elements, they serve
different purposes and require different approaches and priorities.
21
The Current Landscape
1. Data science, as it’s practiced, is a blend of Red-Bull-fueled hacking and espresso-
inspired statistics.
2. But data science is not merely hacking—because when hackers finish debugging
their Bash one-liners and Pig scripts, few of them care about non-Euclidean distance
metrics. In general, the phrase underscores the distinction between hacking and data
science, emphasizing that while they may share some technical elements, they serve
different purposes and require different approaches and priorities.
22
The Current Landscape
1. Data science is the civil engineering of data. Its acolytes possess a practical
knowledge of tools and materials, coupled with a theoretical understanding of what’s
possible.
2. Drew Conway’s Venn diagram of data science
from 2010, shown in Figure 1-1.
23
The Current Landscape
1. Conway's Venn diagram highlights the interdisciplinary nature of data science,
emphasizing that expertise in multiple domains is necessary to excel in the field.
2. It also underscores the importance of collaboration and teamwork, as individuals with
diverse backgrounds and skill sets often work together to address complex data-
related problems and drive innovation in various industries and domains.
3. Overall, the diagram serves as a useful framework for understanding the breadth and
depth of skills required to succeed in the field of data science.
24
A Data Science Profile
1. Skill levels in the following domains:
Computer science; Math; Statistics;
Machine learning; Domain expertise;
Communication and presentation skills
and Data visualization
25
1. A data science team works best when different
skills (profiles) are represented across different
people, because nobody is good at everything
26
Statistical Inference
27
Statistical Thinking in the Age of Big Data
1. When developing the skill set as a data scientist, certain foundational pieces need to
be in place first—statistics, linear algebra, some programming.
2. Even once those pieces are acquired, part of the challenge is that you will be
developing several skill sets in parallel simultaneously—data preparation and
munging, modeling, coding, visualization, and communication—that are
interdependent.
28
Statistical Inference
1. Imagine spending 24 hours looking out the window, and for every minute, counting
and recording the number of people who pass by
2. The point here is that the processes in our lives are actually data-generating
processes
3. Data represents the traces of the real-world processes, and exactly which traces we
gather are decided by our data collection or sampling method.
4. After separating the process from the data collection, we can see clearly that there
are two sources of randomness and uncertainty.
5. Namely, the randomness and uncertainty underlying the process itself, and the
uncertainty associated with your underlying data collection methods 29
Statistical Inference
1. We need a new idea, and that’s to simplify those captured traces into something
more comprehensible, to something that somehow captures it all in a much more
concise way, and that something could be mathematical models or functions of the
data, known as statistical estimators.
2. This overall process of going from the world to the data, and then from the data back
to the world, is the field of statistical inference.
3. More precisely, statistical inference is the discipline that concerns itself with the
development of procedures, methods, and theorems that allow us to extract meaning
and information from data that has been generated by stochastic (random)
processes. 30
Populations and Samples
31
Populations and Samples
1. In classical statistical literature, a distinction is made between the population and the
sample
2. In statistical inference population isn’t used to simply describe only people. It could
be any set of objects or units, such as tweets or photographs or stars
3. If we could measure the characteristics or extract characteristics of all those objects,
we’d have a complete set of observations, and the convention is to use N to
represent the total number of observations in the population.
32
Populations and Samples
1. Suppose your population was all emails sent last year by employees at a huge
corporation, BigCorp.
2. Then a single observation could be a list of things: the sender’s name, the list of
recipients, date sent, text of email, number of characters in the email, number of
sentences in the email, number of verbs in the email, and the length of time until first
reply.
33
Populations and Samples
1. When we take a sample, we take a subset of the units of size n in order to examine
the observations to draw conclusions and make inferences about the population.
2. There are different ways you might go about getting this subset of data, and you want
to be aware of this sampling mechanism because it can introduce biases into the
data, and distort it, so that the subset is not a “mini-me” shrunk-down version of the
population.
3. Once that happens, any conclusions you draw will simply be wrong and distorted
34
Populations and Samples
1. In the BigCorp email example, you could make a list of all the employees and select
1/10th of those people at random and take all the email they ever sent, and that
would be your sample.
2. Alternatively, you could sample 1/10th of all email sent each day at random, and that
would be your sample.
3. Both these methods are reasonable, and both methods yield the same sample size.
35
Populations and Samples of Big Data
36
Populations and Samples of Big Data
1. Sampling solves some engineering challenges
1. In the current popular discussion of Big Data, the focus on enterprise solutions
such as Hadoop to handle engineering and computational challenges caused by
too much data overlooks sampling as a legitimate solution.
2. At Google, for example, software engineers, data scientists, and statisticians
sample all the time.
2. Sampling Let’s rethink what the population and the sample are in various contexts
37
Populations and Samples of Big Data
1. In statistics we often model the relationship between a population and a sample with
an underlying mathematical process.
2. So we make simplifying assumptions about the underlying truth, the mathematical
structure, and shape of the underlying generative process that created the data.
3. We observe only one particular realization of that generative process, which is that
sample.
38
New kinds of data
1. A strong data scientist needs to be versatile and comfortable with dealing a variety of
types of data, including:
2. Traditional: numerical, categorical, or binary
3. Text: emails, tweets, New York Times articles • Records: user-level data,
timestamped event data, jsonformatted log files
4. Geo-based location data: briefly touched on in this chapter with NYC housing data
5. Network
6. Sensor data
7. Images 39
Statistical modeling
1. Exploratory data analysis (EDA) entails making plots and building intuition for your
particular dataset. EDA helps out a lot, as well as trial and error and iteration.
2. There is a trade-off in modeling between simple and accurate. Simple models may be
easier to interpret and understand. Oftentimes the crude, simple model gets 90% of
the way there and only takes a few hours to build and fit, whereas getting a more
complex model might take months and only get to 92%
40
Statistical modeling
1. Probability distributions are the foundation of statistical models
2. Before computers, scientists observed real-world phenomenon, took measurements,
and noticed that certain mathematical shapes kept reappearing.
3. The classical example is the height of humans, following a normal distribution—a
bell-shaped curve, also called a Gaussian distribution, named after Gauss.
4. Natural processes tend to generate measurements whose empirical shape could be
approximated by mathematical functions with a few parameters that could be
estimated from the data. 41
Statistical modeling
1. Natural processes tend to generate measurements whose empirical shape could be
approximated by mathematical functions with a few parameters that could be
estimated from the data.
2. Not all processes generate data that looks like a named distribution, but many do.
3. They are to be interpreted as assigning a probability to a subset of possible
outcomes, and have corresponding functions.
42
Statistical modeling
1. Figure 2-1 as an illustration of the various
common shapes
43
Fitting a model
1. Fitting a model means that you estimate the parameters of the model using the
observed data.
2. You are using your data as evidence to help approximate the real-world mathematical
process that generated the data.
3. Fitting the model often involves optimization methods and algorithms, such as
maximum likelihood estimation, to help get the parameters
4. In fact, when you estimate the parameters, they are actually estimators, meaning
they themselves are functions of the data. 44
Fitting a model
1. Fitting the model is when you start actually coding: your code will read in the data,
and specify the functional form that you wrote down on the piece of paper.
2. Then R or Python will use built-in optimization methods to give the most likely values
of the parameters given the data
3. Dig around in the optimization methods. Initially you should have an understanding
that optimization is taking place and how it works
45
Overfitting
1. Overfitting is the term used to mean that you used a dataset to estimate the
parameters of your model, but your model isn’t that good at capturing reality beyond
your sampled data.
2. You might know this because you have tried to use it to predict labels for another set
of data that you didn’t use to fit the model, and it doesn’t do a good job, as measured
by an evaluation metric such as accuracy.
46
47