0% found this document useful (0 votes)
2 views16 pages

Unit 1 Notes

The document provides an overview of data science, its history, and its foundational disciplines, including statistics and computer science. It discusses the significance of big data, various data types, and the roles of data scientists and engineers in analyzing and interpreting data. Additionally, it outlines the data science lifecycle, including business understanding, data preparation, modeling, evaluation, and deployment.

Uploaded by

paraphraser27
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views16 pages

Unit 1 Notes

The document provides an overview of data science, its history, and its foundational disciplines, including statistics and computer science. It discusses the significance of big data, various data types, and the roles of data scientists and engineers in analyzing and interpreting data. Additionally, it outlines the data science lifecycle, including business understanding, data preparation, modeling, evaluation, and deployment.

Uploaded by

paraphraser27
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Unit I - Introduction

Introduction to Data Science

Data science became the popular field it is today, all thanks to the rise of programming languages like
Python and techniques for collecting, analysing, and interpreting data.

The merging of the established discipline of statistics with a very nascent one—computer science—is
largely the narrative of how data scientists became fashionable. Only lately has the phrase "Data
Science" been coined to describe a new profession tasked with making sense of massive amounts of
data.

Making sense of data, on the other hand, has a long history and has been debated for years by scientists,
statisticians, librarians, computer scientists, and others. The history below shows how the phrase "Data
Science" has evolved over time, as well as attempts to define it and related terms.

As mentioned, Data science finds its foundation and beginning in Statistics. The advancement of Data
science and its evolution has been majorly facilitated by the arrival of Artificial Intelligence, Machine
learning, and the Internet of Things. Data science began to grow in other industries, including medicine,
engineering, and more, as a result of the influx of fresh data and corporations seeking new ways to
improve profit and make better judgments.

Big Data:

• Big Data are data sets so large or so complex that traditional methods of storing, accessing,
and analysing their breakdown are too expensive. However, there is a lot of potential value
hidden in this data, so organizations are eager to harness it to drive innovation and
competitive advantage.
• Big Data technologies and approaches are used to drive value out of data rich environments in
ways that traditional analytics tools and methods cannot.

Theories and techniques from many fields and disciplines are used to investigate and analyse a large
amount of data to help decision makers in many industries such as science, engineering, economics,
politics, finance, and education

– Computer Science
• Pattern recognition, visualization, data warehousing, High performance
computing, Databases, AI
– Mathematics
• Mathematical Modelling
– Statistics
• Statistical and Stochastic modelling, Probability.

Nominal data: Nominal Data is used to label variables without any order or quantitative value.

Ex. Color of hair (Blonde, red, Brown, Black, etc.)

Marital status (Single, Widowed, Married)

Nationality (Indian, German, American)

Gender (Male, Female, Others)

Eye Color (Black, Brown, etc.)

Ordinal data: Ordinal data have natural ordering where a number is present in some kind of order by
their position on the scale.

Ex. Letter grades in the exam (A, B, C, D, etc.)

Ranking of people in a competition (First, Second, Third, etc.)

Economic Status (High, Medium, and Low)

Education Level (Higher, Secondary, Primary)

Discrete data: The term discrete means distinct or separate. The discrete data contain the values that
fall under integers or whole numbers.

Ex. Total numbers of students present in a class

Cost of a cell phone

Numbers of employees in a company

The total number of players who participated in a competition

Days in a week

Continuous data: Continuous data are in the form of fractional numbers. It can be the version of an
android phone, the height of a person, the length of an object, etc. Continuous data represents
information that can be divided into smaller levels. The continuous variable can take any value within
a range.

Ex. Height of a person

Speed of a vehicle
“Time-taken” to finish the work

Wi-Fi Frequency

Market share price

Data Science is a combination of multiple disciplines that uses statistics, data analysis, and machine
learning to analyse data and to extract knowledge and insights from it.

Data Science is about data gathering, analysis and decision-making.

Data Science is about finding patterns in data, through analysis, and make future predictions.

By using Data Science, companies are able to make:

 Better decisions (should we choose A or B)


 Predictive analysis (what will happen next?)
 Pattern discoveries (find pattern, or maybe hidden information in the data)

The data is classified into four categories

• Nominal data.
• Ordinal data.
• Discrete data.
• Continuous data
What is Data Warehousing?

Data warehousing is the process of constructing and using a data warehouse. A data warehouse is
constructed by integrating data from multiple heterogeneous sources that support analytical reporting,
structured and/or ad hoc queries, and decision making. Data warehousing involves data cleaning, data
integration, and data consolidations.

Using Data Warehouse Information

There are decision support technologies that help utilize the data available in a data warehouse. These
technologies help executives to use the warehouse quickly and effectively. They can gather data,
analyse it, and take decisions based on the information present in the warehouse. The information
gathered in a warehouse can be used in any of the following domains −

 Tuning Production Strategies − the product strategies can be well tuned by repositioning the
products and managing the product portfolios by comparing the sales quarterly or yearly.
 Customer Analysis − Customer analysis is done by analysing the customer's buying
preferences, buying time, budget cycles, etc.
 Operations Analysis − Data warehousing also helps in customer relationship management, and
making environmental corrections. The information also allows us to analyse business
operations.

Functions of Data Warehouse Tools and Utilities

The following are the functions of data warehouse tools and utilities −

 Data Extraction − Involves gathering data from multiple heterogeneous sources.


 Data Cleaning − Involves finding and correcting the errors in data.
 Data Transformation − Involves converting the data from legacy format to warehouse format.
 Data Loading − Involves sorting, summarizing, consolidating, checking integrity, and building
indices and partitions.
 Refreshing − Involves updating from data sources to warehouse.

Data mining is the process of extracting knowledge or insights from large amounts of data using
various statistical and computational techniques. The data can be structured, semi-structured or
unstructured, and can be stored in various forms such as databases, data warehouses, and data lakes.

The primary goal of data mining is to discover hidden patterns and relationships in the data that can
be used to make informed decisions or predictions. This involves exploring the data using various
techniques such as clustering, classification, regression analysis, association rule mining, and
anomaly detection.

Data mining has a wide range of applications across various industries, including marketing, finance,
healthcare, and telecommunications. For example, in marketing, data mining can be used to identify
customer segments and target marketing campaigns, while in healthcare, it can be used to identify
risk factors for diseases and develop personalized treatment plans.

Types of Data Mining

Data mining can be performed on the following types of data:

Relational Database:

A relational database is a collection of multiple data sets formally organized by tables, records, and
columns from which data can be accessed in various ways without having to recognize the database
tables. Tables convey and share information, which facilitates data searchability, reporting, and
organization.

Data warehouses:

A Data Warehouse is the technology that collects the data from various sources within the organization
to provide meaningful business insights. The huge amount of data comes from multiple places such as
Marketing and Finance. The extracted data is utilized for analytical purposes and helps in decision-
making for a business organization. The data warehouse is designed for the analysis of data rather than
transaction processing.

Data Repositories:
The Data Repository generally refers to a destination for data storage. However, many IT professionals
utilize the term more clearly to refer to a specific kind of setup within an IT structure. For example, a
group of databases, where an organization has kept various kinds of information.

Object-Relational Database:

A combination of an object-oriented database model and relational database model is called an object-
relational model. It supports Classes, Objects, Inheritance, etc.

One of the primary objectives of the Object-relational data model is to close the gap between the
Relational database and the object-oriented model practices frequently utilized in many programming
languages, for example, C++, Java, C#, and so on.

Transactional Database:

A transactional database refers to a database management system (DBMS) that has the potential to undo
a database transaction if it is not performed appropriately. Even though this was a unique capability a
very long while back, today, most of the relational database systems support transactional database
activities.

Roles & Responsibilities of a Data Scientist

 Management: The Data Scientist plays an insignificant managerial role where he supports the
construction of the base of futuristic and technical abilities within the Data and Analytics field in
order to assist various planned and continuing data analytics projects.

 Analytics: The Data Scientist represents a scientific role where he plans, implements, and
assesses high-level statistical models and strategies for application in the business’s most
complex issues. The Data Scientist develops econometric and statistical models for various
problems including projections, classification, clustering, pattern analysis, sampling, simulations,
and so forth.

 Strategy/Design: The Data Scientist performs a vital role in the advancement of innovative
strategies to understand the business’s consumer trends and management as well as ways to solve
difficult business problems, for instance, the optimization of product fulfilment and entire profit.
 Collaboration: The role of the Data Scientist is not a solitary role and in this position, he
collaborates with superior data scientists to communicate obstacles and findings to relevant
stakeholders in an effort to enhance drive business performance and decision-making.

 Knowledge: The Data Scientist also takes leadership to explore different technologies and tools
with the vision of creating innovative data-driven insights for the business at the most agile pace
feasible. In this situation, the Data Scientist also uses initiative in assessing and utilizing new and
enhanced data science methods for the business, which he delivers to senior management of
approval.
 Other Duties: A Data Scientist also performs related tasks and tasks as assigned by the Senior
Data Scientist, Head of Data Science, Chief Data Officer, or the Employer.

The lifecycle of Data Science

1. Business Understanding: The complete cycle revolves around the enterprise goal. What
will you resolve if you do not longer have a specific problem? It is extraordinarily essential
to apprehend the commercial enterprise goal sincerely due to the fact that will be your
ultimate aim of the analysis. After desirable perception only we can set the precise aim of
evaluation that is in sync with the enterprise objective. You need to understand if the
customer desires to minimize savings loss, or if they prefer to predict the rate of a commodity,
etc.
2. Data Understanding: After enterprise understanding, the subsequent step is data
understanding. This includes a series of all the reachable data. Here you need to intently
work with the commercial enterprise group as they are certainly conscious of what
information is present, what facts should be used for this commercial enterprise problem,
and different information. This step includes describing the data, their structure, their
relevance, their records type. Explore the information using graphical plots. Basically,
extracting any data that you can get about the information through simply exploring the data.
3. Preparation of Data: Next comes the data preparation stage. This consists of steps like
choosing the applicable data, integrating the data by means of merging the data sets, cleaning
it, treating the lacking values through either eliminating them or imputing them, treating
inaccurate data through eliminating them, additionally test for outliers the use of box plots
and cope with them. Constructing new data, derive new elements from present ones. Format
the data into the preferred structure, eliminate undesirable columns and features. Data
preparation is the most time-consuming but arguably the most essential step in the complete
existence cycle. Your model will be as accurate as your data.
4. Exploratory Data Analysis: This step includes getting some concept about the answer and
elements affecting it, earlier than constructing the real model. Distribution of data inside
distinctive variables of a character is explored graphically the usage of bar-graphs, Relations
between distinct aspects are captured via graphical representations like scatter plots and
warmth maps. Many data visualization strategies are considerably used to discover each and
every characteristic individually and by means of combining them with different features.
5. Data Modelling: Data modelling is the coronary heart of data analysis. A model takes the
organized data as input and gives the preferred output. This step consists of selecting the
suitable kind of model, whether the problem is a classification problem, or a regression
problem or a clustering problem. After deciding on the model family, amongst the number
of algorithms amongst that family, we need to cautiously pick out the algorithms to put into
effect and enforce them. We need to tune the hyper parameters of every model to obtain the
preferred performance. We additionally need to make positive there is the right stability
between overall performance and generalizability. We do no longer desire the model to study
the data and operate poorly on new data.
6. Model Evaluation: Here the model is evaluated for checking if it is geared up to be
deployed. The model is examined on an unseen data, evaluated on a cautiously thought out
set of assessment metrics. We additionally need to make positive that the model conforms to
reality. If we do not acquire a quality end result in the evaluation, we have to re-iterate the
complete modelling procedure until the preferred stage of metrics is achieved. Any data
science solution, a machine learning model, simply like a human, must evolve, must be
capable to enhance itself with new data, adapt to a new evaluation metric. We can construct
more than one model for a certain phenomenon, however, a lot of them may additionally be
imperfect. The model assessment helps us select and construct an ideal model.
7. Model Deployment: The model after a rigorous assessment is at the end deployed in the
preferred structure and channel. This is the last step in the data science life cycle. Each step
in the data science life cycle defined above must be laboured upon carefully. If any step is
performed improperly, and hence, have an effect on the subsequent step and the complete
effort goes to waste. For example, if data is no longer accumulated properly, you’ll lose
records and you will no longer be constructing an ideal model. If information is not cleaned
properly, the model will no longer work. If the model is not evaluated properly, it will fail
in the actual world. Right from Business perception to model deployment, every step has to
be given appropriate attention, time, and effort.

Roles in Data Science

 Data Analyst
 Data Engineers
 Database Administrator
 Machine Learning Engineer
 Data Scientist
 Data Architect
 Statistician
 Business Analyst
 Data and Analytics Manager

1. Data Analyst

Data analysts are responsible for a variety of tasks including visualisation, munging, and processing of
massive amounts of data. They also have to perform queries on the databases from time to time. One of
the most important skills of a data analyst is optimization. This is because they have to create and
modify algorithms that can be used to cull information from some of the biggest databases without
corrupting the data.

Few Important Roles and Responsibilities of a Data Analyst include:


 Extracting data from primary and secondary sources using automated tools
 Developing and maintaining databases
 Performing data analysis and making reports with recommendations
 Analysing data and forecasting trends that impact the organization/project
 Working with other team members to improve data collection and quality processes

2. Data Engineers

Data engineers build and test scalable Big Data ecosystems for the businesses so that the data
scientists can run their algorithms on the data systems that are stable and highly optimized. Data
engineers also update the existing systems with newer or upgraded versions of the current technologies
to improve the efficiency of the databases.

Few Important Roles and Responsibilities of a Data Engineer include:

 Design and maintain data management systems


 Data collection/acquisition and management
 Conducting primary and secondary research
 Finding hidden patterns and forecasting trends using data
 Collaborating with other teams to perceive organizational goals
 Make reports and update stakeholders based on analytics

3. Database Administrator

The job profile of a database administrator is pretty much self-explanatory- they are responsible for
the proper functioning of all the databases of an enterprise and grant or revoke its services to the
employees of the company depending on their requirements. They are also responsible for database
backups and recoveries.

Few Important Roles and Responsibilities of a Database Administrator include:

 Working on database software to store and manage data


 Working on database design and development
 Implementing security measures for database
 Preparing reports, documentation, and operating manuals
 Data archiving
 Working closely with programmers, project managers, and other team members

4. Machine Learning Engineer

Machine learning engineers are in high demand today. However, the job profile comes with its
challenges. Apart from having in-depth knowledge of some of the most powerful technologies such
as SQL, REST APIs, etc. machine learning engineers are also expected to perform A/B testing, build
data pipelines, and implement common machine learning algorithms such as classification, clustering,
etc.

Few Important Roles and Responsibilities of a Machine Learning Engineer include:

 Designing and developing Machine Learning systems


 Researching Machine Learning Algorithms
 Testing Machine Learning systems
 Developing apps/products basis client requirements
 Extending existing Machine Learning frameworks and libraries
 Exploring and visualizing data for a better understanding
 Training and retraining systems
 Know the importance of statistics in machine learning

5. Data Scientist

Data scientists have to understand the challenges of business and offer the best solutions using data
analysis and data processing. For instance, they are expected to perform predictive analysis and run a
fine-toothed comb through an “unstructured/disorganized” data to offer actionable insights. They can
also do this by identifying trends and patterns that can help the companies in making better decisions.

Few Important Roles and Responsibilities of a Data Scientist include:

 Identifying data collection sources for business needs


 Processing, cleansing, and integrating data
 Automation data collection and management process
 Using Data Science techniques/tools to improve processes
 Analysing large amounts of data to forecast trends and provide reports with recommendations
 Collaborating with business, engineering, and product teams

6. Data Architect

A data architect creates the blueprints for data management so that the databases can be easily
integrated, centralized, and protected with the best security measures. They also ensure that the data
engineers have the best tools and systems to work with.

Few Important Roles and Responsibilities of a Data Architect include:

 Developing and implementing overall data strategy in line with business/organization


 Identifying data collection sources in line with data strategy
 Collaborating with cross-functional teams and stakeholders for smooth functioning of database
systems
 Planning and managing end-to-end data architecture
 Maintaining database systems/architecture considering efficiency and security
 Regular auditing of data management system performance and making changes to improve
systems accordingly.

7. Statistician

A statistician, as the name suggests, has a sound understanding of statistical theories and data
organization. Not only do they extract and offer valuable insights from the data clusters, but they also
help create new methodologies for the engineers to apply.

Few Important Roles and Responsibilities of a Statistician include:

 Collecting, analysing, and interpreting data


 Analysing data, assessing results, and predicting trends/relationships using statistical
methodologies/tools
 Designing data collection processes
 Communicating findings to stakeholders
 Advising/consulting on organizational and business strategy basis data
 Coordinating with cross-functional teams
8. Business Analyst

The role of business analysts is slightly different than other data science jobs. While they do have a
good understanding of how data-oriented technologies work and how to handle large volumes of data,
they also separate the high-value data from the low-value data. In other words, they identify how the Big
Data can be linked to actionable business insights for business growth.

Few Important Roles and Responsibilities of a Business Analyst include:

 Understanding the business of the organization


 Conducting detailed business analysis – outlining problems, opportunities, and solutions
 Working on improving existing business processes
 Analysing, designing, and implementing new technology and systems
 Budgeting and forecasting
 Pricing analysis

9. Data and Analytics Manager

A data and analytics manager oversees the data science operations and assigns the duties to their team
according to skills and expertise. Their strengths should include technologies like SAS, R, SQL, etc.
and of course management.

Few Important Roles and Responsibilities of a Data and Analytics Manager include:

 Developing data analysis strategies


 Researching and implementing analytics solutions
 Leading and managing a team of data analysts
 Overseeing all data analytics operations to ensure quality
 Building systems and processes to transform raw data into actionable business insights
 Staying up to date on industry news and trends
Applications of Data Science

1. In Search Engines
The most useful application of Data Science is Search Engines. As we know when we want to search
for something on the internet, we mostly used Search engines like Google, Yahoo, Safari, Firefox,
etc. So Data Science is used to get Searches faster.

For Example, When we search something suppose “Data Structure and algorithm courses ” then at
that time on the Internet Explorer we get the first link of GeeksforGeeks Courses. This happens
because the GeeksforGeeks website is visited most in order to get information regarding Data
Structure courses and Computer related subjects. So this analysis is Done using Data Science, and
we get the Topmost visited Web Links.

2. In Transport
Data Science also entered into the Transport field like Driverless Cars. With the help of Driverless
Cars, it is easy to reduce the number of Accidents.

For Example, In Driverless Cars the training data is fed into the algorithm and with the help of Data
Science techniques, the Data is analyzed like what is the speed limit in Highway, Busy Streets,
Narrow Roads, etc. And how to handle different situations while driving etc.

3. In Finance
Data Science plays a key role in Financial Industries. Financial Industries always have an issue of
fraud and risk of losses. Thus, Financial Industries needs to automate risk of loss analysis in order to
carry out strategic decisions for the company. Also, Financial Industries uses Data Science Analytics
tools in order to predict the future. It allows the companies to predict customer lifetime value and
their stock market moves.

For Example, In Stock Market, Data Science is the main part. In the Stock Market, Data Science is
used to examine past behavior with past data and their goal is to examine the future outcome. Data is
analyzed in such a way that it makes it possible to predict future stock prices over a set timetable.

4. In E-Commerce
E-Commerce Websites like Amazon, Flipkart, etc. uses data Science to make a better user experience
with personalized recommendations.

For Example, When we search for something on the E-commerce websites we get suggestions
similar to choices according to our past data and also we get recommendations according to most buy
the product, most rated, most searched, etc. This is all done with the help of Data Science.
5. In Health Care
In the Healthcare Industry data science act as a boon. Data Science is used for:

 Detecting Tumor.
 Drug discoveries.
 Medical Image Analysis.
 Virtual Medical Bots.
 Genetics and Genomics.
 Predictive Modeling for Diagnosis etc.
6. In E-Commerce
E-Commerce Websites like Amazon, Flipkart, etc. uses data Science to make a better user experience
with personalized recommendations.

For Example, when we search for something on the E-commerce websites we get suggestions similar
to choices according to our past data and also we get recommendations according to most buy the
product, most rated, most searched, etc. This is all done with the help of Data Science.

7. In Health Care
In the Healthcare Industry data science act as a boon. Data Science is used for:

 Detecting Tumor.
 Drug discoveries.
 Medical Image Analysis.
 Virtual Medical Bots.
 Genetics and Genomics.
 Predictive Modeling for Diagnosis etc.

8. Image Recognition
Currently, Data Science is also used in Image Recognition. For Example, When we upload our image
with our friend on Facebook, Facebook gives suggestions Tagging who is in the picture. This is done
with the help of machine learning and Data Science. When an Image is Recognized, the data analysis
is done on one’s Facebook friends and after analysis, if the faces which are present in the picture
matched with someone else profile then Facebook suggests us auto-tagging.

9. Targeting Recommendation
Targeting Recommendation is the most important application of Data Science. Whatever the user
searches on the Internet, he/she will see numerous posts everywhere. This can be explained properly
with an example: Suppose I want a mobile phone, so I just Google search it and after that, I changed
my mind to buy offline. Data Science helps those companies who are paying for Advertisements for
their mobile. So everywhere on the internet in the social media, in the websites, in the apps
everywhere I will see the recommendation of that mobile phone which I searched for. So this will
force me to buy online.

10. Airline Routing Planning


With the help of Data Science, Airline Sector is also growing like with the help of it, it becomes easy
to predict flight delays. It also helps to decide whether to directly land into the destination or take a
halt in between like a flight can have a direct route from Delhi to the U.S.A or it can halt in between
after that reach at the destination.

11. Data Science in Gaming


In most of the games where a user will play with an opponent i.e. a Computer Opponent, data science
concepts are used with machine learning where with the help of past data the Computer will improve
its performance. There are many games like Chess, EA Sports, etc. will use Data Science concepts.

12. Medicine and Drug Development


The process of creating medicine is very difficult and time-consuming and has to be done with full
disciplined because it is a matter of Someone’s life. Without Data Science, it takes lots of time,
resources, and finance or developing new Medicine or drug but with the help of Data Science, it
becomes easy because the prediction of success rate can be easily determined based on biological
data or factors. The algorithms based on data science will forecast how this will react to the human
body without lab experiments.

13. In Delivery Logistics


Various Logistics companies like DHL, FedEx, etc. make use of Data Science. Data Science helps
these companies to find the best route for the Shipment of their Products, the best time suited for
delivery, the best mode of transport to reach the destination, etc.

14. Autocomplete
AutoComplete feature is an important part of Data Science where the user will get the facility to just
type a few letters or words, and he will get the feature of auto-completing the line. In Google Mail,
when we are writing formal mail to someone so at that time data science concept of Autocomplete
feature is used where he/she is an efficient choice to auto-complete the whole line.

You might also like