Unit 2 Datascience
Unit 2 Datascience
Data science deals with explorations and Data Analysis makes use
Goals
new innovations. of existing resources.
Steps to explore, preprocess, and condition data prior to modeling and analysis.
It requires the presence of an analytic sandbox, the team execute, load, and transform,
to get data into the sandbox.
Data preparation tasks are likely to be performed multiple times and not in predefined
order.
Several tools commonly used for this phase are – Hadoop, Alpine Miner, Open Refine,
etc.
Phase 3: Model Planning –
Team explores data to learn about relationships between variables and subsequently,
selects key variables and the most suitable models.
In this phase, data science team develop data sets for training, testing, and production
purposes.
Team builds and executes models based on the work done in the model planning phase.
Several tools commonly used for this phase are – Matlab, STASTICA.
Phase 4: Model Building –
The team communicates benefits of project more broadly and sets up pilot project to
deploy work in controlled way before broadening the work to full enterprise of users.
This approach enables team to learn about performance and related constraints of the
model in production environment on small scale  , and make adjustments before
full deployment.
The team delivers final reports, briefings, codes.
Free or open source tools – Octave, WEKA, SQL, MADlib.
Raw data aggregated is data that is not oriented. It requires a thoughtful understanding as
well as the appropriate questions in order to create sense out of it. Many insights fail to
analyse data completely and become difficult for the stakeholders' comprehension. Therefore,
it becomes necessary for a data analyst to define and understand data with the right set of
initial questions and a standardized workflow for the different types of analysis he needs to
perform.
The following words are from Jeff Leek's fascinating book "The Elements of Data Analytic
Style," which broadly categorizes various analysis phases based on the type of question and
the outcome expected to be achieved for the particular business need.
Descriptive Data Analysis
The name suggests that this kind of analysis offers basic "descriptions" or summaries about
the raw data set accumulated and the observations added to the same.
They can be both visual and quantitative, and the data can be depicted using statistics and
simple graphs. This summary does not require any further analysis and is utilized as a
summary to make sense of the information.
The data could be split into various categories such as numbers, gender, residency age, race,
and so on. The information summarizes or groups the data into a fixed set that describes all
students and the specific information. It doesn't suggest anything and only provides specifics.
Thus, it is a type of descriptive analytics.
Analysis of descriptive data output that is further studied for discoveries patterns, trends,
correlations, or inter-relations among different areas of the data in order to develop an
interpretation, an idea, or hypotheses. This is the foundation of Exploratory Data Analysis
(EDA).
In essence, it's expanding over the description data sets and trying to provide a
comprehensive overview of the data. According to Dianne Cook, as well as Deborah F.
Swayne rightly refer to in their book, "(EDA is) a 'play-in-the-sand' to allow us to find the
unexpected and come to some understanding of our data."
The main focus isn't always the result of the problem statement; rather, to look at the various
elements of data in the first place in order to more intimately.
Example: A typical EDA application studies the behaviour of traffic patterns in cities around
the world. Although the data gathered may vary in terms of its nature, various surprising
discoveries may be discovered like the frequency of accidents that occur at traffic signals, the
amount of pollution that is produced on a daily basis because of exhaust emissions from
vehicles, and even the rates of traffic congestion in a week. The outcome of the real issue
isn't always determined by these findings. The information gathered alongside other data may
be helpful to determine the result.
Example: Calculating the mean of marks earned by students taking an exam against the
difficulty index for 100 students can give valuable information on the students of 100.
This data can assist in understanding the quality of the connection between these two
dimensions when studying student performance on exams. Although it's impossible to know
the reasons for these relationships, there is a way to determine the significance of a certain
connection in determining inferential results.
The predictive analysis predicts the outcomes that could be expected from a small subset of
data from the initial population set. This method of predicting new information is mostly
built on quantifiable metrics from the existing data set.
Predictive analysis is not able to quantify the relationship between two dimensions as the
inferential statistical method. Rather it uses probabilities that they share to predict possible
outcomes in the future.
Example: Examining the influence and popularity of the nominees running for election to
determine the outcome of that election.
In this case, we can determine the likelihood of the success of the candidate based on data
about issues he discusses as well as his conservative and liberal views, information on his
popularity in the state of his residence and so on. While we can estimate a potential outcome
based on these data, however, we can't predict the outcome accurately.
Example: A randomized clinical trial to determine whether faecal transfer decreases the
incidence of infections caused by Clostridium di-facile.
Patients in this research were randomly assigned to receive a faecal transfer along with
standard care or regular treatment. Based on the results, the researchers found an
unambiguous relationship between the outcomes of infections and transplants. Therefore, the
study of the causality of patients produced an exact average outcome from raw data.
Although causal data provides an accurate average result, the aim isn't just to comprehend
that there's an impact of the inferences derived from data but also to understand how the
effect is affecting the outcome.
An example: Mechanistic analysis that examines the way in which wing design influences
the flow of air around a wing, which results in less drag. In the absence of any engineering
expertise, mechanical analysis of data is extremely difficult and is rarely done.
Conclusion
As we can see, harnessing big-data analytics can bring huge benefits to companies, providing
the context of data to tell an even more comprehensive story. By converting complex data
sets into actionable intelligence, stakeholders can make better business decisions. If we know
how to make big data accessible to our clients, the value of our service is now ten times
greater.
ADVANCED DATA ANALYTICS
Key Components
Data Collection:
Definition: The process of gathering information from various sources.
Sources: Databases, APIs, web scraping, sensors, social media, transaction logs.
Importance: High-quality data collection is critical for accurate analysis.
Data Cleaning:
Definition: The process of detecting and correcting (or removing) corrupt or inaccurate
records from a dataset.
Techniques: Handling missing values, correcting errors, removing duplicates, standardizing
formats.
Importance: Ensures the reliability and validity of the data analysis.
Data Integration:
Definition: Combining data from different sources to create a unified view.
Techniques: ETL (Extract, Transform, Load), data warehousing, APIs.
Importance: Provides a comprehensive dataset that reflects all relevant information.
Data Transformation:
Definition: The process of converting data into a format suitable for analysis.
Techniques: Normalization, aggregation, feature engineering.
Importance: Enhances data quality and prepares it for effective analysis.
Data Storage:
Definition: Storing data in databases, data warehouses, or data lakes.
Technologies: SQL databases (MySQL, PostgreSQL), NoSQL databases (MongoDB,
Cassandra), data lakes (Amazon S3).
Importance: Efficient storage solutions are crucial for managing large datasets and ensuring
fast access.
Techniques
1. Descriptive Analytics:
Purpose: To describe the main features of a dataset.
Methods:
Summary Statistics: Measures like mean, median, mode, variance.
Data Visualization: Tools like bar charts, histograms, scatter plots to visualize data.
Importance: Helps understand the basic characteristics and trends in data.
2. Predictive Analytics:
Purpose: To make predictions about future outcomes based on historical data.
Methods:
Regression Analysis: Linear and logistic regression to predict continuous or binary
outcomes.
Time Series Analysis: Analyzing time-ordered data to forecast future values.
Machine Learning Models: Algorithms like decision trees, random forests, and neural
networks.
Importance: Provides foresight into future trends and behaviors, aiding in proactive decision-
making.
3. Prescriptive Analytics:
Purpose: To suggest actions that can optimize outcomes.
Methods:
Optimization: Techniques like linear programming to find the best solution under given
constraints.
Simulation: Using models to simulate various scenarios and their potential outcomes.
Importance: Helps determine the best course of action among multiple alternatives.
4. Marketing:
Objective: To understand customer behavior and optimize marketing efforts.
Applications: Customer segmentation, targeted advertising, sentiment analysis.
Examples: Identifying potential customer segments, analyzing customer feedback.
5. Manufacturing:
Objective: To increase efficiency and reduce costs in production processes.
Applications: Predictive maintenance, supply chain optimization.
Examples: Predicting equipment failures, optimizing inventory levels.
DATA ANALYTICS TOOLS
1. Tableau
Tableau is an easy-to-use Data Analytics tool. Tableau has a drag-and-drop interface
which helps to create interactive visuals and dashboards. Organizations can use this to
instantly develop visuals that give context and meaning to the raw data, making the data
very easy to understand. Also, due to the simple and easy-to-use interface, one can easily
use this tool regardless of their technical ability. Furthermore, Tableau comes with a wide
range of features and tools that help you create the best visuals which are easy to
understand.
The advantage of Tableau that overshadows all others is in its Quality Visuals embedded
with Interactive Information. But this doesn’t mean Tableau is perfect. Tableau is only
meant for Data Visualisation, so we can’t preprocess data using this tool. Also, it does have
a bit of a learning curve and is known for its high cost.
Features:
Features:
Features:
Features:
Features:
7. Python
Python is another Programming Language popular for Data Analysis and Machine
[Link] is used extremely in Data analysis tools. Python is widely recognized to
have easy syntax which makes it easy to learn. Along with the easy syntax, the package
manager of Python features a lot of important packages and libraries. This makes it suitable
for Data Analysis and Machine Learning. Another reason to use Python is its scalability.
This doesn’t mean Python is flawless. It is quite slow when we compare it to languages
like Java or C++; this is because Python is an interpreted language while the others are
compiled. Besides, Python is also infamous for its high memory consumption.
Features:
8. SAS
SAS stands for Statistical Analysis System. The SAS Software was developed by the SAS
Institute, and it is widely used for Business Analytics nowadays. SAS has both
a Graphical User Interface and a Terminal Interface. So, depending on the user’s
skillsets, they can choose either one. It also has the ability to handle large datasets. In
addition, SAS is equipped with a lot of Analytical Tools which makes it valid for a lot of
applications.
Although SAS is very powerful, it has a big price tag and a steep learning curve, so it is
quite hard for beginners.
Features:
9. QlikSense
QilkSense is a Business and data analysis Tools that provides support for Data
Visualisation and Data Analysis. QuilkSense supports various Data sources
from Spreadsheets, Databases, and also Cloud Services. You can create amazing
Dashboards and Visualisations. It comes with Machine Learning features and uses AI to
help the user understand the Data. Furthermore, QlikSense also has features like Instant
Search and Natural Language Processing.
But QilkSense does have some drawbacks. The data extraction of QilkSense is quite
inflexible. The Pricing Model is quite complicated, and it is quite sluggish when it comes
to large datasets.
Features:
Features:
Microsoft Excel :
It is an important spreadsheet application that can be useful for recording expenses,
charting data and performing easy manipulation and lookup and or generating pivot tables
to provide the desired summarized reports of large datasets that contain significant data
findings. It is written in C#, C++ and .NET Framework and its stable version were released
in 2016. It involves the use of a macro programming language called Visual Basic for
developing applications. It has various built-in functions to satisfy the various statistical,
financial and engineering needs. It is the industry standard for spreadsheet applications. It
is also used by companies to perform real-time manipulation of data collected from
external sources such as stock market feeds and perform the updates in real-time to
maintain a consistent view of data. It is relatively useful for performing somewhat complex
analyses of data when compared to other tools such as R or python. It is a common tool
among financial analysts and sales managers to solve complex business problems.
RapidMiner :
RapidMiner is an extremely versatile data science platform developed by “RapidMiner
Inc”. The software emphasizes lightning fast data science capabilities and provides an
integrated environment for preparation of data and application of machine learning, deep
learning, text mining and predictive analytical techniques. It can also work with many data
source types including Access, SQL, Excel, Tera data, Sybase, Oracle, MySQL and Dbase.
Here we can control the data sets and formats for predictive analysis.
Approximately 774 companies use RapidMiner and most of these are US-based. Some of
the esteemed companies on that list include the Boston Consulting Group and Dominos
Pizza Inc.
Knime :
Knime, the Konstanz Information Miner is a free and open-source data analytics software.
It is also used as a reporting and integration platform. It involves the integration of various
components for Machine Learning and data mining through the modular data-pipe lining. It
is written in Java and developed by [Link] AG. It can be operated in various
operating systems such as Linux, OS X and Windows. More than 500 companies are
currently using this software for operational purposes and some of them include Aptus Data
Labs and Continental AG.
COVID-19 has changed the business landscape in myriad ways and historical data is no
more relevant. So, in place of traditional AI techniques, arriving in the market are some
scalable and smarter Artificial Intelligence and Machine Learning techniques that can
work with small data sets. These systems are highly adaptive, protect privacy, are much
faster, and also provide a faster return on investment. The combination of AI and Big
data can automate and reduce most of the manual tasks.
Agile data and analytics models are capable of digital innovation, differentiation, and
growth. The goal of edge and composable data analytics is to provide a user-friendly,
flexible, and smooth experience using multiple data analytics, AI, and ML solutions. This
will not only enable leaders to connect business insights and actions but also, encourage
collaboration, promote productivity, agility and evolve the analytics capabilities of the
organization.
One of the biggest data trends for 2024 is the increase in the use of hybrid cloud
services and cloud computation. Public clouds are cost-effective but do not provide high
security whereas a private cloud is secure but more expensive. Hence, a hybrid cloud is a
balance of both a public cloud and a private cloud where cost and security are balanced to
offer more agility. This is achieved by using artificial intelligence and machine learning.
Hybrid clouds are bringing change to organizations by offering a centralized database,
data security, scalability of data, and much more at such a cheaper cost.
4. Data Fabric
A data fabric is a powerful architectural framework and set of data services that
standardize data management practices and consistent capabilities across hybrid multi-
cloud environments. With the current accelerating business trend as data becomes more
complex, more organizations will rely on this framework since this technology can reuse
and combine different integration styles, data hub skills, and technologies. It also reduces
design, deployment, and maintenance time by 30%, 30%, and 70%, respectively, thereby
reducing the complexity of the whole system. By 2026, it will be highly adopted as a re-
architect solution in the form of an IaaS (Infrastructure as a Service) platform.
In 2024, one of the exciting trends in data analysis is edge computing. This approach
basically brings the data processing closer to where it’s generated, like in smart devices or
sensors. This means that instead of sending all the data to a central location for processing,
it’s analyzed right where it’s created. This not only speeds up the analysis but also helps to
keep the data more secure because it doesn’t have to travel long distances over networks.
As businesses look for faster and more reliable ways to make decisions, edge computing is
becoming a key player in helping them get the insights they need, when they need them.
6. Augmented Analytics
Earlier businesses were restricted to predefined static dashboards and manual data
exploration restricted to data analysts or citizen data scientists. But it seems dashboards
have outlived their utility due to the lack of their interactivity and user-friendliness.
Questions are being raised about the utility and ROI of dashboards, leading organizations
and business users to look for solutions that will enable them to explore data on their own
and reduce maintenance costs.
It seems slowly business will be replaced by modern automated and dynamic BI tools that
will present insights customized according to a user’s needs and delivered to their point of
consumption.
8. XOps
XOps has become a crucial part of business transformation processes with the adoption
of Artificial Intelligence and Data Analytics across any organization. XOps started with
DevOps that is a combination of development and operations and its goal is to improve
business operations, efficiencies, and customer experiences by using the best practices of
DevOps. It aims in ensuring reliability, re-usability, and repeatability and also ensure a
reduction in the duplication of technology and processes. Overall, the primary aim of XOps
is to enable economies of scale and help organizations to drive business values by
delivering a flexible design and agile orchestration in affiliation with other software
disciplines.
With evolving market trends and business intelligence, data visualization has captured the
market in a go. Data Visualization is indicated as the last mile of the analytics process and
assists enterprises to perceive vast chunks of complex data. Data Visualization has made it
easier for companies to make decisions by using visually interactive ways. It influences the
methodology of analysts by allowing data to be observed and presented in the form of
patterns, charts, graphs, etc. Since the human brain interprets and remembers visuals more,
hence it is a great way to predict future trends for the firm.
Data Analytics Softwares That You Must Know
There are several popular software options for data analytics, each with its strengths and
suitability for different tasks. Here are some of the most widely used ones:
Microsoft Excel
Excel is a widely used spreadsheet program that may be used for simple computations,
graphing, and data manipulation activities. It also provides basic data analysis features.
Because of its accessibility and familiarity, it is extensively used; yet, it does not have the
sophisticated statistical and visualization tools needed for intricate analysis.
Key Features: Excel is a versatile tool widely used for data analysis due to its familiarity
and ease of use. It offers basic statistical functions, pivot tables, and charting capabilities.
Suitable For: Small to medium-sized datasets and users who prefer a familiar interface
for basic analysis tasks.
Tableau
Users may generate dynamic and interactive representations from a variety of data sources
with Tableau, a sophisticated tool for data visualization. Because of its vast customization
possibilities and user-friendly interface, it is highly regarded for its ability to explore and
convey findings via intuitive dashboards.
Key Features: Tableau is a powerful data visualization software that allows users to
create interactive dashboards and visualizations from multiple data sources. It offers
drag-and-drop functionality and intuitive design tools for creating compelling
visualizations.
Suitable For: Business users, analysts, and data scientists who need to communicate
insights effectively through visually appealing dashboards and reports.
Python
Python provides a rich environment for data analysis with its adaptable libraries,
including NumPy for numerical computation, Pandas for data manipulation,
and Matplotlib for data visualization. It is ideal for a variety of analytical activities, including
machine learning modeling and exploratory data analysis, and it is very adaptable and
scalable.
Key Features: Python, along with libraries like Pandas and NumPy, provides powerful
tools for data manipulation, analysis, and visualization. It offers extensive capabilities for
handling large datasets and performing complex operations.
Suitable For: Data scientists, analysts, and programmers who require flexibility,
scalability, and the ability to integrate data analysis into custom applications.
R
For data analysis, statistical modeling, and visualization, R is a popular statistical
programming language. Because it provides a wide range of packages for different types of
analytical work, statisticians and data scientists find it to be quite popular. R is a powerful
tool for sophisticated statistical analysis and intricate visualizations, but its learning curve
could be more steep than that of other programs.
Key Features: R is a programming language specifically designed for statistical analysis
and data visualization. It offers a vast ecosystem of packages tailored for various
analytical tasks, making it a preferred choice for statistical modeling and advanced data
analysis.
Suitable For: Statisticians, researchers, and analysts working with complex statistical
models and specialized analytical techniques.
Power BI
Power BI is a collection of business intelligence tools that is a component of the Microsoft
Power Platform. It consists of Power Query for data transformation, Power BI for data
visualization, and Power Pivot for data modeling. It is preferred because to its connection
with other Microsoft products and offers complete data analysis capabilities, ranging from
data preparation to interactive dashboard building.
Key Features: Power BI is a business analytics tool by Microsoft that enables users to
visualize and share insights from their data. It offers robust data connectivity, interactive
dashboards, and AI-driven analytics capabilities.
Suitable For: Business users, data analysts, and decision-makers who require self-
service analytics and real-time insights for decision-making.
SAS
SAS is a full-featured software package designed for predictive modeling, data management,
and advanced analytics. The software provides an extensive array of statistical techniques,
machine learning algorithms, and data manipulation capabilities, rendering it appropriate for
intricate analytical assignments in sectors including research, healthcare, and finance.
Key Features: SAS is a comprehensive analytics platform offering a wide range of
statistical analysis, data management, and machine learning capabilities. It is known for
its reliability, scalability, and advanced analytics features.
Suitable For: Enterprises, government agencies, and organizations with complex
analytical needs and stringent data security requirements.
Google Analytics
Google Analytics is a web analytics service offered by Google that tracks and reports website
traffic. It provides valuable insights into user behavior, website performance, and marketing
effectiveness.
Key Features: Google Analytics offers a range of features for analyzing website traffic,
including audience demographics, acquisition sources, and user engagement metrics. It
also supports custom reporting and integration with other Google products.
Suitable For: Businesses and website owners seeking to understand and optimize their
online presence through data-driven insights.
Apache Hadoop
Apache Hadoop is an open-source software framework used for distributed storage and
processing of large datasets. It provides a scalable and fault-tolerant platform for big data
analytics.
Key Features: Hadoop consists of a distributed file system (HDFS) for storage and a
distributed processing framework (MapReduce) for parallel computation. It supports the
processing of large volumes of data across clusters of commodity hardware.
Suitable For: Organizations dealing with massive volumes of data and requiring scalable
solutions for storage and processing.
Apache Spark
Apache Spark is an open-source distributed computing system used for big data processing
and analytics. It provides a fast and general-purpose framework for in-memory data
processing.
Key Features: Spark offers high-level APIs in multiple languages
(e.g., Scala, Java, Python) for building parallel applications. It supports various data
processing tasks, including batch processing, streaming analytics, machine learning, and
graph processing.
Suitable For: Organizations requiring real-time or near-real-time analytics on large
datasets with complex processing requirements.
KNIME (Konstanz Information Miner)
KNIME (Konstanz Information Miner), a powerful open-source platform designed to
streamline and simplify the data analysis and integration process. Born out of the University
of Konstanz in Germany, KNIME has evolved into a leading tool in the data science
community, offering a user-friendly interface coupled with robust functionality.
Key Features: KNIME offers a drag-and-drop interface for building data processing
pipelines, which can include tasks such as data preprocessing, machine learning, and
visualization. It supports integration with various data sources and formats, as well as a
wide range of plugins for extending functionality. KNIME emphasizes collaboration and
scalability, making it suitable for both individual analysts and enterprise-scale data
science teams.
Suitable For: KNIME is suitable for data scientists, analysts, and researchers who prefer
a visual approach to data analysis and workflow creation. It is particularly useful for
organizations requiring flexible and scalable solutions for data analytics and automation.
RapidMiner
RapidMiner is a data science platform that offers an integrated environment for data
preparation, machine learning, predictive analytics, and model deployment. It is designed to
simplify the entire data science workflow, from data ingestion to model deployment.
Key Features: RapidMiner provides a visual workflow designer that allows users to
build, validate, and deploy predictive models without writing code. It offers a wide range
of machine learning algorithms, data preprocessing tools, and model evaluation
techniques. RapidMiner also supports integration with various data sources and systems,
as well as advanced features such as automated machine learning (AutoML) and model
optimization.
Suitable For: RapidMiner is suitable for data scientists, analysts, and business users who
require an end-to-end platform for data science and analytics. It is widely used in
industries such as finance, healthcare, retail, and telecommunications for tasks such as
customer segmentation, fraud detection, and predictive maintenance.
Capabilities of Data Analytics Software
For data analysis, each of these tools has a different set of characteristics and skills, such as:
Microsoft Excel: Well-known for its integrated features, data visualization, and pivot
tables.
Tableau: Provides dashboards, powerful analytics, and interactive data visualization.
Python: Offers flexible statistical analysis, machine learning, and data manipulation
features.
R: Widely used in predictive modeling, data visualization, and statistical computation.
Power BI: Facilitates the preparation, visualization, and intra-organizational sharing of
insights.
SAS: Provides corporate intelligence, data management, and advanced analytics
solutions.
IBM SPSS: Well-known for its capacities in data mining, statistical analysis, and
predictive modeling.
Google Analytics: Concentrates on monitoring user activity, measuring performance,
and web analytics.
RapidMiner: Offers tools for predictive analytics, machine learning, and data
preparation.
KNIME: Provides machine learning, integration, and visual data analytics.
Comparing Data Analytics Tools
Software Language/Platform Pros Cons