Data Visualization Notes
Data Visualization Notes
1. Charts and Graphs: They are used to visualize data, with charts comparing data points
across categories or showing trends over time and graphs analyzing relationships between
variables to identify correlations, trends and outliers. Examples: Bar Charts, Line Charts,
2. Maps: They are used to display geographical data which provides spatial context to trends
3. Dashboards: They combine multiple visualizations into a single interface which provides
1. Simplifies Complex Data: It turns large and complicated data into visual formats like
2. Reveals Patterns and Trends: It helps identify trends, relationships and patterns that are
3. Saves Time: Visuals allow quicker interpretation of data, helping users spot key
5. Tells a Clear Story: Data visuals guide the audience through the information step-by-step,
1. Business Analytics: Used to monitor company performance, track KPIs and make
2. Healthcare: Helps in analyzing patient records, tracking disease outbreaks and managing
3. Sports: Used to visualize player statistics, team performance and match outcomes, helping
4. Retail and E-commerce: Enables tracking of sales, customer preferences and inventory
2. Over-Simplification: Simplifying data too much can lead to important details being lost
like using a pie chart that oversimplifies complex relationships between categories.
3. Choosing the Right Visualization: Using the wrong type of visualization can distort the
message. For example, a pie chart might not work well with many categories which leads to
confusion.
viewers. It's important to focus on key data points and avoid clutter.
Best Practices for Effective Data Visualization
To ensure our visualizations are impactful and easy to understand we follow these best practices:
audience may need detailed graphs while a general audience benefits from simpler charts.
2. Design Clarity and Consistency: Choose the right chart for our data and keep the
design clean with consistent colors, fonts and labels and also avoid clutter to ensure clarity.
3. Provide Context: Always provide context by including labels, titles and data source
acknowledgments. This helps viewers to understand the significance of the data and builds
4. Interactive and Accessible Design: Make visualizations interactive with features like
tooltips and filters and ensure accessibility for all users regardless of device or visual needs.
How Does Data Visualization Simplify Complex Information for Informed Decisions?
Data visualization simplifies complex information by converting data into visual formats such as charts,
graphs, and dashboards. These visual representations make it easier for users to interpret data, identify
trends, and draw insights. By presenting data visually, organizations can communicate complex
information more effectively and facilitate quicker decision-making processes. Visualizations help in
highlighting key data points, comparing different datasets, and identifying patterns that may not be
apparent in raw data. This simplification of data allows stakeholders at all levels of an organization to
understand and act upon information more efficiently, leading to more informed decision-making.
revolutionizing how we understand and interact with data. In the past, static visuals provided a basic
overview, but today's tools offer immersive experiences. Interactive dashboards allow users to
manipulate variables, drill down into details, and uncover insights that were once hidden in static
charts.
This evolution has been driven by advancements in technology, particularly in data processing and user
interface design. Modern data visualization tools leverage these advancements to create engaging and
As a result, decision-makers can now explore data more deeply, make informed decisions more quickly,
and communicate insights more effectively. This evolution continues to shape the field of data
visualization, opening up new possibilities for how we analyze and understand data.
Choosing the Right Data Visualization Tool for Your Business Needs
Selecting the right data visualization tool is crucial for effectively communicating insights. Several
factors should be considered when choosing a tool for your business needs. Firstly, consider the type of
data you're working with. For example, if you're dealing with geospatial data, a tool that specializes in
Secondly, think about your audience and the level of interactivity required. If you're creating
interactive dashboards for business users, a tool with robust dashboarding capabilities would be ideal.
Additionally, consider the scalability and compatibility of the tool with your existing systems.
Lastly, evaluate the ease of use and learning curve of the tool, as well as the support and resources
available from the vendor. By carefully considering these factors, you can select a data visualization tool
that meets your business needs and empowers you to communicate insights effectively.
Effective data visualization is crucial for communicating insights clearly and persuasively. To achieve
this, it's essential to follow several best practices. Firstly, choose the right type of visualization for your
data. Bar charts are ideal for comparing values, while line charts work well for showing trends over
time.
Secondly, keep your visuals simple and uncluttered. Avoid using too many colors or elements that can
distract from the main message. Use color strategically to highlight key points and create visual
hierarchy.
Thirdly, provide context for your data to help viewers understand its significance. Use annotations,
labels, and captions to explain the data and provide additional insights.
Lastly, consider your audience and tailor your visualizations to their needs and level of expertise. By
following these best practices, you can create visualizations that effectively communicate insights and
Understanding and applying color theory can greatly enhance the effectiveness of your visualizations.
Firstly, use colors that are easily distinguishable to avoid confusion. Consider using a color palette with
Secondly, use color to highlight key data points or trends. By using a bright or contrasting color, you
Additionally, consider the psychological effects of color. For example, warm colors like red and orange
can evoke a sense of urgency or importance, while cool colors like blue and green can create a sense of
calmness or stability.
Overall, using color effectively can make your data visualizations more engaging, easier to understand,
understandable, and actionable. Traditional methods of data analysis, such as spreadsheets and reports,
are being replaced by interactive visualizations that allow users to explore data in real-time and gain
deeper insights.
One key way data visualization is transforming business intelligence is by enabling faster
decision-making. With interactive dashboards and visualizations, decision-makers can quickly identify
trends, patterns, and outliers that may not be apparent in raw data, leading to more informed
decisions.
users can now easily create and interpret visualizations, reducing the reliance on data analysts and
Overall, data visualization is revolutionizing business intelligence by providing new ways to analyze and
interpret data, leading to improved decision-making, increased efficiency, and a competitive advantage
Data visualization plays a crucial role in data-driven decision-making by helping users understand
complex data sets and extract valuable insights. By presenting data visually, decision-makers can quickly
identify trends, patterns, and outliers that may not be apparent in raw data. This enables them to make
informed decisions based on data rather than intuition or guesswork, leading to better outcomes for
their organizations.
Visualizations make it easier to share insights and findings, ensuring that everyone is working from the
empowers decision-makers at all levels to make better, data-driven choices that lead to improved
Conclusion
In conclusion, data visualization plays a crucial role in business intelligence by simplifying complex
By presenting data visually, organizations can communicate insights more effectively, leading to
quicker and more informed decision-making processes. Data visualization also promotes collaboration
and alignment across teams, as visualizations can be easily shared and understood by all stakeholders.
Additionally, data visualization enables organizations to identify emerging trends, predict customer
behavior, and spot market opportunities, ultimately driving business growth and fostering innovation.
As organizations continue to collect and analyze large volumes of data, the importance of data
visualization in business intelligence will only continue to grow, helping organizations stay competitive
1. Categorical Data:
Categorical data represent discrete, qualitative information. Examples include Gender, color, marital
status.
Application in Data Science: Categorical data is used for classification tasks and creating meaningful
categories.
Visualization Techniques: Bar charts, pie charts, and stacked bar charts are commonly used to visualize
categorical data.
2. Numerical Data:
Numerical data represents quantitative information. Examples include Age, height, temperature.
Application in Data Science: Numerical data is used for statistical analysis, regression, and prediction
models.
Visualization Techniques: Histograms, box plots, and scatter plots are effective visualization techniques
for numerical data.
Time series data represents data points collected over a period of time. Examples include Stock prices,
temperature readings, website traffic.
Application in Data Science: Time series data is used for forecasting, trend analysis, and anomaly
detection.
Visualization Techniques: Line graphs, area charts, and seasonal decomposition plots help visualize
time series data.
4. Text Data:
Text data comprises unstructured textual information. Examples include Tweets, customer reviews,
news articles.
Application in Data Science: Text data is used for sentiment analysis, natural language processing, and
text mining.
Visualization Techniques: Word clouds, bar charts, and scatter plots with text labels are commonly
employed for visualizing text data.
5. Image Data:
Image data consists of visual information in the form of pixels. Examples include Photographs, medical
scans, satellite imagery.
Application in Data Science: Image data is used in computer vision, object detection, and image
recognition.
Visualization Techniques: Heatmaps, image grids, and image overlays are effective visualization
methods for image data.
6. Geospatial Data:
Geospatial data represents geographic information. Examples include GPS coordinates, city
boundaries, population density.
Application in Data Science: Geospatial data is used for mapping, spatial analysis, and location-based
services.
Visualization Techniques: Choropleth maps, heatmaps, and scatter plots with geographical coordinates
are common visualization techniques for geospatial data.
7. Numerical/Quantiative Data
● Continuous: Data that can take on any value within a specified interval. It can be measured,
and can take on any value, including fractions or decimals.
○ Example: A measure of greenhouse gas emissions in the US.
● Discrete: Data that has distinct values that can be counted. It can only take on integer values.
○ Example: The number of households in a county.
Categorical/Qualitative Data
● Nominal: Data that has distinct values representing different groups or categories with no
inherent ranking or order.
○ Example: The names of different counties.
● Ordinal: Data that has distinct values representing categories with a meaningful order or
ranking.
○ Example: Education levels, such as elementary, middle, high school, college, and
post-graduate.
● Binary: Data that has values that can only be one of two distinct categories.
○ Example: success/failure, yes/no
8. Temporal Data
● Time Series: data points collected over an interval of time and indexed by time.
○ Example: Daily weather data.
Spatial Data
● Raster: Uses pixels to represent geographic information. Each pixel value represents an area on
Earth.
○ Example: Satellite images
● Vector: Represents geographic features as points, lines, and polygons to define the shape of a
spatial object.
○ Example: A polygon representing the boundaries for a state.
1. Bar Charts and Histograms: Suitable for visualizing categorical and numerical data distributions.
2. Scatter Plots: Effective for showcasing relationships between two numerical variables.
3. Line Graphs: Useful for displaying trends and patterns over time.
4. Heatmaps: Ideal for representing matrix-like data with color intensity.
5. Box Plots: Provide a summary of numerical data's distribution and identify outliers.
6. Word Clouds: Visually summarize textual data by displaying frequently occurring words.
- Matching Data Types to Visualization Techniques: Ensure the chosen visualization method aligns
with the data type and the insights you want to convey.
- Considering the Objective and Audience: Tailor the visualization to the intended purpose and the
target audience's level of technical understanding.
- Design Principles for Effective Data Visualization: Pay attention to aspects like color choice, labeling,
and chart layout to create visually appealing and informative visualizations.
Conclusion:
Data is the backbone of Data Science, and understanding the various types of data is essential for
extracting meaningful insights. Through effective data visualization techniques, Data Scientists can
unlock patterns, trends, and relationships, leading to informed decision-making and valuable insights.
By incorporating the appropriate visualization methods based on data type and purpose, Data
Scientists can communicate their findings more effectively and empower stakeholders with actionable
insights.
TYPES OF CHARTS
1. Bar chart
A bar chart visually represents data using rectangular bars or columns. Here, the length of each bar
corresponds proportionally to its value. You can present these bars horizontally or vertically.
Bar charts are excellent for comparing the values of different categories or groups. Apart from that,
these types of charts are also helpful in showing the distribution of data across different categories.
2. Histogram
A histogram visualizes a single, continuous dataset. Although histograms and bar charts are used
interchangeably, they differ in practice. For instance, a bar graph is a plot of a single data point (a sum,
average, or other value) for each category while a histogram is a plot of a range of data.
Think of histograms as a more informative way to view the distribution of values in a dataset. It is ideal
for visualizing the spread and variation of the data, helping you identify outliers or unusual data points
representation
Histograms are less effective with smaller datasets, so make sure you have sufficient data
points.
3. Column chart
Column charts are the simplest, most versatile type of visualization used in data analytics. The
horizontal chart displays your data in bars proportional to the values they represent.
More often than not, column charts effectively compare data across different categories. They are also
helpful in displaying rankings and order in a dataset, allowing viewers to identify trends quickly.
5. Line chart
A line chart connects distinct data points through straight lines. Its best use case is to illuminate trends,
This type of chart helps measure how different groups relate to each other. This type of chart is also
effective for demonstrating progression, making them suitable for scenarios like project timelines,
5. Pie chart
A common but limited type of graph is the pie chart. It is a circular, statistical graphic that divides data
into slices, where each slice represents a percentage or proportion of the whole. You can create your
own pie chart to illustrate how each category contributes to the overall dataset.
This classic chart type is effective when you want to illustrate the proportion of each category in the
dataset. The chart is suitable when you have limited categories, ideally less than six or seven.
What’s hiding in your data? Find out by signing up for your demo
Scatter plot
Scatter plots are types of visualization that show a collection of data points ‘scattered’ around the
Scatter plots are ideal for exploring relationships and patterns between two continuous variables. They
can help you identify trends, correlations, or potential clusters in the data.
Heatmap chart
Heatmap charts are a type of map data visualization that uses a system of color coding to represent
value. Each cell in the matrix is assigned a color based on the value it holds.
This type of chart is commonly used to establish relationships between two variables across a grid. In
the example above, the intensity of the colors in the map clearly demonstrates the variables, making it
Choose an intuitive color palette that effectively conveys the magnitude of values.
Let’s be real—attention spans are short. Like, shorter-than-a-goldfish short. If your chart is confusing,
cluttered, or just plain wrong for the data, you’ll lose your audience before they even get to your
insight.
So how do you choose the right chart? The one that makes your point land and sparks action?
Understand your data: Look at the size, structure, and distribution of your dataset.
Consider your audience: Tailor your visuals to their level of data literacy. Simpler charts
Prioritize clarity: Choose the chart that makes your point instantly clear. Don’t make your
audience second-guess.
Stick to visual best practices: Use consistent colors, intuitive labels, and readable axes.
Use interactivity wisely: When possible, let users explore with filters, drill-downs, or
The data visualization process transforms raw, complex data into actionable insights through a
structured cycle:
collecting relevant data, cleaning and exploring it to identify patterns, analyzing to extract insights,
visualizing through charts/maps, and interpreting to inform decisions. This iterative, analytical-visual
approach helps spot trends, outliers, and relationships.
● Data Collection: Gathering raw data from various sources and ensuring it is accurate and clean.
● Exploration (EDA): Using summary statistics and initial visual plots to understand data
structure, detect outliers, and find initial patterns.
● Analysis
:
Applying statistical methods or models to test hypotheses, identify trends, and derive
meaningful insights from the data
.
● Visualization: Translating analyzed data into graphical representations (charts, maps, graphs)
to make insights easily understandable.
● Interpretation: Drawing conclusions, creating narratives, and making data-driven decisions
based on the visual findings.
● Reduces effort and time spent handling errors during data analysis and reporting
● Missing values: Incomplete records can reduce statistical power and introduce bias into
analysis.
skewed outcomes.
● Incorrect data types: Mismatched formats, such as text stored in numeric fields can cause
● Outliers and anomalies: Extremely high or low values can distort statistical measures and
● Inconsistent formats: Variations in date formats, text casing or measurement units can
● Spelling and typographical errors: Errors in text fields can lead to incorrect grouping,
● Missing Values: Identify any blank or null values in the dataset. Missing values can be due
to various reasons such as incomplete data collection, data entry errors or data loss during
transmission.
● Incorrect Values: Check for values that are outside the expected range or are inconsistent
● Inconsistencies in Data Format: Verify that the data format is consistent throughout the
dataset.
Removing irrelevant or duplicate data ensures the dataset is clean, accurate and meaningful, preventing
skewed analysis and improving overall quality.
● Remove duplicate records to ensure each data point is unique and correctly represented.
● Detect redundant observations that do not add new information to the dataset.
● Eliminate variables or columns that are irrelevant to the analysis and do not provide useful
insights.
Structural errors occur when data formats, naming conventions or variable types are inconsistent
which can affect analysis accuracy. Correcting these issues ensures uniform and reliable data
representation.
● Standardize data formats to maintain consistency in dates, times and other data types
● Correct naming inconsistencies in column names, variable names or labels to ensure clarity
and uniformity.
● Ensure consistent data representation such as using the same units for measurements or the
Missing data can introduce bias and reduce the reliability of analysis. Properly addressing missing
values helps maintain the integrity of your dataset.
● Impute missing values using statistical methods such as mean, median or mode to fill gaps.
● Remove records with missing values when the missing data is extensive or cannot be
accurately imputed.
5. Normalize Data
Data normalization organizes the dataset to reduce redundancy and ensure consistency making it easier
to manage and analyze.
● Split data into multiple tables, with each table storing specific types of information.
● Ensure consistency across the dataset to support efficient querying and accurate analysis.
Outliers are data points that deviate significantly from the rest of the dataset and can affect analysis
accuracy. Properly handling them ensures more reliable insights.
● Remove outliers that result from errors or are not representative of the population.
● Transform extreme but valid outliers to reduce their impact on the analysis.
● Understand the data: Know the source, structure and domain of the data to identify
● Prioritize critical issues: Focus first on major quality problems that could have a systemic
● Automate where possible: Use scripts or tools for repetitive cleaning tasks to improve
● Monitor and maintain: Continuously track data quality and perform cleaning
Limitation
Cleaning data can be difficult due to several factors:
● Large datasets are harder to clean because of their size, requiring efficient tools and
techniques.
● Data from multiple sources often have different structures and formats making integration
● Data cleaning is a continuous process as new data must be regularly assessed and
13.1 INTRODUCTION
In today’s digital era, data is often considered as the new oil—an invaluable
resource powering decision-making across industries like healthcare, finance,
retail, education, and manufacturing. Yet, the true power of data lies not in its
raw form, but in “how effectively we can interpret and apply it?”. That’s
where data science steps in: a multidisciplinary domain blending
mathematics, statistics, computer science, and domain expertise to transform
chaotic data into meaningful knowledge.
This Unit, serves as your foundational guide to four of the most widely
adopted tools in the data science pipeline:
· Tableau – An industy-leading platform for visual storytelling and
dashboard creation.
· Power BI – Power BI – A robust business intelligence tool from
Microsoft that lets you create interactive reports and easily connects with
other Microsoft products for smooth data analysis.
· Python – A versatile and easy-to-learn programming language, Python is
widely favoured by beginners and professionals alike for data tasks—
thanks to powerful libraries like Pandas, NumPy, Matplotlib, and Scikit-
learn that simplify data handling and machine learning.
· R – A statistical powerhouse designed for analytical rigor, favoured in
academia and research-heavy environments.
The sections covered in this unit, introduces not just the technical skills
needed to work with these tools, but also offers practical exposure to their
user interfaces, dataset handling methods, and data visualization techniques.
The goal is to equip students, aspiring data scientists, and professionals with
both the confidence and competence to begin their journey in data science.
13.2 OBJECTIVES
Upon completing this unit, learners will be able to:
These objectives are designed to empower learners to not only master the
basics of data tools, but also apply their skills across academic, research, or
industry settings with clarity and confidence.
Below are four foundational tools widely used in the industry and academia:
· Tableau: A powerful and easy-to-use visualization tool that lets you
build interactive reports and dynamic dashboards—no coding required.
It’s perfect for turning data into clear stories and sharing insights with
people who may not have a technical background.
· Power BI: Created by Microsoft, Power BI is a strong business
intelligence tool that works smoothly with Excel and allows users to
build real-time dashboards—making it ideal for enterprise-level data
analysis and reporting.
· Python (Anaconda/Google Colab): A general-purpose programming
language with extensive libraries like Pandas, NumPy, and Matplotlib
for data manipulation, analysis, and visualization. Anaconda simplifies
local setup, while Colab provides a free cloud-based environment.
· R (RStudio): R is a programming language tailored for statistical
analysis and data visualization. It’s widely used in academic research and
data-centric projects for its strong capabilities in handling complex data
and producing high-quality graphics. RStudio makes coding, visualizing,
and managing data efficient and user-friendly.
These tools are essential to today’s data science process, and mastering them
equips students and professionals with the skills to tackle real-world
challenges efficiently and confidently
Datasets can be sourced from various platforms depending on the need and
application. Popular sources include data-sharing platforms like Kaggle and
the UCI Machine Learning Repository, which host numerous datasets for
academic and practical use. Government portals such as [Link] provide
open-access datasets for public use. Additionally, APIs from platforms like
Twitter or Google Maps offer real-time and structured data. Other methods
include web scraping, where data is extracted from websites, and data
collected through IoT devices or sensors, which continuously generate data
in various domains. Some of the usefil links to gather the datasets are given
below:
· Kaggle — [Link]
· UCI Machine Learning Repository — [Link]
· Google Dataset Search — (search via
[Link]
· Open Government Data Platform India ([Link]) — [Link]
· [Link] (US) — [Link]
· Community/curated lists like “Awesome Public Datasets” on GitHub or
[Link]
We will discuss, the usage of datasets along with the utilization of various
tools mentioned above. Let’s begin with Tableau first, in the next section.
13.5 TABLEAU
Tableau is a user-friendly and a powerful tool for creating interactive data
visualizations. It connects easily to a wide range of data sources and lets
users explore and present their data through clear, engaging visuals with its
simple drag-and-drop interface, even beginners can build charts, graphs, and
dashboards—no coding needed. Widely used by data analysts and scientists,
Tableau is popular in industries like healthcare, tech, and e-commerce to
support smart, data-driven decisions. Once installed, users can get started by
activating their license or signing in with their Tableau account.
448
Tableau Installation: Tableau is a free, easy
easy-to-use tool that lets users build Tools for Data
Science an
engaging
ing and interactive visualizations
visualizations—all without writing any code. It Introduction
helps transform raw data into a more readable and insightful format. Tableau
has become
ome a widely used tool in the business analyt
analytics space.
Userss can connect with others across the globe, share dashboards and data
visualizations,
sualizations, and even publish their work on websites, blogs, or social
media platforms.
Here
re in this section, we will try to cover most of the features of Tableau in
brief.
ef. To begin with the Prerequisites of Using Tableau.
As a Prerequisites
rerequisites for Using Tableau
Tableau, it only requires Basic computer skills,
such
uch as running programs and navigating software interfaces
interfaces, along with
Familiarity
amiliarity with spreadsheet applications
applications, lastly your Willingness to explore
and learn new tools.
Firstly,
irstly, Let’s learn how to Install Tableau, in order to Get Started with
Installation, simply
imply go to the official website at [Link] and
follow the step-by-step
tep instructions provided there.
The stepwise process is also given here, the Fig.1: Tableau Installation shows
the interface available at official
icial website at [Link]
Fig.1:
ig.1: Tableau Installation
Refer
efer to Fig.2: Tableau Desktop Products, it covers Tableau Desktop and
Tableau
ableau Prep
Prep, Tableau Server and Tableau ableau Online as its major
components, a brief
ief description of these components is as follows:
· Tableau Desktop
Desktop: A standalone desktop application designed for
individual
ndividual users to create and analyze data visualizations.
· Tableau Prep
Prep: This includes two components——Tableau Prep Builder,
iss used to design and build data preparation workflows, while Tableau
Prep Conductor allows ows organizations to schedule, manage, and track
these workflows efficiently across teams.
Fig.4:
ig.4: Tableau Server
· Tableau Online iss a public platform, so any content you upload can be
accessed
cessed by other users of the software. As a result, your work is no
longer private or confidential.
Fig.5:
ig.5: Tableau Online
Choose
hoose Tableau Desktop and to start
tart your free trial,
trial perform the following
steps:
· Iff you prefer not to use the trial, you can download Tableau Public
instead. Oncee done, Tableau will be installed on your system.
Once Tableau is installed you will see the user interface as shown in Fig. 7,
below
If you're working with Tableau in a web browser, check out “Creators: Get
Started
tarted with Web Authoring” and “Tour Your Tableau Site” for guidance.
Fig.8:
ig.8: Interface of Tableau
The table (Table 1: Tableau Toolbar) shown below describes what each
button on the toolbar does. Note that certain buttons may not appear in every
version of Tableau. You can also refer to Visual Cues and Icons in Tableau
Desktop for additional details.
13.5.3 How
ow to upload any dataset ?
The stepwise
wise process of uploading any dataset, in Tableau is as follows:
Step – 1: Start Tableau on your system
system, you will see an interface as shown in
Fig 9 below
Fig 9: Tableau
eau Front Page Interface
Step – 2: "Navigate
igate to the 'Connect' section, choose 'To a File', then click on
filee format you want to upload for an instance 'Microsoft Excel'."
Excel'.", as shown
in Fig.10 below.
Fig
ig 10: File Upload Interface in Tableau
Data Science – Allied Step – 3: Clicking
icking this opens a file browser window, allowing you to navigate
Areas
through folders and choose the file you wish to open
open,, as shown in Fig. 11.
Fig
ig 11: File Browser Interface in Tableau
Step – 3(a): Select an appropriate file and open it, as shown in Fig. 12
Finally, Tableau
ableau displays the contents of your Excel file along with all its available
sheets as shown in Fig 14,
Fig
ig 14: Data Connection sheets in Tableau
Byy dragging and dropping a sheet, you'll be able to view its data and proceed with
further tasks, like generating descriptive statistics and many more. The subsequent
section relates to the generation
ration of descriptive statistics.
Let’s
et’s learn how to uuse the Summary Card for generating descriptive statistics,
Following
ollowing are the steps to generate descriptive statistics in Tableau:
Fig
ig 15: Descriptive Statistics in Tableau
Note: It is to be noted that, once enabled, the Summary Card automatically appears Tools for Data
Science an
on the right side of the visualization,, as shown in Fig. 16 below. Introduction
Fig
ig 16: Summary of Descriptive Statistics in Tableau
Till the time we learned, how to upload the data set? and how to generate
descriptive statistics? of the available dataset. Now, let’s extend our
discussion to the plotting of graphs and its analysis. The same is performed in
next section.
B) Creating
ing a Graph (Bar Chart: Category vs Price)
Step-by-Step Graph Plotting:
1. Drag Product_Category to the Columns shelf
shelf.
2. Drag Price to the Rows shelf.
3. From the Marks card, select Bar
Bar.
4. Tableau automatically plots a bar graph showing the total price for each
category.
Data Science – Allied 5. Drag ProductName to the Color option in the Marks card to differentiate
Areas
products.
6. Click on Show Mark Labels to display price values on the bars.
7. Add a title such as: “Category-wise Product Price Distribution”
Iff you hover over the join, you'll see that an inner join has been created using a
common field, namely Product ID.. An inner join means that the two tables share a
common on column, allowing their data to be combined seamlessly
seamlessly. Next, add the
PropertyInfo sheet,
heet, and you'll notice that it also gets joined automatically.
Note:
Table au automatically detects relationships between tables based on matching field
names. Since fields like Product ID share the same name across the connected
sheets,
heets, Tableau recognizes them as common keys and creates the join without
requiring
equiring manual input. This automatic behavior occurs only when field names are
consistent across tables.
The
he current view displays only one measure (the total quantity), which adds up to
10,096. With no dimensions or categories to break it down, the visualization isn't
very
ery insightful. Let's enhance it by adding additional fields to generate a more
detailed and interesting chart.
Fig. 26: Text table Dataset
ataset Connection with Tableau
The Tree
eee Map clearly shows that Furnishings has the highest number of items
ordered.
red. Public Areas and Housekeeping follow closely, with Maintenance
next, and Office Supplies having the lowest quantity.
Power
ower BI is a Microsoft application built to convert raw data into
valuable insights.
Itt enables users to create interactive dashboards, reports, and charts to
simplify data analysis and understanding. Whether you're in business,
research, or any data-driven
riven field, Power BI helps you identify patterns, track
trends,
rends, and make informed decisions more efficiently.
The process
ocess involves three main steps:
1. Power Query Editor – Prepare
repare and clean your data.
2. Dataa Modeling and Relationships – Establish connections between
data tables and organize your data.
3. Visualization –Build
uild interactive graphs and charts to explore and
showcase your data efficiently.
Fig.28: Getting
etting Started with Power BI
Step 2: Connect to Your Data: Open Power BI and click on “Get Data.”
You can connect to a variety of data sources such as Excel files, SQL
databases, or cloud services like Google Analytics and Salesforce. Select
your desired source and load the data. You can also combine data from
multiple sources into one report.
Step 3: Clean and Prepare Your Data: Data often needs cleaning before
analysis. Use Power Query to remove unnecessary rows, correct data
formats, and handle missing values. This ensures your dataset is accurate and
ready for analysis.
Step 4: Create Visualizations: Drag data fields into the report canvas to
start building visualizations. Power BI will automatically generate charts,
tables, or graphs based on the data. Choose visual types that best represent
your insights. Visuals help in identifying patterns and key metrics quickly.
Essential Steps:
1. Extract the Data – Begin by importing your data from various sources.
2. Transform the Data – Use Power Query to clean, modify, and prepare
your data, as well as to manage and adjust relationships between tables
as needed.
3. Apply Calculations with DAX – Use Data Analysis Expressions (DAX)
to perform calculations on your dataset.
4. Start Visualizing – Once the data is prepared, proceed to build your
visualizations.
5. Create
ate and Customize Visuals – Add elements like charts, graphs, and Tools for Data
Science an
cards,
rds, and modify them for clarity and effectiveness. Introduction
6. Publish to the Cloud – Upload your dashboard to the Power BI service
soo your team can access and interact with it
it.
Loading Data - To start handling data in Power BI, select the ‘Get Data’
option under the Home tab. This allows you to connect to a wide range of
data sources, including:
· Files – Import
mport data from formats like Excel or CSV.
· Databases – Establish
ablish connections to databases like SQL Server, Oracle,
or MySQL.
· Direct Query – Access
cess data directly from the source without importing
it.
· Online Services – Connect
onnect with platforms like Google Analytics,
Salesforce, or SharePoint.
· Live Connection – Establish
tablish a real
real-time link to data sources for always
up-to-date information.
After
ter importing the data into the application, you can begi
begin transforming and
modeling
odeling it. You can create new columns with custom configurations and
define your own calculated fields.
Before
efore building meaningful dashboards and analytical reports, it is essential
too prepare and structure the imported data. Raw data often contains missing
fields, formatting inconsistencies, or lacks the calculated values required for
analysis.
lysis. Power BI provides powerful tools for transforming, enriching, and
organizing
nizing data so that it aligns with analytical and business requirements.
Power
ower BI enables users to enhance the existing dataset by creating calculated
columns and measures. These user-definedefined fields help implement business
logic,
ogic, perform advanced calculations, and support accurate data insights that
may
ay not be directly availabl
available from the source. This process forms the core of
data
ata modeling
modeling, which ensures that the data is structured correctly for
visualizations
sualizations and interactions. The steps to create calculated
calcu columns and
measures
asures are listed below:
Step 1: Begin by clicking on the "Data" icon from the three options
available in the left
left-hand panel.
Step 2: In the "Table Tools" tab, go to the "Calculation" group and choose
"New
ew Measure" to create a custom calculation.
Step 3: Next, use the desired functions along with the table field names to
define your logic and generate a new column.
Tools for Data
Science an
Introduction
Step 4: Oncee this step is finished, your required columns or measures will be
set up. Use the "New Column" option when you need the column to remain
partt of the table and update automatically during data refreshes, as it becomes
a permanent part of the data model. On the other hand, measures are
calculated
lculated dynamically and function more like real real-time filters—making
them
hem more efficient since they aren't stored in memory.
Power BI also offers a built-in Power Query Editor that supports advanced
data cleaning and transformation, ensuring your data is well
well-prepared for
analysis and visualization .
Power Query - Power er Query is a robust tool that allows you to shape and
transform
ransform your data through operations such as filtering, grouping, and
pivoting.
voting. It also supports the creation of calculated columns and measures
using
ing the formula bar and a wide range of functions. After the data is cleaned
and prepared, you can move on to creating visualizations
visualizations. Following are the
steps
teps to be performed to prepare the data for its visualization.
Data Science – Allied Step 1: Click on the "Home" tab
ab from the top navigation menu.
Areas
Step
tep 2: In the Queries section, click on "Transform Data."
(This action will launch the Power Query Editor in a new window. If no data
source
ource is connected yet, a blank screen will be displayed.)
Step 3: On this screen, you can view, edit, transform, and clean your data as
required.
quired.
The Home tab
ab provides multiple tools for managing your data sources.
The "Transform" tab contains functions for adding, deleting, splitting, and
changing data types.
Step 4: By right-clicking
cking on your column, you can apply various modifications.
To modify a visualization already added to the report, select it and use the
"Format"
Format" option to customize its appearance.
For
or more advanced customization of your visualizations, you can use DAX
formulas too create calculated columns and measures as shown below in this
section.
Data Analysis Expressions (DAX) is a powerful formula language used in
Power BI (as well as in Excel Power Pivot and SQL Server Analysis
Services) to perform custom calculations on data after it has been imported
into the model. DAX helps users go beyond basic summaries by enabling the
creation of calculated columns, measures, and calculated tables, which
add new analytical meaning to existing data.
A calculated column works row by row and is mainly used to add new
derived data to a table, such as profit calculated from sales and cost. A
measure, on the other hand, is used to perform dynamic calculations such as
totals, averages, percentages, or ratios that automatically change based on
filters and slicers applied in the report. This dynamic behavior is what makes
DAX especially valuable for interactive dashboards.
In short, DAX acts as the analytical brain of Power BI. It allows analysts
to create meaningful metrics, perform complex calculations, compare
performance over time, and build interactive dashboards that support
informed decision-making. Even at an introductory level, understanding basic
DAX concepts like measures, filters, and aggregation functions provides a
strong foundation for advanced business intelligence and data analysis tasks.
Step
tep 1: Importing Data
To begin working in Power BI, the first task is to load the required dataset
into
nto the application. Power BI enables
ables data import from multiple sources,
including
ncluding Excel, CSV, databases, and online ine services. Click on the “Get
Data”
a” option located on the left side of the Home screen to browse and select
your
our data file. Once the desired dataset is chosen, it can be easily loaded into
the
he workspace for further processing and analysis.
· Select
elect the data source (Excel/CSV/Database/etc.)
Depending on the dataset size,, Power BI may require a short processing time
too extract the data. Once the preview is displayed, verify that the data has
beenn imported correctly and then click Load to bring the dataset into the
Power BI workspace.
You can click on any table or individual field to apply formatting options as
required.
quired. For attributes representing values such as dates, time, geographical
locations
ocations (city, state), percentages, or currency, the appropriate data type
andd formatting cann be assigned using the tools available under the
Modeling
ing tab.
Step
tep 3: Selecting an Appropriate Visualization
For
or developing the dashboard, we focus on five key fields: HiredYear,
RecruitmentSource
RecruitmentSource, Position, EmployerID, and Gender (Male/Female).
(Male/Female)
The first visual to be added is a Card visualization.
sualization. To create it, simply select
the Card icon from the Visualizations pane and place the relevant field onto
the
he canvas.
The same procedure wasas repeated to create the remaining Card visuals. These
additional
ional cards represent key summary indicators, including the total
number
umber of Recruitment Sources
Sources, total number of Positions, and the
Maximum Salary valuee present in the dataset. Together, these visuals
provide an at-a-glance
lance statistical overview that supports further analysis
within the dashboard.
Next, a Pie Chart and a Donut Chart are created to represent the
distribution of job positions and the gender-wise employment ratio,
respectively. These charts can be inserted by selecting the corresponding
icons
cons from the Visualizations pane and assigning the required fields to the
chart areas.
Finally,
inally, a Funnel Chart and a Stacked Bar Chart are added to visualize the
year-wise
wise hiring counts and the proportional
roportional contribution of different
recruitment
cruitment sources
sources, respectively. To enhance clarity and presentation
quality, formatting options such as titles,, data labels, axis properties,
legends,
egends, plot area settings, and color schemes are applied. These
adjustments
ments ensure the dashboard communicates insights effectively, as
illustrated
llustrated in the figure below.
Below is an example formula that calculates the Total Sales based on Price ×
Quantity:
, ∗
…………………………………………………………………………….
…………………………………………………………………………….
…………………………………………………………………………….
…………………………………………………………………………….
13.7 PYTHON
This section is an Introduction to Python Interface (Anaconda & Google
Colab) - Python IDLE, Jupyter Notebook, and Google Colab all support
Python development but differ in interface and capabilities, a brief
introduction to each is given below:
· Python IDLE: A simple, built-in IDE with a Python Shell for interactive
code and an Editor for scripts. It’s lightweight and good for beginners
and small projects.
· Jupyter Notebook: A web-based tool with a cell-based layout
supporting live code, text (Markdown), and visualizations. Ideal for data
science, research, and documentation.
· Google Colab: A cloud-hosted version of Jupyter with added benefits
like free GPU/TPU access, Google Drive integration, and real-time
collaboration. Best suited for machine learning tasks.
Before we start diving into Python coding, let's learn how to install Python
IDLE.
Now we can start writing Python code locally using the IDLE editor and the
steps to be followed are as follows:
1. Open Python IDLE
· Windows:
indows: Search for “IDLE” in the Start Menu and open it.
· macOS:
acOS: Use Spotlight to search for IDLE or open from Applications >
Python.
· Linux:
inux: Run idle or idle3 in the terminal.
2. Use the Python Shell
· A window opens
ns called the Python Shell
Shell, as shown below :
Top of Form
You cann enter Python code right here and view the output immediately.
· Example:
Or install
nstall via Anaconda (which includes Jupyter by default).
2. Start Jupyter Notebook
· Open a terminal or command prompt and enter the following command:
The JupyterLab interface features a toolbar on the left that includes a file
browser and a Git tab. The file browser displays the directory from which
JupyterLab
upyterLab was launched. On the right side, you can view and work with
files
es you’ve opened. To enable a split view, simply drag a file tab from the
top
op into another area of the workspace.
7. Create a New Notebook
· Click on “New” and select “Python 3” to launch a fresh notebook.
· A new tab opens with a blank notebook .
8. Write and Run Code
· Write your code in a cell.
Data Science – Allied Example:
ample:
Areas
Fig.66: Python
ython Google Colab Notebook Interface
4. Write and Run Code Tools for Data
Science an
· Enter
er Python code into a cell and press Shift + Enter to execute it. Introduction
Till
ll now we learned how to work with Google Colab, Jupyter Notebook,
IDLE etc., now lets understand
erstand how to work with python using datasets. To
understand the working we used Movie
MovieLens Small Dataset, you can browse it
on internet and download it from the link :
[Link]
tps://[Link]/datasets/movielens/ ; We'll use the file: [Link]
13.7.3 Steps
ps to Upload Dataset in Python IDLE
IDLE
DLE is a basic IDE and doesn’t have a built
built-in way to upload files — but
you can manually
anually place the file in your script folder
folder.
13.7.4 Steps
teps to Upload Dataset in Google Colab
Method 1: Upload from Local System
· Open a new Colab notebook.
· Use this code:
Now, lets apply the understanding and try to explore, How to Perform
Descriptive Statistics in Python?, Let’s use the MovieLens [Link] (has
user ratings of movies) for demonstration.
Step
tep 1: Upload Dataset
(Usee the same file upload method as discussed earlier in IDLE, Jupyter, or
Colab.)
Step
tep 2: Apply Descriptive Statistics
Ø Output Includes:
· count – total number of entries
· mean – average rating
· std – standard deviation (spread)
· min
min, max – lowest and highest values
· 25%
25%, 50%, 75% – percentiles (quartiles)
Ø Additional Analysis
· Mean
ean = average rating = (4 + 5 + 3 + 4.5 + 3.5) / 5 = 4.0
· Median = 4.0 (middle value) Tools for Data
Science an
· Mode = most frequent (if any) Introduction
After
ter generating descriptive statistics, plotting of graphs is an important step,
and we will learn this process of lotting graphs in python in subsequent
section.
13.7.6 Plotting
otting Graphs in Python (Using Matplotlib)
To visualize data in Python, the most widely used library is Matplotlib
Matplotlib,
however
er many other libraries are available.
Follow
ollow the steps below to plot a simple graph:
Step 1: Install Matplotlib (if not already installed)
pip install matplotlib
Step
tep 2: Import Required Libraries
import [Link] as plt
Step 3: Prepare Data
x = [1, 2, 3, 4, 5]
y = [10, 20, 15, 30, 25]
Step 4: Plot the Graph
[Link](x, y)
Step 5: Add Labels and Title
[Link]("X-Axis")
[Link]("Y-Axis")
[Link]("Simple Line Graph")
Step 6: Display the Plot
[Link]()
Let’s
et’s understand the complete process by executing a Hands-on
Example using mtcars dataset: we learned that Descriptive statistics help
highlight
ghlight key characteristics of a dataset, often using just one value. Creating
these
hese statistics is typically the first step after data has been cleaned and
prepared
pared for analysis. We've previously encountered examples like the mean
and median. In this section, we'll revisit those and introduce some additional
statistical
tatistical tools.
Firstly
irstly we begin with the finding of Central
entral tendency of data i.e. Measures of
Center,, which are
re statistical methods used to identify the "middle" or most
typical
ypical value in a numericalerical dataset. They provide insight into what a
standard
tandard or expected value might be. The most common measures of central
tendency
endency include the mean, median, and mode. Wee can compute the mean for
each
ch column in a DataFrame using:
We can also get the
he means of each row by supplying an axis argument: Tools for Data
Science an
Introduction
The median of a dataset is the value that splits the data into two equal parts
parts—
half the values are lower, and half are higher. It represents the center point
when the data is sorted. For this reason, it’s also known as the 50th
percentile.
centile. As shown previously, you can compute the median for each
column in a DataFrame using:
We can also calculate the median across each row by passing the argument
axis=[Link]
though both the mean and median indicate the central tendency of a
dataset,
et, they measure it differently and are not always the same. The median reliably
splits
plits the data into two equal halves, while the mean represents the arithmetic
average,
erage, which can be significantly affected by outliers or extreme values. In a
symmetrical
ymmetrical distribution, the mean and median are generally equal. This concept can
be better understood by examining a density plot.
Fig.67: Output
tput to check Normalization in dataset
Inn the plot shown above, the mean and median are nearly identical—both
identical
aree close to zero
zero—soo the red line representing the median overlaps the thicker
black
ack line that marks the mean.
Inn skewed distributions, the mean is more affected by the skew and is drawn
inn the direction of the skewness, while th the median is less influenced and
tends
ends to stay near the central point of the data.
Tools for Data
Science an
Introduction
The meanan is greatly influenced by outliers because it includes all values, even
the
he extreme ones, in its calculation. On the other hand, the median is more
resistant
sistant to outliers and stays fairly consistent, making it a more reliable
measure
easure of central tendency when outliers are present.
Data Science – Allied
Areas
Since
ince the median is less influenced by skewed data and outliers, it is regarded
as a "robust" measure. In cases where distributions are noticeably skewed or
containn extreme values, the median often gives a more reliable representation
of a typical value.
The mode is the value that appears most often in a dataset. Unlike the m
mean
and median, the mode can also be used with categorical data. Additionally, a
dataset
et may have more than one mode if several values share the highest
frequency.
equency. To identify the mode, you can use:
Tools for Data
Science an
Introduction
If a column has multiple modes—meaningmeaning several values ooccur with the same
highest frequency—then
hen all those values are returned as modes. On the other hand,
if a column has no repeating values,, and thus no mode, the result will be NaN.
Finding the Measures of Spread is also important in data analysis, lets understand
how to determine Measures
easures of spread, also known as dispersion, indicate how much
the values in a dataset differ from each other. While measures of center highlight a
typical
pical or average value, measures of spread show how widely the data points aare
scattered around that center.
Because
ecause these percentile values are frequently used to summarize a dataset,
they
hey are collectively known as the five-number
umber summary" consists of the
minimum,
inimum, 25th percentile (Q1), median (Q2), 75th percentile (Q3), and
maximum.
aximum. These key statistics offer a quick overview of the data’s
distribution
stribution and are also generated by the [Link]() function in pandas.
IQR
QR = Q3 - Q1.
The boxplots shown in fig 57 representing the summary and the interquartile Tools for Data
Science an
range (IQR) offers a clear view of the data’s distribution, central tendency, Introduction
spread, and possible outliers.
Two other commonly used measures of spread are variance and standard
deviation.
· Variance is the average of the squared differences between each data
point
nt and the mean. It provides insight into how much the values in a
dataset
et deviate from the average.
Since
ince variance and standard deviation rely on the mean, they can be heavily
influenced
nfluenced by skewed data and outliers. A more robust alternative is the
Median
edian Absolute Deviation (MAD). MAD AD is computed as the median of the
absolute differences from the median,
edian, making it a dependable measure of
spread
pread for datasets with non
non-normal distributions
tributions or extreme values.
Skewness
kewness and Kurtosis : Inn addition to measures of center and spread,
descriptive
iptive statistics also include metrics that describe the shape of a
distribution
stribution i.e. skewness and kurtosis. A brief introduction to both is given
below:
· Skewness indicates the asymmetry of the distribution—whether
distribution the data
is skewed to the left or right.
· Kurtosis reflects
eflects how much of the data lies in the tails versus the center
of the dist
distribution.
·
While
hile we won’t dive into the exact formulas, these metrics build on the
concept of variance:
· Skewness uses cubed deviations from
om the mean.
· Kurtosis uses fourth power deviations from
om the mean.
Tools for Data
Science an
Introduction
As observed from the results, the normal ormal distribution shows kurtosis close to
zero,, indicating a typical bell
bell-shaped curve. The uniform (flat) distribution has
negative
egative kurtosis
kurtosis, meaning it has lighter tails andnd a flatter peak. In contrast, the
other two distributions
distributions—with more data in the tails ils than in the center—exhibit
center
higher
igher kurtosis
kurtosis, reflecting
eflecting heavier tails and more extreme values.
· Inn Google Colab, import and use the file upload function:
· Then, read
ead the file using Pandas:
3. Handling Data
· Use Pandas: [Link](), [Link](), [Link]()
· Clean
lean data: drop missing, rename columns, filter rows.
· Transformation: df['new']
f['new'] = df['col1'] + df['col2']
Fig.72: R
R-Studio: Interface Layout (4 Panes):
The description of the 4-panes
anes shown in the above figure, is as follows: Tools for Data
Science an
Introduction
1. Source Editor (Top-Left)
· Write and edit R scripts (.R), RMarkdown (.Rmd), and notebooks.
· Supports
upports syntax highlighting, code folding, and auto
auto-completion.
Source Editor Thiss is where you write and edit scripts, RMarkdown
files, and notebooks. You can run code from here directly into the
Console
onsole (Pane 3).Useful for saving and reusing code.
2. Environment / History Pane
· Environment Tab: Shows all the objects, variables, and datasets
you've created.
· History Tab: Keeps a record of previously executed commands
commands.
· Helps
ps manage workspace variables and track command history.
3. Console (Bottom-Left)
· Run
un R commands interactively, and it Displays output, errors, and
messages.
4. Files/Plots/Packages/Help/Viewer
les/Plots/Packages/Help/Viewer (Bottom
(Bottom-Right)
· Files: Navigate your project directory.
· Plots: Display visualizations created in R.
· Packages: Manage installed libraries.
· Help: Access documentation for functions.
Now, after understanding the utility of all 44-panes of the user interface of R-
Studio, let’s extend our discussion
scussion to learn How to uuse RStudio's GUI for
Importing the Dataset. The
he steps to be performed are as follows:
Steps:
1. Open RStudio.
2. In the Environment
nvironment pane (top
(top-right), click “Import Dataset”.
3. Choose from:
o From Text (readr) or From Excel
Excel, or From CSV
4. Browse
rowse and select your file (e.g., [Link]).
5. A preview appears—check the options like separator, header, etc.
6. Click “Import”.
7. The dataset
taset is loaded into your environment (usually as a data frame).
Steps for Alternate Method
thod : Using Code (R Console or Script)
1. To Import a CSV File:
Function
unction Use Package
[Link]()
[Link]() Load CSV files Base R
[Link]()
[Link]() Load general delimited text files Base R
read_excel()
ad_excel() Load Excel .xlsx files readxl
read_csv()
ad_csv() Faster CSV loading readr
Table 4:: F
Functions of R
Note : Use getwd() to check the current
rent working directory,
directory and setwd("path")
to change it:
After
er understanding the basic working structure of R R-Studio, let’s extend our
discussion
cussion a next level, i.e. how to perform Descriptive analysis in R? the same is
addressed in next section.
13.8.2 Descriptive
Deescriptive Analysis in R Programming
Descriptive
criptive statistics in R involve summarizing and organizing data using
various tools such as charts,
s, graphs, tables, and spreadsheets.
spreadsheets This process
helps in presenting the data in a clearr and understandable format,
format allowing
for meaningful insights. Descriptive
ive analysis aims to summarize the main
features
eatures of a dataset. It is typically performed on smaller datasets,
providing
ng a foundational understanding that can also guide future trend
predictions
redictions based on existing data patterns.
Descriptive statistics of
often
en rely on two major types of measures i.e. Measures
of Central Tendency and Measures of Variability (Dispersion)
(Dispersion), the two are
discussed
scussed below
below:
1 . Measures of Central Tendency: Indicates the
he center or typical value of the
dataset and include:
· Mean (Average): Thee arithmetic average of all observations. (R-Syntax:
mean(data vector) ) Tools for Data
Science an
· Median: The middle value when the data is arranged in ascending order. Introduction
(R-syntax: median(data vector))
· Mode: The most frequently occurring value. R does not have a built-in
mode function, so it can be computed as: (R-Syntax: mode value <-
names(sort(table(data vector), decreasing = TRUE))
2. Measures of Variability (Dispersion): Reflects spread out of the data and it
includes:
· Range : Difference between the highest and lowest values.
§ R-Syntax : range(data vector)
§ diff(range(data vector))
· Interquartile Range (IQR) : Measures the spread of the middle 50% of
the data.
§ R-Syntax : IQR(data vector)
· Variance: Average of the squared deviations from the mean.
§ R-Syntax : var(data vector)
· Standard Deviation : Square root of the variance; shows spread in the
same units as the data.
§ R-Syntax : sd(data vector)
Example:
# Sample dataset
data vector <- c(12, 15, 14, 19, 20, 15, 18, 22, 15)
# Central Tendency
mean(data vector)
median(data vector)
mode value <- names(sort(table(data vector), decreasing = TRUE))
# Variability
range(data vector)
IQR(data vector)
var(data vector)
sd(data vector)
This code produces output that clearly shows the descriptive statistics,
Given: data vector <- c(12, 15, 14, 19, 20, 15, 18, 22, 15)
Measure Output
Mean 16.67
Median 15
Mode 15
Range 12 to 22
Range (Difference) 10
IQR 4
Variance 11.25
Standard Deviation 3.354102
Too upload a dataset in R, you can use several methods depending on your
file type
ype (e.g., CSV, Excel) and your preference (command line, GUI, etc.).
The commonly
comm used method are as follows:
1. Uploadi
Uploading a CSV File of dataset can be done by using the [Link]()
[Link](
function:
4. RStudio auto-generates
enerates the code and previews the data.
4. Uploading from the Web
Till
ll the moment we learned how to handle a dataset, now its time to extend
our discussion for data visualization and data interpretation. To understand
these
hese features we are going to refer to the renowned dataset of mtcars, a built
inn dataset in R. you can explore the dataset by applying commands, shown in
the respective
spective screenshots, given below:
Data Science – Allied
Areas
For
or the purpose of Data Visualization and interpretation we are going to use
the
he data visualization library ggplot() over the built-in
in dataset “mtcars”.
Here's how you can systematically perform data visualization and data
interpretation,
nterpretation, using R
Wee know that MPG, or miles per gallon, indicates how far a car can travel
using
ing one gallon of fuel. It’s a standard measure of fuel efficiency—higher
MPG
PG means better fuel economy and less fuel consumption per mile. So, by
plotting
otting the graph between MPG (Mileage) and Number of Cylinders,
Cylinders using
command shown below:
Fig.74
74: Scatter plot
NOTE – (Understanding
derstanding Displacement
Displacement): Engine displacement refers to the total
volume
olume that all cylinders in an engine can move. A larger displacement allows the
Data Science – Allied engine
ine to intake more air and fuel, which increases its potential to generate power.
Areas
However,
ever, the actual power output also depends on other internal components and
the engine's design. Displacement is usually measured in liters. For example, many
modern
odern vehicles have a 2.0
2.0-liter four-cylinder engine, meaning each cylinder has a
capacity of about 0.5 liters or 500 cc
cc.
Task-3:
3: How does cylinder effects mpg and displacement?
splacement? Again, a graph is
plotted
otted between the variables and the same is shown below:
Interpretation
nterpretation - Cars with fewer cylinders tend to have higher mileage and
better
etter fuel efficiency.
On the other hand, vehicles with more cylinders usually have larger engine
displacement,
placement, meaning the engine consumes more air and fuel. As a result, they are
generally
enerally less fuel
fuel-efficient.
Task-4:
4: Understanding the Relation between mpg, disp and hp,
hp to interpret
this
his a graph is plotted between the variables using the command, shown
below:
Tools for Data
Science an
Introduction
Generating
enerating Bar Charts: Barar charts are useful for summarizing and comparing
categorical
egorical variables or aggregated values (like
ke average mileage per cylinder
type).
Data storytelling can be used internally (for instance, to communicate the need for
product improvements based on user data) or externally (for instance, to create a
compelling case for buying your product to potential customers).
Data Storytelling: Best Practices for Dashboard Layout and Interactivity
In the age of data, business dashboards should be designed using data storytelling
techniques . Data storytelling converts metrics into clear narratives that guide the
reader to actionable conclusions.
According to this technique, a dashboard is not just about displaying data in a visually
appealing way, but about structuring the information in such a way that decision
makers and business professionals can quickly understand what is happening, why it
is happening and what actions to take about it.
A well narrated dashboard turns data into insights. By integrating storytelling
techniques into a dashboard, we make the audience remember the information better
and connect emotionally with it.
What is Data Storytelling in Business Dashboards?
Data storytelling is a technique that transforms dashboards into visual narratives that
explain what is happening, why it is happening and what decisions to make. It
combines data visualization, context and narrative structure to facilitate data-
driven decision making.
In a more theoretical sense, data storytelling is the art of constructing a
compelling narrative from complex information, relying on data visualization to
communicate a message to a given audience.
In the context of dashboards, it involve designing interactive dashboards that tell a
story: each visualization should support a key point and all together should take the
reader through a logical path, from a contextual introduction to an actionable
conclusion.
A well-constructed dashboard not only represents information, it tells a story and
conveys a meaningful message. At the end of the journey, the reader should
understand what is going on and be inspired to make an informed decision based on
data.
Benefits of Data Storytelling
In most organizations there is a gap between the abundance of data and the ability
to make decisions with it.
Often, we have so many reports and metrics that we fall into the "data overload
paradox" where, ironically, information overload leads to inaction.
Storytelling with data acts as a bridge to bridge that gap, structuring data into a logical
and persuasive narrative that makes it easier to interpret patterns, gain insights and,
above all, implement concrete actions.
In addition, telling stories with data helps to distils and simplify complex information.
Good data storytelling simplifies the complicated so that the audience can assimilate
it and make decisions more quickly and confidently. It also adds a "human touch" to
the numbers: contextualizing the numbers with stories or examples gives them
relevance and creates connection.
Every story starts with two key elements: what you want to tell and to whom.
Before designing a dashboard, Prepare:
• What is the main message or key insight the user should take away?
• Without a defined objective, we run the risk of building a generic report that
does not solve any specific need.
• Having a clear narrative objective will allow you to focus the dashboard on the
really relevant data, avoiding information overload.
2. Know your audience
Each professional profile interprets data differently. It is not the same to design a
dashboard for a CEO than for a marketing team, as they will have different needs.
When designing a dashboard, we must adapt the complexity, language and
visualization format to the profile of the audience.
Best practices for this technique:
maintain consistency in scales, colors and order between similar visuals. This
prevents the reader from having to "recalibrate" their attention on each graphic.
STORYTELLING WITH DATA
VISUALIZATION
Data visualization is more about storytelling and less about graphs and charts. Having the ability to
transform dense data into short, compelling, and actionable insights is a value. The principles of good data
storytelling are the topic of this chapter, followed by case studies on applying visualization techniques in
different domains and good data storytelling as shown in the given fig.9.1.
Data storytelling is a strong method that integrates data, graphics, and narrative to present insights
effectively as depicted in the above fig 9.2. It serves to convert raw data into actionable stories that can be
used to inform decision-making. The following are the essential components of data storytelling in data
visualization, detailed as follows:
i. Data: At the core of data storytelling is the data itself. Without high-quality, reliable data, even
the most beautiful story has no credibility.
ii. Visuals: Data visualization is very important in telling stories because the human mind absorbs
visuals 60,000 times faster than written words. Nicely designed graphics aid in presenting
complex data more easily.
Selecting the Best Visualization:
a. Bar Charts: Best utilized when comparing distinct categories.
b. Line Charts: Best for presenting trends over some time.
c. Pie Charts: These are utilized to represent proportions but must be utilized judiciously.
d. Heatmaps: Good for detecting patterns in big datasets.
e. Scatter Plots: Assist in demonstrating relationships and correlation between two
variables.
A global company employs an interactive business intelligence dashboard to monitor revenue, costs, and
profitability by region. The dashboard emphasizes financial KPIs using bar charts and line graphs as shown
in the fig 9.4, enabling executives to make fast, smart decisions.
2. Healthcare & Epidemiology:
The medical industry applies data visualization to monitor outbreaks of diseases, compare patient
demographics, and enhance treatment plans using various visualization methods like shown in the fig 9.5.
Strong visualizations assist policymakers, scientists, and healthcare workers in spotting trends and in the
effective distribution of resources.
i. Line Charts: It follows the spread of disease (e.g., new daily COVID-19 cases).
ii. Scatter Plots: It examines the correlation between risk factors (e.g., age and severity of disease).
iii. Bubble Charts: It contrasts the rates of disease prevalence in various regions or populations.
iv. Geospatial Maps (Choropleth Maps): It displays disease outbreaks geographically.
v. Histograms: These present patient age range or hospital admission patterns.
Example:
During the COVID-19 pandemic, Johns Hopkins University's interactive dashboard employed geospatial
maps and time-series graphs to map the virus's global spread. The tool enabled policymakers and the public
to monitor real-time infection rates and health responses.
3. Marketing & Consumer Analytics:
Marketing experts use data visualization to monitor consumer behaviour, measure campaign success, and
optimize ad strategies. Through visual analytics, companies can enhance customer interaction and achieve
maximum return on investment (ROI).
A climate scientist plots global temperature change over the last century with line graphs and
heatmaps and boxplots as shown in fig 9.7 to show the effects of climate change.
The objective of the interactive visualization is to allow users to have a fun and informative
means of viewing Netflix content trends, determining the best genres, evaluating budget patterns,
analyzing ratings
from IMDb, and knowing how different directors and languages affect the content on the streaming
service. The dashboard will be user-friendly, dynamic, and easy on the eyes.
Data visualization plays a vital role in contemporary digital interfaces, with the ability to facilitate
users' interaction with and exploration of large volumes of data. Netflix, being a top international
streaming platform, utilizes interactive data visualization methods to optimize user experience and
data discovery. This chapter discusses Netflix's real-time capabilities, mechanisms of user
interaction, and visual components, highlighting in fig 9.8; how this drive an intuitive and
interactive data-centric interface.
Real-time features and user interactions:
The Persuasive Use of Color in Design: Color in design, data visualization, or user interface is an
important tool to convey meaning, grab attention, and aid understanding. When applied correctly, color
helps the understanding of information; when wrongly applied—laying too much inference on it or using
colors unrelated to the topic or excessive colors—color clouds meaning, reduces readability, and makes
the data or design fail in the first place. With such a vast assortment of colors embedded into
visualization, there seems to be a feeling that such colors obstruct understanding and instead generate
visual clutter in view of the viewer.
One of the major sins in color usage is jamming too many elements into one chart or design as shown
in fig 5.1. The moment too many colors are used, it detracts the focus from key insights that are really
important for the users. A very clear example of this mistake comes in pie charts where differentiation
is attempted between numerous categories having separate colors for each. Introducing such a plethora
of colors to the chart tends to further concentrate on the data rather than clarifying it, confusing the
observer instead
Fig. 5.1: Encode too much color in graph
A fundamental problem connected with excessive color usage in visual representation is that it impedes
value comparison. If anyone attempts a value comparison having a wide palette of colors in their charts
or graphs, they are likely to focus less and less on finding relationships that matter when distractions
keep protruding among the clutter of hues. It is the distractive palette of colors that may be taking away
their attention, preventing them from arriving at useful information. Having a simple structure with few
carefully chosen colors works best for creating .
Fig. 5.2: Different colors for buttons, text, and backgrounds
An organized color scheme facilitates intuitive navigation. Bad color selections, like low contrast
between the background and text, can render content less readable, particularly for people with visual
disabilities. Eye-stressing colors like bright or too saturated ones lower readability. Furthermore,
neglecting color blindness makes important information unavailable to a part of users. Good contrast
and careful use of colors improve readability and provide inclusivity.
For instance: In a weather map depicted with a rainbow scale, the smooth temperature gradient can be
represented as having "hot zones" where the color changes suddenly although the transition might be
gradual. Employ a single-hue sequential gradient that smoothes its transition from light to dark and does
not include sudden changes. Nonmonotonic scales can mislead trends, particularly when audience
members believe equal color distances imply equal numerical differences. For Example: If the blue-
green-yellow-red scale is used to portray air pollution, individuals will be likely to presume that the
distance between blue and green is identical to the distance between yellow and red, although the
numerical difference is much greater.
5.1.2 Pitfall 4: Failing to Design for Color-Vision Deficiency:
Color vision deficiency (CVD) is a condition that causes individuals to have difficulty differentiating
between specific colors as a result of how their eyes perceive light. About 300 million people globally
are affected, which includes 1 in every 12 men and 1 in every 200 women. It affects designs as most
designers take it for granted that color is perceived by all people in the same manner, resulting in visual
elements that are impossible—or at least very hard—for colorblind people to understand. Overlooking
color accessibility can lead to unclear warnings, alerts, and UI elements .
Logical: This model goes one step further. As the name implies, it helps
break data into minor logical elements and build a detailed visual schema
of relationships between them. It aids organizations in streamlining
approaches for data consolidation and segmentation.
Physical: This one derives from a logical concept and helps describe how
data will be structured within a specic database management system. It
serves like a guide for data engineers during the visualization and
implementation of databases.
Common Data Modeling Techniques
#2 Relational Model
Dimensional models are mostly used in data warehouse design with the
goal to optimize a database for faster retrieval of information. They also
help remove redundancy and inconsistencies, thus contributing to better
data quality.
On the other hand, dimension tables are relatively smaller in size. They
contain descriptive information that helps understand data in the context
of the business environment.
Data Mart: The second type of data warehouse modeling is data mart. In
simple words, it is a portion of corporate-wide data that is useful to a
certain group of users. Overall, it focuses on an exact industry,
department, or a particular business area.
Take, for instance, marketing data. It can contain data entities such as
customers, items, campaigns, sales, and website analytics.
#5 Object-Oriented Model
For example, let’s consider a data model for an educational system. In this
case, you might have classes like student, teacher, course, and
department. Students then could have attributes such as ID, name, major,
and a list of enrolled courses. And courses could have attributes like
course ID, title, department, and a list of enrolled students, etc.
Last but not least, is agile data modeling. As the name suggests, it
combines agile software development approaches, like the Agile
Manifesto and Scrum, with the process of data modeling. It aims to
develop database architectures that are 7exible and responsive to
changes in the requirements.