Data Visualization(Data Science)
UNIT I: Introduction to Data Visualization
Definition and importance of data visualization-Role of data visualization in
decision making- Types of data (numerical, categorical, temporal, geographical)-
Data visualization process (data collection, exploration, analysis, visualization,
interpretation)-Challenges and limitations of data visualization
UNIT II: Visualization tools & Data Storytelling
Overview of Visualization Tools (e.g., Excel, Tableau, Power BI, Python)-
Comparing and contrasting features and Use Cases among these tools. Principles
of Data Storytelling: Narrative and Context-Best Practices for Dashboard Layout
and Interactivity Principles of Good Visualization Design - Understanding and
Using Colour in Visualizations – Importance of Data Modelling in Visualization.
Text Books
1. "Storytelling with Data: A Data Visualization Guide for Business
Professionals" Cole Nussbaumer Knaflic, Wiley; 1st edition, 2015.
2. “The Visual Display of Quantitative Information” by Edward Tufte, Graphics
Press USA; 2nd edition, 2001.
Reference Books
1. "Data Visualization: A Practical Introduction" Kieran Healy, Princeton
University Press, 2018.
2. "Analyzing Data with Power BI and Power Pivot for Excel", Alberto Ferrari
and Marco Russo, Microsoft Press; 1st edition, 2017.
3. "Microsoft Power BI Complete Reference", Devin Knight, Brian Knight,
Mitchell Pearson and Manuel Quintana, Packt Publishing; 1st edition, 2018.
UNIT-1
What is Data Visualization and Why is It Important?
Data visualization uses charts, graphs and maps to present information clearly and simply. It turns
complex data into visuals that are easy to understand. With large amounts of data in every industry,
visualization helps spot patterns and trends quickly, leading to faster and smarter decisions.
Common Types of Data Visualization
There are various types of visualizations where each has a unique purpose in data representation. Here
are the most common types:
1. Charts and Graphs: They are used to visualize data, with charts comparing data points
across categories or showing trends over time and graphs analyzing relationships between
variables to identify correlations, trends and outliers. Examples: Bar Charts, Line Charts,
Pie Charts, Scatter Plots, Histograms, Box Plots.
2. Maps: They are used to display geographical data which provides spatial context to trends
and patterns. Examples: Geographic Maps, Heat Maps
3. Dashboards: They combine multiple visualizations into a single interface which provides
real-time insights and interactive features for users to explore data.
Importance of Data Visualization
Data visualization is essential for understanding and communicating information effectively. Here are
some key reasons why it's important:
1. Simplifies Complex Data: It turns large and complicated data into visual formats like
charts and graphs, making the information easier to understand.
2. Reveals Patterns and Trends: It helps identify trends, relationships and patterns that are
not easily seen in raw data or tables.
3. Saves Time: Visuals allow quicker interpretation of data, helping users spot key
information at a glance instead of manually scanning through numbers.
4. Improves Communication: It makes it easier to explain data insights to others, especially
those who may not be familiar with the technical details.
5. Tells a Clear Story: Data visuals guide the audience through the information step-by-step,
making it easier to reach conclusions and make informed decisions.
Real-World Use Cases for Data Visualization
Data visualization is used across various industries to improve decision-making and drive results. Here
are a few examples:
1. Business Analytics: Used to monitor company performance, track KPIs and make
data-driven decisions by visualizing trends, sales and customer metrics.
2. Healthcare: Helps in analyzing patient records, tracking disease outbreaks and managing
hospital operations through easy-to-read charts and dashboards.
3. Sports: Used to visualize player statistics, team performance and match outcomes, helping
coaches and analysts improve strategies and training plans.
4. Retail and E-commerce: Enables tracking of sales, customer preferences and inventory
levels, helping businesses adjust stock and marketing efforts effectively.
Challenges in Data Visualization
1. Data Quality: Accuracy of visualizations depends on the quality of the data. If the data is
inaccurate or incomplete, the insights from the visualization will be misleading.
2. Over-Simplification: Simplifying data too much can lead to important details being lost
like using a pie chart that oversimplifies complex relationships between categories.
3. Choosing the Right Visualization: Using the wrong type of visualization can distort the
message. For example, a pie chart might not work well with many categories which leads to
confusion.
4. Overload of Information: Too much information in a visualization can overwhelm
viewers. It's important to focus on key data points and avoid clutter.
Best Practices for Effective Data Visualization
To ensure our visualizations are impactful and easy to understand we follow these best practices:
1. Audience-Centric Design: Tailor visualizations to our audience’s knowledge. A technical
audience may need detailed graphs while a general audience benefits from simpler charts.
2. Design Clarity and Consistency: Choose the right chart for our data and keep the
design clean with consistent colors, fonts and labels and also avoid clutter to ensure clarity.
3. Provide Context: Always provide context by including labels, titles and data source
acknowledgments. This helps viewers to understand the significance of the data and builds
trust in the results.
4. Interactive and Accessible Design: Make visualizations interactive with features like
tooltips and filters and ensure accessibility for all users regardless of device or visual needs.
ROLE OF DATA VISUALIZATION IN DECISION MAKING
How Does Data Visualization Simplify Complex Information for Informed Decisions?
Data visualization simplifies complex information by converting data into visual formats such as charts,
graphs, and dashboards. These visual representations make it easier for users to interpret data, identify
trends, and draw insights. By presenting data visually, organizations can communicate complex
information more effectively and facilitate quicker decision-making processes. Visualizations help in
highlighting key data points, comparing different datasets, and identifying patterns that may not be
apparent in raw data. This simplification of data allows stakeholders at all levels of an organization to
understand and act upon information more efficiently, leading to more informed decision-making.
The Evolution of Data Visualization: From Charts to Interactive Dashboards
Data visualization has evolved from simple charts and graphs to dynamic, interactive dashboards,
revolutionizing how we understand and interact with data. In the past, static visuals provided a basic
overview, but today's tools offer immersive experiences. Interactive dashboards allow users to
manipulate variables, drill down into details, and uncover insights that were once hidden in static
charts.
This evolution has been driven by advancements in technology, particularly in data processing and user
interface design. Modern data visualization tools leverage these advancements to create engaging and
informative visualizations that cater to a wide range of users.
As a result, decision-makers can now explore data more deeply, make informed decisions more quickly,
and communicate insights more effectively. This evolution continues to shape the field of data
visualization, opening up new possibilities for how we analyze and understand data.
Choosing the Right Data Visualization Tool for Your Business Needs
Selecting the right data visualization tool is crucial for effectively communicating insights. Several
factors should be considered when choosing a tool for your business needs. Firstly, consider the type of
data you're working with. For example, if you're dealing with geospatial data, a tool that specializes in
mapping may be more suitable.
Secondly, think about your audience and the level of interactivity required. If you're creating
interactive dashboards for business users, a tool with robust dashboarding capabilities would be ideal.
Additionally, consider the scalability and compatibility of the tool with your existing systems.
Lastly, evaluate the ease of use and learning curve of the tool, as well as the support and resources
available from the vendor. By carefully considering these factors, you can select a data visualization tool
that meets your business needs and empowers you to communicate insights effectively.
Data Visualization Best Practices for Communicating Insights Effectively
Effective data visualization is crucial for communicating insights clearly and persuasively. To achieve
this, it's essential to follow several best practices. Firstly, choose the right type of visualization for your
data. Bar charts are ideal for comparing values, while line charts work well for showing trends over
time.
Secondly, keep your visuals simple and uncluttered. Avoid using too many colors or elements that can
distract from the main message. Use color strategically to highlight key points and create visual
hierarchy.
Thirdly, provide context for your data to help viewers understand its significance. Use annotations,
labels, and captions to explain the data and provide additional insights.
Lastly, consider your audience and tailor your visualizations to their needs and level of expertise. By
following these best practices, you can create visualizations that effectively communicate insights and
drive informed decision-making.
The Impact of Color Theory in Data Visualization
Color plays a significant role in data visualization, influencing how data is perceived and understood.
Understanding and applying color theory can greatly enhance the effectiveness of your visualizations.
Firstly, use colors that are easily distinguishable to avoid confusion. Consider using a color palette with
contrasting colors for different data categories to ensure clarity.
Secondly, use color to highlight key data points or trends. By using a bright or contrasting color, you
can draw attention to important information and make it stand out.
Additionally, consider the psychological effects of color. For example, warm colors like red and orange
can evoke a sense of urgency or importance, while cool colors like blue and green can create a sense of
calmness or stability.
Overall, using color effectively can make your data visualizations more engaging, easier to understand,
and more impactful.
How Data Visualization is Transforming Business Intelligence?
Data visualization is transforming business intelligence by making data more accessible,
understandable, and actionable. Traditional methods of data analysis, such as spreadsheets and reports,
are being replaced by interactive visualizations that allow users to explore data in real-time and gain
deeper insights.
One key way data visualization is transforming business intelligence is by enabling faster
decision-making. With interactive dashboards and visualizations, decision-makers can quickly identify
trends, patterns, and outliers that may not be apparent in raw data, leading to more informed
decisions.
Additionally, data visualization is democratizing access to data within organizations. Non-technical
users can now easily create and interpret visualizations, reducing the reliance on data analysts and
empowering teams to make data-driven decisions.
Overall, data visualization is revolutionizing business intelligence by providing new ways to analyze and
interpret data, leading to improved decision-making, increased efficiency, and a competitive advantage
in today's data-driven world.
The Role of Data Visualization in Data-driven Decision Making
Data visualization plays a crucial role in data-driven decision-making by helping users understand
complex data sets and extract valuable insights. By presenting data visually, decision-makers can quickly
identify trends, patterns, and outliers that may not be apparent in raw data. This enables them to make
informed decisions based on data rather than intuition or guesswork, leading to better outcomes for
their organizations.
Furthermore, data visualization facilitates communication and collaboration among stakeholders.
Visualizations make it easier to share insights and findings, ensuring that everyone is working from the
same information and helping to build a consensus around key decisions.
Overall, data visualization is an essential tool for organizations looking to leverage their data to drive
informed decision-making. By making data more accessible and understandable, visualization
empowers decision-makers at all levels to make better, data-driven choices that lead to improved
performance and business success.
Conclusion
In conclusion, data visualization plays a crucial role in business intelligence by simplifying complex
information, enabling data exploration, facilitating data-driven decision-making, improving
communication and collaboration, and driving business growth and innovation.
By presenting data visually, organizations can communicate insights more effectively, leading to
quicker and more informed decision-making processes. Data visualization also promotes collaboration
and alignment across teams, as visualizations can be easily shared and understood by all stakeholders.
Additionally, data visualization enables organizations to identify emerging trends, predict customer
behavior, and spot market opportunities, ultimately driving business growth and fostering innovation.
As organizations continue to collect and analyze large volumes of data, the importance of data
visualization in business intelligence will only continue to grow, helping organizations stay competitive
in today's data-driven world.
Types of Data:
1. Categorical Data:
Categorical data represent discrete, qualitative information. Examples include Gender, color, marital
status.
Application in Data Science: Categorical data is used for classification tasks and creating meaningful
categories.
Visualization Techniques: Bar charts, pie charts, and stacked bar charts are commonly used to visualize
categorical data.
2. Numerical Data:
Numerical data represents quantitative information. Examples include Age, height, temperature.
Application in Data Science: Numerical data is used for statistical analysis, regression, and prediction
models.
Visualization Techniques: Histograms, box plots, and scatter plots are effective visualization techniques
for numerical data.
3. Time Series Data:
Time series data represents data points collected over a period of time. Examples include Stock prices,
temperature readings, website traffic.
Application in Data Science: Time series data is used for forecasting, trend analysis, and anomaly
detection.
Visualization Techniques: Line graphs, area charts, and seasonal decomposition plots help visualize
time series data.
4. Text Data:
Text data comprises unstructured textual information. Examples include Tweets, customer reviews,
news articles.
Application in Data Science: Text data is used for sentiment analysis, natural language processing, and
text mining.
Visualization Techniques: Word clouds, bar charts, and scatter plots with text labels are commonly
employed for visualizing text data.
5. Image Data:
Image data consists of visual information in the form of pixels. Examples include Photographs, medical
scans, satellite imagery.
Application in Data Science: Image data is used in computer vision, object detection, and image
recognition.
Visualization Techniques: Heatmaps, image grids, and image overlays are effective visualization
methods for image data.
6. Geospatial Data:
Geospatial data represents geographic information. Examples include GPS coordinates, city
boundaries, population density.
Application in Data Science: Geospatial data is used for mapping, spatial analysis, and location-based
services.
Visualization Techniques: Choropleth maps, heatmaps, and scatter plots with geographical coordinates
are common visualization techniques for geospatial data.
7. Numerical/Quantiative Data
● Continuous: Data that can take on any value within a specified interval. It can be measured,
and can take on any value, including fractions or decimals.
○ Example: A measure of greenhouse gas emissions in the US.
● Discrete: Data that has distinct values that can be counted. It can only take on integer values.
○ Example: The number of households in a county.
Categorical/Qualitative Data
● Nominal: Data that has distinct values representing different groups or categories with no
inherent ranking or order.
○ Example: The names of different counties.
● Ordinal: Data that has distinct values representing categories with a meaningful order or
ranking.
○ Example: Education levels, such as elementary, middle, high school, college, and
post-graduate.
● Binary: Data that has values that can only be one of two distinct categories.
○ Example: success/failure, yes/no
8. Temporal Data
● Time Series: data points collected over an interval of time and indexed by time.
○ Example: Daily weather data.
Spatial Data
● Raster: Uses pixels to represent geographic information. Each pixel value represents an area on
Earth.
○ Example: Satellite images
● Vector: Represents geographic features as points, lines, and polygons to define the shape of a
spatial object.
○ Example: A polygon representing the boundaries for a state.
Common Data Visualization Techniques:
1. Bar Charts and Histograms: Suitable for visualizing categorical and numerical data distributions.
2. Scatter Plots: Effective for showcasing relationships between two numerical variables.
3. Line Graphs: Useful for displaying trends and patterns over time.
4. Heatmaps: Ideal for representing matrix-like data with color intensity.
5. Box Plots: Provide a summary of numerical data's distribution and identify outliers.
6. Word Clouds: Visually summarize textual data by displaying frequently occurring words.
7. Geographic Maps: Visualize spatial data and patterns on a map.
Choosing the Right Visualization Technique:
Selecting an appropriate visualization technique depends on several factors, including:
- Matching Data Types to Visualization Techniques: Ensure the chosen visualization method aligns
with the data type and the insights you want to convey.
- Considering the Objective and Audience: Tailor the visualization to the intended purpose and the
target audience's level of technical understanding.
- Design Principles for Effective Data Visualization: Pay attention to aspects like color choice, labeling,
and chart layout to create visually appealing and informative visualizations.
Conclusion:
Data is the backbone of Data Science, and understanding the various types of data is essential for
extracting meaningful insights. Through effective data visualization techniques, Data Scientists can
unlock patterns, trends, and relationships, leading to informed decision-making and valuable insights.
By incorporating the appropriate visualization methods based on data type and purpose, Data
Scientists can communicate their findings more effectively and empower stakeholders with actionable
insights.
TYPES OF CHARTS
1. Bar chart
A bar chart visually represents data using rectangular bars or columns. Here, the length of each bar
corresponds proportionally to its value. You can present these bars horizontally or vertically.
When to use bar charts?
Bar charts are excellent for comparing the values of different categories or groups. Apart from that,
these types of charts are also helpful in showing the distribution of data across different categories.
Best practices for bar charts:
Clearly label each bar and axis with concise labels
Limit the number of bars and categories to avoid cognitive overload
Purposely use colors to highlight key points and convey meaning
2. Histogram
A histogram visualizes a single, continuous dataset. Although histograms and bar charts are used
interchangeably, they differ in practice. For instance, a bar graph is a plot of a single data point (a sum,
average, or other value) for each category while a histogram is a plot of a range of data.
When to use histograms?
Think of histograms as a more informative way to view the distribution of values in a dataset. It is ideal
for visualizing the spread and variation of the data, helping you identify outliers or unusual data points
that fall far outside the normal range of the data.
Best practices for histograms:
Choose an appropriate number of bins and their sizes
Consider using consistent intervals for bins to maintain uniformity and accuracy in data
representation
Histograms are less effective with smaller datasets, so make sure you have sufficient data
points.
3. Column chart
Column charts are the simplest, most versatile type of visualization used in data analytics. The
horizontal chart displays your data in bars proportional to the values they represent.
When to use column charts?
More often than not, column charts effectively compare data across different categories. They are also
helpful in displaying rankings and order in a dataset, allowing viewers to identify trends quickly.
Best practices for column charts:
Minimize distracting visual elements, such as 3D effects or excessive gridlines
Focus on presenting key data points for a better understanding
Use contrasting colors to highlight specific columns.
5. Line chart
A line chart connects distinct data points through straight lines. Its best use case is to illuminate trends,
patterns, and variable changes.
When to use line charts?
This type of chart helps measure how different groups relate to each other. This type of chart is also
effective for demonstrating progression, making them suitable for scenarios like project timelines,
production cycles, or population growth.
Best practices for line charts:
Make sure the data you're representing has a logical order
Add context through annotations and labels
If the dataset is large, use transparency or spacing to improve visibility
5. Pie chart
A common but limited type of graph is the pie chart. It is a circular, statistical graphic that divides data
into slices, where each slice represents a percentage or proportion of the whole. You can create your
own pie chart to illustrate how each category contributes to the overall dataset.
When to use pie charts?
This classic chart type is effective when you want to illustrate the proportion of each category in the
dataset. The chart is suitable when you have limited categories, ideally less than six or seven.
Best practices for pie chart:
Keep the number of slices limited to maintain clarity
Clearly label each slice with clear text.
Follow consistency so viewers associate colors with specific categories
What’s hiding in your data? Find out by signing up for your demo
Scatter plot
Scatter plots are types of visualization that show a collection of data points ‘scattered’ around the
graph. The data points can be evenly or unevenly distributed.
When to use scatter plots?
Scatter plots are ideal for exploring relationships and patterns between two continuous variables. They
can help you identify trends, correlations, or potential clusters in the data.
Best practices for scatter plots:
Highlight outliers if present in the graph to showcase data distribution.
Add a trendline to highlight the relationship between variables.
Consider using different colors or marker sizes for overlapping points.
Heatmap chart
Heatmap charts are a type of map data visualization that uses a system of color coding to represent
value. Each cell in the matrix is assigned a color based on the value it holds.
When to use heatmap charts?
This type of chart is commonly used to establish relationships between two variables across a grid. In
the example above, the intensity of the colors in the map clearly demonstrates the variables, making it
easy to identify patterns and trends.
Best practices for heatmap charts:
Choose an intuitive color palette that effectively conveys the magnitude of values.
Use visual cues to highlight significant values in the heatmap.
Utilize the design principle of white space to prevent overcrowding.
How do you choose the right type of charts and graphs
Let’s be real—attention spans are short. Like, shorter-than-a-goldfish short. If your chart is confusing,
cluttered, or just plain wrong for the data, you’ll lose your audience before they even get to your
insight.
So how do you choose the right chart? The one that makes your point land and sparks action?
Here are important steps to help you nail it every time:
Know your goal: Are you comparing values, showing trends, highlighting proportions, or
exploring relationships? Your objective should drive the chart type.
Understand your data: Look at the size, structure, and distribution of your dataset.
Different chart types work better for different data formats.
Consider your audience: Tailor your visuals to their level of data literacy. Simpler charts
often resonate more with non-technical viewers.
Prioritize clarity: Choose the chart that makes your point instantly clear. Don’t make your
audience second-guess.
Stick to visual best practices: Use consistent colors, intuitive labels, and readable axes.
Design with intention, not decoration.
Use interactivity wisely: When possible, let users explore with filters, drill-downs, or
tooltips—especially in dashboards and digital reports.
DATA VISUALIZATION PROCESS
The data visualization process transforms raw, complex data into actionable insights through a
structured cycle:
collecting relevant data, cleaning and exploring it to identify patterns, analyzing to extract insights,
visualizing through charts/maps, and interpreting to inform decisions. This iterative, analytical-visual
approach helps spot trends, outliers, and relationships.
● Data Collection: Gathering raw data from various sources and ensuring it is accurate and clean.
● Exploration (EDA): Using summary statistics and initial visual plots to understand data
structure, detect outliers, and find initial patterns.
● Analysis
:
Applying statistical methods or models to test hypotheses, identify trends, and derive
meaningful insights from the data
.
● Visualization: Translating analyzed data into graphical representations (charts, maps, graphs)
to make insights easily understandable.
● Interpretation: Drawing conclusions, creating narratives, and making data-driven decisions
based on the visual findings.
What is Data Cleaning?
Data cleaning is the process of preparing raw data by detecting and correcting errors so it can be
effectively used for analysis. It is a foundational step in data preprocessing that ensures datasets are
suitable for analytical, statistical and machine learning tasks.
● Eliminates duplicates, missing values and inconsistencies that can distort analysis
● Enables organizations to make decisions based on accurate and up to date information
● Reduces effort and time spent handling errors during data analysis and reporting
● Strengthens the reliability of analytics, AI and machine learning outcomes
● Supports adherence to data quality standards and regulatory requirements
Common Data Anomalies
Data quality issues can arise from human errors, system failures or problems during data collection and
integration. These issues if not addressed can significantly affect analysis accuracy and lead to
misleading interpretations. Some of the most common data quality challenges include:
● Missing values: Incomplete records can reduce statistical power and introduce bias into
analysis.
● Duplicate records: Repeated entries may overrepresent certain observations resulting in
skewed outcomes.
● Incorrect data types: Mismatched formats, such as text stored in numeric fields can cause
calculation errors and analysis failures.
● Outliers and anomalies: Extremely high or low values can distort statistical measures and
influence model performance.
● Inconsistent formats: Variations in date formats, text casing or measurement units can
create issues when merging or comparing datasets.
● Spelling and typographical errors: Errors in text fields can lead to incorrect grouping,
classification or interpretation of categorical data
Data Cleaning Process
Data cleaning ensures that your dataset is accurate, consistent and ready for analysis. It involves
assessing quality, correcting errors and preparing data for reliable insights.
1. Assess Data Quality
The first step in data cleaning is to assess the quality of your data. This involves checking for:
● Missing Values: Identify any blank or null values in the dataset. Missing values can be due
to various reasons such as incomplete data collection, data entry errors or data loss during
transmission.
● Incorrect Values: Check for values that are outside the expected range or are inconsistent
with the data type.
● Inconsistencies in Data Format: Verify that the data format is consistent throughout the
dataset.
2. Remove Irrelevant Data
Removing irrelevant or duplicate data ensures the dataset is clean, accurate and meaningful, preventing
skewed analysis and improving overall quality.
● Identify duplicate entries using techniques like sorting, grouping or hashing.
● Remove duplicate records to ensure each data point is unique and correctly represented.
● Detect redundant observations that do not add new information to the dataset.
● Eliminate variables or columns that are irrelevant to the analysis and do not provide useful
insights.
3. Fix Structural Errors
Structural errors occur when data formats, naming conventions or variable types are inconsistent
which can affect analysis accuracy. Correcting these issues ensures uniform and reliable data
representation.
● Standardize data formats to maintain consistency in dates, times and other data types
across the dataset.
● Correct naming inconsistencies in column names, variable names or labels to ensure clarity
and uniformity.
● Ensure consistent data representation such as using the same units for measurements or the
same scales for ratings.
4. Handle Missing Data
Missing data can introduce bias and reduce the reliability of analysis. Properly addressing missing
values helps maintain the integrity of your dataset.
● Impute missing values using statistical methods such as mean, median or mode to fill gaps.
● Remove records with missing values when the missing data is extensive or cannot be
accurately imputed.
5. Normalize Data
Data normalization organizes the dataset to reduce redundancy and ensure consistency making it easier
to manage and analyze.
● Split data into multiple tables, with each table storing specific types of information.
● Ensure consistency across the dataset to support efficient querying and accurate analysis.
6. Identify and Manage Outliers
Outliers are data points that deviate significantly from the rest of the dataset and can affect analysis
accuracy. Properly handling them ensures more reliable insights.
● Remove outliers that result from errors or are not representative of the population.
● Transform extreme but valid outliers to reduce their impact on the analysis.
Data Cleaning Strategies
To ensure high-quality and reliable data follow these best practices:
● Understand the data: Know the source, structure and domain of the data to identify
potential quality issues and determine appropriate cleaning actions.
● Document the process: Keep records of decisions, methods, assumptions and rules
applied during data cleaning.
● Prioritize critical issues: Focus first on major quality problems that could have a systemic
impact on analysis or decision-making.
● Automate where possible: Use scripts or tools for repetitive cleaning tasks to improve
efficiency and consistency.
● Collaborate with domain experts: Engage stakeholders or domain specialists to validate
that the cleaned data meets business requirements.
● Monitor and maintain: Continuously track data quality and perform cleaning
periodically to ensure long-term accuracy and reliability.
Limitation
Cleaning data can be difficult due to several factors:
● Large datasets are harder to clean because of their size, requiring efficient tools and
techniques.
● Data from multiple sources often have different structures and formats making integration
and cleaning complex.
● Data cleaning is a continuous process as new data must be regularly assessed and
maintained for accuracy.
Tools for Data
TOOLS FOR DATA SCIENCE – Science an
Introduction
AN INTRODUCTION
13.1 Introduction
13.2 Objectives
13.3 Tools for Data Science
13.4 Datasets and File Formats
13.5 Basics of Tableau
13.6 Basics of Power BI
13.7 Basics of Python
13.8 Basics of R-STUDIO
13.9 Summary
13.10 Solutions/Answers
13.11 Further Readings
13.1 INTRODUCTION
In today’s digital era, data is often considered as the new oil—an invaluable
resource powering decision-making across industries like healthcare, finance,
retail, education, and manufacturing. Yet, the true power of data lies not in its
raw form, but in “how effectively we can interpret and apply it?”. That’s
where data science steps in: a multidisciplinary domain blending
mathematics, statistics, computer science, and domain expertise to transform
chaotic data into meaningful knowledge.
At the heart of this transformation from data to useful information, there
exists a powerful suite of tools—technologies that enable data professionals
to collect, process, analyze, visualize, and communicate insights with
efficiency and precision. These tools form the foundation of the data science
workflow, streamlining everything from basic data handling to complex
predictive modelling and real-time reporting.
This Unit, serves as your foundational guide to four of the most widely
adopted tools in the data science pipeline:
· Tableau – An industry-leading platform for visual storytelling and
dashboard creation.
· Power BI – Power BI – A robust business intelligence tool from
Microsoft that lets you create interactive reports and easily connects with
other Microsoft products for smooth data analysis.
· Python – A versatile and easy-to-learn programming language, Python is
widely favoured by beginners and professionals alike for data tasks—
thanks to powerful libraries like Pandas, NumPy, Matplotlib, and Scikit-
learn that simplify data handling and machine learning.
· R – A statistical powerhouse designed for analytical rigor, favoured in
academia and research-heavy environments.
The sections covered in this unit, introduces not just the technical skills
needed to work with these tools, but also offers practical exposure to their
user interfaces, dataset handling methods, and data visualization techniques.
The goal is to equip students, aspiring data scientists, and professionals with
both the confidence and competence to begin their journey in data science.
Whether you're analysing trends, forecasting outcomes, or simply making
smarter decisions, mastering these tools is your first step into a data-driven
future. Let’s begin.
13.2 OBJECTIVES
Upon completing this unit, learners will be able to:
1. To understand the core functionalities of Tableau, Power BI, Python, and
R.
2. Identify and work with various dataset formats such as CSV, Excel, and
JSON.
3. Upload, preprocess, and clean datasets for analysis.
4. Design and interpret meaningful data visualizations.
These objectives are designed to empower learners to not only master the
basics of data tools, but also apply their skills across academic, research, or
industry settings with clarity and confidence.
13.3 TOOLS FOR DATA SCIENCE
In the Unit 9 and 10 we had already explored MS-Excel as a tool for Data
Science, but today in the rapidly evolving world of data, many other tools are
available, these tools are essential for converting raw information into
actionable insights. Apart from MS Excel, the following tools for Data
science assist in collecting, cleaning, analyzing, and visualizing data across
various domains.
Below are four foundational tools widely used in the industry and academia:
· Tableau: A powerful and easy-to-use visualization tool that lets you
build interactive reports and dynamic dashboards—no coding required.
It’s perfect for turning data into clear stories and sharing insights with
people who may not have a technical background.
· Power BI: Created by Microsoft, Power BI is a strong business
intelligence tool that works smoothly with Excel and allows users to
build real-time dashboards—making it ideal for enterprise-level data
analysis and reporting.
· Python (Anaconda/Google Colab): A general-purpose programming
language with extensive libraries like Pandas, NumPy, and Matplotlib
for data manipulation, analysis, and visualization. Anaconda simplifies
local setup, while Colab provides a free cloud-based environment.
· R (RStudio): R is a programming language tailored for statistical
analysis and data visualization. It’s widely used in academic research and
data-centric projects for its strong capabilities in handling complex data
and producing high-quality graphics. RStudio makes coding, visualizing,
and managing data efficient and user-friendly.
These tools are essential to today’s data science process, and mastering them
equips students and professionals with the skills to tackle real-world
challenges efficiently and confidently
13.4 DATASET AND FILE FORMATS
A dataset is an organized set of data used for analysis and building models.
It's usually laid out like a table, where rows show individual entries and
columns represent different features or variables. Depending on the type and
use of the data, datasets can be stored in different formats—one of the most
common being CSV (Comma-Separated Values),which is widely used for
storing tabular data; Excel files (.xls, .xlsx), which allow multiple sheets and
more formatting options; JSON and XML, which are suitable for storing
nested or hierarchical data; SQL databases, used for managing structured
data; Parquet and ORC, which are efficient columnar storage formats often
used in big data systems; and TXT files, which store plain text with custom
delimiters.
Format Description / Strengths / Use Cases
A simple, plain-text format where each row is a new line and
CSV (Comma- columns separated by commas. Very widely supported by
Separated Values) spreadsheets (Excel, LibreOffice), data-analysis tools,
databases. Great for tabular data, easy to read and edit.
Spreadsheet format allowing multiple sheets, richer metadata /
formatting than CSV. Handy if you want human-friendly
Excel (.xls / .xlsx) editing, multiple tables in one file, or need features like
formulas or embedded metadata. (Often exported to CSV /
other formats for ML workflows.)
A text-based format supporting nested / hierarchical data
(objects, arrays). Useful for semi-structured or hierarchical data
JSON (JavaScript
(e.g., configuration data, nested records, API outputs).
Object Notation)
Readable by humans and easily parsed by most programming
languages.
Another text-based, hierarchical format. Sometimes used when
XML (eXtensible
data has complex nested structure or needs to follow a schema.
Markup
Less common in modern ML tasks than JSON, but still used in
Language)
legacy systems or data exchange.
SQL Databases / For structured datasets that need efficient querying,
Relational relationships, updates — more dynamic than flat files. Useful
Format Description / Strengths / Use Cases
Databases when data is large, relational, or needs concurrent
access/updates.
Designed for large-scale data (big data / analytics), these
Parquet, ORC,
formats store data in a columnar fashion — which can lead to
and other
better compression, faster read/write especially when only
columnar / big-
some columns are needed. Well-suited for large datasets,
data formats
distributed systems, data warehouses, and big-data frameworks.
Plain text (.txt) /
Sometimes used for very simple or custom data, logs, semi-
delimiter-
structured text. Flexible but requires agreed upon structure or
separated (custom
parsing logic; often not ideal for complex datasets.
delimiters)
Datasets can be sourced from various platforms depending on the need and
application. Popular sources include data-sharing platforms like Kaggle and
the UCI Machine Learning Repository, which host numerous datasets for
academic and practical use. Government portals such as [Link] provide
open-access datasets for public use. Additionally, APIs from platforms like
Twitter or Google Maps offer real-time and structured data. Other methods
include web scraping, where data is extracted from websites, and data
collected through IoT devices or sensors, which continuously generate data
in various domains. Some of the usefil links to gather the datasets are given
below:
· Kaggle — [Link]
· UCI Machine Learning Repository — [Link]
· Google Dataset Search — (search via
[Link]
· Open Government Data Platform India ([Link]) — [Link]
· [Link] (US) — [Link]
· Community/curated lists like “Awesome Public Datasets” on GitHub or
[Link]
We will discuss, the usage of datasets along with the utilization of various
tools mentioned above. Let’s begin with Tableau first, in the next section.
13.5 TABLEAU
Tableau is a user-friendly and a powerful tool for creating interactive data
visualizations. It connects easily to a wide range of data sources and lets
users explore and present their data through clear, engaging visuals with its
simple drag-and-drop interface, even beginners can build charts, graphs, and
dashboards—no coding needed. Widely used by data analysts and scientists,
Tableau is popular in industries like healthcare, tech, and e-commerce to
support smart, data-driven decisions. Once installed, users can get started by
activating their license or signing in with their Tableau account.
448
Tableau Installation: Tableau is a free, easy
easy-to-use tool that lets users build Tools for Data
Science an
engaging and interactive visualizations
visualizations—all without writing any code. It Introduction
helps transform raw data into a more readable and insightful format. Tableau
has become a widely used tool in the business analyt
analytics space.
Users can connect with others across the globe, share dashboards and data
visualizations, and even publish their work on websites, blogs, or social
media platforms.
Here in this section, we will try to cover most of the features of Tableau in
brief.
ief. To begin with the Prerequisites of Using Tableau.
As a Prerequisites for Using Tableau
Tableau, it only requires Basic computer skills,
such as running programs and navigating software interfaces
interfaces, along with
Familiarity with spreadsheet applications
applications, lastly your Willingness to explore
and learn new tools.
Firstly, Let’s learn how to Install Tableau, in order to Get Started with
Installation, simply go to the official website at [Link] and
follow the step-by-step
step instructions provided there.
The stepwise process is also given here, the Fig.1: Tableau Installation shows
the interface available at official website at [Link]
Fig.1: Tableau Installation
At official website at [Link]
[Link] Navigate to the Products
section, where you’ll find four main options as shown in Fig.2: Tableau
Desktop Products:
Data Science – Allied
Areas
Fig.2: Tableau Desktop Products
Refer to Fig.2: Tableau Desktop Products,, it covers Tableau Desktop and
Tableau Prep
Prep, Tableau Server and Tableau Online as its major
components, a brief description of these components is as follows:
· Tableau Desktop
Desktop: A standalone desktop application designed for
individual users to create and analyze data visualizations.
· Tableau Prep
Prep: This includes two components——Tableau Prep Builder,
is used to design and build data preparation workflows, while Tableau
Prep Conductor allows organizations to schedule, manage, and track
these workflows efficiently across teams.
Fig.3: Tableau Prep
· Tableau Server:: This is primarily designed for enterprise
enterprise-level use, Tools for Data
Science an
enabling organizations to securely share, manage, and collaborate on Introduction
data visualizations across teams.
Fig.4: Tableau Server
· Tableau Online is a public platform, so any content you upload can be
accessed by other users of the software. As a result, your work is no
longer private or confidential.
Fig.5: Tableau Online
You can Choose Tableau Desktop and start your free trial
trial, The interface of
Tableau Desktop is shown below in Fig.6.
Data Science – Allied
Areas
Fig.6: Tableau Desktop
Choose Tableau Desktop and to start your free trial,
trial perform the following
steps:
· Choose the software version that matches your system requirements and
download the free trial.
· If you prefer not to use the trial, you can download Tableau Public
instead. Once done, Tableau will be installed on your system.
Once Tableau is installed you will see the user interface as shown in Fig. 7,
below
Fig.7: Tableau Installed
13.5.1 Interface of Tableau - Tableau’s main interface window Tools for Data
Science an
Introduction
The Tableau workspace consists of several components, including menus, a
toolbar, the Data pane, shelves, cards, and one or more sheets. These sheets
may be worksheets, dashboards, or stories as shown in Fig.8 below, but for
detailed guidance on creating dashboards or working with stories, you can
explore the sections dedicated to dashboards and the Story workspace.
If you're working with Tableau in a web browser, check out “Creators: Get
Started with Web Authoring” and “Tour Your Tableau Site” for guidance.
Fig.8: Interface of Tableau
The components of the interface of Tableau are shown in Fig. 8 numbered
from A to I, their respective elaboration and utility is given below:
A. Workbook Title –A A workbook contains multiple sheets, which may
include worksheets, dashboards, or stories. For more information, see the
section on Workbooks and Sheets
Sheets.
B. Cards and Shelves – These are areas where you can drag data fields to
create and personalize your visualizations.
C. Toolbar – The toolbar offers quick access to various commands, as well
as tools for analysis and navigation.
D. View Area – This is the primary workspace where yo
your visualization, or
"viz," is displayed.
E. Home Icon – Click this button to go back to the Start page, where you
can connect to new data sources. For more information, refer to the Start
Page section [1].
F. Side Bar –Located
Located within worksheets, this section contains both the
Data pane and the Analytics pane.
G. Data Source Tab – Click here to navigate to the Data Source page,
where you can view, explore, and manage your data. For more details,
Data Science – Allied refer to the Data Source Page section.
Areas
H. Status Bar – Shows details and status of the current visualization or
view.
I. Sheet Tabs – Each tab represents a different sheet in the workbook, such
as a worksheet, dashboard, or story. More info is available in Workbooks
and Sheets.
13.5.2 Overview of Tableau Toolbar Buttons
While building or modifying a view, the toolbar located at the top provides
quick access to frequently used tools and commands. In Tableau Desktop,
you can choose to show or hide the toolbar by going to Window > Show
Toolbar.
The table (Table 1: Tableau Toolbar) shown below describes what each
button on the toolbar does. Note that certain buttons may not appear in every
version of Tableau. You can also refer to Visual Cues and Icons in Tableau
Desktop for additional details.
Toolbar Button Description
Tableau Icon Opens the Start Page (Desktop only).
Undo Reverses the last action. Multiple undos are allowed.
Redo Restores the last undone action.
Save Saves the workbook (File > Save in Server/Cloud).
New Data Source Connects to a new or existing data source.
Pause Auto Updates Stops automatic updates when changes are made.
Run Update Manually refreshes data (Desktop only).
New Worksheet Adds a new worksheet, dashboard, or story.
Duplicate Creates a copy of the current sheet.
Clear Removes filters, formatting, or entire view elements.
Swap Switches Rows and Columns fields.
Sort Ascending Sorts data in increasing order.
Sort Descending Sorts data in decreasing order.
Totals Adds or removes grand totals or subtotals.
Highlight Enables value highlighting in the sheet.
Group Members Combines selected data into a group.
Show Mark Labels Toggles visibility of data labels.
Fix Axes Locks or unlocks axis range (Desktop only).
Format Workbook Adjust workbook-wide formatting settings.
Fit Controls how the view fits in the window.
Show/Hide Cards Shows or hides worksheet cards.
Presentation Mode Hides all except the view (Desktop only).
Download Exports view/data in image, PDF, or CSV (Server/Cloud only).
Share Publishes workbook to Server or Cloud (Desktop only).
Show Me Suggests best visualizations for selected data.
Table 1: Tableau Toolbar
After understanding the user interface of Tableau, lets understand, how to Tools for Data
Science an
begin the work of data analysis, wherein the first step is to understand “How Introduction
to upload any dataset?”, next section relates to this only.
13.5.3 How to upload any dataset ?
The stepwise process of uploading any dataset, in Tableau is as follows:
Step – 1: Start Tableau on your system
system, you will see an interface as shown in
Fig 9 below
Fig 9: Tableau
eau Front Page Interface
Step – 2: "Navigate to the 'Connect' section, choose 'To a File', then click on
file format you want to upload for an instance 'Microsoft Excel'."
Excel'.", as shown
in Fig.10 below.
Fig 10: File Upload Interface in Tableau
Data Science – Allied Step – 3: Clicking this opens a file browser window, allowing you to navigate
Areas
through folders and choose the file you wish to open
open,, as shown in Fig. 11.
Fig 11: File Browser Interface in Tableau
Now, the sub
sub-steps
steps to be followed for the upload of dataset are as follows:
Step – 3(a): Select an appropriate file and open it,, as shown in Fig. 12
Fig 12: File upload in Tableau
Step – 3(b): Once the file is opened in Tableau, you'll see a screen displaying the Tools for Data
Science an
connected data, as shown in Fig 13 below Introduction
Fig 13:: Data Connection in Tableau
Finally, Tableau displays the contents of your Excel file along with all its available
sheets as shown in Fig 14,
Fig 14: Data Connection sheets in Tableau
By dragging and dropping a sheet, you'll be able to view its data and proceed with
further tasks,, like generating descriptive statistics and many more. The subsequent
section relates to the generation
neration of descriptive statistics.
13.5.4 Generating Descriptive Statistics - Statistical Analysis and
Tableau
Statistical analysis is essential in business intelligence, as it enables
organizations to explore data, identify trends, and gain meaningful insights.
One key area—descriptive
descriptive statistics
statistics—centers on summarizing data through
metrics such as measures of central tendency (like m mean and median) and
variability (such as standard deviation).
Data Science – Allied Descriptive statistics form the starting point for understanding any dataset. It
Areas
summarize
summarizes the main characteristics of numerical values by providing
measures such as averages, ranges, and level
levelss of variation. Tableau includes
several tools that allow users to produce these statistics quickly and interpret
them either numerically or visually. The following steps explain how
descriptive statistics can be generated within Tableau in a practical and
intuitive way
way, by using the Summary card.
A Summary Card in Tableau is a feature used to display quick statistical
information about the data in a worksheet or about a selected field. It
automatically shows key measures such as minimum, maximum, average,
total (sum), and count of values, helping users quickly understand the overall
distribution and range of the data. Summary cards can be applied to the entire
view or to specific dimensions and measures, making them useful for quick
data validation, compari
comparison,
son, and insight generation without performing
manual calculations.
Let’s learn how to uuse the Summary Card for generating descriptive statistics,
Following are the steps to generate descriptive statistics in Tableau:
1) Load your dataset into Tableau and open a new worksheet.
2) Choose the numerical field you want to analyze and drag it to either the
Rows or Columns shelf to create a basic view.
3) On the menu bar, select Worksheet → Show Summary.
4) After opening your worksheet, navigate to the Worksheet menu in the
top--left corner and select Show Summary,, as shown in Fig. 15 below.
Fig 15: Descriptive Statistics in Tableau
Note: It is to be noted that, once enabled, the Summary Card automatically appears Tools for Data
Science an
on the right side of the visualization,, as shown in Fig. 16 below. Introduction
Fig 16: Summary of Descriptive Statistics in Tableau
Note: The Summary Card can be repositioned anywhere on the worksheet
worksheet—whether
to the right, left, top, or bottom—simply
simply by dragging and dropping it as needed
needed, as
shown in fig. 17 below.
Fig. 17: Summary in Tableau
Data Science – Allied In short, Tableau provides an efficient way to generate descriptive statistics
Areas
using the Summary Card. It offers clear, quick insights into key measures
such as count, sum, average, minimum, and maximum values. By using the
Summary Card, users can easily understand data behavior and obtain
essential statistical information. This makes Tableau not only a visualization
tool but also a practical platform for basic descriptive analysis.
Till the time we learned, how to upload the data set? and how to generate
descriptive statistics? of the available dataset. Now, let’s extend our
discussion to the plotting of graphs and its analysis. The same is performed in
next section.
13.5.5 Graph plotting and its analysis
Once the data has been uploaded, it can be used for visualization, which is
the primary goal of using Tableau. This section focuses on demonstrating
how to build a series of visualizations in Tableau. To perform visualization
we need to Connect to the database in Tableau, in this case the referred data
source contains three separate sheets—Product, OrderDetails, and
PropertyInfo—all included within a single Excel file.
Below are the Steps to follow for visualization:
A) Opening Tableau and Connecting to Dataset
1. Launch Tableau Desktop.
2. On the Start Page, click on Text file / Excel depending on the file format.
3. Browse and select the Product dataset file.
Below is the structure of product table (Table-2):
Table 2: Product Table
ProductID ProductName ProductCategory Price
1 Large Towel Housekeeping 9
2 Hand Towel Housekeeping 5
3 Washcloth Housekeeping 3
4 Shampoo Housekeeping 40
5 Moisturizer Housekeeping 40
6 Conditioner Housekeeping 40
7 Hand Soap Housekeeping 35
8 Bath Soap Housekeeping 35
9 Tissues Housekeeping 14
10 Toilet Paper Housekeeping 19
11 Shower Cap Housekeeping 10
12 Bed Sheet (King) Housekeeping 29
13 Bed Sheet (Double) Housekeeping 20
4. Click Open → Tableau loads the data into the Data Source window. Tools for Data
Science an
Introduction
5. Verify all column names and data types.
Fig. 18: Dataset Connection with Tableau
Fig. 19: Open Dataset Connection with Tableau
6. Click on Sheet 1 to begin visualization.
B) Creating a Graph (Bar Chart: Category vs Price)
Step-by-Step Graph Plotting:
1. Drag Product_Category to the Columns shelf
shelf.
2. Drag Price to the Rows shelf.
3. From the Marks card, select Bar
Bar.
4. Tableau automatically plots a bar graph showing the total price for each
category.
Data Science – Allied 5. Drag ProductName to the Color option in the Marks card to differentiate
Areas
products.
6. Click on Show Mark Labels to display price values on the bars.
7. Add a title such as: “Category-wise Product Price Distribution”
13.5.6 Dashboard Creation in Tableau is an important and interesting task,
and to create a dashboard, at least two charts are required.
Here, we create the worksheets for the creation of Dashboard:
Worksheet 1: Category-wise Total Price (Bar Chart)
1. Drag Product_Category → Columns shelf.
2. Drag Price → Rows shelf.
3. From the Marks card, select Bar.
4. Click Show Mark Labels.
5. Rename the sheet as: “Category-wise Price”
Worksheet 2: Product-wise Price Comparison (Bar Chart)
1. Drag ProductName → Columns shelf.
2. Drag Price → Rows shelf.
3. Select Bar from Marks.
4. Click on Sort (Descending) by price.
5. Enable Show Mark Labels.
6. Rename the sheet as: “Product Price Comparison”
Creating the Dashboard
1. Click on the New Dashboard icon at the bottom of Tableau.
2. A blank dashboard page opens.
3. From the Dashboard Pane (left side):
o Drag “Category-wise Price” sheet onto the dashboard.
o Drag “Product Price Comparison” sheet below or beside the first chart.
4. Adjust the size and alignment of both charts.
5. Set Dashboard Size to Automatic for responsive display.
6. Add a Dashboard Title, for example: “Product Pricing Analysis Dashboard”
Further, as an alternative you can create from a Dashboard in Tableau. On using
following these steps.
Ø Open Tableau → Connect to Excel / CSV files
Ø Import the three tables into Data Source Product, OrderDetails, and
PropertyInfo—all included within a single Excel file.
Ø Create manual joins using ProductID
Tools for Data
Science an
Introduction
Fig. 20: Two Dataset Connection with Tableau
If you hover over the join, you'll see that an inner join has been created using a
common field, namely Product ID.. An inner join means that the two tables share a
common column, allowing their data to be combined seamlessly seamlessly. Next, add the
PropertyInfo sheet, and you'll notice that it also gets joined automatically.
Note:
Table au automatically detects relationships between tables based on matching field
names. Since fields like Product ID share the same name across the connected
sheets, Tableau recognizes them as common keys and creates the join without
requiring manual input. This automatic behavior occurs only when field names are
consistent across tables.
Fig. 21: ExcelSheet Dataset Connection with Tableau
Data Science – Allied Hovering over the join reveals that an inner join has been established between
Areas
OrderDetails and PropertyInfo using PropertyID as the common key
Fig. 22: ExcelSheet Dataset Connection with Tableau
Now, that data is completely ready for data visualization through dashboard
Ø Go to Sheet 1 → Create charts:
Histogram / Bar chart for Sales Distribution
Box Plot via Analytics Pane
· Create additional visuals if required (e.g., Region
Region-wise
wise sales)
· Select Dashboard → New Dashboard
· Drag visualizations into the layout and adjust formatting
· Add filters or legends for interactivity
Fig. 23: Excel Sheet Dataset Connection with Tableau
Choose Quantity,, then navigate to the top
top-right corner of the toolbar and click the Tools for Data
Science an
Show Me button to generate a visualization. Introduction
Fig. 24: Excel Sheet Quantity Dataset Connection with Tableau
Go ahead and click on Text Table to display the data in a tabular format.
Fig. 25: Text table Dataset Connection with Tableau
The current view displays only one measure (the total quantity), which adds up to
10,096. With no dimensions or categories to break it down, the visualization isn't
very insightful. Let's enhance it by adding additional fields to generate a more
detailed and interesting chart.
Data Science – Allied
Areas
Fig. 26: Text table Dataset Connection with Tableau
To select multiple non
non-adjacent fields, press and hold the Ctrl key (or
Command key on Mac), then click on each one individually. For example,
select both Quantity and Product Category.. With these fields selected,
Tableau will now suggest a broader range of visualization options.
Now that both Product Category and Quantity are in the view, head to the
Show Me panel and select the Tree Map option. This chart uses both color
and size to visually represent values, making trends easy to identify.
The Tree
ee Map clearly shows that Furnishings has the highest number of items
ordered. Public Areas and Housekeeping follow closely, with Maintenance
next, and Office Supplies having the lowest quantity.
Fig. 27: Tree Map Dataset Connection with Tableau
At first glance, Maintenance and Office Supplies appear to be nearly the
same size in the Tree Map. However, when you revisit the Show Me tab and
explore the visual more closely, it becomes clear that Maintenance has a Tools for Data
Science an
slightly higher quantity compared to Office Supplies. Introduction
This illustrates the fundamental steps for quickly creating visualizations in
Tableau using the options provided in the Show Me panel.
13.5.7 SUM UP: Tableau - Data handling and User Interface
1. Introduction to Tableau Interface
· Workspace: Consists of Data Pane, Shelves (Rows, Columns), and
Worksheets.
· Marks Card: Customize visuals (color, size, label).
· Dashboard Tab: Combine visuals.
· Connect Pane: Connect to CSV, Excel, SQL, etc.
2. Uploading the Dataset
· Launch Tableau Desktop.
· Under "Connect", select CSV/Excel/Database.
· Browse and upload your dataset.
· Drag tables into the workspace if needed.
3. Handling Data
· Use Data Interpreter for auto-cleaning.
· Create Calculated Fields for new variables.
· Use Filters, Groups, and Bins to organize data.
4. Plotting Graphs
· Drag fields to Rows and Columns.
· Use Show Me panel to choose chart types (bar, line, scatter, pie, maps).
· Build dashboards combining multiple visuals.
5. Summary : Generating Descriptive Statistics
Operation Tableau Feature Used
Median, Quartiles & Outliers Box Plot from Analytics Pane
Summary Statistics Show Summary (Worksheet tab)
Distribution Pattern Histogram + Distribution Bands
Quick statistical insights Summary Card (can be dragged anywhere on screen)
Table 3: Brief of generative statistics
These components help quickly identify:
· Data spread and center
Data Science – Allied · Skewness in sales
Areas
· Outlier products with extreme performance
6. Final Output (Expected)
A dashboard visually showing:
· Sales distribution chart
· Box plot with median & quartiles
· Optional supplier/region filter
· Summary card displaying statistics like:
✔ Count
✔ Mean
✔ Minimum
✔ Maximum
Tools for Data
POWER BI Science an
Introduction
Power BI is a Microsoft application built to convert raw data into
valuable insights.
It enables users to create interactive dashboards, reports, and charts to
simplify data analysis and understanding. Whether you're in business,
research, or any data-driven
driven field, Power BI helps you identify patterns, track
trends, and make informed decisions more efficiently.
The process involves three main steps:
1. Power Query Editor – Prepare and clean your data.
2. Data Modeling and Relationships – Establish connections between
data tables and organize your data.
3. Visualization –Build
Build interactive graphs and charts to explore and
showcase your data efficiently.
Fig.28: Getting Started with Power BI
Working with Power BI is straightforward
straightforward, it follows the easy steps, shown in
Fig. 29 below
Fig.29:: Working of Power BI
469
Data Science – Allied Step-by-Step procedure to work by Using Power BI is as follows :
Areas
Step 1: Download and Install Power BI Desktop: Go to Microsoft’s
official website and download Power BI Desktop for free. Install it on your
Windows system. This version enables you to create dashboards and reports.
After installation, you'll get a clean interface to start working with your data.
Step 2: Connect to Your Data: Open Power BI and click on “Get Data.”
You can connect to a variety of data sources such as Excel files, SQL
databases, or cloud services like Google Analytics and Salesforce. Select
your desired source and load the data. You can also combine data from
multiple sources into one report.
Step 3: Clean and Prepare Your Data: Data often needs cleaning before
analysis. Use Power Query to remove unnecessary rows, correct data
formats, and handle missing values. This ensures your dataset is accurate and
ready for analysis.
Step 4: Create Visualizations: Drag data fields into the report canvas to
start building visualizations. Power BI will automatically generate charts,
tables, or graphs based on the data. Choose visual types that best represent
your insights. Visuals help in identifying patterns and key metrics quickly.
Step 5: Build a Dashboard: Arrange multiple visuals into a single-page
dashboard. Dashboards allow users to explore different data perspectives by
interacting with the visuals. You can customize layouts to highlight important
information and enhance the user experience.
Step 6: Share Your Work: After completing your report or dashboard,
publish it to the Power BI Service (online platform). This allows you to
share your insights with others, enabling collaboration and easy access from
anywhere others can access and view it at any time. You can also set up
automatic data refreshes to keep your reports up to date, ensuring they always
display the most recent information.
Using Power BI - Power BI is mainly used to create interactive reports and
dashboards that provide deeper insights into your data. After publishing your
report to the Power BI service, both you and other users can explore and
interact with it while tracking its performance. This guide highlights the
essential steps for building and modifying reports and dashboards.
Essential Steps:
1. Extract the Data – Begin by importing your data from various sources.
2. Transform the Data – Use Power Query to clean, modify, and prepare
your data, as well as to manage and adjust relationships between tables
as needed.
3. Apply Calculations with DAX – Use Data Analysis Expressions (DAX)
to perform calculations on your dataset.
4. Start Visualizing – Once the data is prepared, proceed to build your
visualizations.
5. Create and Customize Visuals – Add elements like charts, graphs, and Tools for Data
Science an
cards, and modify them for clarity and effectiveness. Introduction
6. Publish to the Cloud – Upload your dashboard to the Power BI service
so your team can access and interact with it
it.
Loading Data - To start handling data in Power BI, select the ‘Get Data’
option under the Home tab. This allows you to connect to a wide range of
data sources, including:
· Files – Import data from formats like Excel or CSV.
· Databases – Establish connections to databases like SQL Server, Oracle,
or MySQL.
· Direct Query – Access data directly from the source without importing
it.
· Online Services – Connect with platforms like Google Analytics,
Salesforce, or SharePoint.
· Live Connection – Establish a real
real-time link to data sources for always
up-to-date information.
Fig.30:: Using Power BI
After importing the data into the application, you can begi
begin transforming and
modeling it. You can create new columns with custom configurations and
define your own calculated fields.
Before building meaningful dashboards and analytical reports, it is essential
to prepare and structure the imported data. Raw data often contains missing
fields, formatting inconsistencies, or lacks the calculated values required for
analysis. Power BI provides powerful tools for transforming, enriching, and
organizing data so that it aligns with analytical and business requirements.
Data Science – Allied Power BI enables users to enhance the existing dataset by creating calculated
Areas
columns and measures. These user-defined
defined fields help implement business
logic, perform advanced calculations, and support accurate data insights that
may not be directly availabl
availablee from the source. This process forms the core of
data modeling
modeling,, which ensures that the data is structured correctly for
visualizations and interactions. The steps to create calculated columns and
measures are listed below:
Step 1: Begin by clicking on the "Data" icon from the three options
available in the left
left-hand panel.
Fig.31: Data Tab Power BI
Step 2: In the "Table Tools" tab, go to the "Calculation" group and choose
"New Measure" to create a custom calculation.
Fig.32: Table Tools Power BI
Step 3: Next, use the desired functions along with the table field names to
define your logic and generate a new column.
Tools for Data
Science an
Introduction
Fig.33:: Column Tools Power BI
Fig.34:: Measure Tools Power BI
Step 4: Once this step is finished, your required columns or measures will be
set up. Use the "New Column" option when you need the column to remain
part of the table and update automatically during data refreshes, as it becomes
a permanent part of the data model. On the other hand, measures are
calculated dynamically and function more like real real-time filters—making
them more efficient since they aren't stored in memory.
Power BI also offers a built-in Power Query Editor that supports advanced
data cleaning and transformation, ensuring your data is well
well-prepared for
analysis and visualization .
Power Query - Power wer Query is a robust tool that allows you to shape and
transform your data through operations such as filtering, grouping, and
pivoting. It also supports the creation of calculated columns and measures
using the formula bar and a wide range of functions. After the data is cleaned
and prepared, you can move on to creating visualizations
visualizations. Following are the
steps to be performed to prepare the data for its visualization.
Data Science – Allied Step 1: Click on the "Home" tab from the top navigation menu.
Areas
Fig.35: Home Tab Power BI
Step 2: In the Queries section, click on "Transform Data."
(This action will launch the Power Query Editor in a new window. If no data
source is connected yet, a blank screen will be displayed.)
Fig.36: Transform Data Tab Power BI
Step 3: On this screen, you can view, edit, transform, and clean your data as
required.
The Home tab provides multiple tools for managing your data sources.
Fig.37:: Power Query Editor in Power BI
The "Transform" tab contains functions for adding, deleting, splitting, and
changing data types.
474 Fig.38: Power Query Editor Transform Tab
"The 'Add Column' tab allows you to add new columns and format them Tools for Data
Science an
accordingly." Introduction
Fig.39:: Power Query Editor Adding Column
The 'View' tab manages the way the page is displayed.
Fig.40:: Power Query Editor View Tab
Step 4: By right-clicking
clicking on your column, you can apply various modifications.
Fig.41:: Power Query Editor View Tab
Step 5: After finishing the necessary data transformations, navigate to the
'File' tab and choose 'Close & Apply' to save your changes and return to
Power BI Desktop.
Data Science – Allied
Areas
Fig.42:: Power Query Editor View Tab
Your data is now ready to be used in creating visualization
visualizations and generating
reports.
13.6.2 Visualization in Power BI
In Power BI, the Format and Fields panes play a vital role in customizing and
enhancing the visual appearance of reports and dashboards. The Fields pane
is used to manage and organize the dataset by selecting, adding, or removing
fields for different visualizations, while the Format pane allows users to
control the visual properties of charts and reports such as colors, titles, labels,
backgrounds, data bars, and tooltips. To adjust the overall la layout and
appearance of a report, users can utilize the View tab, where options such as
Themes, Page View, Gridlines, and Snap to Grid are available. The Themes
option provides multiple predefined styling formats that help maintain visual
consistency and im improve
prove report aesthetics according to organizational or
presentation requirements. Together, the Format, Fields, and View options in
Power BI enable users to design professional, interactive, and visually
appealing dashboards and reports.
Fig.43: Power BI - Fields pane, Format Pane & View Tab
The Format and Fields pane is where you customize and improve the Tools for Data
Science an
appearance of your dashboards and reports
reports. To adjust the layout of your Introduction
report, use the "Layout" option under the View tab, which provides various
themee options to match your report’s design.
Fig.44:: Visualization Tab
To modify a visualization already added to the report, select it and use the
"Format" option to customize its appearance.
Fig.45: Visualization Tab
For more advanced customization of your visualizations, you can use DAX
formulas to create calculated columns and measures as shown below in this
section.
Data Science – Allied Data Analysis Expressions (DAX) is a powerful formula language used in
Areas
Power BI (as well as in Excel Power Pivot and SQL Server Analysis
Services) to perform custom calculations on data after it has been imported
into the model. DAX helps users go beyond basic summaries by enabling the
creation of calculated columns, measures, and calculated tables, which
add new analytical meaning to existing data.
A calculated column works row by row and is mainly used to add new
derived data to a table, such as profit calculated from sales and cost. A
measure, on the other hand, is used to perform dynamic calculations such as
totals, averages, percentages, or ratios that automatically change based on
filters and slicers applied in the report. This dynamic behavior is what makes
DAX especially valuable for interactive dashboards.
For example, a simple DAX measure like:
Total Sales = SUM(Sales[Amount])
calculates the total revenue by adding all sales values from the Sales table.
This measure updates instantly when a user applies filters such as product
category, date, or region. In this way, DAX allows reports to become
interactive and responsive, which is essential for business analysis.
DAX also supports more advanced analytical functions, such as time-
intelligence calculations. These functions make it possible to compare
performance across time periods, such as month-to-month growth, year-over-
year growth, or running totals. For example, a Year-over-Year (YoY) Sales
Growth measure compares current year sales with previous year sales and
expresses the change as a percentage. Such calculations are extremely useful
for trend analysis and performance evaluation.
Basic Steps of Working with DAX in Power BI
Step 1: Load Data into Power BI
1. Open Power BI Desktop.
2. Click on Home → Get Data.
3. Select Excel / CSV / Text file.
4. Browse and load your dataset.
5. Click Load to import the data into Power BI.
After loading the dataset:
· Use Data View to see the table.
· Use Report View to create visuals and apply DAX measures.
6. Create a Basic DAX Measure
1. Go to Report View.
2. In the Fields Pane, right-click on the table name (e.g., Sales).
3. Click on New Measure.
4. In the formula bar, type the DAX formula:
Example, Total Sales = SUM (Sales [Price])
5. Press Enter Tools for Data
Science an
Introduction
This measure calculates the total sales amount dynamically.
In short, DAX acts as the analytical brain of Power BI. It allows analysts
to create meaningful metrics, perform complex calculations, compare
performance over time, and build interactive dashboards that support
informed decision-making. Even at an introductory level, understanding basic
DAX concepts like measures, filters, and aggregation functions provides a
strong foundation for advanced business intelligence and data analysis tasks.
This measure dynamically evaluates sales growth relative to the previous
year based on the current filter context. Thus, the use of DAX formulas
significantly enhances analytical flexibility and supports the development of
meaningful business insights within Power BI.
13.6.3 Generation of Descriptive Statistics in Power BI
Descriptive statistics are used to summarize and describe the basic features of
a dataset in a simple and meaningful way. In Power BI, descriptive statistics
are generated using a combination of built-in visual tools and DAX (Data
Analysis Expressions) measures, allowing users to quickly understand
patterns, trends, and overall performance of data.
Power BI provides built-in analytical capabilities to summarize data and
derive descriptive statistics such as totals, averages, minimum and maximum
values, counts, percentages, and distributions. Once a dataset is imported into
Power BI Desktop, visual elements like Cards, Tables, Matrix, and Clustered
Bar/Column Charts can be used to present summary statistics. Additionally,
DAX measures enhance the analytical scope by enabling dynamic
calculations based on user interactions. Common descriptive statistical
measures in Power BI can be generated with simple DAX formulas
To create more flexible and dynamic statistics, Power BI allows users to
define custom DAX measures. Simple DAX formulas such as SUM,
AVERAGE, DISTINCTCOUNT, MIN, and MAX are commonly used to
compute descriptive measures. For instance, Total Quantity =
SUM(OrderDetails[Quantity]) calculates the total number of items sold,
Average Price = AVERAGE(Product[Price]) gives the typical price of
products, and Product Count = DISTINCTCOUNT(Product[ProductID])
returns the number of unique products. Similarly, MIN(Product[Price]) and
MAX(Product[Price]) identify the lowest and highest product prices in the
dataset. These metrics update automatically based on report filters or slicers,
helping users gain real-time insights into product performance, customer
preferences, and revenue trends.
One of the major advantages of using Power BI for descriptive statistics is
that these values are dynamic. This means the results automatically change
whenever a user applies filters or slicers, such as selecting a particular
product category, time period, or region. As a result, users can instantly
observe how key metrics vary across different conditions.
Data Science – Allied We learned that
that, Power BI simplifies thee generation of descriptive statistics
Areas
by combining easy
easy-to-use
use visuals with powerful DAX calculations. This
enables beginners to summarize large datasets, uncover basic patterns, and
gain quick insights into product performance, customer behavior, and oveoverall
business trends in an interactive and user
user-friendly
friendly manner.
A) Creation of Dashboard Steps in Power BI
The dashboard development process in Power BI involves connecting to data,
transforming it as needed, and building interactive visuals. The essentia
essential steps
are:
Step 1: Importing Data
To begin working in Power BI, the first task is to load the required dataset
into the application. Power BI enables data import from multiple sources,
including Excel, CSV, databases, and online services. Click on the “Get
Data” option located on the left side of the Home screen to browse and select
your data file. Once the desired dataset is chosen, it can be easily loaded into
the workspace for further processing and analysis.
· Open Power BI Desktop → Click Get Data
· Select the data source (Excel/CSV/Database/etc.)
· Load the dataset into the Data Model
Fig.46: PowerBI Dashboard
Fig.47: Sample Dataset
The Navigation pane provides an option labeled Files,
Files which allows users to
access data stored locally. Select Files and browse to the folder where the Tools for Data
Science an
Excel workbook or other supported data format is saved. After selecting the Introduction
appropriate file, click Connect to establish the data import into Power BI.
Fig.48:: PowerBI Data Upload
Depending on the dataset size,, Power BI may require a short processing time
to extract the data. Once the preview is displayed, verify that the data has
been imported correctly and then click Load to bring the dataset into the
Power BI workspace.
Fig.49:: HR Dataset Upload
Step 2: Exploring the Data
After the data is loaded, navigate to the Data View to examine the dataset in
a structured tabular format. The Fields Pane on the right side of the interface
displays all available tables along with their respective columns, allowing
you to review and understand the attributes present in the imported data.
Data Science – Allied
Areas
Fig.50: Data View
You can click on any table or individual field to apply formatting options as
required. For attributes representing values such as dates, time, geographical
locations (city, state), percentages, or currency, the appropriate data type
and formatting can be assigned using the tools available under the
Modeling
ing tab.
Step 3: Selecting an Appropriate Visualization
For developing the dashboard, we focus on five key fields: HiredYear,
RecruitmentSource
RecruitmentSource, Position, EmployerID, and Gender (Male/Female).
(Male/Female)
The first visual to be added is a Card visualization. To create it, simply select
the Card icon from the Visualizations pane and place the relevant field onto
the canvas.
Fig.51: Table Tools
Next, choose the fields that you want to include in the visualization from the Tools for Data
Science an
Fields Pane.. You may either click to add them or simply drag and drop the Introduction
selected fields into the corresponding placeholders within the visualization,
as indicated in the interface.
Fig.52: Field Pane
You can customize the visualization further by selecting fields
fields, applying
filters, and adjusting appearance settings using the Format options. The first
Card visualization created displays the total number of Employers,
providing a quick summary metric for the dashboard.
Fig.53:: Visualization Pane
The same procedure was repeated to create the remaining Card visuals. These
additional cards represent key summary indicators, including the total
number of Recruitment Sources
Sources, total number of Positions, and the
Maximum Salary value present in the dataset. Together, these visuals
provide an at-a-glance
glance statistical overview that supports further analysis
within the dashboard.
Next, a Pie Chart and a Donut Chart are created to represent the
distribution of job positions and the gender-wise employment ratio,
respectively. Thesee charts can be inserted by selecting the corresponding
Data Science – Allied icons from the Visualizations pane and assigning the required fields to the
Areas
chart areas.
Fig.54: Types of Charts
Finally, a Funnel Chart and a Stacked Bar Chart are added to visualize the
year-wise
wise hiring counts and the proportional contribution of different
recruitment sources
sources,, respectively. To enhance clarity and presentation
quality, formatting options such as titles, data labels, axis properties,
legends, plot area settings, and color schemes are applied. These
adjustments ensure the dashboard communicates insights effectively, as
illustrated in the figure below.
Fig.55:: Visual demonstration of Recruitment Dashboard
Step 4: Publish and Share
· Publish report to Power BI Service
· Create dashboards by pinning visuals
· Share insights through secure cloud access
13.6.4 SUM UP: Power BI - User Interface and Data Handling Tools for Data
Science an
Introduction
1. Navigating Power BI
· Home Pane: Get Data, Transform Data.
· Fields Pane: Lists all columns and tables.
· Visualizations Pane: Select and customize charts.
· Report View: Main canvas for building reports.
2. Uploading the Dataset
· Launch Power BI Desktop.
· Select "Get Data" and choose your source (Excel, CSV, or Database).
· Use the Power Query Editor to load or clean and shape your data as
needed.
3. Handling Data
· Use Power Query to clean your data by removing rows and adjusting
data types.
· Generate new columns and perform calculations with DAX (Data
Analysis Expressions).
· Establish Relationships between tables.
4. Plotting Graphs
· Use drag-and-drop to build visuals on canvas.
· Available charts: bar, line, scatter, map, waterfall, KPI, donut.
· Interactivity via slicers and filters.
5. Example of a DAX Formula
Power BI allows users to create custom calculations using DAX (Data
Analysis Expressions).
Below is an example formula that calculates the Total Sales based on Price ×
Quantity:
, ∗
This calculated measure dynamically updates based on filters used in the
report and helps in analyzing revenue trends across different dimensions such
as products, categories, or properties.
…………………………………………………………………………….
…………………………………………………………………………….
…………………………………………………………………………….
…………………………………………………………………………….
13.7 PYTHON
This section is an Introduction to Python Interface (Anaconda & Google
Colab) - Python IDLE, Jupyter Notebook, and Google Colab all support
Python development but differ in interface and capabilities, a brief
introduction to each is given below:
· Python IDLE: A simple, built-in IDE with a Python Shell for interactive
code and an Editor for scripts. It’s lightweight and good for beginners
and small projects.
· Jupyter Notebook: A web-based tool with a cell-based layout
supporting live code, text (Markdown), and visualizations. Ideal for data
science, research, and documentation.
· Google Colab: A cloud-hosted version of Jupyter with added benefits
like free GPU/TPU access, Google Drive integration, and real-time
collaboration. Best suited for machine learning tasks.
Before we start diving into Python coding, let's learn how to install Python
IDLE.
Installing Python and Accessing IDLE : To work with Python locally
(instead of Google Colab), we use Python IDLE, which comes bundled with
the official Python installation. The steps to be followed for download and
installation of Python IDLE are as follows:
Steps to Download & Install Python Tools for Data
Science an
Introduction
1. Go to the official Python website: [Link]
2. Click on Downloads and choose the version appropriate for your
operating system (Windows/Mac/Linux).
3. Run the installer.
4. Important: Check the box “Add Python to PATH” before clicking
Install.
5. Complete the installation. IDLE will be installed automatically.
Opening Python IDLE : Once Python is installed
installed, you can start IDLE by
following the steps below:
· On Windows: Search for IDLE (Python GUI) in the Start Menu.
· On Mac/Linux:: Run idle3 from the terminal or locate it in Applications.
Noww we can start writing Python code locally using the IDLE editor and the
steps to be followed are as follows:
1. Open Python IDLE
· Windows: Search for “IDLE” in the Start Menu and open it.
· macOS: Use Spotlight to search for IDLE or open from Applications >
Python.
· Linux: Run idle or idle3 in the terminal.
2. Use the Python Shell
· A window opens called the Python Shell
Shell, as shown below :
Top of Form
You can enter Python code right here and view the output immediately.
· Example:
3. Open a New File (Script Mode)
· Go to File > New File or press Ctrl+N.
· A new editor window opens where you can write multiple lines of code.
4. Write Your Program
· Type your Python program in the new window.
Example:
Data Science – Allied
Areas
5. Save Your File
· Navigate to File > Save or press Ctrl+S.
· Save the file using a .py extension, such as [Link].
[Link]
6. Run Your Program
· In the editor window, go to Run > Run Module or press F5.
· The output appears in the Python Shell window.
7. Use Basic Features
· IDLE also includes:
ü Syntax highlighting in the editor
ü Auto
Auto-indentation
ü Basic debugger (under Debug > Debugger)
Fig .63: Python IDLE Interface
Fig .64: Python IDLE shell Interface
13.7.1 Steps to Use Jupyter Notebook Tools for Data
Science an
Introduction
1. Install Jupyter Notebook (if it isn't already installed)
· You can install it using pip:
Or install via Anaconda (which includes Jupyter by default).
2. Start Jupyter Notebook
· Open a terminal or command prompt and enter the following command:
· A web browser will open showing the Jupyter dashboard (usually at
[Link]
Fig .65:: Python JupyterLab Interface
The JupyterLab interface features a toolbar on the left that includes a file
browser and a Git tab. The file browser displays the directory from which
JupyterLab was launched. On the right side, you can view and work with
files you’ve opened. To enable a split view, simply drag a file tab from the
top into another area of the workspace.
7. Create a New Notebook
· Click on “New” and select “Python 3” to launch a fresh notebook.
· A new tab opens with a blank notebook .
8. Write and Run Code
· Write your code in a cell.
Data Science – Allied Example:
Areas
· Execute the cell by clicking the Run button or pressing Shift + Enter
[5].
9. Use Markdown for Notes
· Change a cell to Markdown (from the dropdown menu).
· Write formatted text using Markdown syntax.
10. Save Your Notebook
· Go to File > Save and Checkpoint,, or use the shortcut Ctrl+S.
· It saves as a. ipynb file.
11. Close the Notebook
· Close the browser tab and shut down the server by pressing Ctrl+C in
the terminal window.
13.7.2 Steps to Use Google Colab
1. Open Google Colab
· Go to: [Link]
2. Sign In with Google Account
· You need a Google account to use Colab.
3. Create or Open a Notebook
· Click “File > New Notebook” to start a new one.
· Or open an existing one from:
o Google Drive
o GitHub
o Upload from your device
Fig.66: Python Google Colab Notebook Interface
4. Write and Run Code Tools for Data
Science an
· Enter Python code into a cell and press Shift + Enter to execute it. Introduction
5. Add Text Cells
· Click + Text to add a Markdown cell for formatted notes or headings.
6. Save the Notebook
· It autosaves to Google Drive.
· To download, click File > Download > .ipynb or .py.
7. Use Extra Features
· You can:
o Upload files using the sidebar or [Link]()
o Mount Google Drive using:
o Use GPUs (under Runtime > Change runtime type > GPU
GPU)
Till now we learned how to work with Google Colab, Jupyter Notebook,
IDLE etc., now lets understand how to work with python using datasets. To
understand the working we used MovieLens Small Dataset
Dataset, you can browse it
on internet and download it from the link :
[Link] ; We'll use the file: [Link]
13.7.3 Steps to Upload Dataset in Python IDLE
IDLE is a basic IDE and doesn’t have a built
built-in way to upload files — but
you can manually place the file in your script folder
folder.
Steps to be followed are:
1. Download [Link] from MovieLens and save it in the same folder as
your .py file.
2. Open IDLE > File > New File.
3. Write and run the code:
4. Save and run the script (F5).
Data Science – Allied Steps to Upload Dataset in Jupyter Notebook are:
Areas
1. Open Jupyter Notebook.
2. Place [Link] in the same folder where your notebook is saved.
3. Use this code to load the file:
4. If you want to browse and upload via Jupyter:
o Click Upload in the Jupyter dashboard > Choose [Link] > Click
Upload button again.
13.7.4 Steps to Upload Dataset in Google Colab
Method 1: Upload from Local System
· Open a new Colab notebook.
· Use this code:
· Select and upload [Link].
· Then use:
Method 2: Load from Google Drive
1. Mount Google Drive:
2. Place [Link] in your Drive, then access it like:
Till this point we learned “how to interact with the python tools like Google Colab, Tools for Data
Science an
Jupyter Notebook etc?”, and also we learned “how to upload the dataset?”. Now, we Introduction
will learn to generate the descriptive statistics in the subsequent section.
13.7.5 Descriptive Statistics in Python
Descriptive statistics summarize and describe the main characteristics of a
dataset in a meaningful way. They help us understand the central tendency,
spread, and overall distribution of the data. Common descriptive measures
include the mean, which represents the average value, and the median, which
identifies the middle value when the data is sorted. The mode indicates the
most frequently occurring value. Measures such as variance and standard
deviation describe how much the data varies from the mean, while minimum,
maximum, and range provide information about the spread of the dataset.
These statistics allow readers to quickly understand the nature of the data
before performing advanced analysis.
So, we learned that Descriptive Statistics, helps us to understand the dataset
before modeling, and it helps to detect patterns, outliers, or missing data,
further it makes us to summarize large data into key insights quickly.
Let’s Learn the process of Generating Descriptive Statistics in Python :
To produce descriptive statistics in Python, the panda’s library is used
extensively. The following steps show how to compute them:
Step 1: Import pandas
import pandas as pd
Step 2: Create or Load a Dataset
data = {
'Marks': [65, 70, 85, 90, 75, 88, 92]
}
df = [Link](data)
Step 3: Generate Descriptive Statistics
stats = [Link]()
print(stats)
Output will include:
· Count
· Mean
· Standard deviation
· Minimum value
· Maximum value
· 25th, 50th (Median), and 75th percentiles
Now, lets apply the understanding and try to explore, How to Perform
Descriptive Statistics in Python?, Let’s use the MovieLens [Link] (has
user ratings of movies) for demonstration.
Data Science – Allied Step 1: Upload Dataset
Areas
(Use the same file upload method as discussed earlier in IDLE, Jupyter, or
Colab.)
Step 2: Apply Descriptive Statistics
Ø Output Includes:
· count – total number of entries
· mean – average rating
· std – standard deviation (spread)
· min
min, max – lowest and highest values
· 25%
25%, 50%, 75% – percentiles (quartiles)
Ø Additional Analysis
Ø Explanation with Example:
Assume these are some ratings:
· Mean = average rating = (4 + 5 + 3 + 4.5 + 3.5) / 5 = 4.0
· Median = 4.0 (middle value) Tools for Data
Science an
· Mode = most frequent (if any) Introduction
· Standard deviation = how much ratings deviate from mean
Ø Sample table example:
After generating descriptive statistics, plotting of graphs is an important step,
and we will learn this process of lotting graphs in python in subsequent
section.
13.7.6 Plotting Graphs in Python (Using Matplotlib)
To visualize data in Python, the most widely used library is Matplotlib
Matplotlib,
however many other libraries are available.
Follow the steps below to plot a simple graph:
Step 1: Install Matplotlib (if not already installed)
pip install matplotlib
Step 2: Import Required Libraries
import [Link] as plt
Step 3: Prepare Data
x = [1, 2, 3, 4, 5]
y = [10, 20, 15, 30, 25]
Step 4: Plot the Graph
[Link](x, y)
Step 5: Add Labels and Title
[Link]("X-Axis")
[Link]("Y-Axis")
[Link]("Simple Line Graph")
Step 6: Display the Plot
[Link]()
Let’s understand the complete process by executing a Hands-on
Example using mtcars dataset: we learned that Descriptive statistics help
Data Science – Allied highlight key characteristics of a dataset, often using just one value. Creating
Areas
these statistics is typically the first step after data has been cleaned and
prepared for analysis. We've previously encountered examples like the mean
and median. In this section,, we'll revisit those and introduce some additional
statistical tools.
Firstly we begin with the finding of Central tendency of data i.e. Measures of
Center,, which are statistical methods used to identify the "middle" or most
typical value in a numerical dataset. They provide insight into what a
standard or expected value might be. The most common measures of central
tendency include the mean, median, and mode. Wee can compute the mean for
each column in a DataFrame using:
We can also get the means of each row by supplying an axis argument: Tools for Data
Science an
Introduction
The median of a dataset is the value that splits the data into two equal parts
parts—
half the values are lower, and half are higher. It represents the center point
when the data is sorted. For this reason, it’s also known as the 50th
percentile. As shown previously, you can compute the median for each
column in a DataFrame using:
We can also calculate the median across each row by passing the argument
axis=[Link]
.Although both the mean and median indicate the central tendency of a
dataset, they measure it differently and are not always the same. The median reliably
splits the data into two equal halves, while the mean represents the arithmetic
average, which can be significantly affected by outliers or extreme values. In a
Data Science – Allied symmetrical distribution, the mean and median are generally equal. This concept can
Areas
be better understood by examining a density plot.
Fig.67: Output to check Normalization in dataset
In the plot shown above, the mean and median are nearly identical—both
identical
are close to zero
zero—so
so the red line representing the median overlaps the thicker
black line that marks the mean.
In skewed distributions, the mean is more affected by the skew and is drawn
in the direction of the skewness, while th thee median is less influenced and
tends to stay near the central point of the data.
Tools for Data
Science an
Introduction
Fig: 68: Mean & Median Output
The mean is greatly influenced by outliers because it includes all values, even
the extreme ones, in its calculation. On the other hand, the median is more
resistant to outliers and stays fairly consistent, making it a more reliable
measure of central tendency when outliers are present.
Data Science – Allied
Areas
Fig.69: Outliers in dataset
Since the median is less influenced by skewed data and outliers, it is regarded
as a "robust" measure. In cases where distributions are noticeably skewed or
contain extreme values, the median often gives a more reliable representation
of a typical value.
The mode is the value that appears most often in a dataset. Unlike the m
mean
and median, the mode can also be used with categorical data. Additionally, a
dataset may have more than one mode if several values share the highest
frequency. To identify the mode, you can use:
Tools for Data
Science an
Introduction
If a column has multiple modes—meaningmeaning several values ooccur with the same
highest frequency—then
then all those values are returned as modes. On the other hand,
if a column has no repeating values,, and thus no mode, the result will be NaN.
Finding the Measures of Spread is also important in data analysis, lets understand
how to determine Measures of spread, also known as dispersion, indicate how much
the values in a dataset differ from each other. While measures of center highlight a
typical or average value, measures of spread show how widely the data points aare
scattered around that center.
A simple method for assessing spread is the range, calculated by subtracting
the minimum value from the maximum value in the dataset.
As noted earlier, the median represents the 50th percentile of a dataset. To
gain a clearer
arer picture of a variable’s spread, it's helpful to examine several
percentiles. These include the minimum (0th percentile), first quartile (25th
percentile), median (50th percentile), third quartile (75th percentile), and
maximum (100th percentile). You can extract these values using the quantile ()
function in pandas.
Data Science – Allied
Areas
Because these percentile values are frequently used to summarize a dataset,
they are collectively known as the five-number
number summary" consists of the
minimum, 25th percentile (Q1), median (Q2), 75th percentile (Q3), and
maximum. These key statistics offer a quick overview of the data’s
distribution and are also generated by the [Link]() function in pandas.
The Interquartile Range (IQR) is a commonly used measure of variability.
It shows the spread of the middle 50% of the data and is determined by
subtracting the first quartile (Q1) from the third quartile (Q3):
IQR = Q3 - Q1.
The boxplots shown in fig 57 representing the summary and the interquartile Tools for Data
Science an
range (IQR) offers a clear view of the data’s distribution, central tendency, Introduction
spread, and possible outliers.
Fig.70:: Boxplot of IQR
Two other commonly used measures of spread are variance and standard
deviation.
· Variance is the average of the squared differences between each data
point and the mean. It provides insight into how much the values in a
dataset deviate from the average.
You can calculate the variance of each column in a DataFrame using:
Data Science – Allied Ø Standard deviation is the square root of the variance. It’s generally
Areas
more intuitive to interpret since it's in the same units as the original data,
unlike variance, which is expressed in squared units. To calculate the
standard deviation for each column in a DataFrame, use use:
Since variance and standard deviation rely on the mean, they can be heavily
influenced by skewed data and outliers. A more robust alternative is the
Median Absolute Deviation (MAD). MAD is computed as the median of the
absolute differences from the median, making it a dependable measure of
spread for datasets with non
non-normal
normal distributions or extreme values.
Note: The Median Absolute Deviation (MAD) is frequently multiplied by a
scaling factor of 1.4826
1.4826.. This adjustment makes the MAD comparable to the t
standard deviation when the data follows a normal distribution, allowing for better
consistency between the two measures.
Skewness and Kurtosis : In addition to measures of center and spread,
descriptive statistics also include metrics that describe the shape of a
distribution i.e. skewness and kurtosis.. A brief introduction to both is given
below:
· Skewness indicates the asymmetry of the distribution—whether
distribution the data
is skewed to the left or right.
· Kurtosis reflects how much of the data lies in the tails versus the center
of the distribution.
·
While we won’t dive into the exact formulas, these metrics build on the
concept of variance:
· Skewness uses cubed deviations from the mean.
· Kurtosis uses fourth power deviations from the mean.
Tools for Data
Science an
Introduction
To better understand skewness and kurtosis,
kurtosis let’s generate some sample (dummy)
data and examine its distribution using these two measures:
Fig.71: Skewness and Kurtosis
Data Science – Allied Next, let’s evaluate the skewness of each dataset. Since skewness reflects the
Areas
asymmetry of a distribution, we expect it to be low for most of the distributions that
are relatively symmetric, and higher only for the one that is intentionally skewed.
Now let's check kurtosis.
As observed from the results, the normal distribution shows kurtosis close to
zero,, indicating a typical bell
bell-shaped curve. The uniform (flat) distribution has
negative kurtosis
kurtosis, meaning it has lighter tails and a flatter peak. In contrast, the
other two distributions
distributions—with more data in the tails than in the center—exhibit
center
higher kurtosis
kurtosis,, reflecting heavier tails and more extreme values.
So we learned that Descriptive statistics provide numerical summaries that
help you understand key aspects of your data, such as its center, spread, and
shape.. These statistics not only guide the direction of your analysis but also
make it easier to communicate findings clearly and efficiently. Moreover,
metrics like mean and variance are foundational in many statistical tests and
predictive models.
SUM UP - User Interface and Data Handling in Python (Anaconda/Google
Colab)
1. Introduction to Anaconda & Colab
· Anaconda Navigator
Navigator:: Launch Jupyter Notebook, Spyder.
· Jupyter Notebook Interface:: Cell
Cell-based execution, Markdown support. Tools for Data
Science an
· Google Colab: Cloud-based,
based, supports GPU, file uploads. Introduction
2. Uploading the Dataset
· In Jupyter:
· In Google Colab, import and use the file upload function:
· Then, read the file using Pandas:
3. Handling Data
· Use Pandas: [Link](), [Link](), [Link]()
· Clean data: drop missing, rename columns, filter rows.
· Transformation: df['new'] = df['col1'] + df['col2']
5. Creating Visualizations Use libraries like Matplotlib and Seaborn to
plot different types of graphs:
R-STUDIO
RStudio is a powerful, user-friendly IDE for R programming. It’s the
preferred environment for most R users. It offers Clean layout for data
science tasks, and quite easy
asy access to documentation and plots with
ntegrated version control (Git). However, Other Interfaces (Optional) are
Jupyter Notebook with IRKernel – Interactive R notebooks, VS Code with
R extension, StatET Plugin for Eclipse,, and Jamovi also.
e will work on R-Studio only,, the user interface of the same is shown
Fig.72: R
R-Studio: Interface Layout (4 Panes):
The description of the 4-panes
panes shown in the above figure, is as follows: Tools for Data
Science an
Introduction
1. Source Editor (Top-Left)
· Write and edit R scripts (.R), RMarkdown (.Rmd), and notebooks.
· Supports syntax highlighting, code folding, and auto
auto-completion.
Source Editor This is where you write and edit scripts, RMarkdown
files, and notebooks. You can run code from here directly into the
Console (Pane 3).Useful for saving and reusing code.
2. Environment / History Pane
· Environment Tab:: Shows all the objects, variables, and datasets
you've created.
· History Tab: Keeps a record of previously executed commands
commands.
· Helps manage workspace variables and track command history.
3. Console (Bottom-Left)
· Run R commands interactively, and it Displays output, errors, and
messages.
4. Files/Plots/Packages/Help/Viewer (Bottom
(Bottom-Right)
· Files:: Navigate your project directory.
· Plots:: Display visualizations created in R.
· Packages:: Manage installed libraries.
· Help:: Access documentation for functions.
Now, after understanding the utility of all 44-panes of the user interface of R-
Studio, let’s extend our discussion to learn How to uuse RStudio's GUI for
Importing the Dataset. The steps to be performed are as follows:
Steps:
1. Open RStudio.
2. In the Environment pane (top
(top-right), click “Import Dataset”.
3. Choose from:
o From Text (readr) or From Excel
Excel, or From CSV
4. Browse and select your file (e.g., [Link]).
5. A preview appears—check the options like separator, header, etc.
6. Click “Import”.
7. The dataset is loaded into your environment (usually as a data frame).
Steps for Alternate Method : Using Code (R Console or Script)
1. To Import a CSV File:
2. Example (Windows path):
Data Science – Allied 3. To Import Excel File:
Areas
Common Import Functions
Function Use Package
[Link]() Load CSV files Base R
[Link]() Load general delimited text files Base R
read_excel() Load Excel .xlsx files readxl
read_csv() Faster CSV loading readr
fread() Fast and efficient CSV loader [Link]
Table 4:: Functions of R
Note : Use getwd() to check the current working directory,
directory and setwd("path")
to change it:
After understanding the basic working structure of R R-Studio, let’s extend our
discussion a next level, i.e. how to perform Descriptive analysis in R? the same is
addressed in next section.
13.8.2 Descriptive Analysis in R Programming
Descriptive statistics in R involve summarizing and organizing data using
various tools such as charts, graphs, tables, and spreadsheets.
spreadsheets This process
helps in presenting the data in a clear and understandable format,
format allowing
for meaningful insights. Descriptive
riptive analysis aims to summarize the main
features of a dataset. It is typically performed on smaller datasets,
providing a foundational understanding that can also guide future trend
predictions based on existing data patterns.
Descriptive statistics of
often
ten rely on two major types of measures i.e. Measures
of Central Tendency and Measures of Variability (Dispersion)
(Dispersion), the two are
discussed below
below:
1 . Measures of Central Tendency: Indicates the center or typical value of the
dataset and include:
· Mean (Average): The arithmetic average of all observations. (R-Syntax:
mean(data vector) ) Tools for Data
Science an
· Median: The middle value when the data is arranged in ascending order. Introduction
(R-syntax: median(data vector))
· Mode: The most frequently occurring value. R does not have a built-in
mode function, so it can be computed as: (R-Syntax: mode value <-
names(sort(table(data vector), decreasing = TRUE))
2. Measures of Variability (Dispersion): Reflects spread out of the data and it
includes:
· Range : Difference between the highest and lowest values.
§ R-Syntax : range(data vector)
§ diff(range(data vector))
· Interquartile Range (IQR) : Measures the spread of the middle 50% of
the data.
§ R-Syntax : IQR(data vector)
· Variance: Average of the squared deviations from the mean.
§ R-Syntax : var(data vector)
· Standard Deviation : Square root of the variance; shows spread in the
same units as the data.
§ R-Syntax : sd(data vector)
Example:
# Sample dataset
data vector <- c(12, 15, 14, 19, 20, 15, 18, 22, 15)
# Central Tendency
mean(data vector)
median(data vector)
mode value <- names(sort(table(data vector), decreasing = TRUE))
# Variability
range(data vector)
IQR(data vector)
var(data vector)
sd(data vector)
This code produces output that clearly shows the descriptive statistics,
Given: data vector <- c(12, 15, 14, 19, 20, 15, 18, 22, 15)
Measure Output
Mean 16.67
Median 15
Mode 15
Range 12 to 22
Range (Difference) 10
IQR 4
Data Science – Allied
Areas Variance 11.25
Standard Deviation 3.354102
Now, lets revisit the Steps for Performing Descriptive Statistics in R
1. Load your dataset using functions like [Link](), read_excel(), or data() for
built
built-in sets.
2. Summarize the data
data:
o Use summary(data) for a quick overview
o Use mean(), median(), sd(), var(), range() for individual measures
3. Visualize the data
data:
o Create graphs such as histograms (hist()), boxplots (boxplot()), and bar
charts (barplot())
4. Interpret results to understand trends, central points, and variability in
your data.
To upload a dataset in R,, you can use several methods depending on your
file type (e.g., CSV, Excel) and your preference (command line, GUI, etc.).
The commonly
common used method are as follows:
1. Uploading a CSV File of dataset can be done by using the [Link]()
[Link](
function:
· Replace "path/to/your/[Link]" with your actual file path.
· If using RStudio, you can drag and drop the file into the Files pane, then
use the path shown.
2. Uploading an Excel File : Use the readxl or openxlsx package:
3. Uploading dataset by Using RStudio GUI (Import Dataset Button)
1. Click on Environment tab > Import Dataset.
2. Choose:
o From Text (readr) – for CSV/TSV
o From Excel – for Excel files Tools for Data
Science an
3. Browse and select your file. Introduction
4. RStudio auto-generates
generates the code and previews the data.
4. Uploading from the Web
Till the moment we learned how to handle a dataset, now its time to extend
our discussion for data visualization and data interpretation. To understand
these features we are going to refer to the renowned dataset of mtcars, a built
in dataset in R. you can explore the dataset by applying commands, shown in
the respective screenshots, given below:
Data Science – Allied
Areas
For the purpose of Data Visualization and interpretation we are going to use
the data visualization library ggplot() over the built-in
in dataset “mtcars”.
Here's how you can systematically perform data visualization and data
interpretation, using R
Task-1: What is the Relation between MPG (Mileage) and Number of
Cylinders ?
We know that MPG, or miles per gallon,, indicates how far a car can travel
using one gallon of fuel. It’s a standard measure of fuel efficiency—higher
MPG means better fuel economy and less fuel consumption per mile. So, by
plotting the graph between MPG (Mileage) and Number of Cylinders,
Cylinders using
command shown below:
We got following output as graph or visualization, showing the re
relationship
between MPG (Mileage) and Number of Cylinders
Tools for Data
Science an
Introduction
Fig.73:: Data Interpretation
Interpretation: The scatter plot shows that as the number of cylinders
increases, the mileage (mpg) decreases.
Task-2: Explore the relationship between mileage (mpg) and engine
displacement: Again, a graph is plotted between the two i.e. between mileage
(mpg) and engine displacement,, and the same is shown below:
Fig.74
74: Scatter plot
Interpretation - Greater engine displacement generally results in lower fuel
efficiency.
NOTE – (Understanding
Understanding Displacement
Displacement): Engine displacement refers to the total
volume that all cylinders in an engine can move. A larger displacement allows the
Data Science – Allied engine to intake more air and fuel, which increases its potential to generate power.
Areas
However, the actual power output also depends on other internal components and
the engine's design. Displacement is usually measured in liters. For example, many
modern vehicles have a 2.02.0-liter four-cylinderr engine, meaning each cylinder has a
capacity of about 0.5 liters or 500 cc
cc.
Task-3:
3: How does cylinder effects mpg and displacement? Again, a graph is
plotted between the variables and the same is shown below:
Fig.75: Colored scatter plot
Interpretation - Cars with fewer cylinders tend to have higher mileage and
better fuel efficiency.
On the other hand, vehicles with more cylinders usually have larger engine
displacement, meaning the engine consumes more air and fuel. As a result, they are
generally less fuel
fuel-efficient.
Task-4:
4: Understanding the Relation between mpg, disp and hp,
hp to interpret
this a graph is plotted between the variables using the command, shown
below:
Tools for Data
Science an
Introduction
Fig.76:: Interpretation of horse power
Interpretation - As horsepower increases, engine displacement also tends to
increase. This indicates a more powerful engine, which typically results in lower
fuel efficiency — meaning the vehicle delivers less mileage.
Task-5: Understanding the Relation between hp, disp, mpg, and numbe number of
cylinders, to interpret this a graph is plotted between the variables using the
command, shown below:
Fig.77:: Interpretation of horse power
Interpretation - As horsepower (hp) increases, there is a corresponding increase in
engine displacement and usually the number of cylinders as well. This combination
points to a more powerful engine.. However, such engines tend to be less fuel-
efficient, which means cars with higher hp, larger displacement, and more cylinders
generally offer lower mileage (mpg).
Addition of shape and color makes it easier to interpret the plot and perform better
analysis. This helps in understanding how various features like displacement,
Data Science – Allied horsepower
horsepower, and mileage interact based on engine configuration.
Areas
The Scatter plot can be generated by using the following syntax in R
Fig.78:: Interpretation of horse power
Generating Bar Charts: Bar charts are useful for summarizing and comparing
categorical variables or aggregated values (like average mileage per cylinder
type).
Fig.79: Bar graph
Tools for Data
Science an
Introduction
Fig.80:: Colored Bar graph
Fig.81:: Colored Bar graph (Motor Cars Cyl & Count)
Data Science – Allied
Areas
Fig.82 : Factor vs. Count Bar graph
Ø From the Analysis of the Bar charts produced above, following can
be interpreted:
1. Most 88-cylinder engines are V-shaped.
2. 4-cylinder
cylinder engines are mainly straight (inline).
3. 6-cylinder
cylinder engines have a nearly equal mix of V and straight
configurations.
Fig.83: Factor vs. Count Bar graph
Ø From the Analysis of the Bar charts produced above, following can
520 be interpreted:
1) No of Cylinders are 4; carb used are 1 and 2 Tools for Data
Science an
2) No of Cylinders are 6; carb used are 1,4, and 6 Introduction
3) No of Cylinders are 8; carb used are 2,3,4, and 8
Ø SUM UP - User Interface and Data Handling in R (RStudio)
1. Introduction to RStudio Interface
· Source Pane: Script writing.
· Console Pane: Code output.
· Environment/History Pane:: Variables and past commands.
· Files/Plots/Packages Pane:: File management, visual output.
2. Uploading the Dataset : df <- [Link]("[Link]")
· Use "Import Dataset" GUI in RStudio.
3. Handling Data
· Use functions like summary(df), str(df), head(df)
· Clean using [Link](), subset(), mutate() (from dplyr).
4. Use Base R plotting functions:
For advanced and customizable visuals, use the ggplot2 package:
· Use base R or ggplot2 for advanced visuals.
☞ Check Your Progress - 4
1. How do you import a CSV file in R?
…………………………………………………………………………….
…………………………………………………………………………….
…………………………………………………………………………….
2. What is the function of summary() in R?
…………………………………………………………………………….
…………………………………………………………………………….
…………………………………………………………………………….
3. Which R function is used to generate a histogram?
…………………………………………………………………………….
…………………………………………………………………………….
…………………………………………………………………………….
521
Data Science – Allied 4. How can you check the structure of a dataset in R?
Areas
…………………………………………………………………………….
…………………………………………………………………………….
…………………………………………………………………………….
5. How does cylinder count affect MPG in mtcars dataset?
…………………………………………………………………………….
…………………………………………………………………………….
…………………………………………………………………………….
This structured approach introduces students to the basic interfaces and
functionalities of Tableau, Power BI, Python (Anaconda/Colab), and R (RStudio)
while ensuring they gain hands-on exposure to uploading datasets, preparing data,
and generating useful visual insights.
13.9 SUMMARY
This unit-13 titled “Tools for Data Science – An Introduction” emphasizes that in
today’s digital age, data is a critical resource, but its value depends on how
effectively it is processed and interpreted. Data science, combining mathematics,
statistics, computing, and domain expertise, transforms raw data into actionable
knowledge. To do this, professionals rely on powerful tools that enable collection,
cleaning, analysis, and visualization of data. This unit introduces four essential
tools—Tableau, Power BI, Python, and R—which together form the foundation of
modern data science practice.
Tableau is presented as an intuitive visualization platform requiring little or no
coding. It allows learners to connect datasets from various sources, create interactive
dashboards, and generate descriptive statistics. Through drag-and-drop features and
built-in analytics, Tableau helps communicate insights to both technical and non-
technical audiences. Power BI, developed by Microsoft, is introduced as a robust
business intelligence tool, closely integrated with Excel and cloud services. Learners
are guided through steps of connecting to data, cleaning it using Power Query,
building relationships, applying calculations with DAX, and creating dashboards
that can be shared via the Power BI service.
The unit then shifts to Python, a versatile programming language widely used in
data science for its rich libraries such as Pandas, NumPy, Matplotlib, and Seaborn.
Learners are introduced to different interfaces like IDLE, Jupyter Notebook, and
Google Colab, with hands-on steps for uploading datasets, performing descriptive
statistics (mean, median, mode, standard deviation, skewness, kurtosis), and
visualizing data through plots. Practical emphasis is laid on using Python for data
preparation and exploratory analysis.
Finally, R (via RStudio) is introduced as a statistical powerhouse favored in
research environments. The unit explains the RStudio interface and dataset import
methods, followed by performing descriptive analysis using summary statistics,
measures of central tendency, dispersion, and visualization techniques such as
522 histograms, scatter plots, and boxplots. Learners are encouraged to explore
relationships between variables using examples like the mtcars dataset. Tools for Data
Science an
Introduction
Overall, the unit equips learners with a comparative understanding of four
foundational tools in data science. Tableau and Power BI provide user-friendly, no-
code solutions for visualization and business intelligence, while Python and R offer
greater flexibility and analytical depth through coding. By the end of the unit,
learners gain the ability to import, clean, analyze, and visualize datasets, thereby
building confidence to apply these tools in academic, research, or professional
contexts.
13.10 SOLUTIONS/ANSWERS
☞ Check Your Progress - 1
1. What are the key components of the Tableau workspace?
Answer: The key components include Data Pane, Shelves (Rows,
Columns), Cards, Toolbar, Worksheet Tabs, and the View Area.
2. How do you connect to a Microsoft Excel file in Tableau?
Answer: Go to the 'Connect' section, select 'To a File', and then choose
'Microsoft Excel' to upload the dataset.
3. What is the purpose of the 'Show Me' panel in Tableau?
Answer: It suggests the most suitable visualizations based on selected
data fields.
4. How can you calculate the median in Tableau?
Answer: You can use the median aggregation function or drag the
Median option from the Analytics pane.
5. How are outliers identified using Tableau?
Answer: Outliers are identified using distribution bands in visualizations
like box plots or line charts.
☞ Check Your Progress-2
1. What is Power Query used for?
Answer: Power Query is used to clean, transform, and shape data before
analysis.
2. How can you create a new measure in Power BI?
Answer: Use the 'New Measure' option under the 'Table Tools' tab and define
logic using DAX.
3. What are calculated columns in Power BI?
Answer: Calculated columns are fields created with DAX that become part of
the data model and update with refreshes.
4. How do you upload a dataset in Power BI?
Answer: Click on 'Get Data' under Home tab, select the source (e.g., Excel),
and load the file.
5. How do you publish a report in Power BI?
523
Data Science – Allied Answer: Use the 'Publish' button to upload your report to the Power BI Service
Areas
for sharing.
☞ Check Your Progress - 3
1. What function in pandas shows basic statistics of a dataset?
Answer: `[Link]()` provides count, mean, std deviation, min, max, and
quartiles.
2. How can you read a CSV file in pandas?
Answer: Use `pd.read_csv('[Link]')` to read the file into a DataFrame.
3. What library in Python is used for plotting graphs?
Answer: Matplotlib and Seaborn are widely used for data visualization.
4. How do you calculate mean and median in pandas?
Answer: Use `[Link]()` for mean and `[Link]()` for median.
5. How can you identify outliers using visualization?
Answer: Use boxplots from Seaborn or matplotlib to highlight outliers.
☞ Check Your Progress - 4
1. How do you import a CSV file in R?
Answer: Use `[Link]('[Link]')` or use the Import Dataset GUI in
RStudio.
2. What is the function of summary() in R?
Answer: `summary(data)` gives min, max, mean, median, and quartile
information.
3. Which R function is used to generate a histogram?
Answer: `hist()` is used to create histograms in R.
4. How can you check the structure of a dataset in R?
Answer: Use `str(data)` to see the data types and structure.
5. How does cylinder count affect MPG in mtcars dataset?
Answer: More cylinders typically result in lower MPG due to higher fuel
consumption.
13.11 FURTHER READINGS
· Official Tableau Learning Resources [Link]
· Tableau Your Data! – Daniel G. Murray A comprehensive book covering Tableau
techniques and real-world applications.
· Visual Analytics with Tableau – Alexander Loth Focuses on building dashboards and
interpreting data effectively.
· Microsoft Power BI Documentation [Link]
· The Definitive Guide to DAX – Marco Russo & Alberto Ferrari, Excellent for learning
data modeling and DAX functions in Power BI.
· Power BI for the Excel Analyst – Wyn Hopkins, Helpful for Excel users transitioning
to Power BI.
524
· Python Official Documentation, [Link] Tools for Data
Science an
· Automate the Boring Stuff with Python – Al Sweigart (free to read online), Great Introduction
beginner book with hands-on examples.
· Python for Data Analysis – Wes McKinney, Focuses on NumPy, pandas, and data
wrangling in Python.
· Google’s Python Class, [Link]
· R for Data Science – Hadley Wickham & Garrett Grolemund, [Link]
A great introductory book for learning R in the context of data analysis.
· The Art of R Programming – Norman Matloff, Ideal for students who want to
understand R at a deeper programming level.
· Hands-On Programming with R – Garrett Grolemund, Interactive and beginner-
friendly for learning R syntax and logic.
· CRAN R Manuals and Tutorials, [Link]
· Tableau Software. (2024). Tableau Training and Tutorials. Retrieved from
[Link]
· Loth, A. (2019). Visual Analytics with Tableau. Wiley.
· Murray, D. G. (2016). Tableau Your Data! Fast and Easy Visual Analysis with
Tableau Software. Wiley.
· Microsoft Corporation. (2024). Power BI Documentation. Retrieved from
[Link]
· Russo, M., & Ferrari, A. (2020). The Definitive Guide to DAX: Business
Intelligence for Microsoft Power BI, SQL Server Analysis Services, and Excel.
Microsoft Press.
· Hopkins, W. (2021). Power BI for the Excel Analyst: A Step-by-Step Guide to
Effective Data Analysis. Apress.
· Python Software Foundation. (2024). Python 3 Documentation. Retrieved from
[Link]
· Sweigart, A. (2019). Automate the Boring Stuff with Python. No Starch Press.
· McKinney, W. (2022). Python for Data Analysis: Data Wrangling with pandas,
NumPy, and Jupyter. O'Reilly Media.
· Google Developers. (2024). Google's Python Class. Retrieved from
[Link]
· Wickham, H., & Grolemund, G. (2016). R for Data Science. O'Reilly Media.
Retrieved from [Link]
· Matloff, N. (2011). The Art of R Programming. No Starch Press.
· Grolemund, G. (2014). Hands-On Programming with R. O'Reilly Media.
· R Core Team. (2024). R: The R Project for Statistical Computing. Retrieved
from [Link]
525
What Is Data Storytelling?
Data storytelling is the ability to effectively communicate insights from a dataset using
narratives and visualizations. It can be used to put data insights into context for and
inspire action from your audience.
There are three key components to data storytelling:
Data: Thorough analysis of accurate, complete data serves as the foundation of your
data story. Analyzing data using descriptive, diagnostic, predictive,
and prescriptive analysis can enable you to understand its full picture.
Narrative: A verbal or written narrative, also called a storyline, is used to
communicate insights gleaned from data, the context surrounding it, and actions you
recommend and aim to inspire in your audience.
Visualizations: Visual representations of your data and narrative can be useful for
communicating its story clearly and memorably. These can be charts, graphs,
diagrams, pictures, or videos.
Data storytelling can be used internally (for instance, to communicate the need for
product improvements based on user data) or externally (for instance, to create a
compelling case for buying your product to potential customers).
Data Storytelling: Best Practices for Dashboard Layout and Interactivity
In the age of data, business dashboards should be designed using data storytelling
techniques . Data storytelling converts metrics into clear narratives that guide the
reader to actionable conclusions.
According to this technique, a dashboard is not just about displaying data in a visually
appealing way, but about structuring the information in such a way that decision
makers and business professionals can quickly understand what is happening, why it
is happening and what actions to take about it.
A well narrated dashboard turns data into insights. By integrating storytelling
techniques into a dashboard, we make the audience remember the information better
and connect emotionally with it.
What is Data Storytelling in Business Dashboards?
Data storytelling is a technique that transforms dashboards into visual narratives that
explain what is happening, why it is happening and what decisions to make. It
combines data visualization, context and narrative structure to facilitate data-
driven decision making.
In a more theoretical sense, data storytelling is the art of constructing a
compelling narrative from complex information, relying on data visualization to
communicate a message to a given audience.
In the context of dashboards, it involve designing interactive dashboards that tell a
story: each visualization should support a key point and all together should take the
reader through a logical path, from a contextual introduction to an actionable
conclusion.
A well-constructed dashboard not only represents information, it tells a story and
conveys a meaningful message. At the end of the journey, the reader should
understand what is going on and be inspired to make an informed decision based on
data.
Benefits of Data Storytelling
In most organizations there is a gap between the abundance of data and the ability
to make decisions with it.
Often, we have so many reports and metrics that we fall into the "data overload
paradox" where, ironically, information overload leads to inaction.
Storytelling with data acts as a bridge to bridge that gap, structuring data into a logical
and persuasive narrative that makes it easier to interpret patterns, gain insights and,
above all, implement concrete actions.
In addition, telling stories with data helps to distils and simplify complex information.
Good data storytelling simplifies the complicated so that the audience can assimilate
it and make decisions more quickly and confidently. It also adds a "human touch" to
the numbers: contextualizing the numbers with stories or examples gives them
relevance and creates connection.
Data Storytelling Techniques and Best Practices to Create Effective Dashboards
:
An effective dashboard narrative does not emerge by chance. It requires deliberately
applying principles of visual design, data communication and data-driven analytics to
make data-driven decisions.
Here are the key data-driven storytelling techniques that will enable you to transform
your reports into interactive dashboards that drive data-driven business decisions.
1. Define the purpose of the dashboard
Every story starts with two key elements: what you want to tell and to whom.
Before designing a dashboard, Prepare:
• What is the main message or key insight the user should take away?
• Without a defined objective, we run the risk of building a generic report that
does not solve any specific need.
• Having a clear narrative objective will allow you to focus the dashboard on the
really relevant data, avoiding information overload.
2. Know your audience
Each professional profile interprets data differently. It is not the same to design a
dashboard for a CEO than for a marketing team, as they will have different needs.
When designing a dashboard, we must adapt the complexity, language and
visualization format to the profile of the audience.
Best practices for this technique:
• Use simple language familiar to non-technical users.
• Include only key metrics for the business.
• Highlight what really matters and discard the irrelevant.
• Make sure that anyone, regardless of their technical level, can understand the
dashboard.
3. Structure the dashboard story: introduction, development and conclusion.
One of the most powerful data storytelling techniques is to apply the classic narrative
structure to data visualization.
An effective dashboard is not a group of isolated graphs, but a story with a beginning,
middle and end. This organization makes it easy for decision makers to follow the
thread and extract actionable insights effortlessly.
Introduction: provide context
Start the dashboard with an overview that situates the reader.
Example: an "Annual sales vs. target" card on a financial dashboard allows you to
quickly understand the overall status (above or below target?).
Development: explain "why"
This is where you break down the information. Add visualizations that show trends,
comparisons and segmentations that explain the causes behind the initial data.
Example: graphs by region, product line or time periods that show what factors are
driving - or holding back - results.
Conclusion: highlight the action
Close with a key insight or recommendation. It can be a summary graphic, a final
indicator or a text that makes the main message clear.
Example: a "Year-over-year cumulative growth" indicator along with a note:
"Exceeded target thanks to Q3 performance".
This narrative structure turns the dashboard into a coherent visual story. Each
visualization should have a clear purpose within the narrative. Avoid logical jumps
and arrange the information in a way that answers the user's questions:
What's happening, why, and now what do we do?
4. Contextualize data
Isolated data can be misleading. Be sure to accompany each visualization with the
necessary information to interpret it: units, dates, goals, historical evolution or
relevant segmentation.
Example: don't just show "Total sales Q2", but "Sales Q2 2025 (€) vs. annual target".
Best practices for contextualizing your data:
Always include units and time ranges in titles, axes or legends.
Clear example:
✔ "Sales Q1 2025 (in million €)"
✘ "Sales"
Compare each performance indicator with a target or previous period.
This way the reader understands whether the current value represents a breakthrough,
a drop or an alert.
Example:
"+10% vs. previous year" provides much more than a raw number.
Support your graphs with brief explanations or comments.
Simple text can make a difference:
"Demand fell in March due to supply issues."
Take advantage of functionalities such as tooltips or pop-up boxes to provide
additional details without cluttering the visualization.
This is especially useful for corporate or executive dashboards, such as the balanced
scorecard, where clear, quick-to-consume information is valued.
5. Apply the "less is more" principle in data visualization.
One of the pillars of effective data storytelling is simplicity. In data visualization, less
is more: showing only relevant information helps the key message stand out without
noise or distractions.
A clean, focused dashboard is much more effective than one overloaded with
visualizations. Each additional graphic or metric competes for the reader's attention.
If you present too many elements, you risk creating confusion, slowing down
interpretation and diluting key insights.
Best practices to apply this principle:
Avoid dashboards with dozens of small or ir-relevant charts .
Select only the essential data that support your narrative or help you make decisions.
Eliminate the ancillary: if a visualization doesn't change any decision, it probably
shouldn't be there.
Relegate technical or exploratory details to secondary pages, tabs or supplemental
reports.
6. Highlight key information with visual hierarchy and clear design.
One of the golden rules of storytelling with data is that not all data has the same
weight. A fundamental part of effective dashboard design is to guide the user's
attention to the most important insights.
Techniques to highlight key information:
Main KPIs in large size and prominent position (top left corner).
Corporate or contrasting color for key figures.
Smaller typography and neutral tones for secondary data.
Example: 14.089.323,98 € → show as 14,1 M€ and leave the detail in a tooltip.
For example:
"Summer promotion here → sales peak."
These notes function as the "voice of the narrator" within the dashboard and reinforce
the visual narrative.
8. Summarize insights with notes, labels and annotations.
It is important not to assume that the reader will correctly interpret each graphic.
Add brief text within or next to the visuals to explain the message when necessary.
Example: an annotation such as "Here begins the summer campaign → sales
increase".
These micro-narratives act as the narrator's voice within the dashboard and reinforce
understanding without the need for external explanations.
9. Maintain narrative consistency in colors, fonts and formats.
Visual design in data storytelling is not decoration: it is part of the message. A poorly
executed design can distract, confuse or even lead to erroneous interpretations of the
data. On the other hand, a coherent and clear design reinforces the visual narrative
and the understanding of key insights.
Color palette: communicate with intent.
Use consistent color for the same concepts throughout the dashboard.
Example: blue for "sales", red for "costs".
Establish corporate color or primary and secondary tones in advance.
Avoid the "`": too many different colors make reading difficult.
Prioritize color with sufficient contrast to ensure accessibility and legibility.
Typographies and formats: consistency above all.
Use the same font for titles, labels and body text.
Use hierarchical sizes to highlight elements (KPI > graphic > annotation).
Avoid random mixtures of fonts or styles: they give an unprofessional image.
In corporate dashboards, respect the visual identity of the brand.
10. Choose the right visual for each data
Prioritize clarity over originality: a well-done horizontal bar is worth more than a
confusing 3D graphic.
Avoid pie charts if they make reading difficult.
Use the right type of visual according to the purpose:
Bars or columns: comparison of values.
Lines: evolution over time
Maps: geographical distribution
Cards/KPIs: summarized key indicators
maintain consistency in scales, colors and order between similar visuals. This
prevents the reader from having to "recalibrate" their attention on each graphic.
STORYTELLING WITH DATA
VISUALIZATION
Data visualization is more about storytelling and less about graphs and charts. Having the ability to
transform dense data into short, compelling, and actionable insights is a value. The principles of good data
storytelling are the topic of this chapter, followed by case studies on applying visualization techniques in
different domains and good data storytelling as shown in the given fig.9.1.
Fig. 9.1: Storytelling using Data Visualization
Designing Effective Data Stories:
What is Data Storytelling?
Data storytelling is the process of combining data, visuals, and storytelling to present insights in an effective
way. A good data story answers critical questions, leads the audience to conclusions, and employs well-
crafted narratives and visuals.
Fig. 9.2: What is Data Storytelling?
Key Elements of a Data Story:
Data storytelling is a strong method that integrates data, graphics, and narrative to present insights
effectively as depicted in the above fig 9.2. It serves to convert raw data into actionable stories that can be
used to inform decision-making. The following are the essential components of data storytelling in data
visualization, detailed as follows:
i. Data: At the core of data storytelling is the data itself. Without high-quality, reliable data, even
the most beautiful story has no credibility.
Best Practices for Data Usage in Storytelling:
a. Ensure Data Accuracy: Cross-validate sources to prevent misinformation.
b. Use Relevant Data: Utilize only the data points that add to the story.
c. Provide Context: Raw figures are misleading; always provide their context.
d. Reveal Trends and Patterns: Data storytelling must emphasize relationships, trends,
and key findings instead of mere numbers.
A narrative based on incorrect or incomplete data may result in misinterpretation and poor
decision-making.
ii. Visuals: Data visualization is very important in telling stories because the human mind absorbs
visuals 60,000 times faster than written words. Nicely designed graphics aid in presenting
complex data more easily.
Selecting the Best Visualization:
a. Bar Charts: Best utilized when comparing distinct categories.
b. Line Charts: Best for presenting trends over some time.
c. Pie Charts: These are utilized to represent proportions but must be utilized judiciously.
d. Heatmaps: Good for detecting patterns in big datasets.
e. Scatter Plots: Assist in demonstrating relationships and correlation between two
variables.
Guidelines for Effective Visualization:
a. Keep it Simple: Refrain from clutter and extraneous design components.
b. Use Color Judiciously: Apply contrasting colors to emphasize crucial aspects, and
make it accessible (e.g., colorblind-friendly color palettes).
c. Be Consistent: Employ similar scales, labels, and formats for a consistent
appearance.
d. Include Labels and Annotations: Highlight axes, data points, and trends with clear
labels for enhanced understanding.
Good visual storytelling decreases cognitive load and makes the message readily
understandable
iii. Narrative: An effective narrative gives meaning and context to data. It takes the reader
through the findings and makes sure they get the message from the numbers.
Elements of a Great Narrative:
It gives the background on the data and why it is significant.
Conflict or Challenge: States the problem or question the data is trying to solve.
Insights and Findings: Includes the main conclusions, backed up by data visuals.
Ends with something actionable or suggestive.
A strong narrative aid in making data more relatable and ensuring the audience can relate to
the message.
Principles for Effective Data Storytelling: Design Principles for effective data visualization are
given in the figure
Fig. Design Principles
i. Start with a Clear Purpose:
i. Establish the tone of the narrative: Narrative, descriptive, or expository?
ii. Identify the target audience (executives, analysts, general public).
ii. Know Your Audience:
i. What is the audience's level of knowledge?
ii. What conclusions will they make from this information?
iii. Organize the Story:
i. Begin: Indicate the problem or question.
ii. Middle: Display the data with analysis.
iii. End: Key takeaways and call to action.
iv. Have Good Visualization Skills:
i. Trends & comparisons → Line graphs, bar charts.
ii. Proportions → Pie charts, stacked bars.
iii. Relationships → Scatter plots, bubble charts.
v. It should be concise and easy:
i. Avoid unnecessary complexity in charts.
ii. Highlight the main points to avoid information overload.
vi. Highlight Key Takeaways:
i. Employ annotating, coloring, and labelling to direct attention.
ii. Remove distracting items that are not adding anything.
9.2 Applying Visualization Methods to Different Domains:
Data visualization is a key tool in the majority of fields, assisting professionals in understanding complex
data sets, recognizing patterns, and making sound decisions. Certain visualization methods are employed
by industries based on the nature of data they have and the purpose. The following are key areas where
visualization plays a significant role, with the right methods and examples.
1. Business & Finance:
Companies and financial institutions are dependent on data visualization to track performance, monitor
financial trends, and make improved decision-making. Executives, investors, and analysts require
transparent, real-time insights to evaluate business health and market dynamics.
Key Visualization methods:
i. Line Charts: It employed to illustrate trends in expenses, revenues, or stock prices over some
time.
ii. Bar Charts: It compares financial performance between various periods or business segments.
iii. Pie Charts: This illustrates the percentage of revenue contributed by different products or
services.
iv. Heatmaps: They determine the most profitable and underperforming areas or business segments.
v. Dashboard Reports: It combines several visualizations for an overall business overview.
Fig. 9.4: Visualization methods for business & finance
Example:
A global company employs an interactive business intelligence dashboard to monitor revenue, costs, and
profitability by region. The dashboard emphasizes financial KPIs using bar charts and line graphs as shown
in the fig 9.4, enabling executives to make fast, smart decisions.
2. Healthcare & Epidemiology:
The medical industry applies data visualization to monitor outbreaks of diseases, compare patient
demographics, and enhance treatment plans using various visualization methods like shown in the fig 9.5.
Strong visualizations assist policymakers, scientists, and healthcare workers in spotting trends and in the
effective distribution of resources.
Major Visualization Techniques:
Fig. 9.5: Visualization methods for healthcare
i. Line Charts: It follows the spread of disease (e.g., new daily COVID-19 cases).
ii. Scatter Plots: It examines the correlation between risk factors (e.g., age and severity of disease).
iii. Bubble Charts: It contrasts the rates of disease prevalence in various regions or populations.
iv. Geospatial Maps (Choropleth Maps): It displays disease outbreaks geographically.
v. Histograms: These present patient age range or hospital admission patterns.
Example:
During the COVID-19 pandemic, Johns Hopkins University's interactive dashboard employed geospatial
maps and time-series graphs to map the virus's global spread. The tool enabled policymakers and the public
to monitor real-time infection rates and health responses.
3. Marketing & Consumer Analytics:
Marketing experts use data visualization to monitor consumer behaviour, measure campaign success, and
optimize ad strategies. Through visual analytics, companies can enhance customer interaction and achieve
maximum return on investment (ROI).
Visualization techniques used:
i. Funnel Charts: It monitors the conversion of visitors into customers.
ii. Sankey Diagrams: They illustrate customer journeys and determine drop-off points in sales
funnels.
iii. Heatmaps: They display user interaction on websites with high-engagement areas.
iv. Word Clouds: It draws important themes from social media conversations and customer reviews.
v. Pie Charts: It divides customer demographics by age, income, or interests.
Example:
A digital marketing firm examines website heatmaps to see where visitors click most, thereby enhancing
webpage design and user experience.
4. Social & Political Sciences:
Social scientists and political analysts use data visualization to study public sentiment, voting behaviour,
and demographic changes. These insights are crucial for policy-making, electoral strategies, and media
analysis.
Key Visualization Methods:
Fig. Visualization for Social & Political Science
i. Choropleth Maps: It displays election results by state, district, or county.
ii. Bar Charts: It compares survey results on social issues across different populations.
iii. Word Clouds: It summarizes key themes from public debates, speeches, or news articles.
iv. Network Graphs: They show relationships between political entities, influencers, or social
media users.
v. Stacked Bar Charts: They Visualize demographic distributions (e.g., education levels of voters).
Example:
During a presidential election, news agencies like The New York Times use interactive maps to display
real-time voting results, allowing users to explore voting patterns across different regions using above
mentioned graphs shown in fig 9.6.
5. Engineering & Scientific Research:
Scientists and engineers apply data visualization to examine experiment outcomes, track system behaviour,
and improve designs. Large datasets need to be shown in a straightforward, relevant manner so that research
and development can be enabled.
Major Visualization Techniques:
i. Scatter Plots: It illustrates variable relationships (e.g., material strength as a function of
temperature).
ii. Box Plots: It illustrates statistical distributions of experimental data.
iii. Time-Series Graphs: It examines system behaviour over time (e.g., sensor values within an IoT
network).
iv. 3D Surface Plots: These visualize intricate physical or chemical processes.
v. Histograms: They visualize frequency distributions of experimental data.
Fig. 9.7: Visualization for Engineering
Example:
A climate scientist plots global temperature change over the last century with line graphs and
heatmaps and boxplots as shown in fig 9.7 to show the effects of climate change.
9.3 Design an Interactive data visualization storyboard for
real-time data:
With the current big data age, interactive data visualization is indispensable in deriving sensible
insights from giant datasets. Well-crafted storytelling allows users to spot trends, engage with
information, and make quick decisions based on facts. Below is an interactive data visualization
storyboard with the aim of analyzing Netflix data using director, release year, budget, language,
IMDb rating, genre, and other attributes.
Fig. 9.8: Netflix Story Board in Tableau
Storyboard Objectives:
The objective of the interactive visualization is to allow users to have a fun and informative
means of viewing Netflix content trends, determining the best genres, evaluating budget patterns,
analyzing ratings
from IMDb, and knowing how different directors and languages affect the content on the streaming
service. The dashboard will be user-friendly, dynamic, and easy on the eyes.
Data visualization plays a vital role in contemporary digital interfaces, with the ability to facilitate
users' interaction with and exploration of large volumes of data. Netflix, being a top international
streaming platform, utilizes interactive data visualization methods to optimize user experience and
data discovery. This chapter discusses Netflix's real-time capabilities, mechanisms of user
interaction, and visual components, highlighting in fig 9.8; how this drive an intuitive and
interactive data-centric interface.
Real-time features and user interactions:
Fig. 9.9: Netflix visualization using different graphs
1. Filters and Dropdown Menus for Dynamic Data Filtering:
Netflix uses filters and dropdown menus to enable users to dynamically filter datasets according to
their interests. This feature makes it possible for users to browse through particular categories like:
i. Genre: Action, Drama, Comedy, Thriller, etc.
ii. Year of Release: Search for films and television shows by particular time ranges.
iii. IMDb Rating: Choose content on the basis of user ratings and critic scores.
iv. Language and Country of Production: Filter by language preferences or regional content. These
filtering tools offer personalized and pertinent results so that users are able to examine and explore content
effectively as illustrated in the fig 9.9.
2. Hover Tooltips for Contextual Information:
Hover tooltips improve data visualization by offering extra information when users hover over
objects in the interface. Hover tooltips normally show:
i. The quantity of films or television programs related to a given data point.
ii. Title, genre, and rating of the movie on a mouse-over a data marker.
iii. Other metadata, e.g., release date, director, or runtime.
Through hover tooltips, Netflix enhances usability and understanding without overloading users with
too much information on the primary display.
3. Real-Time Data Updates:
Netflix constantly updates and refreshes its information in real-time to keep it accurate and current.
This feature is especially valuable for:
i. Trending content: The most recent popular television shows and films are updated according to
viewership metrics.
ii. Genre rankings: The ranking of genres is automatically modified by the system based on recent
viewing patterns.
iii. IMDb ratings: Any fluctuation in IMDb ratings is updated dynamically within the visualization.
Real-time updates guarantee that the latest information is always accessible, promoting users' and
analysts' decision-making.
4. Drill-Down Capability for In-Depth Exploration:
Netflix also features drill-down functionality, enabling users to click on visual components to
get additional data. The functionality is applied(shown in fig 9.10):
Fig. 9.10: Netflix visual components of data
visualization
i. Geospatial maps: A click on a nation reveals the number of movies or TV shows produced there.
ii. Trend charts by genre: Users can view genre-based trends by choosing specific genres or years.
iii. Best-performing content: A click on a particular show or film gives additional information
regarding its popularity, rating, and reviews.
Drill-down functionality supports a tiered data exploration strategy, where users can move from
macro- level observations to in-depth analysis.
9.4 Netflix's Key Visual Elements for Data Representation:
1. Key metrics, or KPI cards:
Netflix employs Key Performance Indicator (KPI) cards to summarize high-level summary statistics
in a readable format. Such KPI cards illustrated in fig 9.11 show key data points including:
i. Total movies and TV shows on the platform.
ii. Most watched genre, based on viewing statistics.
iii. Top three languages, which refer to the most commonly available content languages.
iv. Average IMDb rating for all titles accessible, giving the quality of the content.
Fig. 9.11: KPI cards
KPI cards provide an instant glance at key measures, enabling people to understand the key
information.
2. Interactive World Map for Geospatial Analysis
Netflix's interactive world map depicts the worldwide distribution of its movie and TV show
productions visually. This map allows users to:
i. Identify the locations of film and TV show productions.
ii. Click on a country to see the total number of movies or TV shows produced there.
iii. Examine geographic content trends to enable users to discover country-by-country entertainment
economies.
Fig. 9.12: Geospatial Analysis
Fig 9.12 represents geospatial visualization offers a natural means of investigating worldwide
media trends and regional production centres.
3. Content Trend by Year (Line Chart):
Netflix uses line charts to demonstrate the trends of content production across time. The visualization
emphasizes:
i. The number of movies and series released each year to monitor the growth of the industry.
ii. Filtering options so that users can analyze trends in terms of:
• Director: Display content created by directors.
• Genre: Evaluate which genres are on the rise or decline.
• Language: Examine content availability in various languages.
• IMDb Rating: Review the development of highly rated content.
Fig. 9.13: Trends using line graph
By allowing interactive filtering, this chart (as shown in fig 9.13) enables monitoring content
trends by multiple parameters, allowing valuable insights to be gained regarding industry trends.
Netflix's interactive data visualization platform is built to give real-time insights, dynamic filtering,
and profound data exploration. With functionalities like dropdown filters, hover tooltips, real-time
refresh, and drill-down, the users can efficiently analyze content trends and find worthwhile
insights. The important visual aspects like KPI cards, interactive maps, and trend charts make data
storytelling easy and provide better access to complicated information.
By combining these sophisticated data visualization methods, Netflix provides smooth,
informative, and engaging user experience, creating a standard for data-driven websites.
Interactive Maps and Geospatial Representation:
One of the most characteristic aspects of NYT's election visualizations is its interactive geospatial
maps, which present election results at various levels, including:
• State level (for presidential and gubernatorial elections)
• County level (for detailed voting patterns)
• District level (for congressional elections)
Fig. 9.14: Election data visualization
These are color-coded, usually blue for the Democrats and red for Republicans, to mirror real-time vote
totals as shown in fig 9.14. Interactive features enable the viewer to move their cursor over or click on an
area to access detailed figures, including overall votes, percentage of margin, and past voting behaviour.
This helps to better present election trends and allows the reader to make comparisons between current and
previous elections.
One of the strongest aspects of NYT's geospatial visualizations is their capacity to display electoral changes
across time. An example is that users can look at how particular counties or states have politically changed
from the past elections to the present, giving indications of swing states as well as changes in voting
behaviour.
1. Principles of Good Visualization Design
Good design isn't about making things "pretty"; it’s about reducing the cognitive load for the viewer.
• Clarity & Simplicity: Remove "chart junk" (unnecessary grid lines, 3D effects, or decorative icons)
that doesn't represent data.
• Integrity: Ensure the visual proportionalities match the numerical proportions. For example,
starting a bar chart axis at something other than zero can misleadingly exaggerate differences.
• Hierarchy: Use size, bolding, and placement to lead the eye to the most important insight first.
• Accessibility: Design with all users in mind, ensuring text is legible and layouts work for people
with visual impairments.
2. Understanding and Using Color
Color is one of the most powerful—and most abused—tools in visualization.
• Types of Color Scales:
o Sequential: Use for ordered data (e.g., light blue to dark blue for low to high sales).
o Diverging: Use for data with a meaningful midpoint, like zero (e.g., red for negative, white
for neutral, blue for positive).
o Categorical: Use distinct hues for unrelated groups (e.g., different colors for "Apples,"
"Oranges," and "Bananas").
• Consistency: Keep colors the same across different charts if they represent the same category.
• Color Blindness: Avoid red-green combinations. Use tools like ColorBrewer to find palettes that
are safe for color-blind viewers.
COLOR USAGE IN DATA VISUALIZATION
The Role of Color in Communication
Color plays a crucial role in visual communication. It deepens comprehension, captures interest, conveys
feelings, and improves accessibility. In areas like branding, user interfaces, data visualization, and
printed materials, color directs perception and influences choices.
As much as color can stir emotions, it also has the strength to confuse, misinterpret, and leave those who
should otherwise access materials inaccessible. It is perhaps one of the most potent weapons in the
arsenal of a designer and data visualizer. But it is only as good as it is well applied. To draw color
appropriately, you have to understand the merits as well as the demerits that come with it
Perhaps one of the biggest advantages is that color carries meaning. Accordingly, it should enhance
reading more effectively by drawing attention to the important parts when used correctly. For example,
if the report is stressed in a different color, this can help people see critical points more quickly. Also,
colors show the distinctions between different categories, groups, or trends. For example, in data
visualization, you can give a specific color to different segments of the bar chart or the slice of a pie
chart so that users can easily understand the patterns and relationships without reading comprehensive
labels. The most common and well-illustrated example of this effective color use is the system of traffic
lights. Red shows "stop," yellow signifies "warning," green indicates "go." This universal applicability
of color means that a driver simply recognizes and reacts to a signal without anything needing to be
added. Similarly, in consistent patterns, the user can understand more and improve the interaction when
applying color consistently in design and data presentation.
Explaining Pitfalls in Data Visualization:
Pitfall 1: Encoding Too Much Information with Color
The Persuasive Use of Color in Design: Color in design, data visualization, or user interface is an
important tool to convey meaning, grab attention, and aid understanding. When applied correctly, color
helps the understanding of information; when wrongly applied—laying too much inference on it or using
colors unrelated to the topic or excessive colors—color clouds meaning, reduces readability, and makes
the data or design fail in the first place. With such a vast assortment of colors embedded into
visualization, there seems to be a feeling that such colors obstruct understanding and instead generate
visual clutter in view of the viewer.
One of the major sins in color usage is jamming too many elements into one chart or design as shown
in fig 5.1. The moment too many colors are used, it detracts the focus from key insights that are really
important for the users. A very clear example of this mistake comes in pie charts where differentiation
is attempted between numerous categories having separate colors for each. Introducing such a plethora
of colors to the chart tends to further concentrate on the data rather than clarifying it, confusing the
observer instead
Fig. 5.1: Encode too much color in graph
A fundamental problem connected with excessive color usage in visual representation is that it impedes
value comparison. If anyone attempts a value comparison having a wide palette of colors in their charts
or graphs, they are likely to focus less and less on finding relationships that matter when distractions
keep protruding among the clutter of hues. It is the distractive palette of colors that may be taking away
their attention, preventing them from arriving at useful information. Having a simple structure with few
carefully chosen colors works best for creating .
Fig. 5.2: Different colors for buttons, text, and backgrounds
An organized color scheme facilitates intuitive navigation. Bad color selections, like low contrast
between the background and text, can render content less readable, particularly for people with visual
disabilities. Eye-stressing colors like bright or too saturated ones lower readability. Furthermore,
neglecting color blindness makes important information unavailable to a part of users. Good contrast
and careful use of colors improve readability and provide inclusivity.
5.1.1 Pitfall 2: Irrelevant Colors
Color is intrinsically important in thinking about data interpretation; misapplying it according to the
normal understandings often results in confusion and misinterpretation. Irrelevant or misleading
coloring may distort meaning as well and thus hinder the viewer's ability to understand and reach
conclusions based on trends.
For example, green would commonly represent such positive results as profits, growth, and success,
whereas red usually indicates caution, loss, or negative trends. By inverting these associations, it
confuses viewers who are conditioned with the normal color coding.
EarnPitfall 3: Employing Nonmonotonic Color Scales to
Represent Data Values
Color is a critical data visualization tool that allows the audience to rapidly understand complicated
information. However, the misuse of color scales can result in misinterpretation, visual disorientation,
and data distortion
There are two types of color scales i.e. Monotonic Color Scale and Nonmonotonic Color Scale. A
monotonic color scale is a logical, perceptually smooth progression that allows viewers to easily relate
increasing or decreasing color intensity to a corresponding change in data values. A nonmonotonic color
scale does not exhibit a smooth or consistent progression.
balanced color scheme ensures emphasis is obtained consciously and consistently, guiding users to the
appropriate conclusions.
Fig. 5.4: The rainbow color scale is non-monotonic
The Nonmonotonic Color Scales shown in fig 5.4 are Problematic because it produce visual aberrations
that mislead the interpretation of the data. Human eyes do not see all colors with the same intensity.
Some colors, for example, yellow, are brighter than others, like blue, even though they might be showing
the same data value. For Example: A rainbow color scale heatmap can misleadingly imply that yellow
areas are "higher" than neighboring colors, even if the values are equal. Utilize perceptually uniform
color scales like Viridis, Inferno, or Cividis that provide equal visual weighting. If nonmonotonic scales
jump from one color to another abruptly (e.g., green → yellow → red), they create artificial visual
boundaries that are not present in the data.
For instance: In a weather map depicted with a rainbow scale, the smooth temperature gradient can be
represented as having "hot zones" where the color changes suddenly although the transition might be
gradual. Employ a single-hue sequential gradient that smoothes its transition from light to dark and does
not include sudden changes. Nonmonotonic scales can mislead trends, particularly when audience
members believe equal color distances imply equal numerical differences. For Example: If the blue-
green-yellow-red scale is used to portray air pollution, individuals will be likely to presume that the
distance between blue and green is identical to the distance between yellow and red, although the
numerical difference is much greater.
5.1.2 Pitfall 4: Failing to Design for Color-Vision Deficiency:
Color vision deficiency (CVD) is a condition that causes individuals to have difficulty differentiating
between specific colors as a result of how their eyes perceive light. About 300 million people globally
are affected, which includes 1 in every 12 men and 1 in every 200 women. It affects designs as most
designers take it for granted that color is perceived by all people in the same manner, resulting in visual
elements that are impossible—or at least very hard—for colorblind people to understand. Overlooking
color accessibility can lead to unclear warnings, alerts, and UI elements .
Fig. 5.5: Color-vision deficiency (CVD) simulation of the sequential
color scale Heat
Fig. 5.7: Custom icons on map
Icons provide an intuitive means of augmenting data representation. Small graphic
symbols can stand in or be used in addition to color coding as shown in fig 5.7, making
visual information more understandable briefly. For instance, on a weather map, a sun
symbol for sunny and a cloud for cloudy weather allows the forecast to be comprehensible
even in black and white. Similarly, in financial reporting, using upward or downward
arrows, in addition to red-green color coding, allows users to easily identify positive or
negative trends.
3. Importance of Data Modelling
Data modelling is the foundation. If the underlying data structure is messy, the visualization
will be confusing or incorrect.
• Data Shaping: Visual tools often require data to be "Long" (tall) rather than "Wide."
Converting spreadsheets into a format where each row is a single observation is crucial.
• Performance: A well-modelled dataset (using star schemas or joined tables) allows the
visualization software to filter and aggregate data quickly without lagging.
• Accuracy: Proper modelling defines the relationships between variables, ensuring that
when you "drill down" into a chart, the numbers remain mathematically sound.
DATA MODELING Ref: DA | Unit-
II
Data modeling helps companies structure and organize data by
designing a visual representation of it.
What Is Data Modeling?
Data modeling is the process of evaluating and defining different sources
and types of data that your company works with. Simply put, it establishes
connections between pieces of information and categorizes them into
logical groups. That is achieved by creating a visual representation of data
with all of its attributes, relationships, and storage locations.
In general, data modeling acts as a well-defined roadmap for data
management, helping organizations plan their data architecture more
efficiently. On top of that, it supports stakeholders in better decision-
making by providing ground for data analytics and facilitating it.
Concepts of Data Modeling
Three main data modeling concepts.
Conceptual: It is typically used at an early stage of the project when we
analyze requirements. It provides a high-level overview of what the
system will include and how it will be organized.
Logical: This model goes one step further. As the name implies, it helps
break data into minor logical elements and build a detailed visual schema
of relationships between them. It aids organizations in streamlining
approaches for data consolidation and segmentation.
Physical: This one derives from a logical concept and helps describe how
data will be structured within a specific database management system. It
serves like a guide for data engineers during the visualization and
implementation of databases.
Common Data Modeling Techniques
As we’ve already mentioned, when well-structured, a data model serves
as a basis for subsequent analysis and informed decision-making.
#1 Entity-Relationship (ER) Model
Though introduced in 1976, the entity-relationship data model still
remains a thing these days. It illustrates the connections between entities
in a database using formal diagrams. They assist in understanding the
fundamentals of the information that will be located within the data store.
The ER model contains three main components: entities, attributes, and
relationships. Entity symbolizes a real-world entity, such as a person or
location, and is displayed on tables. As for attributes, they explain the
features of each entity. And the relationship is a connection between two
or more entities that can take several forms, such as one-to-one, one-to-
many, or many-to-many.
#2 Relational Model
Another technique that was introduced a long time ago is a relational
model. It is still widely used in database architecture, connecting data in
tables through rows and columns.
It aims to simplify the data and offer a clear perspective, as well as
facilitate efficient storage and analysis. Moreover, it utilizes the principles
of set theory and predicate logic.
For example, relational databases can help handle massive amounts of
essential customer information, monitor stocks, conduct e-commerce
transactions, and much more.
#3 Dimensional Model
Dimensional models are mostly used in data warehouse design with the
goal to optimize a database for faster retrieval of information. They also
help remove redundancy and inconsistencies, thus contributing to better
data quality.
These models have numerous benefits, including the capacity to organize
data for analysis in a simple, clear, and customizable manner.
Furthermore, they enable the quick and easy addition of new data, which
can support companies with dynamic business environments.
In a dimensional model, information is divided into two types of tables:
fact and dimension. Fact tables are often large and contain millions or
billions of rows. For example, they can store data about sales
transactions, customer orders, and more.
On the other hand, dimension tables are relatively smaller in size. They
contain descriptive information that helps understand data in the context
of the business environment.
#4 Data Warehouse Modeling
The next technique in our list is data warehouse modeling. It is the
process of building and arranging data models within your data
warehouse platform.
Overall, there are 3 types of data warehouse modeling:
Enterprise Warehouse: An enterprise warehouse integrates data from
diverse sources, as a result it reduces data silos and ensures accurate and
reliable information throughout the organization. Generally, it comprises
both extensive and summary information and can range in size from a few
gigabytes to hundreds of gigabytes, terabytes, and even more.
Here the warehouse serves as a repository for both structured and
unstructured data. It helps enterprise software developers use this data to
build various solutions for reporting and analysis.
Data Mart: The second type of data warehouse modeling is data mart. In
simple words, it is a portion of corporate-wide data that is useful to a
certain group of users. Overall, it focuses on an exact industry,
department, or a particular business area.
Take, for instance, marketing data. It can contain data entities such as
customers, items, campaigns, sales, and website analytics.
Virtual Warehouse: The third one is a virtual warehouse. It is a cloud-
based storage and processing environment where organizations can store,
manage, and analyze their data without the need for physical
infrastructure. Virtual warehouse takes data management a step further
by promoting agility, scalability, and cost-effectiveness.
#5 Object-Oriented Model
Another technique of data modeling is called object-oriented, where data
is stored in the form of objects. This model is based on the object-oriented
programming approach and involves designing data models that mirror
real-world objects and their relationships. Each object has its own set of
attributes and behaviors, or methods.
For example, let’s consider a data model for an educational system. In this
case, you might have classes like student, teacher, course, and
department. Students then could have attributes such as ID, name, major,
and a list of enrolled courses. And courses could have attributes like
course ID, title, department, and a list of enrolled students, etc.
The object-oriented model is flexible and adaptable since it allows for
quick model modification and enhancement as requirements change. Thus
you can benefit from it when you need, let’s say, to add new object types
without changing the old ones.
#6 Hierarchical Database Model
This approach organizes data in the form of a tree, maintaining a parent-
child relationship within records. A parent can have more than one child,
while a child record is limited to having only one parent.
Hierarchical model helps maintain data consistency since changes to a
parent record are automatically transmitted to its children. Furthermore,
by limiting access to certain levels of the hierarchy, you can gain
comprehensive control over data.
#7 Network Database Model
The network model is an expanded version of the hierarchical one. It
enables a child record to have one or more parents. When compared to
the hierarchical approach, it allows for more flexible data access.
By incorporating the network model technique into the data modeling
process, you may capture the intricate relationships and interactions
between various data points, enhancing the accuracy of the overall data
model.
#8 Big Data Modeling
When it comes to analyzing big data, it’s quite challenging to implement it
using the traditional methods due to large volumes of data. In order to
manage this process smoothly, you can leverage big data modeling.
This technique is tailored to handle the unique characteristics of big data,
such as volume, velocity, variety, and variability. By using visual models,
like diagrams, charts, and graphs, you can easily generate insights from
massive and complicated datasets.
Big data modeling is often empowered by machine learning and artificial
intelligence technologies that help businesses comprehend and evaluate
vast amounts of data. AI and ML can be used for many purposes such as
automated data collection, making predictions, deriving insights, and so
on.
#9 Agile Data Modeling
Last but not least, is agile data modeling. As the name suggests, it
combines agile software development approaches, like the Agile
Manifesto and Scrum, with the process of data modeling. It aims to
develop database architectures that are flexible and responsive to
changes in the requirements.
In general, by using this approach, it is possible to refine the data model
based on the feedback from end users and evolving business needs. This
dynamic method enables your development team to quickly adapt the
data model and incorporate new elements