Unit 2 Data Visualization and Data Preprocessing
Unit 2 Data Visualization and Data Preprocessing
Data Visualization
In the world of big data, data visualization tools and technologies are essential to
IT
analyze massive amounts of information and make data-driven decisions. Our culture is
visual, including everything from art and advertisements to TV and movies, and our eyes
BP
are drawn to colors and patterns. Our interaction with data should reflect this reality.
n our increasingly data-driven world, it’s more important than ever to have accessible
i,
ways to view and understand data. After all, the demand for data skills in employees is
an
steadily increasing each year. Employees and business owners at every level need to
have an understanding of data and of its impact.
w
That’s where data visualization comes in handy. With the goal of making data more
accessible and understandable, data visualization in the form of dashboards is the go-to
as
visual elements like charts, graphs, and maps, data visualization tools provide an
accessible way to see and understand trends, outliers, and patterns in data. Additionally,
rti
In the world of Big Data, data visualization tools and technologies are essential to
analyze massive amounts of information and make data-driven decisions.
r.
visualization?
Something as simple as presenting data in graphic format may seem to have no
downsides. But sometimes data can be misrepresented or misinterpreted when placed
in the wrong style of data visualization. When choosing to create a data visualization,
it’s best to keep both the advantages and disadvantages in mind.
Advantages
Our eyes are drawn to colors and patterns. We can quickly identify red from blue, and
squares from circles. Our culture is visual, including everything from art and
advertisements to TV and movies. Data visualization is another form of visual art that
grabs our interest and keeps our eyes on the message. When we see a chart, we quickly
see trends and outliers. If we can see something, we internalize it quickly. It’s
storytelling with a purpose. If you’ve ever stared at a massive spreadsheet of data and
couldn’t see a trend, you know how much more effective a visualization can be.
IT
Some other advantages of data visualization include:
BP
● Easily sharing information.
● Interactively explore opportunities.
● Visualize patterns and relationships.
i,
Disadvantages
an
w
as
While there are many advantages, some of the disadvantages may seem less obvious.
For example, when viewing a visualization with many different datapoints, it’s easy to
make an inaccurate assumption. Or sometimes the visualization is just designed wrong
H
Big data has become more than a buzzword as information has grown more complex
rti
and vast in quantity and organizations struggling to gather, curate, understand, and use
data effectively. It also describes challenges in IT, business, as well as emerging
Aa
analytics technologies. But where did the term come from, how can you use big data at
your organization, and how can you advance your big data analytics strategies? We’ll
address these questions and provide tips to get started using your big data.
r.
In the 1960s, the United States created a large data center to store millions of tax
records. This data center was the first real use case of digital data management.
Through the 1990s and 2000s, leaders in the data space worried that existing
technology would not be able to store large amounts of information produced by
businesses, by government, and by people around the world. An even larger worry:
would anyone be able to make sense of that much data. Today, thanks to technology
innovations and more sophisticated analytics capabilities, it’s become cheaper and
easier to store data and then analyze it. Now, concerns have shifted to effectively using
big data in many formats such as structured, unstructured, and semi-structured, to
inform business decisions.
IT
make big data a big deal. The four Vs distinguish and define big data and describe its
challenges.
BP
1. Volume
The most well-known characteristic of big data is the volume generated. Businesses
have grappled with the ever-increasing amounts of data for years. However, now it’s
i,
possible to store data for pennies on the dollar using data lakes or data warehouses like
an
Snowflake. Businesses prioritize data organization with platforms like Hadoop, but it’s
important to develop policies that standardize how long users keep data, and then a
procedure for deleting or archiving it. Ultimately, it doesn’t matter how big the data is if
w
you can’t use it. With increasing volume, users’ hands will be tied behind their backs
unless information is stored and governed by an agile, accessible framework. Without
as
strategies to use and access huge volumes of data, that data will sit stagnant, lose
value, and fail to surface insights.
H
2. Velocity
rti
Not only are businesses producing a lot of data, but they are also doing it at an
ever-increasing rate. Customers and employees use many applications to complete
Aa
data-driven tasks. Technologies have to be ready for the speed and volume of that data
to keep up with the pace of business. Because the volume is high, velocity becomes
more and more difficult to manage as it becomes more important. Speed to insight is a
serious consideration in both data software as well as data structure.
r.
D
3. Value
Velocity, volume, and variety were the three original characteristics associated with big
data. However, leaders in the space have added value, recognizing the opportunity and
need to use big data for business transformation—both to fulfill goals and metrics,
uncover risks and opportunities that no one realized, and much more. Analytics
technologies have evolved and can now augment analysts’ abilities to find correlations,
identify outliers, and predict outcomes with data. For example, sales teams can use a
platform to connect data from social media, eCommerce, and sales for a full picture of
the customer journey—from awareness to conversion. As another example, AI
technologies like Explain Data can suggest possible explanations for outliers for
analysts to drill into.
4. Variety
IT
organizations, that means there are more types of data to store and analyze, plus more
sources to pull from. The variety of structured has exploded and it has become more
BP
important to leverage metadata as well as ways to wrangle unstructured data. Location
data In the following section, we will review the types of big data that exist.
i,
an
Big data is crucial because of its untapped potential, but recent technology such as
visual analytics finally allows businesses to discover critical, even surprising insights
that give us a clearer view into processes and human behaviors. On some occasions,
w
these processes and behaviors must be refined for the sake of the business and its
future success. And recognizing that no organization is exactly the same, the
as
technology providers that invest ahead of the curve in big data—whether that’s creating
partnerships in the data ecosystem or developing capabilities that improve data access
H
Structured data
rti
Structured data is the neatly organized data you keep in databases, datasets, and
Aa
spreadsheets. It’s easy for traditional analytics tools to read this data. Organizing
unstructured data into structured data is time-consuming, but possible with the right
solution. It involves data cataloging, data mapping, and data transformation. You can
learn more about these processes here.
r.
D
Unstructured data
Semi-structured
Semi-structured data has some organizational structure, but isn’t easy to analyze as-is.
With some organizing or cleaning, semi-structured data could be imported into a
relational database just like structured data. Semi-structured data and structured data
can be analyzed and visualized with solutions like Tableau. With a combination of
IT
solutions like Hadoop and Tableau, all three of these types of data can be used for
analysis.
BP
Three best practices to follow with big data
It’s easy to be overwhelmed by big data, but the good news is that technologies and
i,
analytics platforms are becoming more efficient and comfortable to use. Industries and
an
teams such as sales, IT, and government agencies have used their big data to discover
trends and reduce analysis time. To effectively use big data, focus on the following best
practices:
w
1. Build with flexibility and long-term sustainability in mind
as
A big data best practice is to always think about long-term solutions. As discussed, big
data is growing at a fast, steep trajectory; so your data management solution and
H
strategy should scale with it. Technology will evolve and work together more easily.
Don’t be afraid to upgrade to new innovative solutions. Example: Abercrombie & Fitch
rti
saw the benefits of upgrading from spreadsheets and deployed Tableau fast so they
could use match shopper insights with inventory.
Aa
Secondly, be aware that companies often need more than one solution to manage their
r.
big data from first ingestion to final data visualizations. This isn’t a bad thing. The best
big data platforms can talk to each other and form a symbiotic relationship. Example:
D
PepsiCo worked with Tableau and Trifacta to wrangle disparate data and uncover
insights.
Third, keep in mind that data cultures empower their business teams to use and play
with data. With big data, all hands are on deck. That’s the only way to keep up with the
volume and velocity of data. Example: Charles Schwab knew this, so they democratized
data analysis across hundreds of branch locations.
IT
● Simplifies Complex Data: It turns large and complicated data into visual formats
like charts and graphs, making the information easier to understand.
BP
● Reveals Patterns and Trends: It helps identify trends, relationships and patterns
that are not easily seen in raw data or tables.
● Saves Time: Visuals allow quicker interpretation of data, helping users spot key
i,
information at a glance instead of manually scanning through numbers.
● Improves Communication: It makes it easier to explain data insights to others,
an
especially those who may not be familiar with the technical details.
● Tells a Clear Story: Data visuals guide the audience through the information
w
step-by-step, making it easier to reach conclusions and make informed
decisions.
as
Business Analytics: Used to monitor company performance, track KPIs and make
data-driven decisions by visualizing trends, sales and customer metrics.
Sports: Used to visualize player statistics, team performance and match outcomes,
helping coaches and analysts improve strategies and training plans.
Retail and E-commerce: Enables tracking of sales, customer preferences and inventory
levels, helping businesses adjust stock and marketing efforts effectively.
Common Types of Data Visualization
There are various types of visualizations where each has a unique purpose in data
representation. Here are the most common types:
Charts and Graphs: They are used to visualize data, with charts comparing data points
across categories or showing trends over time and graphs analyzing relationships
IT
between variables to identify correlations, trends and outliers. Examples: Bar Charts,
Line Charts, Pie Charts, Scatter Plots, Histograms, Box Plots.
BP
Maps: They are used to display geographical data which provides spatial context to
trends and patterns. Examples: Geographic Maps, Heat Maps
Dashboards: They combine multiple visualizations into a single interface which provides
i,
real-time insights and interactive features for users to explore data.
an
Types of Data Visualization Charts: From Basic to Advanced
w
Data visualization includes a variety of charts, each designed to present data in a clear
and meaningful way. From simple bar and line charts to advanced visuals like heatmaps
as
and scatter plots, the right chart helps turn raw data into useful insights.
Let’s explore some common types of charts from basic to advanced and understand
H
Basic charts are best suited for displaying simple comparisons, trends over time and
Aa
basic relationships within the data. These charts are easy to understand and ideal for
communicating insights to a broad audience.
r.
D
1. Bar Charts
Bar charts are used to compare values across different categories using rectangular
bars. X-axis shows categories while Y-axis represents values. Common types include
horizontal, stacked and grouped bar charts.
Below is the Example of Bar Chart:
bar-chart
IT
BP
i,
an
w
as
H
rti
Aa
When to Use:
r.
2. Line Charts
Line charts show how values change over time by connecting data points with lines.
They help visualize trends like increases, decreases or stability.
Below is the example of line chart:
line-chart
IT
BP
i,
an
w
as
H
rti
Aa
When to Use:
r.
D
3. Pie Charts
Pie charts are round charts divided into slices, where each slice shows a part of the
whole. The size of each slice represents its percentage.
IT
pie-chart
BP
i,
an
w
as
H
rti
Aa
r.
D
When to Use:
Scatter charts use dots to show relationship between two numerical variables. X-axis
shows the independent variable and Y-axis shows the dependent variable.
IT
scatter-plot
BP
i,
an
w
as
H
rti
Aa
r.
D
When to Use:
A histogram displays the distribution of numerical data by grouping values into intervals
(bins) and showing their frequency as bars. It helps reveal the shape, spread and
patterns in the data.
IT
BP
histogram
i,
an
w
as
H
rti
Aa
r.
D
When to Use:
Advanced charts are designed to handle more complex data. They help analyze multiple
variables, uncover deeper insights and reveal patterns that might be missed with basic
visuals.
IT
1. Heatmap
BP
A heatmap displays data in a matrix format using color to represent values. It's ideal for
spotting patterns, correlations and variations in large datasets.
i,
Below is the example of heatmap:
an
w
heatmap
as
H
rti
Aa
r.
D
IT
BP
i,
an
w
as
H
When to Use:
2. Area Chart
r.
An area chart shows trends over time by filling the space beneath a line. It's ideal for
D
When to Use:
rti
A box plot displays distribution of numerical data, showing median, quartiles and
outliers. It’s useful for understanding variability and detecting unusual values.
When to Use:
rti
Aa
4. Bubble Chart
A bubble chart displays data points as circles, where size and color of each bubble
represent additional variables. It’s useful for visualizing three or more dimensions in a
single chart.
Below is the example of bubble chart:
IT
BP
i,
an
w
as
H
rti
This bubble chart shows the number of spacewalks that happened every year between
Aa
the years 1998 and 2019. Each bubble is proportional to the number of spacewalks and
shown in chronological order.
r.
● The bubbles overlap, showing trends over time. It is obvious that between 2005
D
and 2011 there were more spacewalks than any other time period.
● The bubbles are all one color, leaving the viewer aware that each bubble
measures the same dimension
● The bar chart above shows the data in the same way as it is presented in the
bubble chart. However, the relationship between years is not as obvious in the
bar chart.
When to Use:
5. Tree Map
IT
A tree map visualizes hierarchical data using nested rectangles, where the size of each
represents a value. It’s useful for showing structure and comparing proportions within a
BP
hierarchy.
i,
Below is the example of tree map:
an
w
as
H
rti
Aa
r.
D
When to Use:
6. Network Graph
IT
A network graph shows relationships between entities as nodes and links (edges). It
helps visualize complex networks like social media, transport routes or biological
BP
systems.
i,
Below is the example of a network graph:
an
w
as
H
rti
Aa
r.
D
When to Use:
Relationship Visualization: Helps show connections or interactions between entities.
IT
A donut chart is a circular chart like a pie chart but with a hole in the center. Each slice
shows a category’s contribution to the whole, making it visually cleaner and ideal for
BP
comparisons.
i,
Below is the example of a donut chart:
an
w
as
H
rti
Aa
r.
D
When to Use:
8. Gauge Chart
A Gauge chart shows progress toward a goal using a dial-like arc, similar to a
speedometer. It is useful for tracking a single value like a KPI or project status.
IT
Below is the example of Gauge chart:
BP
i,
an
w
as
H
When to Use:
rti
Aa
9. Sunburst Chart
A sunburst chart shows hierarchical data as nested rings. Each ring represents a level in
the hierarchy making it useful for visualizing multi-level data structures like categories,
subcategories or organizational hierarchies.
Below is the example of sunburst chart:
IT
BP
i,
an
w
as
IT
● Ensures the accuracy and consistency of the dataset.
BP
Steps in Data Preprocessing
Some key steps in data preprocessing are Data Cleaning, Data Integration, Data
Transformation, and Data Reduction.
i,
an
w
as
H
rti
Aa
r.
D
IT
● Binning Method: The data is sorted into equal segments, and each segment is
smoothed by replacing values with the mean or boundary values.
BP
● Regression: Data can be smoothed by fitting it to a regression function, either
linear or multiple, to predict values.
● Clustering: This method groups similar data points together, with outliers either
being undetected or falling outside the clusters. These techniques help remove
i,
noise and improve data quality.
an
● Removing Duplicates: It involves identifying and eliminating repeated data entries
to ensure accuracy and consistency in the dataset. This process prevents errors
and ensures reliable analysis by keeping only unique records.
w
2. Data Integration: It involves merging data from various sources into a single, unified
as
● Record Linkage is the process of identifying and matching records from different
rti
datasets that refer to the same entity, even if they are represented differently. It
helps in combining data from various sources by finding corresponding records
Aa
3. Data Transformation: It involves converting data into a format suitable for analysis.
Common techniques include normalization, which scales data to a common range;
standardization, which adjusts data to have zero mean and unit variance; and
discretization, which converts continuous data into discrete categories. These
techniques help prepare the data for more accurate analysis.
● Data Normalization: The process of scaling data to a common range to ensure
consistency across variables.
● Discretization: Converting continuous data into discrete categories for easier
analysis.
● Data Aggregation: Combining multiple data points into a summary form, such as
averages or totals, to simplify analysis.
● Concept Hierarchy Generation: Organizing data into a hierarchy of concepts to
IT
provide a higher-level view for better understanding and analysis.
BP
4. Data Reduction: It reduces the dataset's size while maintaining key information. This
can be done through feature selection, which chooses the most relevant features, and
feature extraction, which transforms the data into a lower-dimensional space while
preserving important details. It uses various reduction techniques such as,
i,
an
● Dimensionality Reduction (e.g., Principal Component Analysis): A technique that
reduces the number of variables in a dataset while retaining its essential
information.
w
● Numerosity Reduction: Reducing the number of data points by methods like
sampling to simplify the dataset without losing critical patterns.
as
Data preprocessing is utilized across various fields to ensure that raw data is
transformed into a usable format for analysis and decision-making. Here are some key
Aa
ensures the data is consistent and reliable for future queries and reporting.
2. Data Mining: Data preprocessing in data mining involves cleaning and transforming
raw data to make it suitable for analysis. This step is crucial for identifying patterns and
extracting insights from large datasets.
3. Machine Learning: In machine learning, preprocessing prepares raw data for model
training. This includes handling missing values, normalizing features, encoding
categorical variables, and splitting datasets into training and testing sets to improve
model performance and accuracy.
IT
5. Web Mining: In web mining, preprocessing helps analyze web usage logs to extract
BP
meaningful user behavior patterns. This can inform marketing strategies and improve
user experience through personalized recommendations.
i,
data to create dashboards and reports that provide actionable insights for
an
decision-makers.
● Improved Data Quality: Ensures data is clean, consistent, and reliable for
analysis.
● Better Model Performance: Reduces noise and irrelevant data, leading to more
rti
● Organizes categorical data into a matrix where each cell shows the frequency or
IT
count of a category combination.
● Enables clear comparison between different groups making associations and
BP
patterns easier to identify.
● Commonly applied in data analysis, survey research and machine learning
preprocessing to explore variable relationships before advanced modeling.
i,
Cross-tabulation analysis, also known as contingency table analysis, is most often used
an
to analyse categorical (nominal measurement scale) data.
At their core, cross-tabulations are simply data tables that present the results of the
w
entire group of respondents, as well as results from subgroups of survey respondents.
With them, you can examine relationships within the data that might not be readily
as
a cross-tabulation is a two- (or more) dimensional table that records the number
H
(frequency) of respondents that have the specific characteristics described in the cells
of the table. Cross-tabulation tables provide a wealth of information about the
rti
Cross-tabulation analysis has its unique language, using terms such as “banners”,
Aa
The cells of the table report the frequency counts and percentages for the number of
respondents in each cell.
D
IT
BP
Tabulation professionals call the column variables in these multiple tables “Banners”
i,
and row variables “Stubs”.
an
When should you use cross-tabulation?
w
You typically use cross tabulation when you have categorical variables or data – e.g.
information that can be divided into mutually exclusive groups.
as
For example, a categorical variable could be customer reviews by region. You divide this
information into reviews per geographical area: North, South, East, West, or state, and
H
Another example of when to use cross-tabulation is with product surveys – you could
ask a group of 50 people “Do you like our products?” and use cross-tabulation to get a
Aa
more insightful answer. Rather than just recording the 50 responses, you can add
another independent variable, such as gender, and use cross-tabulation to understand
how the male and female respondents view your product.
r.
With this information, you might see that your female customers prefer your products
D
more than your male customers. You can then use these insights to improve your
products for your male customers.
As such, using two variables – gender and product likability – along with
cross-tabulation, you can get a more comprehensive breakdown of data sets to identify
patterns, trends, or other useful information.
As a result, cross-tabulation is excellent for assessing categorical variables in market
research or survey responses, as you can readily compare data sets to discover the
relationship between two (or more) seemingly unrelated items.
IT
variables to get a better understanding.
● Exit interviews: why are employees leaving and what could you do differently? Is
BP
there a correlation between employees leaving and progression, for example?
● Departmental issues: are issues caused by one or more variables and if so, are
these issues more prevalent at specific job levels?
● Evaluation surveys at schools: how do students feel about the course material
i,
and is the time spent on the material sufficient enough?
an
● Product research and feedback: how do certain demographics feel about your
product? What would they like to see and how does feedback differ from region
to region? What about customer satisfaction?
w
What are the benefits of cross-tabulation?
as
As a statistical analysis method that allows categorical evaluation across a data set,
cross-tabulation can help to uncover variables or multiple variables that affect a specific
H
With the examples above, you should now have a good idea of how to cross-tabulation
can be used in certain contexts to glean insights. But there are several other benefits to
Aa
cross-tabulation:
● Error reduction: analysing data sets can be confusing, let alone accurately pulling
insights from them. Using cross-tabulation, you can make your data sets more
r.
manageable at scale (as they simplify them and divide them into representative
D
subgroups).
● More insights: cross-tabulation looks at the relationships between one or more
categorical variables to uncover more granular insights. These insights might go
unnoticed with standard approaches (or require more work to reveal).
● Actionable information: as cross-tabulation simplifies data sets and allows you
to quickly compare the relationships between them, you can uncover insights
faster and apply new strategies as necessary.
How to do cross-tabulation analysis in Microsoft Excel
In Microsoft Excel, cross-tabulation tables (or crosstabs) can be automated using the
Pivot Table. You can either use the Pivot Table icon in the toolbar or click on Data >
Pivot Table and Pivot Chart Report.
Alternatively, you can press Insert and then click the PivotChart button. Once clicked, the
PivotChart dialog box will open. Select the data that should be used in your crosstab
analysis and select where you want it to be placed.
IT
You can experiment with the individual parameters of the crosstab, but select the
BP
columns that you want to compare and analyse to get the information you need.
i,
an
The Chi-square statistic is the primary statistic used for testing the statistical
significance of the cross-tabulation table. Chi-square tests determine whether or not
the two variables are independent. If the variables are independent (have no
w
relationship), then the results of the statistical test will be “non-significant” and we are
not able to reject the null hypothesis, meaning that we believe there is no relationship
as
between the variables. If the variables are related, then the results of the statistical test
will be “statistically significant” and we can reject the null hypothesis, meaning that we
H
The chi-square statistic, along with the associated probability of chance observation,
rti
may be computed for any table. If the variables are related (i.e., the observed table
relationships would occur with very low probability, say only 5%) then we say that the
Aa
This means that the variables have a low chance of being independent. Students of
statistics will recall that the probability values (.05 or .01) reflect the researcher’s
r.
willingness to accept a type I error, or the probability of rejecting a true null hypothesis
D
(meaning that we thought there was a relationship between the variables when there
wasn’t).
The chi-square statistic is computed by first computing a chi-square value for each
individual cell of the table and then summing them up to form a total chi-square value
for the table. The chi-square value for the cell is computed as: (Observed Value –
Expected Value)2 / (Expected Value). The chi-Square computations are highlighted in
grey.
In this example table, we observe that the chi-square value for the table is 19.35, and
has an associated probability of occurring by chance less than one time in 1000. We,
therefore, reject the null hypothesis of no difference and conclude that there must be a
IT
relationship between the variables. We can observe the relationship in two places in the
table.
BP
The most obvious is in the chi-square value computed for each cell. We observe that the
i,
cells “Red Socks and Boston”, “Blue Jays and Montreal” and “Red Socks and Montpellier,
Vermont” were the three cells where the number of observed respondents was greater
an
than expected. We further note that when we examine the expected and observed
frequencies, the “Yankees and Montreal”, “Red Socks and Montpellier, Vermont”, and
w
“Red Socks and Montreal” frequencies were fewer than expected.
as
H
rti
Aa
r.
D
IT
BP
i,
an
w
as
H
Because the cell chi-square and the expected values are often not displayed, these
same relationships can be observed by comparing the column total percent to the cell
rti
percent (of the row total). In the cell “Red Socks and Boston” we would compare 41.10%
with 64.71% and observe that more Red Socks fans liked Boston than expected. Caution
Aa
In the current table, we observe that “Red Socks and Boston” had the greatest delta
D
between the number of observed and expected respondents, for any team preference
and city of residence. However, we must be careful in concluding that the Red Socks
caused respondents to move to Boston, or that Boston as a city of residence causes fan
loyalty. Red Socks and Boston are the most observed fan and city relationship, but are
most likely totally independent when considering other concepts or relationships.
Crosstabs and chi-square are powerful ways to analyse your survey data.
Data Modeling Explained: Techniques, Examples, and Best Practices
Data modeling is a crucial step in database design and management. Although it may
initially appear to be only a technical activity, it plays an important role in ensuring that
data is properly structured, easily accessible, and ready for analysis.
A well-designed data model helps organizations manage their data efficiently and
supports effective decision-making. In the absence of a strong data model, even
IT
advanced databases can become difficult to handle, leading to data inconsistencies and
operational inefficiencies.
BP
Data modeling is equally important when designing a database from scratch as well as
when refining or improving an existing system. A clear understanding of data modeling
enables organizations to make better use of their data for analysis, reporting, and
business insights.
i,
an
Data modeling is the process of creating a structured and visual representation of data,
defining how data elements relate to one another within a system. It helps translate
business requirements into organized data structures that support accurate analysis
w
and effective decision-making.
as
Defining data elements and their relationships helps teams organize information to
support efficient storage, retrieval, and analysis—improving both performance and
decision-making.
Data Model
A data model is a visual and logical representation of an organization data elements
and the relationships between them. It helps structure and organize data in alignment
with business processes, enabling effective communication between business users
and technical teams. Data models define how data is stored, accessed, shared and
maintained across information systems.
IT
BP
i,
an
w
There are three main types of data models. Let’s explore them in this section.
as
A conceptual model provides a high-level view of the data. This model defines key
business entities (e.g., customers, products, and orders) and their relationships without
getting into technical [Link] is primarily used in the early stages of a project to define
rti
● Purpose: Define tables, columns, relationships and constraints that form the data
structure.
● Focus: Structure of the data without depending on any specific database
IT
management system (DBMS).
● Audience: Data architects and analysts.
BP
● Example Use: Outlining the schema, relationships and rules for customer and
order data, which later guides the physical database design.
i,
an
3. Physical data model
A physical data model represents how data is actually stored in a database. This model
w
defines the specific table structures, indexes, and storage mechanisms required to
optimize performance and ensure data integrity. It translates the logical design into a
as
columns, keys and constraints (primary key, foreign key, NOT NULL, etc.).
● Focus: Actual implementation of the database using queries and the chosen
rti
DBMS features.
● Audience: Developers and database administrators (DBAs).
Aa
● Example Use: Creating the database schema and ensuring that all constraints
and relationships are enforced in the physical database.
.
r.
The first step is to identify and examine data sources, both internal and
external to the organization.
IT
contribute to the overall information landscape.
BP
gathering all relevant data, laying the foundation for an accurate and
complete representation of the data ecosystem.
i,
2. Defining Entities and Attributes
an
At this stage, data modelers identify entities (objects or concepts) and their
attributes (characteristics).
w
Entities: Represent the main subjects of the data (e.g., Customer, Product).
as
Product Price).
3. Mapping Relationships
Purpose: Identify and describe the links, including nature and cardinality
D
IT
creating a model suited for the type of data being handled.
BP
5. Implementing and Maintaining
i,
an
Implementation Tasks: Creating tables, defining constraints, adding
database-specific details.
w
Maintenance: Updating the model to reflect changes in business
requirements or technology.
as
ER modeling is one of the most common techniques used to represent data. It’s
concerned with defining three key elements:
IT
Consider an online store. You might have the following entities:
BP
2. Orders (with Order_ID, Order_Date, Total_Amount)
3. Products (with Product_ID, Product_Name, Price)
4. The relationships could be:
i,
5. "Customers place Orders" (One-to-Many)
6. "Orders contain Products" (Many-to-Many)
i,
ER Model in Database Design Process
an
We typically follow the below steps for designing a database for an application.
w
● Gather the requirements (functional and data) by asking questions to the
as
database users.
● Create a logical or conceptual design of the database. This is where ER
model plays a role. It is the most used graphical representation of the
H
modeling of objects which makes them intently useful. Unlike technical schemas,
D
IT
BP
i,
an
w
Entity
as
Example of entities:
rti
The entity type defines the structure of an entity, while individual instances of that
r.
Entity Set
An entity refers to an individual object of an entity type, and the collection of all
entities of a particular type is called an entity set. For example, E1 is an entity
that belongs to the entity type "Student," and the group of all students forms the
entity set.
In the ER diagram below, the entity type is represented as:
IT
BP
i,
an
w
as
depicted with a single rectangle (e.g., EMPLOYEE). Weak entities depend on a strong
D
entity for existence, use a partial key (discriminator), and are shown with a double
rectangle (e.g., DEPENDENT). Weak entities require total participation in relationships,
while strong entities do not.
Key Differences Between Strong and Weak Entity Sets
IT
entity.
● ER Diagram Representation: Strong entities use a single rectangle. Weak
BP
entities use a double rectangle.
● Relationship: A strong entity-to-strong entity relationship uses a single diamond.
A weak entity-to-strong entity relationship uses a double diamond (identifying
i,
relationship).
an
● Participation: Strong entities may have partial participation. Weak entities always
have total participation.
w
Examples
as
Identifying Relationship: A weak entity is connected to its owner entity via a double
r.
diamond in ER diagrams. Total Participation: A weak entity cannot exist without its
D
i,
Student. In ER diagram,
Types of Attributes
1. Key Attribute
an
w
The attribute which uniquely identifies each entity in the entity set is called the key attribute. For
example, Roll_No will be unique for each student. In ER diagram, the key attribute is
as
2. Composite Attribute
Aa
An attribute composed of many other attributes is called a composite attribute. For example, the
Address attribute of the student Entity type consists of Street, City, State, and Country. In ER
diagram, the composite attribute is represented by an oval comprising of ovals.
r.
D
3. Multivalued Attribute
An attribute consisting of more than one value for a given entity. For example,
Phone_No (can be more than one for a given student). In ER diagram, a
multivalued attribute is represented by a double oval.
IT
4. Derived Attribute
BP
An attribute that can be derived from other attributes of the entity type is known
as a derived attribute. e.g.; Age (can be derived from DOB). In ER diagram, the
derived attribute is represented by a dashed oval.
i,
an
w
The Complete Entity Type Student with its Attributes can be represented as:
as
H
rti
Aa
r.
D
D
r.
Aa
rti
H
as
w
an
i,
BP
IT
Relationship Type and Relationship Set
IT
BP
i,
Entity-Relationship Set
an
A set of relationships of the same type is known as a relationship set. The
w
following relationship set depicts S1 as enrolled in C2, S2 as enrolled in C1, and
S3 as registered in C3.
as
H
rti
Aa
r.
D
Relationship Set
The number of different entity sets participating in a relationship set is called the
degree of a relationship set.
IT
person is married to only one person.
BP
i,
an
w
Unary Relationship
as
Binary Relationship
i,
the relationship is called an n-ary relationship.
an
w
as
H
rti
Aa
Cardinality in ER Model
r.
D
1. One-to-One
When each entity in each entity set can take part only once in the relationship,
the cardinality is one-to-one. Let us assume that one person can be issued only
one passport, and one passport is issued to only one person. So, the relationship
will be One-to-One (1 : 1), meaning that each person has a single passport, and
each passport belongs to a single person.
IT
BP
i,
an
one to one cardinality Using Sets, it can be represented as:
w
as
H
rti
Aa
r.
D
2. One-to-Many
i,
an
w
as
H
rti
Aa
3. Many-to-One
r.
D
When entities in one entity set can take part only once in the relationship set and
entities in other entity sets can take part more than once in the relationship set,
cardinality is many to one.
Let us assume that multiple surgeries can be performed by one surgeon, but one
surgery is performed by only one surgeon. So, the cardinality will be M to 1,
meaning that many surgeries can be done by a single surgeon, but each surgery
is done by only one surgeon.
IT
BP
i,
many to one cardinality
an
Using Sets, it can be represented as:
w
as
H
rti
Aa
r.
D
In this case, each student is taking only 1 course but 1 course has been taken by
many students.
4. Many-to-Many
When entities in all entity sets can take part more than once in the relationship
cardinality is many to many. Let us assume that an employee can work on
multiple projects and each project can have multiple employees working on it.
So, the relationship will be many-to-many (M:N), meaning that one employee
may be associated with several projects, and one project may involve several
employees.
IT
BP
i,
an
many to many cardinality
w
Using Sets, it can be represented as:
as
H
rti
Aa
r.
D
In this example, student A1 is enrolled in B1,B2 and B3, and course B3 is taken
by A1, A2, and A3. Therefore, this represents a many-to-many relationship.
Participation Constraint
Total Participation: Each entity in the entity set must participate in the
IT
relationship. If each student must enroll in a course, the participation of students
will be total. Total participation is shown by a double line in the ER diagram.
BP
Partial Participation: The entity in the entity set may or may NOT participate in
the relationship. If some courses are not enrolled by any of the students, the
participation in the course will be partial.
i,
The diagram depicts the 'Enrolled in' relationship set with Student Entity set
an
having total participation and Course Entity set having partial participation.
w
as
H
rti
Aa
Identify Entities: The very first step is to identify all the Entities. Represent these
H
them and represent them accordingly using the Diamond shape. Ensure that
relationships are not directly connected to each other.
Aa
Add Attributes: Attach attributes to the entities by using ovals. Each entity can
have multiple attributes (such as name, age, etc.), which are connected to the
respective entity.
r.
D
Define Primary Keys: Assign primary keys to each entity. These are unique
identifiers that help distinguish each instance of the entity. Represent them with
underlined attributes.
Link- [Link]
IT
1️. Retail Store Management System
A retail store maintains data about customers, products, suppliers, and sales
BP
transactions.
Tasks:
i,
Identify all possible entities
1. Customer
H
● Customer_ID (PK)
● Name
rti
Phone
Aa
● Email
● Address
r.
2. Product
D
● Product_ID (PK)
● Product_Name
● Price
● Stock_Quantity
● Category
3. Supplier
● Supplier_ID (PK)
● Supplier_Name
● Contact_No
● Address
● GST_No
IT
4. Sales
BP
● Sales_ID (PK)
● Date
● Customer_ID (FK)
i,
● Total_Amount
● Payment_Mode
an
w
Relationships
as
Associative Entity:
rti
Sales_Details
● Sales_ID (FK)
Aa
● Product_ID (FK)
● Quantity
r.
● Price
D
Dimensional modeling
Dimensional modeling is widely used in data warehousing and analytics, where data is
often represented in terms of facts and dimensions. This technique simplifies complex
data by organizing it into a star or snowflake schema, which helps in efficient querying
and reporting.
IT
Example: Sales reporting
Imagine you need to analyze sales data. You would structure it as follows:
BP
● Fact table:
● Sales (stores transactional data, e.g., Sales_ID, Revenue,
Quantity_Sold)
i,
● Dimension tables:
an
● Time (e.g., Date, Month, Year)
● Product (e.g., Product_ID, Category, Brand)
w
● Customer (e.g., Customer_ID, Location, Segment)
as
H
In a star schema, the Sales fact table directly links to the dimension tables, allowing
analysts to efficiently generate reports such as total revenue per month or top-selling
products by category. Here’s how the schema looks like:
rti
Aa
r.
D
D
r.
Aa
rti
H
as
w
an
i,
BP
IT
Object-oriented modeling
Object-oriented modeling is used to represent complex systems, where data and the
functions that operate on it are encapsulated as objects. This technique is useful for
modeling applications with complex, interrelated data and behaviors – especially in
software engineering and programming.
IT
Suppose you're designing a library management system. You might define objects like:
BP
● Member (Name, Membership_ID, Checked_Out_Books)
● Librarian (Name, Employee_ID, Role)
i,
Each object includes both attributes (data fields) and methods (functions). For example,
an
a Book object might have a method .check_out() that updates the book’s status
when borrowed.
w
This approach is particularly beneficial in object-oriented programming (OOP)
languages like Java and Python, where data models can be directly mapped to classes
as
and objects.
H
rti
Aa
r.
D
IT
BP
i,
an
w
as
H
rti
Aa
Each data modeling technique aligns with different stages of database design, from
high-level planning to physical implementation. Here’s how they connect with the types
r.
IT
BP
i,
an
w
as
H
rti
Aa
r.
D