lOMoARcPSD|63293293
UNIT I DV
Computer Science SL (The International School Bangalore)
Scan to open on Studocu
Studocu is not sponsored or endorsed by any college or university
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
UNIT I Introduction : Introduction- context of data visualization- definition
methodology, visualization design objectives. Key factors-purpose, visualization
function and tone, visualization design options- data representation, data
presenation, seven stages of data visualization, widgets, data visualization tools
Introduction
Data visualization is the representation of data through use of
common graphics, such as charts, plots, infographics and even
animations.
Data visualization convert large and small data sets into visuals, which is
easy to understand and process for humans.
Data visualization tools provide accessible ways to understand outliers,
patterns, and trends in the data.
Data visualizations are common in your everyday life, but they always appear in
the form of graphs and charts
What makes Data Visualization Effective?
Effective data visualization are created by
communication, data science, and design collide.
Importance of Data Visualization
Data visualization is an easy and quick way to convey concepts universally. You
can experiment with a different outline by making a slight adjustment
o Data visualization can identify areas that need improvement or
modifications.
o Data visualization can clarify which factor influence customer
behavior.
o Data visualization helps you to understand which products to place
where.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
o Data visualization can predict sales volumes.
Why Use Data Visualization?
1. To make easier in understand and remember.
2. To discover unknown facts, outliers, and trends.
3. To visualize relationships and patterns quickly.
4. To ask a better question and make better decisions.
5. To competitive analyze.
6. To improve insights.
Visualization Design Objectives
Here ares some established guidelines followed during the design process. They come
from the book "Data Visualization: A Successful Design Process" by Andy Kirk.
1. Strive for form and function. One should strive to find a balance between
both. People are more tolerant of beautiful things that don't work well, but they
also get attached to things who are very useful, despite their appearance. The
WindMap animation we saw at the beginning is an example of something both
beatiful and very informative. However, if you are just starting in the field of
visualization, start with functionality first (build the house), and then enhance
its form (decorate the house).
2. Justify the selection of everthing you do. Just because we can do certain
things (add shadows, make 3-D bar charts, have fancy animations, etc.), that
doesn't mean we have to do it. The guiding principle should be: what does this
add to the visualization, is it necessary? Everything that you do should be
planned and have a reason. Challenge your decisions. Constantly ask yourself,
why am I using this color, this shape, this animation, and so on. The Literary
Organism visualization we saw above is an example where every visible
property is used to communicate data. Visualization should not be a platform to
showcase technical competence.
3. Creating accessibility through intuitive design. Intuitive design is very
difficult. Think of how difficult is sometime to figure out how to open a door
(push or pull). The same applies to designing visualization. If we are making
the users think very hard simply to understand how to interpret something, we
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
have failed. That doesn't mean that we should avoid novelty and only stick to
old and established means. The novelty is often something that will draw the
user in. But, make the user's effort worthwhile. Once they understand your
design, reward them with insights. Don't use complex visualization for
conveying simple information. The work of Edward Tufte gives a
comprehensive overview into accessible designs and how often, less is more.
4. Never deceive the receiver. You might have heard at some point in your life
the phrase Lies, damned lies, and statistics. It is possible to lie and deceive with
visualization too. Sometimes this might be unintentional, because of optical
illusion, but in other occasions it can be done in purpose. Especially our
political views and other life choices can bias our work. As an example,
consider this article that dissects visualizations about the politically divise
debate "pro-life" vs. "pro-choice". There is ethics to visualization, and it's
important to question not only our technical reasons for the design choices, but
our motivations too.
Key factors-purpose, visualization function and tone.
Key factors:
Audience. It’s important to adjust data representation to the specific target
audience. For example, fitness mobile app users who browse through their
progress would prefer easy-to-read uncomplicated visualizations on their
phones.
Content. The type of data you are dealing with will determine the
tactics. For example, if it’s time-series metrics, you will use line
charts to show the dynamics in many cases.
To show the relationship between two elements, scatter plots
are often used. In turn, bar charts work well for comparative
analysis.
Context. You can use different data visualization approaches and read data
depending on the context.
To emphasize a certain figure, for example, significant profit
growth, you can use the shades of one color on the chart and highlight the
highest value with the brightest one.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Dynamics. There are various types of data, and each type has a different rate
of change. For example, financial results can be measured monthly or yearly,
while time series and tracking data are changing constantly.
Purpose. The goal of data visualization affects the way it is implemented. In
order to make a complex analysis, visualizations are compiled into dynamic
and controllable dashboards equipped with different tools for visual data
analytics (comparison, formatting, filtering, etc.).
visualization function and tone
Data visualization is a powerful tool to communicate complex
information, reveal patterns and insights. Data visualizations
that match your purpose, tone, and audience.
Know your goal
Before you start creating your data visualization, you need to
have a clear goal in mind. What is the main message or story
you want to convey with your data? How do you want your
audience to react or act upon your data?
For example, if your goal is to inform your audience about a
trend or comparison, you might use a line chart or a bar
chart.
Know your audience
Another key factor to align your data visualization with your
writing style and audience engagement is to know your
audience.
For example, if your audience is experts or decision-makers,
you might use more technical terms, advanced charts, and
interactive features.
Know your tone
Your tone is the way you express your attitude and emotion
through your data visualization. Your tone can be formal or
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
informal, serious or humorous, objective or subjective,
depending on your goal and audience. Your tone will affect
your choice of words, colors, fonts, and images.
For example,If your tone is informal and humorous, you might
use bright or contrasting colors
Choose the right format
The format of your data visualization is the way you present
and organize your data in a visual form. There are many
types of formats, such as charts, graphs, maps, tables,
dashboards, infographics, and more.
For example,If you have a small amount of data and want to
show a simple comparison or distribution, you might use a
table or a chart.
visualization design options- data representation, data
presenation.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Unit 2 Data Literacy IX - notes
Artificial Intelligence Class 9 (Maharaja Agrasen Model School)
Scan to open on Studocu
Studocu is not sponsored or endorsed by any college or university
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Unit 2: Data Literacy and AI
2.1 - Basics of Data Literacy
2.1.1 Introduction to Data Literacy
Data literacy is about understanding, working with, and communicating data. It involves knowing how to
collect, analyze, and present data in meaningful ways. It enables informed decision-making and critical
thinking.
Why is Data Literacy Important?
Data literacy is important in everyday life as it helps people make decisions based on facts, think
critically, solve problems, and innovate. It also helps in recognizing data privacy and security concerns.
Key Concepts
1. Quantitative Data: Data that consists of numbers, like scores, measurements, or readings.
2. Qualitative Data: Data made up of words and phrases, like opinions or descriptions.
3. Data Privacy: Protecting sensitive information such as personal data.
4. Data Security: Safeguarding digital data from unauthorized access, corruption, or theft.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Data Pyramid
The Data Pyramid has four stages:
1. Data: Raw information, not very useful on its own.
2. Information: Processed data that gives insights.
3. Knowledge: Information that leads to understanding.
4. Wisdom: Understanding why things happen the way they do.
Let’s understand Data Pyramid with a simple Traffic Light example:
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
2.1.2 Impact of Data Literacy
Activity: Impact of News Articles (Select any trending news) [ To be pasted in AI practical file]
Session Preparation Logistics: For a class of 40 Students [Pair Activity]
Materials Required:
ITEM QUANTITY
Online Data Sources Clues NA
Computers 20
Purpose: The purpose of this activity is to engage participants in various scenarios that involve collecting
data and analyzing its sources. By understanding how authentic data sources contribute to reliable and
unbiased decision-making.
Brief: {Pair Activity} Participants will search the internet for data sources, extracting key information to
support their decisions.
Use the given template:
Author of the Weblink to the How was the Key figures/points
Source Source situation described in the source
by the Source
You have to rank the sources of the news articles from most accurate to least, state reasons for your
choice.
Rank Data Source Remarks
So, we can conclude that every data tells a story, but we must be careful before believing the story. Data
literacy is essential because it enables individuals to make informed decisions, think critically, solve
problems, and innovate.
2.1.3 How to Become Data Literate
Being data literate means understanding how to research, filter, and analyze data for everyday tasks like
online shopping or reading reviews. It involves using data to make informed decisions and recognizing the
differences between reliable and unreliable sources of information.
Scenario: Buying a Video game online
Data literacy helps people research about products while shopping over the Internet
How do you decide the following things when we are shopping online?
● Which is the cheapest product available?
● Which product is liked by the users the most?
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
● Does a particular product meet all the requirements?
A data literate person can –
● Filter the category as per the requirement – If the budget is low, select the price filter as low to
high
● Check the user ratings of the products
● Check for specific requirements in the product
2.1.4 What are Data Security and Privacy? How are they related to AI?
Data Privacy and Data Security are often used interchangeably but they are different from each other.
What is Data Privacy?
Data privacy referred to as information privacy is concerned with the proper handling of sensitive data
including personal data and other confidential data, such as certain financial data and intellectual property
data, to meet regulatory requirements as well as protecting the confidentiality and immutability of the data.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
The following best practices can help you ensure data privacy:
● Understanding what data, you have collected, how it is handled, and where it is stored.
● Necessary data required for a project should only be collected.
● User consent while data collection must be of utmost importance.
What is Data Security?
Data security is the practice of protecting digital information from unauthorized access, corruption, or
theft throughout its entire lifecycle.
Why is it important?
Due to the rising amount of data in the cloud there is an increased risk of cyber threats. The most
appropriate step for such an amount of traffic being generated is how we control and protect the transfer
of sensitive or personal information at every known place.
The most possible reasons why data security is more important now are:
● Cyber-attacks affect all the people
● The fast-technological changes will boom cyber attacks
2.1.5 Best Practices for Data Security
Data Security protects computers, servers, mobile devices, electronic systems, networks and data from harmful
attacks.
A few practices are as follows:
1. Use strong passwords with a mix of characters.
2. Enable Two-Factor Authentication (2FA) for extra security.
3. Download files only from trusted sources.
4. Always lock your screen when not in use.
5. Be cautious while sharing personal information online.
Activity - Exploring Data Online [ To be pasted in AI practical file]
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
1. Choose a product to research online.
2. Compare products based on price, ratings, and reviews.
3. Write down how the data helps you make a decision.
2.2 Acquiring Data, Processing, and Interpreting
2.2.1 Data Types of data
Artificial Intelligence is crucial, with data serving as its foundation. We come across different types of
information every day. Some common types of data include:
Textual Data (Qualitative Data) Numeric Data (Quantitative Data)
● It is made up of words and ● It is made up of numbers
phrases ● It is used for Statistical Data
● It is used for Natural ● Any measurements, readings, or
Language Processing (NLP) values would count as numeric
● Search queries on the internet data
are an example of textual data ● Example: Cricket Score,
● Example: “Which is a good Restaurant Bill
park nearby?”
Numeric Data is further classified as:
● Continuous data is numeric data that is continuous. E.g., height, weight, temperature, voltage
● Discrete data is numeric data that contains only whole numbers and cannot be fractional. E.g. the
number of students in the class – it can only be a whole number, not in decimals
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Types of Data used in three domains of AI:
2.2.2 Data Acquisition/Acquiring Data
Data Acquisition also known as Acquiring data , refers to the procedure of gathering data. It involves
searching of datasets suitable for training AI models.
The process typically comprises 3 key steps:
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Acquiring Data – Sample Data Discovery
● Let’s say we want to collect data for making a CV
model for a self-driving car
● We will require pictures of roads and the objects
on roads
● We can search and download this data from the
internet
● This process is called data discovery
Acquiring Data – Sample Data Augmentation
● Data augmentation means increasing the amount of
data by adding copies of existing data with small changes
● The image given here does not change, but we get
data on the image by changing different parameters like
color and brightness
● New data is added by slightly changing the existing data
Acquiring Data – Sample Data Generation
● Data generation refers to generating or recording
data using sensors
● Recording temperature readings of a building is an
example of data generation
● Recorded data is stored in a computer in a suitable
form
Sources of Data
Various Sources for Acquiring Data:
● Primary Data Sources — Some of the sources for primary data include surveys, interviews,
experiments, etc. The data generated from the experiment is an example of primary data.
● Secondary Data Sources—Secondary data collection obtains information from external sources,
rather than generating it personally. Some sources for secondary data collection include:
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
2.2.3 Best Practices for Acquiring Data
Checklist of factors that make data good or bad
Data acquisition from Websites:
The process of collecting data from websites using software is called Web Scraping.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Ethical concerns in data acquisition
While gathering data and choosing datasets, certain ethical issues can be addressed before they
occur.
2.2.4 Data Processing and Data Interpretation
Data processing and interpretation have become very important in today’s world
Can you answer this?
▪ Niki has 7 candies, and Ruchi has 4 candies
▪ How many candies do Niki and Ruchi have in total?
▪ We can answer this question using data processing
▪ Who should get more candies so that both Niki and Ruchi have an equal number of candies?
▪ How many candies should they get?
▪ We can answer this question using data interpretation
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Data Processing
▪ Data processing helps computers understand raw data.
▪ Use of computers to perform different operations on data is included under data processing.
Data Interpretation
▪ It is the process of making sense out of data that has been processed.
▪ The interpretation of data helps us answer critical questions using data.
Types of Data Interpretation
There are three ways in which data can be presented:
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Textual DI
▪ The data is mentioned in the text form, usually in a paragraph.
▪ Used when the data is not large and can be easily comprehended by reading.
▪ Textual presentation is not suitable for large data.
Tabular DI
▪ Data is represented systematically in the form of rows and columns.
▪ Title of the Table (Item of Expenditure) contains the description of the table content.
▪ Column Headings (Year; Salary; Fuel and Transport; Bonus; Interest on Loans; Taxes) contains
the description of information contained in columns.
Graphical DI
Bar Graphs
In a Bar Graph, data is represented using vertical and horizontal bar.
Pie Charts
▪ Pie Charts have the shape of a pie and each slice of the pie represents the portion of the entire
pie allocated to each category
▪ It is a circular chart divided into various sections (think of a cake cut into slices)
▪ Each section of the pie chart is proportional to the corresponding value
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Assignment
1. What is Data Literacy?
2. Differentiate between Data Privacy and Data Security.
In what ways are data privacy and data security different from each other?
3. Explain the Data Pyramid. Describe the four stages of the data pyramid and their
significance.
4. What is Quantitative Data? Give an example of quantitative data and explain how it is used.
5. What is Qualitative Data? Provide an example of qualitative data and its relevance in data
analysis.
6. List Two Best Practices for Data Security.
7. What is Data Processing? Explain the role of data processing in understanding raw data.
8. How does Data Interpretation help us? Describe how data interpretation helps us answer
critical questions.
9. What are the Types of Data Interpretation?
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Ch-2 Data literacy question and answers
Social and Cultural Anthropology HL (Saint Paul Higher Secondary School)
messages.pdf_cover_qr_code_label
messages.studocu_not_sponsored_or_endorsed_by_college
messages.downloaded_by
lOMoARcPSD|63293293
Chapter 2: Data Literacy
SUBJECTIVE TYPE QUESTIONS
Unsolved Questions:
1. Data literacy is the ability to read, understand, create and communicate data as
information. It is important in the modern world because:
(i) It helps us understand that different types of data may have good or poor
quality,i.e., the data may be reliable or unreliable.
(ii) It helps us understand that different types of data carry different values or
significance.
(iii) It makes it easier to understand how data is collected and presented.
2. Data refers to raw facts and figures that are unprocessed and don’t possess any
meaning, e.g., a list of numbers (1, 2, 3, 4, 5…). On the other hand, information is
processed data that has been organized, structured and presented in a meaningful
context, e.g., if we add ‘roll number’ as context to the list of numbers, then the data
becomes information.
3. The DIKW model is a pyramid-shaped model of knowledge management
whichrepresents the relationship between Data, Information, Knowledge and Wisdom.
(i) Datais the main component of the DIKW model and consists of raw facts,
figures, etc.
(ii) Information is processed data and is used to convey meaning within a specific
context.
(iii) Knowledge provides a deeper understanding that is gained through
interpreting the information.
(iv) Wisdom is the highest level of the DIKW model and helps in decision-
making.
4. Quantitative data refers to the data values that can be measured, counted and
compared with each other on the basis of quantity. For example, the age of a person,
the number of pages in a book, the mileage of a vehicle, etc.
On the other hand, qualitative data refers to the data values that can depict the quality,
properties or the distinguishing characteristics to be classified into one out of the two
or more categories. For example, the color of eyes, the name of the author of a book,
the type of vehicle, etc.
5. Discrete data consists of discrete values that are whole numbers and cannot be
subdivided further. For example,the number of students in a classroom, the number of
items sold in a store, etc. On the other hand, continuous data represents measurements
that can take any value within a range and can be infinitely subdivided into smaller
increments. For example, the height and weight of individuals, temperature readings,
etc.
11. Critical thinking helps in effective data literacy because it enables users to analyze
data objectively, identify biases and evaluate the validity of sources. It also helps in
interpreting data accurately, making informed decisions and solving problems effectively.
It further ensures that data is used responsibly and accurately, leading to better outcomes.
messages.downloaded_by
lOMoARcPSD|63293293
14. Describing data is the ability to understand and describe data, including understanding
concepts such as how data is categorized into data types and the way it is stored using
variables, datasets and data structures.
It helps in a better understanding of data by identifying the trends and outliers present
in the data. It also simplifies large datasets by summarizing them.
19. To protect sensitive information, some common data security measures that can be
taken are:
(i) Adding two-factor authentication to your accounts.
(ii) Usingstrong passwords with a combination of digits, letters and symbols.
(iii) Using an advanced encryption algorithm to encrypt data.
(iv) Applying firewalls to the outgoing and incoming network traffic.
(v) Ensuring the use of updated software.
(vi) Using the latest anti-virus software.
22. Government initiatives provide different rules, regulations and frameworks that help
in enhancing data privacy and security, thus protecting the personal data of individuals.
One such initiative is the Digital Personal Data Protection Act, 2023. This act emphasizes
obtaining informed consent from individuals before processing their personal data.
Through this, we have the right to know how data is being collected, used and stored.
Based on this information, we can choose to give or withhold our consent.
24. The following steps are involved in the process of data interpretation:
(i) Understanding the data
(ii) Cleaning and preprocessing the data
(iii) Exploring the data by applying exploratory data analysis
(iv) Visualization
(v) Pattern recognition
Data interpretation is important because it helps to transform raw data into meaningful
insights, which help in informed decision-making and planning.
messages.downloaded_by
lOMoARcPSD|63293293
Data Visualisation Techniques and Tools for Effective
Analysis
Enterprise Computing (Prairiewood High School)
Scan to open on Studocu
Studocu is not sponsored or endorsed by any college or university
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
The Purposes Of Data Visualisations
Data visualisations are tools used to transform raw data into a visual format,
enabling users to quickly and effectively understand complex information. By
presenting data graphically, such as through charts, graphs and infographics,
visualisations provide clarity and insights that might otherwise be difficult to
discern in text-based or numerical forms. Data visualisations are essential in
fields like business, science and education, where interpreting large datasets
accurately and efficiently can drive decision-making and innovation.
Simplify Understanding
A primary purposes of data visualisations is to simplify the interpretation of
complex datasets.
By translating data into visual formats like bar graphs, pie charts and heat
maps, visualisations allow users to identify patterns, relationships and trends,
reducing the cognitive load required to process data and makes insights
accessible to a broader audience.
Telling a Story
Data visualisations serve as a narrative tool, helping to tell a compelling story
about datasets.
By carefully selecting the visual format, such as a graph or infographic, and
sequencing the data presentation, creators can highlight the journey or
evolution of a dataset, revealing trends, changes and causations that can
influence viewers opinions.
Highlighting Results
Data visualisations should draw attention to significant results or key insights
within a dataset.
By using visual emphasis, such as colours, shapes or annotations,
visualisations can direct the audience’s focus to the most critical aspects of
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
the data. For instance, a pie chart illustrates larger segments to show areas of
success or concern.
Features Of Software Supporting Data
Visualisations
The features of software used for data visualisation play a crucial role in
enhancing the understanding of datasets. These tools provide functionalities
such as data aggregation, filtering and the ability to generate interactive
visualisations. They enable users to explore datasets from multiple
perspectives, uncover trends and identify patterns that may not be
immediately obvious in raw data.
Additionally, features like real-time updates, customisation options and
seamless integration with other tools make these software solutions
supportive in analysing and presenting data effectively.
Features of software supporting data visualisation include:
● Spreadsheets
● Creative Design Applications
● Combining Applications
Spreadsheets
The most widely used tools for data visualisation, offering features like charts,
pivot tables and conditional formatting. They allow users to quickly
summarise large datasets, identify trends and highlight key metrics. The
flexibility of formulas and data analysis tools in spreadsheets makes them
essential for building dashboards that provide clear insights.
Creative Design Applications
Applications such as Adobe Illustrator or Canva, provide advanced tools for
creating polished and visually appealing data visualisations. These tools are
ideal for tailoring infographics and presentations to a specific audience,
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
allowing for the inclusion of brand colours, logos and other design elements
that make data more engaging and accessible.
Combining Applications
The integration of multiple software applications enhances the ability to track
trends and forecast outcomes. Different applications can connect to diverse
data sources, enabling real-time analysis and the combination of different
datasets into unified visualisations. Essentially different applications can
conduct different processes on datasets, where data can be analysed and
exported between programs based on the strengths of specific applications.
Interpreting & Comparing Datasets
Analysing data involves interpreting and comparing datasets to uncover
trends, relationships and anomalies. Through techniques such as
aggregation, filtering and statistical analysis, users can transform raw data
into actionable insights. For example, line charts can reveal upward or
downward trends over time, while pie charts can can show values associated
with different demographics. This process is vital for predictive data analytics,
where historical data is used to forecast future outcomes.
Support Enterprise
Analysis can uncover business performance trends and opportunities for
growth. An analysis of sales data across regions can reveal underperforming
markets or highlight seasonal fluctuations.
Analysed patterns can inform inventory management, marketing strategies
and resource allocation. Predictive analytics, such as forecasting customer
demand, can help enterprises optimise supply chains and reduce costs.
Support Social Issues
Data visualisation can identify patterns that reveal societal trends or
disparities. For instance, comparing datasets on public health indicators
across different communities can highlight areas of inequity. Governments
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
can use these insights to allocate resources more effectively or develop
targeted interventions, while using predictive analytics to forecast potential
public health crises based on historical trends.
Support Ethical Issues
Interpreting data to address ethical issues can highlight areas of concern
such as data privacy, discrimination or bias. Predictive analytics can identify
potential ethical risks in new technologies, such as unintended biases in
machine learning algorithms, enabling enterprises to address them
proactively.
The Evolution Of Hardware Supporting
Data Analytics
The advancement of hardware and software has transformed the field of data
analytics, enabling the processing and interpretation of increasingly large and
complex datasets. Modern hardware innovations such as high-performance
processors have drastically reduced the time required to analyse data, while
sophisticated software tools provide more versatile platforms for visualising
and modelling data. This evolution has expanded the potential applications of
data analytics, from real-time decision-making to predictive modelling,
driving progress across industries.
Processing Power
The evolution of processing power has been a cornerstone of advancements
in data analytics. Innovations like multi-core CPUs, GPUs, OS and dedicated AI
accelerators have enabled the rapid processing of massive datasets and the
execution of complex algorithms.
Storage / Memory
Development of more efficient and scalable storage and memory solutions
has enhanced data analytics capabilities. Technologies such as solid-state
drives (SSDs), cloud storage and random-access memory (RAM) have enabled
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
the storage and retrieval of vast amounts of data at high speeds. OS also play
a role.
Communications Media
Improvements in communication media have facilitated the seamless
transfer of data across networks. High-speed internet, 5G networks, optic fibre
and communication protocols have enabled the real-time sharing of large
datasets among teams and systems, critical for applications like IoT analytics.
NOS also support transfer of data.
Types of Hardware: Processing Devices
The components of a system that assist in physically transforming data into
information.
Central Processing Unit (CPU)
Used for processing data, the CPU is the brain of any system. Fast processors
allow many rapid calculations.
Examples of CPU’s include Intel’s i series (i5, i7, i9), AMD Ryzen (5,7,9) and
Apples M series (M1, M2). CPU’s are made of multiple cores, which
independently support the processing of different tasks.
Graphics Processing Unit (GPU)
Specifically processes graphical task required by a system. This may relate to
applications presenting video data or displaying graphics in a video game. A
GPU’s processing power is also required when editing video media. Processing
can be upgraded with graphics cards.
Random Access Memory (RAM)
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Storage for data sent for processing. RAM assists in the processing of Data. It
does this though providing a storage location for data and instructions to be
sent prior to processing. The larger the capacity of RAM, the larger the file
sizes of data and the increased amount of data that can be stored for
processing.
The CPU works with RAM in order to process data through the fetch-execute
cycle. The CPU retrieves instructions from RAM, decodes it, processes it, then
stores the processed information back in RAM.
Types of Hardware: Storage Devices
Storage devices allow for data to be saved in a location on a system both
during use (Primary Storage), as well as so the data can be accessed again and
retrieved at a later date (Secondary Storage).
Magnetic Disks
Hard drives make use of a magnetic disk to store data. Data is stored on
within specific sectors of the magnetic disk, which are located on the disks
tracks. The hard disk drive contains a needle which reads data off the tracks
as the disk is rotated. Each sector of the disk can usually store up to 512 bytes
of data.
Optical Disks
The data on an optical disk is stored as small notches on the disc and is read
by a laser from an optical drive, which translates the notches read by the
laser in to usable data. Data stored on optical disks is also contained in tracks
and sectors, though with greater capacity. Types of disks include CD’s, DVD’s
and Blu-Ray.
Network Storage
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
It is common practice these days for storage mediums such as hard drives to
be contained on a network file server which multiple systems can access.
Essentially the user connects to these files through a private network or the
internet, accessing the systems file server (hard drives stored on servers) to
view data online.
Flash Memory
Contains a chipset that connects directly to a motherboard. Allows for
high-speed storage and retrieval of data. Flash memory is found in USB
storage devices, SD cards and solid-state drives (SSD). RAM, which is primary
storage, uses flash memory.
Types Of Hardware: Communication
Devices
Communication devices allow for data to be transmitted (sent) and
received between devices on a network. A variety of connectivity devices
and mediums are required to establish a communications infrastructure
for an enterprise.
Central Nodes
Devices used to connect multiple nodes (networked devices) together. A
switch has multiple communication channels, allow for data to be sent to
a specific node on a local network. Routers have wireless capabilities and
can connect devices to other external networks and the internet. Wireless
Access Points (WAP) may also be used to connect multiple users to a
network.
Communication Mediums
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Wired mediums used in order physically connect networked devices
together, such as twisted pair, coaxial and optic fibre cabling. Wireless
mediums are also incorporated to connect devices using different
wireless frequencies, such as microwave, satellite and radio waves.
Servers
Storage mediums such as hard drives being available through a network
to provide resources which multiple devices can access through a
client-server architecture.
Types of servers include:
● File servers for documents and media.
● Mail servers for email.
● Print servers for accessing printers.
● Web servers for hosting websites.
Other Technology
A modem modulates (MO) and demodulates (DEM) between analogue
and digital formats for transmission across different mediums. Mobile
devices can support communications through cellular, Wi-Fi and
Bluetooth connectivity.
Transmission Media
Transmission media refers to the pathways through which
communication signals are sent from one device to another. Each type of
media has distinct characteristics and is suited for different applications
based on factors like bandwidth, distance, cost and environmental
influences. Categories include wired and wireless media.
Wired Media
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Wired (physical) mediums include twisted pair, coaxial and optic fibre
cabling.
Twisted Pair
A cable made up of two copper wires that are individually insulated, then
twisted to form a spiral. The wires are twisted to reduce the amount of
interference from other cabling. Ethernet cabling used for networking is
made up of twisted pair wires.
Coaxial Cabling
Contains a single copper wire that is covered by an insulator, a grounded
copper mesh and an outside insulator. The layers shielding allow for data
packets to be transmitted with minimal distortion. When manipulating a
coaxial cable it can feel quite rigid due the copper mesh layer. Coaxial
cable is commonly used as arial cables which connect to antenna's on the
roof of a house. You probably have one behind you television.
Optic Fibre
Uses a laser of light to carry data in very thin glass fibres that are about
the diameter of a human hair, protected by an insulation layer. It is free
from electromagnetic and radio interference is very secure and can
transmit data for long distances at very high speed without errors. A
single cable contains multiple glass fibres, which are all capable of
carrying data. This one of the reasons optic fibre has such great
bandwidth!
Wireless Media
Wireless mediums include radio waves, microwave and infrared signals.
Radio Waves
Radio transmission uses radio waves that can be line of site or wide area.
Wifi networks operate using radio transmission, as well as technologies
such as RFID (radio frequency identification), its subset NFC (near field
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
communication) and Bluetooth. Devices need to include a wireless
adapter that will translate data through the radio signal.
Microwaves & Satellites
Microwaves are a high frequency radio signal that can send wireless
communications in line of sight through transponders which are
stationed 40 to 50 km in distance from each other, as well as to satellites
around the Earth. The devices receive a signal, amplify the signal, then
retransmit the signal to the next device. For this reason, transponders
need to be placed strategically and are usually setup on higher ground,
such as upon tall buildings. Satellites move with the rotation of the Earth.
Microwave is used by telephone networks, internet service providers
(ISP’s), global positioning systems (GPS).
Infrared
Infrared has a wavelength slightly greater than the red end of the visible
light spectrum but shorter than those of radio waves. Its frequencies are
higher than those of microwaves, but lower than those of visible light.
Infrared is used for short range, line of site data transmission, meaning
that the medium is unable to penetrate objects such as walls. Its
short-range data transmission means no security measures are applied to
the transmission in most cases.
Areas where Infrared is used include remote controls, security sensors
and fire sensors.
Online Analytical Processing (OLAP)
Online Analytical Processing (OLAP) is a technology used to organise
large datasets into multidimensional structures that facilitate complex
queries and data analysis. OLAP systems allow users to slice and dice
information across different dimensions, generating summaries or
detailed views. This capability enables faster and more flexible data
exploration, making OLAP essential for decision support systems and
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
data visualisation tasks. OLAP may be used as a data mining tool on
accumulated data stored in a data warehouse.
Using OLAP in Data Visualisations
In the context of data visualisations, OLAP can be used to dynamically
explore and present datasets, revealing patterns and trends across
multiple dimensions. For example, a business analyst might use OLAP to
visualise sales data, breaking it down by region, product line and time
period to identify seasonal trends or high-performing markets.
The ability to pivot and manipulate data interactively makes OLAP an
invaluable tool for creating dashboards and infographics that offer both
high-level overviews and detailed insights.
Advantages of OLAP in Data Visualisations
OLAP’s multidimensional approach allows users to analyse data from
various perspectives, uncovering hidden relationships that might not be
evident in flat data representations. OLAP also supports rapid querying
and aggregation, enabling real-time updates to visualisations as data is
adjusted or filtered. This responsiveness is beneficial in dynamic
environments where timely insights are crucial.
The hierarchical structure of OLAP data simplifies the creation of layered
visualisations, which help audiences intuitively navigate from summary
information to more detailed data.
Data Integrity In The Development
Of Data Visualisations
Data integrity refers to the accuracy, consistency and reliability of data
throughout its lifecycle, ensuring that it remains unaltered and
trustworthy. In the context of developing data visualisations, maintaining
data integrity is critical to producing meaningful and accurate
representations of the information. Poor data integrity can lead to
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
misleading visualisations, undermining decision-making processes and
eroding trust in the analysis. Ensuring data integrity involves addressing
key aspects such as ownership, source reliability, validation processes,
and risk management.
Ownership
Identifying and recognising the individuals or organisations responsible
for the data. Clear ownership ensures accountability for data accuracy and
ethical handling. Ownership also ensures that proper acknowledgements
are made, respecting IP and ICIP rights.
Source
Using credible and authentic sources ensure data’s legitimacy. Proper
documentation of data sources also allows users to trace the origin of the
data, which is essential for verifying findings and the credibility of data.
Validation
Cross-checking and verifying data to confirm its accuracy and relevance.
This step is crucial to prevent errors or inconsistencies from distorting
insights. Validating datasets across multiple sources ensures that the
visualisations reflect accurate and up-to-date information.
Risk
Maintaining data integrity involves identifying and addressing risks, such
as potential errors, data manipulation or breaches of security. Any
limitations should stated to an audience, support the transparency of
information.
Big Data & Data Warehousing
Big data and data warehousing are critical components in modern data
management and analysis. Big data refers to vast, complex datasets that
are too large for traditional data processing systems to handle. These
datasets are collected from a wide range of sources, including social
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
media, sensors and transactions, providing valuable insights when
properly analysed.
Data warehousing is a system used to store, manage and organise large
volumes of historical data from multiple sources to be archived in a
centralised location, allowing for efficient retrieval and analysis, though
processes such as data mining.
Together, big data and data warehousing enable organisations to make
data-driven decisions using larger scales of datasets. Below are areas that
relate to the characteristics of data when using these types of systems.
Volume
Refers to the sheer amount of data generated and stored.
Managing and storing enormous datasets requires scalable infrastructure
and efficient storage solutions. Processing large volumes is
resource-intensive. Companies like Facebook or Google process petabytes
of data daily, requiring vast storage systems and real-time analysis tools.
Variety
Refers to the different types and formats of data collected (e.g.,
structured, unstructured, or semi-structured data).
Handling diverse data sources, such as combinations of text, image and
video-based data, requires specialised tools for integration and analysis.
Retail businesses collect data from sales transactions, customer feedback
and inventory systems, all in different formats.
Velocity
Refers to the speed at which data is generated and needs to be
processed.
Rapid data streams require real-time processing technologies to derive
actionable insights before the data becomes outdated. Financial markets
require real-time processing of data to make split-second trading
decisions.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
The Impact Of Enterprise Data
Warehousing On Data Visualisation
Enterprise data warehousing consolidates large volumes of data from
multiple sources into a centralised location, enabling streamlined access
for analysis and visualisation. This centralisation supports data
consistency and reliability, which are critical for creating accurate and
insightful visualisations.
With data warehousing, organisations can efficiently handle vast
datasets, integrate historical and real-time data, and apply advanced
analytics. This significantly enhances the depth and scope of
visualisations, making them more informative and actionable for
decision-makers.
Analysis and Use of Historical Data Trends & Patterns
Data warehousing retains extensive records from various periods and
sources. By visualising this historical data, organisations can identify
long-term trends, seasonality and anomalies.
The ability to compare historical and current performance is invaluable for
strategic planning.
Data Refinement / Optimisation
Data warehouses ensure that data is cleaned, standardised and
structured before it is visualised, eliminating inconsistencies and
redundancies.
This refinement improves the clarity and accuracy of visual outputs,
making them more reliable for users.
Correlation with Current Data
Enterprise data warehousing enables the integration of historical data
with current datasets, facilitating comparative analysis and correlation.
This capability allows organisations to track ongoing performance relative
to past benchmarks and identify emerging trends.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
For example, in financial services, comparing current stock performance
with historical market trends through visual dashboards can uncover
correlations and inform investment decisions. Such integration enhances
real-time decision-making and predictive analytics.
How Big Data Affects The Design &
Development Of Data Visualisation
The sheer volume, variety and velocity of big data requires advanced tools
and techniques to create visualisations that are both meaningful and
accessible. Designers must prioritise clarity and focus to avoid
overwhelming users, while developers must implement scalable and
responsive systems capable of processing and presenting large datasets
in real time.
This evolution has driven innovation in visualisation techniques, enabling
organisations to extract actionable insights from ever-expanding
datasets.
Scope Of Visible Information
Big data significantly expands the scope of visible information available in
data visualisations. With access to vast datasets from diverse sources,
visualisations can offer a comprehensive view of complex systems,
relationships and patterns. A global supply chain dashboard can
aggregate and display data from multiple regions, tracking shipments,
inventory and demand in real time. However, this breadth of information
necessitates careful design choices, to ensure users can focus on the
most relevant data.
Types & Depth of Insight Provided by the Data
The diversity of big data allows for deeper and more nuanced insights in
visualisations. Advanced analytics, such as clustering, sentiment analysis
or predictive modelling, can uncover patterns that were previously
undetectable. The depth of insights afforded by big data enhances
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
decision-making and enables proactive responses to emerging
challenges.
Evaluating Bias When Developing
Data Visualisations
Bias in data visualisation arises when distortions occur during data
collection, storage or analysis, leading to misrepresentations of the
underlying information.
● In data collection, bias can emerge from sampling methods or the
exclusion of certain groups.
● During storage, inconsistencies in data organisation or omissions
can skew results.
● In the analysis phase, improper techniques or assumptions can
exacerbate these issues, influencing the narrative conveyed by the
visualisation.
Addressing bias at each stage is critical to ensure the integrity and
fairness of the final representation.
Accuracy
Bias can compromise the accuracy of a visualisation, leading to incorrect
conclusions. Inaccurate visualisations may arise from incomplete data,
errors in measurement or deliberate manipulation. This could include
excluding data points to simplify a graph will distort trends. Ensuring
accuracy requires rigorous validation of data and careful adherence to
ethical principles in designing the visualisation.
Audience
The intended audience plays a significant role in shaping how data is
visualised, and bias can arise if their needs and interpretations are not
considered. Visualisations must be tailored to the audience's level of
expertise and cultural context to avoid miscommunication. For example,
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
using overly technical jargon or culturally ambiguous symbols can
alienate or mislead viewers, reducing the effectiveness of the
visualisation.
Data Source
The choice of data source can introduce bias, particularly if the data lacks
diversity or represents a narrow perspective. Relying on a single or
unverified source may result in visualisations that overlook critical
viewpoints. Using data from a single demographic to represent a broader
population will lead to skewed conclusions. Ensuring a variety of credible
data sources can minimise this bias and provide a more balanced
representation.
Unconscious Bias
Bias can unintentionally occur from selecting datasets to designing visual
outputs. It can also stem from personal or institutional assumptions that
influence what is excluded or emphasised. For example, even selection of
particular colour schemes or labels can perpetuate stereotypes or create
unintended associations. Incorporating diverse perspectives can help
create fairer and more inclusive visualisations.
Software Tools Used To Develop Data
Visualisations
Software tools enable enterprises to transform raw data into meaningful
insights for decision-making. These tools range from spreadsheets for
basic charts and dashboards to advanced business analytics services that
provide real-time and predictive data analysis. By leveraging these tools,
enterprises can create dynamic, accessible and visually engaging
representations of complex data to drive strategic outcomes.
Spreadsheets
Allow users to transform raw data into meaningful insights through
charts, graphs, formulas and dashboards. Built-in functions enable users
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
to create dynamic visual representations that highlight trends, patterns
and relationships.
Presentations
Used for developing and delivering data visualisations in an engaging
format. Incorporates animations and interactive elements, transforming
complex data into visually compelling stories that enhance audience
understanding.
Business Analytics Services
Provide tools for developing data visualisations that transform raw data
into actionable insights. Offers tools such as real-time analytics,
predictive modelling and AI-driven insights specific to a business.
Custom Software
Provides tailored data visualisation capabilities, allowing an enterprise to
design visual tools that meet their specific needs. It enables complete
control over areas such as data integration, user interface design and
interactive features.
Types of Software: Spreadsheet
Spreadsheet software allows users to create and manipulate electronic
spreadsheets. In a spreadsheet application, each piece of data sits in a cell.
Each sheet is a table of data arranged into rows and columns, with rows being
identified numerically and columns referenced alphabetically.
Each piece of data can have a predefined relationship to data contained in
other cells established using cell references (e.g. A1, B6, CJ13). If you change
one value in a particular cell you more than likely see changes in other cells
to which the cell is being referenced.
Spreadsheet Tools
Spreadsheets contain the following tools:
● Formatting Tools: Edit cells colour, fonts and borders, as well as
conditional formatting.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
● Functions & Formulas: Program equations into cells that will
automate calculations.
● Graphical Tools: Generate graphs and charts that will illustrate
spreadsheet data, supporting analysis of data.
● Sorting Tools: re-order spreadsheet based on columns data.
Spreadsheets data can also be exported for use with other applications,
such as a sheets data being used as data source for a mail merge within
word processing Software. Examples of spreadsheet software includes
Microsoft Excel, Google Sheets and Apple’s Numbers.
Spreadsheet Tools Used To Develop
Data Visualisations
Spreadsheet software allows users to transform raw data into meaningful
insights through charts, graphs and dashboards. With built-in functions
for organising, analysing and formatting data, spreadsheets enable users
to create dynamic visual representations that highlight trends, patterns,
and relationships. Whether for business analytics, financial reporting or
academic research, spreadsheet software provides an accessible and
flexible platform for data-driven decision-making.
Spreadsheet Tools
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Spreadsheet tools used to develop data visualisations include:
● Charts & Graphs: Line charts, bar charts, pie charts and other charts
for representing data visually.
● Pivot Tables & Pivot Charts: Summarising datasets and creating
interactive charts to explore trends.
● Conditional Formatting: Highlighting specific cells or ranges using
colours based on defined criteria.
● Slicers & Filters: Enabling users to interactively filter data within
tables and charts for focused analysis.
● Data Linking Across Sheets: Creating visualisations that dynamically
update by linking data from multiple sheets or workbooks.
● What-If Analysis Tools: Help explore potential outcomes and
visualise impacts on datasets.
● Form Controls: Dropdown menus, sliders, and buttons to allow users
to interact with data visualisations.
● Predictive Analysis: Adding trendlines to charts for analysing
patterns and projecting future values.
● Data Validation Tools: Ensuring accurate data input for creating
reliable visualisations.
● Templates: Using pre-designed templates for specific layouts or
dashboards.
Spreadsheet Tools In Action
The following dashboard was made using spreadsheet software, relating
to car sales made by a business. Data within the spreadsheet relates to
the amount of sales of each car made, as well as data relating to the
manufacturers of each car. The dashboard makes use data that has been
linked from other worksheets contained within the spreadsheet file.
Spreadsheet Tools
The following spreadsheet tools were used to develop this data
visualisations:
● Charts & Graphs: Line charts and column charts for representing
data visually.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
● Pivot Tables & Pivot Charts: Summarising datasets and creating
interactive charts to explore trends in car sales.
● Conditional Formatting: Highlighting specific cells or ranges using
colours based on defined criteria, which in this dashboards context,
was the top and bottom 10% of sales over the years.
● Slicers & Filters: Enabling users to interactively filter data within
tables and charts for focused analysis. In this dashboard slicers
enabled us to focus on specific manufacturers cars and the sales,
we made on them, allowing us to look at different manufacturers
against each other.
● Data Linking Across Sheets: Creating visualisations that dynamically
update by linking data from multiple sheets or workbooks. This
dashboard was created using
Types of Software: Presentation
Presentation software is used to create multimedia that is intended to be
displayed to an audience and / or support a speaker. While presentations may
be viewed in many ways, quite often they are used in conjunction with a
projector in order to enlarge a presentation for an audience to see. For this
reason, presentation software must provide tools that make presentations
colourful, engaging and supportive of concepts a speaker is presenting about.
Presentation Software Tools
Presentation software has the following characteristics:
● Slides: Use master slides to create templates for an entire
presentation, including colour schemes, headings and layouts.
● Print Options: Making paper-based copies of slides, which may
include multiple slides per sheet, supporting a presentation.
● Animations: The ability to make objects move within slides, in order
to catch the attention of the audience. Slide transitions too!
● Use of Media: Insert and display text, images, audio, video and
animations, as well as embed objects such as charts and tables.
● Narrations and timings can also be applied to slides.
Presentation Software Tools
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Presentation tools used to develop and display data visualisations include:
● Embedded Charts & Graphs: Importing charts from spreadsheets or
creating custom visuals directly within software.
● Slide Templates: Pre-designed, professional templates with layouts
optimised for data presentations.
● SmartArt & Infographics: Creating visually appealing
representations of data (e.g. flowcharts).
● Animations & Transitions: Adding motion to slides to sequentially
reveal data or trends for audience engagement.
● Text & Image Overlays: Adding descriptive annotations to
complement and clarify data visualisations.
● Interactive Elements: Hyperlinking to specific slides or external data
for deeper dives into analysis. Embedding videos or animations.
● Data Tables: Present raw or summarised data alongside charts.
● Exporting Options: Converting presentations to formats like PDF or
video to distribute the analysis effectively to varied audiences.
● Speaker Notes: Adding contextual details or talking points for the
presenter to guide the explanation of data visualisations.
Business Analytics Services Tools
Business Analytics Services that can support the development of data
visualisations that support enterprise decision making include:
● Interactive Dashboards: Real-time, customisable dashboards that
provide an overview of KPI’s and metrics.
● Data Integration & Connectivity: Importing and integrating data
from multiple sources, such as databases, cloud services and APIs.
● Advanced Chart Types: A wide range of visualisation options,
including heat maps, waterfall charts and geographic maps.
● Predictive Analytics Tools: Built-in features to forecast trends and
simulate future scenarios using ML algorithms.
● AI and Natural Language Querying: Using artificial intelligence to
generate visualisations automatically or answer queries in plain
language.
● Drill-Down Capabilities: Allowing users to interactively explore
detailed data behind high-level visualisations.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
● Data Modelling & Transformation: Tools to structure and manipulate
datasets before visualisation.
● Collaboration & Sharing Features: Sharing resources with team
members or clients via cloud platforms.
● Exporting & Embedding Options: Exporting visualisations to various
formats or embedding them in other tools.
● Security & Permissions Management: Configuring access levels and
protecting sensitive data.
IaaS
Infrastructure as a Service (IaaS) provides virtualised computing resources
over the internet. It's the most basic category of cloud services, offering a
virtualised hardware platform.
● Components: This includes virtual server space, network
connections, bandwidth, IP addresses and load balancers.
● Usage: Enterprises use IaaS for temporary, experimental or
unexpected workload increases. It offers significant scalability
options for businesses looking to outsource their infrastructure.
SaaS
Software as a Service (SaaS) delivers software applications over the
internet, on a subscription basis. It allows users to connect to and use
cloud-based apps over the Internet.
● Components: Applications, ranging from office software to unified
communications among a wide range of other business apps that
are delivered as a service.
● Usage: SaaS is used by end-users who want to run and operate
software without the complexity of installation, maintenance and
scaling. It’s beneficial for applications that require web or mobile
access.
PaaS
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Platform as a Service (PaaS) provides a platform allowing customers to
develop, run and manage applications without the complexity of building
and maintaining the infrastructure typically associated with developing
and launching an app.
● Components: Development tools, database management systems,
business analytics services, and more, all hosted on the cloud
provider’s infrastructure.
● Usage: PaaS is used by developers who want to build applications
without spending time on configuring servers, networks, storage
and database management systems. It’s ideal for rapid development
and deployment of applications.
Business Analytics Services that support users through YouTube include:
● Interactive Dashboards: Real-time, customisable dashboards that
provide an overview of creators content (video) and metrics.
● Charts: A wide range of visualisation options, showing trends and
patterns in video view and playlist data, as well as global
demographic data.
● Predictive Analytics Tools: Built-in innovation tools to analyse trends
in viewership and channel views to suggest (inspire) new videos a
content creator could make using ML algorithms.
● Drill-Down Capabilities: Allowing users to interactively explore data
visualisations and go into more detail about viewership data and
trends.
● Collaboration & Sharing Features: Sharing resources with team
members or clients via cloud platforms.
Advantages Of Developing Custom Software
An enterprises decision to invest in custom software can support their
development of data visualisations in the following ways:
● Tailored Features & Functionality: Designed to meet specific
enterprise needs, incorporating unique features that off-the-shelf
tools may lack.
● Scalability: Custom solutions can be developed to grow with an
enterprise.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
● Integration with Proprietary Systems: Ability to integrate with
in-house systems or specialised software.
● Branding & Custom Design: Allows full control over the aesthetics of
data visualisations and enterprise branding.
● Optimised Performance: Built for specific data processing
requirements, ensuring high performance and responsiveness even
for large or complex datasets.
● Enhanced Security: Prioritise custom security features like
encryption and role-based access controls.
● Cost-Effectiveness Over Time: While initially a higher initial
investment, it eliminates recurring licensing costs.
● Flexibility & Control: Complete control over the software, avoiding
reliance on third-party providers.
● Support for Niche Use Cases: Addresses specialised requirements
that standard tools may not support.
● Enhanced User Experience: User interfaces and workflows can be
designed based on the preferences and expertise of the intended
users.
Interrogate Data From A Data
Visualisation
Interrogating data from a visualisation involves critically analysing the
information presented to uncover patterns, trends or anomalies and find
meaningful insights. This process ensures that the visualisation is not just
viewed at face value but thoroughly examined for its accuracy, relevance
and implications. Interrogation allows decision-makers to validate
findings, look deeper into data and use the visualisation as a tool to guide
informed actions.
Data may be interrogated from a data visualisation using the following
tools and techniques:
● Interpreting What You See: Understanding the relationships,
patterns or trends depicted. This involves identifying key features
such as peaks, dips or consistent patterns in the visualisation and
relating them to the underlying dataset.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
● The Effect of Outliers: Outliers are data points that deviate
significantly from the rest of the dataset. Identifying and examining
these outliers is a critical step in interrogating data, as they can
represent errors, anomalies or unique insights.
● Aggregation: The process of summarising data to make it more
comprehensible and actionable. In visualisations, this might involve
combining data points into groups, such as summarising sales by
region or time period.
● Filtering: Allows users to refine the scope of the data being
visualised by focusing on specific fields or criteria. This process
helps narrow the analysis to areas of interest, enabling users to
explore specific questions or scenarios without the distraction of
irrelevant data.
● Reasoning: Drawing logical conclusions from a data visualisation by
synthesising the observed trends, patterns and anomalies with
external knowledge or context. Requires critical thinking to
understand causation, impacts and potential actions.
Principles of engaging UX design include:
● User-Centred Design: A system should prioritise end-users
throughout the design process, to be tailored to a specific target
audience, as well as providing a specific feel.
● Navigation: Systems are easy to use, learn and navigate, reducing
the learning curve and making the user experience intuitive.
● Consistency: Maintaining uniformity in visual elements,
functionalities and terminology across a system to prevent
confusion and support user navigation.
● Simplicity: Using basic design to avoid overwhelming users with
unnecessary information or actions, making system interaction
straightforward and efficient.
● Interactivity: How users engage with a system through peripheral
devices. Systems also need to respond accordingly to user input,
providing a satisfying and engaging experience to keep user’s
attention.
● Contextual Design: The scenario in which the system will be used,
including the environment, device or user’s specific needs, to
design solutions that fit into the user’s situation seamlessly.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
● Accessibility: Systems are accessible to people with a wide range of
abilities, ensuring that everyone can access the system regardless
of their physical or cognitive abilities.
● Feedback / Response Time: Providing immediate and clear feedback
for actions to keep users informed about the system’s
responsiveness to their inputs.
Graphic design tools can support the development of data visualisations
in the following areas:
● Colour Palettes: Designing consistent and audience-appropriate
colour schemes to highlight key data points and trends.
● Typography Options: Customising fonts, text alignment and styles
for annotations, labels and headings to improve readability.
● Iconography: Incorporating high-quality icons, shapes and
illustrations to enhance visual storytelling.
● Integration Tools: Importing data visualisations from spreadsheet or
analytics software for further refinement.
● Exporting Visualisations: Supporting high-resolution exports for use
in presentations, reports or digital platforms.
● Animation Tools: Adding motion to visual elements to highlight
changes or tell a sequential story.
● Reusable Templates: Creating reusable templates for consistent
visualisation designs across different projects or presentations.
User Experience (UX) Influencing The
Development Of Data Visualisations
User experience (UX) plays a crucial role in the development of effective
data visualisations by ensuring that they are intuitive, engaging and
tailored to the needs of the audience. A well-designed visualisation
considers factors such as clarity, accessibility and interactivity to provide
users with a supportive experience in interpreting data. By prioritising UX
principles, developers can create visualisations that not only convey
insights but also empower users to explore and interact with data,
enhancing their decision-making processes.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Relevance to the Audience
Addressing the specific needs and interests of the target audience,
employees or the general public. Ensuring relevance fosters engagement
and supports the visualisation.
Audience Interpretation
Supporting interpretation by the intended audience. This involves
selecting the most appropriate visual formats, such as line graphs or pie
charts, with annotations to help ensure understanding.
Customisation
Allowing users to tailor visualisations to their specific preferences or
requirements. Features such as interactive filters, adjustable chart
parameters, and personalised dashboards enable users to focus on the
most relevant aspects of the data.
Live Analysis
Provide real-time updates and dynamic interactions within a data
visualisation. These features allow users to monitor changing trends or
conditions as they occur, enhancing their ability to respond quickly.
Criteria For Evaluating User
Experience (UX)
Developing a clear and structured criteria for evaluating the effectiveness
of user experiences in data visualisations is essential to ensure that visual
representations of data are accessible, engaging and meaningful to their
intended audience. An evaluation criteria helps assess how effectively a
visualisation communicates insights, supports decision-making, and
enhances user interaction. By considering specific factors, enterprises can
refine their visualisation strategies to improve comprehension and
engagement.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Key areas that would be a part of a criteria for evaluating effective user
experiences include:
● Usability: Ease of navigation, interaction and comprehension of
data.
● Clarity: A visualisation clearly communicating trends, patterns and
insights.
● Relevance: Data aligning with an audience’s needs and
expectations.
● Customisation: Inclusion of interactive features like filters,
drill-down options, and user preferences.
● Accessibility: Compatibility with assistive technologies and
adherence to UX best practices.
● Visual Appeal: Effectiveness of design elements such as colour
schemes, layout and contrast.
● Performance & Responsiveness: The speed and efficiency of updates
and interactions.
● Contextual Support: Labels and explanatory elements to aid
understanding.
The Impact Of Emerging
Technologies
Emerging hardware and software technologies are reshaping user
interface (UI) and user experience (UX) design by introducing new levels
of interactivity, efficiency and personalisation. Advances in technology
enable more intuitive and adaptive interfaces that enhance user
engagement. Modern UI/UX development is increasingly driven by
automation, real-time data processing and cross-platform accessibility,
allowing for more seamless and responsive digital experiences.
Examples of Emerging Technologies
Examples of emerging technologies impacting user experience include:
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
● Artificial Intelligence (AI) & Machine Learning: Enables predictive
interfaces, personalisation and automated design adjustments
based on user behaviour.
● Augmented Reality (AR) & Virtual Reality (VR): Enhances immersive
experiences, for users to interact with data in both virtual and
real-world environments.
● Gesture Recognition & Haptic Feedback: Provides touchless control
and enhances user interaction with sensory responses.
● Cloud Computing: Improves real-time UI responsiveness and
enables more powerful web-based applications.
● Wearable & IoT Devices: Expands UI/UX design for smart device
interaction in both traditional and new contexts, supporting
automation and decision-making.
● Blockchain for Secure UX: Strengthens trust in digital transactions
and identity verification through decentralised interfaces.
Data Appropriate For A Visualisation
Creating data visualisations involves transforming raw data into graphical
representations that communicate insights effectively. This process
requires careful planning, from researching relevant datasets to selecting
the appropriate visual format. A well-designed data visualisation
enhances comprehension, highlights trends and supports data-driven
decision-making.
Key steps in preparing data for visualisation include research, souring
data, organising data and storing data.
Research
Identify the purpose of the visualisation and gather relevant datasets
from credible sources. Understanding the target audience and key
insights to be conveyed to determine the most effective way to structure
visualisations.
Source Data
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Collect structured or unstructured data from databases, APIs, surveys or
public sources. Ensuring data accuracy and reliability at this stage is
crucial, as poor-quality data can lead to misleading visual representations.
Organise Data
Clean, filter and format the data to ensure consistency and remove errors
or redundancies. Data transformation techniques help prepare the
dataset for meaningful analysis and visual representation.
Store Data
Use spreadsheets, cloud databases or data warehouses, to maintain
accessibility and security. Storage management ensures that data
remains up-to-date and can be efficiently retrieved for use.
Data Visualisation Tools
(Spreadsheets)
Spreadsheets are essential tools for managing and summarising large
datasets, enabling users to extract key insights and identify trends
efficiently. By leveraging built-in features, individuals and businesses can
streamline data analysis, transforming raw numbers into actionable
information. Whether through formulas, visual tools, or data organisation
techniques, spreadsheets enhance clarity and improve decision-making.
Below are some of the most useful tools for summarising data effectively.
Functions
Formulas play a crucial role in automating calculations and summarising
key aspects of a dataset. Commonly used functions include:
● SUM: Adds up numerical values within a selected range to provide
totals.
● AVERAGE (MEAN): Calculates the mean value, offering insights into
central trends.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
● MAXIMUM (MAX): Identifies the highest numerical value in a
selected range of data.
● MINIMUM (MIN): Finds the lowest numerical value within a given
dataset.
● COUNT: Determines how many numerical values exist in a selected
range, ignoring blank cells and text.
● ABSOLUTE VALUE (ABS): Returns the positive equivalent of any
given number, removing negative signs.
● SQUARE ROOT (SQRT): Calculates the square root of a given
number.
● INTEGER (INT): Rounds a decimal number down to the nearest
whole number.
● PART (MOD): Returns the remainder when one number is divided by
another (useful for cyclic calculations).
● STANDARD DEVIATION (STDEV): Measures the amount of variation
or dispersion in a dataset, indicating how spread out the values are.
● IF: Applies logical conditions to return specific values based on
predefined criteria.
● LOOKUP: Searches for specific values within a dataset and retrieves
related information, simplifying data retrieval.
These functions help users analyse and interpret numerical data
efficiently, reducing manual work while improving accuracy.
Conditional Formatting
Conditional formatting allows users to highlight key trends, patterns, or
outliers based on custom rules. By applying different colours, icon sets, or
data bars, spreadsheets visually emphasise important figures, such as low
inventory levels, high-performing sales, or overdue tasks. This feature
improves data readability and enables quick identification of critical
insights.
Sorting & Filtering
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Sorting and filtering functions enhance data organisation, making it
easier to locate and focus on relevant information.
● Sorting arranges data in ascending or descending order based on
numerical values, dates, or text.
● Filtering allows users to display only specific data that meets
selected criteria, temporarily hiding irrelevant rows.
These tools improve efficiency by enabling users to isolate specific trends
without altering the original dataset.
Charts & Graphs
Visual representations of data help make complex information easier to
interpret.
● Bar Chart: Displays categorical data using rectangular bars, making
it easy to compare different groups or categories.
● Column Chart: Similar to a bar chart but with vertical bars, often
used to show trends or comparisons over time.
● Line Graph: Uses data points connected by lines to show trends over
time, ideal for tracking growth, sales, or performance.
● Pie Chart: Represents data as proportional slices of a circle, best for
showing percentage breakdowns of a whole.
● Scatter Plot: Uses points to display relationships between two
numerical variables, helping identify correlations or trends.
● Histogram: Displays the distribution of numerical data by grouping
values into ranges (bins), useful for frequency analysis.
● Area Chart: Similar to a line graph but with the area beneath the line
filled in, often used to show cumulative trends.
● Bubble Chart: A variation of a scatter plot where data points are
represented as circles, with size indicating an additional variable.
● Radar Chart (Spider Chart): Displays multivariable data on a circular
graph, useful for comparing different attributes across multiple
categories.
● Waterfall Chart: Helps visualize cumulative changes in a dataset,
often used for financial statements or progressive data analysis.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
By converting raw figures into visual summaries, charts and graphs make
data-driven insights more accessible to decision-makers.
Data Consolidation
Data consolidation allows users to merge information from multiple
sources into a central worksheet for a more comprehensive analysis. This
tool is particularly useful for businesses combining sales reports, financial
records, or departmental data into one structured format, ensuring a
holistic view of performance metrics.
Pivot Tables & Charts
Pivot tables and pivot charts are powerful spreadsheet tools used to
summarise, analyse, and visualise large datasets efficiently. Pivot tables
allow users to dynamically reorganise and filter data without modifying
the original dataset, making it easier to identify trends and patterns. Pivot
charts complement pivot tables by presenting summarised data in a
visual format, such as bar charts or pie charts, for clearer interpretation.
These tools are widely used in business and data analysis to simplify
complex information and support informed decision-making.
To specify:
● Pivot Tables allow users to drag and drop fields to quickly
reorganise and analyse data without modifying the original dataset.
● Pivot Charts visually represent pivot table summaries, making
trends and relationships clearer.
These tools are particularly valuable in enterprise settings, where rapid
data analysis is essential for informed decision-making.
Charts & Graphs
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Sometimes data is more meaningful when it is displayed graphically.
Quite often people’s minds can comprehend information on a graph
easier than a bunch of accurate numbers. Charts and graphs are used to
summarise information generated on a spreadsheet, so that it is easier for
the user or a target audience to understand.
There are a variety of graphs and charts available in spreadsheet software.
It is important to select the right type of graph to suit the data being
displayed.
Types of Charts & Graphs:
Column / Bar
This is a type of chart, which contains labelled horizontal (Column) or
vertical (Bar) bars displaying information on an axis. The numbers along
the side of bar graph compose the axis.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
These graphs are useful when there is a numerical comparison between
data series, such as student scores in a test. The choice whether to use a
column or a bar graph may relate to your page orientation, Column being
better for portrait orientation and Bar better for landscape orientation.
Line
A line graph is a way of representing changes in one or more series of
data over time, with each line representing a different data series. This is
useful when comparisons are needed between different data series, as
the lines intersect and overlap each other.
Pie
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
A chart that displays data as subsections of a whole quantity, which is
usually represented as a circle. The pie is split into different sized parts
based on the values of data being depicted. The different pieces of the pie
in the chart are called sectors.
A pie chart would be useful in representing productivity within different
departments in a company, or statistics on different demographics in
society. A Donut Chart is similar to a pie chart, it just has a hole in the
middle.
Area
A type of graph used to display how data may change with respect to
time. An area chart is similar to a line chart, though the different data
series usually completely overlap each other in order to distinguish the
series’ in order of greatest values based on data values.
Radar
A type of chart that displays different data values of a data series relative
to a centre point. Radar charts can be displayed with markers for each
data point. The greater the data value, the further it will be from the
centre point of the chart.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Tables, Filters & Slicers
Tables, filters, and slicers are essential spreadsheet tools that help users
organise, analyse and interact with data efficiently. Whether summarising
information for reporting or refining data for analysis, these tools improve
accessibility and usability, ensuring that only relevant information is
displayed to a user when needed.
Tables
Tables transform raw datasets into structured, dynamic ranges that make
data easier to manage and analyse. When a dataset is converted into a
table, features like automatic formatting, column sorting, filtering options
and structured references become available. Tables also allow for easy
expansion, meaning new data entries are automatically included in
calculations and formatting, streamlining data organisation.
Filters
Filters enable users to display only specific data that meets certain
criteria while temporarily hiding irrelevant information. Users can filter by
text, numbers or date ranges, making it easier to focus on relevant details
within large datasets. This is particularly useful in business settings where
users need to extract specific information, such as all sales above a
certain value or products from a specific category.
Slicers
Slicers provide an interactive way to filter data visually, making it easier to
refine and analyse datasets. Unlike traditional filters, slicers present
buttons that users can click to filter tables or pivot tables without
navigating dropdown menus. They are especially useful when working
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
with large reports, dashboards or complex data models, offering a more
user-friendly and visually appealing way to control data views.
Slicers may also be used in conjunction with Pivot Tables and Pivot
Charts.
Pivot Tables
Pivot tables are a spreadsheet tool used to summarise and analyse datasets efficiently.
The tables allow users to quickly extract meaningful insights from a dataset by
reorganising and filtering data without modifying the original data source. providing a
structured summary of information,
A pivot table enables users to dynamically group, filter and summarise large datasets in
a structured format. By simply dragging and dropping fields, users can calculate totals,
averages, counts, and percentages, providing a clear and interactive way to explore data.
Pivot tables make it easy to identify trends and patterns, such as total sales by region,
best-selling products, or performance over time. They are especially valuable for
businesses and organisations that need to process large amounts of information
efficiently.
By combining pivot tables and pivot charts, users can effectively process, summarise
and visualise data, transforming complex datasets into valuable insights for better
decision-making.
Pivot Charts
A pivot chart is a visual representation of a pivot table, displaying summarised data in an
easy-to-read format such as a bar chart, line graph, or pie chart.
These charts automatically update when changes are made to the pivot table, allowing
users to interactively analyse data. Pivot charts help highlight key trends and
comparisons, making them an essential tool for business reporting, financial analysis
and performance tracking.
Developing Data Visualisations
Designing and developing a data visualisation for a specific scenario
requires careful planning to ensure that the representation effectively
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
communicates key insights. This process involves selecting appropriate
visual elements, structuring data logically, and tailoring the design to
meet the needs of the intended audience. An effective data visualisation
transforms raw data into meaningful and actionable information.
Trends, Patterns & Relationships
A visualisation should highlights trends, patterns and relationships within
a dataset, making complex information easier to interpret. Charts and
graphs can be used to showcase changes over time, correlations between
variables or anomalies in data. By structuring data in a visually intuitive
way, users can quickly identify key insights and make informed decisions.
Predictive Analysis Incorporating Big Data
Predictive analysis forecasts future outcomes and trends to support
strategic planning. Forecasting charts and AI-driven dashboards help
illustrate predictive analytics in an intuitive format. A company could use
predictive visualisations to anticipate demand fluctuations based on
historical sales and external factors like economic trends.
Incorporating big data and data warehoused data allow for deeper
insights, as vast amounts of structured and unstructured data can be
processed and visualised dynamically to support decision-making.
Data Security
Implementing a specific combination of effective security protocols within an enterprise
is essential. Enterprises must conduct comprehensive risk assessments to identify
vulnerabilities within their systems and prioritise areas for immediate attention. Based on
cyber risk assessment data, an organisation can develop and implement a layered
security strategy that includes both preventive and reactive measures.
Security measures might include the use of:
● Login Procedures – Passwords, Biometrics and MFA
● Establish different access levels and user permissions.
● Firewalls to filter incoming and outgoing data for a network
● Encryption to protect data stored on drives, as well as in transit.
● Intrusion detection systems to identify unauthorised access.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
● Regular software updates.
● Data storage and backup procedures to protect against data loss.
Employee training is also required, as human error often leads to security breaches.
Workers must understand that maintaining data integrity and confidentiality is essential!
Furthermore, businesses should have an incident response plan in place that outlines
specific steps to be taken in the event of a security breach. Compliance with relevant
laws and regulations must also be implemented.
Strategies Used to Protect Data Part 1:
Password Protection, Biometrics,
Multi-Factor Authentication (MFA) &
Permissions
A combination of security strategies are required in order to protect data
stored on systems. In this section we will look at strategies that protect
data from being access by unauthorised users on a system or network.
Password Protection
A password is a secret word or string of characters made up of a
combination of letters, symbols and numbers. Passwords are entered on a
login screen along with a username to authenticate a user accessing a
system.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Elements related to passwords include strength, complexity, change
frequency and reuse.
Biometrics
A system may require the entry of biometric data for users to gain access.
This is data that is obtained from an individual’s biology, such as their
thumbprint being scanned or the use of facial recognition software.
Biometric elements are obviously very specific to individual users, making
them extremely hard to replicate.
Multi-Factor Authentication (MFA)
MFA requires uses to authenticate themselves, using two (2FA) or more
(MFA) sources of information.
This process may include a user entering in their login details to access a
website as the first method of authentication, then either being sent an
email, or SMS message that includes a pin to be entered as the next form
of authentication.
Permissions
Permissions need to be established in order to determine which users of
an organisation have the right to view specific information.
Specific permissions may be assigned to authenticated users to
determine what they can actually do with the information, such as Edit,
Read-Only, Comment or Share.
Strategies Used to Protect Data Part 2:
Encryption, Firewalls, Antivirus &
Anti-malware
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
A combination of security strategies are required in order to protect data
stored on systems. In this section we will look at data protection tools
that focus on data entering and exiting a system or network.
Encryption
This process scrambles data using a public or private encryption key, so
that if it is read by an unauthorised user during transmission or storage, it
will appear as incoherent symbols.
The same encryption key is then required for decrypting the data by an
authorised user, so that the message can be put back into legible form to
comprehend and read.
Firewall
Software used to protect data stored on a network. This is achieved
through implementing and using a series of predetermined security rules.
Packets of data can be assessed both entering and exiting a network
against existing rules, helping determine whether a specific packet or
source may be deemed malicious.
Anti-Malware & Antivirus
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Malware includes any type of software that may be deemed dangerous or
malicious to a system or network. Anti-malware and virus detection
(Antivirus) software are used to scan systems for known malicious
software based on known virus signatures.
A virus signature is a sequence of code that is known to be dangerous or
harmful when executed on a system.
Intrusion detection software detects when an unauthorised user or bot is
trying or has gained access to a secure network. The software provides
tools for administrators to be notified and respond to the intrusion.
Strategies Used to Protect Data Part
3: Isolation, Physical Security,
Backup & Disaster Recovery
A combination of security strategies are required to protect data stored
on systems. In this section we will look at strategies that focus on
physically protecting data and the mediums information is stored on.
Backup & Disaster Recovery
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Backup is the process of making a copy of data to a separate medium or
site to prevent the loss of data should the original copy be damaged or
lost, such as in a natural disaster or a cyber attack. Backups maybe stored
on an external device or networked location such as the cloud.
Recovery is the opposite process, whereby a backup copy of the data is
restored using a separate backup storage medium and copied back into
the operational system.
By conducting regular backups, developing and testing disaster recovery
plans, organisations can ensure minimal downtime and swift recovery in
critical situations.
Physical Security
Physical security focuses on safeguarding the physical infrastructure,
such as servers and data centres, from unauthorised intrusion, damage or
physical theft. Implementing secure access controls, surveillance
systems, and environmental protections ensures the physical safety of
data.
Isolation
Isolation involves segregating sensitive data from other network areas to
minimise unauthorised access or the spread of threats. This can be
achieved through encrypted mediums for highly confidential data,
network segmentation to create secure zones, and the use of
containerisation technologies for isolating applications.
System Access Security Measures
A combination of security strategies are required in order to protect data
stored on systems. In this section we will look at strategies that protect
data from being access by unauthorised users on a system or network.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Completely Automated Public Turing Test To Tell Computers and
Humans Apart (CAPTCHA)
A type of challenge-response test used in computing to determine
whether the user is human. It is commonly used to prevent bots from
accessing websites or using online services.
CAPTCHA tests might include identifying objects in photos, word
problems or equations. That are easy for humans but difficult for
automated systems to solve.
Automatic Software Updates
Automatic updates are used by developers to quickly distribute security
patches that fix vulnerabilities, which could be exploited by malware or
hackers.
Keeping software up-to-date is one of the most effective ways to protect
networks and devices from emerging security threats as the support is
provided directly from the software’s developers.
Trusted Platform Module (TPM)
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
A TPM is a specialised chip on an endpoint device (such as a desktop
computer) that provides hardware-based security functions. It is designed
to secure hardware through integrated cryptographic keys. The TPM can
be used to securely boot or provide full disk encryption for a system. It
stores encryption keys specific to the host system for hardware
authentication, making it difficult to access data without authorisation.
TPMs are used in various devices for secure computing, enhancing the
security of operating systems and protecting sensitive data. The
hardware-based approach to supports user authentication and network
access.
Data Backup & Recovery
Backup is the process of making a copy of data to separate media to
prevent the loss of data should the original copy be damaged or lost. This
involves duplicating data in each directory of a system’s drives and
duplicating it to an external device or networked location such as the
cloud. Recovery is the opposite process, whereby a backup copy of the
data is restored or copied back into the operational system.
Backups are necessary for restoring data that may have been
compromised or lost for the following reasons:
● Loss of data through electronic failure
● Mechanical fault of Hardware
● Software errors
● Corruption of data
● User-generated error
● Cybersecurity Incidents
Full & Partial Backups
A full backup copies all data within a system, every file in each directory
of the systems. A partial backup is completed to save time and resources.
Only files changed since the last full backup are stored. When performing
partial backups, there needs to be a full backup performed first, then on a
more frequent basis, partial backups are made.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Backup Software & Services
There is a variety of software and online cloud services available for the
backup and restoring of data for systems. Regardless of the developer,
these tools usually include the following features:
● The scheduling of automated backups
● The selection of different types of backup, full and partial to be used
in combination with each other.
● The selection of files for different backups
● File compression options, for reducing the backup file sizes.
● Features for efficient restoration of data, potentially from different
historical versions of files.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
DV Unit-3 - Data Visualization
Data analysis (Race Concept School)
Scan to open on Studocu
Studocu is not sponsored or endorsed by any college or university
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
10
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
11
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
12
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
13
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
14
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
15
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
16
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
17
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
18
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
19
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
20
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Comprehensive Guide to Geospatial Data and Cultural
Landscapes
Geography (Lafayette High School)
Scan to open on Studocu
Studocu is not sponsored or endorsed by any college or university
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Comprehensive Guide to Geospatial
Data and Cultural Landscapes
Geospatial Data
Geospatial data refers to any information that can be linked to a specific
geographic location on Earth, enabling a deeper understanding of spatial
relationships and patterns in our environment. This type of data is
fundamental in analyzing how different phenomena are distributed across
the landscape, revealing insights into natural and human-made features.
Examples include the location of roads, land cover types, population
distributions, and environmental conditions. By associating data with
precise coordinates, geospatial data allows us to visualize and interpret
complex spatial interactions, facilitating informed decision-making in
various fields such as urban planning, environmental management, and
public health.
Methods of Gathering Geospatial Data
Collecting accurate geospatial data involves multiple methods, each suited
to different contexts and objectives:
Fieldwork: Direct observation and recording at specific locations, such
as land surveys measuring distances, elevations, or environmental
conditions; census interviews gathering demographic data; or capturing
photographs and informal observations about land use. Fieldwork
provides firsthand, detailed data but can be time-consuming and
limited in scale.
Technology:
Global Positioning Systems (GPS): Devices receive signals from
satellites to determine exact coordinates, essential for navigation,
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
boundary delineation, and locating features like fire hydrants or hiking
trails.
Remote Sensing: Utilizes sensors on satellites or aircraft to capture
images or videos of Earth's surface, providing large-scale data on land
cover, vegetation health, and atmospheric conditions without physical
presence.
Secondary Sources: Government documents, treaties, environmental
regulations, news reports, and videos often contain geospatial
information. Smartphone data, when anonymized and aggregated, can
reveal movement patterns and activity hotspots. Photos with
embedded geotags also serve as geospatial data points, linking visual
information to specific locations.
Geospatial Technologies: GPS, Remote Sensing,
and GIS
Three key geospatial technologies are instrumental in acquiring, analyzing,
and visualizing spatial data:
Global Positioning System (GPS): Uses a constellation of satellites to
provide precise location data. Commonly integrated into smartphones
and vehicle navigation systems, GPS is vital for navigation, boundary
delineation, and emergency response. For example, emergency services
rely on GPS to locate individuals in need swiftly.
Remote Sensing: Involves collecting data from sensors mounted on
satellites or aircraft. It enables monitoring of environmental changes,
such as deforestation, glacier melting, urban expansion, and natural
disasters like floods or wildfires. Remote sensing provides critical data
for climate studies and disaster management.
Geographic Information Systems (GIS): Digital platforms that store,
analyze, and visualize layered spatial data. GIS allows for complex spatial
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
analysis, such as overlaying crime data with demographic information to
identify high-risk areas or combining transportation networks with
environmental data to improve urban mobility. GIS enhances decision-
making by revealing relationships that are not apparent in raw data.
Uses of Geospatial Data
Geospatial data plays a crucial role across numerous sectors:
Urban Planning: Designing efficient transportation systems, zoning, and
land use planning.
Environmental Monitoring: Tracking deforestation, pollution levels,
climate change effects, and natural resource management.
Public Health: Monitoring disease outbreaks, planning healthcare
facilities, and analyzing environmental health risks.
Disaster Response: Mapping affected areas during emergencies to
coordinate relief efforts.
Global Challenges: Addressing water scarcity, famine risks, and conflict
zones by analyzing spatial patterns and resource distribution.
Geovisualization
Geovisualization involves creating interactive maps and visual
representations of spatial data, often in two or three dimensions. Tools like
Google Earth, ESRI 3D GIS, and OpenStreetMap allow users to explore
detailed geographic information dynamically. For instance, during the
COVID-19 pandemic, Johns Hopkins University developed real-time maps
tracking infection spread worldwide, illustrating how geovisualization can
communicate complex data effectively. High-quality geovisualizations help
identify spatial patterns, support decision-making, and communicate
findings to diverse audiences, making abstract data accessible and
actionable.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Cultural Landscapes
Cultural landscapes are the visible imprints of human activity on the
environment, reflecting a culture's values, practices, and history. They
encompass physical structures, land use patterns, and spatial
arrangements that reveal how people adapt to and modify their
surroundings. These landscapes serve as tangible expressions of cultural
identity and social organization, illustrating the interaction between
humans and their environment over time.
Built Environment and Architectural Styles
The built environment includes all human-made structures such as
buildings, roads, bridges, and urban layouts. Architectural styles within
these structures provide insights into cultural values, historical periods,
and environmental adaptations:
Traditional architectures often utilize locally available materials and
techniques, reflecting the environment and cultural heritage. For
example, adobe homes in the southwestern U.S. are designed to suit hot,
dry climates.
Modern and postmodern architecture feature innovative designs,
often influenced by globalization, with styles like glass skyscrapers and
playful forms that symbolize technological advancement and
interconnectedness.
Contemporary architecture pushes boundaries with daring shapes
and sustainable designs, emphasizing environmental consciousness and
aesthetic experimentation.
Toponyms and Place Names
Toponyms, or place names, encode cultural, historical, and geographical
information:
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Names honoring historical figures or events reflect cultural memory and
societal values.
Geographical descriptors like "Riverbend" or "Mountain View" relate
directly to physical features, helping people orient themselves.
Changes in toponyms can indicate shifts in political power, cultural
identity, or social priorities, serving as markers of historical
transformation.
Ethnic Enclaves
Ethnic enclaves are neighborhoods with a high concentration of residents
sharing a common ethnicity, often forming for mutual support, cultural
preservation, or economic reasons:
They serve as cultural hubs where traditions, languages, and cuisines are
maintained.
Examples include "Little Italy" or "Chinatown," which showcase unique
architectural styles, businesses, and community events.
These enclaves often emerge due to historical segregation, migration
patterns, or economic opportunities, and they contribute to the cultural
diversity of urban landscapes.
Geography of Gender
The geography of gender examines how cultural norms influence spatial
organization and gender roles:
Traditional societies often assign specific spaces or activities as
gender-specific, shaping the physical landscape.
For example, certain areas might be designated for men or women
based on societal expectations.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Modern shifts, driven by increased education and social change,
challenge traditional gendered spaces, leading to more integrated and
equitable spatial arrangements.
Analyzing who has access to various spaces reveals power dynamics,
gender roles, and cultural values within a society.
Origins of Culture and Cultural Hearths
A cultural hearth is a geographic area where a specific culture or cultural
trait first develops, acting as a center of innovation and diffusion:
These hearths are crucial for understanding the origins and spread of
practices, technologies, and beliefs.
The Fertile Crescent, for example, is recognized as a cultural hearth for
agriculture, where domestication of plants and animals revolutionized
human society.
Innovations like writing originated in Mesopotamia and diffused across
civilizations, shaping cultural landscapes globally.
Cultural traits such as food preparation techniques, artistic styles, or
social norms often trace back to these hearths.
Traditional and Folk Cultures
Traditional cultures emphasize the transmission of long-held beliefs,
values, and practices across generations, often resisting rapid change.
They focus on continuity, community cohesion, and adherence to
customs.
Folk cultures are typically smaller, rural, and isolated groups with
distinct practices and strong community bonds. They tend to change
slowly, preserving unique traditions, languages, and crafts.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Examples include indigenous tribes in the Amazon or rural communities
in remote areas. Both types of cultures highlight the diversity of human
cultural expression and adaptation, often rooted in local environments
and histories.
This comprehensive guide synthesizes the key concepts from geospatial
data to cultural landscapes, providing a detailed understanding of how
humans interact with and shape their environments, and how these
interactions are documented, analyzed, and visualized through various
technologies and cultural expressions.# Comprehensive Guide to
Geospatial Data and Cultural Landscapes
Geospatial Data
Geospatial data refers to any information that can be linked to a specific
geographic location on Earth, enabling a deeper understanding of spatial
relationships and patterns in our environment. This type of data is
fundamental in analyzing how different phenomena are distributed across
the landscape, revealing insights into natural and human-made features.
Examples include the location of roads, land cover types, population
distributions, and environmental conditions. By associating data with
precise coordinates, geospatial data allows us to visualize and interpret
complex spatial interactions, facilitating informed decision-making in
various fields such as urban planning, environmental management, and
public health.
Methods of Gathering Geospatial Data
Collecting accurate geospatial data involves multiple methods, each suited
to different contexts and objectives:
Fieldwork: Direct observation and recording at specific locations, such
as land surveys measuring distances, elevations, or environmental
conditions; census interviews gathering demographic data; or capturing
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
photographs and informal observations about land use. Fieldwork
provides firsthand, detailed data but can be time-consuming and
limited in scale.
Technology:
Global Positioning Systems (GPS): Devices receive signals from
satellites to determine exact coordinates, essential for navigation,
boundary delineation, and locating features like fire hydrants or hiking
trails.
Remote Sensing: Utilizes sensors on satellites or aircraft to capture
images or videos of Earth's surface, providing large-scale data on land
cover, vegetation health, and atmospheric conditions without physical
presence.
Secondary Sources: Government documents, treaties, environmental
regulations, news reports, and videos often contain geospatial
information. Smartphone data, when anonymized and aggregated, can
reveal movement patterns and activity hotspots. Photos with
embedded geotags also serve as geospatial data points, linking visual
information to specific locations.
Geospatial Technologies: GPS, Remote Sensing,
and GIS
Three key geospatial technologies are instrumental in acquiring, analyzing,
and visualizing spatial data:
Global Positioning System (GPS): Uses a constellation of satellites to
provide precise location data. Commonly integrated into smartphones
and vehicle navigation systems, GPS is vital for navigation, boundary
delineation, and emergency response. For example, emergency services
rely on GPS to locate individuals in need swiftly.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Remote Sensing: Involves collecting data from sensors mounted on
satellites or aircraft. It enables monitoring of environmental changes,
such as deforestation, glacier melting, urban expansion, and natural
disasters like floods or wildfires. Remote sensing provides critical data
for climate studies and disaster management.
Geographic Information Systems (GIS): Digital platforms that store,
analyze, and visualize layered spatial data. GIS allows for complex spatial
analysis, such as overlaying crime data with demographic information to
identify high-risk areas or combining transportation networks with
environmental data to improve urban mobility. GIS enhances decision-
making by revealing relationships that are not apparent in raw data.
Uses of Geospatial Data
Geospatial data plays a crucial role across numerous sectors:
Urban Planning: Designing efficient transportation systems, zoning, and
land use planning.
Environmental Monitoring: Tracking deforestation, pollution levels,
climate change effects, and natural resource management.
Public Health: Monitoring disease outbreaks, planning healthcare
facilities, and analyzing environmental health risks.
Disaster Response: Mapping affected areas during emergencies to
coordinate relief efforts.
Global Challenges: Addressing water scarcity, famine risks, and conflict
zones by analyzing spatial patterns and resource distribution.
Geovisualization
Geovisualization involves creating interactive maps and visual
representations of spatial data, often in two or three dimensions. Tools like
Google Earth, ESRI 3D GIS, and OpenStreetMap allow users to explore
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
detailed geographic information dynamically. For instance, during the
COVID-19 pandemic, Johns Hopkins University developed real-time maps
tracking infection spread worldwide, illustrating how geovisualization can
communicate complex data effectively. High-quality geovisualizations help
identify spatial patterns, support decision-making, and communicate
findings to diverse audiences, making abstract data accessible and
actionable.
Cultural Landscapes
Cultural landscapes are the visible imprints of human activity on the
environment, reflecting a culture's values, practices, and history. They
encompass physical structures, land use patterns, and spatial
arrangements that reveal how people adapt to and modify their
surroundings. These landscapes serve as tangible expressions of cultural
identity and social organization, illustrating the interaction between
humans and their environment over time.
Built Environment and Architectural Styles
The built environment includes all human-made structures such as
buildings, roads, bridges, and urban layouts. Architectural styles within
these structures provide insights into cultural values, historical periods,
and environmental adaptations:
Traditional architectures often utilize locally available materials and
techniques, reflecting the environment and cultural heritage. For
example, adobe homes in the southwestern U.S. are designed to suit hot,
dry climates.
Modern and postmodern architecture feature innovative designs,
often influenced by globalization, with styles like glass skyscrapers and
playful forms that symbolize technological advancement and
interconnectedness.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Contemporary architecture pushes boundaries with daring shapes
and sustainable designs, emphasizing environmental consciousness and
aesthetic experimentation.
Toponyms and Place Names
Toponyms, or place names, encode cultural, historical, and geographical
information:
Names honoring historical figures or events reflect cultural memory and
societal values.
Geographical descriptors like "Riverbend" or "Mountain View" relate
directly to physical features, helping people orient themselves.
Changes in toponyms can indicate shifts in political power, cultural
identity, or social priorities, serving as markers of historical
transformation.
Ethnic Enclaves
Ethnic enclaves are neighborhoods with a high concentration of residents
sharing a common ethnicity, often forming for mutual support, cultural
preservation, or economic reasons:
They serve as cultural hubs where traditions, languages, and cuisines are
maintained.
Examples include "Little Italy" or "Chinatown," which showcase unique
architectural styles, businesses, and community events.
These enclaves often emerge due to historical segregation, migration
patterns, or economic opportunities, and they contribute to the cultural
diversity of urban landscapes.
Geography of Gender
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
The geography of gender examines how cultural norms influence spatial
organization and gender roles:
Traditional societies often assign specific spaces or activities as
gender-specific, shaping the physical landscape.
For example, certain areas might be designated for men or women
based on societal expectations.
Modern shifts, driven by increased education and social change,
challenge traditional gendered spaces, leading to more integrated and
equitable spatial arrangements.
Analyzing who has access to various spaces reveals power dynamics,
gender roles, and cultural values within a society.
Origins of Culture and Cultural Hearths
A cultural hearth is a geographic area where a specific culture or cultural
trait first develops, acting as a center of innovation and diffusion:
These hearths are crucial for understanding the origins and spread of
practices, technologies, and beliefs.
The Fertile Crescent, for example, is recognized as a cultural hearth for
agriculture, where domestication of plants and animals revolutionized
human society.
Innovations like writing originated in Mesopotamia and diffused across
civilizations, shaping cultural landscapes globally.
Cultural traits such as food preparation techniques, artistic styles, or
social norms often trace back to these hearths.
Traditional and Folk Cultures
Traditional cultures emphasize the transmission of long-held beliefs,
values, and practices across generations, often resisting rapid change.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
They focus on continuity, community cohesion, and adherence to
customs.
Folk cultures are typically smaller, rural, and isolated groups with
distinct practices and strong community bonds. They tend to change
slowly, preserving unique traditions, languages, and crafts.
Examples include indigenous tribes in the Amazon or rural communities
in remote areas. Both types of cultures highlight the diversity of human
cultural expression and adaptation, often rooted in local environments
and histories.
This comprehensive guide synthesizes the key concepts from geospatial
data to cultural landscapes, providing a detailed understanding of how
humans interact with and shape their environments, and how these
interactions are documented, analyzed, and visualized through various
technologies and cultural expressions.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Human Geography: Spatial Concepts and Geographic Data
Analysis
AP Human Geography (Vista High School)
messages.pdf_cover_qr_code_label
messages.studocu_not_sponsored_or_endorsed_by_college
messages.downloaded_by
lOMoARcPSD|63293293
Name:__________________________________________________ Date: _________________________________
Unit 1: Thinking Geographically
Human Geography study of people AND places
how we make places
how we organize space and society
how we interact with each other in places and across spaces
how we make sense of others and ourselves in our localities, regions, & the world
types of maps reference maps: maps used to show landforms and/or places
physical map: reference map that shows identifiable natural landmarks such as
mountains, rivers, oceans, elevation
political map: reference map that shows political boundaries
e.g. countries, cities, capitals, etc.
thematic maps: maps used to display specific types of information (theme) pertaining
to an area
cartogram: thematic map that shows statistical data by transforming space e.g. population
choropleth map: thematic map that uses shading or coloring to show statistical data e.g. population
dot density map: thematic map that uses dots to indicate a feature or occurrence
e.g. population
graduated symbols map (proportional symbols map): thematic map that indicates relative
magnitude of some value for a geographic region in which the symbol varies in proportion to data
e.g. population
types of spatial absolute distance: measurement using a standard unit of length
messages.downloaded_by
lOMoARcPSD|63293293
patterns represented on e.g. mile, kilometer
maps
relative distance: measurement of the social, cultural, and/or economic connectivity
between places (how connected or disconnected)
e.g. USA and Iran vs USA and China
absolute direction: finding a location using compass direction
e.g. north, south, east, west
relative direction: finding a location not using compass direction
e.g. left, right, forward, backward, up, down
spatial pattern: the way things are laid out and organized on the surface of the Earth
clustering: objects that form a group e.g. coastal population
dispersal: objects that are scattered e.g. rural population
elevation: height above sea level
spatial scale: hierarchy of spaces
e.g. location of French speakers:
global: in the world
regional: in North America
national: in Canada
local: in Quebec
map projections map distortion: all maps are distorted as a result of projecting a 3-dimensional surface onto a 2-dimensional surface
inevitably distort in area, distance, shape, and/or direction
spatial relationships in
shape, area, distance,
and direction map projection: a way to transfer the 3-dimensional earth onto a 2-dimensional map to reduce distortion
in area, distance, shape, and/or direction
data may be gathered geographic data: information that identifies the geographic location of features and boundaries on earth (natural and
messages.downloaded_by
lOMoARcPSD|63293293
in the field by constructed)
organizations or by
individuals
geospatial geospatial technologies: technology that provides geographic data that is used for personal (navigation), business
technologies (marketing), and governmental (environmental planning) purposes
GIS (Geographic Information System):
- map created by a computer that can combine layers of spatial data
- data is displayed and analyzed to gain insights into geographical patterns/relationships
e.g. vulnerability of the Florida Aquifer, school boundaries, crime rates
satellite navigation systems: system of satellites that provide geo-spatial positioning
e.g. GPS
remote sensing: collecting data with instruments that are distant from the area of study
types of Remote Sensors: satellites, planes, aircraft, spacecraft, ships, buoys
uses of Remote Sensing:
Track storm systems
Search for natural resources
Military surveillance
Monitor volcanoes
Monitor deforestation/glacier melting
online mapping and visualization: compilation and publication of web sites that provide
graphical and text information in the form of maps/visuals
e.g. homicide statistics
spatial information can spatial information can also come from written accounts (not just technology): field observations, media reports, travel
come from written narratives, policy documents, personal interviews, landscape analysis, and photographic evidence
accounts
geospatial and census data: systematically acquiring and recording information about the members of a given population
messages.downloaded_by
lOMoARcPSD|63293293
geographical data are
used at all scales
(personal, business, satellite imagery: images of earth collected by satellites operated by governments and businesses around the world
governmental decision
making)
spatial concepts absolute location: describes the precise location of a place using the Earth’s Graticule (latitude & longitude)
e.g. Palm Beach Gardens = 26°49′43″N 80°06′36″W
/ /
relative location: describes the location of a place relative to other human and physical features
e.g. Palm Beach Gardens = north of West Palm Beach, south of Jupiter
space (geography): relational concept that acquires meaning and sense when related to other concepts
e.g. geographers study phenomena across space
place: describes an area on the surface of the Earth with distinguishing human & physical characteristics
(place is space with meaning) e.g. Agra, India
pattern: an arrangement of objects on earth, including the space in between those objects
human-environment interaction: describes the ways humans modify or adapt to
the natural world e.g. bridges, dams, houses, roads
distance decay: the idea that the likelihood of interaction diminishes with increasing distance
time-space compression: term that refers to the increasing sense of connectivity that seems to be bringing people closer
together even though their distances are the same
time space convergence: term that refers to the greatly accelerated movement of goods, information, and ideas during the
20th century made possible by technological innovations e.g. TV, internet, satellite communication
movement (geography): describes the ways in which people, goods, and ideas move from place to place
flows (geography): movement in a steady stream e.g. migration
globalization: the process of increased interconnectedness among countries most notably in the areas of economics, politics,
and culture
network: a system of interconnected people or things e.g. transportation, communication, financial, governmental
concepts of nature and sustainability: meeting an increased demand for resources (energy, food, fuel) in a way that protects the ability of future
messages.downloaded_by
lOMoARcPSD|63293293
society generations to meet their own needs
natural resources: something found in nature and is necessary or useful to humans
e.g. forest, mineral deposit, water
land use: the function of land
e.g. agricultural, commercial, residential, transportation, recreation
theories regarding the environmental determinism: theory that a society is formed and determined by the physical environment, especially the
interaction of the climate; the physical environment predisposes societies towards particular development; human society development is
natural environment controlled by the environment
with human societies
possibilism: theory that the environment sets certain constraints or limitations but people use their creativity to decide how
to respond to the conditions of a particular natural environment
scales of analysis spatial scale: analyzing data at a variety of scales-global, regional, national, local
e.g. location of French speakers:
global: in the world
regional: in North America
national: in Canada
local: in Quebec
patterns and processes spatial scale: analyzing data at different scales reveal variations/different interpretations of data
at different scales e.g. fertility rate
global: in the world (2.4)
regional: in Sub-Saharan Africa (4.7)
national: in Tunisia (2.1)
regions region: describes an area on Earth marked by similarity in some way (a way to organize space)
messages.downloaded_by
lOMoARcPSD|63293293
regionalism: refers to a group’s perceived identification with a particular region
e.g. the South
types of regions formal region: region marked by a shared trait (cultural, physical, etc.)
e.g. The Keys, The Caribbean
functional region: region marked by a particular set of activities that occur
e.g. Southwest Airlines, newspaper
perceptual/vernacular region: region that exists as an idea
e.g. the South, Kurdistan
regional boundaries regional boundaries: transitional and often contested and overlapping
e.g. Kurdistan in Turkey and Northern Iraq
regional analysis regional analysis: analyzing regions at a variety of scales-global, national, local
e.g. Muslim population
global: in the world
national: in Turkey
local: in Kurdistan
messages.downloaded_by
lOMoARcPSD|63293293
PDF document-F2775823 A8A1-1
Mathematics: Applications and Interpretation SL (Maharashtra State Secondary High
School & Junior College)
Scan to open on Studocu
Studocu is not sponsored or endorsed by any college or university
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
UNIT 5
Data Visualization - Basic principles, ideas and tools for data visualization 3 - Examples
of inspiring (industry) projects - Exercise: create your own visualization of a complex
dataset - Data Science and Ethical Issues - Discussions on privacy, security, ethics - A
look back at Data Science - Next-generation data scientists
Module V: Data Visualization and Data Science
This module delves into the principles, ideas, and tools of data visualization, explores
industry applications, and addresses ethical and future considerations for data science. The
goal is to equip students with theoretical and practical knowledge to create insightful
visualizations and ethically sound data-driven decisions.
This module provides foundational skills for data visualization, awareness of data science
applications, and an understanding of ethical issues in the evolving landscape of data science.
1. Data Visualization: Basic Principles, Ideas, and Tools
1.1 Definition and Purpose of Data Visualization
Definition: Data visualization is the art and science of displaying data in a graphical
or visual form to make complex information more understandable, insightful, and
actionable.
Purpose:
o Insights and Patterns: Data visualization helps discover trends, correlations,
and outliers.
o Accessibility: Visualization breaks down complex information, making it
accessible to both technical and non-technical audiences.
o Decision-Making: Well-constructed visuals can drive more informed and
strategic decision-making in business, policy, and research.
1.2 Basic Principles of Data Visualization
Clarity
o Visualization must convey information in an immediately understandable way.
o Avoid unnecessary decorative elements that do not contribute to the
understanding of the data.
Simplicity
o Eliminate non-essential details to avoid clutter and focus on key insights.
o Use white space effectively and avoid overly complex charts for simpler
insights.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Accuracy
o Represent data precisely to prevent misinterpretation; for instance, starting
axis values at zero or clearly defining them.
o Avoid scaling or coloring that may distort data or imply correlations not
present in the dataset.
Aesthetics
o Visual appeal helps hold the viewer's attention and enhances comprehension.
o Use balanced colors, fonts, and alignment for a clean, professional look.
Contextuality
o Provide sufficient labels, legends, and explanations so that the data can be
understood in its correct context.
o Incorporate descriptive titles, annotations, and explanations as needed.
1.3 Key Ideas in Data Visualization
Choosing the Right Chart: Selection should align with the nature and goals of the
data:
o Bar Charts: For comparing quantities across categories (e.g., sales across
regions).
o Line Charts: For trends over time (e.g., stock prices, population growth).
o Scatter Plots: To identify relationships or correlations between two variables
(e.g., age vs. income).
o Heat Maps: For density or intensity in a spatial format (e.g., geographic data).
Color Usage:
o Use colors strategically to group similar data points or highlight key aspects.
o Avoid excessive color use, which can distract; stick to 2-4 key colors with
meaningful contrasts.
Storytelling with Data:
o The best visualizations tell a story, guiding the viewer through a logical
sequence of insights.
o Utilize annotations to emphasize critical points and provide necessary
background or context.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
1.4 Popular Tools for Data Visualization
Tableau: Known for its user-friendly interface, it allows for highly interactive, drag-
and-drop style visualization, making it ideal for business intelligence.
Power BI: Microsoft’s visualization tool is highly integrated with Office and suitable
for enterprise reporting.
Matplotlib & Seaborn (Python): Matplotlib is a robust library for static
visualizations, while Seaborn offers additional statistical functions to enhance these
visuals.
ggplot2 (R): A powerful package for creating detailed custom visualizations,
particularly in the field of data science.
[Link]: A JavaScript library for creating highly customizable web-based visualizations,
useful for dynamic, interactive displays.
2. Inspiring Industry Examples
2.1 Netflix - User Recommendations System
Description: Netflix leverages data visualization to provide insight into how
recommendations are tailored. By displaying how viewing history and ratings
influence suggested titles, they increase user engagement and trust in the platform.
Impact: This transparency builds user loyalty and satisfaction, as viewers can better
understand the personalized content offered.
Netflix - User Recommendations System
Netflix’s recommendation system is a cornerstone of its success, helping drive user
engagement by suggesting personalized content based on individual viewing habits.
Netflix uses data visualization to illustrate the mechanics of its recommendation
engine, showing users how their viewing history, ratings, and preferences shape the
titles recommended to them. This transparency encourages user engagement and
loyalty, as users can better understand and trust the process that brings personalized
content to their screens.
In terms of implementation, a recommendation system in R would typically involve
several key steps, including data collection, preprocessing, creating a recommendation
model, and visualizing the output. Here’s a breakdown of how Netflix’s
recommendation system can be prototyped in R:
Step 1: Data Collection and Preprocessing
In a typical Netflix-style recommendation setup, you would have a dataset with:
1. User viewing history (which shows which users have watched which movies or
shows),
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
2. Ratings (ratings users have given to watched titles), and
3. Movie metadata (details about each title, such as genre, release year, etc.).
For example, we might use a dataset like MovieLens, which contains information
about movies, users, and their ratings.
# Load required libraries
library(dplyr)
library(ggplot2)
library(recommenderlab) # for building a recommendation model
# Load the MovieLens data (assuming it's in CSV format)
movie_data <- [Link]("[Link]")
rating_data <- [Link]("[Link]")
# Data Preprocessing
# Filter out unnecessary columns and handle missing data
rating_data <- rating_data %>% select(userId, movieId, rating)
movie_data <- movie_data %>% select(movieId, title, genres)
# Merge data for a complete dataset
merged_data <- merge(rating_data, movie_data, by = "movieId")
Step 2: Building a Recommendation Model
Using collaborative filtering, a common recommendation method, we can predict
which movies users may enjoy based on their historical preferences. We’ll use the
recommenderlab package in R to build a model.
# Convert data to a recommendation-friendly matrix format
rating_matrix <- as(merged_data, "realRatingMatrix")
# Build a user-based collaborative filtering model
rec_model <- Recommender(rating_matrix, method = "UBCF")
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
# Generate recommendations for users
# For instance, recommending top 5 titles for the first 5 users
recommended <- predict(rec_model, rating_matrix[1:5], n = 5)
Step 3: Visualizing the Recommendation Influence
Netflix-style visualization can show how a user’s viewing history and ratings
influence their recommendations. Here’s how we could visualize, in a simple way, the
most influential genres based on a user’s watch history.
R Code for Visualization:
# Calculate genre influence based on watched movies
user_history <- merged_data %>% filter(userId == 1) # Assuming user ID 1
genre_influence <- user_history %>%
separate_rows(genres, sep = "\\|") %>%
group_by(genres) %>%
summarise(watch_count = n()) %>%
arrange(desc(watch_count))
# Plotting the genre influence
ggplot(genre_influence, aes(x = reorder(genres, -watch_count), y = watch_count)) +
geom_bar(stat = "identity", fill = "skyblue") +
theme_minimal() +
labs(title = "Influence of Viewing History on Recommendations",
x = "Genres",
y = "Watch Count") +
coord_flip()
Explanation of Each Step:
1. Data Preprocessing: Cleaned and merged user ratings and movie information to
create a comprehensive dataset that reflects viewing habits and preferences.
2. Modeling: Used collaborative filtering to make user-based recommendations,
reflecting a typical approach in recommendation systems.
3. Visualization: Displayed a bar chart that reflects a user's genre preferences, helping
the user understand the types of content they’re likely to see recommended.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
This R-based example captures the essence of Netflix’s approach by providing
insights into how user data drives recommendations, building transparency, and
enhancing user trust in the system.
The output for the Netflix-style recommendation system in R would consist of two
primary components:
1. Recommendations for Users: A list of recommended movies for specific users,
based on their viewing history and ratings.
2. Visualization of Genre Influence: A bar plot showing the genres most frequently
watched by a specific user, illustrating how their viewing history influences the
recommendations.
Output 1: Recommendations for Users
The recommendation output is based on collaborative filtering. Here’s what the output
might look like for the first five users, assuming we’re recommending five titles per
user:
# Viewing the recommendations
as(recommended, "list")
Example Output (assuming user IDs 1 to 5)
$`User 1`
[1] "Inception" "The Matrix" "Interstellar" "The Dark Knight" "Pulp Fiction"
$`User 2`
[1] "Forrest Gump" "Titanic" "Gladiator" "Braveheart" "The Godfather"
$`User 3`
[1] "The Shawshank Redemption" "Fight Club" "Seven" "Gone Girl" "The
Silence of the Lambs"
$`User 4`
[1] "The Avengers" "Iron Man" "Thor" "Guardians of the Galaxy" "Spider-
Man"
$`User 5`
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
[1] "Toy Story" "The Lion King" "Finding Nemo" "Shrek" "Monsters, Inc."
Each user receives a list of five titles recommended based on their viewing patterns.
Output 2: Visualization of Genre Influence
This visualization is a bar chart showing the influence of different genres on the user’s
viewing history, helping them see the genre preferences that likely contribute to their
recommendations.
Code Output (Bar Plot):
The code provided generates a bar plot like this:
Genre Influence on Viewing History
----------------------------------------------------------
| Genre | Watch Count |
----------------------------------------------------------
| Action | 12 |
| Drama |8 |
| Adventure |7 |
| Sci-Fi |5 |
| Comedy |4 |
Here's how it looks visually:
Y-Axis: Genres
X-Axis: Number of times a genre was watched
Bar Color: The bars represent the frequency (watch count) in "sky blue."
This bar plot makes it easy for users to see which genres they've watched most,
helping them understand why certain types of content might be recommended to
them.
2.2 Airbnb - Diversity Dashboard
Description: Airbnb’s Diversity Dashboard displays employee diversity statistics to
promote workplace inclusivity and transparency.
Impact: This dashboard encourages diversity initiatives by making diversity metrics
accessible to the public, driving internal and external accountability.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
2.3 COVID-19 Tracking by Johns Hopkins University
Description: A real-time data dashboard providing global COVID-19 statistics, such
as case counts, recoveries, and fatalities.
Impact: This data visualization helped both policymakers and the public make
informed decisions throughout the pandemic, serving as a crucial resource in
managing public health responses.
2.4 Spotify - Year in Music
Description: Spotify generates personalized annual music summaries for users,
showing listening trends, top artists, and favorite genres.
Impact: This enhances user experience by providing a personal recap and makes data
engaging and relatable, promoting Spotify’s brand through social sharing.
2.5 Uber Movement
Description: Uber’s Movement tool visualizes billions of anonymized rides to
highlight traffic patterns and urban mobility insights.
Impact: This data helps city planners, policymakers, and researchers understand and
improve urban mobility, which can lead to smarter city infrastructure.
3. Exercise: Create Your Own Visualization of a Complex Dataset
This exercise will encourage hands-on practice, allowing students to apply visualization
principles to create a compelling visual from raw data.
1. Dataset Selection:
o Students will choose from datasets in domains like healthcare, finance, retail,
or e-commerce. Suggested datasets include public sources like government
databases or open-source repositories such as Kaggle.
2. Tools: Students are free to use any tool—Tableau, Power BI, or programming libraries
(Python or R).
3. Visualization Process:
o Data Exploration: First, perform an exploratory data analysis to understand
the dataset's characteristics.
o Defining Key Insights: Identify 1-3 key insights or trends to focus the
visualization on.
o Chart Selection: Select charts and visual elements that best represent the
dataset’s complexity and findings.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
o Design: Ensure adherence to principles of clarity, simplicity, and contextuality.
4. Presentation:
o Present the visualization to the class, explaining how each design choice
supports the insights.
o Highlight the tool used, any challenges encountered, and the final takeaway
insights.
4. Applications of Data Science
Healthcare:
o Applications range from predicting patient outcomes to detecting diseases
through image recognition.
o Example: Machine learning models predict patient readmission rates to
improve care and reduce hospital costs.
Finance:
o In finance, data science powers fraud detection, credit risk scoring, and
investment portfolio optimization.
o Example: Credit card companies use anomaly detection to identify fraudulent
transactions.
Retail:
o Retailers use data science for inventory management, personalized marketing,
and demand forecasting.
o Example: E-commerce platforms recommend products based on past purchase
behavior.
Social Media:
o Analysis of social media data enables companies to understand trends,
sentiment, and user engagement.
o Example: Social listening tools help brands manage reputation and engage
with trends in real-time.
Manufacturing:
o Predictive maintenance in manufacturing minimizes downtime by anticipating
equipment failures.
o Example: Sensor data analysis enables proactive maintenance schedules,
reducing unplanned outages.
5. Data Science and Ethical Issues
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Data scientists must be aware of ethical considerations that influence public trust and the
integrity of the field.
Privacy:
o Ethical data science practices must respect user consent and the sensitivity of
data.
o Example: GDPR-compliant data usage policies that inform users of how their
data will be used.
Security:
o Data must be stored and processed securely to prevent breaches and
unauthorized access.
o Example: Encrypting sensitive financial data and limiting access to authorized
personnel only.
Bias in Algorithms:
o Algorithms can reinforce existing biases if training data isn’t representative or
if biases are embedded.
o Example: Facial recognition systems must be trained on diverse datasets to
avoid racial and gender biases.
Transparency:
o Making algorithmic decisions transparent is essential for user trust.
o Example: Transparency in AI decision-making helps users understand
automated outcomes.
Data Ownership:
o Ownership issues arise when third-party data is repurposed without consent.
o Example: Social media platforms must obtain explicit permission to use user-
generated data in research.
6. A Look Back at Data Science and Next-Generation Data Scientists
Evolution of Data Science:
o It has evolved from simple statistical analysis to advanced machine learning
and AI, influenced by big data advancements.
Below is an outline of its evolution:
1. Pre-Digital Era (Before 1950s)
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Key Characteristics:
Data collection and analysis were manual.
Mathematical and statistical methods, such as regression analysis and probability
theory, formed the basis of data analysis.
Tools: Hand-written records, basic calculators, and rudimentary statistical tools.
Notable Events:
1662: John Graunt published the first known statistical analysis of population data in
his book, Natural and Political Observations Made upon the Bills of Mortality.
1800s: The advent of probability and inferential statistics.
2. Early Computer Age (1950s–1970s)
Key Characteristics:
Emergence of computing power allowed faster calculations.
Use of programming languages like FORTRAN and COBOL for data processing.
Focus on database management systems (DBMS).
Notable Events:
Development of relational databases by Edgar F. Codd in the 1970s.
Introduction of statistical software like SPSS (1968).
3. Big Data and Internet Revolution (1980s–1990s)
Key Characteristics:
Widespread adoption of personal computers and internet connectivity.
Increased generation of structured and unstructured data.
Development of advanced machine learning (ML) algorithms.
Notable Events:
Birth of the World Wide Web in 1989 by Tim Berners-Lee.
Introduction of data warehousing and data mining techniques.
4. The Era of Data-Driven Decisions (2000s–2010s)
Key Characteristics:
Explosion of data due to social media, IoT, and mobile devices.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Rise of cloud computing, enabling scalable data storage and analysis.
Democratization of AI and machine learning tools.
Notable Events:
Introduction of Hadoop (2006) and the MapReduce framework.
Emergence of Python and R as leading programming languages for data analysis.
5. Modern Data Science (2010s–Present)
Key Characteristics:
Integration of AI, deep learning, and neural networks into data science workflows.
Real-time data analysis through edge computing.
Focus on ethical AI, data privacy, and interpretability.
Notable Events:
Growth of self-service analytics platforms like Tableau and Power BI.
Advances in natural language processing (NLP) with models like GPT.
6. Future Directions
Trends to Watch:
Quantum computing’s potential impact on complex data analysis.
Increased use of generative AI and large language models (LLMs).
Enhanced tools for data collaboration and visualization.
Challenges:
Addressing ethical concerns and biases in AI.
Managing and securing vast amounts of sensitive data.
Technological Advancements:
o Emerging fields include deep learning, augmented analytics, and real-time
data analysis, all of which demand versatile skill sets.
The evolution of Data Science has been powered by key technological
breakthroughs across various periods. Below is an overview of these
advancements within the outlined stages:
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
1. Pre-Digital Era (Before 1950s)
Technological Milestones:
Development of mechanical calculators like the Pascaline and the Difference
Engine.
Statistical tools like frequency tables and hand-drawn graphs.
2. Early Computer Age (1950s–1970s)
Technological Milestones:
Mainframe Computers: Early machines like IBM 704 for data processing.
Programming Languages: Introduction of FORTRAN (1957) and COBOL (1959),
simplifying data handling and processing.
Relational Databases: Edgar F. Codd's Relational Database Model (1970)
revolutionized data storage and querying.
Batch Processing: Used for handling large amounts of data in jobs or sets.
3. Big Data and Internet Revolution (1980s–1990s)
Technological Milestones:
Data Warehousing: Tools like Teradata emerged for structured data storage and
analysis.
Data Mining Tools: Techniques like decision trees, clustering, and association
rules.
Internet and Web: The creation of the World Wide Web (1989) enabled
unprecedented data generation and sharing.
First Statistical Software: Tools such as SAS (introduced in the 1970s, matured in
the 80s) and SPSS became widespread.
4. The Era of Data-Driven Decisions (2000s–2010s)
Technological Milestones:
Big Data Frameworks: Tools like Hadoop (2006) and MapReduce enabled
distributed computing for massive datasets.
Cloud Computing: Platforms like AWS, Azure, and Google Cloud offered scalable
data storage and analytics capabilities.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Open Source Tools: Languages like Python and R became dominant, supported by
libraries like Pandas, NumPy, TensorFlow, and scikit-learn.
Visualization Tools: Emergence of Tableau, Power BI, and Matplotlib for
accessible data storytelling.
5. Modern Data Science (2010s–Present)
Technological Milestones:
Deep Learning Frameworks: Development of TensorFlow, PyTorch, and Keras
for advanced AI applications.
Natural Language Processing: Breakthroughs in NLP with transformer models
like BERT and GPT.
Edge Computing: Real-time data analysis at the source, reducing latency and
enhancing IoT applications.
Self-Service Analytics: Tools like Qlik and Looker empower non-technical users to
explore data.
6. Future Directions
Technological Milestones to Watch:
Quantum Computing: Promises exponential speed-ups for complex computations.
Automated Machine Learning (AutoML): Simplifies model building for non-
experts.
Generative AI: Advancements like DALL-E, Stable Diffusion, and enhanced
LLMs.
Blockchain for Data Integrity: Ensures data provenance and tamper-proof storage.
Skills for Next-Generation Data Scientists:
The rapid evolution of Data Science has created a demand for professionals equipped with a
diverse skill set that goes beyond traditional data analysis. Below is a detailed explanation of
the core skills required for next-generation data scientists:
1. Technical Skills
Overview:
Next-generation data scientists must have strong technical expertise to handle increasingly
complex data and leverage cutting-edge tools.
Key Areas of Technical Expertise:
1. Proficiency in Programming:
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
o Mastery of languages like Python, R, and SQL is essential for data
manipulation, analysis, and visualization.
o Knowledge of C++ and Java may be necessary for performance-intensive
applications.
o Familiarity with frameworks like Spark or Hadoop for handling big data.
2. Advanced Machine Learning and AI:
o Deep understanding of machine learning algorithms, including supervised,
unsupervised, and reinforcement learning.
o Expertise in deep learning frameworks like TensorFlow, PyTorch, and
Keras.
o Hands-on experience in natural language processing (NLP), computer
vision, and time-series analysis.
3. Data Engineering:
o Ability to design and optimize pipelines for data collection, transformation,
and storage.
o Proficiency in cloud technologies like AWS, Azure, or Google Cloud for
scalable solutions.
o Skills in database management (relational and NoSQL databases).
4. Visualization and Communication:
o Use of tools like Tableau, Power BI, and libraries such as Matplotlib and
Seaborn for creating compelling visualizations.
o Communicating insights effectively through storytelling.
2. Ethical Awareness
Overview:
As the impact of data-driven decisions grows, ethical considerations are vital for ensuring
fairness, accountability, and transparency in data science applications.
Key Areas of Ethical Awareness:
1. Bias and Fairness in Algorithms:
o Awareness of biases in datasets and machine learning models.
o Implementing techniques to mitigate biases and ensure equitable outcomes.
2. Data Privacy and Security:
o Understanding global data protection laws such as GDPR, CCPA, and HIPAA.
o Ensuring secure handling and storage of sensitive data.
3. Social Impact of Algorithms:
o Assessing the societal effects of deploying predictive models, especially in
sensitive areas like hiring, credit scoring, and law enforcement.
o Building interpretable models that provide explanations for their decisions
(e.g., using SHAP or LIME).
4. Transparency and Accountability:
o Advocating for open communication with stakeholders about the limitations
and potential risks of data-driven solutions.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
3. Domain Knowledge
Overview:
Beyond technical skills, a deep understanding of specific industries enables data scientists to
craft impactful, context-aware solutions.
Key Areas of Domain Knowledge:
1. Finance:
o Expertise in risk modeling, fraud detection, and portfolio optimization.
o Familiarity with market behavior, financial instruments, and economic trends.
2. Healthcare:
o Understanding of medical data formats, terminology, and regulatory
requirements like HIPAA.
o Experience with predictive modeling for patient diagnosis, treatment
recommendations, and resource allocation.
3. Urban Planning and Smart Cities:
o Knowledge of geospatial data analysis and infrastructure optimization.
o Applications such as traffic management, energy optimization, and
environmental monitoring.
4. Retail and E-Commerce:
o Insights into consumer behavior analysis, inventory optimization, and
personalized recommendations.
5. Other Emerging Fields:
o Autonomous vehicles, climate science, and education technology are also
becoming crucial areas where domain-specific knowledge enhances data
science applications.
Downloaded by Aseneth Jepchirchir (asenethjepchirchir96@[Link])
lOMoARcPSD|63293293
Data Visualization Techniques: Identifying Misleading Charts
and Graphs
Statistics. (Montclair State University)
messages.pdf_cover_qr_code_label
messages.studocu_not_sponsored_or_endorsed_by_college
messages.downloaded_by
lOMoARcPSD|63293293
Class 2
Data Visuals, Good
and Bad: Part I
by Amrit Parmar
messages.downloaded_by
lOMoARcPSD|63293293
Introduction
In a world where data is so important, we all want to create clear and effective charts.
However, constructing accurate data visuals isnʼt something that is usually taught in school or during on-the-job
training. Instead, we often pick it up as we go, which means we often make choices that can result in misleading
visuals.
Similarly, incorrect data visualization has been frequently used in traditional media (especially TV and
newspaper) and social media to mislead the audience, whether intentionally or unintentionally.
Goals
In this chapter, we will learn about data visualization techniques, and call out BS on various misleading data
visualizations that have been published or presented in the media the years.
messages.downloaded_by
lOMoARcPSD|63293293
Visualizing Categorical Data
For categorical data, we typically want to compare the frequency or proportion of categories.
Pie Charts Bar Charts Stacked Bar Charts
Great for showing the proportion Useful for comparing quantities Useful for comparing the
of categories within a whole across different categories composition of categories across
groups
messages.downloaded_by
lOMoARcPSD|63293293
Visualizing Numerical Data
For numerical data, we typically want to show distributions, trends, or relationships
Line Charts Histograms Scatterplots
Ideal for showing trends over time Great for showing the distribution Useful for showing the
or continuous data of a single continuous numerical relationship between two
variable numerical variables
messages.downloaded_by
lOMoARcPSD|63293293
Pie Charts
Pie charts are commonly used to represent categorical data, with each slice showing a category's contribution to
the whole. They work best with a small number of categories and for simple, straightforward visualizations.
Example: Let's visualize the distribution of pet types in a community by percentage, where the data shows 40% dogs,
30% cats, 15% fishes, 10% birds, and 5% others.
Distribution of Types of Pets Owned in a Community
Dogs Cats Fish Birds Others
messages.downloaded_by
lOMoARcPSD|63293293
What do you think about this chart?
Point 1 Point 2 Point 3
Percentage of Americans who The individual slices of pie do Misleading data visualization
smoke marijuana in different not add up to 100%, this is the to convince the audiences that
years is a trend over time (i.e. most common way to mislead the percentage of Americans
quantitative data). A line the audience with a pie chart. who have smoked marijuana
graph should be used here has drastically increased
instead. since 1997.
messages.downloaded_by
lOMoARcPSD|63293293
What do you think about this chart?
Point 1 Point 2 Point 3 Point 4
'Biggest COVID-19 The individual slices Even if we assume Thereʼs no mention
worries' is of pie do not add up that individual slices of the sample size
categorical data, to 100%, this is the added up to 100%, or source. Did they
and the author is most common way red, orange, and sample 10 people in
trying to show the to mislead the yellow are too your studio or 1000
proportion of audience with a pie similar as colors. Itʼs random Americans?
categories within a chart. always better to use What are we looking
whole. So a pie- three easily at?
chart is the correct distinguishable
graphical colors.
representation.
messages.downloaded_by
lOMoARcPSD|63293293
Bar Charts: Best Practices
Bar charts are effective for comparing quantities across different categories.
They should have a vertical axis starting at 0 and consistent calibration.
Common mistakes include not starting the y-axis at 0, inconsistent scaling,
and improper ordering of categories.
1 Proper Axis Calibration
Always start the vertical axis at 0 for accurate comparisons
2 Consistent Scaling
Use uniform increments on both axes
3 Logical Ordering
Arrange categories in a meaningful order (e.g. chronological,
alphabetical, or by value)
4 Clear Labeling
Include clear labels and a legend if necessary
messages.downloaded_by
lOMoARcPSD|63293293
Bar Chart Example
Average Monthly
Rainfall
This bar chart shows the
average monthly rainfall in mm
across major US Cities.
messages.downloaded_by
lOMoARcPSD|63293293
What do you think about this graph?
Point 1 Point 2 Point 3
No calibration of vertical axis No consistency in calibration Clearly a misleading bar chart
(y-axis). A good bar chart of percentages on the vertical to forcefully show that NBC2
always has the vertical axis axis. 13% has the highest bar viewers are “not at allˮ
starting from 0. where in reality it should be concerned about the zika
the lowest. 28% doesnʼt even virus
have a bar to it
messages.downloaded_by
lOMoARcPSD|63293293
What do you think about this graph?
Point 1 Point 2
A good bar chart always has the vertical axis Here the y-axis starts at 58% instead of 0%,
starting from 0 so that we have an even exaggerating the difference between the two bars.
comparison. Not starting the vertical axis at 0 is The difference between 77% and 65% is
the most common ways to mislead the audience significant, but the chart makes it look far more
and make a flawed bar chart. drastic than it actually is.
messages.downloaded_by
lOMoARcPSD|63293293
How it should really look…
messages.downloaded_by
lOMoARcPSD|63293293
What do you think about this graph?
Point
Here the y-axis starts at 155,000 instead of 0. This makes the change look much more drastic than it actually is.
messages.downloaded_by
lOMoARcPSD|63293293
What do you think about this graph?
Point 1 Point 2
In this graph we can see that, x-axis is not Both axes should have consistent, uniform
organized chronologically. This is one of the most increments. For example, if the vertical axis starts
commonly used technique to make a deceptive with increments of 5 (0, 5, 10, 15, 20), it shouldn't
graph. It gives the false appearance of an overall suddenly jump to larger values like 70, 80, 90, 100.
decrease in cases across the board. In this case, the vertical axis is well-calibrated.
messages.downloaded_by
lOMoARcPSD|63293293
From the March 12, 2020 broadcast of
The Rachel Maddow Show…
Point 1 Point 2 Point 3
COVID-19 cases by day is Take a look at the horizontal What effect do you think this
numerical data, and so a line axis (x-axis). It's jumping from spacing has on the
graph is the ideal graphical January 21 to February 21st to representation of the data?
representation to highlights February 28 and an increment What are the potential
trends over time. In this case, of 4 days. NEVER do this! motives of the ones who
you should be using a line There should always be an designed this graph?
graph instead of a bar chart. even spacing between the
dates on for correct data
visualization.
messages.downloaded_by
lOMoARcPSD|63293293
Line Charts: Effective Use and Potential
Pitfalls
Line charts are ideal for displaying trends or changes over time. While it's generally acceptable not to start the y-axis at
0 for line charts, caution is needed to avoid exaggerating differences. Common errors include inconsistent axis
increments and improper scaling.
Choose Appropriate Data
1 Use line charts for continuous data or trends over time
Consider Y-Axis Start Point
2 Starting at 0 isn't always necessary, but be cautious about exaggerating differences
Use Consistent Increments
3 Ensure both axes have uniform, logical increments
Avoid Misleading Scales
4 Be wary of manipulated scales that distort the data's true nature
messages.downloaded_by
lOMoARcPSD|63293293
Line Chart Example
messages.downloaded_by
lOMoARcPSD|63293293
'Stand Your Ground', Effective Law?
Point 1
It looks like the gun death in Florida was on a rise
in 1990s and once the ‘Stand Your Groundʼ law
was enacted in 2005, then the murders by
firearm dropped drastically. Right?
Point 2
Wrong! The vertical axis is inverted! It doesnʼt
start at 0. What we would normally perceive as a
decreasing trend is actually an increasing trend,
and vice-versa.
messages.downloaded_by
lOMoARcPSD|63293293
What Standing Your Ground really did
messages.downloaded_by
lOMoARcPSD|63293293
What do you think about this report?
Point 1 Point 2 Point 3
Always start by asking However, look carefully at the It gives the graph a flattened
yourself: "Is the graphical y-axis and the uneven graph look, essentially making the
representation technique a intervals. What's the problem growth during the later dates
good one to represent the here? What effect does it seem less drastic and more
kind of data weʼre dealing have on the data? stabilized.
with?" COVID-19 Cases per
day is numerical data, so a
line graph is the ideal
graphical representation to
highlights trends over time in
this case.
messages.downloaded_by
lOMoARcPSD|63293293
Conclusion: Spotting and
Avoiding Misleading
Visualizations
Misleading data visualizations come in many forms. To spot them, question
the source, examine the axes and scales, and consider if the visualization
method is appropriate. Practice and vigilance are key to identifying
deceptive graphs and charts.
Assess Appropriateness Examine Closely
Evaluate if the visualization method Look carefully at axes, scales, and
fits the data type data representation
Question Sources Think Critically
Consider the origin and potential Apply logical reasoning to interpret
biases of the data the presented information
messages.downloaded_by
lOMoARcPSD|63293293
Student assignment - class
Computer Science (Catholic University of Eastern Africa)
messages.pdf_cover_qr_code_label
messages.studocu_not_sponsored_or_endorsed_by_college
messages.downloaded_by
lOMoARcPSD|63293293
Data Visualization Principles - Student Assignment
This assignment will assess your ability to apply the principles of effective data visualization,
including clarity, accuracy, and storytelling. You will use Power BI or Tableau to create high-quality
visualizations and provide a written critique explaining your design decisions.
Tasks:
• Using the provided dataset, create at least five (5) different visualizations that follow best
practices:
• 1. Line Chart – Show a trend over time.
• 2. Bar Chart – Compare categories.
• 3. Scatter Plot – Show relationships between two variables.
• 4. Stacked Bar Chart – Display composition within categories.
• 5. Map Visualization – Show geographic distribution (if applicable).
• Ensure your visuals are clear, accurate, and tell a story.
• Write a 500–700 word critique explaining your design choices using the provided critique
template.
• Submit both the visualization file (Power BI or Tableau) and your written critique document.
Critique Template:
• Visualization Title
• Chart Type Used
• Reason for Choice
• How it Ensures Clarity
• How it Maintains Accuracy
• How it Supports Storytelling
• Possible Improvements
Grading Criteria:
• Application of principles – 30%
• Chart quality and formatting – 30%
• Clarity and depth of critique – 20%
• Creativity & storytelling – 20%
Submission Deadline: Please submit all files by the date specified by your instructor.
messages.downloaded_by
An Introduction to Analysis and Data
Visualization using Tableau Software
DEPLOY FOR GROWTH 1
Presentation 01 What is Tableau Software?
02 Benefits for Teachers & Researchers
Overview 03 What is Data Visualization?
04 General Overview of Tableau
05 Use for Reporting - Examples
06 Use for Storytelling - Examples
07 Use for Analysis - Examples
08 Advanced Features - Example
09 Resources (Public, WMTUG, Books)
DEPLOY FOR GROWTH 2
What is Tableau Software?
• Software company Founded in 2003 from
Stanford research
• Intent is to bring ‘data to the people’
through easy to use data visualization
software
• Would be classified as a hybrid business
intelligence (BI) / analytics software
company
• Used by many of the largest companies in
the world and most large companies in
West Michigan
3
What is Tableau Software?
• Similar tools to Tableau include Microsoft
Power BI, Qlik, Tibco Spotfire, and Looker –
these are all data visualization tools
4
What is Tableau Software?
The main focus of Tableau software is
for you to better understand your
datasets, especially large datasets.
BI software in the past required
highly technical IT skills and took a
long time to build dashboards.
Tableau has changed that paradigm.
Courtesy: [Link]
Tableau invests a lot of research time
into developing intuitive software.
They approach software design from
the human perspective.
5
Benefits for Researchers & Teachers
• Free course licenses for
students
• Pre-built curriculum for
teaching Tableau and data
analysis
• Use of powerful ‘big’ data
platform for large datasets
• Provides skills needed in
industry (various professions)
[Link]
6
Benefits for Researchers
• Ability to handle ‘big’ data
(hundreds of millions of rows)
that Excel cannot
• Ability to share (link) your
research articles to datasets
and results through Tableau
Public
• Access to online help forums
& local users groups
• Ability to connect to “R” and
Python for more advanced
analytics and analysis
7
What is Data
Visualization?
8
What is Data Visualization
What is the Purpose of Data Visualizations?
Drive
Inform Persuade Entertain
Action
Communicate
What guides the design process?
How do we judge success?
9
What is Data Visualization?
Matthew Fontaine Maury
• Unfit for duty due to a leg injury
• Sent to Depot of Charts and Instruments
• Vault of logs from every ship in US Navy
• Hundreds of thousands of observations available in
written logs
• Manual ‘data mining’ with his team
Standardized collection moving forward (form)
[Link]
Matthew_Fontaine_Maury •
Ref. (The Clipper Ships – Time Life Books)
Ref. (Wind & Current Charts -1847)
10
What is Data Visualization?
Wind & Current Charts - 1847
• Visualization of his team’s findings
• Use of symbols and colors to highlight best routes
• Findings were counter-intuitive (heading west to go
faster east)
Results
• Roundtrip from Virginia to Rio 75 days instead of 112 days
• Found the Gulf Stream’s full shape
• Cut time from Cape Horn to California by a third
• Reduced ship lost due to storms
11
What is Data Visualization
A Basic Framework – Rhetoric for Data Visualization Methodology
• Who will be using the tool? 1. Identify Purpose (Intended Use)
• What level in the organization? 2. Consider Audience
• Strategic, tactical, operational? 3. Research
• Multiple user types?
• Global?
i. Identify Available Datasets
ii. Identify Data Elements
iii. Benchmark Designs
• Informative, persuasive 4. Design
• What action will result? i. Sketch
• Guided, static, decision support ii. Iterate
iii. Collect Feedback
5. Execute Design
i. Collect Feedback
6. Document – Deploy
• Summary data (<10,000 records) • Microsoft Excel, PowerPoint 7. Sustain
• 1 million records ? • Adobe Illustrator
• 10 million records ? • Tableau, Qlikview, MSBI
• “Big Data” ? • SAS Visual Analytics
12
What is Data Visualization
Example – Decision Support
Those looking to
catch big fish in
Michigan
Provide decision
Michigan DNR
support to
Database;
increase chances
Public Use
of catching big
Pictures
fish
Tableau
13
What is Data Visualization
Consistent Color (lack)
Elements of Design - Unity
Unity is the application of Simplified
methods that ensure that Images
elements in the design
appear to ‘go together’ -
(color, font, & shape
consistency)
Consistent
Font
[Link] Author: Mike Moore
14
What is Data Visualization
Elements of Design - Hierarchy
Level 1
Hierarchy is the Level 2
application of design Level 3
methods to indicate
importance and ‘flow’ Level 4
within the visual (size,
placement)
[Link]
precipitation Author: Matt Chambers
15
Elements of Design - Color
Use of color provides
contrast for data points
in opposition and brings
attention to relevant
elements within the
visual.
[Link]
Author: Oliver Linder
16
What is Data Visualization
Elements of Design – Balance
& Alignment
Balance and alignment
are used to create
harmonious visuals that Alignment
do not distract from the
message being
communicated. Balance
[Link] Author: George Gorczynski
17
What is Data Visualization
Elements of Design – Grouping / Spacing
Grouping and spacing
can be used to associate
similar elements and
provide a narrative or
visual flow within the
visualization.
[Link] Author: Shine Pulikathara
18
What is Data Visualization
The Iterative Design Process
19
What is Data Visualization
Detailed Example - Design
Hierarchy
Grouping
Grouping
Balance
20
Now . . .
Back to Tableau
General Overview
21
Tableau – General Overview
• All worksheets &
dashboards start with data Files
(Excel, CSV,
• Tableau connects to almost JSON, SAS…)
every type of data file
imaginable
• You can join across
Servers
different type of data (Databases)
sources!
22
Tableau – General Overview – simple example
• A simple table with 15 rows
of data in an Excel
spreadsheet
• Build an interactive
dashboard in under three
minutes
23
24
Tableau – General Overview
Calculated Fields
25
Tableau – General Overview
Basic Analytics
26
Tableau – General Overview: Bringing it all together
• Many different
worksheets, text boxes, Text Box
parameters, and filters Text Box Parameter
come together to create
Worksheet
a dashboard
#3
Text Box
• Multiple dashboards can
Worksheet Worksheet #6
be ‘chained’ together so
#2
that users are guided
through multiple
analytical paths Worksheet #5
Text Box
Worksheet Worksheet #4
#1
27
Use for Reporting
- Examples
28
Tableau – Reporting Example
• The results of detailed statistical analysis can
be made available freely on Tableau Public
where individuals can interact with data
visualizations to view results – to supplement
published research or publicly available
reports
• Expands the audience for consuming research
and provides a visual and interactive
experience.
[Link]
policy-changes-snap?tab=featured&type=featured 29
[Link]
Tableau – Reporting Example
• Story Points – (a Tableau feature)
provides a user experience similar to
PowerPoint but with interactive data
visualizations
• This allows for guided analytics where
you create a general narrative and
allow users to interact with
visualizations to ‘deep dive’ into key
points.
30
Use for
Storytelling
- Examples
31
Tableau – Storytelling Example (Story Points)
32
Tableau – Storytelling Example (K-MAX)
33
Advanced Features
- Examples
34
Advanced Features – Connecting Tableau to “R”
• Step #1
• Install “R” or “R” Studio on
your computer
• Load the Rserve library
package
• Start Rserve
35
Advanced Features – Connecting Tableau to “R”
• Step #2
• Connect Tableau
to your Rserve
instance
36
Advanced Features – Connecting Tableau to “R”
INT(SCRIPT_Str("library(xml2);
• Step #3 dater <- [Link]([Link]()-.arg2);
year <- paste('year_', format(dater, '%Y'), '/', sep = '');
• Write “R” script month <- paste('month_', format(dater, '%m'), '/', sep = '');
within a day <- paste('day_', format(dater, '%d'), '/', sep = '');
calculated field in xmlFile <-
Tableau paste('[Link] year,
month, day, '[Link]', sep = '');
x <- read_xml(toString(xmlFile));
Note: This is also games=xml_children(x);
generally the same way to ns <- xml_ns(x);
connect Tableau to awayruns <-xml_attr(games,'away_team_runs',ns);
Python in Anaconda – awayrunsdf <- [Link](awayruns);
with a few small awayrunsdf$ID <- [Link](nrow(awayrunsdf));
configuration differences. toString(awayrunsdf[.arg1, 1]);
",MAX([Idvalue]),max([zz_date])))
37
Advanced Features – Example
• Example that queries Major
League Baseball’s open API for
statistics
• “R” script downloads data as an
XML file, parses the data and
returns the results to Tableau for
visualization.
38
Available Resources
39
Books
The Functional Art Envisioning Information Design Basics Index
Alberto Cairo Edward Tufte Jim Krause
Information Beautiful Evidence
Dashboard Design Visual Explanations Edward Tufte
Stephen Few Edward Tufte
Universal Principles Information Design
of Design The Visual Display of Workbook
William Lidwell Quantitative Kim Baer
Information
Edward Tufte
40
Tableau Public & Other Resources
[Link]
• Daily inspiration through ‘viz of
the day’
• A place to upload your work to
the cloud
• Open environment to share
visualizations and data (don’t
post confidential data here ☺ )
[Link]
[Link]
[Link]
[Link]
[Link]
National Geographic Magazine
Bloomberg Businessweek
41
West Michigan Tableau Users Group (WMTUG)
[Link]
• Meet three to four
times a year in
Kalamazoo or Grand
Rapids
• 100-150 participants
• Sharing tips, tricks, and
case studies
• Develops a strong
network with other
analytics focused
individuals
42
Tableau Conference
• 15,000 of your best
data visualization
friends in the same
place
• One week of in-depth
sessions on data
visualization and
Tableau software
43
lOMoARcPSD|63293293
Exp-8
computer science (Chaitanya Bharathi Institute of Technology)
messages.pdf_cover_qr_code_label
messages.studocu_not_sponsored_or_endorsed_by_college
messages.downloaded_by
lOMoARcPSD|63293293
Aim:
Creating Dashboards & Storytelling, creating your first dashboard and
Story, Design for different displays, adding interactivity to your
Dashboard, Distributing & Publishing your Visualization.
Before you start creating the dashboard with Tableau, it is a good practice to design it by
hand with pen and paper. That means you must story down first. You should have a clear
picture of what you want to show on the dashboard before you open up and start designing
the visualization.
In this dashboard, I am going to show four things:
1. The Sales numbers by Category.
2. Sales numbers over time.
3. Performance of Sales by States
4. Sales by City
So, these are the four visualizations that I want on my superstore sales dashboard. Let’s start
creating the dashboard with Tableau.
Sheet 1: Sales by Category
The first visualization is sales by category. For this, I will select Category and Sales from the
data pane and drop them in Row and Column shelf respectively.
Remember to rename all the sheets, as they come in very handy while designing a dashboard.
Sheet 2: Sales Overtime
messages.downloaded_by
lOMoARcPSD|63293293
The second sheet is for Sales over time. That means how the sales number performed over
the years or months. For this, select Sales and drag to the Rows and drop Order Date into the
columns.
To move this visualization to a more granular level. Click on the + on year in the columns
shelf. On clicking first, the Quarter will appear and on another click on the Quarter, Month
will appear. Now drop the Year and Quarter. It will create a Month view of sales.
But what I want is Month wise sales comparison over the years. That means how was the
performance in January 2016 in comparison to January 2017. To do this, pick Order Date
from the data pane and drop it on the colors card.
Sheet 3: Sales Across States
Now let’s create our third sheet. In this chart, we will represent the sales distribution across
the states. That means it’s going to be a geospatial analysis.
Select Sales and State by pressing ctrl+ click and click on show me. Out of the recommended
visualizations select the map. Click on the 49 unknown messages. Go to “Edit Locations”
messages.downloaded_by
lOMoARcPSD|63293293
then go to country/Region, select “from field”. The United States will appear itself and click
ok. and we have Sales by State.
Sheet 4: Sales by City
Finally, let’s create our fourth visualization. Here, we will show Sales by City distribution.
similar to the last chart, select Sales and City and click on the show me. Then click on the
recommended visualization. Again we have 530 unknowns, click on that then edit locations.
Go to the country and click on them from the field. You will see the United States poped out,
click OK.
Here, we have city sales numbers. For a better idea of cities, select State from the data pane
and drop it on the details card.
Now our all sheets are ready, let’s create the Dashboard.
Create a Dashboard
To create a new dashboard click on the icon given on the bottom bar for showing the text ”
New Dashboard”. When your dashboard appears, you can see all the worksheets you have
created on a side panel.
In a dashboard, you have an option to change the view of the dashboard. It can be a default or
a phone view. I don’t like the current half view of my dashboard, to change this, click on the
size drop-down and select automatically. It may seem a small thing but trust, it can break
your entire dashboard experience.
messages.downloaded_by
lOMoARcPSD|63293293
Now just drag and drop all the sheets from the sidebar to the dashboard. Here, you have your
first dashboard.
Still, there are some small things that you can work with, As I don’t like the space we left on
the right due to the legends. We can utilize that space creatively. Tableau allows us to either
completely remove these legends or move them somewhere else by making them floating.
To remove the sales legend as it’s not contributing anything to the dashboard, click on the
cross icon and it will be gone. For others, click on the legend and then to the arrow and select
floating as shown below. Now you can drag and drop it anywhere you want on the dashboard.
Similarly, you can move the Year of order date to the respective visualization. So we have
our superstore sales number ready for the leadership to infer and plan accordingly. Here is the
final dashboard.
messages.downloaded_by
lOMoARcPSD|63293293
Create a Story
Let’s see the various steps required to create a Story in Tableau. This story uses the
Superstore data set that is available as a sample on Tableau Desktop.
Step 1: Click on the new Story tab to create a new story. You can then add various sheets
and dashboards to create a story point.
Step 2: You can double-click on the sheets and dashboards on the left to add them to a
story point. You can also drag the sheets into your story point on the Tableau desktop. All
the sheets and dashboards that are added to a story are connected to their original forms. So
any changes made to the original sheets or dashboards are reflected in the story. For
example, let’s add a dashboard containing the relation between Discounted Sales and Profit
by Category to the story.
messages.downloaded_by
lOMoARcPSD|63293293
Step 3: We can also add a caption to summarize the story point by clicking on “Add a
caption” and then writing it. Let’s add the caption “Relation between Discounted Sales and
Profit by Category and Subcategory” to our example.
Step 4: It is possible to add another story point by 2 methods. You can either click on the
Blank tab to use a blank sheet for the next story point or click on the Duplicate tab to obtain
a duplicate sheet as the current story point. Let’s click on the blank option.
Step 5: You can change the size of your story by clicking on the Size option in the lower-
left corner. You can choose from one of the predefined sizes or set your custom size in
pixels. You can also change the name of your story by right-clicking on your Story tab and
choosing rename.
messages.downloaded_by
lOMoARcPSD|63293293
Step 6: Now, let’s see a complete story on the relationship between the discounted sales
and profit
messages.downloaded_by
lOMoARcPSD|63293293
Add interactivity
1. Select Profit Map in the dashboard, and click the Use as filter icon in the upper
right corner.
2. Select a state within the Southern region of the map.
messages.downloaded_by
lOMoARcPSD|63293293
The Sales in the South bar chart automatically updates to show just the sub-category
sales in the selected state. You can quickly see which sub-categories are profitable.
3. Click an area of the map other than the colored Southern states to clear your selection.
You also want viewers to be able to see the change in profits based on the order date.
4. Select the Year of Order Date filter, click its drop-down arrow, and select Apply to
Worksheets > Selected Worksheets.
5. In the Apply Filter to Worksheets dialog box, select All in dashboard, and then
click OK.
This option tells Tableau to apply the filter to all worksheets in the dashboard that use
this same data source.
Explore state performance by year with your new, interactive dashboard!
Check your work! Watch "Add interactivity" in action
Here, we filter Sales in the South to only items sold in North Carolina, and then explore year
by year profit.
Click the image to replay it
Rename and go
You show your boss your dashboard, and she loves it. She's named it "Regional Sales and
Profit," and you do the same by double-clicking the Dashboard 1 tab and typing Regional
Sales and Profit.
In her investigations, your boss also finds that the decision to introduce machines in the North
Carolina market in 2021 was a bad idea.
messages.downloaded_by
lOMoARcPSD|63293293
Your boss is glad she has this dashboard to explore, but she also wants you to present a clear
action plan to the larger team. She asks you to create a presentation with your findings.
Before you publish your workbook, make sure you know the following:
The name of the server and how you sign in to it. If your organization uses Tableau
Cloud, you can click the Quick Connect link.
Any publishing guidelines your Tableau administrator might have, such as the name
of the project you should publish to.
Publish your workbook
1. With the workbook open in Tableau Desktop, click the Share button in the toolbar.
If you aren’t already signed in to Tableau Server or Tableau Cloud, do so now. If you
don’t have a site yet, you can create one on Tableau Cloud.
2. In the Publish Workbook dialog box, select the project to publish to.
3. Name the workbook according to whether you’re creating a new one or publishing
over an existing one.
4. Under Data Sources, select Edit. For Authentication, select Allow refresh
access or Embed password.
For some data connections, only one authentication option appears. If None shows,
leave it set to that.
5. Click Publish.
If this is your first time publishing a workbook, test it on the server and work out any
glitches before letting other users know the workbook is available.
messages.downloaded_by
lOMoARcPSD|63293293
Comprehensive Overview of Data Visualisation & Intelligent
Systems
Enterprise Computing (Patrician Brothers' College Blacktown)
messages.pdf_cover_qr_code_label
messages.studocu_not_sponsored_or_endorsed_by_college
messages.downloaded_by
lOMoARcPSD|63293293
Comprehensive Guide to Data
Visualisation, Data Management, and
Intelligent Systems in Enterprise
Computing
Data Visualisation Definition and Types
Data visualisation involves the visual presentation of data to communicate
stories within the dataset through static, dynamic, and interactive
visualisations, making complex information easier to interpret.
Types of Data Visualisation
Static Visualisation: Uses graphs, charts, and maps to provide
snapshots of data at a point in time.
Dynamic Visualisation: Incorporates animations to emphasise key
information and show movements or changes over time.
Interactive Visualisation: Allows users to manipulate graphics—
changing variables, filtering data, zooming—which promotes active
exploration and deeper insights.
Purposes of Data Visualisation
Simplify Understanding: Transforms overwhelming datasets into
digestible visuals, enabling quick grasp of trends (e.g., scatter plots
revealing correlations).
Tell a Story: Connects multiple visuals to narrate insights, guiding
viewers through data-driven conclusions.
Highlight Significant Results: Draws attention to outliers or key findings
(e.g., heatmaps pinpointing high-traffic website areas) for focused
messages.downloaded_by
lOMoARcPSD|63293293
analysis.
Software Features Enhancing Data
Understanding
Spreadsheets (Excel, Google Sheets): Basic charting, pivot tables,
conditional formatting for quick summaries.
Creative Design Applications (Adobe Illustrator, Canva): Custom
infographics combining visuals, icons, and text.
Data Integration Tools: Combining datasets from multiple sources to
reveal hidden patterns.
Forecasting & Trend Lines: Software features like trend analysis predict
future behaviour (e.g., rising sales trends).
Identifying Patterns in Data
Visual comparison and interpretation help uncover:
Trends: e.g., seasonal sales fluctuations.
Relationships: e.g., income vs. education levels.
Outliers: e.g., unusually high crime rates in specific areas.
Correlations: e.g., social media sentiment linked to product sales.
These insights support strategic decisions, social interventions, and ethical
considerations.
Impact of Technology Evolution on Data
Analytics
Advances in hardware/software have transformed data analysis:
messages.downloaded_by
lOMoARcPSD|63293293
Processing Power: Faster CPUs and GPUs enable real-time, complex
computations.
Storage Capacity: Cloud and data lakes facilitate vast data
management.
Communication: High-speed internet supports seamless data sharing
globally.
Real-time Visualisation: Instantaneous insights from streaming data
improve responsiveness.
Effect on Data Analysis
Enables handling of big datasets.
Facilitates sophisticated visualisations and machine learning integration.
Drives operational efficiency and timely decision-making.
Online Analytical Processing (OLAP) and OLAP
Cubes
OLAP systems allow multidimensional data analysis:
Features: Drill down, roll up, slice and dice, pivot.
OLAP Cubes: Organise data into dimensions (e.g., time, region, product),
measures (sales, profit), and hierarchies (year > quarter > month).
Functionality
Pre-calculated aggregations for quick retrieval.
Enable exploration of data from multiple perspectives.
Support decision-making by revealing patterns over various dimensions.
Data Integrity in Visualisation Development
messages.downloaded_by
lOMoARcPSD|63293293
Ensuring data quality is vital:
Ownership & Source: Confirm data is from a reliable, authoritative
source.
Validation: Cross-check for errors, missing values, and consistency.
Risks: Inaccurate data leads to misleading visuals, faulty conclusions,
and poor decisions.
Mitigation: Data cleaning, validation processes, and critical evaluation
before visualisation.
Common Data Visualisation Misleading
Techniques
Truncated Y-Axis: Exaggerates differences; e.g., starting y-axis above
zero.
Cherry-Picking Data: Selective presentation; e.g., only successful
products.
Misleading Scales: Using non-linear axes; e.g., logarithmic scales
distorting perceptions.
Omitting Labels/Units: Causes misinterpretation; e.g., missing time
frames.
3D Effects: Distort proportions; e.g., 3D pie charts exaggerate slice sizes.
Visual Clutter: Overuse of colours, graphics distract from key info.
Manipulative Colours: Evoke emotions or biases; e.g., red for negative
data.
Critical analysis is essential to identify and counteract these techniques.
Enterprise Data Warehousing's Role in
Visualisation
messages.downloaded_by
lOMoARcPSD|63293293
Data warehouses centralise data from multiple sources, enabling:
Historical Analysis: Identify long-term trends (e.g., sales over years).
Current Data Integration: Correlate real-time and historical data (e.g.,
customer behaviour).
Data Quality: Standardisation and cleaning improve reliability.
Enhanced Decision-Making: Rich visuals reveal insights, support
operational and strategic choices.
Operational Efficiency & Customer Experience: Personalisation and
trend analysis improve service and loyalty.
Big Data's Impact on Visualisation Design
Handling massive datasets requires:
Scope Management: Sampling, aggregation, filtering, zooming.
Representation of Diverse Data Types: Hierarchical (treemaps),
relational (network graphs), spatial (maps), temporal (time-series).
Advanced Techniques: Machine learning for deeper insights, real-time
dashboards for live monitoring.
Challenges: Avoiding overload, managing bias, ensuring relevance.
Effective big data visualisation uncovers complex stories within vast,
varied datasets.
Evaluating Bias in Data Visualisation
Critical for ethical, accurate insights:
Data Accuracy: Verify source credibility, validate data, check for errors.
Audience & Cultural Sensitivity: Tailor visuals to audience literacy and
cultural context.
messages.downloaded_by
lOMoARcPSD|63293293
Selection Bias: Ensure representative sampling.
Designer Bias: Be aware of personal biases influencing visual choices.
Viewer Bias: Recognise cognitive biases like confirmation bias.
Countermeasures: Transparent data, balanced visuals, contextual
explanations.
This ensures visuals are trustworthy and ethically sound.
Software Tools for Data Visualisation
Basic: Spreadsheets (Excel, Google Sheets) for quick charts and
dashboards.
Intermediate: Presentation software (PowerPoint, Google Slides) for
storytelling.
Advanced: Business analytics platforms (Tableau, Power BI, Google Data
Studio) for interactive, real-time dashboards.
Custom Solutions: Tailored software for specific needs—highly flexible
but costly.
Selection Criteria: Data complexity, audience, budget, technical skill,
desired interactivity.
Choosing the right tool enhances clarity, engagement, and insight delivery.
Interrogating Data from Visualisations
Critical analysis involves:
Understanding Visual Type: What story does it tell? Trends,
relationships, distributions.
Assessing Aggregation & Filtering: How data is grouped or subsetted
influences conclusions.
messages.downloaded_by
lOMoARcPSD|63293293
Outlier Impact: Outliers can skew insights; determine if they are errors
or meaningful.
Reasoning & Context: Draw logical conclusions, considering data
limitations and purpose.
Question assumptions: Ensure visuals accurately reflect data and
avoid misinterpretation.
This promotes informed, ethical decision-making.
User Experience (UX) Principles in Data
Visualisation
Effective UX ensures visuals are:
Clear & Concise: Proper labels, legends, and minimal clutter.
Accessible: Suitable for diverse users, including those with disabilities.
Relevance: Tailored to audience needs and context.
Interactive & Customisable: Filters, drill-downs, real-time data.
Engaging: Visually appealing, intuitive, and informative.
Supports Decision-Making: Highlights key insights, facilitates
exploration.
Good UX leads to better comprehension and actionable outcomes.
Big Data and Predictive Analysis
Big data's volume, variety, and velocity enable:
Pattern Recognition: Detecting trends, correlations, anomalies.
Real-Time Insights: Dynamic visualisations for immediate decision-
making.
Applications: Fraud detection, traffic forecasting, demand prediction.
messages.downloaded_by
lOMoARcPSD|63293293
Challenges: Data privacy, bias, managing vast, unstructured data.
Predictive analytics enhances strategic planning and operational
efficiency.
Data Security Practices
Safeguarding data involves:
Access Control: User authentication, role-based permissions.
Encryption: Data protected in transit and at rest.
Backups: Regular, offsite copies to prevent data loss.
Network & Physical Security: Firewalls, secure facilities.
Employee Training: Raising awareness on threats and best practices.
Regular Assessments: Vulnerability scans, penetration testing.
Ensuring data integrity and privacy is critical for trust and compliance.
Quantitative vs. Qualitative Data
Quantitative Data: Numerical, measurable, analysed statistically (e.g.,
sales figures, test scores).
Qualitative Data: Descriptive, interpretive, thematic (e.g., customer
opinions, interview transcripts).
Use: Quantitative supports trend analysis; qualitative provides context
and insights into perceptions and experiences.
Both are essential for comprehensive understanding.
Levels of Measurement: Nominal, Ordinal,
Interval, Ratio
Nominal: Categorised without order (e.g., gender, country).
messages.downloaded_by
lOMoARcPSD|63293293
Ordinal: Ordered categories, unequal intervals (e.g., satisfaction ratings).
Interval: Equal intervals, no true zero (e.g., temperature in Celsius).
Ratio: Equal intervals with a true zero (e.g., height, income).
Understanding these levels guides appropriate data analysis techniques.
Data Sampling and Collection Methods
Sampling: Selecting subsets (random, stratified) to represent
populations efficiently.
Active Collection: Direct interaction (surveys, experiments).
Passive Collection: Indirect, background data (analytics, logs).
Manual vs. Computerised: Manual (interviews, observations) or
automated (sensors, online forms).
Ethics & Bias: Critical to ensure representativeness, privacy, and data
quality.
Effective sampling underpins valid, reliable insights.
Assessing Data Quality
Relevance: Data must address the research question.
Accuracy: Data free from errors; validated through cross-referencing.
Validity: Measures what it claims to measure.
Reliability: Consistent results over time and across methods.
Ensures data-driven decisions are trustworthy and sound.
Informatics Supporting Data Understanding
Databases & Data Warehouses: Organise and store large datasets.
Data Mining & Machine Learning: Discover patterns, predict outcomes.
messages.downloaded_by
lOMoARcPSD|63293293
Knowledge Representation: Use ontologies, semantic networks.
Decision Support: Facilitate informed, data-driven decisions.
Informatics transforms raw data into actionable knowledge.
Data Presentation Methods
Graphs & Charts: Trends, relationships, distributions.
Infographics: Summarise complex info visually.
Dashboards: Real-time, interactive summaries.
Reports: Detailed, structured analysis.
Network Diagrams & Maps: Show relationships and geographic data.
Selection depends on audience, purpose, and data type.
Structured, Semi-structured, and Unstructured
Data
Structured Data: Rigid, tabular, easily queried (e.g., relational
databases).
Semi-structured Data: Flexible, with tags or metadata (e.g., JSON, XML).
Unstructured Data: Diverse formats, complex to analyse (e.g., videos,
social media posts).
Understanding these types guides storage, processing, and analytics
strategies.
Alternative Feedback Data: Likes, Emoticons,
Memes
Likes: Quantitative approval, also reflects emotional response.
Emoticons: Indicate sentiment, clarify tone.
messages.downloaded_by
lOMoARcPSD|63293293
Memes: Cultural commentary, social insights.
Uses: Gauge public opinion, sentiment analysis, cultural trends.
Challenges: Subjectivity, cultural differences, potential manipulation.
These data sources provide rich, real-time feedback but require careful
interpretation.
Impact of Errors, Uncertainty, and Limitations
Data flaws—errors, biases, incomplete info—can lead to:
Misleading visualisations.
Faulty decisions.
Financial or reputational damage. Mitigation includes validation,
validation, and critical evaluation of data sources and methods.
Blockchain for Data Management and
Verification
Blockchain offers:
Decentralisation: No single point of control.
Immutability: Data cannot be altered retroactively.
Transparency: Shared ledger accessible to stakeholders.
Applications: Voting, supply chain, medical records, land titles.
Ensures data integrity, traceability, and trustworthiness.
Software Features Affecting Data Privacy &
Security
Autofill & Forms: Risk of data leaks.
messages.downloaded_by
lOMoARcPSD|63293293
Connection Types: Public Wi-Fi vs. secure VPNs.
Checkboxes & Terms: Unintentional data sharing.
Security Protocols: Encryption, strong authentication, regular updates.
User awareness and secure design are critical.
Big Data Characteristics
Volume: Massive data quantities.
Variety: Multiple formats—structured, semi-structured, unstructured.
Velocity: Rapid data generation.
Veracity: Data quality and trustworthiness. Handling these requires
scalable storage, fast processing, and advanced analytics.
Data Mining Risks and Benefits
Benefits: Discover patterns, improve decision-making, personalise
services.
Risks: Privacy breaches, discrimination, misuse. Responsible practices
include anonymisation, transparency, and ethical guidelines.
Impact of Data Scale on Analytics
Large datasets enable:
Deep pattern recognition.
Machine learning applications.
Real-time insights. But also raise ethical issues—privacy, bias, and data
overload.
Data Storage Methods
Local Storage: Fast but limited capacity.
messages.downloaded_by
lOMoARcPSD|63293293
Cloud Storage: Scalable, remote, accessible.
Portable Media: Convenient but less secure.
Data Warehouses: Centralised, supports analytics, costly to maintain.
Choice depends on data volume, security needs, and access
requirements.
Ethical, Social, and Legal Issues
Bias & Discrimination: Algorithmic bias from training data.
Privacy: Data collection, consent, and control.
Transparency & Accountability: Explainability of AI decisions.
Legal Frameworks: GDPR, privacy laws.
Cultural Sensitivity: Respect for indigenous data rights.
Responsible data use ensures fairness, trust, and compliance.
Data Literacy & Social Impact
Critical for evaluating data credibility.
Helps detect manipulation.
Guides ethical and informed decisions.
Reduces susceptibility to misinformation.
Education in data literacy empowers responsible data use.
Spreadsheet Data Summarisation & Analysis
Aggregate data (sum, average, count).
Visualise with charts.
Use "what-if" scenarios.
messages.downloaded_by
lOMoARcPSD|63293293
Organise and filter data for clarity.
Build dashboards for quick insights.
Facilitates effective, quick decision-making.
Databases vs. Spreadsheets
Databases: Handle large, complex, relational data; support advanced
queries and security.
Spreadsheets: Suitable for small datasets, ad-hoc analysis, and quick
visualisations.
Choose based on data complexity and scalability needs.
Enterprise System Development & Project
Management
Problem Definition: Clear scope and goals.
Iterative Development: Continuous testing and refinement.
Tools: Gantt charts, flowcharts, decision trees, collaboration platforms.
Approach: Manage scope, time, cost, quality, risks, and stakeholders.
Team Roles: Clear responsibilities, communication, and documentation.
Structured planning ensures successful deployment.
Computational, Design, and Systems Thinking
Decomposition: Break down complex problems.
Pattern Recognition: Identify trends for decision-making.
Abstraction: Focus on core features, ignore details.
Algorithms: Step-by-step procedures for tasks.
messages.downloaded_by
lOMoARcPSD|63293293
Design Thinking: Human-centered approach, empathetic, iterative,
prototype, test.
Systems Thinking: Holistic view, interconnections, feedback loops,
emergent properties.
These skills underpin effective enterprise system design.
System Implementation & Testing
Planning: Feasibility, risk, stakeholder input.
Testing Methods: Functional, volume, beta, acceptance, regression,
security.
Validation: Ensure system meets requirements.
Evaluation: Monitor performance, gather user feedback.
Maintenance: Corrective, adaptive, perfective, preventative.
Continuous Improvement: Regular updates, security patches, user
training.
Ensures longevity, performance, and relevance.
Decision Support & Expert Systems
DSS: Aid human decision-making with data, models, scenario analysis.
Expert Systems: Mimic human expertise via knowledge bases,
inference engines, and user interfaces.
Knowledge Representation: Rules, frames, semantic networks.
Inference Techniques: Forward/backward chaining, truth maintenance,
hypothetical reasoning, fuzzy logic, ontologies.
Applications: Diagnosis, monitoring, process control, scheduling.
Core to automating complex, knowledge-intensive tasks.
messages.downloaded_by
lOMoARcPSD|63293293
Hardware in Intelligent Systems
Biometrics: Fingerprint, facial, iris, voice recognition.
Haptics: Tactile feedback in devices.
Touch & Gesture: Touchscreens, gesture recognition.
VR/AR: Immersive environments, overlays.
Microcontrollers: Arduino, Raspberry Pi for embedded control.
Sensors & Actuators: Temperature, proximity, motors for automation.
These components enable physical interaction and control.
Computational Thinking in System Design
Decomposition: Divide problems into manageable parts.
Pattern Recognition: Detect trends for predictive insights.
Abstraction: Simplify complexity.
Algorithms: Create procedures for automation and decision-making.
Essential for developing intelligent, efficient systems.
Flowcharts & Data Flow Diagrams
Flowcharts: Visualise process steps, decisions, control flow.
DFDs: Map data movement, sources, storage, transformations.
Use: Clarify system logic, identify gaps, facilitate communication.
Aid in designing, analysing, and refining intelligent systems.
Disruptive Effects of Intelligent Systems
Multitasking vs. Distraction: Rapid task-switching reduces
productivity (~40%), impairs memory.
messages.downloaded_by
lOMoARcPSD|63293293
Changing Work Practices: Automate and augment tasks, transforming
roles.
Automation in Manufacturing: Hard, programmable, flexible
automation improves efficiency but may cause job displacement.
Impact on Employment: Automation displaces routine jobs; creates
new roles in AI, data analysis, and system management.
Societal Changes: Need for continuous upskilling, managing ethical
issues like bias and privacy.
Understanding these effects helps adapt strategies for future work
environments.
Social & Ethical Issues of Intelligent Systems
Algorithmic Bias: Perpetuates societal inequalities; e.g., facial
recognition inaccuracies.
Privacy: Data collection, consent, and control concerns.
Transparency & Accountability: Explaining AI decisions, assigning
responsibility.
Legal & Cultural Considerations: Regulations like GDPR, indigenous
data rights.
Societal Impact: Discrimination risks, surveillance, loss of autonomy.
Responsible development ensures equitable, trustworthy AI deployment.
Technologies Underpinning Intelligent Systems
AI & Machine Learning: Automate tasks, identify patterns.
Deep Learning & Neural Networks: Model complex non-linear
relationships.
Natural Language Processing: Enable human-like communication.
messages.downloaded_by
lOMoARcPSD|63293293
Generative AI: Create new content (text, images, video).
Webometrics: Leverage web data for probabilistic inference.
Edge Computing: Process data locally for efficiency.
Data Mining & Data Modelling: Extract insights, structure data.
These technologies drive the capabilities of modern intelligent systems.
Enterprise Value from Intelligent Systems
Operational Efficiency: Automate repetitive tasks, optimise workflows.
Cost Reduction: Minimise waste, reduce errors.
Customer Personalisation: Tailored experiences (e.g., Netflix
recommendations).
Enhanced Decision-Making: Data-driven insights.
Competitive Edge: Innovation, agility, market responsiveness.
Case studies:
Netflix: Personalised content boosts engagement.
Amazon: AI-driven supply chain reduces costs.
Retail: Demand forecasting improves inventory management.
Strategic integration of AI techniques creates measurable business value.
Components of an IoT Enterprise Network
Servers: Central processing and management.
Local & Cloud Storage: Immediate and long-term data handling.
End-Point Devices: Sensors, actuators, smart gadgets.
Communication Links: Wi-Fi, Ethernet, Zigbee, LoRaWAN, Cellular—
chosen based on range, bandwidth, power needs.
messages.downloaded_by
lOMoARcPSD|63293293
Data Lifecycle: Collection, storage, processing, application,
transmission.
Ensures seamless, real-time data flow for intelligent decision-making.
Data Lifecycle & Relevance
Collection: Sensors gather raw data.
Type: Structured (organized), unstructured (media), semi-structured
(JSON).
Storage: On-device, local, or cloud.
Processing: Filtering, aggregation, analysis.
Application: Automation, monitoring, insights.
Transmission: Wired, wireless, cellular.
Relevance vs. Surplus: Focus on essential data; discard noise to
optimise resource use.
Efficient data management supports accurate, timely insights.
Simulation & Data Modelling
Simulation: Virtual testing of real-world scenarios (e.g., disaster
planning, medical training).
Data Modelling: Structuring data relationships (e.g., student info, supply
chain).
Automation: Reduces human effort, increases safety (e.g., factory
robots, autonomous vehicles).
Supports risk assessment, process optimisation, and decision support.
Intelligent Systems in Surveillance
messages.downloaded_by
lOMoARcPSD|63293293
Video Analysis: AI-driven CCTV detects anomalies, recognises
individuals.
Biometrics: Fingerprints, facial, iris for access control.
Customer Loyalty & Business Analytics: Tracking purchases for
insights.
Fraud Detection: Real-time transaction monitoring.
Network Sniffing & Trolling: Data interception and analysis for security
and intelligence.
Enhances security, operational efficiency, and data-driven insights.
AI Supporting IoT Efficiency
Edge Computing: Local processing reduces latency.
Predictive Maintenance: Anticipate failures via sensor data.
Resource Management: Optimise energy, water, and materials.
Pattern Recognition: Detect anomalies, optimise operations (e.g., traffic
flow, energy use).
Maximises system responsiveness and reduces operational costs.
Developing Facts, Rules, and Conclusions for
Expert Systems
Facts: Input data about the user or environment.
Rules: IF-THEN statements representing expert knowledge.
Certainty Factors: Quantify confidence (e.g., +0.8 for high likelihood).
Inference Engine: Applies rules to facts, combines evidence, and
produces conclusions.
messages.downloaded_by
lOMoARcPSD|63293293
Example: Student profile analysis for university course
recommendation, with rules like:
IF High Distinction in Math AND interest in problem-solving AND field
is Technology THEN recommend Bachelor of Engineering.
Certainty Factors allow nuanced conclusions, e.g., confidence levels in
recommendations.
Verifying Data Sources in Decision Support
Systems
Why Verify? Ensures decisions are based on accurate, credible data,
avoiding costly errors.
How to Verify:
Internal Data: Check for completeness, consistency, accuracy, and
recency.
External Data: Assess source credibility, methodology, and
timeliness.
Examples:
Sales data validation by cross-referencing reports.
Market reports from reputable agencies.
Customer feedback sampling and sentiment validation.
Consequences of Unverified Data: Faulty strategies, financial loss,
reputational damage.
Rigorous source verification underpins trustworthy decision-making.
Developing IF-Then Rules via Flowcharts
Purpose: Visualise decision logic before coding.
messages.downloaded_by
lOMoARcPSD|63293293
Process:
Define the problem.
Interview experts.
Map decision points as diamonds (questions).
Connect with arrows.
Convert decision paths into IF-THEN rules.
Example: Laptop troubleshooting flowchart translating into rules:
IF (Laptop Off) THEN check power.
IF (No display) THEN connect external monitor.
IF (External works) THEN internal display fault, etc.
Flowcharts clarify logic and facilitate validation.
Designing & Modelling Automated Smart
Systems
Steps:
Define problem & goals: e.g., energy-efficient climate control.
Identify inputs: sensors, user preferences, external data.
Specify outputs: actuator controls, alerts.
Architect system: sensors, edge devices, cloud, actuators.
Model logic: rules, algorithms, ML models.
Plan data flow: collection, processing, storage.
Design UI: apps, voice control, dashboards.
Implement security: encryption, authentication.
Test & evaluate: functional, performance, user feedback.
messages.downloaded_by
lOMoARcPSD|63293293
Iterative Process: Revisit steps to refine functionality, usability, and
robustness.
This systematic approach ensures reliable, effective intelligent systems.
This comprehensive overview synthesises core concepts, techniques, and
applications in data visualisation, data management, and intelligent
systems, equipping students with the understanding necessary for
advanced enterprise computing.# Comprehensive Guide to Data
Visualisation, Data Management, and Intelligent Systems in Enterprise
Computing
Data Visualisation Definition and Types
Data visualisation involves the visual presentation of data to communicate
stories within the dataset through static, dynamic, and interactive
visualisations, making complex information easier to interpret.
Types of Data Visualisation
Static Visualisation: Uses graphs, charts, and maps to provide
snapshots of data at a point in time.
Dynamic Visualisation: Incorporates animations to emphasise key
information and show movements or changes over time.
Interactive Visualisation: Allows users to manipulate graphics—
changing variables, filtering data, zooming—which promotes active
exploration and deeper insights.
Purposes of Data Visualisation
Simplify Understanding: Transforms overwhelming datasets into
digestible visuals, enabling quick grasp of trends (e.g., scatter plots
revealing correlations).
messages.downloaded_by
lOMoARcPSD|63293293
Tell a Story: Connects multiple visuals to narrate insights, guiding
viewers through data-driven conclusions.
Highlight Significant Results: Draws attention to outliers or key findings
(e.g., heatmaps pinpointing high-traffic website areas) for focused
analysis.
Software Features Enhancing Data
Understanding
Spreadsheets (Excel, Google Sheets): Basic charting, pivot tables,
conditional formatting for quick summaries.
Creative Design Applications (Adobe Illustrator, Canva): Custom
infographics combining visuals, icons, and text.
Data Integration Tools: Combining datasets from multiple sources to
reveal hidden patterns.
Forecasting & Trend Lines: Software features like trend analysis predict
future behaviour (e.g., rising sales trends).
Identifying Patterns in Data
Visual comparison and interpretation help uncover:
Trends: e.g., seasonal sales fluctuations.
Relationships: e.g., income vs. education levels.
Outliers: e.g., unusually high crime rates in specific areas.
Correlations: e.g., social media sentiment linked to product sales.
These insights support strategic decisions, social interventions, and ethical
considerations.
messages.downloaded_by
lOMoARcPSD|63293293
Impact of Technology Evolution on Data
Analytics
Advances in hardware/software have transformed data analysis:
Processing Power: Faster CPUs and GPUs enable real-time, complex
computations.
Storage Capacity: Cloud and data lakes facilitate vast data
management.
Communication: High-speed internet supports seamless data sharing
globally.
Real-time Visualisation: Instantaneous insights from streaming data
improve responsiveness.
Effect on Data Analysis
Enables handling of big datasets.
Facilitates sophisticated visualisations and machine learning integration.
Drives operational efficiency and timely decision-making.
Online Analytical Processing (OLAP) and OLAP
Cubes
OLAP systems allow multidimensional data analysis:
Features: Drill down, roll up, slice and dice, pivot.
OLAP Cubes: Organise data into dimensions (e.g., time, region, product),
measures (sales, profit), and hierarchies (year > quarter > month).
Functionality
Pre-calculated aggregations for quick retrieval.
messages.downloaded_by
lOMoARcPSD|63293293
Enable exploration of data from multiple perspectives.
Support decision-making by revealing patterns over various dimensions.
Data Integrity in Visualisation Development
Ensuring data quality is vital:
Ownership & Source: Confirm data is from a reliable, authoritative
source.
Validation: Cross-check for errors, missing values, and consistency.
Risks: Inaccurate data leads to misleading visuals, faulty conclusions,
and poor decisions.
Mitigation: Data cleaning, validation processes, and critical evaluation
before visualisation.
Common Data Visualisation Misleading
Techniques
Truncated Y-Axis: Exaggerates differences; e.g., starting y-axis above
zero.
Cherry-Picking Data: Selective presentation; e.g., only successful
products.
Misleading Scales: Using non-linear axes; e.g., logarithmic scales
distorting perceptions.
Omitting Labels/Units: Causes misinterpretation; e.g., missing time
frames.
3D Effects: Distort proportions; e.g., 3D pie charts exaggerate slice sizes.
Visual Clutter: Overuse of colours, graphics distract from key info.
Manipulative Colours: Evoke emotions or biases; e.g., red for negative
data.
messages.downloaded_by
lOMoARcPSD|63293293
Critical analysis is essential to identify and counteract these techniques.
Enterprise Data Warehousing's Role in
Visualisation
Data warehouses centralise data from multiple sources, enabling:
Historical Analysis: Identify long-term trends (e.g., sales over years).
Current Data Integration: Correlate real-time and historical data (e.g.,
customer behaviour).
Data Quality: Standardisation and cleaning improve reliability.
Enhanced Decision-Making: Rich visuals reveal insights, support
operational and strategic choices.
Operational Efficiency & Customer Experience: Personalisation and
trend analysis improve service and loyalty.
Big Data's Impact on Visualisation Design
Handling massive datasets requires:
Scope Management: Sampling, aggregation, filtering, zooming.
Representation of Diverse Data Types: Hierarchical (treemaps),
relational (network graphs), spatial (maps), temporal (time-series).
Advanced Techniques: Machine learning for deeper insights, real-time
dashboards for live monitoring.
Challenges: Avoiding overload, managing bias, ensuring relevance.
Effective big data visualisation uncovers complex stories within vast,
varied datasets.
Evaluating Bias in Data Visualisation
Critical for ethical, accurate insights:
messages.downloaded_by
lOMoARcPSD|63293293
Data Accuracy: Verify source credibility, validate data, check for errors.
Audience & Cultural Sensitivity: Tailor visuals to audience literacy and
cultural context.
Selection Bias: Ensure representative sampling.
Designer Bias: Be aware of personal biases influencing visual choices.
Viewer Bias: Recognise cognitive biases like confirmation bias.
Countermeasures: Transparent data, balanced visuals, contextual
explanations.
This ensures visuals are trustworthy and ethically sound.
Software Tools for Data Visualisation
Basic: Spreadsheets (Excel, Google Sheets) for quick charts and
dashboards.
Intermediate: Presentation software (PowerPoint, Google Slides) for
storytelling.
Advanced: Business analytics platforms (Tableau, Power BI, Google Data
Studio) for interactive, real-time dashboards.
Custom Solutions: Tailored software for specific needs—highly flexible
but costly.
Selection Criteria: Data complexity, audience, budget, technical skill,
desired interactivity.
Choosing the right tool enhances clarity, engagement, and insight delivery.
Interrogating Data from Visualisations
Critical analysis involves:
Understanding Visual Type: What story does it tell? Trends,
relationships, distributions.
messages.downloaded_by
lOMoARcPSD|63293293
Assessing Aggregation & Filtering: How data is grouped or subsetted
influences conclusions.
Outlier Impact: Outliers can skew insights; determine if they are errors
or meaningful.
Reasoning & Context: Draw logical conclusions, considering data
limitations and purpose.
Question assumptions: Ensure visuals accurately reflect data and
avoid misinterpretation.
This promotes informed, ethical decision-making.
User Experience (UX) Principles in Data
Visualisation
Effective UX ensures visuals are:
Clear & Concise: Proper labels, legends, and minimal clutter.
Accessible: Suitable for diverse users, including those with disabilities.
Relevance: Tailored to audience needs and context.
Interactive & Customisable: Filters, drill-downs, real-time data.
Engaging: Visually appealing, intuitive, and informative.
Supports Decision-Making: Highlights key insights, facilitates
exploration.
Good UX leads to better comprehension and actionable outcomes.
Big Data and Predictive Analysis
Big data's volume, variety, and velocity enable:
Pattern Recognition: Detecting trends, correlations, anomalies.
messages.downloaded_by
lOMoARcPSD|63293293
Real-Time Insights: Dynamic visualisations for immediate decision-
making.
Applications: Fraud detection, traffic forecasting, demand prediction.
Challenges: Data privacy, bias, managing vast, unstructured data.
Predictive analytics enhances strategic planning and operational
efficiency.
Data Security Practices
Safeguarding data involves:
Access Control: User authentication, role-based permissions.
Encryption: Data protected in transit and at rest.
Backups: Regular, offsite copies to prevent data loss.
Network & Physical Security: Firewalls, secure facilities.
Employee Training: Raising awareness on threats and best practices.
Regular Assessments: Vulnerability scans, penetration testing.
Ensuring data integrity and privacy is critical for trust and compliance.
Quantitative vs. Qualitative Data
Quantitative Data: Numerical, measurable, analysed statistically (e.g.,
sales figures, test scores).
Qualitative Data: Descriptive, interpretive, thematic (e.g., customer
opinions, interview transcripts).
Use: Quantitative supports trend analysis; qualitative provides context
and insights into perceptions and experiences.
Both are essential for comprehensive understanding.
messages.downloaded_by
lOMoARcPSD|63293293
Levels of Measurement: Nominal, Ordinal,
Interval, Ratio
Nominal: Categorised without order (e.g., gender, country).
Ordinal: Ordered categories, unequal intervals (e.g., satisfaction ratings).
Interval: Equal intervals, no true zero (e.g., temperature in Celsius).
Ratio: Equal intervals with a true zero (e.g., height, income).
Understanding these levels guides appropriate data analysis techniques.
Data Sampling and Collection Methods
Sampling: Selecting subsets (random, stratified) to represent
populations efficiently.
Active Collection: Direct interaction (surveys, experiments).
Passive Collection: Indirect, background data (analytics, logs).
Manual vs. Computerised: Manual (interviews, observations) or
automated (sensors, online forms).
Ethics & Bias: Critical to ensure representativeness, privacy, and data
quality.
Effective sampling underpins valid, reliable insights.
Assessing Data Quality
Relevance: Data must address the research question.
Accuracy: Data free from errors; validated through cross-referencing.
Validity: Measures what it claims to measure.
Reliability: Consistent results over time and across methods.
Ensures data-driven decisions are trustworthy and sound.
messages.downloaded_by
lOMoARcPSD|63293293
Informatics Supporting Data Understanding
Databases & Data Warehouses: Organise and store large datasets.
Data Mining & Machine Learning: Discover patterns, predict outcomes.
Knowledge Representation: Use ontologies, semantic networks.
Decision Support: Facilitate informed, data-driven decisions.
Informatics transforms raw data into actionable knowledge.
Data Presentation Methods
Graphs & Charts: Trends, relationships, distributions.
Infographics: Summarise complex info visually.
Dashboards: Real-time, interactive summaries.
Reports: Detailed, structured analysis.
Network Diagrams & Maps: Show relationships and geographic data.
Selection depends on audience, purpose, and data type.
Structured, Semi-structured, and Unstructured
Data
Structured Data: Rigid, tabular, easily queried (e.g., relational
databases).
Semi-structured Data: Flexible, with tags or metadata (e.g., JSON, XML).
Unstructured Data: Diverse formats, complex to analyse (e.g., videos,
social media posts).
Understanding these types guides storage, processing, and analytics
strategies.
messages.downloaded_by
lOMoARcPSD|63293293
Alternative Feedback Data: Likes, Emoticons,
Memes
Likes: Quantitative approval, also reflects emotional response.
Emoticons: Indicate sentiment, clarify tone.
Memes: Cultural commentary, social insights.
Uses: Gauge public opinion, sentiment analysis, cultural trends.
Challenges: Subjectivity, cultural differences, potential manipulation.
These data sources provide rich, real-time feedback but require careful
interpretation.
Impact of Errors, Uncertainty, and Limitations
Data flaws—errors, biases, incomplete info—can lead to:
Misleading visualisations.
Faulty decisions.
Financial or reputational damage. Mitigation includes validation,
validation, and critical evaluation of data sources and methods.
Blockchain for Data Management and
Verification
Blockchain offers:
Decentralisation: No single point of control.
Immutability: Data cannot be altered retroactively.
Transparency: Shared ledger accessible to stakeholders.
Applications: Voting, supply chain, medical records, land titles.
Ensures data integrity, traceability, and trustworthiness.
messages.downloaded_by
lOMoARcPSD|63293293
Software Features Affecting Data Privacy &
Security
Autofill & Forms: Risk of data leaks.
Connection Types: Public Wi-Fi vs. secure VPNs.
Checkboxes & Terms: Unintentional data sharing.
Security Protocols: Encryption, strong authentication, regular updates.
User awareness and secure design are critical.
Big Data Characteristics
Volume: Massive data quantities.
Variety: Multiple formats—structured, semi-structured, unstructured.
Velocity: Rapid data generation.
Veracity: Data quality and trustworthiness. Handling these requires
scalable storage, fast processing, and advanced analytics.
Data Mining Risks and Benefits
Benefits: Discover patterns, improve decision-making, personalise
services.
Risks: Privacy breaches, discrimination, misuse. Responsible practices
include anonymisation, transparency, and ethical guidelines.
Impact of Data Scale on Analytics
Large datasets enable:
Deep pattern recognition.
Machine learning applications.
messages.downloaded_by
lOMoARcPSD|63293293
Real-time insights. But also raise ethical issues—privacy, bias, and data
overload.
Data Storage Methods
Local Storage: Fast but limited capacity.
Cloud Storage: Scalable, remote, accessible.
Portable Media: Convenient but less secure.
Data Warehouses: Centralised, supports analytics, costly to maintain.
Choice depends on data volume, security needs, and access
requirements.
Ethical, Social, and Legal Issues
Bias & Discrimination: Algorithmic bias from training data.
Privacy: Data collection, consent, and control.
Transparency & Accountability: Explainability of AI decisions.
Legal Frameworks: GDPR, privacy laws.
Cultural Sensitivity: Respect for indigenous data rights.
Responsible data use ensures fairness, trust, and compliance.
Data Literacy & Social Impact
Critical for evaluating data credibility.
Helps detect manipulation.
Guides ethical and informed decisions.
Reduces susceptibility to misinformation.
Education in data literacy empowers responsible data use.
messages.downloaded_by
lOMoARcPSD|63293293
Spreadsheet Data Summarisation & Analysis
Aggregate data (sum, average, count).
Visualise with charts.
Use "what-if" scenarios.
Organise and filter data for clarity.
Build dashboards for quick insights.
Facilitates effective, quick decision-making.
Databases vs. Spreadsheets
Databases: Handle large, complex, relational data; support advanced
queries and security.
Spreadsheets: Suitable for small datasets, ad-hoc analysis, and quick
visualisations.
Choose based on data complexity and scalability needs.
Enterprise System Development & Project
Management
Problem Definition: Clear scope and goals.
Iterative Development: Continuous testing and refinement.
Tools: Gantt charts, flowcharts, decision trees, collaboration platforms.
Approach: Manage scope, time, cost, quality, risks, and stakeholders.
Team Roles: Clear responsibilities, communication, and documentation.
Structured planning ensures successful deployment.
Computational, Design, and Systems Thinking
messages.downloaded_by
lOMoARcPSD|63293293
Decomposition: Break down complex problems.
Pattern Recognition: Identify trends for decision-making.
Abstraction: Focus on core features, ignore details.
Algorithms: Step-by-step procedures for tasks.
Design Thinking: Human-centered approach, empathetic, iterative,
prototype, test.
Systems Thinking: Holistic view, interconnections, feedback loops,
emergent properties.
These skills underpin effective enterprise system design.
System Implementation & Testing
Planning: Feasibility, risk, stakeholder input.
Testing Methods: Functional, volume, beta, acceptance, regression,
security.
Validation: Ensure system meets requirements.
Evaluation: Monitor performance, gather user feedback.
Maintenance: Corrective, adaptive, perfective, preventative.
Continuous Improvement: Regular updates, security patches, user
training.
Ensures longevity, performance, and relevance.
Decision Support & Expert Systems
DSS: Aid human decision-making with data, models, scenario analysis.
Expert Systems: Mimic human expertise via knowledge bases,
inference engines, and user interfaces.
Knowledge Representation: Rules, frames, semantic networks.
messages.downloaded_by
lOMoARcPSD|63293293
Inference Techniques: Forward/backward chaining, truth maintenance,
hypothetical reasoning, fuzzy logic, ontologies.
Applications: Diagnosis, monitoring, process control, scheduling.
Core to automating complex, knowledge-intensive tasks.
Hardware in Intelligent Systems
Biometrics: Fingerprint, facial, iris, voice recognition.
Haptics: Tactile feedback in devices.
Touch & Gesture: Touchscreens, gesture recognition.
VR/AR: Immersive environments, overlays.
Microcontrollers: Arduino, Raspberry Pi for embedded control.
Sensors & Actuators: Temperature, proximity, motors for automation.
These components enable physical interaction and control.
Computational Thinking in System Design
Decomposition: Divide problems into manageable parts.
Pattern Recognition: Detect trends for predictive insights.
Abstraction: Simplify complexity.
Algorithms: Create procedures for automation and decision-making.
Essential for developing intelligent, efficient systems.
Flowcharts & Data Flow Diagrams
Flowcharts: Visualise process steps, decisions, control flow.
DFDs: Map data movement, sources, storage, transformations.
Use: Clarify system logic, identify gaps, facilitate communication.
messages.downloaded_by
lOMoARcPSD|63293293
Aid in designing, analysing, and refining intelligent systems.
Disruptive Effects of Intelligent Systems
Multitasking vs. Distraction: Rapid task-switching reduces
productivity (~40%), impairs memory.
Changing Work Practices: Automate and augment tasks, transforming
roles.
Automation in Manufacturing: Hard, programmable, flexible
automation improves efficiency but may cause job displacement.
Impact on Employment: Automation displaces routine jobs; creates
new roles in AI, data analysis, and system management.
Societal Changes: Need for continuous upskilling, managing ethical
issues like bias and privacy.
Understanding these effects helps adapt strategies for future work
environments.
Social & Ethical Issues of Intelligent Systems
Algorithmic Bias: Perpetuates societal inequalities; e.g., facial
recognition inaccuracies.
Privacy: Data collection, consent, and control concerns.
Transparency & Accountability: Explaining AI decisions, assigning
responsibility.
Legal & Cultural Considerations: Regulations like GDPR, indigenous
data rights.
Societal Impact: Discrimination risks, surveillance, loss of autonomy.
Responsible development ensures equitable, trustworthy AI deployment.
messages.downloaded_by
lOMoARcPSD|63293293
Technologies Underpinning Intelligent Systems
AI & Machine Learning: Automate tasks, identify patterns.
Deep Learning & Neural Networks: Model complex non-linear
relationships.
Natural Language Processing: Enable human-like communication.
Generative AI: Create new content (text, images, video).
Webometrics: Leverage web data for probabilistic inference.
Edge Computing: Process data locally for efficiency.
Data Mining & Data Modelling: Extract insights, structure data.
These technologies drive the capabilities of modern intelligent systems.
Enterprise Value from Intelligent Systems
Operational Efficiency: Automate repetitive tasks, optimise workflows.
Cost Reduction: Minimise waste, reduce errors.
Customer Personalisation: Tailored experiences (e.g., Netflix
recommendations).
Enhanced Decision-Making: Data-driven insights.
Competitive Edge: Innovation, agility, market responsiveness.
Case studies:
Netflix: Personalised content boosts engagement.
Amazon: AI-driven supply chain reduces costs.
Retail: Demand forecasting improves inventory management.
Strategic integration of AI techniques creates measurable business value.
Components of an IoT Enterprise Network
messages.downloaded_by
lOMoARcPSD|63293293
Servers: Central processing and management.
Local & Cloud Storage: Immediate and long-term data handling.
End-Point Devices: Sensors, actuators, smart gadgets.
Communication Links: Wi-Fi, Ethernet, Zigbee, LoRaWAN, Cellular—
chosen based on range, bandwidth, power needs.
Data Lifecycle: Collection, storage, processing, application,
transmission.
Ensures seamless, real-time data flow for intelligent decision-making.
Data Lifecycle & Relevance
Collection: Sensors gather raw data.
Type: Structured (organized), unstructured (media), semi-structured
(JSON).
Storage: On-device, local, or cloud.
Processing: Filtering, aggregation, analysis.
Application: Automation, monitoring, insights.
Transmission: Wired, wireless, cellular.
Relevance vs. Surplus: Focus on essential data; discard noise to
optimise resource use.
Efficient data management supports accurate, timely insights.
Simulation & Data Modelling
Simulation: Virtual testing of real-world scenarios (e.g., disaster
planning, medical training).
Data Modelling: Structuring data relationships (e.g., student info, supply
chain).
messages.downloaded_by
lOMoARcPSD|63293293
Automation: Reduces human effort, increases safety (e.g., factory
robots, autonomous vehicles).
Supports risk assessment, process optimisation, and decision support.
Intelligent Systems in Surveillance
Video Analysis: AI-driven CCTV detects anomalies, recognises
individuals.
Biometrics: Fingerprints, facial, iris for access control.
Customer Loyalty & Business Analytics: Tracking purchases for
insights.
Fraud Detection: Real-time transaction monitoring.
Network Sniffing & Trolling: Data interception and analysis for security
and intelligence.
Enhances security, operational efficiency, and data-driven insights.
AI Supporting IoT Efficiency
Edge Computing: Local processing reduces latency.
Predictive Maintenance: Anticipate failures via sensor data.
Resource Management: Optimise energy, water, and materials.
Pattern Recognition: Detect anomalies, optimise operations (e.g., traffic
flow, energy use).
Maximises system responsiveness and reduces operational costs.
Developing Facts, Rules, and Conclusions for
Expert Systems
Facts: Input data about the user or environment.
messages.downloaded_by
lOMoARcPSD|63293293
Rules: IF-THEN statements representing expert knowledge.
Certainty Factors: Quantify confidence (e.g., +0.8 for high likelihood).
Inference Engine: Applies rules to facts, combines evidence, and
produces conclusions.
Example: Student profile analysis for university course
recommendation, with rules like:
IF High Distinction in Math AND interest in problem-solving AND field
is Technology THEN recommend Bachelor of Engineering.
Certainty Factors allow nuanced conclusions, e.g., confidence levels in
recommendations.
Verifying Data Sources in Decision Support
Systems
Why Verify? Ensures decisions are based on accurate, credible data,
avoiding costly errors.
How to Verify:
Internal Data: Check for completeness, consistency, accuracy, and
recency.
External Data: Assess source credibility, methodology, and
timeliness.
Examples:
Sales data validation by cross-referencing reports.
Market reports from reputable agencies.
Customer feedback sampling and sentiment validation.
Consequences of Unverified Data: Faulty strategies, financial loss,
reputational damage.
messages.downloaded_by
lOMoARcPSD|63293293
Rigorous source verification underpins trustworthy decision-making.
Developing IF-Then Rules via Flowcharts
Purpose: Visualise decision logic before coding.
Process:
Define the problem.
Interview experts.
Map decision points as diamonds (questions).
Connect with arrows.
Convert decision paths into IF-THEN rules.
Example: Laptop troubleshooting flowchart translating into rules:
IF (Laptop Off) THEN check power.
IF (No display) THEN connect external monitor.
IF (External works) THEN internal display fault, etc.
Flowcharts clarify logic and facilitate validation.
Designing & Modelling Automated Smart
Systems
Steps:
Define problem & goals: e.g., energy-efficient climate control.
Identify inputs: sensors, user preferences, external data.
Specify outputs: actuator controls, alerts.
Architect system: sensors, edge devices, cloud, actuators.
Model logic: rules, algorithms, ML models.
Plan data flow: collection, processing, storage.
messages.downloaded_by
lOMoARcPSD|63293293
Design UI: apps, voice control, dashboards.
Implement security: encryption, authentication.
Test & evaluate: functional, performance, user feedback.
Iterative Process: Revisit steps to refine functionality, usability, and
robustness.
This systematic approach ensures reliable, effective intelligent systems.
This comprehensive overview synthesises core concepts, techniques, and
applications in data visualisation, data management, and intelligent
systems, equipping students with the understanding necessary for
advanced enterprise computing.
messages.downloaded_by
lOMoARcPSD|63293293
Data Visualization and Ethics
Ethical and Legal Issues in Business (Grand Canyon University)
messages.pdf_cover_qr_code_label
messages.studocu_not_sponsored_or_endorsed_by_college
messages.downloaded_by
lOMoARcPSD|63293293
Data Visualization and Ethics
Hayden Ellis
Colangelo College of Business, Grand Canyon University
BIT-301, Fundamentals in Business Analytics
Professor Canada
March 23, 2025
messages.downloaded_by
lOMoARcPSD|63293293
Introduction
This paper highlights the ethical challenges of selection bias and data privacy violations
in data analytics. These issues can distort truth, violate trust, and mislead decision-makers. When
data is mishandled or misrepresented, it can have far-reaching consequences. A biblical
worldview, grounded in integrity and transparency, encourages us to ensure data is accurate and
handled responsibly, in alignment with Psalms 90:17 and GCU’s Statement on the Integration of
Faith and Work.
Two Ethical Issues
Ethical Issue 1: Selection Bias
Selection bias occurs when data is collected in a way that overrepresents or
underrepresents certain groups. This can lead to incorrect conclusions and discriminatory
outcomes. For example, a Wall Street Journal article (Mitchell, 2023) examined how hiring
algorithms favored certain candidates based on biased training data, reinforcing systemic
inequalities. When companies rely on flawed data inputs, their decisions can unintentionally
exclude qualified individuals.
To combat this, ethical practitioners must actively review and adjust data sources to
ensure balanced representation. Psalms 90:17 reminds us to seek the Lord's favor and to ensure
the work of our hands is established with purpose and righteousness. Upholding fairness in data
collection is one way to fulfill this biblical principle. Faith integration reminds us that every
person holds God-given dignity and must not be unfairly excluded by faulty data practices.
messages.downloaded_by
lOMoARcPSD|63293293
Ethical Issue 2: Data Privacy Violations
Data privacy violations occur when personal information is collected, shared, or stored
without proper consent. A New York Times article (Singer, 2024) detailed how a fitness app
tracked user locations and sold the data to third parties without explicit user permission. This
undermines user trust and raises concerns about exploitation.
Faith and ethics guide us to treat others with respect and integrity. GCU’s Statement on
the Integration of Faith and Work calls on us to live out our Christian worldview in every
professional setting. Misusing someone’s private information contradicts biblical teachings on
stewardship and honoring others. Ethical data professionals must establish safeguards, remain
transparent, and always gain informed consent to protect user rights.
Conclusion
Ethical issues like selection bias and data privacy violations challenge the integrity of
data analytics. By applying faith-based principles of fairness, honesty, and respect, professionals
can ensure their work benefits others and glorifies God. The modeling of ethical behavior in data
handling reflects the commitment described in Psalms 90:17 to establish the work of our hands
with purpose and grace.
messages.downloaded_by
lOMoARcPSD|63293293
References
Mitchell, R. (2023). AI Hiring Tools Show Bias in Candidate Selection. Wall Street Journal.
[Link]
Singer, N. (2024). Your Workout App Is Selling Your Location. New York Times.
[Link]
Grand Canyon University. (n.d.). Statement on the Integration of Faith and Work.
[Link]
messages.downloaded_by
Data visualization for M&E practitioners
Starting shortly, please wait!
1
Meet your instructor
Eliza Avgeropoulou
Senior Monitoring and Evaluation Implementation
Specialist
BeDataDriven
2
Presented by the ActivityInfo Team
All in one information
management software for
humanitarian and development
operations.
● Track activities, outcomes
● Beneficiary management
● Surveys
● Work offline/online
3
BeDataDriven Mission
Provide the UN and NGOs with a standard, easy-to-use and
comprehensive data management platform so that as many
organizations as possible can become data-driven to achieve
better outcomes for rights holders worldwide.
BeDataDriven pursues this mission by building and
helping organizations implement ActivityInfo.
4
ActivityInfo
An end-to-end solution for M&E data management
Data collection Data management Data analysis
Easily collect the data you Organize your information Generate actionable insights
need from anywhere according to your workflow in real-time
….built on a relational data model 5
ActivityInfo is your integrated solution for managing your data across the data lifecycle.
Diagram adapted from Harvard Business Review 6
ActivityInfo Users
7
Outline
● Introduction
○ Importance of data visualization
● Principles of good data visualization
○ Understanding your target audience and the purpose
○ Choosing the right chart
○ Best practices for clarity and consistency
○ Identifying good vs. bad visualizations
● Data visualization examples
○ Analyzing real-world data visualization examples
● QandAs
Introduction
Introduction
"The profile of a curve reveals in a flash a whole situation — the life history of an epidemic, a panic, or an era of
prosperity. The curve informs the mind, awakens the imagination, convinces."
- Henry D. Hubbard, National Bureau of Standards
10
Importance
Explore Explain
Data visualization
Think
effectively
11
The starting point
How
03 Which is the
Who appropriate visual?
Understand
target audience
01 02 What
Are you trying to
communicate?
12
12
Starting point
M&E plan Data Model
✓ Identification of data needs Visual representation of:
✓ Identification of analysis ✓ Information flows from data
✓ Identification of reports collection to data use
needed ✓ Association amongst the Effective data
various data sources visualization
Identification of the correct
questions that we need to ask! Consistency across data
sources and higher data quality
13
Example
Program team report
M&E plan Data Model
Indicator: Number of ✓ Avoid double counting
registered participants ✓ Data source beneficiary
registration
✓ We collect daily ✓ Structure your data into
✓ We analyze monthly usable formats
✓ Program teams needs the
per month calculation. We
disaggregate internally per
partner
✓ Donor needs the quarter
Donor report
calculation.
14
Principles of good data visualization
Understand the audience and
purpose
16
Understanding the target audience and purpose
Country
team
HQ
Program
team
Audience Partner
Other
departments
Match your visualization to
Donor their needs and
Researchers understanding
17
Understanding the target audience and purpose
Audience level of Purpose
understanding
● Which stakeholders need to have timely information?Who are the
stakeholders?
● Do I need different reports depending on the audience?
● Why am I designing the report? (quarterly progress to donor? Yearly
progress to HQ? Monthly monitoring for field supervisors?)
18
Example
We created a report for the field coordinators based
on a survey that answers the question:
“How many beneficiaries had their basic need met as a
results of a cash distribution project?”
The visual that I choose needs to match my
audience level of understanding - do not
complicate it!
19
Example
SUDAN - Multi purpose cash assistance
Audience needs to understand quickly rather than
spending too much time!
20
Choosing the right graph
21
Choosing the right graph
Before we can create an effective visualization, we need to define what we’re trying to understand
What is the main question that you want to answer?
Choose a graph that matches your question!
22
Choosing the right chart
Relationship Data type
Main question
● Comparison per category?
● Track over time? ● Data that can be
● Correlate two or more counted or measured?
variables? ○ A range value?
● Data distribution? ○ Finite number of
● Compare a subset of data options?
to a whole amount? ● Data can be grouped
● Examine deviation? per category?
● Rank variable?
23
Common chart types
Two or more
Proportions and categories Over time Distribution
variables
Scatter plot
24
Pie charts
Pie charts work well for questions about proportions. They work best when all your categories sum to a meaningful whole 100%
E.g. What proportion of the total does each category represent?
Group age
Group age
Pie chart Donut chart
25
Bar charts
When our data question is about comparing discrete categories (distinct groups or types), a bar chart is often the best choice - no
more than 15 categories
E.g. How do different categories of X compare in terms of a value?
26
Stacked bar charts
Stacked bar charts work well when you need to show how different subcategories contribute to a total. Each bar clearly
represents the total value, with segments showing the contribution of each subcategory.
E.g. How do subcategories contribute to each category total?
27
Line plot
If a data question involves understanding how data changes over a continuous period, especially time, a line chart is a great
visualization. Line charts illustrate trends, patterns, or fluctuations.
E.g. How has beneficiary number changed over the years?
Distribution of beneficiaries across partners and years
P1
P2
P3
P4
P5
P6
P7
P8
P9
2000 2005 2010 2020
28
Histogram
Histograms are a great choice when we’re asking about the distribution or frequency of numerical data
E.g. What is the wage distribution amongst project participants?
Wage distribution
wage 29
Scatter plot
When we have a question about how two numeric variables relating to each other, we should immediately think of scatter plots.
E.g. How does [numeric variable A] relate to [numeric variable B]? value in the 20’ compared to the 2015’?
Wage gap
30
Bubble map
When we have geographical information and we wish to showcase the variation of numeric values across regions
E.g. How does [numeric variable A] varies per region?
Incidence of violation
per province
31
Clarity and consistency
32
Keep it clear and consistent
Carefully select only the data that will
support your clarity of intent, so that
your main message isn’t lost.
Keep your designs simple and clear
SUDAN - Multi purpose cash
assistance
33
Clarity and consistency
Inclusivity
Clarity and
Meaning consistency Culture
34
Color choice
Sufficient contrast and separation between elements Color to convey meaning - inclusivity
Color in culture
Pink: Feminine in West
But in Japan: equally used
for masculine and feminine
Storytelling with data
35
Labels and descriptions
Each data point has the callout for the amount so a user Clear text that labels the significant parts of the data
doesn't have to guess or rely on color to identify different slices
Font size is
important! Rule of
thumb over 12
36
Labels and descriptions
Consider Alt Text Provide a chart description
37
White divider
Consider white space
38
Good Vs Bad visualization
39
Examples - pie chart
Depicting too many slices decreases the Place the largest section at 12 o’clock, going clockwise.
impact of the visualization Place the second largest section at 12 o’clock, going
counterclockwise or clockwise.
Data_Visualization_101_How_to_Design_Charts_and_Graphs
40
Examples - bar chart
Space between bars should be ½ bar width when Starting at a value above zero truncates the bars
tools provide that option. and doesn’t accurately reflect the full value.
Data_Visualization_101_How_to_Design_Charts_and_Graphs
41
Examples - line chart
Plot all data points so that the line chart
If you need to display more, break them out into
takes up approximately two-thirds of the y-axis’
separate charts for better comparison.
total scale.
Data_Visualization_101_How_to_Design_Charts_and_Graphs
42
Examples - scatter plot
Use trend lines: These help draw correlation between Too many lines make data difficult to interpret.
the variables to show trends.
Data_Visualization_101_How_to_Design_Charts_and_Graphs
43
Examples - bubble map
Avoid adding too much detail or using shapes
Bubbles should be scaled according to area,
that are not entirely circular; this can lead to
not diameter.
inaccuracies.
Data_Visualization_101_How_to_Design_Charts_and_Graphs
44
Data visualization examples
Information management system and visualization
Data collection system Data visualization system Integration is needed
ActivityInfo PowerBI
● Dedicated system for data visualization - more data visualization options
● More time in integration and higher level of capacity building is needed
46
Information management system and visualization
Data collection system and data visualization in the same system No Integration is needed
ActivityInfo
Less time, people and budget needed for integration
Development assistance project Feedback complaint and
Cash based interventions
response mechanism
47
Key messages
● Start always with who will read your report and what is the message that you want
to convey!
● If you want to confirm the type of reports and audience look back at your M&E plan
and your data model!
● The chart type depends on questions (i.e. relationships) and data type.
● Always consider font size, colors and text descriptions in your data visualizations.
48
Resources
● Development assistance project
● Sudan - Multi purpose assistance
● Harvard University - Data accessibility
● Storytelling with data
● From data to viz
● Data_Visualization_101_How_to_Design_Charts_and_Graphs
● Contrast ratio
● Development assistance project
● Feedback complaint and response mechanism
● Cash based interventions
49
Questions?
Follow us:
LinkedIn page: [Link]
LinkedIn group: [Link]
50