0% found this document useful (0 votes)
32 views23 pages

Understanding Data Analytics Basics

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
32 views23 pages

Understanding Data Analytics Basics

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Written By: Mr.

Wasi Abbas (Aspire College Jhang)

Chapter 5: Data Analytics


Define: Data analytics is the process of collecting, organizing, and analyzing data to discover
useful information, patterns, and insights that help in making better decisions.

OR

Data analytics means using statistical methods, tools, and technologies to examine raw data and
turn it into meaningful conclusions.

Basic Statistical concepts:

Statistics is the branch of mathematics that deals with collecting, organizing, analyzing,
interpreting, and presenting data to draw meaningful conclusions.

Example

If you survey 1,000 people to find out how many like tea vs. coffee, the process of collecting and
analyzing that information is statistics.

Uses:

We use Statistics because:

• It helps us to condense large amount of data into simple for.


• It is used to understand pattern and tools.
• It is used to make decision.
• It also provides us the concise output of result.

Measures of Central Tendency


It is process, method or measure that is used to find the typical or central value from the data
set. Measures of central tendency are statistical tools that describe the center or average value
of a set of data.
They show what is typical or most representative of all the data values.

Main Types of Measures of Central Tendency

1. Mean (Arithmetic Average)

It is the sum of all values divided by the total number of values. We can say that it is just like
average or it is average.
Written By: Mr. Wasi Abbas (Aspire College Jhang)

2. Median

The middle value when all observations are arranged in ascending or descending order.

If the number of values is even, the median is the average of the two middle values.

o Example:
For 5, 10, 15 → Median = 10
For 5, 10, 15, 20 → Median = (10 + 15) / 2 = 12.5

3. Mode

The value that occurs most frequently in the data.

There can be one mode (unimodal), two modes (bimodal), or more (multimodal).

o Example:
Data: 2, 3, 3, 5, 7 → Mode = 3

Measures of Dispersion
It is a measure that helps us to understand how much the data values are differ from the
average or mean. They also give us sense of variation within the dataset. There are two
measures of dispersion: 1. Variance 2: Standard Deviation.

Variance:

Variance is a measure of how much the values in a dataset differ from the mean (average).
It shows the degree of spread or dispersion of data values around the mean.

In simple words, it tells us how far the data points are from the average value.
Written By: Mr. Wasi Abbas (Aspire College Jhang)

Formulas:

Example 1:
Written By: Mr. Wasi Abbas (Aspire College Jhang)

Example 2:

Standard Deviation
Standard deviation is a measure of dispersion that shows how much the values in a dataset
deviate (differ) from the mean on average.

It indicates whether the data values are close to the mean (low deviation) or spread out over a
wide range (high deviation).

Formula:
Written By: Mr. Wasi Abbas (Aspire College Jhang)

Formula:

Example 1

Standard Deviation = 3.54

Example 2

Standard Deviation = 21.33

Probability
Probability is the branch of mathematics that deals with the chance of an event is to occur.
It shows the chance or likelihood of a specific outcome happening in a given situation.

The probability of any event is always between 0 and 1:

• 0 means the event cannot happen.

• 1 means the event will definitely happen.

The formula for probability is:

Example 1

If you toss a coin, there are two possible outcomes: Head (H) or Tail (T).
The probability of getting a head is:
Written By: Mr. Wasi Abbas (Aspire College Jhang)

This means there’s a 50% chance of getting a head.

Example 2 (with Data)

A box contains 5 red balls, 3 blue balls, and 2 green balls.


Total balls = 5 + 3 + 2 = 10

If we randomly pick one ball:

This means there’s a 50% chance to pick a red ball, 30% for blue, and 20% for green.

Applications of Probability

1. Weather Forecasting: Meteorologists use probability to predict rain chances (e.g., 80%
chance of rain).

2. Business and Risk Analysis: Companies use probability to estimate risks, losses, or
profits.

3. Games and Gambling: Probability determines winning chances in games like cards, dice,
or lotteries.

4. Medical Research: Used to calculate chances of disease recovery or side effects.

5. Insurance: Companies use probability to calculate premium rates based on risk.

Conclusion

Probability is an essential concept in mathematics and real life.


It helps in decision-making, risk management, and predicting uncertain outcomes.
Understanding probability allows individuals and organizations to make more informed and
logical choices in uncertain situations.

Data Collection and Preparation


Data Collection Methods
Data collection methods are the techniques and tools used to gather information from various
sources to answer research questions, test hypotheses, and evaluate outcomes. These methods
ensure that the collected data is accurate, reliable, and relevant for analysis.
Written By: Mr. Wasi Abbas (Aspire College Jhang)

Surveys

A survey is a method of data collection in which information is gathered from a group of people
(called respondents) using a set of structured questions. Surveys are commonly used to collect
opinions, behaviors, preferences, or facts about a specific topic.

Surveys can be conducted through questionnaires, interviews, phone calls, or online forms.

Examples of Surveys

1. Student Satisfaction Survey:


Used by schools or colleges to find out how satisfied students are with teaching,
facilities, or administration.
Example: “How satisfied are you with your computer lab facilities?”

2. Employee Feedback Survey:


Conducted by companies to measure employee morale, engagement, and job
satisfaction.
Example: “Do you feel valued at your workplace?”

3. Customer Preference Survey:


Used by businesses to understand what customers like, dislike, or expect from products
or services.

Example – Customer Preference Survey

Purpose: To find out customer choices and expectations about a product.


Example Scenario: A mobile phone company wants to design a new smartphone.

Sample Questions:

1. Which brand of smartphone do you currently use?

2. What features do you value the most?

o Camera quality

o Battery life

o Storage capacity

o Price

3. How much are you willing to spend on a new smartphone?

4. Would you recommend our brand to others? (Yes/No)


Written By: Mr. Wasi Abbas (Aspire College Jhang)

Observations
Observation is a data collection method in which the researcher watches, listens, and records
behaviors, events, or situations as they naturally occur — without asking questions or
interfering.

It helps in collecting real and accurate information about how people actually behave rather
than what they say they do.

Examples of Observation

1. Classroom Observation:
A teacher observes how actively students participate in class discussions.
Example: Recording how many students raise their hands during a question-answer
session.

2. Customer Behavior Observation:


A store manager observes how customers move through the store to improve product
placement.
Example: Watching which shelves attract the most attention in a supermarket. *

3. Traffic Observation:
A researcher observes traffic flow at an intersection to study rush-hour patterns.

Experiment
An experiment is a scientific method of data collection in which the researcher manipulates one
or more variables to study their effect on another variable, under controlled conditions.

It helps in finding cause-and-effect relationships between factors.

Examples of Experiments

1. Educational Experiment:
A teacher introduces a new teaching method to one group of students and uses the old
method for another group to compare performance.
Example: Testing whether using multimedia lessons improves test scores compared to
traditional lectures. *

2. Marketing Experiment:
A company tests two different advertisements (A and B) to see which one attracts more
customers.
Example: Ad A shown on Facebook, Ad B on Instagram — results compared based on
sales. *
Written By: Mr. Wasi Abbas (Aspire College Jhang)

Handling Missing Data – Explanation

Definition:
Handling missing data means using suitable techniques to manage or replace the values that are
missing in a dataset so that analysis or decision-making remains accurate and reliable.

Missing data can occur due to:

• Human error (e.g., skipping a question on a form)

• System errors (e.g., data not recorded o

Common Methods to Handle Missing Data

1. Deletion Method

• Explanation: Remove rows or columns that contain missing values.

• Example:

Student Age Marks

A 16

B 17 —

C 16 90

• → Remove B’s row because Marks is missing.

2. Imputation Method

• Explanation: Replace missing values or calculated ones like mean, median, or mode.

• Example (Mean Imputation):


Marks = (85 + 90) / 2 = 87.5
So, B’s Marks = 87.5

Student Age Marks

A 16 85

B 17 87.5
Written By: Mr. Wasi Abbas (Aspire College Jhang)

Student Age Marks

C 16 90

• Use When:
The missing data is not large, and you want to keep the dataset complete.

3. Model-Based (Predictive) Imputation

• Explanation: Predict missing values using relationships between other variables (e.g.,
regression, decision trees, or KNN).

• Example:
If "Marks" depends on "Hours Studied," use regression to predict missing Marks.

Student Hours Studied Marks

A 2 85

B 4 —

C 3 90

• Regression predicts Marks for B = 88.

• Use When:
Data shows strong relationships between variables.

Data Cleaning and Transformation

Introduction

In data analysis, raw data collected from different sources often contains errors, missing values,
or inconsistencies. To make it useful for analysis, we need to clean and transform the data. Data
cleaning and transformation are essential steps in the data preprocessing stage to ensure
accuracy, consistency, and reliability.

1. Data Cleaning

Definition:
Data cleaning (or data cleansing) is the process of detecting and correcting (or removing) errors
and inconsistencies from data to improve its quality.

Common issues handled in data cleaning:

• Missing values
• Duplicate records
Written By: Mr. Wasi Abbas (Aspire College Jhang)

• Inconsistent data formats


• Typographical errors
• Outliers (extreme or unusual values)

Example:

Name Age City Salary

Ali 25 Lahore 50000

Ahmad Karachi 55000

Ali 25 Lahore 50000

Sana 27 lahore 58000

Cleaning Steps:

• Fill missing Age (e.g., using average or median value).


• Remove duplicate rows.
• Standardize city names (“lahore” → “Lahore”).

Cleaned Data:

Name Age City Salary

Ali 25 Lahore 50000

Ahmad 26 Karachi 55000

Sana 27 Lahore 58000

2. Data Transformation

Definition:
Data transformation is the process of converting data from one format or structure into another
suitable format for analysis.

Common transformation methods:

• Normalization: Scaling values into a specific range (e.g., 0–1).

• Encoding: Converting categorical data into numeric form.


Written By: Mr. Wasi Abbas (Aspire College Jhang)

• Aggregation: Summarizing data (e.g., total sales per month).

• Date formatting: Converting dates into a consistent format.

Example:

Product Sales Month

A 1000 Jan-2024

B 2000 01/02/2024

Transformation Steps:

• Convert “Jan-2024” → “2024-01” for consistent date format.

• Add a new column for Quarter (e.g., “Q1”).

• Normalize sales data between 0 and 1 for comparison.

Transformed Data:

Product Sales Month Quarter Normalized Sales

A 1000 2024-01 Q1 0.33

B 2000 2024-02 Q1 0.66

Building Statistical Models


Definition

A statistical model is a mathematical representation of real-world data that shows the


relationship between different variables.
It helps in understanding patterns, making predictions, and supporting decisions based on data.

Five Steps in Building a Statistical Model:


Written By: Mr. Wasi Abbas (Aspire College Jhang)

1. Define a Problem:
Clearly state what you want to analyze or predict.
Example: Predict students’ exam scores based on study hours and attendance.

2. Collect Data:
Gather relevant and reliable data from surveys, databases, or observations.
Example: Collect data of students’ study hours, attendance, and exam results.

3. Choose an Algorithm:
Select a suitable statistical or machine learning algorithm (e.g., linear regression,
decision tree) based on the problem type.
Example: Use linear regression to find the relationship between study hours and exam
scores.

4. Train a Model:
Feed the collected data into the algorithm so it can learn patterns and relationships
among variables.
Example: The model learns that more study hours lead to higher scores.

5. Evaluate a Model:
Test the model’s performance using new data and accuracy measures (like R² or error
rate).
Example: Check how well the model predicts exam scores for new students.

Linear Regression

Introduction

Linear regression is a statistical method used to study the relationship between two variables —
one independent (cause) and one dependent (effect). It helps us predict the value of one
variable based on the value of another.

For example, if you own a fruit stall, you may want to know how your daily earnings depend on
the number of customers you get each day. Linear regression helps you find this relationship
and predict future earnings.

Step 1: Collecting Data

To build a linear regression model, we need data.


Suppose you record your number of customers and daily earnings for 5 days:

Number of Customers Daily Earnings (Rs.)

10 500
Written By: Mr. Wasi Abbas (Aspire College Jhang)

Number of Customers Daily Earnings (Rs.)

15 700

20 900

25 1100

30 1300

Here:

• Independent Variable (X): Number of Customers

• Dependent Variable (Y): Daily Earnings

Step 2: Understanding the Linear Regression Formula

The general formula for simple linear regression is:

Where:

• Y = Dependent variable (Earnings)

• X = Independent variable (Customers)

• β₀ = Intercept (value of Y when X = 0)

• β₁ = Slope (how much Y changes for each 1-unit change in X)

• ε = Error term (difference between actual and predicted values)

Step 3: Building the Linear Regression Model

We now find the slope (β₁) and intercept (β₀).

Finding the slope (β₁):


From the table, when customers increase by 5, earnings increase by 200 Rs.

This means each new customer adds 40 Rs to your daily earnings.


Written By: Mr. Wasi Abbas (Aspire College Jhang)

Finding the intercept (β₀):


Using one data point (10 customers = 500 Rs):

So, when there are no customers, you’d still earn 100 Rs, maybe from regular customers or
fixed sales.

Final Equation

This means:

• You’ll always earn 100 Rs even if no one comes.

• Each new customer adds 40 Rs to your total earnings.

Step 4: Interpreting the Model

You can now use this equation to predict future earnings.


If tomorrow you expect 22 customers, the prediction is:

So, you can expect to earn 980 Rs with 22 customers.

Step 5: Testing the Model

To check how accurate the model is, test it with new data.
If on the 6th day, 28 customers visit, predicted earnings are:

But your actual earnings were 1,250 Rs.

The error (difference) is:


Written By: Mr. Wasi Abbas (Aspire College Jhang)

This small difference shows that the model is fairly accurate, though real-world data may vary
slightly.

2nd Example:

Step 1: Collecting Data

Suppose a teacher wants to know how students’ study hours affect their exam marks.
She collects data from 5 students:

Study Hours (X) Exam Marks (Y)

2 50

4 60

6 70

8 80

10 90

Here:

• X (Independent variable) = Number of hours studied

• Y (Dependent variable) = Marks obtained in the exam

Step 2: Understanding the Relationship

By looking at the table, you can see that every time study hours increase by 2, marks increase by
10.

So, the slope (β₁) is:

This means that for every extra hour studied, marks increase by 5.

Step 3: Finding the Intercept (β₀)

To find β₀, use one data point, for example (X=2, Y=50):
Written By: Mr. Wasi Abbas (Aspire College Jhang)

So, intercept (β₀) = 40


That means even if a student studies 0 hours, they might still get 40 marks, maybe due to
classroom learning or guessing.

Step 4: Final Equation

Step 5: Making a Prediction

If a student studies for 7 hours, then:

The model predicts that if a student studies for 7 hours, they will score 75 marks.

Logistic Regression
Logistic Regression is a statistical method used when we want to predict an outcome that can be
categorized as “Yes” or “No.”
It helps us find the probability that an event will happen based on given input data.

Example:
Suppose we want to predict whether a student will pass or fail an exam depending on how
many hours they studied.
Instead of giving a specific score prediction, logistic regression tells us the probability (like 0.8 or
80%) that the student will pass.

Understanding Logistic Regression

Logistic Regression is different from Linear Regression because:

• Linear regression predicts continuous numbers (like marks or prices).


Written By: Mr. Wasi Abbas (Aspire College Jhang)

• Logistic regression predicts categories (like Pass/Fail, Yes/No).


It gives results between 0 and 1, showing how likely it is that something will happen.

Example:
If logistic regression predicts a probability of 0.85, it means there’s an 85% chance the student
will pass the exam.

Clustering Techniques
Definition:
Clustering is a method of grouping similar items together based on their characteristics.
It helps in identifying patterns or groups in data.

Example (Clustering of Students by Performance):


Let’s say we have the following scores:

Student Math Score English Score

Basim 85 70

Umer 90 65

Anie 80 85

Tallat 40 60

Maliha 60 65

We can use clustering to group students who have similar performance.


For example:

• Group 1: Basim and Umer (good at Math)

• Group 2: Anie and Tallat (good at English)

• Group 3: Maliha (average in both)

This helps teachers understand which students perform similarly and where extra help may be
needed.

K-Means Clustering
Definition:
K-Means is one of the simplest and most used clustering methods. It divides data into K groups
(clusters) based on similarities.
Written By: Mr. Wasi Abbas (Aspire College Jhang)

Example:
If we set K=2 for the above table:

• One cluster may include Basim and Umer (strong in Math)

• Another cluster may include Anie and Tallat (strong in English)

The algorithm calculates distances between student scores and groups the most similar ones
together.

Evaluating and Interpreting Models

Once a model is built, we must test how well it performs.


This is called model evaluation, and it helps us understand how accurate and reliable our model
is.

Performance Metrics

Performance metrics measure how well a model is performing.

Error Metrics

Error metrics show how much the model’s predictions differ from the actual values.

Example:
If a model predicts a grocery bill of 8,000 rupees but the actual bill is 10,000 rupees,
then the error = 2,000 rupees.

Accuracy Metrics

Accuracy metrics tell us how many predictions were correct.

Example:
If a model predicts whether students will pass or fail,
and it correctly predicts 9 out of 10 results, then its accuracy is 90%.

Interpreting Outputs
Interpreting a model’s output means understanding what the results actually mean and how
they can be used to make real-life decisions.

Drawing Conclusions from Insights

After analyzing a model’s results, we can draw conclusions or insights that help improve
performance or decision-making.
Written By: Mr. Wasi Abbas (Aspire College Jhang)

Example:
If a linear regression model shows that the number of study hours strongly affects exam scores,
we can conclude that:

“Students who study more hours are likely to get higher marks.”
This means that increasing study time can help improve performance.

Ethical Considerations

When building statistical or machine learning models, it’s important to think about ethics —
that is, making sure the model is fair, unbiased, and respects people’s privacy.

Fairness and Bias

A good model should be fair and unbiased. It should not favor one person or group unfairly.

Example:
If a bank uses a model to decide who gets a loan, the model should not prefer people from a
certain area or gender — it should only decide based on financial data and repayment history.

Data Privacy

When personal data is used to build a model, it is essential to keep that data secure and private.
The data should not be shared or misused.

Example:
If a company uses customer data to predict buying behavior, it must make sure that all customer
information is protected and not shared with others without permission.

Introduction to Data Visualization

Data visualization is the process of presenting information in a visual form such as graphs or
charts. It helps us quickly identify patterns, trends, and insights from data.

Types of Visualizations

Data visualization makes it easier to understand and analyze complex information. There are
many types of visualizations, each designed for a specific purpose. Below are some common
examples:
Written By: Mr. Wasi Abbas (Aspire College Jhang)

Bar Charts

A bar chart is used to compare different categories. Each bar represents a category, and the
height (or length) of the bar shows the value for that category.

Example:
If you want to compare sales figures for different products in a store, a bar chart can be used to
display this data clearly.

(Figure 5.1: Bar chart showing sales figures for different products)

Line Graphs

A line graph is used to show how data changes over time. It connects individual data points with
a line, making it easy to spot trends or patterns.

Example:
If you track the temperature throughout a week, a line graph can show how the temperature
rises and falls each day.

(Figure 5.2: Line graph showing variation of temperature over time)

Histograms

A histogram shows how data is distributed across different ranges or intervals. It helps you see
how frequently certain values appear in a dataset.

Example:
If you want to analyze how students scored in a math exam, a histogram can display the
distribution of their scores.

(Figure 5.3: Histogram showing the distribution of exam scores)

Scatterplots

A scatterplot shows the relationship between two variables. Each point represents an
observation, with its position indicating the values of both variables.

Example:
You can use a scatterplot to explore the relationship between the number of hours studied and
exam scores.

(Figure 5.4: Scatterplot showing the relationship between study hours and exam scores)

Boxplots
Written By: Mr. Wasi Abbas (Aspire College Jhang)

Boxplots (also called whisker plots) are used to summarize the distribution of a dataset. They
display important statistical values such as the median, quartiles, and outliers. A boxplot
provides a quick visual summary of how data is spread out and how variable it is.

Example:
A boxplot can be used to compare exam scores of different classes to see which class performed
better overall.

(Figure 5.5: A boxplot showing class performance in exam scores)

Tools for Data Visualization


Data visualization tools help us turn large sets of numbers into easy-to-understand charts and
graphs. These visual tools make it simpler to identify trends and patterns in data.

There are many tools available for creating visualizations, but some of the easiest to use are
Microsoft Excel and Google Sheets. These tools are accessible, user-friendly, and allow you to
create bar charts, line graphs, pie charts, and more.

Using Excel and Google Sheets for Visualization

Excel and Google Sheets make it easy to enter your data and generate visualizations with just a
few clicks.

Example:
If you run a small business and want to track monthly sales, you can record your sales figures in
Excel or Google Sheets. Then, by using the chart options, you can create a bar chart to see
which month had the highest sales.

Creating and Interpreting Visualizations

Here’s a simple step-by-step guide to creating a visualization in Excel or Google Sheets:

1. Enter Your Data:


Type your data into the spreadsheet. For example, one column could have the months
(January, February, etc.), and another column could have the sales figures for each
month.

2. Select the Data:


Highlight the cells that contain the data you want to visualize.

3. Choose a Chart Type:


Go to the “Insert” tab and select the type of chart you want to create — such as a bar
chart, line graph, or pie chart.
Written By: Mr. Wasi Abbas (Aspire College Jhang)

4. Customize the Chart:


Add labels, titles, and axis names (e.g., months on the x-axis and sales on the y-axis) to
make the chart easier to understand.

5. Understand the Visualization:


Always interpret what the chart shows. It’s important to understand the story your
visualization is telling — such as identifying patterns, trends, or outliers.

You might also like