Module 1: Introduction to
Statistics
Dr. Abdul Wajid
Origin and Development of Statistics
Origin
The word "Statistics" originates from the Latin word "Status" or the Italian word "Statista," meaning state or
political affairs.
Historically, statistics was used by ancient governments to collect data on population, taxation, and military
strength.
Examples of early use:
Egyptians (3000 BC): Maintained records for building Romans: Used census data for administrative Kautilya’s Arthashastra (Ancient India): Mentioned
pyramids. purposes. methods of record-keeping for governance
Development
16th–17th Century: Emergence of statistical Introduction of descriptive and inferential
thinking as a science. statistics.
•John Graunt (1662): Laid the foundation of •Development of statistical tools like regression
demography through population studies. analysis by Sir Francis Galton.
18th Century 20th Century to Present
17th Century 19th Century
Statistics began expanding into economics Expanded applications in business,
and other fields. healthcare, technology, and social sciences.
•Contributions from Laplace and Gauss (probability Integration of computing and software for
theory).
advanced analytics (e.g., Python, R, and
SPSS).
Importance of Statistics
Importance of Statistics
Importance of Statistics
Limitations
Not Applicable to Qualitative Data:
• Statistics cannot measure emotions, feelings, or abstract concepts directly.
Misleading Results:
• Poor data quality or inappropriate methods can lead to incorrect conclusions.
Dependence on Assumptions:
• Many statistical models rely on assumptions that may not always hold true.
Lack of Universality:
• Statistical findings may vary with changing contexts and conditions.
Challenges
Inaccurate, incomplete, or biased data
Data Quality Issues: can compromise results.
Complexity of Advanced Modern methods require specialized
Techniques: training and software.
Misuse of statistics for manipulation or
Ethical Concerns: misleading conclusions.
Managing and analyzing vast datasets
Big Data and Privacy: while safeguarding personal
information.
Distrust of Statistics
Misrepresentation of Data:
• Statistics can be intentionally manipulated to mislead (e.g., selective sampling, biased graphs).
Complexity of Methods:
• Advanced techniques can confuse non-experts, leading to skepticism.
Cherry-Picking Results:
• Presenting only favorable outcomes while ignoring contradictory data.
Overgeneralization:
• Applying findings from small or biased samples to larger populations.
How to Address Distrust
Transparency: Clearly disclose data sources and methods.
Education: Improve statistical literacy among the general public.
Ethical Practices: Avoid misuse of statistics for propaganda or manipulation.
Techniques of Data Collection
Primary Data Collection
Surveys and Questionnaires:
• Structured methods to collect data from individuals or groups.
• Example: Customer satisfaction surveys.
Interviews:
• Face-to-face or virtual discussions to gather in-depth information.
• Example: Interviews with focus groups for product development.
Observation:
• Recording behavior or events as they occur.
• Example: Observing consumer behavior in retail stores.
Experiments:
• Controlled studies to test hypotheses.
Secondary Data Collection
Published • Government reports, research
Sources: papers, and industry publications.
Online • Sources like World Bank, IMF, or
Databases: company annual reports.
Historical • Archived data from previous
Records: studies or documents.
Classroom
Activity
Under what circumstances should a researcher use
primary data, and when is it more appropriate to rely on
secondary data?
Primary Data
Definition: Data collected firsthand by the researcher for a specific purpose.
When to use:
1. When specific information is required that is not available elsewhere.
2. When the research problem is unique or new.
3. When the researcher needs up-to-date, original, and reliable data.
4. When high accuracy is important.
In field studies, surveys, experiments, interviews, or direct observations.
Secondary Data
Definition: Data collected by someone else for a different purpose but used by the
researcher for analysis.
When to use:
• When data is already available and reliable.
• When time and cost constraints make it difficult to collect primary data.
• For background information, literature review, or identifying trends.
• When doing comparative or historical studies.
• When large-scale data is required (e.g., census, company annual reports, government
publications).
Example: If you use government census data to study population growth in your country,
that’s secondary data
Tabulation
Definition: Systematic arrangement of data in rows and columns for easy
understanding.
Types:
• Simple Table: Summarizes data for one variable (One way table).
• Complex Table: Summarizes data for multiple variables.
[Link]-way table
[Link]-way table
[Link] table
Advantages:
• Easy comparison of data.
• Facilitates statistical analysis.
Simple Table
Complex Tables
Graphical Presentation
Bar Chart: Displays categorical data
through bars.
• Example: Monthly sales comparison across regions.
Pie Chart: Represents proportions
within a dataset.
• Example: Market share of different companies.
Histogram: Displays the frequency
distribution of numerical data.
• Example: Distribution of student grades.
Line Graph: Shows trends over
time.
• Example: Stock price movements over a year.
Scatter Plot: Shows relationships • Example: Correlation between advertising expenses and
between two variables. sales revenue.
Bar Chart
Pie Chart
Line Graph
Scatter Plot
MEASUREMENT SCALES OF A DATA
1. Nominal-level Data (no order or no comparing values)
• Equality, Categories, No mathematical meaning
• The nominal-level of measurement classifies data into mutually exclusive (non
Overlapping), exhaustive categories in which no ordering or ranking can be
imposed on the data.
Some examples of variables that can be measured on a nominal scale include:
•Gender: Male, female
•Eye color: Blue, green, brown
•Hair color: Blonde, black, brown, grey, other
•Blood type: O-, O+, A-, A+, B-, B+, AB-, AB+
•Political Preference: Republican, Democrat, Independent
•Place you live: City, suburbs, rural
Variables that can be measured on a nominal scale have the following
properties:
•They have no natural order. For example, we can’t arrange eye colors in order
of worst to best or lowest to highest.
•Categories are mutually exclusive. For example, an individual can’t
have both blue and brown eyes. Similarly, an individual can’t live both in the
city and in a rural area.
•The only number we can calculate for these variables are counts. For example,
we can count how many individuals have blonde hair, how many have black
hair, how many have brown hair, etc.
•The only measure of central tendency we can calculate for these variables is the
mode. The mode tells us which category had the most counts. For example,
we could find which eye color occurred most frequently.
MEASUREMENT SCALES OF A DATA
2. Ordinal-Level Data – Order, Rank
• The ordinal-level of measurement classifies data into categories that can be ordered
or ranked.
(Only before and after no bigger or less...)
• However, precise differences between the ranks do not exist
Some examples of variables that can be measured on an ordinal scale include:
•Satisfaction: Very unsatisfied, unsatisfied, neutral, satisfied, very satisfied
•Socioeconomic status: Low income, medium income, high income
•Workplace status: Entry Analyst, Analyst I, Analyst II, Lead Analyst
•Degree of pain: Small amount of pain, medium amount of pain, high amount of
pain
Variables that can be measured on an ordinal scale have the following properties:
•They have a natural order. For example, “very satisfied” is better than “satisfied,”
which is better than “neutral,” etc.
•The difference between values can’t be evaluated. For example, we can’t exactly
say that the difference between “very satisfied and “satisfied” is the same as the
difference between “satisfied” and “neutral.”
•The two measures of central tendency we can calculate for these variables are the
mode and the median. The mode tells us which category had the most counts and
the median tells us the “middle” value.
3. Interval-level Data (Quantitative data)
• The interval level of measurement ranks data, and precise differences between
units of measure do exist. (Equal distances between 2 points) However, there is
no meaningful zero (i.e., starting point)
Some examples of variables that can be measured on an interval scale include:
•Credit Scores: Measured from 300 to 850
These variables have a natural order.
•We can measure the mean, median, mode, and standard deviation of
these variables.
•These variables have an exact difference between values. Recall that ordinal
variables have no exact difference between variables – we don’t know if the
difference between “very satisfied” and “satisfied” is the same as the
difference between “satisfied” and “neutral.” For variables on an interval scale,
though, we know that the difference between a credit score of 850 and 800 is
the exact same as the difference between 800 and 750.
•These variables have no “true zero” value. For example, it’s impossible to have
a credit score of zero. And for temperatures, it’s possible to have negative
values (e.g. -10° F) which means there isn’t a true zero value that values can’t
go below.
[Link]-level Data (Quantitative data)
Possesses all the characteristics of interval
measurement (i.e., data can be ranked, and
there exists a true zero or starting point). In
addition, true ratios exist between different
units of measure.
What makes data “interval”?
Equal differences matter: The difference between two values is meaningful.
No true zero: Zero doesn’t mean “nothing.”
Example: Temperature in Celsius
10°C → 20°C → 30°C
Difference between 10°C and 20°C = 10°C
Difference between 20°C and 30°C = 10°C
So the intervals are equal.
But: 0°C doesn’t mean “no temperature.” It’s just a point on the scale.
That means we can say 20°C is 10°C warmer than 10°C, but we cannot say 20°C is
twice as hot as 10°C.
Some examples of variables that can be measured on a ratio scale include:
•Height: Can be measured in centimeters, inches, feet, etc. and cannot have a
value below zero.
•Weight: Can be measured in kilograms, pounds, etc. and cannot have a value
below zero.
•Length: Can be measured in centimeters, inches, feet, etc. and cannot have a
value below zero.
Variables that can be measured on a ratio scale have the following properties:
•These variables have a natural order.
•We can calculate the mean, median, mode, standard deviation, and a variety of other
descriptive statistics for these variables.
•These variables have an exact difference between values.
•These variables have a “true zero” value. For example, length, weight, and height all
have a minimum value (zero) that can’t be exceeded. It’s not possible for ratio
variables to take on negative values. For this reason, the ratio between values can
be calculated. Someone who is 6 feet tall is 1.5 times taller than someone who is 4
feet tall.
Major International Statistical Organizations
1. United Nations Statistics Division (UNSD)
• Global coordination of statistical systems
• Publishes UN Data and Statistical Yearbook
2. International Monetary Fund (IMF)
• International Financial Statistics, Balance of Payments
3. World Bank (WDI- World Development Indicators)
• Open data on development and poverty
4. OECD Statistics (Organization of Economic Co-operation and Development)
• Economic, social, and environmental indicators
Major International Statistical Organizations
5. ILO (International Labor Organization)
• Employment, wages, and labor conditions
6. WHO (Global Health Observatory)
• Mortality, disease, and healthcare access
7. Eurostat
• EU member states statistics
Major Statistical Organizations in Qatar
1. Planning and Statistics Authority (PSA)
• National census, surveys, economic & demographic data
2. Qatar Central Bank (QCB)
• Monetary and financial statistics
3. Qatar Financial Centre (QFC) & Qatar Stock Exchange (QSE)
• Financial market and investment statistics
4. Ministries & Specialized Agencies
• Ministry of Public Health – health stats
• Ministry of Education – education stats
• Qatar Energy & Ministry of Environment – energy & environment
Classroom Activity
Find out the major statistical
organizations in your country.
Cross-Sectional Data
Meaning: Data collected at one point in time. It studies different individuals,
groups, firms, or countries at the same time.
Focus: Comparing differences across units.
Example:
Suppose you survey 100 students in September 2025 about their monthly
pocket money.
Student A: 200 QAR
Student B: 300 QAR
Student C: 250 QAR
Here, you are not tracking over time, just taking a snapshot at one time across
many units (students).
Time-Series Data
Meaning: Data collected for one unit (individual, company, country, etc.) over different
points in time.
Focus: Observing trends, growth, or changes over time.
Example:
Monthly inflation rate of Qatar from Jan 2020 to Dec 2024.
Jan 2020: 1.5%
Feb 2020: 1.7%
…
Dec 2024: 2.8%
Here, the data is for one country (Qatar) but across different time periods.
Panel Data
Meaning: A combination of cross-sectional and time-
series data.
•It observes multiple units over multiple time periods.
Student Sept 2025 Oct 2025 Nov 2025
•Focus: Allows comparison across units and over time.
A 200 QAR 220 QAR 250 QAR
Example:
Suppose you collect data on the monthly pocket money B 300 QAR 310 QAR 330 QAR
of 3 students over 3 months:
C 250 QAR 260 QAR 270 QAR
•Here, multiple students (cross-section) are observed
across different months (time-series).