0% found this document useful (0 votes)
4 views35 pages

Basic Statistics

The document provides a comprehensive overview of basic statistics concepts, including definitions of statistics, data types, and key statistical measures such as mean, median, and mode. It also covers probability, variability, percentages, ratios, and data visualization techniques, explaining their importance and applications. Additionally, it discusses correlation and its types, emphasizing the relationships between variables.

Uploaded by

yasirshakeel575
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views35 pages

Basic Statistics

The document provides a comprehensive overview of basic statistics concepts, including definitions of statistics, data types, and key statistical measures such as mean, median, and mode. It also covers probability, variability, percentages, ratios, and data visualization techniques, explaining their importance and applications. Additionally, it discusses correlation and its types, emphasizing the relationships between variables.

Uploaded by

yasirshakeel575
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

📊 Basic Statistics

Statistics Fundamentals
1. What is statistics?

Answer: Statistics is the study of collecting, organizing, analyzing, and understanding


information from different types of data.​
Example: We can use statistics to analyze students' marks and find their average.

2. Why is statistics important?

Answer: Statistics helps us understand data, find patterns, compare information, and make
better decisions.​
Example: A company can analyze sales data to understand which product sells more.

3. What is data?

Answer: Data is a collection of facts, numbers, measurements, or information gathered for


analysis.​
Example: Students' names, ages, and marks can form a dataset.

4. What are the main types of data?

Answer: The main types include qualitative data and quantitative data, which describe different
kinds of information.​
Example: Eye color is qualitative, while height is quantitative.

5. What is quantitative data?

Answer: Quantitative data contains numerical information that can be measured, counted, or
used in calculations.​
Example: Age, height, marks, and number of students are quantitative data.

6. What is qualitative data?

Answer: Qualitative data describes qualities or categories, such as colors, names, opinions, or
types.​
Example: Student feedback such as “good” or “poor” is qualitative information.

7. What is discrete data?


Answer: Discrete data consists of separate countable values, such as the number of students
in a class.​
Example: A class can have 30 students, but not 30.5 students.

8. What is continuous data?

Answer: Continuous data can take different values within a range, such as height, weight, or
temperature.​
Example: A person's height could be 170.5 centimeters.

9. What is a dataset?

Answer: A dataset is a collection of related information organized for studying, analyzing, or


understanding data.​
Example: A spreadsheet containing student names, marks, and grades is a dataset.

10. What is an observation?

Answer: An observation is one individual piece of information recorded from a person, object,
or event.​
Example: One student's mark of 85 can be one observation.

11. What is a variable?

Answer: A variable is a characteristic that can have different values among people, objects, or
observations.​
Example: Age is a variable because different people can have different ages.

12. What is a population?

Answer: A population is the complete group of people or objects that we want to study.​
Example: All students in a university could be the population.

13. What is a sample?

Answer: A sample is a smaller group selected from a population for collecting and studying
information.​
Example: We could select 100 students from a university to represent its students.

14. What is the difference between population and sample?

Answer: A population includes the complete group, while a sample represents only part of that
group.​
Example: All university students are the population, while 100 selected students are the
sample.

15. What is a parameter?

Answer: A parameter is a numerical value that describes a characteristic of an entire


population.​
Example: The average age of every student in a university could be a population parameter.

16. What is a statistic?

Answer: A statistic is a numerical value calculated from sample data to describe that sample.​
Example: The average age of 100 selected students is a statistic.

17. What is data collection?

Answer: Data collection is the process of gathering information for a specific study or analysis.​
Example: We can collect students' marks using a questionnaire or school records.

18. What is data analysis?

Answer: Data analysis involves examining information to find patterns, relationships,


differences, and useful results.​
Example: We can analyze marks to find the highest and average scores.

19. Why is data visualization important?

Answer: Data visualization makes information easier to understand by presenting numbers and
patterns using charts or graphs.​
Example: A graph can show changes in sales more clearly than a long table.

20. Why is statistics useful in Computer Science?

Answer: Statistics helps Computer Science students understand data, analyze results, and
work with data-based problems.​
Example: Statistics can help analyze datasets used in programming and data analysis.

Mean, Median, and Mode


21. What is the mean?
Answer: The mean is the average value calculated by adding all values and dividing by their
number.​
Example: For 10, 20, and 30, the mean is 20.

22. How do you calculate the mean?

Answer: To calculate the mean, add all values together and divide the total by their count.​
Example: 10 + 20 + 30 = 60, then 60 ÷ 3 = 20.

23. What is the median?

Answer: The median is the middle value when data is arranged from smallest to largest.​
Example: In 2, 4, 6, 8, 10, the median is 6.

24. How do you calculate the median?

Answer: Arrange the values in order and select the middle value or average the two middle
values.​
Example: For 2, 4, 6, the middle value is 4.

25. What is the mode?

Answer: The mode is the value that appears most frequently in a dataset.​
Example: In 2, 3, 3, 4, the mode is 3.

26. How do you find the mode?

Answer: Count how often each value appears, then identify the value appearing most
frequently.​
Example: If 80 appears three times and other marks appear once, 80 is the mode.

27. What is the range?

Answer: The range is the difference between the highest value and the lowest value.​
Example: If the highest mark is 90 and lowest is 50, the range is 40.

28. How do you calculate the range?

Answer: Subtract the smallest value from the largest value to calculate the range.​
Example: 90 − 50 = 40.

29. What is the difference between mean and median?


Answer: The mean uses all values for calculation, while the median identifies the middle value.​
Example: An unusually high value can affect the mean more than the median.

30. What is the difference between median and mode?

Answer: The median is the middle value, while the mode is the value appearing most
frequently.​
Example: In 2, 3, 3, 4, 5, the median and mode are both 3.

31. When is the mean useful?

Answer: The mean is useful when you want to find the overall average of numerical data.​
Example: We can calculate the average marks of a class.

32. When is the median useful?

Answer: The median is useful when data contains extreme values that could affect the average.​
Example: Median can be useful when comparing salaries with a few very high salaries.

33. When is the mode useful?

Answer: The mode is useful when you want to identify the most common value or category.​
Example: A shop can find which product size is purchased most often.

34. Can a dataset have more than one mode?

Answer: Yes, a dataset can have multiple modes when different values appear with equal
highest frequency.​
Example: In 2, 2, 3, 3, 4, both 2 and 3 are modes.

35. Can a dataset have no mode?

Answer: Yes, a dataset can have no mode when every value appears only once.​
Example: In 1, 2, 3, 4, every value occurs once.

36. What happens to the mean when an extreme value is added?

Answer: An extreme value can significantly increase or decrease the mean depending on its
size.​
Example: Adding 100 to 10, 20, and 30 will increase the average considerably.

37. What is an outlier?


Answer: An outlier is a value that is unusually different from most other values in the dataset.​
Example: If most marks are between 60 and 80, a mark of 5 could be an outlier.

38. Why can outliers affect the mean?

Answer: Outliers can strongly change the mean because every value is included in its
calculation.​
Example: One extremely high salary can increase the average salary of a small group.

39. Is the median affected by outliers?

Answer: The median is usually less affected by outliers because it depends mainly on the
middle position.​
Example: A very high salary may change the mean but have little effect on the median.

40. Give an example of mean, median, and mode.

Answer: For 2, 3, 3, 4, and 8, the mean is 4, median is 3, and mode is 3.​


Reason: The total is 20, the middle value is 3, and 3 appears most often.

Probability
41. What is probability?

Answer: Probability measures how likely an event is to happen, using values between zero and
one.​
Example: A fair coin has a 0.5 probability of landing on heads.

42. What is an experiment in probability?

Answer: An experiment is an action or process performed to observe a possible outcome.​


Example: Rolling a dice is a probability experiment.

43. What is an outcome?

Answer: An outcome is a possible result that can occur from a probability experiment.​
Example: Getting a 4 after rolling a dice is one outcome.

44. What is an event?


Answer: An event is one outcome or a group of outcomes that we are interested in.​
Example: Getting an even number when rolling a dice is an event.

45. What is a probability scale?

Answer: A probability scale ranges from zero to one, showing how likely an event is.​
Example: 0 means impossible, while 1 means certain.

46. What does probability 0 mean?

Answer: Probability zero means the event is impossible and cannot happen under the given
conditions.​
Example: Getting a 7 from a normal six-sided dice has probability zero.

47. What does probability 1 mean?

Answer: Probability one means the event is certain and will happen under the given conditions.​
Example: A normal dice will definitely produce a number from 1 to 6.

48. What does a probability of 0.5 mean?

Answer: A probability of 0.5 means the event has a fifty percent chance of happening.​
Example: A fair coin has a 0.5 chance of landing heads.

49. What is a random event?

Answer: A random event is an event whose exact outcome cannot be known beforehand.​
Example: We cannot know the exact result before rolling a dice.

50. What is the difference between probability and statistics?

Answer: Probability studies possible outcomes, while statistics mainly uses collected data to
understand and describe information.​
Example: Probability can predict possible results, while statistics analyzes actual collected
results.

Variability
51. What is variability?
Answer: Variability describes how much values in a dataset differ from each other.​
Example: Marks of 50, 51, and 52 have low variability.

52. What is variance?

Answer: Variance measures how far data values are spread from the mean.​
Example: A larger variance generally means values are more spread out.

53. What is standard deviation?

Answer: Standard deviation measures how much values usually differ from the dataset's mean.​
Example: It can show whether students' marks are close to their average.

54. Why is standard deviation important?

Answer: Standard deviation helps us understand whether data values are closely grouped or
widely spread.​
Example: A small standard deviation means marks are generally close to the average.

55. What does a small standard deviation mean?

Answer: A small standard deviation means most values are relatively close to the mean.​
Example: Marks of 70, 71, and 72 would have low spread.

56. What does a large standard deviation mean?

Answer: A large standard deviation means the values are more widely spread around the
mean.​
Example: Marks of 20, 60, and 95 show greater spread.

57. What is the difference between variance and standard deviation?

Answer: Variance measures squared differences from the mean, while standard deviation uses
the square root.​
Reason: Standard deviation is often easier to understand because it uses the original data
units.

58. What is the range used for?

Answer: The range shows the overall spread between the smallest and largest values.​
Example: Scores from 40 to 90 have a range of 50.

59. What is dispersion?


Answer: Dispersion describes how widely data values are spread around a central value.​
Example: Standard deviation and range are measures that help describe dispersion.

60. Why do we measure variability?

Answer: We measure variability to understand how different or spread out values are within
data.​
Example: It helps compare whether two groups have similar or different data spread.

Percentages and Ratios


61. What is a percentage?

Answer: A percentage represents a value as a part of one hundred.​


Example: 50% means 50 parts out of 100.

62. How do you calculate a percentage?

Answer: Divide the required value by the total value and multiply the result by one hundred.​
Example: 40 out of 50 gives (40 ÷ 50) × 100 = 80%.

63. What is a ratio?

Answer: A ratio compares two quantities to show their relationship with each other.​
Example: A ratio of 2:3 compares two quantities.

64. What is a proportion?

Answer: A proportion shows that two ratios or relationships are equal.​


Example: 1:2 and 2:4 represent the same proportion.

65. What is percentage change?

Answer: Percentage change shows how much a value has increased or decreased compared
with its original value.​
Example: It can show how much sales increased from one month to another.

66. How do you calculate percentage increase?


Answer: Subtract the original value from the new value, divide by original, and multiply by one
hundred.​
Example: Increasing from 100 to 120 gives a 20% increase.

67. How do you calculate percentage decrease?

Answer: Subtract the new value from the original value, divide by original, and multiply by one
hundred.​
Example: Decreasing from 100 to 80 gives a 20% decrease.

68. What is an average percentage?

Answer: An average percentage summarizes several percentage values, usually by calculating


their mean.​
Example: We can calculate the average percentage of several students' test results.

69. Why are percentages useful in data analysis?

Answer: Percentages make comparisons easier by showing values as parts of a common total.​
Example: 60% is easier to compare across groups than different raw totals.

70. How can percentages be misleading?

Answer: Percentages can be misleading when the total size, comparison group, or original
information is ignored.​
Example: 50% of 10 people and 50% of 10,000 people are very different numbers.

Data Visualization
71. What is data visualization?

Answer: Data visualization means presenting information using charts, graphs, or other visual
methods.​
Example: A graph can show changes in student marks clearly.

72. What is a bar chart?

Answer: A bar chart uses bars to compare values between different categories.​
Example: It can compare the number of students in different departments.

73. What is a histogram?


Answer: A histogram shows how numerical data is distributed across different value ranges.​
Example: It can show how many students scored within each mark range.

74. What is a line chart?

Answer: A line chart shows changes or trends in data, especially across different time periods.​
Example: It can show monthly changes in sales.

75. What is a pie chart?

Answer: A pie chart shows how different categories contribute to a complete total.​
Example: It can show different percentages of a total budget.

76. What is a scatter plot?

Answer: A scatter plot uses points to show the relationship between two numerical variables.​
Example: It can show the relationship between study hours and exam marks.

77. What is a box plot?

Answer: A box plot summarizes data using its middle values, spread, and possible unusual
values.​
Example: It can help identify the spread and possible outliers in marks.

78. What is the difference between a bar chart and histogram?

Answer: A bar chart compares categories, while a histogram shows the distribution of
numerical data ranges.​
Example: A bar chart can compare subjects, while a histogram can show mark ranges.

79. When would you use a line chart?

Answer: I would use a line chart when I need to show changes over time.​
Example: I could show how monthly sales changed during a year.

80. When would you use a scatter plot?

Answer: I would use a scatter plot to examine the relationship between two numerical
variables.​
Example: I could compare students' study hours with their examination marks.
Basic Data Analysis
81. What is correlation?

Answer: Correlation describes how two variables change in relation to each other.​
Example: Study hours and exam marks may have a positive relationship.

82. What does positive correlation mean?

Answer: Positive correlation means two variables generally increase or decrease together.​
Example: More study hours may be associated with higher exam marks.

83. What does negative correlation mean?

Answer: Negative correlation means one variable generally increases while the other
decreases.​
Example: As price increases, demand may decrease in some situations.

84. What does zero correlation mean?

Answer: Zero correlation means there is no clear linear relationship between two variables.​
Example: Two unrelated variables may show little or no linear relationship.

85. Does correlation always mean causation?

Answer: No, correlation does not always mean one variable directly causes changes in another
variable.​
Example: Two variables may change together because another factor affects both.

86. What is an outlier?

Answer: An outlier is an unusual value that is noticeably different from most other observations.​
Example: If most marks are 60–80, a mark of 10 could be unusual.

87. How can outliers be identified?

Answer: Outliers can be identified using charts, ranges, or statistical methods such as the
interquartile range.​
Example: A box plot can help identify unusually high or low values.

88. What is data cleaning?


Answer: Data cleaning involves finding and correcting errors, duplicates, missing values, or
incorrect information.​
Example: We can remove duplicate student records before analyzing the data.

89. Why is data cleaning important?

Answer: Data cleaning improves data quality and helps produce more reliable results from
analysis.​
Example: Incorrect marks could produce incorrect averages if they are not corrected.

90. What is missing data?

Answer: Missing data means some expected information is not available or recorded in the
dataset.​
Example: A student's name may be present while their mark is missing.

Practical / Interview Questions


91. How would you calculate the average marks of students?

Answer: I would add the students' marks and divide the total by the number of students.​
Example: For 60, 70, and 80, the average is (60+70+80) ÷ 3 = 70.

92. How would you find the highest score?

Answer: I would compare the scores or use the MAX function to identify the highest score.​
Example: In Excel, MAX() can quickly find the highest mark.

93. How would you find the lowest score?

Answer: I would compare the scores or use the MIN function to identify the lowest score.​
Example: In Excel, MIN() can quickly find the lowest mark.

94. How would you find the most common score?

Answer: I would count how often each score appears and identify the most frequently occurring
score.​
Example: If 70 appears five times, it may be the mode.

95. How would you compare two datasets?


Answer: I would compare their averages, ranges, distributions, and other useful statistical
measures.​
Example: I could compare the average marks and spread of two classes.

96. How would you identify unusual values?

Answer: I would examine the data using graphs, ranges, or suitable statistical methods to
identify unusual values.​
Example: A box plot can help identify possible outliers.

97. How can Excel be used for basic statistics?

Answer: Excel can calculate averages, totals, minimums, maximums, percentages, and create
useful charts.​
Example: I can use AVERAGE(), MIN(), and MAX() to analyze marks.

98. How can Python be used for statistics?

Answer: Python can process datasets, perform calculations, create charts, and support
statistical analysis using different libraries.​
Example: Python can be used to calculate averages and visualize data with charts.

99. How can SQL help with statistical analysis?

Answer: SQL can retrieve, filter, group, and summarize database information before performing
statistical analysis.​
Example: SQL can find the average marks for students in a particular class.

100. Why did you include Basic Statistics on your CV?

Answer: I included Basic Statistics because I have learned fundamental concepts useful for
understanding and analyzing data.​
Example: I understand concepts such as mean, median, mode, probability, range, and basic
data visualization.

Yes. I compared these 20 Statistics questions with the 100 Basic Statistics questions you
gave earlier.

Already covered — removed

●​ #2 What is Statistics? — already covered.


●​ #18 Difference between sample and population? — already covered.

So below are the 18 new/uncovered questions, with simple interview answers, reasons, and
examples.
📊 Statistics — New Interview Questions
1. What do you think of the phrase "p-value"?

Answer:​
A p-value is a statistical measure used in hypothesis testing. It tells us how strong the evidence
is against the null hypothesis.

Example:​
If the p-value is 0.03, it means the observed result would be relatively unlikely if the null
hypothesis were true.

Reason:​
A commonly used significance level is 0.05. If p < 0.05, we often consider the result statistically
significant.

Interview tip: Don't say "p-value is the probability that the null hypothesis is true."
That is incorrect.

2. What is the Central Limit Theorem?

Answer:​
The Central Limit Theorem (CLT) says that when we take sufficiently large random samples
from a population, the distribution of the sample means tends to become approximately normal,
even if the original population is not normally distributed.

Example:​
Suppose we repeatedly take samples of 50 students and calculate the average marks of each
sample. The distribution of those averages will tend to look approximately normal.

Reason:​
The CLT is important because it allows us to use many statistical methods based on the normal
distribution.

3. What is a hypothesis test? How is statistical significance determined?


Answer:​
A hypothesis test is a method used to determine whether there is enough statistical evidence
to support a claim about a population.

Usually, we have:

●​ Null hypothesis (H₀): There is no effect or difference.


●​ Alternative hypothesis (H₁): There is an effect or difference.

We then calculate a p-value and compare it with a significance level, often 0.05.

Example:​
Suppose we want to know whether a new teaching method improves student marks.

●​ H₀: The teaching method has no effect.


●​ H₁: The teaching method improves marks.

If p = 0.02, which is less than 0.05, we would generally reject H₀.

4. Why are statistical data referred to as observational and experimental?

Answer:​
Data can be classified based on how it is collected.

Observational data: We observe what happens without deliberately changing conditions.

Example:​
Recording the height and weight of students.

Experimental data: We deliberately change something and observe the effect.

Example:​
Giving one group a new teaching method and another group the normal method, then
comparing their results.

Reason:​
The main difference is whether the researcher controls or changes a condition.

5. What is an inlier?

Answer:​
An inlier is a data point that is unusual or incorrect but still falls within the expected range of the
dataset.
Example:​
Suppose most students have marks between 60 and 90, but a data-entry error records a
student's mark as 65 instead of 85. The value 65 may look normal, so it could be an inlier
caused by incorrect data.

Reason:​
Outliers are usually easy to notice because they are far away from other values. Inliers can be
more difficult to detect because they look normal.

6. What is Six Sigma in statistics?

Answer:​
Six Sigma is a data-driven approach used to reduce errors, variation, and defects in processes.

The term refers to a very high level of process quality, where the goal is to have very few
defects.

Example:​
A manufacturing company can use Six Sigma methods to reduce the number of defective
products.

Reason:​
It is mainly used for quality improvement and process optimization.

7. What does KPI stand for?

Answer:​
KPI stands for Key Performance Indicator. It is a measurable value used to evaluate how well
a person, team, process, or organization is achieving a specific goal.

Example:

For an online store:

●​ Monthly sales
●​ Customer retention rate
●​ Conversion rate
●​ Number of new customers

can be KPIs.
Reason:​
KPIs help organizations measure performance and make data-based decisions.

8. Why is the Pareto principle famous?

Answer:​
The Pareto principle, also called the 80/20 rule, suggests that a large portion of results often
comes from a relatively small portion of causes.

Example:​
A business might find that approximately 80% of its sales come from 20% of its customers.

Reason:​
It helps identify the most important factors that have the greatest impact.

The 80/20 relationship is a rule of thumb, not a universal mathematical law.

9. What are the characteristics of large numbers in statistics?

Answer:​
In statistics, the Law of Large Numbers says that as the number of observations or trials
increases, the average result tends to get closer to the expected or true value.

Example:​
If you toss a fair coin 10 times, you might get 7 heads. But if you toss it 10,000 times, the
percentage of heads will generally get closer to 50%.

Reason:​
Larger samples generally provide more stable and reliable estimates.

10. What is the difference between Data Science and Statistics?

Answer:​
Statistics mainly focuses on collecting, analyzing, interpreting, and drawing conclusions from
data.

Data Science is broader. It combines statistics with programming, databases, machine learning,
data visualization, and other technologies.

Example:
A statistician might analyze whether two variables have a significant relationship.

A data scientist might collect the data, clean it using Python, analyze it, build a machine-learning
model, and visualize the results.

11. What are cherry-picking, p-hacking, and significance chasing?

Answer:​
These are practices that can lead to misleading statistical conclusions.

Cherry-picking: Selecting only the data or results that support a particular conclusion.

P-hacking: Trying many analyses or statistical tests until obtaining a statistically significant
result.

Significance chasing: Repeatedly changing analyses or conditions simply to obtain statistical


significance.

Example:​
If a researcher performs 20 different tests and reports only the one with p < 0.05, the result may
appear stronger than it really is.

Reason:​
These practices can increase the chance of false or unreliable conclusions.

12. What is the difference between Type I and Type II errors?

Answer:

Error Meaning

Type I error Rejecting a true null hypothesis

Type II error Failing to reject a false null hypothesis

Example:

Suppose a medical test is checking whether a person has a disease.

●​ Type I: Test says the person has the disease when they actually don't.
●​ Type II: Test says the person doesn't have the disease when they actually do.
Easy way to remember:

●​ Type I = false positive


●​ Type II = false negative

13. What is statistical interaction?

Answer:​
A statistical interaction occurs when the effect of one variable on an outcome depends on the
value of another variable.

Example:​
Suppose we study the effect of a training program on performance.

The training might improve performance more for beginners than for experienced workers.

Here, the effect of training depends on experience, so there is an interaction between training
and experience.

14. What are examples of datasets with non-Gaussian distributions?

Answer:​
A non-Gaussian distribution is a distribution that does not follow the normal or bell-shaped
distribution.

Examples include:

●​ Income data
●​ House prices
●​ Number of website visits
●​ Waiting times
●​ Insurance claims
●​ Number of customers per day

Example:​
Income data can be strongly right-skewed because a small number of people may have
extremely high incomes.

Reason:​
Many real-world datasets are not normally distributed.
15. How does linear regression work?

Answer:​
Linear regression is a statistical method used to model the relationship between a dependent
variable and one or more independent variables.

For simple linear regression, we commonly write:

y = a + bx

where:

●​ y = predicted value
●​ x = input variable
●​ a = intercept
●​ b = slope

Example:​
We could use hours studied to predict exam marks.

If the model finds:

Marks = 40 + 5 × Hours

then a student who studies 4 hours would have a predicted mark of:

40 + (5 × 4) = 60

Reason:​
Regression can help us understand relationships and make predictions.

16. What are the necessary conditions for a Binomial Distribution?

Answer:​
A binomial distribution generally requires:

1.​ A fixed number of trials.


2.​ Each trial has only two possible outcomes, such as success or failure.
3.​ The probability of success remains the same for each trial.
4.​ The trials are independent.

Example:​
Flipping a fair coin 10 times and counting the number of heads.

Each flip has:


●​ Fixed number of trials = 10
●​ Two outcomes = head/tail
●​ Same probability of head = 0.5
●​ Independent trials

Therefore, the number of heads follows a binomial distribution.

17. What are the different kinds of variables or levels of measurement?

Answer:​
There are four common levels of measurement:

1.​ Nominal – categories with no order.


2.​ Ordinal – categories with an order.
3.​ Interval – numerical values with meaningful differences but no true zero.
4.​ Ratio – numerical values with meaningful differences and a true zero.

Examples:

●​ Nominal → eye color


●​ Ordinal → satisfaction: poor, average, good, excellent
●​ Interval → temperature in Celsius
●​ Ratio → weight, height, age

Reason:​
The measurement level determines which statistical operations are appropriate.

18. What is the difference between Descriptive and Inferential Statistics?

Answer:

Descriptive statistics summarizes and describes the data we already have.

Examples:

●​ Mean
●​ Median
●​ Mode
●​ Range
●​ Standard deviation
●​ Charts
Inferential statistics uses sample data to make conclusions or predictions about a larger
population.

Example:

If we calculate the average marks of 100 students, that is descriptive statistics.

If we use those 100 students to estimate the average marks of all students in a university,
that is inferential statistics.

Easy interview answer:

"Descriptive statistics tells us what the data shows, while inferential statistics helps
us make conclusions about a larger population based on sample data."

Yes. I compared these 30 Statistics questions with the Statistics questions you already gave
me earlier.

You already covered several topics, so I would remove the duplicates and study these new
questions:

📊 Statistics — New/Remaining Questions


1. What is sampling?

Answer:​
Sampling is the process of selecting a smaller group, called a sample, from a larger population
to study the population.

Example:​
If a university has 10,000 students and we survey 500 students, the 500 students are the
sample.

Why?​
Studying the entire population may be expensive or time-consuming, so sampling makes data
collection easier.

2. How do you determine the statistical significance of an insight?


Answer:​
Statistical significance is commonly determined using a hypothesis test and p-value. We
compare the p-value with a chosen significance level, such as 0.05.

Example:​
If p = 0.03 and our significance level is 0.05, the result is statistically significant.

Why?​
It helps determine whether an observed result provides enough evidence against the null
hypothesis.

3. How do you define Exploratory Data Analysis (EDA)?

Answer:​
Exploratory Data Analysis is the process of examining and understanding a dataset before
performing detailed analysis or modeling.

It can involve:

●​ Checking missing values


●​ Finding outliers
●​ Calculating summary statistics
●​ Creating charts
●​ Looking for patterns and relationships

Example:​
Before analyzing student marks, I might check the average, highest and lowest marks, missing
values, and distribution of scores.

Why?​
EDA helps us understand the data and identify problems before drawing conclusions.

4. What is selection bias?

Answer:​
Selection bias occurs when the way a sample is selected causes it to be unrepresentative of
the population.

Example:​
If we want to know how all university students feel about online classes but only survey
computer science students, our sample may not represent all students.
Why?​
Selection bias can produce misleading conclusions.

5. What are the applications of long-tailed distributions?

Answer:​
A long-tailed distribution occurs when a small number of observations have very large values
compared with most observations.

Examples:

●​ Income
●​ Website traffic
●​ Book sales
●​ Social media followers
●​ Insurance claims

Example:​
Most websites may have relatively few visitors, while a small number of websites receive
millions of visitors.

Why?​
Understanding long-tailed distributions is useful for business analysis, risk analysis,
recommendation systems, and understanding extreme values.

6. When is the median more suitable than the mean?

Answer:​
The median is more suitable when the data contains extreme values or is highly skewed.

Example:

Salaries:

30k, 35k, 40k, 45k, 500k

The 500k salary greatly increases the mean, while the median remains more representative of
the typical salary.

Why?​
The median is less affected by extreme values.
7. How might root cause analysis be applied to a real-life situation?

Answer:​
Root cause analysis is a process used to identify the underlying reason why a problem
occurred.

Example:​
Suppose a company has many late deliveries.

We could investigate:

Late deliveries → transportation problems → vehicle breakdowns → poor vehicle


maintenance

The root cause might be inadequate maintenance.

Why?​
Fixing the root cause can prevent the problem from happening repeatedly.

8. How does Design of Experiments (DOE) work?

Answer:​
Design of Experiments is a structured method for planning experiments to understand how
different factors affect an outcome.

Example:​
A company wants to determine which factors affect product quality. It might test:

●​ Temperature
●​ Material
●​ Machine speed

and measure the resulting product quality.

Why?​
DOE helps identify which factors have important effects while reducing unnecessary
experiments.

9. What does standard deviation mean?


Answer:​
Standard deviation measures how much values typically vary from the mean.

Example:

Dataset A:

49, 50, 51

has a small standard deviation.

Dataset B:

20, 50, 80

has a larger standard deviation.

Why?​
It helps us understand how spread out the data is.

10. What are the characteristics of a bell-curve distribution?

Answer:​
A bell curve usually refers to a normal distribution. Its main characteristics are:

●​ It is symmetric.
●​ It has one central peak.
●​ Mean, median, and mode are equal.
●​ Most observations are near the center.
●​ Values become less common toward the tails.

Example:​
Some naturally occurring measurements, such as certain measurement errors, can
approximately follow a bell-shaped distribution.

11. What is skewness?

Answer:​
Skewness describes the asymmetry of a probability distribution.

There are three common cases:


●​ Positive/right skew: longer tail on the right.
●​ Negative/left skew: longer tail on the left.
●​ Approximately symmetric: both sides are similar.

Example:​
Income distributions are often right-skewed because a small number of people have very high
incomes.

12. What is kurtosis?

Answer:​
Kurtosis is a statistical measure related to the shape of a distribution, particularly the heaviness
of its tails and tendency to produce extreme values.

Example:​
A distribution with heavier tails can have more extreme observations than a normal distribution.

Interview-level answer:

"Kurtosis helps describe the tail behavior of a distribution and its tendency to
produce extreme values."

13. What is correlation?

Answer:​
Correlation measures the strength and direction of a relationship between two variables.

It is commonly represented by a correlation coefficient ranging from -1 to +1.

●​ +1 → perfect positive linear relationship


●​ 0 → no linear relationship
●​ -1 → perfect negative linear relationship

Example:​
If study time and exam marks tend to increase together, they may have a positive correlation.

14. What are left-skewed and right-skewed distributions?

Answer:
Right-skewed:​
The distribution has a longer tail toward larger values.

Example: Income.

Left-skewed:​
The distribution has a longer tail toward smaller values.

Example: If most students score very highly but a few score very low, the distribution may be
left-skewed.

Easy memory trick:​


The tail tells you the direction of the skew.

15. What is covariance?

Answer:​
Covariance measures how two variables change together.

●​ Positive covariance → variables tend to increase together.


●​ Negative covariance → one tends to increase while the other decreases.
●​ Near-zero covariance → little linear co-movement.

Example:​
If study hours increase and marks also tend to increase, they may have positive covariance.

Difference from correlation:​


Covariance indicates the direction of the relationship, while correlation also provides a
standardized measure of its strength.

16. What is Bessel's correction?

Answer:​
Bessel's correction uses n − 1 instead of n when calculating the sample variance.

For a sample:

[​
s^2 = \frac{\sum(x_i-\bar{x})^2}{n-1}​
]
Example:​
If we calculate variance from a sample of 10 observations, we divide by 9, not 10.

Why?​
It helps reduce the bias when using a sample to estimate the population variance.

17. What are inferential statistics used for?

Answer:​
Inferential statistics are used to make conclusions or estimates about a population based on
sample data.

Examples:

●​ Hypothesis testing
●​ Confidence intervals
●​ Estimating population parameters
●​ Regression analysis

Example:​
We survey 500 voters and use the results to estimate the preferences of a much larger
population.

18. How are mean and median related in a normal distribution?

Answer:​
In a perfectly normal distribution:

Mean = Median = Mode

They are all located at the center of the distribution.

Example:

2, 3, 4, 5, 6

Mean = 4​
Median = 4

19. What is the relationship between standard error and margin of error?
Answer:​
Standard error (SE) measures the variability of a sample statistic, such as a sample mean,
across repeated samples.

The margin of error (MOE) is generally calculated by multiplying the standard error by a critical
value.

For example:

[​
MOE = Critical\ Value \times SE​
]

Example:​
For a 95% confidence interval, the critical value is often approximately 1.96 when using a
normal approximation.

Why?​
A smaller standard error generally produces a smaller margin of error and therefore a more
precise estimate.

20. What does degree of freedom (DF) represent?

Answer:​
Degrees of freedom represent the number of independent pieces of information available for
estimating a statistical quantity.

Simple example:​
Suppose we have 5 values and calculate their mean. Once the mean is fixed, only 4 values can
vary freely because the fifth value is determined by the required total.

Therefore:

DF = 5 − 1 = 4

Why?​
Degrees of freedom appear in many statistical calculations, including sample variance and
t-tests.

21. How do you explain the Law of Large Numbers?


Answer:​
The Law of Large Numbers says that as the number of observations or trials increases, the
average result tends to get closer to the expected value.

Example:​
If you toss a fair coin 10 times, you might get 7 heads. If you toss it 10,000 times, the proportion
of heads will generally get closer to 50%.

22. How does TF-IDF vectorization relate to meaning?

Answer:​
TF-IDF (Term Frequency-Inverse Document Frequency) is a technique used in text analysis
to represent words numerically.

It gives greater importance to words that:

●​ Appear frequently in a particular document


●​ But are relatively uncommon across the whole collection of documents

Example:​
In a collection of technology articles, the word "technology" might appear frequently across
many documents, so it may receive less weight than a more specific term.

Why?​
TF-IDF helps computers identify important words in documents.

This is mainly a Natural Language Processing/Data Science concept rather than


basic statistics.

23. What is the purpose of hash tables in statistics?

Answer:​
A hash table is a data structure used to store key-value pairs and retrieve information
efficiently.

It is not specifically a statistical method, but it can be useful when processing statistical data.

Example:​
We could use a dictionary/hash table to count how frequently different categories occur:

counts = {
"A": 10,

"B": 15,

"C": 7

Why?​
Hash tables can make lookup and counting operations very efficient.

24. Do symmetric distributions need to be unimodal?

Answer:​
No.

A distribution can be symmetric but have more than one peak.

Example:​
Imagine a distribution with two equally sized peaks positioned symmetrically around the center.
It can be symmetric and bimodal.

Key point:

Symmetric means the two sides mirror each other; unimodal means there is one
main peak. They are different properties.

25. Is there any significance to outliers?

Answer:​
Yes. Outliers can be important because they may represent:

●​ Data-entry errors
●​ Measurement errors
●​ Rare events
●​ Genuine extreme observations

Example:​
If most customers spend $20–$100 but one customer spends $10,000, we should investigate
that value.
Why?​
Outliers can affect the mean, standard deviation, correlation, and regression results.

26. What is central tendency?

Answer:​
Central tendency describes the central or typical value of a dataset.

The three common measures are:

●​ Mean
●​ Median
●​ Mode

Example:

2, 3, 3, 4, 8

Mean = 4​
Median = 3​
Mode = 3

Why?​
Central tendency gives us a simple way to describe a dataset using a representative value.

27. What is a normal distribution?

Answer:​
A normal distribution is a continuous probability distribution with a symmetric, bell-shaped
curve.

Its main characteristics include:

●​ Mean = median = mode


●​ Symmetric around the mean
●​ Most observations are near the center
●​ Fewer observations occur toward the tails

Example:​
Some measurement-related datasets can be approximately normally distributed.
28. What is the empirical rule?

Answer:​
The empirical rule, also called the 68–95–99.7 rule, describes approximately how data is
distributed in a normal distribution.

●​ About 68% of values are within 1 standard deviation of the mean.


●​ About 95% are within 2 standard deviations.
●​ About 99.7% are within 3 standard deviations.

Example:​
If exam scores follow a normal distribution with:

Mean = 70​
Standard deviation = 10

Approximately 68% of scores would fall between:

60 and 80

You might also like