0% found this document useful (0 votes)
7 views39 pages

Teenagers' Fizzy Drink Consumption Study

The document contains various statistical problems and concepts, including calculations of mean, median, mode, and measures of dispersion for different datasets. It also discusses data types, visualization techniques, and the application of the Bayes classifier in decision-making. Additionally, it includes questions about probability, skewness, and the interpretation of statistical measures in real-world scenarios.

Uploaded by

Soham Durge
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views39 pages

Teenagers' Fizzy Drink Consumption Study

The document contains various statistical problems and concepts, including calculations of mean, median, mode, and measures of dispersion for different datasets. It also discusses data types, visualization techniques, and the application of the Bayes classifier in decision-making. Additionally, it includes questions about probability, skewness, and the interpretation of statistical measures in real-world scenarios.

Uploaded by

Soham Durge
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

[Link]

v=RbyUuJNd50E

1.​ The number of cans of fizzy drinks consumed by teenagers each day is the subject of an empirical
study. The following data have been recorded:

Cans per day 0 1 2 3 4 5

Number of teenagers 2 3 26 20 1 10

Assume that no teenager drinks more than five cans per day.

(i) Calculate the mean, median and mode for this sample.

●​ Associate english sentence while writing each central tendency and dispersion like
“teenagers prefer drinking 2 cans each day”
●​ Observe variation in the consumption of fizzy drink with different methods

(ii) Comment on the symmetry of the observed data, using your answer to part (i) and without
making any further calculations.

1) Data collected on the eye colors (blue, brown, green, hazel) of a population represents which
type of data scale?

2) The most appropriate measure of central tendency for a highly skewed distribution with several
extreme outliers is the ________.

3) A dataset contains the annual salaries (in dollars) of all employees at a tech firm. Which
measure of dispersion, **Range** or **Interquartile Range (IQR)**, would be more robust to the
influence of a few executive salaries that are significantly higher than the rest?

4) Explain the primary difference in the visual insight gained from a **boxplot** compared to a
**histogram** when analyzing a univariate, continuous dataset.

5) When analyzing the relationship between two continuous variables, a **scatter plot** is
generally preferred over a **stacked bar chart**. (True/False) Justify your answer.

6) Which of the following data types supports the calculation of a meaningful **Mean** and
**Variance**?
a) Ordinal
b) Nominal
c) Interval
d) Categorical

7) You are tasked with analyzing the distribution of commute times (in minutes) for employees.
The resulting plot shows a long tail to the right. This suggests the distribution is exhibiting:
a) Negative skewness
b) Zero skewness
c) Positive skewness
d) Platykurtosis
8) Describe the calculation used to define the upper fence for outlier detection based on the
**Interquartile Range (IQR) rule**.

9) To visually compare the market share percentage of five different smartphone brands, the most
effective visualization is typically a **pie chart**.

10) A **stacked bar chart** is most effective for visualizing the relationship between:
a) Two continuous variables.
b) A nominal variable and a ratio variable.
c) Two categorical variables.
d) A time series variable and a continuous variable.

11) Explain the rationale for why the **median** is a more appropriate measure of central
tendency than the **mean** for describing house prices in a neighborhood where a few luxury
homes are disproportionately expensive compared to the majority.

12) A researcher is analyzing customer ratings (1 to 5 stars) for a product. These ratings constitute
an example of an **ordinal** data type. (True/False)

13) Consider a university dataset recording the temperature (in Celsius) of laboratory samples and
the type of experiment (A, B, C) performed. Justify the selection of a **heatmap** as a primary
visualization tool for exploring the relationship between these two attributes.

14) Describe a scenario where a **geographical heatmap** would be the optimal visualization
choice, specifying the attribute being mapped and the insight it aims to communicate to a research
audience.

15) A marketing team is comparing the distribution of purchase amounts (in dollars) for
customers acquired through two different campaigns (Campaign X and Campaign Y). Outline the
steps and appropriate visualizations/statistics required to synthesize a comprehensive summary of
both datasets' characteristics, variability, and structure for a research audience.

2) Given the following dataset of wait times (in minutes) at a clinic: 5,7,10,12,12,15,20. Calculate
the **range** and the **median** for this univariate dataset.

3) A small company reports employee salaries (in thousands of dollars) as:


40,45,50,52,55,60,200. Compare and contrast the **mean** and **median** for this dataset and
explain which measure better represents the typical salary, justifying your choice based on data
distribution.

4) Describe a scenario where a **stacked bar chart** would be a more effective visualization for
communicating insights than a simple **bar chart** or a **pie chart**. Specify the type of data
attributes you would need.

5) A survey asks participants to rank their satisfaction with a new service on a scale of 1 (Very
Dissatisfied) to 5 (Very Satisfied). What data type is this, and why would the **Interquartile
Range (IQR)** be a preferred measure of dispersion over the **variance**?
6) Describe the steps and components required to construct a **boxplot** for a dataset. Explain
how a boxplot visually communicates the **skewness** of the data distribution.

7) You have a dataset of house prices. If you discover a strong positive skew, what does that
imply about the relationship between the **mean**, **median**, and **mode**, and how
should this affect your choice of a measure of central tendency for reporting?

8) A scientist records the temperature (in Celsius) of a chemical reaction at one-minute intervals.
What data type is this, and what is the most appropriate visualization (out of **histogram** and
**bar chart**) to display the frequency distribution of the temperature readings? Justify your
choice.

9) The Q1 (first quartile) for a list of student exam scores is 65, and the Q3 (third quartile) is 85.
Use the **$1.5 \times \text{IQR}$ rule** to calculate the lower and upper bounds for outlier
detection. Explain the significance of an outlier score of 120.

10) You are analyzing the relationship between advertising spend and sales revenue. Describe the
two variables and the appropriate visualization (**scatterplot**, **pie chart**, or
**histogram**) to determine if a linear relationship exists between them.

11) Explain the difference between **variance** and **standard deviation** as measures of
dispersion. Why is the standard deviation generally preferred over the variance when presenting
descriptive statistics to a general research audience?

12) Describe a research scenario involving two categorical variables where a **heatmap** would
be a superior visualization choice compared to a **bar chart**. Explain what the color intensity
in the heatmap would represent.

13) Imagine you have a dataset detailing COVID-19 case rates per capita for every county in a
state. Explain how you would generate a **geographical heatmap** by hand, detailing the
necessary data and the interpretation of the visual output for a spatial analysis.

14) A dataset contains customer ages, product preferences (low, medium, high), and annual
income. Select one measure of central tendency and one measure of dispersion for **annual
income** (ratio data). Justify your choices assuming the income distribution is roughly
symmetrical.

15) A political scientist uses a **pie chart** to display the percentage of voters for each of five
candidates in an election. Critically evaluate the appropriateness and effectiveness of the pie chart
for this purpose, proposing an alternative visualization and justifying the replacement.

2.​ Which of the following statement is TRUE about the Bayes classifier?
a.​ Bayes classifier works on the Bayes theorem of probability.
b.​ Bayes classifier is an unsupervised learning algorithm.
c.​ Bayes classifier is also known as maximum apriori classifier.
d.​ It assumes the independence between the independent variables or features.

3.​ True or False: In a naive Bayes algorithm, when an attribute value in the testing record has no
example in the training set, then the entire posterior probability will be zero.
a.​ True
b.​ False
c.​ Can’t determined

4.​ If P(A) = 0.25, P(B) = 0.6 and P(A∩B) = 0.5. Find P(A/B).
5.​ If P(A) = 0.3, P(B/A) = 0.4 and P(A∪B) = 0.5 then find P(B).
6.​ A dice is rolled. If X = {3, 5, 6}, Y = {3, 4} and Z = {1, 3, 6} then, find P(X/Y), P(Y/Z), P(X/Z).
7.​ A coin is tossed 4 times. Find P(C/D) in C = Head on second Toss and D = Tail on third toss
8.​ Given that the two numbers appearing on throwing two dice are different. Find the probability of
the event the sum of numbers appearing on dice is 8 and the number 4 appears once.
9.​ Evaluate P(X∪Y), if 2P(X) = P(Y) = 4/11 and P(X/Y) = 2/9.
10.​ If P(A) = 5/8, P(B) = 2/3 and P (A B) = 3 / 8. Find P (Ac/ Bc).
11.​ In a school there are 120 students out of which 60 are boys. It is known that out of 40, 10% of
boys study in class 11. What is the probability that a student chosen randomly studies in class 11 ,
given that the chosen student is a boy.
12.​ A coin is tossed 5 times. Find (X/Y) if X = at least two heads and Y = at most one tail.
13.​ A coin is tossed 4 times. Find (X/Y) if X = at most one head and Y = at most 3 tails.
14.​ The following table shows hypothetical data concerning student characteristics and whether or
not each student should be hired.

Name GPA Effort Hirable

Sarah poor lots Yes

Dana average some No

Alex average some No

Annie average lots Yes

Emily excellent lots Yes

Pete excellent lots No

John excellent lots No

Kathy poor some No


Using the Naïve Bayes classifier, answer the following questions:

1.​ Compute the prior probabilities of hiring (Yes/No).

2.​ Compute the conditional probabilities for GPA and Effort given Hirable = Yes/No.​

3.​ For a student with the following characteristics:​

○​ GPA = poor​

○​ Effort = lots

determine whether they are more likely to be Hirable = Yes or Hirable = No.
4.​ Show your step-by-step calculations.

5.​ What is your final decision — should the student be hired?

15.​ A feature F1 can take a certain value: A, B, C, D, E, or F, which represents the grades of students
from a college.

Which of the following statement is true in the following case?

a.​ Feature F1 is an example of a nominal variable.


b.​ Feature F1 is an example of an ordinal variable.
c.​ It doesn’t belong to any of the above categories.
d.​ Both of these

Ordinal variables are the variables that have some order in their categories. For example, grade A
should be considered a high grade than grade B.

16.​ Match List I with List II and choose the correct answer from the options given below:

List I (Scale List II (Data type measured)


system)

(A) Nominal (I) Altitude from mean sea level

(B) Ordinal (II) Soil type

(C) Interval (III) Monthly income (in Rs.)

(D) Ratio (IV) Town size

a.​ (A) – (II), (B) – (IV), (C) – (I), (D) – (III)


b.​ (A) – (III), (B) – (II), (C) – (I), (D) – (IV)
c.​ (A) – (III), (B) – (IV), (C) – (I), (D) – (II)
d.​ (A) – (II), (B) – (III), (C) – (I), (D) – (IV)

17.​ For a negatively skewed distribution, the median would be equal to which of the following
measures?

A. Mean​
B. Mode​
C. Second quartile​
D. Fifth quartile​
E. 50th percentile
Choose the most appropriate answer from the options given below:

a.​ A, B and C only


b.​ B, C and D only
c.​ C and E only
d.​ A, B, C, D and E

18.​ Identify the type of data (nominal, ordinal, interval, or ratio) represented by each of the following.
Confirm your answers by giving your own examples.
a.​ What type of data is Blood group?
b.​ What type of data is Temperature (Celsius)?
c.​ What type of data is Ethnic group?
d.​ What type of data is Job satisfaction index (1–5)?
e.​ What type of data is Number of heart attacks?
f.​ What type of data is Calendar year?
g.​ What type of data is Serum uric acid (mg/100 ml)?
h.​ What type of data is Number of accidents in a 3-year period?
i.​ What type of data is Number of cases of each reportable disease reported by a health
worker?​

19.​ A man travels from Jaipur to Agra by a car and takes 4 hours to cover the whole distance. In the
first hour he travels at a speed of 50 km/hr, in the second hour his speed is 64 km/hr, in third hour
his speed is 80 km/hr and in the fourth hour he travels at the speed of 55 km/hr. Find the average
speed of the motorist.

20.​ Three rotten apples are mixed accidently with seven good apples and four apples are drawn one
by one without replacement. Let the random variable X denote the number of rotten apples. If μ
and σ² represent the mean and variance of X, respectively, then 10(μ^2+σ^2) is equal to:
a.​ 20
b.​ 250
c.​ 25
d.​ 30
21.​

22.​ Ervin bowled 7 games last weekend. His scores are: 155, 165, 138, 172, 127, 193, 142. What is
the sample standard deviation of Ervin's scores?
a.​ 511.33
b.​ 22.61
c.​ 438.48
d.​ 20.94
23.​ Ervin bowled 7 games last weekend. His scores are: 155, 165, 138, 172, 127, 193, 142. What is
the population variance of Ervin's scores?
a.​ 511.33
b.​ 22.61
c.​ 438.29
d.​ 20.94
24.​ What is the mathematical relationship between the variance and the standard deviation?
a.​ The variance is the difference between the standard deviation and the average.
b.​ The standard deviation is the square of the difference between the variance and the
average.
c.​ The variance is the square root of the standard deviation.
d.​ The standard deviation is the square root of the variance.
25.​ If you want to invest in a new stock but get very stressed out when you see large fluctuations in
the price, even if the price is generally increasing, which of the following stocks (all with an
average of around $30) is best for you?
a.​ DGH, with a standard deviation of 1.7
b.​ JKL, with a standard deviation of 3.6
c.​ OPQ, with a standard deviation of -1.3
d.​ WXY, with a standard deviation of 1.1
26.​ What is the general category of metrics that describe numeric data?
a.​ Standard deviation
b.​ Dependent variables
c.​ Descriptive statistics
d.​ Variance
27.​ What is the ordinary average of a group of numbers called?
a.​ Median
b.​ Mode
c.​ Standard Deviation
d.​ Mean
28.​ In which type of distribution is the median greater than the mean?
a.​ Normal
b.​ Skewed to the right
c.​ Skewed to the left
d.​ Symmetrical
29.​ Which statistic is usually used to describe the representative value for a nominal variable such as
religious affiliation?
a.​ Mode
b.​ Median
c.​ Outlier
d.​ Mean
30.​ Construct a box plot for the following data set.

3,5,8,8,9,11,12,12,13,13,16

31.​ The reaction times (in milliseconds) of a group of 20-year-olds and a group of 30-year-olds were
tested. The reaction times for the 20-year-olds has been plotted below:

The reaction times for the 30-year-olds are as follows:

220,252,256,312,332,332,400
Construct a box plot for this set of the data and note two differences between the two groups.

32.​ Which of the following about Naive Bayes is incorrect?


a.​ Attributes can be nominal or numeric
b.​ Attributes are equally important
c.​ Attributes are statistically dependent of one another given the class value
d.​ Attributes are statistically independent of one another given the class value
e.​ All of above
33.​ Given all the previous patients I’ve seen (below are their symptoms and diagnosis):
Chills Runny Nose Headache Fever Flu?

Y N Mild Y Y

Y Y No N Y

Y N Strong Y Y

Y Y Mild N N

N N No N N

N Y Strong Y Y

N Y Strong N Y

N Y Mild Y Y

Do I believe that a patient with the following symptoms has the flu?

Chills = Y, Runny Nose = N, Headache = Mild, Fever = Y, Flu = ?​

34.​ If the distribution is negatively skewed, then:


a.​ Mean is more than the mode
b.​ Median is at right to the mode
c.​ Mean is less than the mode
d.​ Mean is at right to the median
35.​ For a frequency distribution of a variable xxx, mean = 32, median = 30.​
The distribution is:
a.​ Positively skewed
b.​ Negatively skewed
c.​ Mesokurtic
d.​ Platykurtic
36.​ Which of these events are not mutually exclusive?
a.​ Rolling a 3 and a 4 on a die.
b.​ Rolling an even number and a 5 on a die.
c.​ Rolling an even number and a prime number on a die.
d.​ Rolling an even number and a 3 on a die.
37.​ Which of these events are mutually exclusive?
a.​ Selecting a prime number and a factor of 12 from the numbers 1 to 12.
b.​ Selecting a consonant and a vowel from the alphabet.
c.​ It raining and being sunny on a given day.
d.​ Rolling a multiple of 2 and a multiple of 4 on a fair six-sided die.
38.​ What is the probability of the following spinner landing on a 1,2 or 3?

39.​ Lucy has a box of chocolates containing milk, white and dark chocolates. The probability of
picking a milk chocolate from the box is ½ and the probability of picking a milk chocolate or a
white chocolate is 4/5. What is the probability of picking a white chocolate from the box?
40.​ Which sequences correctly represent the relationships between the mean, median, and mode in
right-skewed and left-skewed distributions?​ [0.5]
a.​ Right Skewed: Mean < Median < Mode; Left Skewed: Mean < Mode < Median
b.​ Right Skewed: Mode < Median < Mean; Left Skewed: Mean < Median < Mode
c.​ Right Skewed: Mean < Median < Mode; Left Skewed: Mode < Median < Mean
d.​ Right Skewed: Mode < Mean < Median; Left Skewed: Mode < Median < Mean
41.​ The probability that a smoker gets cancer is 0.8, and a non-smoker gets cancer is 0.1. What are
the chances that a person actually smokes if he has cancer? Numerous research confirms that 30%
of people smoke.
42.​ For the given data 85, 72, 90, 68, 88, 75, 92, 78, 80, 85, 70, 95, 82, 76, 88, 91, 74, 83, 87, 79
a.​ Given a dataset of numbers, calculate the range, mean(expected value), median, sample
skewness, population skewness, kurtowsis and mode.
b.​ Calculate the standard deviation, variance and coefficient of variation of the given
dataset.
c.​ Find the quartiles (Q1, Q3) and IQR of a given dataset.
d.​ Determine if any outliers exist in a given dataset using the IQR method.
e.​ Calculate the skewness of a given dataset.
f.​ Given a dataset of grouped data (in a frequency table), calculate the mean.
g.​ Calculate the weighted mean of a set of values with their corresponding weights.
43.​ A professor recorded marks of 100 students. The mean score is 60, and the median score is 58.​
Using the empirical relation between mean, median, and mode, estimate the mode of the
distribution.
a.​ 56
b.​ 54
c.​ 62
d.​ 64
44.​ The weights of 10-year-old girls are known to be normally distributed with a mean of 70 pounds
and a standard deviation of 13 pounds. Find the percentage of 10-year-old girls with weights
between 60 and 90 pounds.​
45.​ A multiple-choice test has 10 questions, each with 4 options. A student guesses all answers.
1. What is the probability that the student gets exactly 3 correct answers?
2. What is the expected number of correct answers?
46.​ A factory produces steel rods with lengths normally distributed with a mean of 100 cm and a
standard deviation of 2 cm. The factory rejects rods shorter than 96 cm or longer than 104 cm.
1. What percentage of rods are rejected?
2. If the factory wants to reduce the rejection rate to 5%, what should the new acceptable range of
lengths be?
47.​ Consider following joint distribution mass function

Y=4 Y=10

X=1 0.1 0.01

X=2 0.25 0.4

X=3 0.09 0.15

a.​ Find Marginal mass functions


b.​ Correlation between variables
c.​ Covariance matrix
d.​ What would be the PMF, if they are independent variables?
48.​ For the given probability distribution find population mean, variance and standard deviation

Outcome(X) 1 2 3 4 5 6
PMF(X); P(x) 0.2 0.1 0.2 0.2 0.2 0.1
CDF(X);P(X<=x) 0.2 0.3 0.5 0.7 0.9 1

49.​What is the probability density function values at X = 0.8, if all possible outcomes
(0.5<=X<=0.9) of a stochastic event are equally likely to occur.
50.​ Mimic the numbers coming on a dice which exhibit following probability mass function

Outcome(X) 1 2 3 4 5 6
PMF(X); P(x) 0.2 0.1 0.2 0.2 0.2 0.1
CDF(X);P(X<=x) 0.2 0.3 0.5 0.7 0.9 1

a.​ Generate 100 numbers from such a dice. Check how close your distribution is to the above
probability distribution
b.​ Calculate sample and population variance
c.​ For a fixed distribution, repeat the above experiments 100 times, check how many times which
formula comes better for calculating variance (with denominator n or n-1)
51.​ If a fair coin is tossed twice, what is the probability of getting two heads?
52.​ A bag contains 5 red balls and 3 green balls. What is the probability of drawing a red ball?
53.​ Given P(A) = 0.4, P(B) = 0.3, and P(A ∩ B) = 0.12, are events A and B independent?
54.​ Calculate the mean and variance of a binomial distribution with n=10 and p=0.5.
55.​ If the probability of rain on any given day is 0.2, what is the probability that it will rain on exactly
2 out of 5 days?
56.​ Given P(A) = 0.6, P(B|A) = 0.8, find P(A ∩ B).
57.​ A box contains 5 red balls and 3 blue balls. Two balls are drawn without replacement. What is the
probability that both balls are red? And what is the probability when balls are drawn without
replacement?
58.​ Calculate the expected value of a random variable X with the following probability distribution:
a.​ | X | 1 | 2 | 3 |
b.​ |------|-----|-----|-----|
c.​ | P(X) | 0.2 | 0.5 | 0.3 |
59.​ If a continuous random variable X follows a uniform distribution between 0 and 1, find the
probability that 0.2 ≤ X ≤ 0.8.

60.​ What is the sum of probabilities of all possible outcomes in a sample space?

a) 0

b) 1

c) Infinity

d) It varies depending on the experiment

61.​ Two events are independent if:


a) The occurrence of one event affects the probability of the other

b) The occurrence of one event does not affect the probability of the other

c) They are mutually exclusive

d) Their probabilities sum to 1

62.​ The Naive Bayes algorithm assumes that:

a) All features are independent of each other

b) All features are dependent on each other

c) Only two features are relevant

d) The data is normally distributed


63.​ What is the expected value of a binomial distribution where n=16 and p=0.85?
a)​ 6
b)​ 7.4
c)​ 12.4
d)​ 13.6
64.​ Which of the following is FALSE about Normal Distribution?
1.​ Normal Distribution is applied for discrete Random Distribution​

2.​ The shape of the Normal Curve is Bell Shaped​

3.​ The area under a standard normal curve is 1​

4.​ The standard normal curve is symmetric about the value 0

65.​ Which one of the following is not the correct property of Normal Distribution?
1.​ Continuous distribution​

2.​ Equality of central values (Mean, Mode, and Median)​

3.​ Standard deviation is the sole parameter of the distribution ❌​


4.​ Uni-modal distribution

66.​ Match List I with List II:

List I (Applications)​
A. Height of individuals in a population​
B. Daily sales of a retail store​
C. Lifetimes of electronic devices before they fail​
D. Number of defects in a production line​
E. Number of selections from a group without replacement

List II (Distributions)​
I. Normal distribution​
II. Hypergeometric distribution​
III. Exponential distribution​
IV. Poisson distribution

a)​
b)​
A–I, B–III, C–III, D–IV, E–II
A–III, B–I, C–IV, D–II, E–I

c)​ A–I, B–II, C–III, D–IV, E–III
d)​ A–IV, B–III, C–I, D–II, E–I

67.​ A basketball game is played for 30 minutes. A coach claims that his team's players commit, on
average, no more than 10 fouls per game. Let µ represent the team's average number of fouls per
game. Another coach thinks that these players create more fouls. And in the next game the team
fouled 100 times."
What type of distribution do the fouls follow? What is the probability associated with 100 or
more fouls per game given the team’s coach’s statement is true?

68.​ A company has 1000 employees, consisting of: 700 Developers, 100 Sales & Marketing
employees, 200 QA Engineers. A random sample of 30 employees is selected. Let X be the
number of QA Engineers in this sample. Solve the following questions.
1.​ What is the underlying distribution? & What is the probability of selecting exactly 10 QA
Engineers?​ [0.5]
2.​ What is the probability of selecting at least 5 QA Engineers? & What is the probability of
selecting at most 8 QA Engineers?​ [0.5]

Given:

●​ Population mean (μ): 8.8


●​ Population variance (σ²): 0.22
●​ Sample variance (s²): 0.275
69.​ What is the population standard deviation?​
A) 0.4690​
B) 0.5244​
C) 0.22​
D) 0.275

Answer: A) 0.4690​
Explanation: Population standard deviation = √(population variance) = √0.22 ≈ 0.4690.

70.​ 2. What is the sample standard deviation?​


A) 0.4690​
B) 0.5244​
C) 0.22​
D) 0.275

Answer: B) 0.5244​
Explanation: Sample standard deviation = √(sample variance) = √0.275 ≈ 0.5244.

71.​ 3. Why might the sample variance (0.275) be larger than the population variance (0.22)?​
A) Due to sampling error​
B) Because the sample size is small​
C) Both A and B​
D) It should always be smaller

Answer: C) Both A and B​


Explanation: Sample variance can be larger or smaller than the population variance due to
sampling variability, especially with small sample sizes.

72.​ 4. If the sample size is n=20, what is the unbiased estimate of the population variance?​
(Note: The given sample variance is likely already unbiased, but this tests understanding.)​
A) 0.22​
B) 0.275​
C) 0.259​
D) 0.290

Answer: B) 0.275​
Explanation: The problem states "sample variance is 0.275", which is typically the unbiased
estimator (s²) for the population variance.

73.​ 5. What is the coefficient of variation (CV) for the population?​


A) 5.33%​
B) 5.96%​
C) 6.25%​
D) 7.12%

Answer: A) 5.33%​
Explanation: CV = (σ / μ) × 100 = (0.4690 / 8.8) × 100 ≈ 5.33%.

74.​ 6. What is the coefficient of variation (CV) for the sample?​


A) 5.33%​
B) 5.96%​
C) 6.25%​
D) 7.12%
Answer: B) 5.96%​
Explanation: Sample CV = (s / x̄) × 100. But note: the sample mean is not given, so we assume it
is close to μ=8.8. Thus, CV ≈ (0.5244 / 8.8) × 100 ≈ 5.96%.

75.​ 7. If the sample size is large, the sample variance should be:​
A) Equal to the population variance​
B) Larger than the population variance​
C) Smaller than the population variance​
D) Unrelated

Answer: A) Equal to the population variance​


Explanation: As sample size increases, the sample variance converges to the population variance.

76.​ 8. Which of the following is true about standard deviation?​


A) It is the square root of variance​
B) It is measured in the same units as the data​
C) Both A and B​
D) Neither A nor B

Answer: C) Both A and B​


Explanation: Standard deviation is indeed the square root of variance and has the same units as
the original data.

77.​ The mileage which car owners get with a certain kind of radial tyre is a random variable having
an exponential distribution with mean 40,000 km. Find the probabilities that one of these tires
will last (i) at least 20,000 km and (ii) at most 30,000 km.
78.​ A call center receives customer calls at a random rate throughout the day. The number of calls
received per hour follows which probability distribution?
a) Normal
b) Poisson
c) Binomial
d) Exponential
79.​ The following dataset contains information about customer usage of phone services, including the
total minutes used, the number of calls made, and the total charges incurred during the day.
However, some values are missing (NaN). The value of k for KNN Imputation is fixed at k=2.
Write only important values that will be used to arrive the solution.​ [1]
No TotalDayMinutes TotalDayCalls TotalDayCharge

1 100.0 30.0 69.0

2 90.0 45.0 40.0

3 97.5 56.0 80.0


4 95.0 NaN 98.0

80.​ The cumulative frequency graph below shows the weight of 100 people who attend Weight
Watchers.

The weight of the lightest member was 61 kilograms and the weight of the heaviest member was 135. Draw
a box plot to show the distribution of the Weight Watchers members.
81.​ Which graph can utilise both quantitative and categorical variables at the same time? Justify your
answer.​
c. Boxplot
a. Histogram
d. Contingency table
b. Pie chart
82.​ A company has three offices across India: Delhi, Mumbai, and Kolkata. Each office has sent a
compiled list sharing sales in quarter 1, quarter 2, quarter 3, and quarter 4. Provide a visualisation
that presents the variation of sales within each city's quarters and also between the cities. Justify
your answer.​

b. Pie chart
c. Boxplot
d. Heatmap
a. Stacked bar chart
83.​ Which of the following visualizations is most effective at illustrating the distribution of a single
quantitative variable?​

a. Scatter plot
c. Histogram
b. Pie chart
d. Line chart
84.​ A soccer player successfully scores a penalty 75% of the time. If they take 10 penalties in a
match, what is the probability that they score at least 8 goals?
a) 0.2543
b) 0.3894
c) 0.6127
d) 0.7589

-----------------------------
More questions related to the subject will be added soon.

Feature Selection
{Feature Selection, Extraction, Curse of dimensionality, Filter method based on
correlation methods such as correlation coefficient, Spearman and Kendall Tau
coefficients, Null Space, Information Gain, Decision Tree, Wrapper Methods, Recursive
Feature Elimination method, Feature importance, Embedded methods, Regularization,
Overfitting, Underfitting}

1) Differentiate between feature selection and feature extraction in the context of


high-dimensional data, focusing on the transformation of the original feature space.

Correct answer: Feature selection chooses a subset of the original features, maintaining
their interpretability, while feature extraction creates new, lower-dimensional features
(e.g., principal components) that are combinations or transformations of the original
features, often sacrificing interpretability.

2) The primary mechanism by which both feature selection and feature extraction
mitigate the curse of dimensionality is by reducing the effective dimensionality of the
data space.

a) True
b) False

Correct answer: a) True

3) Which of the following is an example of a filter method for feature selection?

a) Recursive Feature Elimination (RFE)


b) L1 Regularization (Lasso)
c) Calculating the Spearman correlation coefficient between each feature and the target
variable
d) Using a greedy search algorithm to test model performance with different feature
subsets

Correct answer: c) Calculating the Spearman correlation coefficient between each


feature and the target variable
4) Explain how the Spearman rank correlation coefficient differs from the Pearson
correlation coefficient in its assessment of the relationship between a feature and the
target variable.

Correct answer: Pearson measures the linear relationship between the raw values of two
variables, assuming a Gaussian distribution. Spearman measures the monotonic
relationship between the ranks of the variables, making it more robust to non-linear
relationships and less sensitive to outliers, without assuming a specific distribution.

5) In the context of classification, Information Gain is often used as a feature selection


metric because a feature with high information gain is expected to significantly reduce
the dataset's overall ______.

Correct answer: Entropy

6) A data scientist is analyzing a dataset for a binary classification task. She finds that
Feature X₁ has a high Pearson correlation (0.85) with the target, while Feature X₂ has a
high Information Gain. Briefly explain why she might prefer to select Feature X₂ over X₁
based on these metrics.

Correct answer: Pearson correlation only measures linear relationships, and its
interpretation is complicated for non-continuous target variables (like binary
classification). Information Gain, derived from entropy reduction, is a non-linear,
non-parametric measure that directly assesses how well a feature splits the data based
on class labels, making it generally more appropriate and powerful for feature selection
in classification problems, especially when relationships are non-linear.

7) Recursive Feature Elimination (RFE) is considered a wrapper method because it


iteratively trains a specific machine learning model and uses the model's performance
and feature importance to guide the selection of the optimal feature subset.

a) True
b) False

Correct answer: a) True

8) When implementing RFE with a Support Vector Machine (SVM) as the base model,
what criterion is typically used at each step to determine which feature to eliminate?

a) The feature with the lowest Information Gain.


b) The feature with the smallest absolute magnitude of its corresponding weight vector
component.
c) The feature that, when removed, results in the highest increase in the model's
cross-validation score.
d) The feature with the highest p-value in a statistical significance test.

Correct answer: b) The feature with the smallest absolute magnitude of its corresponding
weight vector component.
9) The Null Space of a design matrix X consists of all vectors 𝐯 such that the
matrix-vector product X𝐯=𝟎. If a non-zero vector exists in the Null Space, what does this
imply about the features (columns) of X?

Correct answer: It implies that the features are linearly dependent (or redundant).
Specifically, the non-zero vector 𝐯 represents the coefficients of a linear combination of
the features that sums to zero, meaning at least one feature can be perfectly
represented as a linear combination of the others.

10) Which of the following regularization techniques inherently drives the coefficients of
irrelevant features to exactly zero, thus performing intrinsic feature selection?

a) L₂ Regularization (Ridge)
b) L₁ Regularization (Lasso)
c) Elastic Net Regularization
d) Dropout Regularization

Correct answer: b) L₁ Regularization (Lasso)

11) Explain the primary difference in the penalty applied to model coefficients by L₁
(Lasso) versus L₂ (Ridge) regularization and how this leads to different outcomes in
feature selection.

Correct answer: L₁ adds a penalty proportional to the sum of the absolute values of the
coefficients, which results in sparsity, driving the coefficients of less important features to
exactly zero (feature selection). L₂ adds a penalty proportional to the sum of the squares
of the coefficients, which shrinks the magnitude of all coefficients toward zero but rarely
makes them exactly zero, thus performing regularization but not explicit feature
selection.

12) A model with very high variance and low bias, which performs exceptionally well on
the training data but poorly on unseen test data, is suffering from ______.

Correct answer: Overfitting

13) If a predictive model exhibits high bias and low variance, failing to capture the
underlying relationship in the data for both training and test sets, applying L₂
regularization with a very large penalty term would likely:

a) Increase model variance and decrease bias, mitigating the problem.


b) Decrease both model variance and bias.
c) Worsen the problem by further increasing bias.
d) Act as an effective embedded feature selection method.

Correct answer: c) Worsen the problem by further increasing bias.

14) Briefly describe the mechanism by which feature importance values derived from a
trained Decision Tree (or ensemble methods like Random Forest) are calculated, and
how a practitioner might use these values for feature selection.
Correct answer: Feature importance in a Decision Tree is typically calculated based on
the total reduction in the loss function (e.g., Gini impurity or entropy for classification, or
mean squared error for regression) attributed to that feature across all splits in the tree.
A practitioner uses these values by setting a threshold and selecting only the features
whose importance score exceeds that threshold, as they contribute the most to the
model's predictive power.

15) Consider two features, F_(A) and F_(B), in a predictive model. If F_(A) and F_(B) are
highly correlated with each other, but only F_(A) is selected by an L₁ regularized model
while $F_B$'s coefficient is driven to zero, this illustrates the L₁ regularization's tendency
to:

a) Distribute the weight evenly between correlated features.


b) Select an arbitrary feature from a highly correlated group and discard the others.
c) Overfit the model by eliminating too many features.
d) Always select the feature with the highest individual correlation to the target.

Correct answer: b) Select an arbitrary feature from a highly correlated group and discard
the others.

Linear Regression
{Linear model, prediction, model, extending to multiple attributes, bias, weights,
optimizing model parameters, projection method, model hyper-parameters, model
parameters}

1) In the standard mathematical formulation of a linear regression model for a single


feature x, the predicted output y is given by y=w₀+w₁x. The term w₀ is conventionally
referred to as the __________.

Correct answer: bias (or intercept)

2) True or False: The primary purpose of the bias parameter (w₀) in a linear regression
model is to control the slope (steepness) of the regression line.

a) True
b) False

Correct answer: b) False

3) Which of the following mathematical expressions represents the Mean Squared Error
(MSE) objective function for a dataset with N samples, where y_(i) is the actual value
and y_(i) is the predicted value?

a) 1/N ∑(i=1)^(N)|y(i)-y_(i)|
b) ∑(i=1)^(N)(y(i)-y_(i))
c) 1/N ∑(i=1)^(N)(y(i)-y_(i))²
d) 1/N sqrt(∑(i=1)^(N)(y(i)-y_(i))²)

Correct answer: c) 1/N ∑(i=1)^(N)(y(i)-y_(i))²

4) Consider a simple linear regression model where the optimal weight w₁ is 3.5 and the
bias w₀ is -2. If a new data point has a feature value x=10, what is the model's prediction
for the output y?

Correct answer: 33

5) True or False: When extending univariate linear regression to multiple attributes, the
feature vector 𝐱 is typically augmented with an initial component of '1' to incorporate the
bias term in the matrix-vector multiplication.

a) True
b) False

Correct answer: a) True

6) For a multiple linear regression model with D features, the prediction y is represented
in vector-matrix notation as y=𝐰^(T)𝐱. The vector 𝐰 (weights and bias) and the
augmented feature vector 𝐱 must both have a dimensionality of __________.

Correct answer: D+1

7) Which of the following is the primary reason for using the squared difference
(y_(i)-y_(i))² rather than the absolute difference |y_(i)-y_(i)| in the Mean Squared Error
objective function for linear regression optimization?

a) The absolute difference is computationally more expensive to calculate.


b) The squared difference is always positive, which simplifies the loss calculation.
c) The squared difference is continuously differentiable, which is required for using
gradient-based optimization methods.
d) The squared difference penalizes large errors less severely than the absolute
difference.

Correct answer: c) The squared difference is continuously differentiable, which is


required for using gradient-based optimization methods.

8) True or False: The closed-form solution for the optimal weights in linear regression is
often referred to as the Normal Equation and involves the pseudoinverse of the feature
matrix.

a) True
b) False

Correct answer: a) True

9) In a multivariate linear regression problem, you have a design matrix 𝐗 (including the
bias column) and a target vector 𝐲. If 𝐗^(T)𝐗 is invertible, the Normal Equation solution
for the optimal weight vector 𝐰 (the projection method result) is given by:
a) 𝐰=(𝐗𝐗^(T))^(-1)𝐗𝐲
b) 𝐰=𝐗^(T)𝐲(𝐗𝐗^(T))^(-1)
c) 𝐰=(𝐗^(T)𝐗)^(-1)𝐗^(T)𝐲
d) 𝐰=𝐗(𝐗^(T)𝐗)^(-1)𝐲

Correct answer: c) 𝐰=(𝐗^(T)𝐗)^(-1)𝐗^(T)𝐲

10) Given the design matrix 𝐗=((1,2),(1,4)) and the target vector 𝐲=((6),(10)). Calculate
𝐗^(T)𝐲, which is a component of the Normal Equation calculation.

Correct answer: ((16),(52)) (or ((1,1),(2,4))((6),(10))=((6+10),(12+40))=((16),(52)))

11) The learning rate used in Gradient Descent to train a linear regression model is an
example of a model _________, while the optimal weight vector 𝐰 found by the Normal
Equation is an example of a model _________. (Separate your two answers with a
comma).

Correct answer: hyper-parameter, parameter

12) True or False: In linear regression, model parameters are learned directly from the
training data during the optimization process, whereas model hyper-parameters are
typically set prior to or outside of this optimization.

a) True
b) False

Correct answer: a) True

13) Which of the following is the best description of a model parameter in the context of
linear regression?

a) A setting that controls the complexity of the model, chosen before training begins.
b) The cost function used to measure the model's performance on the training data.
c) A value internal to the model (like a weight or bias) whose value is estimated from the
data.
d) The type of regularization (e.g., L1 or L2) applied to the objective function.

Correct answer: c) A value internal to the model (like a weight or bias) whose value is
estimated from the data.

Optimization methods including Gradient


Descent
{ L1, L2 loss function, Brute force method, Grid Search, Basic idea of Soft Computing
Techniques, equating gradient to be zero}

{convergence, learning rate, stopping criteria, batch size, vanishing gradient issue}
1) Explain the primary mathematical difference between the L₁ (Lasso) and L₂ (Ridge)
loss functions when used as regularization terms, specifically focusing on the effect each
has on model weights.

Correct answer: The L₁ loss (absolute value of weights) promotes sparsity by driving
some weights exactly to zero (feature selection), while the L₂ loss (squared magnitude of
weights) shrinks all weights towards zero but rarely makes them exactly zero.

2) True or False: The L₁ loss function is generally more robust to outliers in the training
data than the L₂ loss function.

a) True
b) False

Correct answer: a) True

3) A regression model uses the L₂ loss function for regularization with a regularization
parameter λ=0.5. If a specific model weight is w_(j)=2.0, calculate the L₂ regularization
penalty term contributed by this single weight.

Correct answer: 2.0

4) Which of the following is the primary limitation that makes the Brute Force
optimization method generally impractical for high-dimensional, continuous-variable
problems?

a) Susceptibility to local minima.


b) High risk of the vanishing gradient problem.
c) Computational complexity grows exponentially with the number of variables.
d) Inability to handle non-convex cost functions.

Correct answer: c) Computational complexity grows exponentially with the number of


variables.

5) In the context of hyperparameter tuning, Grid Search is considered appropriate when


the total number of hyperparameter combinations is relatively ______ and the
computational cost of evaluating each combination is ______ .

Correct answer: small, manageable (or low)

6) A research team is optimizing a non-linear objective function f(x) where x is a vector of


5 discrete parameters, each with 10 possible values. Calculate the total number of
function evaluations required for a complete Grid Search.

Correct answer: 100,000

7) Define the high-level philosophical difference between "Hard Computing" (e.g.,


deterministic algorithms like Brute Force) and "Soft Computing Techniques."
Correct answer: Hard computing demands precision, certainty, and truth, while Soft
Computing tolerates imprecision, uncertainty, partial truth, and approximations to
achieve tractability, robustness, and lower cost solutions for complex problems.

9) True or False: The primary goal of Soft Computing is to find the exact, verifiable global
optimum for a problem, unlike traditional optimization methods.

a) True
b) False

Correct answer: b) False

10) The fundamental update rule for Gradient Descent on a parameter θ with respect to
a cost function J(θ) is θ_(new)=θ_(old)-α⋅ (𝜕J(θ))/(𝜕θ). By setting the gradient (𝜕J(θ))/(𝜕θ)
to zero, you are attempting to find the ______ of the cost function.

Correct answer: minimum (or stationary point/local extremum)

11) A machine learning model is being trained using Gradient Descent. The cost function
is J(w)=w²-4w+5. Find the value of the weight w when the gradient is exactly zero.

Correct answer: 2

12) Which of the following best describes the role of the learning rate (α) in the Gradient
Descent algorithm?

a) It determines the batch size for the gradient calculation.


b) It sets the number of training epochs before the algorithm stops.
c) It controls the magnitude of the step taken in the direction opposite to the gradient.
d) It defines the total number of iterations before convergence.

Correct answer: c) It controls the magnitude of the step taken in the direction opposite to
the gradient.

13) True or False: Using a large batch size in Stochastic Gradient Descent (SGD)
generally leads to faster convergence per epoch but also requires more computation
time per update step.

a) True
b) False

Correct answer: b) False

14) The vanishing gradient issue in Gradient Descent is primarily caused by which of the
following?

a) Choosing a learning rate that is too high, leading to overshooting.


b) Using an insufficient number of training samples in the batch.
c) Repeated multiplication of small gradients through many layers, resulting in near-zero
updates.
d) The cost function being highly non-convex, leading to many local minima.
e) loss curve to be flat like ReLU

Correct answer: c) Repeated multiplication of small gradients through many layers,


resulting in near-zero updates. And e) loss curve to be flat like ReLU

15) A common stopping criterion for Gradient Descent is to halt the process when the
change in the cost function between consecutive iterations falls below a small
predetermined threshold ϵ. If the cost function changes from J_(k)=1.0003 to
J_(k+1)=1.0001, and ϵ=10^(-4), should the algorithm stop based on this criterion? Justify
your answer.

Correct answer: No. The absolute change is |1.0001-1.0003|=0.0002=2×10^(-4), which


is greater than the threshold ϵ=1×10^(-4).

Linear Algebra
Vector and Matrix Operations
Vectors and Basic Operations{ scaling, adding, linear combination, unit vector, vector
magnitude}

Matrices and their basic Operations {Matrix multiplication as linear combination of columns or
rows}

1) Given a vector 𝐯=((2),(-1)), the resulting vector after scaling by a factor of 3 is 𝐮=((6),(-3)).
Geometrically, what transformation does this scaling operation perform on 𝐯?
a) Rotation
b) Translation
c) Change in length only
d) Change in length and possibly direction (if the scalar is negative)

2) If 𝐮=((1),(3)) and 𝐯=((-2),(4)), calculate the sum 𝐮+𝐯.

3) The magnitude of a vector 𝐰=((3),(-4)) is 7.


a) True
b) False

4) Determine the unit vector in the direction of 𝐚=((5),(0),(-12)).

5) Express the vector 𝐰=((1),(8)) as a linear combination of 𝐮=((1),(2)) and 𝐯=((-1),(4)). The
scalar coefficients (c₁,c₂) are:
a) (c₁,c₂)=(3,2)
b) (c₁,c₂)=(2,3)
c) (c₁,c₂)=(1,7)
d) (c₁,c₂)=(5,-4)
6) Geometrically, the vector addition 𝐮+𝐯 can be visualized using the **parallelogram rule** or
the **triangle rule**.
a) True
b) False

7) Let 𝐀 be a 2×3 matrix and 𝐁 be a 3×4 matrix. The product 𝐁𝐀 is well-defined.


a) True
b) False

8) If matrix 𝐌 is multiplied by a vector 𝐱, the resulting vector 𝐌𝐱 is a linear combination of the


**rows** of 𝐌, with the entries of 𝐱 as the weights.
a) True
b) False

9) The geometric interpretation of the linear combination c₁𝐯₁+c₂𝐯₂ involves reaching a point by
scaling the initial vectors and then performing vector addition via the ______ rule.

10) Given 𝐀=((1,2),(0,3)) and 𝐁=((4,1),(-1,5)), compute the matrix product 𝐀𝐁.

11) A 3×2 matrix 𝐀 is multiplied by a 2×1 column vector 𝐱. The resulting 3×1 vector 𝐲=𝐀𝐱 is a
linear combination of which components of 𝐀?
a) The rows of 𝐀
b) The columns of 𝐀
c) The entries on the main diagonal of 𝐀
d) The transpose of $\mathbf{A}$'s rows

12) Explain why the magnitude of a vector 𝐯 is sometimes referred to as its "$L_2$ norm."

13) Consider the matrix multiplication 𝐂=𝐀𝐁, where 𝐀 is 2×3 and 𝐁 is 3×4. If you want to
compute the i-th column of 𝐂 without calculating the entire product, you would express it as a
linear combination of the columns of ______ weighted by the entries of the i-th column of
______.

14) Let 𝐌=((2,1),(3,0)) and 𝐱=((-1),(4)). The product 𝐌𝐱 is equivalent to the linear combination:
a) 4((2),(3))+(-1)((1),(0))
b) (-1)((2),(3))+4((1),(0))
c) 2((-1),(4))+1((3),(0))
d) 3((2),(1))+0((-1),(4))

15) A vector 𝐯 has a unit vector 𝐯=((1/ sqrt(2)),(-1/ sqrt(2))). If the magnitude of 𝐯 is 6, what is 𝐯?

---------------
Simultaneous system of linear equations
A simultaneous system of linear equations form a matrix equation {Row and Column
view of System of equations, Gauss Elimination, Gauss Jordon, pivot element, rank, REF,
RREF, Particular solution, homogeneous solution, total solution, Properties of Rank,
Elimination matrix, Elementary matrix that changes only one row, LU decomposition}

1) Consider the system of equations:

2x-3y=5
$4x + y = 1$
Write the corresponding matrix equation in the form A𝐱=𝐛, where A is the coefficient matrix, 𝐱 is
the vector of variables, and 𝐛 is the constant vector.

2) In the column view of the system A𝐱=𝐛, the vector 𝐛 must be a linear combination of the
______ of the matrix A.

3) The process of Gauss Elimination always results in a unique Row Echelon Form (REF) for a
given matrix.

a) True
b) False

4) Consider the matrix:

A=((1,2,3),(0,4,5),(0,0,6))
Identify the pivot elements after applying Gauss Elimination to the matrix A.

5) Find the rank of the matrix:

A=((1,3),(2,6))

6) The elementary matrix that transforms ((1,0),(3,1)) into the identity matrix I by subtracting 3
times the first row from the second row is an elimination matrix.

a) True
b) False

7) Determine the Row Echelon Form (REF) of the augmented matrix for the system:

x+y=2
$2x + 3y = 5$
a) ((1,1,|,2),(0,1,|,1))
b) ((1,1,|,2),(0,0,|,1))
c) ((1,0,|,1),(0,1,|,1))
d) ((1,2,|,2),(0,1,|,1))

8) What is the fundamental property that distinguishes a Reduced Row Echelon Form (RREF)
from a general Row Echelon Form (REF)?

9) Given a system A𝐱=𝐛, if the rank of the coefficient matrix A is less than the rank of the
augmented matrix [A|𝐛], what is the structure of the solution set?

10) Find the Reduced Row Echelon Form (RREF) of the matrix:

B=((1,2),(3,4))

11) For a system of m equations and n variables, A𝐱=𝐛, if Rank(A)=Rank([A|𝐛])=k and k<n, the
homogeneous solution 𝐱_(h) has n-k free variables.

a) True
b) False

12) The total solution to a non-homogeneous system A𝐱=𝐛 is written as the sum of the particular
solution 𝐱_(p) and the homogeneous solution 𝐱_(h), i.e., 𝐱_(total)=𝐱_(p)+𝐱_(h). The particular
solution 𝐱_(p) is found by setting all free variables to zero. Which of the following defines the
homogeneous solution 𝐱_(h)?

a) The set of all solutions to A𝐱=𝐛


b) The set of all solutions to A𝐱=𝟎
c) A single, non-zero vector that solves A𝐱=𝐛
d) The zero vector 𝟎 only

13) The matrix E=((1,0,0),(-2,1,0),(0,0,1)) is an elimination matrix that, when multiplied on the
left, performs which elementary row operation on a 3×n matrix A?

14) Find the L and U factors in the LU decomposition (A=LU) of the matrix:

A=((1,2),(3,8))
a) L=((1,0),(3,1)), U=((1,2),(0,2))
b) L=((1,2),(0,2)), U=((1,0),(3,1))
c) L=((1,0),(-3,1)), U=((1,2),(0,2))
d) L=((1,0),(3,1)), U=((1,2),(0,6))

15) What is the rank of an n×n matrix A if and only if A is invertible?

a) n
b) 0
c) n-1
d) Any integer k<n
-----------------------------------------

Vector spaces
{space, subspace, span, linearly dependent vectors, linearly independent vectors,
determinants, , basis, dimension of a vector space, vector space of matrices}, 4 vector spaces
related to a matrix A of size m x n {Row space of A, R(A), a subspace of R^n, Column space of
A, C(A), subspace of R^m, Null Space of A, N(A), a subspace of R^n, Null Space of A^T,
N(A^T), a subspace of R^m), Fundamental Theorem of Linear Algebra (dim(R(A))+dim(N(A)) =
n and dim(R(A^T))+dim(N(A^T)) = m}

1) Define a vector space V over a field F and list the two main properties that must hold for a
subset W of V to be considered a subspace.

2) The set of all polynomials of degree exactly 2, P₂={ax²+bx+c ∣ a≠0,b,c∈ℝ}, forms a vector
space under the usual polynomial addition and scalar multiplication.

a) True
b) False

3) Given the set of vectors S={((1),(0)),((0),(1))} in ℝ², describe the span span(S).

4) Determine if the set of vectors S={((1),(2)),((-2),(-4))} is linearly dependent or linearly


independent.

5) For a set of n vectors in ℝ^(n), what must the determinant of the matrix formed by these
vectors as columns be equal to for the vectors to be linearly independent?

a) Exactly 0
b) Any non-zero real number
c) Any non-negative real number
d) 1 or -1

6) Calculate the determinant of the matrix A=((3,1),(4,2)).

7) Find a basis for the subspace W=span{((1),(1),(0)),((0),(1),(1)),((1),(0),(-1))} of ℝ³.

8) What is the dimension of the vector space of all 2×3 matrices, denoted as M_(2×3)(ℝ)?

9) The Row Space R(A) of a matrix A is the subspace spanned by the ______ of A, and is a
subspace of ℝ^(n), where A is m×n.
10) For a 3×4 matrix A, the Null Space N(A) is a subspace of which Euclidean space?

a) ℝ³
b) ℝ⁴
c) ℝ^(3×4)
d) ℝ¹

11) Find the Null Space N(A) for the matrix A=((1,2),(2,4)).

12) The Column Space C(A) of a matrix A is defined as the range of the linear transformation
T(𝐱)=A𝐱.

a) True
b) False

13) If a 4×5 matrix A has a rank of 3, what is the dimension of the Null Space of $A$ dim (N(A)),
according to the Fundamental Theorem of Linear Algebra (Rank-Nullity Theorem)?

14) For a 5×7 matrix A, if dim (R(A))=4, what is dim (N(A^(T)))?

a) 7
b) 5
c) 3
d) 1

15) State the two main dimensional relationships given by the Fundamental Theorem of Linear
Algebra concerning a matrix A of size m×n.

16) Which of the following statements is a consequence of the Fundamental Theorem of Linear
Algebra's Orthogonality relationships?

a) The Row Space R(A) and the Column Space C(A) are orthogonal complements.
b) The Null Space N(A) is orthogonal to the Column Space C(A).
c) The Row Space R(A) is orthogonal to the Null Space N(A).
d) The Row Space R(A) and the Null Space N(A^(T)) are the only orthogonal subspaces of the
four fundamental subspaces.

17) Consider the matrix A=((1,0,-1),(0,1,1)). Find a basis for the Null Space of the Transpose
N(A^(T)).

18) Let V be the vector space of all continuous functions on the interval [0, 1]. Is the set of
vectors W={f(x)∈V∣f(0)=4} a subspace of V? Justify your answer briefly.
---------------------

Linear mapping & Orthonormal vectors & matrices


{Matrix as Linear map or Transformation Operator, Operators like- Scaling, Stretching, Shear,
Projection, Mirror, Rotation, Translation Operator that is Shifting from origin is not a linear
operator.}

{Orthogonal vectors, Unit vectors, Angle and length of a vector does not change on the
application of Orthonormal Linear Operators. These properties are also followed by rotation
operators.}

1) Define a linear map T:V→W between two vector spaces V and W. Specify the two axiomatic
properties that T must satisfy.

2) Let T:ℝ²→ℝ² be a function defined by T(x,y)=(x²,x+y). The property of homogeneity,


T(c𝐯)=cT(𝐯), fails for this function.

a) True
b) False

3) Which of the following matrices represents a shear transformation in ℝ² that adds twice the
second component to the first component?

a) ((1,0),(2,1))
b) ((1,2),(0,1))
c) ((2,0),(0,1))
d) ((1,1),(1,1))

4) Consider the vector 𝐯=((3),(-4)). Apply the scaling transformation matrix A=((5,0),(0,5)) to 𝐯.
The resulting vector is ______.

5) The operator T(𝐯)=𝐯+𝐚, where 𝐚 is a non-zero fixed vector (a translation), is a linear operator
because it preserves the operation of vector addition.

a) True
b) False

6) Let P:ℝ³→ℝ³ be the linear operator that projects a vector onto the xy-plane (i.e., z=0). What is
the 3×3 matrix representation of P?

7) Which condition must be satisfied by a set of vectors {𝐯₁,𝐯₂,…,𝐯_(k)} to be considered an


orthogonal set?
a) 𝐯_(i)⋅𝐯_(j)=0 for all i≠j.
b) 𝐯_(i)⋅𝐯_(j)=δ_(ij) (Kronecker delta).
c) 𝐯_(i)⋅𝐯_(i)=1 for all i.
d) The vectors span the entire space ℝ^(k).

8) Given the set of orthogonal vectors 𝐮₁=((1),(1)) and 𝐮₂=((-1),(1)), transform this set into an
orthonormal set {𝐞₁,𝐞₂}.

9) The Gram-Schmidt process is used to transform a basis for a vector space into an ______
basis.

10) Find the 2×2 matrix R that performs a counter-clockwise rotation of a vector in ℝ² by an
angle of θ=π/3 radians.

11) Let A be an n×n orthogonal matrix. For any vector 𝐯∈ℝ^(n), the length of the transformed
vector A𝐯 is equal to the length of 𝐯 (i.e., ‖A𝐯‖=‖𝐯‖).

a) True
b) False

12) Consider the orthogonal matrix Q=((0,1),(-1,0)), which represents a rotation by 270^(∘)
clockwise. Calculate the cosine of the angle between 𝐮=((1),(0)) and 𝐯=((0),(1)). Then calculate
the cosine of the angle between Q𝐮 and Q𝐯. The value of the cosine of the angle does not
change.

a) True
b) False

13) Which of the following is NOT a property of an n×n orthogonal matrix Q?

a) Q^(T)Q=I, where I is the identity matrix.


b) The columns of Q form an orthonormal basis for ℝ^(n).
c) The determinant of Q is always 1.
d) The inverse of Q is its transpose, Q^(-1)=Q^(T).

14) Let T:ℝ²→ℝ² be the linear map defined by the matrix M=((4,0),(0,1/2)). This transformation is
best categorized as a ______ operator.

15) Prove that an orthonormal linear operator T (represented by an orthogonal matrix Q)


preserves the dot product, i.e., T(𝐮)⋅T(𝐯)=𝐮⋅𝐯 for all vectors 𝐮,𝐯. This preservation of the dot
product is the reason why orthonormal operators preserve both angle and length.

----------------------
Eigenvalues & Eigenvectors
{eigenvectors if multiplied by any scalar remains the eigenvector for the operator. For a
symmetric matrix Eigenvectors are orthogonal if the eigenvalues are distinct and if two
eigenvectors share the same eigenvalue then each vector in the vector space spanned by these
vectors will be eigenvector sharing the same eigenvalue.}

1) Define, in precise mathematical terms, what an eigenvector of a linear operator T acting on a


vector space V is.

Correct answer: A non-zero vector 𝐯∈V is an eigenvector of a linear operator T if T(𝐯)=λ𝐯 for
some scalar λ.

2) If 𝐯 is an eigenvector of a matrix A with corresponding eigenvalue λ, and c is a non-zero


scalar, the vector c𝐯 is also an eigenvector of A with the same eigenvalue λ.

a) True
b) False

Correct answer: a) True

3) The eigenvalues of the matrix A=((2,0),(0,3)) are:

a) 0 and 1
b) 2 and 3
c) 1 and 2
d) 2 and 0

Correct answer: b) 2 and 3

4) Explain the process used to find the eigenvalues of a square matrix A.

Correct answer: The eigenvalues λ are found by solving the characteristic equation, which is
given by det(A-λI)=0, where I is the identity matrix.

5) Compute the eigenvalues of the matrix B=((1,1),(0,2)).

Correct answer: λ₁=1, λ₂=2

6) If 𝐯 is an eigenvector of matrix A with eigenvalue λ=4, what is the result of the matrix
multiplication A(5𝐯)?

a) 4𝐯
b) 5𝐯
c) 20𝐯
d) 9𝐯

Correct answer: c) 20𝐯

7) Determine the eigenvector corresponding to the eigenvalue λ=3 for the matrix A=((2,1),(1,2)).

Correct answer: 𝐯=k((1),(1)) (or any non-zero scalar multiple of ((1),(1)))

8) For a symmetric matrix, if two eigenvectors 𝐯₁ and 𝐯₂ correspond to the same eigenvalue λ,
they must be orthogonal.

a) True
b) False

Correct answer: b) False

9) Which of the following conditions guarantees that two eigenvectors 𝐯₁ and 𝐯₂ of a symmetric
matrix A are orthogonal?

a) The eigenvalues λ₁ and λ₂ are equal.


b) The matrix A is invertible.
c) The eigenvectors 𝐯₁ and 𝐯₂ correspond to distinct eigenvalues λ₁≠λ₂.
d) The determinant of A is non-zero.

Correct answer: c) The eigenvectors 𝐯₁ and 𝐯₂ correspond to distinct eigenvalues λ₁≠λ₂.

10) Prove that if 𝐯 is an eigenvector of A with eigenvalue λ, then for any scalar c≠0, c𝐯 is also an
eigenvector of A with eigenvalue λ.

Correct answer: A(c𝐯)=c(A𝐯)=c(λ𝐯)=λ(c𝐯). Since c𝐯≠𝟎 (because 𝐯≠𝟎 and c≠0), c𝐯 is an


eigenvector with eigenvalue λ.

11) A matrix M is symmetric. 𝐮 is an eigenvector with eigenvalue λ₁=5, and 𝐰 is an eigenvector


with eigenvalue λ₂=8. What is the value of the dot product 𝐮⋅𝐰?

Correct answer: 0

12) In the context of eigenvectors and eigenvalues, what is a degenerate eigenvalue?

Correct answer: A degenerate eigenvalue is an eigenvalue whose algebraic multiplicity is


greater than 1 (or, equivalently, the dimension of its corresponding eigenspace is greater than
1).
13) If a symmetric matrix S has an eigenvalue λ with algebraic multiplicity 2, the eigenspace
E_(λ) corresponding to λ is spanned by exactly one linearly independent eigenvector.

a) True
b) False

Correct answer: b) False

14) Consider a 3×3 symmetric matrix A with a degenerate eigenvalue λ. If 𝐯₁ and 𝐯₂ are two
linearly independent eigenvectors for λ, what can be said about any vector 𝐮 in the span of
{𝐯₁,𝐯₂}?

a) 𝐮 is the zero vector.


b) 𝐮 is an eigenvector of A corresponding to a different eigenvalue.
c) 𝐮 is also an eigenvector of A corresponding to the eigenvalue λ.
d) 𝐮 is linearly dependent on 𝐯₁ but not 𝐯₂.

Correct answer: c) 𝐮 is also an eigenvector of A corresponding to the eigenvalue λ.

15) If a 2×2 symmetric matrix has a repeated eigenvalue λ=4, what is the minimum dimension of
the eigenspace E₄?

Correct answer: 2

-----------------

Matrix Diagonalization, Principal Component Analysis, Singular


value Decomposition

1) If A is an n×n matrix, its singular values are the square roots of the eigenvalues of the matrix
A^(T)A. This is a necessary step when computing the ______ decomposition.

Correct answer: Singular Value

2) The dimensionality reduction method that finds a lower-dimensional subspace which


maximizes the variance of the projected data is called ______.

Correct answer: Principal Component Analysis (PCA)


3) The Spectral Theorem guarantees that any symmetric n×n matrix A can be diagonalized by
an orthogonal matrix Q, such that A=QDQ^(T).

a) True
b) False

Correct answer: a) True

4) Explain the fundamental distinction in applicability between the Spectral Theorem for matrix
diagonalization and the Singular Value Decomposition (SVD).

Correct answer: Spectral Theorem applies only to symmetric matrices (or more generally,
normal matrices) and diagonalizes A. SVD applies to any matrix, factoring A into UΣV^(T).

5) Given a data matrix X, the eigenvectors of the covariance matrix C= 1/(m-1) X^(T)X
correspond to the principal components.

a) True
b) False

Correct answer: a) True

6) Which of the following conditions is necessary and sufficient for an n×n matrix A to be
diagonalizable via A=PDP^(-1)?

a) A has n distinct eigenvalues.


b) A is symmetric.
c) A has n linearly independent eigenvectors.
d) The characteristic polynomial of A has n real roots.

Correct answer: c) A has n linearly independent eigenvectors.

7) When performing Singular Value Decomposition on a matrix A=UΣV^(T), what information is


encoded by the singular values, which are the diagonal entries of the matrix Σ?

Correct answer: The "magnitudes" or "strengths" of the principal components/directions,


representing the variance or energy along the singular vector directions.
8) Consider a 1000×1000 dense matrix A. Which of the following decomposition methods would
generally be the most computationally efficient for finding a low-rank approximation?

a) Standard Eigenvalue Decomposition (Diagonalization)


b) QR Decomposition
c) Singular Value Decomposition (SVD)
d) LU Decomposition

Correct answer: c) Singular Value Decomposition (SVD)

9) Describe the mathematical role of the principal components in the context of dimensionality
reduction via PCA.

Correct answer: Principal components are the eigenvectors of the covariance matrix that define
a new, orthogonal basis for the data, ordered by the amount of variance in the data they
capture.

10) For a non-square matrix A, which of the following decompositions is always possible?

a) Diagonalization A=PDP^(-1)
b) A=QR
c) Singular Value Decomposition A=UΣV^(T)
d) Eigendecomposition A=VΛV^(-1)

Correct answer: c) Singular Value Decomposition A=UΣV^(T)

11) A dynamical system is described by the first-order linear system of ODEs d𝐱/(dt) =A𝐱, where
A is a diagonalizable n×n matrix. Outline how matrix diagonalization is used to find the general
solution 𝐱(t) in terms of eigenvalues and eigenvectors.

Correct answer: Diagonalize A=PDP^(-1). Define 𝐲=P^(-1)𝐱. The ODE transforms to d𝐲/(dt) =D𝐲,
which decouples into n independent scalar equations dy_(i)/(dt) =λ_(i)y_(i). The solution is
𝐲(t)=[c₁e^(λ₁t),…,c_(n)e^(λ_(n)t)]^(T), and the final solution is 𝐱(t)=P𝐲(t).

12) In the context of low-rank approximation using SVD, if a matrix A has singular values
σ₁≥σ₂≥…≥σ_(r)>0, the best rank-$k$ approximation (k<r) is obtained by setting all singular
values σ_(k+1) through σ_(r) to zero and reconstructing the matrix.

a) True
b) False
Correct answer: a) True

13) What is the fundamental connection between the Principal Component Analysis (PCA)
algorithm and the eigendecomposition of the data's covariance matrix?

Correct answer: The principal components (directions of maximal variance) are the eigenvectors
of the covariance matrix, and the variance along those components is given by the
corresponding eigenvalues.

14) A key limitation of applying matrix diagonalization to high-dimensional data is that the matrix
must be square and must have a full set of linearly independent eigenvectors. For Singular
Value Decomposition (SVD), a key practical limitation for large datasets is the high
computational cost, specifically the O(n³) complexity for a dense n×n matrix.

a) True
b) False

Correct answer: a) True

15) A researcher applies PCA to a dataset and finds that the first principal component accounts
for 85% of the total variance. If they use only the first principal component for dimensionality
reduction, what is the most significant practical limitation they must be aware of regarding their
ability to reconstruct the original data?

Correct answer: By retaining only 85% of the variance, the remaining 15% (which may contain
important, albeit smaller, features or noise) is lost, leading to an unavoidable information loss
and distortion in the reconstructed data.

Common questions

Powered by AI

A positive skew in a dataset, such as house prices, indicates that the mean is greater than the median, which is typically greater than the mode. The tail on the right side pulls the mean higher than central measures, suggesting consideration of the median for reporting central tendencies as it is less affected by outliers .

The Interquartile Range (IQR) is more robust than the range because it is not affected by the extreme values, like high executive salaries, as it focuses on the middle 50% of data. This makes it particularly useful for understanding the typical variability in a dataset where outliers could skew the range significantly, providing a clearer picture of the central distribution of salaries .

A heatmap is ideal for visualizing the relationship between temperature measurements and types of experiments because it effectively conveys the frequency and intensity of different temperature values across experiment types. The color intensity in the heatmap can represent the range of temperatures associated with each experiment type, allowing researchers to easily identify patterns or anomalies in the data distribution .

Using only the first principal component for dimensionality reduction, which captures 85% of the variance, implies that 15% of the variance is unaccounted for. This leads to an information loss, potentially overlooking smaller yet significant features or noise in the data, thus distorting the original data's full complexity .

Standard deviation is generally preferred over variance for presentation because it is expressed in the same units as the original data, making it more intuitive for understanding and comparing the variability directly. This aids in communication, particularly to non-statistical audiences, making the information more accessible and meaningful .

The Bayes classifier is referred to as a maximum a posteriori classifier because it relies on Bayes' theorem to estimate the posterior probabilities of classes. The decision rule involves choosing the class with the highest posterior probability, effectively maximizing a posteriori probability and hence the name .

In Singular Value Decomposition, singular values represent the 'strengths' or 'magnitudes' of the principal component directions of a matrix, where larger singular values indicate directions of higher variance. These singular values relate to principal components as they define the scaling along each principal axis, thus capturing the essence of data variability .

A geographical heatmap is optimal when visualizing spatial data, such as COVID-19 case rates per capita across counties. It displays variations in data intensity based on geographical locations. The insights come from recognizing patterns of concentration or spread over the area, enabling researchers to interpret high-intensity zones as areas with more cases and thus prioritize interventions .

Principal Component Analysis (PCA) employs the eigendecomposition of a covariance matrix to identify eigenvectors, which form the principal components. This helps in transforming the dataset to a new orthogonal basis where components capture maximum variance, enabling dimensionality reduction by prioritizing components with the largest eigenvalues, effectively summarizing the data's variance with fewer dimensions .

The L₁ regularization (Lasso) applies absolute value to weights, encouraging sparsity by potentially driving some weights to zero, thus assisting in feature selection. Conversely, L₂ regularization (Ridge) uses squared magnitudes of weights, which shrinks all weights uniformly towards zero but does not typically zero them, helping to reduce overfitting while retaining all features .

You might also like