0% found this document useful (0 votes)
3 views18 pages

Sec Assignment

The document discusses various mathematical concepts, including the vertical line test for functions, types of functions (linear, quadratic, cubic, etc.), scatter diagrams, correlation coefficients, hypothesis testing, and bivariate regression analysis. It explains definitions, applications, and examples for each concept, emphasizing the importance of understanding relationships between variables and the implications of statistical tests. Additionally, it addresses potential errors in hypothesis testing and the significance of regression models in predicting outcomes.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views18 pages

Sec Assignment

The document discusses various mathematical concepts, including the vertical line test for functions, types of functions (linear, quadratic, cubic, etc.), scatter diagrams, correlation coefficients, hypothesis testing, and bivariate regression analysis. It explains definitions, applications, and examples for each concept, emphasizing the importance of understanding relationships between variables and the implications of statistical tests. Additionally, it addresses potential errors in hypothesis testing and the significance of regression models in predicting outcomes.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

IT Skills and Data Analysis – II

(Assignment)

PRANAV RAMCHANDANI
23/33561
[Link] Hons
Q1. Answer the following:
(a) What is a vertical line test? Explain geometrically and
algebraically

Vertical Line Test: The vertical line test is a method used to determine
whether a given graph represents a function. It states that if any vertical
line drawn on the coordinate plane intersects the graph at more than one
point, then the graph does not represent a function. Conversely, if every
vertical line intersects the graph at exactly one point, the graph
represents a function.

Geometrical Explanation:

1. Plot the graph of the given equation on the coordinate plane.


2. Draw a vertical line (parallel to the y-axis), for example x=a.
3. Observe the points of intersection:
o If the line intersects the graph at only one point, the relation
is a function.
o If the line intersects at more than one point, the relation is not
a function.

Algebraic Explanation:

 A vertical line is represented as x=a.


 Substitute x=a into the given equation.
 If the substitution yields only one value of y, the relation is a
function.
 If it yields more than one value of y, the relation is not a function.

Conclusion: The vertical line test provides a straightforward way to


verify whether a relation satisfies the definition of a function.

(b) Application to the Equation y=2x+5

The given equation is:


y=2x+5

This is a linear equation, and its graph is a straight line. Applying the
vertical line test:

 Any vertical line drawn at x=a intersects the straight line at exactly
one point.
 Therefore, for each input x, there is only one output y.

Result: The equation passes the vertical line test and hence represents
a function.

(c) Importance of the Vertical Line Test

The vertical line test is important because:

 It provides a quick visual method to determine whether a curve


represents a function.
 It ensures that each input (x-value) corresponds to exactly one
output (y-value).
 It prevents misinterpretation of relations that do not meet the
definition of a function.

(d) Misunderstanding the Vertical Line Test: Real-World Example

Misunderstanding the vertical line test can lead to errors in interpreting


data. If a relation is incorrectly assumed to be a function, conclusions
drawn from it may be invalid.

Example: Consider the relationship between a person’s age and height.

 At age 10, suppose two different height values (140 cm and 150
cm) are recorded.
 Drawing a vertical line at x=10 intersects the graph at two points.
 This indicates multiple outputs (heights) for a single input (age),
which violates the definition of a function.

Conclusion: Incorrectly applying or misunderstanding the vertical line


test can result in flawed data analysis. It is essential to ensure that each
input corresponds to only one output when identifying functions.

Q2. Answer the following:


(a) Define linear, quadratic, cubic, reciprocal, exponential, and log
functions. Explain standard forms and general shapes. Give a
suitable example.

(a) Types of Functions: Definitions, Standard Forms, Shapes, and


Examples

1. Linear Function
o Definition: A function where the highest power of x is 1.
o Standard Form: y=mx+c
o Shape: Straight line (increasing if m>0, decreasing if m<0)
o Example: y=2x+3
2. Quadratic Function
o Definition: A function where the highest power of x is 2.
o Standard Form: y=ax2+bx+c
o Shape: Parabola (U-shaped if a>0, inverted U-shaped if a<0)
o Example: y=x2−4
3. Cubic Function
o Definition: A function where the highest power of x is 3.
o Standard Form: y=ax3+bx2+cx+d
o Shape: S-shaped curve with possible inflection points.
o Example: y=x3
4. Reciprocal Function
o Definition: A function where x appears in the denominator.
o Standard Form: y=1x
o Shape: Hyperbola with two separate branches.
o Example: y=1x
5. Exponential Function
o Definition: A function where the variable appears in the
exponent.
o Standard Form: y=ax
o Shape: Rapidly increasing (growth) or decreasing (decay)
curve.
o Example: y=2x
6. Logarithmic Function
o Definition: The inverse of an exponential function.
o Standard Form: y=log⁡(x)
o Shape: Slowly increasing curve, undefined for x≤0.
o Example: y=log⁡(x)

(b) Effect of Slope on Linear Functions

 Definition: The slope (m) measures the steepness of a line.


 Formula:
 Interpretation:
o Positive slope → line rises from left to right.
o Negative slope → line falls from left to right.
o Zero slope → horizontal line.
o Undefined slope → vertical line.

(c) Effect of the Leading Coefficient

The leading coefficient is the number multiplied by the highest power of


x. It influences the graph’s direction and steepness:

 Positive coefficient → graph opens upwards (quadratic) or rises


overall.
 Negative coefficient → graph opens downwards or reflects across
the x-axis.
 Large coefficient → graph becomes steeper/narrower.
 Small coefficient → graph becomes wider/flatter.

(d) Difference Between Quadratic and Cubic Functions

 Quadratic Function: Has only one turning point (vertex). The graph
changes direction once.
 Cubic Function: Can have up to two turning points. The graph is S-
shaped and may change direction more than once.

(e) Exponential Growth and Decay

 Exponential Growth: Value increases continuously at a constant


rate.
o Example: Population growth, compound interest.
 Exponential Decay: Value decreases continuously at a constant
rate.
o Example: Radioactive decay, cooling of a hot object.

(f) Slope as Rate of Change in Real Life

Slope represents how one quantity changes relative to another.

 Distance-Time Graph: Slope = speed.


 Cost vs Quantity Graph: Slope = rate at which cost increases with
quantity.
 Interpretation:
o Steep slope → rapid change.
o Gentle slope → slow change.

Q3. Answer the following:


(a) What is a scatter diagram? When do we use scatter diagrams, and
how do we interpret them? Give examples with graphs.

A scatter diagram (or scatter plot) is a graph used to show the relationship
between two variables.
In this diagram, data points are plotted on a graph using the x-axis and
the y-axis, where each point represents a pair of values.

Scatter diagrams are used when we want to:


● Study the relationship between two variables
● Check whether there is any correlation
● Identify patterns or trends in data
The pattern of points helps us understand the type of relationship:
● If points move upward (left to right) → Positive correlation
● If points move downward, → Negative correlation
● If points are scattered randomly, → No correlation
Examples:
● Height and weight → Positive correlation
● Price and demand → Negative correlation
(b) Draw a Scatter Diagram for the following data and state the type
of correlation between the given two variables X and Y.
Height (m) 3.14 3.87 2.84 4.34
Diameter (cm) 4.20 5.55 3.33 6.91

After plotting the points, we observe that as height increases, diameter


also increases. The scatter diagram shows a positive correlation between

height and diameter. (c) What are the drawbacks of a scatter diagram? (At
least 5) The drawbacks of scatter diagrams are:
1. It does not show the exact numerical value of correlation.
2. It can be difficult to interpret if there are many data points.
3. It does not show a cause-and-effect relationship.
4. It may yield misleading results if the data is inaccurate.
5. It is not suitable for large datasets.
QUES4

(a) Correlation Coefficient

Definition: The correlation coefficient (r) measures the strength and


direction of the relationship between two variables.

 Positive correlation: Both variables increase together.


 Negative correlation: One variable increases while the other
decreases.
 Near zero: No significant relationship exists.

Visualisation: Linear correlations are visualised using a scatter


diagram:

 Upward trend of points → Positive correlation.


 Downward trend of points → Negative correlation.
 Random scatter → No correlation.

Range of r:
−1≤r≤+1

 r=+1: Perfect positive correlation.


 r=−1: Perfect negative correlation.
 r=0: No correlation.

Real-life Examples:

1. Height and weight of a person (positive correlation).


2. Price and demand of a product (negative correlation).

(b) Correlation vs. Causation

Correlation and causation are not the same:

 Correlation: Indicates a relationship between two variables.


 Causation: Means one variable directly influences the other.

Example: Ice cream sales and temperature are correlated (both rise
together), but ice cream sales do not cause temperature to increase.

(c) Pearson’s r and Spearman’s rho

 Pearson’s Correlation Coefficient (r):


o Measures the linear relationship between two numerical
variables.
o Suitable for continuous, normally distributed data.
o Range: −1 to +1.
 Spearman’s Rank Correlation Coefficient (ρ):
o Based on ranks rather than raw values.
o Useful for ordinal data or when data is not normally
distributed.
o Range: −1 to +1.

(d) Coefficient of Determination

The coefficient of determination (r2) shows how much of the variation in


one variable is explained by the other.

Example: If r=−0.60:
r2=(−0.60)2=0.36

This means 36% of the variation in one variable is explained by the


other, while the remaining 64% is due to other factors.

Interpretation: Although the correlation is negative, the strength of the


relationship is moderate, and a significant portion of variation is
accounted for.

Q5. The following data shows the number of hours studied and
marks obtained by students:
Hours Studied (X) 2 4 6 8 10
Marks Obtained (Y) 35 50 65 80 95

(i) Calculate the correlation coefficient (r). Interpret the result.


(ii) Calculate the coefficient of determination (r²). Explain what the
value of r² indicates in this context.

Formula: r² = (1)² = 1
The value of r² = 1means:
● 100% of the variation in marks is explained by hours studied
● There is a perfect relationship between the two variables

Q6. Answer the following:

(a) Define a statistical hypothesis

A statistical hypothesis is a statement or assumption made about a


population parameter (such as mean or proportion). It is tested using
sample data to determine whether the assumption holds true. In
essence, it allows us to make decisions and draw conclusions based on
evidence from data analysis.

(b) Z, t and P: when to use in hypothesis testing?

 Z-test: Used when the sample size is large (n>30) and the
population standard deviation is known.
 t-test: Applied when the sample size is small (n<30) and the
population standard deviation is unknown.
 p-value: Helps decide whether to reject or accept the null
hypothesis. A small p-value indicates strong evidence against the
null hypothesis, suggesting that the observed result is unlikely
under H0.

(c) Differentiate between the null hypothesis (H₀ ) and the


alternative hypothesis (H₁ )
Null Hypothesis (H₀ ) Alternative Hypothesis (H₁ )
States there is no effect or States there is an effect or
difference difference
Assumed true at the start of
Represents what we aim to prove
testing
Example: “No change in marks” Example: “Marks have improved”

(d) What is meant by a one-tailed and a two-tailed test?

 One-tailed test: Checks for an effect in only one direction (either


increase or decrease).
o Example: Testing if marks increased.
 Two-tailed test: Checks for an effect in both directions (increase
or decrease).
o Example: Testing if marks changed, regardless of whether
they increased or decreased.

(e) Define Type I and Type II errors. What is the level of


significance?

 Type I Error: Rejecting a true null hypothesis (false positive).


 Type II Error: Accepting a false null hypothesis (false negative).
 Level of significance (α): The probability of making a Type I
error. Commonly set at 0.05 (5%), meaning there is a 5% risk of
rejecting a true null hypothesis.

Q7. Jeffrey, as an eight-year-old, established a mean time of 16.43


seconds for swimming the 25- yard freestyle, with a standard
deviation of 0.8 seconds. His dad, Frank, thought that Jeffrey could
swim the 25-yard freestyle faster using goggles. Frank bought
Jeffrey a new pair of expensive goggles and timed Jeffrey for 15 25-
yard freestyle swims. For the 15 swims, Jeffrey's mean time was 16
seconds. Frank thought that the goggles helped Jeffrey swim faster
than the 16.43 seconds. Conduct a hypothesis test using a preset α
= 0.05. Assume that the swim times for the 25- yard freestyle are
normal. What are the Type I and Type II errors for this problem?

Step 1: Given Data

 Population mean (μ) = 16.43 seconds


 Sample mean (xˉ) = 16 seconds
 Standard deviation (σ) = 0.8 seconds
 Sample size (n) = 15
 Level of significance (α) = 0.05

Step 2: Hypotheses

Since we are testing if the time has become faster (less than the
population mean):

 Null Hypothesis (H0): μ=16.43 (no improvement).


 Alternative Hypothesis (H1): μ<16.43 (time has decreased →
faster).

This is a one-tailed (left-tailed) test.

Step 3: Test Statistic Formula

For a Z-test:
Z=xˉ−μσ/n

Step 4: Substitution of Values


Z=16−16.430.8/15≈−2.09

Step 5: Critical Value

At α=0.05 (one-tailed test):

 Critical Z-value = -1.645

Step 6: Decision

 Calculated Z = -2.09
 Critical Z = -1.645
Since −2.09<−1.645, we reject H0. This means there is enough
evidence to conclude that the time has decreased (performance is
faster).

Type I and Type II Errors

 Type I Error: Rejecting H0 when it is actually true.


o Example: Concluding that goggles improve performance
when in reality they do not.
 Type II Error: Failing to reject H0 when it is false.
o Example: Concluding that goggles do not help, when in fact
they do improve performance.

Q8. Explain bivariate regression analysis. Write the general form of


a simple linear regression model.

What is the purpose of the error term in a regression model?


Bivariate regression analysis is a method used to study the relationship
between two variables, where one variable depends on the other.
● One variable is called the independent variable (X)
● The other is called the dependent variable (Y)
It helps us to predict the value of one variable based on the other.
Simple Example:
If we study how hours studied (X) affect marks (Y), we can use regression
to predict marks.
General Form of Simple Linear Regression
Y=a+bX+e
Where:
● Y = dependent variable
● X = independent variable
● a = intercept (value of Y when X = 0)
● b = slope (rate of change)
● e = error term
The error term (e) represents the difference between actual and predicted
values.
It is included because:
● Not all factors affecting Y are included in the model
● There may be random variations or measurement errors
So, the error term accounts for unexplained variation in the data.
Q9. Find the line of best fit for the following data of heights and
weights of students of a school using the Least Squares method,
and interpret the slope.
Height (cm) 160 162 164 166 168
Weight (kg) 52 55 57 60 61

Calculate sums: ΣX=820 , ΣY=285 , ΣX^2=134520 , ΣXY=46786


Slope: b = {nΣXY−(ΣX)(ΣY) }/ nΣX^2 -(ΣX)^2 , where 𝑛 = 5

b=1.15
Intercept a = ΣY−bΣX/n = −131.6

Final Equation : Y = −131.6+1.15X

For every 1 cm increase in height, weight increases by approximately 1.15


kg.

Q10. Explain multiple linear regression and goodness of fit. (In


detail) State the importance of adjusted R² as compared to R².

(a) Multiple Linear Regression

Definition: Multiple linear regression is a statistical technique used to


examine the relationship between one dependent variable and two or
more independent variables. It helps us understand how several factors
together influence a single outcome.

General Form:
Y=a+b1X1+b2X2+⋯+bnXn+e

Where:

 Y = dependent variable (outcome)


 X1,X2,…,Xn = independent variables (predictors)
 a = intercept (value of Y when all X are zero)
 b1,b2,…,bn = regression coefficients (showing the effect of each
independent variable)
 e = error term (variation not explained by the model)

Interpretation: Each coefficient tells us how much the dependent


variable changes when the corresponding independent variable
changes, while keeping other variables constant.

(b) Goodness of Fit

Definition: Goodness of fit measures how well the regression model


explains the observed data. It shows the accuracy of the model in
capturing the relationship between variables.

Common Measure – Coefficient of Determination (R2):

 R2 ranges from 0 to 1.
 A higher value indicates a better fit.
 Example: R2=0.8 means 80% of the variation in the dependent
variable is explained by the independent variables in the model.

(c) Importance of Adjusted R2 Compared to R2

Why not rely only on R2?

 R2 always increases when more variables are added, even if


those variables do not contribute meaningfully.
 This can give a false impression of a better model.

Adjusted R2:

 Adjusted R2 modifies R2 by considering the number of


independent variables.
 It increases only if the newly added variable actually improves the
model.
 If the variable is irrelevant, adjusted R2 may decrease.

Importance:

 Provides a more accurate measure of model performance.


 Helps avoid overfitting (adding too many variables just to increase
R2).
 Useful for comparing models with different numbers of predictors.
Conclusion

Multiple linear regression is a powerful tool for analyzing the combined


effect of multiple variables on a single outcome. While R2 shows how
well the model explains variation, adjusted R2 is more reliable because it
accounts for the number of predictors, ensuring that only meaningful
variables improve the model’s fit.

You might also like