100% found this document useful (1 vote)
43 views3 pages

Statistical Analysis and Assessments

Uploaded by

shanmuselvam
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
100% found this document useful (1 vote)
43 views3 pages

Statistical Analysis and Assessments

Uploaded by

shanmuselvam
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Life Long Learning Assessment

21CSC529T - INFERENTIAL STATISTICS

Answer any 10 questions

1. Calculate mean, median, mode, range, quartile deviation, standard deviation, variance
and the coefficient of quartile deviation, standard deviation, coefficient of variation, Karl
Pearson’s coefficient of skewness, Bowley’s coefficient of skewness, measures of
skewness using γ1 and measures of kurtosis using γ2 for the following data:

Sales 15 18 25 27 30 35

Expenditu 50 65 82 95 110 120


re

Also find the Pearson’s and Spearman’s Correlation Co-efficient

2. A random variable X has the following distribution

X -2 -1 0 1 2 3

P[X=x] 0.1 k 0.2 0.2k 0.3 0.3k

Find (i) the value of k, (ii) the Distribution Function (CDF) (iii) P (0 < X < 3/X <
2) and
(iv) the smallest value of α for which P (X ≤ α) >½

3. Box 1 contains 1000 bulbs of which 10% are defective. Box 2 contains 2000 bulbs of
which 5% are defective. Two bulbs are drawn (without replacement) from a randomly
selected box. (i) Find the probability that both balls are defective and (ii) assuming that
both are defective, find the probability that they came from box 1.

4. The chance of a doctor D will diagnose a disease Z correctly is 60%. The chance that
the patient will survive by his treatment after a wrong diagnosis is 40% and the chance
of surviving for a correct diagnosis is 70%. Find the probability of the patient surviving. A
patient of doctor D, who has disease Z, survives. What is the chance that his disease
was diagnosed wrongly?

5. Out of 500 companies with 4 secretaries each, how many companies would be expected
to have
(a) 2 Men and 2 Women executives
(b) More than one men executive
(c) At least 2 women executives
(d) Executives of both genders
(Assume both men and women have equal probabilities)

6. The number of defective pins in a box of 100 follows the binomial probability law. The
question is to find the probability that more than 4 pins are defective in a box. What is
the probability that a box will fail to meet the guaranteed quality.

7. The life time of a certain brand of tube light may be considered as a random variable
with mean 1200 hours and S.D. 250 hours. Find the probability using CLT that the
average life time of 60 lights exceeds 1250 hours.

8. A salesman in a departmental store claims that at most 60% of the shoppers entering
the store leaves without making a purchase. A random sample of 50 shoppers showed
35 of them without making a purchase. Are these sample results consistent with the
salesman’s claim?

9. The marks of 100 students in an exam are found to be normally distributed with mean 70
and S.D. 5. Estimate the number of students whose marks will be
(a) between 60 and 75.
(b) more than 75 marks.
(c) less than 65 marks.

10. Determine whether the average weight of a sample of 20 mangoes is significantly


different from the population’s average weight of 70 grams. The sample mean weight is
70.55 grams, and the sample standard deviation is 2.82 grams.

11. Determine if there is a significant difference in the average scores between the two
teams. The following data is given:
Team A: Score: 65, 68, 70, 63, 67
Team B: Score: 62, 66, 69, 64, 68

12. You need to assess the effectiveness of a new teaching scheme by comparing the test
scores of the same group of students before and after the implementation of the
scheme. The following data is given:
Before (new teaching scheme scores): 76, 88, 65, 56, 76
After (new teaching scheme scores): 85, 95, 75, 60, 81
Determine if there is a significant difference in the average test scores before and after
the implementation of the scheme.

13. Customers are surveyed by a company to determine whether their age group (under 20,
20-40, over 40) and their preferred product category (food, apparel, or electronics) are
related. The information gathered is:
• Under 20: Electronic - 50, Clothing - 30, Food - 20
• 20-40: Electronic - 60, Clothing - 70, Food - 50
• Over 40: Electronic - 30, Clothing - 40, Food - 80

14. A researcher wants to determine whether three different teaching methods (A, B, and C)
have different effects on students' test scores. She randomly assigns students to one of
the three teaching methods and records their test scores. The data is as follows:
● Method A: 85,90,78,92,88
● Method B: 72,75,80,68,74
● Method C: 88,85,84,90,86

15. In an experiment to see whether the amount of coverage of light-blue interior latex paint
depends either on the brand of paint or on the brand of roller used, one gallon of each
of four brands of paint was applied using each of three brands of roller, resulting in the
following data (number of square feet covered). (a) Test whether the Roller Brands differ
with respect to treatment. (b) Test whether the Paint Brands differ with respect to
treatment. (c) Test whether the roller brands are the same for the different paint brands.

Soil Type
Coating
A B C
I 454 446 451
II 446 444 447
III 439 442 444
IV 444 437 443

Common questions

Powered by AI

Interpret results of statistical hypothesis testing by comparing the p-value to a predetermined significance level (α, often 0.05). If the p-value is less than α, reject the null hypothesis, indicating sufficient evidence for the alternative hypothesis about the population parameter. If the p-value exceeds α, fail to reject the null, suggesting insufficient evidence. Also, consider confidence intervals to support decisions, offering a range within which the true parameter likely lies. This methodology helps decide if sample observations support claims about population parameters, guiding actions such as product releases or treatment plans .

Assess claims about population proportions using hypothesis testing by comparing sample data to the claimed proportion. Use a z-test for proportions, where the null hypothesis states the population proportion equals the claim. Calculate the test statistic: z = (p̂ - p0) / √(p0(1-p0)/n), where p̂ is the sample proportion, p0 is the claimed proportion, and n is the sample size. Compare the test statistic to critical values from the standard normal distribution to decide on the hypothesis. This approach assesses whether sample results are consistent with population claims, as seen in checking the consistency of purchase likelihood claims in store shoppers .

The Central Limit Theorem (CLT) assists in estimating probabilities for sample means by allowing one to assume that the distribution of sample means approximates a normal distribution, regardless of the population distribution, given a large enough sample size. For instance, when determining the probability that the average lifetime of 60 tube lights exceeds 1250 hours, use CLT. Compute the sample mean and standard deviation, then convert to a standard normal distribution using Z-scores to find probabilities. This application provides practical estimation of probabilities for sample means, crucial for planning and quality assurance in production .

Correlation analysis quantifies the strength and direction of the linear relationship between variables. For example, Pearson's and Spearman's correlation coefficients measure how one variable, like sales, relates to another, like expenditure. Regression analysis goes further by modeling the relationship, predicting one variable based on another. It provides coefficients representing the change in dependent variable per unit change of the independent. These analyses help identify significant predictors, assess trends, and make informed decisions based on the strength and nature of variable relationships .

Statistical methods to determine educational intervention effects include paired t-tests and ANOVA. To test if a new teaching scheme significantly affects student performance, use a paired t-test. Calculate the mean difference between pre- and post-intervention scores, standard deviation of differences, and test statistic: t = (mean difference) / (SD/√n). If testing multiple interventions, ANOVA is used by comparing variances between groups to within groups to calculate an F-statistic. Significance is quantified by p-values, with values below a threshold (e.g., 0.05) indicating significant differences. These tests reveal whether interventions lead to meaningful performance changes .

To calculate survival probabilities in medical diagnosis, use conditional probabilities. First, determine the probability of correct diagnosis (P(D)) and wrong diagnosis (P(W)). Given survival probabilities with correct (P(S|D)) and wrong (P(S|W)) diagnoses, apply the law of total probability: P(S) = P(S|D)P(D) + P(S|W)P(W). This calculates the overall survival probability. To assess the likelihood of a wrong diagnosis given survival, use Bayes' theorem: P(W|S) = [P(S|W)P(W)] / P(S). These probabilities interpret survival chances and error likelihood, influencing treatment evaluations .

When using ANOVA to test differences among group means, consider assumptions such as normality, homogeneity of variance, and independence. Data must not deviate significantly from these assumptions for valid results. ANOVA reveals whether there are statistically significant differences in means across groups, using the F-statistic to compare variance among groups to variance within groups. A significant F-test suggests at least one group mean differs. This analysis is essential for comparing multiple groups without inflating Type I error, suitable for experiments like determining the effect of different teaching methods on test scores .

The key statistical measures calculated for evaluating sales and expenditure data include the mean, median, mode, range, quartile deviation, standard deviation, variance, coefficient of variation, and skewness coefficients (Karl Pearson's and Bowley's). These measures provide insights into central tendency, variability, and distribution shape. The mean gives the average, the median represents the central value, and the mode indicates the most frequent observation. Range and standard deviation highlight variability. Skewness coefficients measure asymmetry: Karl Pearson's focuses on mean and mode differences, while Bowley's uses quartiles. Understanding these aspects helps to interpret the dataset's overall characteristics and identify trends or outliers .

To determine the distribution function (CDF) for a discrete random variable, sum up the probabilities of all possible outcomes less than or equal to each value. For the given probabilities P[X=x], calculate cumulative probabilities: F(x) = P(X ≤ x). For the conditional probability P (0 < X < 3 | X < 2), use the formula P(A|B) = P(A ∩ B) / P(B). Compute P(X < 2), which sums relevant probabilities, and find P(0 < X < 3 ∩ X < 2) by identifying common values between the events. Substitute these into the formula to calculate the desired probability .

The binomial probability model applies to quality assessment by evaluating the likelihood of defects. Each item produced is a trial with two outcomes: defective or not. The binomial model calculates the probability of observing a specific number of defects in a sample. For instance, finding the probability that more than 4 pins are defective uses the cumulative binomial probability formula: P(X > 4). By assessing these probabilities, manufacturers can estimate defect rates and decide if processes meet quality standards. If P(X > 4) is high, it indicates potential process issues, prompting further investigation .

You might also like