0% found this document useful (0 votes)
8 views27 pages

Data Analytics

The document discusses various statistical concepts, focusing on normal distribution, descriptive statistics, and inferential statistics. It provides examples of calculating mean, variance, and probabilities related to normally distributed data, as well as the differences between descriptive and inferential statistics. Additionally, it covers hypothesis testing and methods for making inferences about populations based on sample data.

Uploaded by

yobrohi415
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF or read online on Scribd
0% found this document useful (0 votes)
8 views27 pages

Data Analytics

The document discusses various statistical concepts, focusing on normal distribution, descriptive statistics, and inferential statistics. It provides examples of calculating mean, variance, and probabilities related to normally distributed data, as well as the differences between descriptive and inferential statistics. Additionally, it covers hypothesis testing and methods for making inferences about populations based on sample data.

Uploaded by

yobrohi415
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF or read online on Scribd
34 Data Analytics Prob.39. For a certain normal distribution the first momen, is 40 and the fourth moment about 50 is 48. What is the arith mes My and variance of normal distribution ? yp Sol. Suppose m is the arithmetic mean and 6? is the Variance t normal distribution foy n'(0) = E(x — 10) = E(x) — 10 = m— 10 = 40 (given) Thus m= 40+ 10 = 50. Again since, mean = 50, we have 114 (50) = My = 304 = 48 (given) ot=16>07=4 | Hence, mean = 50 and o? = 4 Ay Prob.40. If the probability of committing an error of magnitude a given by h enix? yy sir | Compute the probable error from the following data — m, = 1.305, m, = 1.301, m; = 1.295, m,= 1.286, ms; = 1.318, mg, = 1.321, m, = 1.283, mg = 1.289, my, = 1.300, my, = 1.286. Sol. From the given data which is normally distributed, we have . _ ty _ 12.984 Bil Mean = 10 mj = i oe 1.2984 and 6? = paca — Mean) 1 ary [(0.007)? + (0.003)? + (0.003)? + (0.012)? + (0.0292 + (0.023)? + (0.015)? + (0.009)? + (0.002)? + (0.012)3] or, 07 = 0,0001594 where o = 0.0126 2 Probable error = ae = 0.0084 (Approx.) A Prob.4l, A factory turns out an article by mass production and itis fo that 10% of the product is rejected, Find the S.D. of the number of rejects the equation to the normal curve to represent the number of rejects. «Sol. Here Di Too =0.1 “gq = 1-p=1-0.1=0.9, andn = 100 ~. Binomial distribution of rejects gives, mean = np = 10, S.D.= Jnpq =3 Descriptive Statistics 35 binomial distribution is approximated bya normal distribution, then to the normal curve is i = —100_,-(«-m)?/262 —(x-10)2/5¢4)2 . e109? 203) — 100 -0-10)7/18 ain Ans. . The mean height of 500 students in 151 cm and the standard vial is 1S cm. Assuming that the heights are normally distributed. Fi i how many students have heights between 120 and 155 cm. ([Link].V., Dec. 2012, June 2015) ‘Sol. Here m= 1Slem, o = 15 cm, N = 500, we have X=m_— x=151 Phe o 15 ‘When x = 120, : 120-151 = =-2. eZ) 15 1 when x = 155, 155-151 pre See ie i aan sane 0.3 Fig. 1.5 P(120 0002 =— 1.75, where x, = 0.748 ZS Fa 0.756 — 0.7515 Seto) 0°790, 29 = = 2.25 0.002 Area under, z, = —1.75 to z) = 2.25 = (Area from z = 0 to z; =— 1.75) + (Area from z = 0 to z) = 2.25) = 0.4599 + 0.4878 = 0.9477 ber of plugs likely to be rejected = 1000 (1 — 0.9477) = 1000 x 0.0523 = 52.3 “Approximately 52 plugs are likely to be rejected. Ans. Prob.48. A sample of 100 dry battery cells tested to find the length of produced the following results — : Mean (X)= 12 hours, S.D. (o) = 3 hours. Assuming data to be normally distributed. What percentage of battery are expected to have life | . (i). More than 15 hours j (ii) Less than 6 hours { (iii) Between 10 and 14 hours. ' Suppose x denotes the length of the life of dry battery cells. i _ xX=m.x-12 ig a ep (i) When x=15,z=1 P(x > 15) = P(z> 1) = (Area to the right of z = 0) — (Area between z = 0 and z = =i): = 0.5 — 0:3413 7 2=0 z=1 = 0.1587 = 15.87% Ans. ' ; Fig. 1.8 40 Data Analytics (ii) When x=6,2==2 P(x < 6) = P(z <-2) = (Area to the left of z = 0) (Area between z = 0 and z = 2) = 0.5 — 0.4772 = 0.0228 7 =2.28 % Ans. (iii) When x= 10, Ee) =-067 3 When x= 14, mela 2 =067 3 P(10 g 2S eo 8 Se ae ig gS a ce Ss Ss & Ww a % Co = > + 5 BS cs Ss 3 co % 3 & ss 3 a a < an on = ! SS ae iS oles S e a|SSs a Ss S/23s a & 238 Sos } > S Sas gq & Ns a Equal variances assumed not assumed HEADACHE Equal variances 46 Data Analytics the slope, f), is exactly 0. This is sometimes called the “intercept-on} a for obvious reasons, The alternative is that all of the simple linear pay assumptions hold with B; ¢ R. The alternative, non-zero-slope me always fit the data better than the null, intercept-only model; the F aes Wi the improvement in fit is larger than we would expect under the null 9 There are situations where it is useful to know about this Precise ua and so run an F test on the regression. It is hardly ever, however, a 200, “t to check whether the simple linear regression model is correctly because neither retaining nor rejecting the null gives us informati what we really want to know. Suppose first that we retain the null hypothesis, i.e., we do not find 4 Significant share of variance associated with the regression. This could ; because (i) the intercept-only model is right; (ii) B; # 0 but the test does , have enough power to detect departures from the null. We do not know Whi itis. There is also Possibility that the real relationship is nonlinear, but the be linear approximation to it has slope (nearly) zero, in which case the F test W have no power to detect the nonlinearity. Suppose instead that we reject the null, intercept-only hypothesis, Tr does not mean that the simple linear model is right. It means that the latter mod predicts better than the intercept-only model — too much better to be due chance. The simple linear Tegression model can be absolute garbage, with eve single one of its assumptions flagrantly violated, and yet better than the mo which makes all those assumptions and thinks the optimal slope is zero. Neither the F test of B, = 0 vs. B, #0 nor the Wald/t test of the sar hypothesis tell us anything about the correctness of the simple linear regressi model. All these tests presume the simple linear regression model with Gaussi Moise is true, and check a special case (flat line) against the general one (tit! line). They do not test linearity, constant variance, lack of correlation, Gaussianity. Pecifig, ON abo, Q.22. Write short note on t-test. Ans. A t-test is used to compare the mean scores obtained by two grou on a single variable. The critical ratio test or t-test is used for two samy difference of means. Here it is applied to determine the differences betwe means of two scores obtained from the one group based on the two variabl It is very useful when the population variance is not known and when t sample size is small. The formula for estimating the ratio following ANO\ test is — Descriptive Statistics 47 Mean of the first sample M; \ Mean of second sample M> o; | wheres Standard deviation of first sample oO) Standard deviation of second sample Ni Sample size of the first sample. N, = Sample size of the second sample etation of t-ratio — If the calculated t is less than the tabulated of tat 0,05 or 0.01 levels then the null hypothesis is accepted. If the dt is greater than the tabulated t at 0.05 or 0.01 levels then the null jeulate® is rejected. In the present study, if the ANOVA value is significant thesis ! i ! s is further subjected to t-test. nypothes's 9.23. Write the procedure for performing an inferential test. ‘Ans. Here is @ step-by-step procedure for performing inferential statistics. @ Start with a theory (TV violence reduces children’s ability to lent behaviour because it blurs the distinction between real and fantasy InterPr viol jolence)- , f (i) Make a research hypothesis (Children who experience TV jolence will fail to detect violent behaviour in other children). (iii) Operationalize the variables. (iy) Identify the population to which the study results should apply jds in developed nations). (v) Forma null hypothesis for this population (Ho : uy = ul, where is mean accurate detection of violence by children with high prior exposure TV violence; u, is low prior exposure). (vi) Collecta sample of children from the population and run the study. (vii) Perform statistical tests to see if the obtained sample acteristics are sufficiently different from what would be expected under null hypothesis to be able to reject the null hypothesis. (viii) Publish the paper, get famous, find a job at Harvard. 0.24. What do you mean by Hypotheses ? Explain its types. Ans. Hypotheses are educated guesses about possible difference, lationships or causes. Hypotheses are statements of expectation about some haracteristics of a population. Etymologically, hypothesis are made up of ‘o words, “hypo” (less than) and “thesis” (less certain than thesis). It is the ‘sumptive statement of a proposition or a reasonable guess, based upon the Wailable evidence, which the researcher seeks to prove through his study. Hypothesis is a formal affirmative statement predicting a single research utcome, a tentative explanation of the relationship between two or more lables, 48 Data Analytics Simply stated, a hypothesis is an assumption or SUPPOSition top, or disproved. It is a guiding idea, a tentative explana Probabilities which serves to initiate and guide observatio | tion or a ici Pro, le ny edict results 88 4 el data or considerations to predict results or consequences Hypothe measurable and testable, They are of various types based on the mac’ Mey ' N, Search for y which they are tested. Hypotheses are of two types Directional hypothesis @ (ii) Non-directional hypothesis ~ This hypothesis States a elation, bh SI @ Directional Hypothesis between the variables being studied or a difference between experimen treatments that the researcher expects to emerge. Directional hypothesis also be tested as a statistical hypothesis. However, a statistical hypothesis a be stated in the directional form only when there is a complete certainty th the findings will show a relationship or difference in the expected directig This is because the directional hypothesis can be tested using one-tailed test, significance. (ii) Non-directional Hypothesis — If a given hypothesis do i the nature of the relationship between two variables (i.e. whet}, Positive or negative) or it does not indicate the nature/direction of differenc, between two or more groups on a variable (i.e. which roup will perfor better) then it is known as the non-directional hypothesis. othesis are formulated in order to study th In the present study, null hyp. information literacy skills of student teachers and effect of interventiy Programs. directional in nature, as it does not specify th A null hypothesis is non direction of differences between relationships among variables. The ny observed in the sample. Hypotheses are formed to study the existing conditions. Thus, the nul hypothesis is individually tested Statistically in order to decide whether it shoul be accepted or rejected, 9.25. Write short note on tests of hypothesis. Ans, A test of statistical hypothesis is a procedure or a tule for decidin whether to accept or reject the hypothesis on the basis of sample values obtained A hypothesis which is tested under the assumption that it is true is called null hypothesis and is denoted by Ho. Thus a hypothesis which is tested {0 »ssible rejection under the assumption that it is true is known as ml! pothesis. Descriptive Statistica 49 thesis W hich differs from a given null hypothesis, Hy and when mpjected iS called an alternative hypothesis and is denoted by H, a » Write down the rules for testing a hypothesis. pene procedure for testing a hypothesis is as follows 4 (i) Mention the null hypothesis Hp to be tested along with an ative hypothesis Hy. ; (i) Make some assumption such as the sample is random, the pulation js normal, the variances of two different population are equal or own. (ii) Then find the most appropriate test statistic together with its ling distribution. A statistic whose primary role is that of providing a test some hypothesis is called a test statistic. (ivy) On the basis of the sampling distribution make a decision to accept or reject the null hypothesis Hy. Let a die be thrown. Then the hypothesis Ho is that the die is unbiased, i.e. the proportion of aces is 1/6 any number of throws. Let H, be the proportion of aces 1/7. Then the following tule be suggested for testing our hypothesis — ‘Accept Hy if there are more than 17 aces occur in 100 throws. Reject Hy i.e. accept H, if there are more than 17 aces occur in 100 ws. Thus the number 17 has separated the sample sapce into two regions ne for the acceptance and second for the rejection of the null hypothesis (v) Take a random sample and compute the test statistic. If the ted value of the test statistic falls in the acceptance region, then accept null hypothesis Hg. If it falls in the region of rejection, reject the null pothesis and accept Hy. 9.27. Discuss the various techniques used for testing hypothesis. Ans. There are two types of statistical techniques which are used for ing of hypothesis. They are parametric and non-parametric techniques. (i) Parametric Techniques — Parametric techniques can be applied the purpose of testing the hypotheses if the following conditions are satisfied. (a) When the sample is randomly selected. ral (b) When the variances of the various groups are equal or near (c) When the data are in the form of interval scale or ratio scale. (d) When the observations are independent. (e) When the sample size is more than 30. (f) When the data follow a normal distribution. 50 Data Analytics (ii) Non-parametric Techniques ~ When the above ¢, dit \ os have to be us, i Not satistied, the non-parametric techniques have to be used. The n Oh, ea ara, | tests Me act cify normally distributed Population i OF 6, Mp; fe population free tests, as they are not based on the char, the population They do not sp: Varian Sop. 7 The techniques which enable us to compare samples and Make OF tests of significance without having to assume normalit: are known as Non-parametric techniques. Some of the non-parametric techniques are the chi 5 difference Correlation coefficient, the sign test, the media of-ranks test. The non-parametric techniques do not ha Parametric tests, that is, they are less able to detect a t Such is present. Non parametric tests should not be u: other more exact tests are applicable. Data have been c independently, Since the following techni i 'y in the Popa " quare test, ef In test and es, ve the «, Wer rue difference A sed, therefore, Wh ollected randomly wherein every individual had to Tes po all the conditions required for parametric tes . ques are employed. (a) t-test (b) ANOVA (c) w? estimate, ts are Satis 0.28. Explain in detail about multiple hypothesis testing. [[Link]., May 2019 ( VIL-Sem problem is the situation when neously. For example, suppose ; levels for each gene among healt Ans. The multiple hypothesis testing Wish to consider many hypotheses simulta have n genes and data about expression | individuals and those with prostate cance: | Healthy (k Patients) | Prostate Cancer (/ patients Expression Levelof Genei | x1 Sj

You might also like