0% found this document useful (0 votes)
2 views186 pages

Probability Set Solution

The document provides an overview of statistics, including its introduction, applications in engineering, types of variables, data sources, and methods for presenting and classifying data. It covers key concepts such as measures of central tendency, probability, random variables, sampling, estimation, hypothesis testing, and regression analysis. Additionally, it includes various statistical techniques and their applications, along with examples and solutions to important questions.

Uploaded by

Ramesh Karki
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF or read online on Scribd
0% found this document useful (0 votes)
2 views186 pages

Probability Set Solution

The document provides an overview of statistics, including its introduction, applications in engineering, types of variables, data sources, and methods for presenting and classifying data. It covers key concepts such as measures of central tendency, probability, random variables, sampling, estimation, hypothesis testing, and regression analysis. Additionally, it includes various statistical techniques and their applications, along with examples and solutions to important questions.

Uploaded by

Ramesh Karki
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF or read online on Scribd
sous Chapt 11 12 13 1.4 15 1.6 1.7 18 19 CONTENTS INTRODUCTION OF STATISTICS AND PRESENTATION OF DATA Introduction to Statistics. Application of Statistics in Engineering... Variable, Types of Variable: Numerical and Categorical Variable Sources of Data: Primary and Secondary Source. Presentation and Classification of Data: Stem-and-Leaf Displays Frequency Distribution... Pie-Diagram, Histogram, Frequency Curve, and Frequency Polygon 5 Cumulative Frequency Curve or Ogive Curve. SOLUTION TO IMPORTANT QUESTIONS... SUMMARIZING AND DESCRIBING THE NUMERICAL 2.1 22: 2.3 24 3.1 hapter DATA Measure of Central Tendency, Partition Values. Measure of Variation... Coefficient of Variation. Box and Whisker Plot... SOLUTION TO IMPORTANT QUESTIONS. PROBABILITY Random Experiment, Sample Space, Event and Types of Events, Counting Rule...... es) 32 33 34 35 Various Approaches to Probal Laws of Probability - Additive, Multiplicative Conditional Probebility and Independence Baye's Theorem. SOLUTION TO IMPORTANT QUESTIONS. ‘ EnmEE Gomme oS BIVARIATE RANDOM VARIABLES AND JOINT ty PROBABILITY DISTRIBUTION 5,1 _ Joint Probability Mass Function, Joint Probability Density Function, Joint Probability Distribution Function ....178 5.2 Marginal Probability Mass Function, Marginal it it ional Probabilit RANDOM VARIABLE AND PROBABILITY Marie et eee 41 42 43 44 AS 46 DISTRIBUTION 53. Sums and Average of Random Variables Random Variable: Discrete and Continuous Random SOLUTION TO IMPORTANT QUESTIONS. Variable Probability Mass Function. Expectation, Laws of Expectation (Addition and Product Law) on ter = 7 SAMPLING AND ESTIMATION gj 6.1 Population and Sample. Discrete Probability Distribution, gy 62. Sampling Distribution of Sample Mean.. 4.4.1 Binomial Distrbutio oy 63, Types of Sampling . 442 Poisson Distribution. it 64 Determination of Sample Size. 443 Hypergeometric Distribution in) 65. Central Limit Theorem and its Application 44.4 Negative Binomial Distribution io AS ecmatlon... 6.6.1 Concept of Point Estimation and Interval Estimation ..... 103 6.6.2 Criteria of Good Estimator. 6.6.3 Maximum Likelihood Estimation. 10 6.7 Confidence Interval for Population Mean and Population Proportio SOLUTION TO IMPORTANT QUESTIONS. Probability Density Function, Cumulative Distribution Funetion, Expected Values of Continuous Random Variables... Continuous Probability Distribution 4.6.1 Rectangular Distribution 4.62. Exponential Distribution 46.3 Gamma Distribution 464 Normal Distribution... 4.65 Log-Normal Distribution 4.6.7 Beta Distribution TESTING OF HYPOTHESIS 7.1 Hypothesis... 7.2 One Sample Test for Mean and Propo 233 235 souueasune BBY 73 Two Sample Test for Mean (Independent and Dependeny and Proportion sre SOLUTION TO IMPORTANT QUESTIONS... 24 SIMPLE LINEAR REGRESSION AND CORRELATION 8.1 Simple Correlation and its Properties. 82 Concept of Simple Regression Analysis. 8.2.1 Estimation of Regression Coefficient by using Least Square Estimation Method . 8.3 _ Standard Error and Coefficient of Determination 84 Inference Concerning Least Square Method SOLUTION TO IMPORTANT QUESTIONS... INTRODUCTION OF STATISTICS AND PRESENTATION OF DATA 1.1 Introduction to Statistics Everything dealing with the collection, processing, analysis and Interpretation of numerical data belongs to the domains of statistics. Answers provided by statistical approaches can provide the basis for making decisions or choosing actions. There are two branches of statistics: a. Descriptive statistic: It consists of procedure used to summarize and describe the important characteristics of a set of measurement. It is used to describe whole population because of which it becomes too expensive or too time consuming, b, Inferential statistic: It consists of procedures used to make inferences about population characteristics from inforniation contained in a sample drawn from this Population. It is easier, faster, cheaper but less accurate than descriptive statistic. [1.2 “Application of Statistics in Engineering In engineering, statistics can be used to do different diversified tasks. The importance of statistics can be felt in different fields of engineering. Some of the importance can be listed as below: 1. Collection and presenting data of [Link] product for analysis and management. 2. Showing the relationship between two different engineering component. 3. For checking the quality of product and controlling the production. 4. Accurate estimation of time, quantity, qualities of different parameters. Introduction of Statistics and Presentation of Data |1) es) To understand phenomena subject (0 variety ang , effectively predict or control them. 13 Variable, Types ‘of Variable: Numerical ay Categorical Variable Variable is the number of quantity that can be measured 9, Counted. Age, sex country of birth income, education level ete arg Variables. Some of-them can be counted whereas some arg measurable. Numerical variables have quantitative attributes. Height, weight, age, volume, voltage etc are numerical variable Categorical variable is also called qualitative variables which cay be categorized only. They do not have mathematical properties, Voting preference, color of things, breeds of dogs etch are categorical variable. 14 Sources of Data: Primary and Secondary Source Primary data are generated by researcher himself/herself from ‘surveys, interview and experiments, Sources of primary data are population census, mailed questionnaire, direct personal interview ete. ‘Secondary data are second hand information, which is not originally collected but collected from already published or unpublished sources. Sources of secondary data are books, journals, newspapers, websites, government records etc. pa al lait 15 Presentation and Classification of Data: Stem-and- Leaf Displays Es few data, Which sn general huge and unwieldy, need tobe pst Presented in meaningful and readily an form in order to facilitate further statistical * Data can be presented in three broad ways: . Textual presentation b. Tabular presen n © Graphic Graphic or diagrammatic presentation 121 Insights on Probability and Staumice hile presenting data in table, data having some similarity and resemblance should be arranged into groups or classes. This process is known as classification of data. The bases of classification are a, Geographical: arranging data according to geographical b. ranging data according to the order of Quantitative: arranging data according to its numerical magnitude. Qualitative: arranging data on the basis of some attribute orquality Stem and leaf display is a technique that is used to present ‘quantitative data in condensed form. We do not lose information ‘on individual observation which is an advantage over a frequency distribution. In this type of presentation, each value Is divided {into two portions ~ a stem and a leaf. The leaves for each stem are shown separately in a display. For example, to construct stem leaf for two digit number, we split each score into two parts. The first part contains the first digit which is called the stem. The second part contains the second digit, which is called leaf. So for a score ‘of 52, 5 is stem and 2 is leaf. We draw a vertical line and write the stems on the left side and arrange it in increasing order. Then leaves for al scores are put in right side of vertical line. s|207 1 1 1 5 Fig: Stem anc ware oeuan Ruwasr oreo leaf display 16) Frequency Distribution After the classification of data according to quantitative magnitude, the items are classified into groups or classes Introduction of Statistics and Presentation of Data’ |3| 2 5 2 3 according to their incre e, and increasing order in terms of ma Frequency curver magnitude an, Fiistogram, Frea (8 Ple-Diagram, umber of items falling into each group is determined, wr? ined, which own as Frequency Distribution. So the numberof repo : ng in particular group is known asthe frequency ofthat grog For an ideal frequency distribution, number of ees intervals x e determined with Sturge's Rule. According to this rule was of classes (K) is given by oe K=1+3322 logioN Where, N= total number of observations Log = logarithm of a number K= number of classes The size ofthe class interval is then determined Size of lass intra = ES0EE Largest value - Smallest value K 1.7 Diagrammatic and Graphical Presentation of Data: Pareto Diagram Frequency distribution are easier to visualize and comprehend when they are represented graphically or diagrammatically. It also helps to make comparisons between two or more than two sets of data. ‘A Pareto diagram is a bar graph for qualitative data, with the bars arranged according to frequencies. The tallest bar is at left and the smaller bars are farther at the right. Restaurant Complaints UEveEBERS / if 7 a7 MYLIYF Fig: Pareto diagram TA] Insights on Probability and statistics i Frequency PolyBOR —_— rt is the familiar circular graph ‘that shows ‘x ple chal i Prsurements are distributed. rai setae gm tows owe mene Fe of the ar a ayared among the categories with the OTT in his dist ow often a particular category WAS CPT ye an ‘or options oF a arth alfleren categories of OPUONs OF Ate TN oval rawn rm rates are on horizontal axis and ars ae eral axis, The bars inthe istogram ae : her with no gaps between them. adjacent to each ot mee obtained by joining the ‘A frequency polygon Is obiaine points of upper horizontal side of corresponding recian histogram. if there are large number of ‘A frequency curve is obtained Goservations and if the class intervals are taken to be small cnough, it may be possible to have sizeable frequency for most of the classes, Then the frequency polygon will closely approximate a curve, which is called frequency curve. ighboring mid gles of the [lo Cumulative Frequency Curve or Ogive Curve When frequencies are added, they are called cumulative frequencies. The curve obtained by plotting cumulative frequencies is called a cumulative frequency curve or Ogive curve. There are two methods of constructing Ogive, i._Less than ogive i, More than ogive In ‘less than’ method, we start with the upper limits of the classes and go on adding the frequencies. When these frequencies are plotted, we geta rising curve. In ‘more than’ method, we start with the lower limits of the lasses and from the frequencies we subtract the frequency of Be class. When these frequencies are plotted, we get a declining ve. re gs AE {Introduction of Statistics and Presentation of Data 15] es) — SS ao_—hsest—mW Importance _ limitation © Easy to understand J- Provide ; vague complex i Simplified presentation Je Limited information © Reveals hidden facts |» Low precision © Quick to grasp J+ Restrict further data analysis | © Easy to compare l= Possibility of misuse Universally accepted Careful usage | SOLUTION TO IMPORTANT QUESTIONS 1. The amount of money expended in fiscal year 2019p, Curtin University in various categories is shown ig table. Category 1 | Amount (in Millions) Teacher's salary 55 Staffs salary 12 Maintenance 15.01 Research 11 ‘Taxes and others 56 Constructa bar chart. Solution: 16] Insights on Probability and Statistics jon of the Graph Interpretation © je largest amount of money was SPEDE &D WS One athe institute invests small amount " \dful of money on its It also invests han rount of taxes to the government. raw ple chart. ‘The graph sho ‘maintenance. It seem: the research projects: staffs. It also pays huge am« 2, Forthe same data of example 1-1, d we know, complete circle is 360°. So, we divide the circle in required proportion. For teacher's salary, For staff's salary, 12 2 369° 15.2~15° Baz *360°= 152 ‘similarly, for maintenance = 190° Research = 14° Taxes and other = 71° Now, Taxes and others Research| Maintenance 3. Draw a line chart for the following population growth projections and interpret the result. Year __|2010|2020]2030]2040|2050) 85 and older (millions)| 6.1 | 7.3 | 9.6 | 15.4] 209 Introduction of Statistics and Presentation of Data 17] es) 6 —OV—V—_~_——sSs—a Mean (X) 22 where m= middle value 2. Median i. Individual series n+ tn, Median (Mz) =("3--)" item, when n is odd 2n+ 1) . AnT =) item, when nis even il, Discrete series N+ 1) Median (Ma) =("—)" item ii, Continuous series Lee Median (Mg) = T xh [16] insights on Probability and Statistics a _—_ — este 3 Mode 1, Individual series Mode = Maximum repetition Discrete series Mode= Maximum frequency iii. Continuous series Ai Mode=L+a,4a3 *P Where, A: = fi-fo.d2=fi-f2 * fris highest frequency fois preceding frequency “fis succeeding frequency his difference of class Merits and Demerits of Mean ‘The merits and demerits of mean are. Merits: : 1. Itcan be easily understood and calculated simply. It is rigidly defined. Variation of data affects the mean least. It is based on/all the observations. Arrangement of data are not necessary. ‘As the mathematical formula is rigid one, therefore the result remains the same. 7. It “is used when further statistical treatment is necessary. avepy Demerits: 1, It cannot be located graphically. 2, Asingle item can bring big change in the result. 3, Mean. of qualitative forms of data such as cleverness, riches ete. can’t be calculated, 4. It cannot be computed when class intervals have open als ha\ ‘Summarizing and Describing the Numerical Data 117] souueasune BBY extreme valUes. red by 5, Ieisaffecte y formation about spregg 6 It doesn't convey an} tread of data Meritsand DemeritsofMedian ‘The merits and demerits of median are: 1. Ieis simple to understand and easy to calculate, 2 tisnotaffeced by the extreme items in the series, 3. Itcanbe determined graphically. 44. Foropen-ended classes, median can be calculated, 5, Itcan be located by inspection, after arranging the day inorder of magnitude even ifthe data are incomplete 6. It can be located for qualitative factor such as abiliy genius, beauty, etc. Demerits: 1. It does not consider all: variables because it is 3 positional average. 2. The value of median is affected more by sampling fluctuations. 3. It is not capable of further algebraic treatment. Like ‘mean, combined median cannot be calculated. 4. It cannot be computed precisely when it lies between ‘wo items. 5. A slight change in the series may bring drastic change it median value. Merits and Demerits of Mode The merits and demerits of mode are: Merits: 1. Modes the term that occur most in the series, hence it not an isolated value like median nor it is value 't ‘mean that may not be there in the series. 2 It Is not afected by extreme values hence is 6% representative of the Series, ‘Tt tess on Probab aea pgp bility and Statistics 3. Itcan also be found graphically. 4. For open end intervals, it is not necessary to know the length of open intervals. 5. Itcan also be used in case of quantitative phenomenon. 6. With only just a single glance on data, we can find its value. Itis simplest. 7. It is the most used average in day today life, such as average marks of a class, average number of students in ‘a section, average size of shoes, etc. Demerits: 1. Mode cannot be determined if the series is bimodal or ‘multimodal. 2. Mode is based only on concentrated values; other values are not taken into account in-spite of their big difference with the mode. In continuous series only the lengths of class intervals are considered. 3. Mode is most affected by fluctuation of sampling. 4. Mode is not so rigidly defined. Solving the problem by different methods we won't get the same results as in case of mean. 5. It is not capable of further algebraic treatment. It is impossible to find the combined mode of some series as is in case of mean. 6. Also we can’t find the total of whole series from value of mode as is in case of mean. 7. If the number of terms is too large, only then we can call itas the representative value. 8. It is also’ said that sometimes mode is ill-defined, definite and indeterminate. 2.2 Measure of Variation It is the measure of spread of the data. It shows the variation in sizes oF quantities of the items of a group or series. Summarizing and Describing the Numerical Data 119] souueasune BBY of he measares of variably some ange = Largest ~Smalest= eee LS | Coeficient of range *"L+5 | {quartile Deviation cae | % J Qo. 2a = Coefficient of QD-= 9,4 Q, e ‘Mean Deviation (M.D.) Xx | ‘Mean deviation from mean Beal | MD. from mean Coefficient of M.D. ‘mean IIX-Ma| | ‘Mean deviation from median = 5 | Mean deviation from mode = =AGM Coeficient of MD - MP from mode ‘Standard Deviation (SD) Variance = o2 4. Individual series o=9 [Ow N 'l. Discrete and continuous series ox UAE ae Ee N N Ay “ya 120} insi ns on Probability and Stag ahaa ie Measure of central i ident Measure of dispersion It is used to quantify the size of| average behavior. the difference of variables. |s Itrepresents a given set off» It represents extent of data, heterogeneity of given set of data. J+ Mean, mode, median are|* Range, quartile deviation, mean| measures of central| deviation, standard derivation| ‘© It measures variable’ tendency. are measures of dispersion. 2.3 _ Coefficient of Variation. For a given set of data, coefficient of variation (CV) is the ratio of standard deviation to the mean expressed as percentage. It is a unit less quantity, which is used to,compare the variability of two or more than two series. The series having less CV is said to be more uniform, consistent or homogeneous series. CV= Ffean 100% 2.4 Box and Whisker Plot ‘A box plot (also known as box and whisker plot) is a type of chart often used in explanatory data analysis to visually show the istribution of numerical data and skewness through displaying the data quartiles (or percentiles) and average on a rectangular " box aligned either horizontally or vertically. Box plots show the five number summary of a set of data: minimum score, first quartile, median, third quartile and maximum score. OiPin Q. Q Min Max whisker whisker, lower ~~ Median upper guattile quanil Tnterquarile range (IQR) (QO) Outliers ee ee ‘Summarizing and Describing the Numerical Data |21| souueasune BBY eae ee eee ee ‘Outliers are observations that are usually far away from, ma ™, of data. Uperwhisher= 6 15» 00) poner mum vate of date, whichever i lower, Lowerwhisker=Q-15*(1QR) oe minimura value data, whichever is greater, Box plot covers following: 1 Skewness of dataset 2. Dispersion ofa data set 3. Center, average etc of data 4, Identification of outliers” 5, Extent and nature of symmetry. ‘Tipsand Tricks 1. Variation is calculated by calculating variance or deviation (ie.c20r0) ie 2 Use formula to calulate mean, varia ; | variation, st eviation, and check using calculator to recheck the va | ‘SOLUTION TO IMPORTANT QUESTIONS 1. The number of mim lan ried that a person had to wait fr 269,17305 4andag, PSA 110.15, 128 a. Find the values constitu ». Construct a box plot, Solution: ‘Arranging the given d 12,4,5,6,8,9,10,1 Minimum value= 1 Maximum value= 39 a 138 4 =325% number =44025% (5.4) +025 £4.25 {2| insights on Probatm i Statist es ting the 5-number summary latain ascending order, 2,13,15,17,30 a aL 2 8 = 6.5" number gy = 328 -9.750 number 2+ 078 x(13- 12) = 12.75 Upper whisker = Q3+15 *10R = 275+15%(Q-) = 12.75 +1.5 « (1275-425) = 2555 (which is lower than the maximum value hence the value greater than 25.5 are taken as outliers. Upper whisker = 255 »wer whisker Qi- 1.5 «IQR 425-15 *(Q-Q) 4,25 - 15 x (12.75 -4.25) = -85 (which is lower than the minimum value hence, lower whisker = 1 (minimum value) Box plot (td 1 ams 85. ‘12.75 2, Prepare the Box-plot for the following data of daily registration of workers in the construction site 34, 42, 66, 40, 59, 36, 41, 35, 36, 62, 43, 30, 43, 32, 44, 48, 53,50, 48, 38 Solution: Arranging in ascending order, 30, 32, 34, 35, 36, 36, 38, 40, 41, 42, 43,43, 44, 48, 50, 53, 58, 59, 62, 66 ‘Summarizing and Describing the Numerical Data [23] Max. value= 66 Min. value = 30 Upper whisker = Qs+ 1.5 «IQR = Qs+15 x (50-36) =50+15%14 =71 (which shigher than maximum value) | *. Upper whisker is 66 (Which is lower than minimum value) «+ Lower whisker is 30 + Nooutliers below and above whiskers — 30 6 a2 50 A techs #7625 given. Draw a box plot and interpret be 47,51, 48, 52,50, 46,49, 45, 52, 46, 51, 49, 46,51, 49,45, 44, 5 50 | st » 50,48, 49,50, 50 | Arranging in ascending order, 43,44, 44,45, 45, " eke 6 4646, 47,4848, 49, 49,49, 49, 50,50, Max. value= 52 Min. value = 43 240 124] insights on Probability ang Statistics souuvoguies BBY = ak " y= 2 48 Lamas 3n _ 72th =F =7F = 184250 Upperwhisker =Q3+1.5* IQR 0+ 1.5 x (50-46) OF 1S x4 = 56 (which is upper than max. value) + Upper whisker is 52. Lower whisker = Q:~ 1.5 *1QR =46-6 = 40 (which is lower than min. value) :. Lower whisker is 43 valued. -. No outliers below lower whisker B 52 a a) 4. The following table gives the frequency distribution of marks of 800 candidates. fmarks[o-10]10-20/20-30]30-10]40-50[50-e0]60-70]70-60| 80-50 | 90-100, 70] 40 | 50 [40170 130] 100] 70 | «0 | 20 a. Find the median and interpret the result. b. Ifthe minimum mark to pass is 35, what percentage of candidate pass the exam? 3 . What is the no. of candidates who get marks between 34and 79? Solution: a 2 Marks) Freq (Qn mi eh 0-10 10 5 10 10-20 40 15 50 20-30, 80 25 130 30-40) 140 35 270 40-50 170 45 440 ‘Summarizing and Describing the Numerical Data 125] 2 5 2 5 a Se c 57 50-60 st 5 oe 5. An analysis of monthly wages paid to workers in two wo 75 Tae firms A and B belonging to the same industry gives the [70-80 | 85 sae following results: ae 88 280 i FirmA | Firm 90100 | No. of workers 500 | 600 me [Average monthly wages 186_| 175 cian esin = = 400 Variance of distribution of wages 1100 Nets i scbussb dea i, Which firm, A or B, has a larger wage bill? ‘cf just greatér than 400 is 440 an POnding of ii. In which firms, A or B, is there greater variability in 4001s 40-50. individual wage? : Now, fil, Calculate combined mean and combiied variance of : the wage of firm A and firm B? ae Solution: MgzL+—o*h For firm A ' 400-270 Total wages = 500 x 186 = Rs, 93000 =404 Gy 10 For firm B 247667 Total wages = 600 « 175 = Rs. 105000 Interpretation ofthe result Hence, company B pays larger wages bill than company '50% of students has screed more mark 47.67. é re 67. s m on b. [Link] student scoring more than 35 marks Cust 100% Cve= f,* 100% 140 =] +170+130+ 100+70+40+20 = 600 vei 2100 2 = 49g * 100% = 75 100% = 4.84% =5.7143% + The total percentage = $00 . 596 = 75%, candidate, ‘pass the examination, © Theno. of candidate having marks between 34and 79. 140 ="2 +170+130+ 100+70 40. For Guestion b and c, we can use Ogive curve for the curate suitable) Dt the given marks above method is 126] Insights on Probabiity ang Statisticg Hence, company B shows greater variability since Cye > Cvs. (Cy= coefficient of variation) ili, Mean together = M404 * Hens Ta* Ma 186 x 500+ 175 x 600 =" 500+ 600 I = 180 Standard deviation together (0) [(Ga2 + wa?) na + (ou? + jun?) np — (He a+ ne Summarizing and Describing the Numerical Data [27] es) (100+ 175)? x 600 - (189); 0 ar 108) #50 500+ 600 1 se (000 Families is given as het, ruency distribution is 87, feat [g0-99100-119] a I ian of ee are |40-59160-79 roof Families) 50 | ~ Find the missing frequencies Solution: Converting incusive classes into exclusive class as give, 500 [= 50 below. sa — offamilles (9) cf 395-595 50, 50 595-795 xsay) 50+ 795-995 500 550 +2| 1195 400-2 950 1195 1395 50 1000 17 Sf 1000 Given, median (Ma) = 87 So,median lies in (795 - 995) class ‘We have, N 7-cF Ma=Le=F— xh 187 =793 250 (2 + 50) 01,87 = 795 +e) 20 or,z= 263 and 400-2= 400-263 =137 families © ++ Missing frequencies are 263 and 137. 7. Calculate approxi from the fon owing measures of central tendenl [Way Rs/week Thess than 35] 35-37 [30-37 alee| 13 LU Mootwayaseamed] 14 |-ap 195 128] insights on Probability ang: Statisties Solution: [Wages /weete nsy]oa]%@-98W28=5) im | et 315-345 [a3] 14 762 4 345-375 82 2952 | 96 375-405 99 3861 | 195 405-435 18 756 | 213 435 - 465, 7 315__| 220 220 _|Sfm = 8346] Mean (X) = "FF Position of median =(y Item = 110% ftem Ge, just greater than 110 is 195 and class is 375 - 405 class. a Ma = =375 4795 3 =37.924 Mode class is 37.5 - 40.5, since it has highest frequency. Here, f\=99 f= 18 2 Aa=fi-=99-18=81 Ay it de * + Mode = L+ h 17 =375 +7738] «3 =38.02 A civil engineering monitors water quality by measuring the amount of suspended solids in a sample of river water. Over 11 week days, he observed 14, 21, 21, 28, 30, 63, 29, 65, 55, 19, 20 suspended solids (parts per. million) find the third quartile and interpret its meaning. ‘Summarizing and Describing the Numerical Data [29] es) Te oe tutor So equate aan 3419,20.21, 21.28 Position of #3 xf x er 4) |= 9% position she data in position is 55- Interpretation: jean 75¥ of data is below oF equal to SS pars py million suspended solids. A eat , conducted by the department at mechanicy * Se at Virginia Tech, the steel reds supplied by tub different companies were compared. Ten samy Spring were made out of the steel rods. Supplied by sh vcompany and a measure of flexibility was recorded for each. The data are as follows. Company A:93,88, 68,87,85, 6.7, 8.0, 6.5, 9.2,7.0. Company B : 11.0, 98,99, 10.2, 10.1, 9.7, 11.0, 11.1, 102 96 4. Calculate the sample mean and median for the date for the two companies. b, Calculate the sample standard deviation. ‘€ Which company is more consistent? Solution: Sample mean of company A, $.5+6.7168+7+8+8518.748,8+9.249.3 SITET BUESHAT+86092893 | ‘Arranging in ascending order, 65,67,8,7.0,80,85,87,88,92,93 20+ 1 Median ti oa Heson\ 7) item when nis even 5.25titem= Similarly, sample mean of company B 0-2 = 10.26 ‘Arranging in ascending order 9.6,97, 98,9, 10.1, 10.2, 10.2, 11.0, 11.0, 11.1 ‘an+ 1, Median =(“"7) “item when nis even =10.125 . Forstindard deviation [company A company B eas x 65 | 4225 | 96 92.16 67 |, 4499 | 97 94.09 68 |" 4624 | 98 96.04 7 49 99 98.01 8 64 | 101 | © 102.01 es | 7225 | 102} 10404 a7 | 7569 | 102} 104.04 es | 7744 | 110 121 92 | ses | 110 121 93 | 649 | 11a 123.21 EX? = 642.89 EX?= 10556 on=\ ewe m= \ Be 642.83 fp - (7.95)? = 1.0423 130] Insights on Probabity and Stacy ‘Summarizing and Describing the Numerical Data 132] es) on SE oma! =054 ‘c. For consistency we calculate CV x 100% 25. 26% 1.0423 “795 100% = fag * 100% ‘As Cre < Cra, we can conclude company B is mon nt. 10. The following table represents the ‘marks of 1% students. [Marks + |0-20]20-40|40-60[60-80[80-100] ([Link] students] 14 | 18 | 27 26 15 Find the mean, median and standard deviation of all 100 students. Solution: arks {2 No-of «|v. % | Marks Smadeats (0) mcf} mf | m? | tm 0-20 4 10/14] 140 | 100 | 1400 20-40 18 30} 32} 540 | 900 | 16200 40-60) 27° 50/59] 1350 | 2500| 67500 60-80) 26 | 70).85 | 1820 | 4900| 127400 60-100] 15 | 90 |100} 1350 | e100 | 121500) Bren ‘Imf= Sf ps oe 5200 [334008 Mean Se 8200 ("Sp = Toy =52 ‘ Position of median (M) =. CE just greater th. is40-60 100 = 2-50" item I an 50 is z and class corresponding o* 132] Insights on Probatieyand Stace =25.22 11. In two companies A and B engaged in similarly type of industry, the average weekly wage and standard deviation are given below: [LE [company [Company | [Average weekly (Rs)|__ 460 490) [Standard deviation [50 40 No. of wage earners | 100 80. i. Which company pays larger amount as weekly wage? il, Which company shows greater variability in the distribution weekly wages? . What is the mean and standard deviation of all the workers in two companies taken together? Solution: For company A For company B Total wages = 460x100 Total wages = 490x80 = 46000 9200 Hence, company A pays larger amount that company B. Hh CVn =a % 100% ‘Summarlang and Describing the Numerical Dats 1331 eae company Ashows greater VAD Since Cy, (G= coefficient of variation) gna sya ‘ii, Mean together=" nye ne 100 +80 ee 473.363 Standard deviation together (9) oe wat) na + (52+ as?) rg = (1)? a+ SOE 4602) 100 + (402 + 4902) 80 - (47333; : 100480 = 48.22 12, The mean and standard deviation of 20 items is found to be 10 and 2 respectively. At the time of checking it was found that one item 8 was incorrect. Calculate the mean and standard deviation. If (a) the wrong item is ‘omitted, (b) itis replaced by 12, Solution: X=10 Standard deviation (0) = 2 | +2 When wrong item is omitted (Le. wrong item = 8)n = 19 Now, 10*20-8 Mean="Tq = 10.10 Standard deviation g =~ [22+ 10?) x 20 - 6? 2 ets 10 = 20- ~ (10.10) +2023 b. When tis replaced by 12, Mean = 10220-8412 102 134] insights on Probabiity ang Statistics es) New standard devi: __, [EMO 20 BE o 20 = V108- 104.04 £1.99 2 13, Ina moderately asymmetrical distribution the value of mean and median are 20 and 24 respectively: Find the values of mode. Solution: Given, ina moderately asymmetrical distribution: Mean (X) = 20 Median (Ma) = 24 Mode = 3 x median - 2 * mean =3x24-2%20=32 3 14, Two different of a statistics class take the same quiz, and the scores are recorded below. a. Find the, range and standard deviation for each section, b. What do the range valves lead you to conclude about the variation in the two sections? ¢ Whyis the range misleading in this case? . What do the standard deviation valves lead you to conclude about the variation in two sections? Section | 12020] 20] 20[20]20| 20] 20| 20] 20 Section2|2/3 [4 [5 [6 [14]15]16]17[ 18] 19 Solution: “a, Forsection1 b. For section 2 Range (R2) ‘Summarizing and Describing the Numerical Data 135] es) “1, Standard deviation (62) ee @y 6421 ™ Range in the both sections concludes thatthe var, eipoth secons is near about the same, Hoye’ verication in first section 1S TeSS COMPATE to seeye section. ce. Range is misleading because it only depends on smal, and target valve and doesn't encounter the intermedi valves. 4. Though there is single huge variation in section first ayy there is alot variation in section second, the standix deviations are close. 15, The mean weight of 100 students in a certain class iss) kg. The mean weight of boys in the class is 65 kg ant that of girls is 50 kg, Find the number of boys and gi: inthe class. 4 Solution: Given, Total student () = 100 ‘Mean weight (X) =59 ‘i Forboys Mean weight (X .) = 65kg Forgirls Mean weight (X .) = 50kg We know, my #n¢= 100 Now, mE meFyrnak, F, 100 « 59 = (109 - 2 m4 and,ns= 100-40 = 60 M)* 65 +n,x 50 ‘ 1961 nse on Prob ana tmacy In statistics paper, five candidates obtained the marks 16. as 33, 38, 48, 59 and 72, Calculate the mean and standard deviation of these marks. If 10 marks are added for each student, what will be mean and standard deviation. Solution: Xo fe x 33 1089 38 1444 48 2304 59 3481 2 5184 x= 250 | Se=13502 =\ [=e —- (50)? = 14.156 ‘After adding 10 marks to each student, 43, 48, 58, 69 and 82 TEER | se. 4B 1849 48 2304 58 3364 69 4761 82 6724 =X =300 DX? = 19002 2 Meany BE 3 Itshows that standard deviation does not change. ~ Summariing and Describing the Numerical Data 1371 es) ; — marks obtained in an exam by a groy, ee mea? ts was found to be 49.96. The mean on Mi obained in the same exam BY another groyy ma 52.32, Find the mean of then 200 students was Jorained by both the group of the students {a4 together. Solution: For 1 group For 24 group m= 100 m= 200 X= 49.96 X2= 5232 Now, =) pik + meXe Mean ofboth group (X )= 100+ 200 1.53 18. Over a period of 40 days the percentage relative ‘humidity ina vegetable storage building was measured, Mean daily values were recorded as shown below: 60 | 63 | 64 [71 [6773779 | 80 | 63 | 81 86 | 9096 | 98 | 98 | 99139 | 80} 771 78 ‘ 71 | 79 {74 [4 | 85 | @2 [90 | 78] 79 | 79 76 | 80 [62/83 [86/81 [80] 76 | 66 | 74 ‘- Prepare a stem-and leaf display for these data. Show the leaves on each ste I” order of increasing magnitude fi. Draw ‘a box plot for in practical manner. Solution: Sorting the give data: 60, 63, 64, 66, 67,71, 73, 64, 66, 67, 71, 71, 7 79, 79, 79, 80,80, 99 3,74, 74, 76, 7, 78, 78, 78, 7% 80,8 ; 86,89, 80, 90,9656 9,65 1 82 82, 83,83, 04, 85,95 these data and interpret the data 138 Insights on Probability and Statice if (2017 Spring} at Stem & leaf diagram: 6 |0,3,4,6,7 7 |1,1,3,4,4,6,7,8,8,8,9,9,9,9 8B [0,0,0,0,1,1,2,2,3,3,4,5,6,6.9 9 |0,0,6,8,8,9 For box plot, we need to find three quartiles, minimum and maximum value. Minimum value = 60 Maximum value = 99 40+ 1) _(4iyh sequarie=(2225)"=(22)* = ras = 74+ (76-74) x 025 2745 (n+ 240+)" 9 ih. 20e Ue 205 0 obs. + (21+ obs. - 20% obs.) 0.5 = 80+ (80-80) «05 0 3(n+1)"_3(40+19% 3 quartile =S= NU" _ 340+ NY 39758 = 30H obs. + (31% obs. - 30% obs) « 0.75 = 84+ (85-84) x 0.75 = 84.75 To check outliers, Upper whisker, Qs+1.5 (Qs- Q:) = 84.75 + 1.5 x 10.25 = 100.125 Lower whisker, Q-15 (Qs- Qi) = 745-15 * 10.25 = 59.125 Since no values of observed data are above upper whisker and below lower whisker, there don't exist any outliers. Box plot (in graph) Interpretation: From graph, the length of box i.e. Qs- Qu is very short, which means there is less variation in data. Also Q lies-in the middle of Q; & Qs, so data is symmetrical ie. 204 quartil ‘Summarizing and Describing the Numerical Data [39] es) are concentrated atthe Cmte Since there most of data ar® me observations are present. no outliers, 10 ex woof pMax.99 of} |g, 80 2, a 70 ot Min. 60 $0: 40 30) 20; 10 0 19. From the following distribution of 500 students of a college find the marks if only 20% students had failed and also minimum marks obtained by the top 25% of the students. [Marks [0-20/20-40]40-50|50-60]60-80|80-100) (No. of students| 50 [100 | 150 | 90 | 60 | 50 Also represent the data by histogram to locate the mode, Making class intervals equal width, Marks Frequency cf, 0-20 50 50 20-40 100 150 40-60 240 390, 60-80 60 450 80-100 50 500 [2020 Fall] Solution: Marks {[Link] students (f)| Cumulative frequency (cf) 0-20 50 0 2040. 100 150 40-50 150 300 50-60 90 390 60-80 Py a a 500 ar, 500 2 If 20% of the students failed, then pass marks is given by Pap ie. 20% percentile, th = 500)" Paotesin NS)” te. (2252)" «1008 position ive, in 20-40 interval cf 100~ © $0, Pi9=14=—G—xh 100-50 2208 9p * 20 =30 © 30 is the pass marks, To find minimum marks of top 25% students, we have to find Ps. 0) ; Prslies in (75 x39) "= 375% position Which lies in 40-60 interval. 375-150 2. Prs = 40+ 23180 20 = 58.75 is the minimum marks obtained by top 25% students To locate the mode by histogram, Summarizing and Describing the Numerical Data [44] es) 220 x 2 HY oO 8 100 Mode=49 We join the right corner of modal class rectangle by st line to top right comer of preceding rectangle and top let comer of modal class rectangle is joined to top left corners! rectangle on the right. From the point of intersection of these two [Link],a perpendicular is drawn on x-axis. Tt Fag tithe medal value, From graph, the mode valued xis 49, 20. pe Measurements that follow are furnatt semiconductor cried on successive batches in # 953, 950 oy manufacturing process (units are): i calcu, 55951 989,957,954, 955 te sa E é Standard device = 80, sample variance, # il Constru Solution, | *P*Plot ofthe data, 20z0 Fall ‘Sample mean «2% _ 953 +9504 ple mean =, s. 1988 + [42] Insights on. Probability ang Stain 8164430 3 x (952.44)? = 9.527 ‘Standard deviation = +S? =-9.527 = 3.086 Construction of Box Plot: Arranging data, 948, 949, 950, 951, 953, 954, 955, 955, 957 Minimum value = 948, N=9 = 204 obs, + (3"!- 2 obs) x 0.25, += 949 + (950-949) «0.25 = 949.25 : yh oy &=(2«3)"=(2)"=45e = 951 + (953-951) «0.5 =952 gyh we(oe8)*- (ora = 954 + (955-954) x 0.75 = 954.75 For upper whisker, Qs + 1.5 (Qs~ Qi) = 963 For lower whisker, Qi - 1.5 (Qs - Q1) = 941 Outliers does not exist. ‘Summarizing and Describing the Numerical Data 143]

You might also like