Module01 Slides MATH7016
Module01 Slides MATH7016
Module 1
Outline
1 Motivation
2 This Subject
3 R
1 Motivation
2 This Subject
3 R
3 / 44
Introduction
4 / 44
What is Data Science?
The term “Data Science” has existed for about a decade. In the early days, many people working with
data called themselves data scientists, which lead to confusion about what data science is.
5 / 44
What is Data Science?
6 / 44
Applications of Data Science
The world collects more and more data every day and in every field.
Business — banking and insurance records, transaction databases, loyalty cards, checkout data,
advertising
Social Science — Social networks, crime data
Science — Astronomy, Genomics, Physics
7 / 44
Data Science Problems
Biotechnology
Mapping genes and relating them to disease
Understanding metabolites and their relationship to the environment
Finance
Credit Risk management
Customer Segmentation
Predictive Modelling
Fraud detection
Energy
Forecasting demand
Government
Tax Evasion
Security
Health care
8 / 44
Data Science Problems
Retail
Recommendations
Loyalty programs
Supply Chain
Telecommunications
Customer behaviour modelling
Network Optimisation
Social Media and Viral campaigns
Churn
9 / 44
Outline
1 Motivation
2 This Subject
3 R
10 / 44
This Unit
11 / 44
Unit Philosophy
This is about practical Data Science with computers. We assume very little prior knowledge — just
some High School maths. The assessments are practical based, and so there is no written formal
exam.
We will use an open source statistical package called R, and a front end called RStudio. They should
be on all lab machines in the University and you can download and install them for free on Windows,
Mac, and Linux.
[Link]
[Link]
12 / 44
Topics
Lecture Topic
1 Introduction and “Are the digits of pi random?”
2 Counting Eels and Iraqi Refugees
3 Maternal smoking and birth weight
4 Maternal Smoking and Birth Weight (cont)
5 Mapping disease
6 Observation or Experiment?
7 Do taller people earn more?
8 No really, do taller people earn more?
9 Do redheads have a lower pain threshold?
10 What is Normal anyway?
11 Normality as opposed to being deviant, eccentric or unusual
12 When it all goes wrong.
Subject Coordinator:
Lecturer:
Lab demonstrator:
Franco Ubaudi [Link]@[Link]
Shaira Viaje [Link]@[Link]
14 / 44
Emailing Teaching Staff
All email communications must be sent from your student email account.
14 / 44
Examining the Learning Guide
The learning guide is found in vUWS > MATH7016 > Subject Information
15 / 44
Assessment
16 / 44
Text Book
There is no direct text book for this subject, all of the material will be provided in the lecture
notes and lab notes.
The topics we will be covering can be found in the following text books.
Dalgaard, P. (2008). Introductory statistics with R (2nd ed.). New York: Springer.
Reinhart, A. (2015). Statistics Done Wrong. San Francisco, CA: No Starch Press.
Lock, R. H. (Ed.). (2013). Statistics : unlocking the power of data. Hoboken, N.J.: Wiley.
Zumel, N., & Mount, J. (2014). Practical Data Science with R. Shelter Island, NY: Manning
Publications.
17 / 44
Outline
1 Motivation
2 This Subject
3 R
18 / 44
Introducing R
19 / 44
What is R?
R is a software environment for statistical computing and graphics. It runs on just about any
platform (except iPad!) and is completely free (in the GNU sense).
It is used extensively by academic statisticians for research and teaching and is gaining ground in
business.
It has 18800 extension packages available.
Pros
Its free and open source. It has a large active communinity of contributors, meaning that many
classical and modern statistical methods are available for it. It has the publication quality graphics.
It extendable.
Cons
It has a steep learning curve. No GUI by default. Poor (but improving) memory management;
difficulty with very large data sets.
20 / 44
R Resources
21 / 44
Are the digits of π random?
MATH7016 The Nature of Data
4 The digits of pi
5 What is random?
4 The digits of pi
5 What is random?
24 / 44
Pi
Pi or π is usually defined to be the ratio of the circumference of a circle to its diameter.
Figure: Depiction of pi
25 / 44
Pi
Pi is equal to 3.141593 to six decimal places. But pi is an irrational (it cannot be expressed as a
fraction, although 22/7 is used as an approximation).
Pi is transcendental (it cannot be expressed using a polynomial equation).
Its decimal representation has an infinite number of digits.
If a sequence of digits follows a pattern, then we can fully describe the sequence by describing the
pattern.
If we were to see part of the sequence of π, could we predict the next number in the sequence?
26 / 44
Pi to 500 decimal places
3.141592653589793238462643383279502884197169399375105820974944592307816406286208
99862803482534211706798214808651328230664709384460955058223172535940812848111745
02841027019385211055596446229489549303819644288109756659334461284756482337867831
65271201909145648566923460348610454326648213393607260249141273724587006606315588
17488152092096282925409171536436789259036001133053054882046652138414695194151160
94330572703657595919530921861173819326117931051185480744623799627495673518857527
248912279381830119492
Can we see a pattern or does any sub-sequence appear to be random? What does it mean to be
random?
27 / 44
Some other numbers to 500 decimal places
1/9 =
0.111111111111111111111111111111111111111111111111111111111111111111111111111111
11111111111111111111111111111111111111111111111111111111111111111111111111111111
11111111111111111111111111111111111111111111111111111111111111111111111111111111
11111111111111111111111111111111111111111111111111111111111111111111111111111111
11111111111111111111111111111111111111111111111111111111111111111111111111111111
11111111111111111111111111111111111111111111111111111111111111111111111111111111
1111111111111111111111
28 / 44
Some other numbers to 500 decimal places
1/81 - 10/89999999991 =
0.012345678901234567890123456789012345678901234567890123456789012345678901234567
89012345678901234567890123456789012345678901234567890123456789012345678901234567
89012345678901234567890123456789012345678901234567890123456789012345678901234567
89012345678901234567890123456789012345678901234567890123456789012345678901234567
89012345678901234567890123456789012345678901234567890123456789012345678901234567
89012345678901234567890123456789012345678901234567890123456789012345678901234567
89012345678901234567890
29 / 44
Some other numbers to 500 decimal places
e=
2.718281828459045235360287471352662497757247093699959574966967627724076630353547
59457138217852516642742746639193200305992181741359662904357290033429526059563073
81323286279434907632338298807531952510190115738341879307021540891499348841675092
44761460668082264800168477411853742345442437107539077744992069551702761838606261
33138458300075204493382656029760673711320070932870912744374704723069697720931014
16928368190255151086574637721112523897844250569536967707854499699679468644549059
879316368892300987931
30 / 44
Outline
4 The digits of pi
5 What is random?
31 / 44
What it means to be random
It seems that random in this context is not all that well defined.
For this type of example there are generally two aspects to random.
1 The digits should occur the same proportion of the time (uniformity)
2 They should not be predictable, given part of the sequence/expansion we can’t predict the next
digit.
So is it true for π? So far, no one has been able to prove that it is not true.
32 / 44
Are the digits uniformly distributed?
A set is uniformly distributed if each object is expected to appear the same number of times
(e.g. rolling a fair dice). For the first 500 digits of π we get the following table of digit frequencies (if
we ignore the leading “3.”)
0 1 2 3 4 5 6 7 8 9
45 59 54 50 53 50 48 36 53 52
If they are uniform, then we should expect to see equal numbers of each digit.
33 / 44
Are the digits uniformly distributed?
Since there are ten possible digits and we looked at the first 500 we expect to see 50 of each if they
are uniform…
0 1 2 3 4 5 6 7 8 9
45 59 54 50 53 50 48 36 53 52
34 / 44
Random Digits.
In the following table, there are the counts of truly ¹ random digits. Each row is one replication of
500 random digits. What do you notice?
0 1 2 3 4 5 6 7 8 9
48 50 48 57 51 60 42 45 61 38
60 52 40 47 42 52 51 55 55 46
41 47 50 47 38 62 63 52 48 52
58 43 45 47 64 51 52 45 54 41
57 45 59 37 53 55 49 48 47 50
52 63 47 58 42 46 49 41 52 50
55 58 49 50 41 48 43 45 54 57
50 56 48 47 54 44 48 53 55 45
57 55 44 38 52 48 53 51 52 50
44 54 49 52 55 47 45 52 53 49
There are quite a few differences from the expected count of 50. In fact, the largest is 64 and the
smallest is 37
But how do these numbers compare to the digits of π? It is difficult to compare sets of numbers, so
instead, we want to summarise the sets into single measurements (numbers). These summaries are
called statistics.
Statistics that you might have used before are mean (average) and standard deviation.
We want a statistic that tells us how different the numbers are from what we expect.
36 / 44
Random Digits
It is a bit difficult to see the differences in the table from the expected count of 50. So we can subtract
it. Remember each row is different set of 500 digits.
0 1 2 3 4 5 6 7 8 9
-2 0 -2 7 1 10 -8 -5 11 -12
10 2 -10 -3 -8 2 1 5 5 -4
-9 -3 0 -3 -12 12 13 2 -2 2
8 -7 -5 -3 14 1 2 -5 4 -9
7 -5 9 -13 3 5 -1 -2 -3 0
2 13 -3 8 -8 -4 -1 -9 2 0
5 8 -1 0 -9 -2 -7 -5 4 7
0 6 -2 -3 4 -6 -2 3 5 -5
7 5 -6 -12 2 -2 3 1 2 0
-6 4 -1 2 5 -3 -5 2 3 -1
37 / 44
Random Digits
Some of these differences are negative and some are positive, but they are all distances to the
expected count. We can get rid of the sign by squaring.
0 1 2 3 4 5 6 7 8 9 Total
4 0 4 49 1 100 64 25 121 144 512
100 4 100 9 64 4 1 25 25 16 348
81 9 0 9 144 144 169 4 4 4 568
64 49 25 9 196 1 4 25 16 81 470
49 25 81 169 9 25 1 4 9 0 372
4 169 9 64 64 16 1 81 4 0 412
25 64 1 0 81 4 49 25 16 49 314
0 36 4 9 16 36 4 9 25 25 164
49 25 36 144 4 4 9 1 4 0 276
36 16 1 4 25 9 25 4 9 1 130
The total column is just the sum across each row of the squared values.
38 / 44
Random Digits
The totals of the squared differences from the expected count of 50 for 10 sets of 500 random digits
again are;
512 348 568 470 372 412 314 164 276 130
39 / 44
1/81
0.012345679012345679012345679012345679012345679012345679012345679012345679012345
67901234567901234567901234567901234567901234567901234567901234567901234567901234
56790123456790123456790123456790123456790123456790123456790123456790123456790123
45679012345679012345679012345679012345679012345679012345679012345679012345679012
34567901234567901234567901234567901234567901234567901234567901234567901234567901
23456790123456790123456790123456790123456790123456790123456790123456790123456790
12345679012345679012346
0 1 2 3 4 5 6 7 8 9
56 56 56 56 56 55 55 55 0 55
40 / 44
1/81
In fact for 1/81 this total is around 2.5 times the largest value obtained when looking at random
uniform digits (at least for 10 sets)
sum of squared differences = 2780
512 348 568 470 372 412 314 164 276 130
41 / 44
Is ten samples enough?
We only compared to 10 samples of random numbers. Maybe we should look at a few more.
So for 1/81 the total squared difference is still larger than largest from 10000 sets of random digits.
But for π, its well under.
This suggests that the digits of π is consistent with the uniform random digits, whereas 1/81 is not.
42 / 44
Summary
43 / 44
Postscript
In this lecture we looked at one aspect of whether the digits of pi are random: Uniformity of the
distribution of single digits.
We computed the distance of digit frequencies from an expected set of frequencies under
uniformity.
And compared that distance to distances similarly achieved using known random digits.
If the distance is comparable (it was) there is no evidence that the digits of pi are not uniformly
distributed.
This is rudimentary Hypothesis Testing which you may have seen before, and we will be covering in
depth in this Subject.
44 / 44