0% found this document useful (0 votes)
3 views45 pages

Module01 Slides MATH7016

The document introduces the MATH7016 course, 'The Nature of Data', which focuses on data science and its applications across various fields. It outlines the course structure, including topics, assessments, and the use of the R programming language for practical data analysis. The course aims to equip students with the skills to analyze data and build predictive models, assuming minimal prior knowledge.

Uploaded by

Thế Đạt
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views45 pages

Module01 Slides MATH7016

The document introduces the MATH7016 course, 'The Nature of Data', which focuses on data science and its applications across various fields. It outlines the course structure, including topics, assessments, and the use of the R programming language for practical data analysis. The course aims to equip students with the skills to analyze data and build predictive models, assuming minimal prior knowledge.

Uploaded by

Thế Đạt
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Introduction to the Subject

MATH7016 The Nature of Data

School of Computer, Data and Mathematical Sciences

Module 1
Outline

1 Motivation

2 This Subject

3 R

Introduction to the Unit 2 / 44


Outline

1 Motivation

2 This Subject

3 R

3 / 44
Introduction

What is Data Science?


About this subject
Topics
Assessments

4 / 44
What is Data Science?
The term “Data Science” has existed for about a decade. In the early days, many people working with
data called themselves data scientists, which lead to confusion about what data science is.

5 / 44
What is Data Science?

There are many thoughts on what data science is.

It has something to do with data and computer science


It might involve machine learning and statistics
It might involve programming and databases

It is pretty clear it involves data.


The modern definition of data science is the science of extracting information from data using
statistical and computational methods.

6 / 44
Applications of Data Science

The world collects more and more data every day and in every field.

Business — banking and insurance records, transaction databases, loyalty cards, checkout data,
advertising
Social Science — Social networks, crime data
Science — Astronomy, Genomics, Physics

7 / 44
Data Science Problems

Biotechnology
Mapping genes and relating them to disease
Understanding metabolites and their relationship to the environment
Finance
Credit Risk management
Customer Segmentation
Predictive Modelling
Fraud detection
Energy
Forecasting demand
Government
Tax Evasion
Security
Health care

8 / 44
Data Science Problems

Retail
Recommendations
Loyalty programs
Supply Chain
Telecommunications
Customer behaviour modelling
Network Optimisation
Social Media and Viral campaigns
Churn

9 / 44
Outline

1 Motivation

2 This Subject

3 R

10 / 44
This Unit

Everything in Data Science depends on Data


In this subject we will,

Learn what data is and what forms it can take


Use data to answer questions about a population
Use data estimate parameters of a population
Build simple predictive models
Look at common errors in data analysis

11 / 44
Unit Philosophy

This is about practical Data Science with computers. We assume very little prior knowledge — just
some High School maths. The assessments are practical based, and so there is no written formal
exam.
We will use an open source statistical package called R, and a front end called RStudio. They should
be on all lab machines in the University and you can download and install them for free on Windows,
Mac, and Linux.

[Link]
[Link]

12 / 44
Topics
Lecture Topic
1 Introduction and “Are the digits of pi random?”
2 Counting Eels and Iraqi Refugees
3 Maternal smoking and birth weight
4 Maternal Smoking and Birth Weight (cont)
5 Mapping disease
6 Observation or Experiment?
7 Do taller people earn more?
8 No really, do taller people earn more?
9 Do redheads have a lower pain threshold?
10 What is Normal anyway?
11 Normality as opposed to being deviant, eccentric or unusual
12 When it all goes wrong.

The above table should be used as a guide only, as it is subject to change.


13 / 44
Teaching Team

Subject Coordinator:

Gizem Intepe [Link]@[Link]

Lecturer:

Gizem Intepe [Link]@[Link]

Lab demonstrator:
Franco Ubaudi [Link]@[Link]
Shaira Viaje [Link]@[Link]

14 / 44
Emailing Teaching Staff

All email communications must be sent from your student email account.

Emails must contain:


your full name,
student number,
subject name,
clearly state the purpose of the message in a professional standard.

• Don't expect to get an immediate reply.


• Teaching staff is not only teaching this subject, they might be in another lecture or
in a meeting.

14 / 44
Examining the Learning Guide

The learning guide contains a description of:

the content of the subject


what is expected from each student
delivery of the subject
the assessment

The learning guide is found in vUWS > MATH7016 > Subject Information

15 / 44
Assessment

The assessment for the subject is:

5 online quizzes [5 × 4 = 20 marks]


Written project [40 marks]
2 hour lab based exam [40 marks]

You must obtain


• at least 50 of the 100 marks and
• at least 12 marks (out of 40) from the final exam,
to pass this subject.

16 / 44
Text Book

There is no direct text book for this subject, all of the material will be provided in the lecture
notes and lab notes.
The topics we will be covering can be found in the following text books.

Dalgaard, P. (2008). Introductory statistics with R (2nd ed.). New York: Springer.
Reinhart, A. (2015). Statistics Done Wrong. San Francisco, CA: No Starch Press.
Lock, R. H. (Ed.). (2013). Statistics : unlocking the power of data. Hoboken, N.J.: Wiley.
Zumel, N., & Mount, J. (2014). Practical Data Science with R. Shelter Island, NY: Manning
Publications.

17 / 44
Outline

1 Motivation

2 This Subject

3 R

18 / 44
Introducing R

This semester, we will be using R to perform our data analysis.

19 / 44
What is R?

R is a software environment for statistical computing and graphics. It runs on just about any
platform (except iPad!) and is completely free (in the GNU sense).
It is used extensively by academic statisticians for research and teaching and is gaining ground in
business.
It has 18800 extension packages available.
Pros
Its free and open source. It has a large active communinity of contributors, meaning that many
classical and modern statistical methods are available for it. It has the publication quality graphics.
It extendable.

Cons
It has a steep learning curve. No GUI by default. Poor (but improving) memory management;
difficulty with very large data sets.

20 / 44
R Resources

[Link] — Main R website.


CRAN — [Link] — Comprehensive R Archive Network — base software and
add-on packages.
RStudio — [Link] — is a powerful IDE for R
R Commander — [Link](Rcmdr) — is a partial GUI interface to R — requires TclTk.
R Graph Gallery — [Link] — loads of pretty pictures.
[Link] — “A (very) short
Introduction to R”
“Introductory Statistics with R”, Peter Dalgaard, Springer 2008.

21 / 44
Are the digits of π random?
MATH7016 The Nature of Data

School of Computer, Data and Mathematical Sciences

Module 1 - Examining Randomness


Outline

4 The digits of pi

5 What is random?

Are the digits of π random? 23 / 44


Outline

4 The digits of pi

5 What is random?

24 / 44
Pi
Pi or π is usually defined to be the ratio of the circumference of a circle to its diameter.

Figure: Depiction of pi

25 / 44
Pi

Pi is equal to 3.141593 to six decimal places. But pi is an irrational (it cannot be expressed as a
fraction, although 22/7 is used as an approximation).
Pi is transcendental (it cannot be expressed using a polynomial equation).
Its decimal representation has an infinite number of digits.

If a sequence of digits follows a pattern, then we can fully describe the sequence by describing the
pattern.
If we were to see part of the sequence of π, could we predict the next number in the sequence?

26 / 44
Pi to 500 decimal places

3.141592653589793238462643383279502884197169399375105820974944592307816406286208
99862803482534211706798214808651328230664709384460955058223172535940812848111745
02841027019385211055596446229489549303819644288109756659334461284756482337867831
65271201909145648566923460348610454326648213393607260249141273724587006606315588
17488152092096282925409171536436789259036001133053054882046652138414695194151160
94330572703657595919530921861173819326117931051185480744623799627495673518857527
248912279381830119492

Can we see a pattern or does any sub-sequence appear to be random? What does it mean to be
random?

27 / 44
Some other numbers to 500 decimal places

1/9 =

0.111111111111111111111111111111111111111111111111111111111111111111111111111111
11111111111111111111111111111111111111111111111111111111111111111111111111111111
11111111111111111111111111111111111111111111111111111111111111111111111111111111
11111111111111111111111111111111111111111111111111111111111111111111111111111111
11111111111111111111111111111111111111111111111111111111111111111111111111111111
11111111111111111111111111111111111111111111111111111111111111111111111111111111
1111111111111111111111

28 / 44
Some other numbers to 500 decimal places

1/81 - 10/89999999991 =

0.012345678901234567890123456789012345678901234567890123456789012345678901234567
89012345678901234567890123456789012345678901234567890123456789012345678901234567
89012345678901234567890123456789012345678901234567890123456789012345678901234567
89012345678901234567890123456789012345678901234567890123456789012345678901234567
89012345678901234567890123456789012345678901234567890123456789012345678901234567
89012345678901234567890123456789012345678901234567890123456789012345678901234567
89012345678901234567890

29 / 44
Some other numbers to 500 decimal places

e=

2.718281828459045235360287471352662497757247093699959574966967627724076630353547
59457138217852516642742746639193200305992181741359662904357290033429526059563073
81323286279434907632338298807531952510190115738341879307021540891499348841675092
44761460668082264800168477411853742345442437107539077744992069551702761838606261
33138458300075204493382656029760673711320070932870912744374704723069697720931014
16928368190255151086574637721112523897844250569536967707854499699679468644549059
879316368892300987931

30 / 44
Outline

4 The digits of pi

5 What is random?

31 / 44
What it means to be random

It seems that random in this context is not all that well defined.
For this type of example there are generally two aspects to random.

1 The digits should occur the same proportion of the time (uniformity)
2 They should not be predictable, given part of the sequence/expansion we can’t predict the next
digit.

It can be shown that this is equivalent to

single digits occur uniformly,


pairs of digits occur uniformly,
triples of digits occur uniformly
… etc.

So is it true for π? So far, no one has been able to prove that it is not true.

32 / 44
Are the digits uniformly distributed?

A set is uniformly distributed if each object is expected to appear the same number of times
(e.g. rolling a fair dice). For the first 500 digits of π we get the following table of digit frequencies (if
we ignore the leading “3.”)

0 1 2 3 4 5 6 7 8 9
45 59 54 50 53 50 48 36 53 52

If they are uniform, then we should expect to see equal numbers of each digit.

33 / 44
Are the digits uniformly distributed?

Since there are ten possible digits and we looked at the first 500 we expect to see 50 of each if they
are uniform…

0 1 2 3 4 5 6 7 8 9
45 59 54 50 53 50 48 36 53 52

Only two digits have exactly 50 occurrences (3 and 5).


The digit 7 occurs only 36 times.

Why might this be?

They may not be uniform


We only looked at 500 digits
We only looked at the first 500 digits

34 / 44
Random Digits.
In the following table, there are the counts of truly ¹ random digits. Each row is one replication of
500 random digits. What do you notice?

0 1 2 3 4 5 6 7 8 9
48 50 48 57 51 60 42 45 61 38
60 52 40 47 42 52 51 55 55 46
41 47 50 47 38 62 63 52 48 52
58 43 45 47 64 51 52 45 54 41
57 45 59 37 53 55 49 48 47 50
52 63 47 58 42 46 49 41 52 50
55 58 49 50 41 48 43 45 54 57
50 56 48 47 54 44 48 53 55 45
57 55 44 38 52 48 53 51 52 50
44 54 49 52 55 47 45 52 53 49

¹As generated by a computer at least


35 / 44
Random Digits.

There are quite a few differences from the expected count of 50. In fact, the largest is 64 and the
smallest is 37
But how do these numbers compare to the digits of π? It is difficult to compare sets of numbers, so
instead, we want to summarise the sets into single measurements (numbers). These summaries are
called statistics.
Statistics that you might have used before are mean (average) and standard deviation.
We want a statistic that tells us how different the numbers are from what we expect.

36 / 44
Random Digits
It is a bit difficult to see the differences in the table from the expected count of 50. So we can subtract
it. Remember each row is different set of 500 digits.

0 1 2 3 4 5 6 7 8 9
-2 0 -2 7 1 10 -8 -5 11 -12
10 2 -10 -3 -8 2 1 5 5 -4
-9 -3 0 -3 -12 12 13 2 -2 2
8 -7 -5 -3 14 1 2 -5 4 -9
7 -5 9 -13 3 5 -1 -2 -3 0
2 13 -3 8 -8 -4 -1 -9 2 0
5 8 -1 0 -9 -2 -7 -5 4 7
0 6 -2 -3 4 -6 -2 3 5 -5
7 5 -6 -12 2 -2 3 1 2 0
-6 4 -1 2 5 -3 -5 2 3 -1

37 / 44
Random Digits
Some of these differences are negative and some are positive, but they are all distances to the
expected count. We can get rid of the sign by squaring.

0 1 2 3 4 5 6 7 8 9 Total
4 0 4 49 1 100 64 25 121 144 512
100 4 100 9 64 4 1 25 25 16 348
81 9 0 9 144 144 169 4 4 4 568
64 49 25 9 196 1 4 25 16 81 470
49 25 81 169 9 25 1 4 9 0 372
4 169 9 64 64 16 1 81 4 0 412
25 64 1 0 81 4 49 25 16 49 314
0 36 4 9 16 36 4 9 25 25 164
49 25 36 144 4 4 9 1 4 0 276
36 16 1 4 25 9 25 4 9 1 130

The total column is just the sum across each row of the squared values.
38 / 44
Random Digits

The totals of the squared differences from the expected count of 50 for 10 sets of 500 random digits
again are;

512 348 568 470 372 412 314 164 276 130

For the first 500 digits of π the equivalent total is 344.


Is this consistent with the totals obtained for random digits?
If anything this number is on the small side. That is, counts of the digits of π seems to be similar to
that for uniform random digits.

39 / 44
1/81

0.012345679012345679012345679012345679012345679012345679012345679012345679012345
67901234567901234567901234567901234567901234567901234567901234567901234567901234
56790123456790123456790123456790123456790123456790123456790123456790123456790123
45679012345679012345679012345679012345679012345679012345679012345679012345679012
34567901234567901234567901234567901234567901234567901234567901234567901234567901
23456790123456790123456790123456790123456790123456790123456790123456790123456790
12345679012345679012346

0 1 2 3 4 5 6 7 8 9
56 56 56 56 56 55 55 55 0 55

sum of squared differences = 2780

40 / 44
1/81

In fact for 1/81 this total is around 2.5 times the largest value obtained when looking at random
uniform digits (at least for 10 sets)
sum of squared differences = 2780

512 348 568 470 372 412 314 164 276 130

41 / 44
Is ten samples enough?

We only compared to 10 samples of random numbers. Maybe we should look at a few more.

Number of sets Maximum squared difference


10 568
100 1024
1000 1388
10000 1912

So for 1/81 the total squared difference is still larger than largest from 10000 sets of random digits.
But for π, its well under.
This suggests that the digits of π is consistent with the uniform random digits, whereas 1/81 is not.

42 / 44
Summary

Pi is an irrational number — its decimal expansion doesn’t terminate


The expected digit count for 500 uniformly random digits is 50 of each
Measuring the distance of the first 500 digits of pi from this expected counts gives a number
that is not unusual for randomly generated digits
Doing the same for 1/81 gives a very unusual distance
1/81 is not consistent with random digits, pi (probably) is.
We haven’t considered pairs of digits, or triples etc.
We have only looked at the first 500 digits of pi.

43 / 44
Postscript

In this lecture we looked at one aspect of whether the digits of pi are random: Uniformity of the
distribution of single digits.
We computed the distance of digit frequencies from an expected set of frequencies under
uniformity.
And compared that distance to distances similarly achieved using known random digits.
If the distance is comparable (it was) there is no evidence that the digits of pi are not uniformly
distributed.

This is rudimentary Hypothesis Testing which you may have seen before, and we will be covering in
depth in this Subject.

44 / 44

You might also like