0% found this document useful (0 votes)
2 views14 pages

Module01 Notes

This document introduces the concept of Data Science, defining it as the science of extracting information from data using statistical and computational methods. It outlines the applications of Data Science across various fields, such as business, social science, and government, and details the unit's structure, topics, and assessments. The document also emphasizes the use of R for data analysis and discusses the nature of randomness in the context of the digits of pi.

Uploaded by

Thế Đạt
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views14 pages

Module01 Notes

This document introduces the concept of Data Science, defining it as the science of extracting information from data using statistical and computational methods. It outlines the applications of Data Science across various fields, such as business, social science, and government, and details the unit's structure, topics, and assessments. The document also emphasizes the use of R for data analysis and discusses the nature of randomness in the context of the digits of pi.

Uploaded by

Thế Đạt
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Chapter 0

Introduction to the Unit

0.1 Motivation

Introduction

• What is Data Science?


• About this unit
• Topics
• Assessments

What is Data Science?

The term “Data Science” has existed for about a decade. In the early
days, many people working with data called themselves data scien-
tists, which lead to confusion about what data science is.

What is Data Science?

There are many thoughts on what data science is.

• It has something to do with data and computer science


• It might involve machine learning and statistics
• It might involve programming and databases

It is pretty clear it involves data.

The modern definition of data science is the science of extracting


information from data using statistical and computational methods.
2

Figure 1: A business analyst/statistician that


lives in california?

Applications of Data Science

The world collects more and more data every day and in every field.

• Business — banking and insurance records, transaction databases,


loyalty cards, checkout data, advertising
• Social Science — Social networks, crime data
• Science — Astronomy, Genomics, Physics

Data Science Problems

• Biotechnology
– Mapping genes and relating them to disease
– Understanding metabolites and there relationship to the envi-
ronment
• Finance
– Credit Risk management
– Customer Segmentation
– Predictive Modelling
– Fraud detection
• Energy
– Forecasting demand
• Government
– Tax Evasion
– Security
– Health care
3

Data Science Problems

• Retail
– Recommendations
– Loyalty programs
– Supply Chain
• Telecommunications
– Customer behaviour modelling
– Network Optimisation
– Social Media and Viral campaigns
– Churn

0.2 This Unit

This Unit

Everything in Data Science depends on Data


In this unit we will,

• Learn what data is and what forms it can take


• Use data to answer questions about a population
• Use data estimate parameters of a population
• Build simple predictive models
• Look at common errors in data analysis

Unit Philosophy

This is about practical Data Science with computers. We assume


very little prior knowledge — just some High School maths. The
assessments are practical based, and so there is no written formal
exam.
We will use an open source statistical package called R, and a
front end called RStudio. They should be on all lab machines in
the University and you can download and install them for free on
Windows, Mac, and Linux.

• [Link]
• [Link]

Topics
4

Lecture Topic
1 Introduction and “Are the digits of pi random?”
2 Counting Eels and Iraqi Refugees
3 Maternal smoking and birth weight
4 Maternal Smoking and Birth Weight (cont)
5 Mapping disease
6 Observation or Experiment?
7 Do taller people earn more?
8 No really, do taller people earn more?
9 Do redheads have a lower pain threshold?
10 What is Normal anyway?
11 Normality as opposed to being deviant, eccentric or unusual
12 When it all goes wrong.

The above table should be used as a guide only, as it is subject to


change.

Teaching Team

Unit Coordinator:

• Gizem Intepe [Link] @ [Link]

Examining the Learning Guide

The learning guide contains a description of:

• the content of the unit


• what is expected from each student
• delivery of the unit
• the assessment

The learning guide is found in vUWS > 301108 > Unit Information
5

Assessment

The assessment for the unit is:

• 5 online quizzes [5 × 4 = 20 marks]


• Written project [40 marks]
• 2 hour lab based exam [40 marks]

• You must obtain at least 50 of the 100 marks to pass this subject and must take at least 12 marks from
the final exam.

Text Book

There is no text book for this unit, all of the material will be provided
in the lecture notes and lab notes.

The topics we will be covering can be found in the following text


books.

• Dalgaard, P. (2008). Introductory statistics with R (2nd ed.). New


York: Springer.
• Reinhart, A. (2015). Statistics Done Wrong. San Francisco, CA: No
Starch Press.
• Lock, R. H. (Ed.). (2013). Statistics : unlocking the power of data.
Hoboken, N.J.: Wiley.
• Zumel, N., & Mount, J. (2014). Practical Data Science with R. Shel-
ter Island, NY: Manning Publications.

0.3 R

Introducing R

This semester, we will be using R to perform our data analysis.

What is R?

R is a software environment for statistical computing and graphics. It


runs on just about any platform (except iPad!) and is completely free
(in the GNU sense).

It is used extensively by academic statisticians for research and


teaching and is gaining ground in business.

It has 18800 extension packages available.


6

Pros Its free and open source. It has a large active communinity
of contributors, meaning that many classical and modern statistical
methods are available for it. It has the publication quality graphics. It
extendable.

Cons It has a steep learning curve. No GUI by default. Poor (but


improving) memory management; difficulty with very large data
sets.

R Resources

• [Link] — Main R website.


• CRAN — [Link] — Comprehensive R Archive
Network — base software and add-on packages.
• RStudio — [Link] — is a powerful IDE for R
• R Commander — [Link](Rcmdr) — is a partial GUI
interface to R — requires TclTk.
• R Graph Gallery — [Link] — loads
of pretty pictures.
• [Link]
pdf — “A (very) short Introduction to R”
• “Introductory Statistics with R”, Peter Dalgaard, Springer 2008.
Chapter 1

Are the digits of π random?

1.1 The digits of pi

Pi

Pi or π is usually defined to be the ratio of the circumference of a


circle to its diameter.

Figure 1.1: Depiction of pi

Pi

• Pi is equal to 3.141593 to six decimal places. But pi is an irrational


(it cannot be expressed as a fraction, although 22/7 is used as an
approximation).

• Pi is transcendental (it cannot be expressed using a polynomial


equation).
8

• Its decimal representation has an infinite number of digits.

If a sequence of digits follows a pattern, then we can fully describe


the sequence by describing the pattern.
If we were to see part of the sequence of π, could we predict the
next number in the sequence?

Pi to 500 decimal places

3.141592653589793238462643383279502884197169399375105820974944592307816406286208
99862803482534211706798214808651328230664709384460955058223172535940812848111745
02841027019385211055596446229489549303819644288109756659334461284756482337867831
65271201909145648566923460348610454326648213393607260249141273724587006606315588
17488152092096282925409171536436789259036001133053054882046652138414695194151160
94330572703657595919530921861173819326117931051185480744623799627495673518857527
248912279381830119492

Can we see a pattern or does any sub-sequence appear to be ran-


dom? What does it mean to be random?

Some other numbers to 500 decimal places

1/9 =

0.111111111111111111111111111111111111111111111111111111111111111111111111111111
11111111111111111111111111111111111111111111111111111111111111111111111111111111
11111111111111111111111111111111111111111111111111111111111111111111111111111111
11111111111111111111111111111111111111111111111111111111111111111111111111111111
11111111111111111111111111111111111111111111111111111111111111111111111111111111
11111111111111111111111111111111111111111111111111111111111111111111111111111111
1111111111111111111111

Some other numbers to 500 decimal places

1/81 - 10/89999999991 =

0.012345678901234567890123456789012345678901234567890123456789012345678901234567
89012345678901234567890123456789012345678901234567890123456789012345678901234567
89012345678901234567890123456789012345678901234567890123456789012345678901234567
89012345678901234567890123456789012345678901234567890123456789012345678901234567
89012345678901234567890123456789012345678901234567890123456789012345678901234567
89012345678901234567890123456789012345678901234567890123456789012345678901234567
89012345678901234567890
9

Some other numbers to 500 decimal places

e=

2.718281828459045235360287471352662497757247093699959574966967627724076630353547
59457138217852516642742746639193200305992181741359662904357290033429526059563073
81323286279434907632338298807531952510190115738341879307021540891499348841675092
44761460668082264800168477411853742345442437107539077744992069551702761838606261
33138458300075204493382656029760673711320070932870912744374704723069697720931014
16928368190255151086574637721112523897844250569536967707854499699679468644549059
879316368892300987931

1.2 What is random?

What it means to be random

It seems that random in this context is not all that well defined.
For this type of example there are generally two aspects to ran-
dom.

1. The digits should occur the same proportion of the time (unifor-
mity)
2. They should not be predictable, given part of the se-
quence/expansion we can’t predict the next digit.

It can be shown that this is equivalent to

• single digits occur uniformly,


• pairs of digits occur uniformly,
• triples of digits occur uniformly
• . . . etc.

So is it true for π? So far, no one has been able to prove that it is


not true.

Are the digits uniformly distributed?

A set is uniformly distributed if each object is expected to appear the


same number of times (e.g. rolling a fair dice). For the first 500 digits
of π we get the following table of digit frequencies (if we ignore the
leading “3.”)
10

0 1 2 3 4 5 6 7 8 9
45 59 54 50 53 50 48 36 53 52

If they are uniform, then we should expect to see equal numbers of


each digit.

Are the digits uniformly distributed?

Since there are ten possible digits and we looked at the first 500 we
expect to see 50 of each if they are uniform. . .

0 1 2 3 4 5 6 7 8 9
45 59 54 50 53 50 48 36 53 52

• Only two digits have exactly 50 occurrences (3 and 5).


• The digit 7 occurs only 36 times.

Why might this be?

• They may not be uniform


• We only looked at 500 digits
• We only looked at the first 500 digits

Random Digits.

In the following table, there are the counts of truly 1 random dig- 1
As generated by a computer at least
its. Each row is one replication of 500 random digits. What do you
notice?

0 1 2 3 4 5 6 7 8 9
48 50 48 57 51 60 42 45 61 38
60 52 40 47 42 52 51 55 55 46
41 47 50 47 38 62 63 52 48 52
58 43 45 47 64 51 52 45 54 41
57 45 59 37 53 55 49 48 47 50
52 63 47 58 42 46 49 41 52 50
55 58 49 50 41 48 43 45 54 57
50 56 48 47 54 44 48 53 55 45
57 55 44 38 52 48 53 51 52 50
44 54 49 52 55 47 45 52 53 49
11

Random Digits.

There are quite a few differences from the expected count of 50. In
fact, the largest is 64 and the smallest is 37

But how do these numbers compare to the digits of π? It is diffi-


cult to compare sets of numbers, so instead, we want to summarise
the sets into single measurements (numbers). These summaries are
called statistics.
Statistics that you might have used before are mean (average) and
standard deviation.
We want a statistic that tells us how different the numbers are
from what we expect.

Random Digits

It is a bit difficult to see the differences in the table from the expected
count of 50. So we can subtract it. Remember each row is different set
of 500 digits.

0 1 2 3 4 5 6 7 8 9
-2 0 -2 7 1 10 -8 -5 11 -12
10 2 -10 -3 -8 2 1 5 5 -4
-9 -3 0 -3 -12 12 13 2 -2 2
8 -7 -5 -3 14 1 2 -5 4 -9
7 -5 9 -13 3 5 -1 -2 -3 0
2 13 -3 8 -8 -4 -1 -9 2 0
5 8 -1 0 -9 -2 -7 -5 4 7
0 6 -2 -3 4 -6 -2 3 5 -5
7 5 -6 -12 2 -2 3 1 2 0
-6 4 -1 2 5 -3 -5 2 3 -1

Random Digits

Some of these differences are negative and some are positive, but
they are all distances to the expected count. We can get rid of the sign
by squaring.

0 1 2 3 4 5 6 7 8 9 Total
4 0 4 49 1 100 64 25 121 144 512
100 4 100 9 64 4 1 25 25 16 348
81 9 0 9 144 144 169 4 4 4 568
12

0 1 2 3 4 5 6 7 8 9 Total
64 49 25 9 196 1 4 25 16 81 470
49 25 81 169 9 25 1 4 9 0 372
4 169 9 64 64 16 1 81 4 0 412
25 64 1 0 81 4 49 25 16 49 314
0 36 4 9 16 36 4 9 25 25 164
49 25 36 144 4 4 9 1 4 0 276
36 16 1 4 25 9 25 4 9 1 130

The total column is just the sum across each row of the squared
values.

Random Digits

The totals of the squared differences from the expected count of 50


for 10 sets of 500 random digits again are;

512 348 568 470 372 412 314 164 276 130

For the first 500 digits of π the equivalent total is 344


Is this consistent with the totals obtained for random digits?

If anything this number is on the small side. That is, counts of the
digits of π seems to be similar to that for uniform random digits.

1/81

0.012345679012345679012345679012345679012345679012345679012345679012345679012345
67901234567901234567901234567901234567901234567901234567901234567901234567901234
56790123456790123456790123456790123456790123456790123456790123456790123456790123
45679012345679012345679012345679012345679012345679012345679012345679012345679012
34567901234567901234567901234567901234567901234567901234567901234567901234567901
23456790123456790123456790123456790123456790123456790123456790123456790123456790
12345679012345679012346

0 1 2 3 4 5 6 7 8 9
56 56 56 56 56 55 55 55 0 55

sum of squared differences = 2780


13

1/81

In fact for 1/81 this total is around 2.5 times the largest value ob-
tained when looking at random uniform digits (at least for 10 sets)

sum of squared differences = 2780

512 348 568 470 372 412 314 164 276 130

Is ten samples enough?

We only compared to 10 samples of random numbers. Maybe we


should look at a few more.

Number of sets Maximum squared difference


10 568
100 1024
1000 1388
10000 1912

So for 1/81 the total squared difference is still larger than largest
from 10000 sets of random digits. But for π, its well under.

This suggests that the digits of π is consistent with the uniform


random digits, whereas 1/81 is not.

Summary

• Pi is an irrational number — its decimal expansion doesn’t termi-


nate
• The expected digit count for 500 uniformly random digits is 50 of
each
• Measuring the distance of the first 500 digits of pi from this ex-
pected counts gives a number that is not unusual for randomly
generated digits
• Doing the same for 1/81 gives a very unusual distance
• 1/81 is not consistent with random digits, pi (probably) is.
• We haven’t considered pairs of digits, or triples etc.
• We have only looked at the first 500 digits of pi.
14

Postscript

• In this lecture we looked at one aspect of whether the digits of pi


are random: Uniformity of the distribution of single digits.
• We computed the distance of digit frequencies from an expected
set of frequencies under uniformity.
• And compared that distance to distances similarly achieved using
known random digits.
• If the distance is comparable (it was) there is no evidence that the
digits of pi are not uniformly distributed.

This is rudimentary Hypothesis Testing which you may have seen


before, and we will be covering in depth in this Unit.

You might also like