0% found this document useful (0 votes)
10 views81 pages

GEA1000 Tutorial Overview and Sampling Methods

Uploaded by

浓糖NaNa
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
10 views81 pages

GEA1000 Tutorial Overview and Sampling Methods

Uploaded by

浓糖NaNa
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

GEA1000

TUTORIAL 1
About myself
Afiq Daniel
Teaching Assistant
adaniel@[Link]

Note:
At the end of every even week, I will send the
slides used in tutorial to all of you through e-mail
(not Canvas).
Recap the most recent chapter
(30-45mins)
•It is a recap, it will be fast-paced

Break (10mins)
How will
Go through tutorial materials in
tutorials groups (60mins)

be like? •Discuss actively for tutorial participation

Break (10mins)

In-Class Quiz (ICQ) (15mins)


Population vs Sample
Population: Sample:
•The ENTIRE group that you want to draw •The SPECIFIC group that you will collect the
conclusions about data from, proportion of the population
•May not be people, it could be countries,
species, events, objects, etc.

Parameter: Estimate:
•Numerical fact about population, constant •Inference about population’s parameter,
based on sample
A researcher surveys 32 people on their smoking
habits. He wants to know whether people in
Singapore are willing to stop smoking entirely. The
researcher finds that 41% of them are willing to
Population stop entirely.

vs Sample
What is the population?
What is the sample?
What is the 41% referred to as?
Questions that we have about
some characteristic of a population

Research Types:
Questions
•Making an estimate about the population
•Testing a claim about the population
•Compare two sub-populations /
Investigate a relationship between two
variables in the population
What proportion of students in this class
passed GEA1000?

Type of Did more than 75% of students in this class


passed GEA1000?
Research Is there a higher proportion of males than
Questions females in this class?
Is there a higher proportion of males than
females in this class who passed GEA1000?
Explore the data

Create summaries
Exploratory
Data Evaluate and analyse
Analysis
Tweak our question or data or
methods
Rinse and repeat until something
useful is found
Source from where we
gather our sample such as
a list
Sampling
Frame
We need our sampling
frame to cover our
population of interest if we
want to generalise fully
A census is an attempt to
reach out to every unit from
the population
A sample is selecting only a
Census proportion

Pros and cons?


Bias
Selection: Non-response:
•Individuals left out •Individuals not responding
•Biased towards those included •Out of fear, inconvenience or
simply unwilling
•Non-probability sampling in
selection
Bias (some others – not tested)
Social desirability/conformity: Extreme responding:

•Individuals feel the need to conform to societal •In scale questions such Likert scale, certain
norms cultures more prone to responding more
extremely
•Might not fill truthful answers
•Similarly, for people with lower IQ
Acquiescence/agreement:
•Survey results can be quite drastic
•Individuals tend to select positive responses
more frequently Question order/order-effects:

•Could be due to politeness or just getting tired •Individuals may want to give internally
of giving thoughtful answers in a long survey consistent answers
Sampling process via
a known randomized
mechanism
Probability
Sampling
Eliminates selection
biases using the
element of chance
Types of Probability Sampling
•Simple random sampling (without replacement)
• Units are randomly selected from sampling frame
• Every set of units has equal chance
• Variability is due to chance
• Pros: good representation of population
• Cons: limited flexibility

•Systematic sampling
• Selecting units using a selection interval, K, so that every Kth unit from the random starting point in the
first interval is selected
• Pros: easy to execute
• Cons: not suitable if list is not random
Types of Probability Sampling
•Stratified sampling
• Sampling frame broken down into strata
• A stratum contains similar characteristics
• Apply simple random sampling to each stratum
• Pros: good representation for each stratum
• Cons: complicated and time-consuming,
sometimes hard to define strata

•Cluster sampling
• Population broken down into clusters
• Sample a fixed number of clusters and sample all units in those clusters
• Pros: less tedious, less costly, less time-consuming
• Cons: clusters need to be reasonably heterogeneous (all clusters share similar characteristics)
01 02
Non-
Sampling process Sampling by
Probability not random human discretion
Sampling
Types of Non-Probability Sampling

Convenience sampling Volunteer sampling


Choosing subjects that are most convenient Subjects volunteer to participate
Pros: convenient Pros: convenient
Cons: both selection and non-response bias Cons: might be slow, both selection and non-
response bias
Decide on a sampling frame

Decide what kind of sampling


General method is feasible
Process of
Sampling
Sample

Removal of unwanted units can


happen at several different steps
Sampling frame should be
larger or equal to population
Large sample sizes reduce
Some ways we variability, hence reduces error
can help raise
generalisability Probability sampling minimizes
selection bias
Minimise non-response bias,
or high response rate
An attribute that can be
measured or labelled

Variable
Data sets contain
individuals and such
variables pertaining to
those individuals
Independent vs Dependent
Independent: Dependent:
•Subject to manipulation in a study •Hypothesised to change depending
on how independent variable
•Can be deliberate or spontaneous
changes
Independent vs Dependent
What are the independent and dependent variables in the research questions
below?

How does the amount of sleep impact test scores?


What is the effect of caffeine on sleep?
Categorical vs Numerical
Categorical: Numerical:
•Take category or label values •Take numerical values
•Each observation can only hold •Arithmetic operations make sense
one label on them
•Labels must be mutually exclusive
Categorical: Ordinal vs Nominal
Ordinal: Nominal:
•Take natural ordering •No intrinsic ordering
•Usually shown with numbers
Numerical: Discrete vs Continuous
Discrete: Continuous:
•Possible values form a set of •Can take on all possible numerical
numbers with gaps values in a given range or interval
•Not necessary to be whole
numbers
What type of variables?
100m sprint race positions, from 1 to 8
Types of pets owned, where 1 represents cats, 2 represent dogs, and 3
represents rabbits
Weight of a person
Weight of a person measured by a digital scale of up to 1 decimal place
Summary Statistics
Measures of central tendencies: Measures of dispersion:
•Mean •Standard deviation
•Median •Interquartile range
•Mode
•Average value for a numerical variable
𝑥𝑥1 +𝑥𝑥2 +⋯+𝑥𝑥𝑛𝑛 ∑𝑛𝑛
𝑖𝑖=1 𝑥𝑥𝑖𝑖
•𝑥𝑥̅ = =
𝑛𝑛 𝑛𝑛

•𝑛𝑛 is the number of data points, while 𝑥𝑥𝑖𝑖


refers to the value of the 𝑥𝑥 of the 𝑖𝑖-th data
Mean point
•Adding a constant 𝑐𝑐 to all data points will
change mean by 𝑐𝑐
•Multiplying by constant 𝑐𝑐 to all data points
will multiply mean by 𝑐𝑐
•“Spread” of data about the mean
𝑥𝑥1 −𝑥𝑥̅ 2 + 𝑥𝑥2 −𝑥𝑥̅ 2 +⋯+ 𝑥𝑥𝑛𝑛 −𝑥𝑥̅ 2
•𝑆𝑆𝑆𝑆𝑆𝑆𝑆𝑆𝑆𝑆𝑆𝑆 𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉 =
𝑛𝑛−1
and 𝑠𝑠𝑥𝑥 = 𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉𝑉
Standard •𝑠𝑠𝑥𝑥 is always non-negative
Deviation •Adding a constant 𝑐𝑐 to all data points will not
change 𝑠𝑠𝑥𝑥
•Multiplying by constant 𝑐𝑐 to all data points will
multiply 𝑠𝑠𝑥𝑥 by 𝑐𝑐
Let’s look at these 2 groups of income:
Is standard
deviation
easy to Group 1: $5, $6, $7, $8, $9, $10
interpret
when Group 2: $995, $996, $997, $998,
comparing $999, $1000
between Both have 𝑠𝑠𝑥𝑥 = 1.871 but when
groups? comparing between the groups, it
seems that 1.871 has a higher impact
for Group 1
•A way to quantify the degree of spread
Coefficient of relative to the mean
Variation •𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶𝐶 𝑜𝑜𝑜𝑜 𝑣𝑣𝑣𝑣𝑣𝑣𝑣𝑣𝑣𝑣𝑣𝑣𝑣𝑣𝑣𝑣𝑣𝑣 =
𝑠𝑠𝑥𝑥
𝑚𝑚𝑚𝑚𝑚𝑚𝑚𝑚 𝑜𝑜𝑜𝑜 𝑥𝑥
Let’s look at these 2 groups of income:

Group 1: $5, $6, $7, $8, $9, $10

Using Group 2: $995, $996, $997, $998, $999, $1000


coefficient of
Both have 𝑠𝑠𝑥𝑥 = 1.871 but when comparing between
variation the groups, it seems that 1.871 has a higher impact
for Group 1

Coefficient of variation is 0.249 and 0.00188 for the


groups respectively, showing us that there is a higher
degree of spread relative to the mean for Group 1
•Middle value for a sorted numerical variable
•For an odd number of data points, the
middle value will be the median
•For an even number of data points, the

Median average of the middle 2 values will be the


median
•Adding a constant 𝑐𝑐 to all data points will
change median by 𝑐𝑐
•Multiplying by constant 𝑐𝑐 to all data points
will multiply median by 𝑐𝑐
Median does not consider the precise
value of each observation and hence
does not use all information available in
the data.

Why mean Unlike mean, median is not amenable to


further mathematical calculation and
over hence is not used in many statistical
median? tests.
If we pool the observations of two
groups, median of the pooled group
cannot be expressed in terms of the
individual medians of the pooled groups.
Mean is not
robust
Why median
over mean?
Median is robust
to outliers
•First quartile, or 𝑄𝑄1 , is the 25th percentile of the data-
values
•Third quartile, or 𝑄𝑄3 , is the 75th percentile of the data-
values
•The 𝑖𝑖-th percentile of the data-values means that 𝑖𝑖𝑖
of the data is equal or less than this value
Interquartile •The interquartile range is the difference between the

Range
two
•𝐼𝐼𝐼𝐼𝐼𝐼 = 𝑄𝑄3 − 𝑄𝑄1
•𝐼𝐼𝐼𝐼𝐼𝐼 is always non-negative
•Adding a constant 𝑐𝑐 to all data points will not change
𝐼𝐼𝐼𝐼𝐼𝐼
•Multiplying by constant 𝑐𝑐 to all data points will multiply
𝐼𝐼𝐼𝐼𝐼𝐼 by 𝑐𝑐
Mean and standard deviation tend to be paired up

Median and IQR tend to be paired up, preferred


when data is not symmetrical

Using Summary statistics are useful but loads of


Summary information are left behind precisely because it is
only a form of summary.
Statistics
Can use different statistics together to give you a
fuller picture.

Best if you can have the summaries together with


appropriate data visualisation.
•Most frequent value
•Unlike mean and median, mode can be used
Mode for both numerical and categorical variables
•The “peak” of a distribution
To answer a research question
that we have, we would
conduct a study
Types of
Study
Designs
Studies can be designed using:

Experimental Observational
We intentionally manipulate one
variable in an attempt to cause an
effect on another

Our goal is to provide evidence for


Experimental a cause-and-effect relationship
between two variables

Independent Variable →
Dependent Variable
In our experiments, we can split our
subjects into two groups, treatment and
control
Treatment group: receives the treatment

Treatment vs
Control Control group: does not receive the
treatment we are targeting/testing

Control group provides a baseline for


comparison with the treatment group
To establish such a relationship, we
need to make sure that the
independent variable is the only factor
that is affecting the dependent variable
We want to remove all other “hidden”
Cause-and- effects
Effect

We can do so using random


assignment
An impartial procedure that uses
chance to select subjects for the
treatment and control groups

If the subject count is large enough,


Random the laws of probability state that the
Assignment two groups will tend to be similar in all
aspects
As such, we have resolved our concern
of any “hidden” effects from other
variables
Random assignment solves “hidden” variables but
what about bias?

Subjects who know that they are from specific


group may behave a certain way, creating bias in
our experiment
We can use a placebo: a treatment that has no
Placebo effect

Another problem comes in, the placebo effect

This refers to the response observed when subjects


received a placebo treatment but still showed
some positive effects because they thought they
received the actual treatment
Blinding is a way to make sure subjects do not
know which group they are in

Effective blinding uses a placebo that is very


similar to the treatment

The subjects are “blind”, so personal ideas or


Blinding beliefs prevent them from affecting the results

Sometimes, some experiments require us to


blind the assessors as well to prevent bias

Blinding both parties is known as double


blinding
Experiments can be useful,
but we need to ensure that
they are ethical
Ethical
Issues
If we cannot find ethical
means for our experimental
study designs, we can
consider observational ones
We observe individuals and
measure variables of interest

We do not make attempts to


Observational directly manipulate one variable to
cause an effect on another

We do not provide evidence for a


cause-and-effect relationship
between two variables
In our observations, we can split our subjects
into two groups, treatment and control
although we did not do any actual treatment

We can view it as exposure instead, whether


the subject was exposed to something or not
Treatment vs
Control Treatment group: exposure

Control group: non-exposure


Experimental vs Observational
Experimental: Observational:
•Assigned by researcher •Assigned by subjects
•Can provide evidence of cause- •Cannot provide evidence of cause-
and-effect and-effect
•Can provide evidence of
association
Tutorial
Section
Sampling frame: Movies listed on the website

Sampling method: Simple random sampling since every sample of size N has the same chance of
being chosen
Categorical: Genre, Release_Year, MPAA_Rating

Numerical: Production_Budget, Worldwide_Gross, Duration, CPI, Voter_Numbers

Title is neither but belongs to another type known as “identifier” or “ID” variables. Names used to
identify objects or people are typical examples of such variables.

IMDb_Rating is another variable that is difficult to classify, why? While it is a rating from 1 to 10, an
average is calculated using an unrevealed algorithm. We should not take the average of categorical
variables, but some still do it and treat the result as either numerical or categorical. If treated as
categorical, it might be better to group the scores into broader categories and not let it be as finely
divided as it is now since the number of labels is too large to say anything meaningful about it.

Note that such a contentious discussion will not be used for any form of assessment and what is
written here is for students to merely start thinking about such questions.
Release_Year
CPI
Duration
MPAA_Rating
IMDb_Rating
Voter_Numbers
Genre
To count missing values, you can use “=COUNTBLANK”. You cannot use “=COUNTBLANK(A:A)” as it
will count the entire column, even past the last row, giving you missing values. You must indicate the
cutoff point, “=COUNTBLANK(A2:A1092)”.
You can also turn your data into a Table by selecting all the cells with data and then Insert > Table.

You can then select a column to see its “Count” at the bottom. Missing values can be equated as
1092 − Count. For example, Release_Year has 1092 − 1073 = 19
Release_Year, CPI, Genre

We can try to fill up information where possible to make the data set more complete.

For example, a Google search would help us identify “The Three Stooges” as a “Comedy”, or “Dude,
Where’s My Dog?” was released in 2014, allowing us to update the Release_Year and CPI.

If it is not possible to find the information, and we happen to be working with these variables, we
may have to ignore these data points instead.

We should not resort to dealing to deleting/ignoring data as our first choice of dealing with dirty
data if it is not too tedious to try and fix them.
Duration

It is tough and may be considered impractical to fill up the 112 missing values manually.

In such scenarios, if you are well versed in coding or certain software, you can run a web scraper to
scrape information on these movies from websites such as IMDb and fill in the missing data
partially.

If there are no efficient way of filling in the missing information, we may have no choice but to ignore
these rows when working with the variable. We need to bear in mind that the lack of information on
the duration for 112 movies may result in a form of bias when working on this variable.
The minimum of 0 is suspicious as it is tough to imagine a movie with no revenue.

The description claims that it only considers the revenue generated from
screening in public theatres. Other sources such as movie rentals in the past or
streaming platforms in the present are not included. If we are only interested in
movies that had a public theatrical release, we can justifiably ignore these movies
that generated $0.
For Excel, it is best to install the Data Analysis Toolpak
add-in. For some of you, you can install it here:

File > Options > Add-ins > Go… (at Manage: Excel Add-ins)
> Tick Analysis Toolpak > Ok

Or, Tools > Excel Add-ins > Tick Analysis Toolpak > Ok

After doing that, “Data Analysis” will now appear under the
Data tab.

Go to Data > Data Analysis > Descriptive Statistics > Ok.


This gives us everything except for Q1 and Q3, which we can derive using =QUARTILE(C:C, 1) and
=QUARTILE(C:C, 3) respectively.

If you are not using the Data Analysis Toolpak, here are the
remaining functions needed:

Mean: =AVERAGE(C:C)
Median: =MEDIAN(C:C)
SD: =STDEV.S(C:C)
[Do not use STDEV.P as we are working with a sample]
Minimum: =MIN(C:C)
Maximum: =MAX(C:C)
IQR: =QUARTILE(C:C, 3)-QUARTILE(C:C, 1)
In Excel, create a new column K with header “Adjusted_Production_Budget”. The formula to use for
cell K2 would be =294.4/E2*B2, and then double click the bottom right of the cell to fill the entire
column (you should see a black plus sign before double clicking).
Do this and do
remember to click
the green “Store”
button.
As mentioned in the Appendix, do change Release_Year to categorical or “factor” before using. Then, you will need
to filter out data from 2012 to 2022. Here are some ways for Radiant:

1) Create a new data set in Excel by filtering there, and then upload it to Radiant.

2) In “View”, just under the column name “Release_Year”, you can adjust the values to only show 2012 to 2022.
After doing that, you can store filtered data as Movies_filtered and use that data set instead.

3) You can immediately “filter data” and key in (Release_Year %in% c("2012", "2013", "2014", "2015", "2016",
"2017", "2018", "2019", "2020", "2021", "2022")) or (Release_Year %in% [Link](2012:2022)).

4) Instead of “factor”, you transform Release_Year to “character”. After that, select “filter data” just below your
data set selection. Input: (Release_Year>=2012 & Release_Year<=2022)

5) You can duplicate Release_Year as Release_Year_v2 and convert that one to numerical. After that, select
“filter data” just below your data set selection. Input: (Release_Year_v2>=2012 & Release_Year_v2<=2022)

6) Be creative (such as “remove levels” under transform, but be careful if you save it under the same variable
name as you will make all the other years blank permanently).
The general trend is the same
for both.

However, we do observe the


differences between the
adjusted production budget
of 2022 to other years to be
significantly lower than the
differences for production
budget without adjustment.
For Excel, time to introduce PivotTable. You can do so by Insert > PivotTable. Play around with the
PivotTable to get familiar with what is being shown. For this question, what we need is shown below:

Your PivotTable should look something like this:


Filter to only include 2012 to 2022:

With this PivotTable, we can create a line graph


by going to Insert > PivotChart > Line > Ok:
You can tweak the PivotTable to include both lines:
Any findings obtained from this data would be difficult to generalize to the movie industry.

It is good that probability sampling is used, and sample size is not a big source of concern.

However, sampling frame of around 6000 movies covers the target population of millions of movies
only to a very small extent. We are also unsure of the selection process of these 6000 movies and
whether it was random.

There is also missing data within the chosen 1091 movies, questioning the accuracy of results.

It is also worth noting that the website highlights that getting accurate information about movie
budgets is not easy despite trying their best. Therefore, the accuracy of data cannot be taken for
granted as well.
To investigate if the new instructional/assessment system is better at improving the mathematic
proficiency of students aged 3-5 years old as compared to current methods that do not involve this
new system.
Although the researchers may have done a good study demonstrating the effectiveness of the new
instructional/assessment system, it is generally not a good idea to simply implement the new
instructional/assessment regime to our chosen target population since the conditions in which the
study was conducted can be very different from the conditions here.

For example, the 3-5 years old children who participated in the original study were predominantly
from low-income families and which may not be the case for the 3-5 years old children in the school
that we are working with.

So, we cannot simply assume that the new instructional/assessment system has the same level of
effectiveness amongst all kindergarteners aged 3-5 years old. We should exercise prudence and still
conduct a pilot study to determine its effectiveness for the profile of students that we are working
with.

However, we can certainly take the research results as an encouraging sign and try to replicate the
study while making adjustments/tweaks as and when necessary, based on practical constraints.
An experimental study of treatment and control groups would ideally be suitable.
Method of Assessment

For students coming from both the experimental and control groups, we can administer a post-study
test similar to what the researchers have done. For all students in the same level, we can administer
the same test.

Possibility of random assignment

It would be possible (using RNG or any appropriate mechanism) to assign students into control
(without the new system) and treatment (with the new system). This random assignment will ensure
that the characteristics of students in both groups are similar (socioeconomic status, age
composition, sex ratio, etc.), provided that sample size (final pool of students after obtaining consent)
is large enough.
Possibility of Blinding

If is almost impossible to blind the subjects since students will be aware of their learning mode.
Also, discussions among friends outside of class will make it impossible for them to be kept unaware
of the difference in modes provided. Furthermore, when obtaining consent, parents will be aware of
the existence of the two groups and may know which group their child is in.

The teachers who are working with the students also cannot be blinded since they know which group
they are with and what kind of methods they are using.

On the other hand, similar to the paper, we can still blind the assessors who are administering the
test at the end of the experiment such that they will not know which exam scripts were done by
students who used the new learning mode, and which were not.
Limitations

Students and parents are aware of the different modes of learning given to different students.

Getting parental consent would be tough. Although the original study managed to do a randomised
assignment, we cannot rule out parents may consent on the condition that they decide which group
to enroll their child in, which would make random assignment impossible. They may also choose to
intervene if they feel midway that their child is not performing well enough which introduces bias into
the study.

For the parents that do not consent, we need to decide if they should just be part of the control
group (making random assignment impossible) or should they be excluded completely. If too many
parents do not consent to let their children participate, there may simply not be enough students to
conduct any meaningful experiment.

Ethical side of things. Suppose the new method helps, students who were not given it would not be
given a change to perform better. Suppose the new method has a negative impact, it may not be fair
to the affected students. Also, we may not be able to have a long-term study as it will implicate the
student’s long term educational development.

You might also like