0% found this document useful (0 votes)
5 views10 pages

Understanding Bootstrapping in Statistics

The document discusses various data analysis techniques, including linear regression and KNN (K-Nearest Neighbors), highlighting their strengths and weaknesses. It emphasizes the importance of understanding model accuracy, overfitting, and underfitting in predictive modeling. Additionally, it touches on concepts like sampling distributions and the significance of sample size in statistical estimates.

Uploaded by

hejjdh
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views10 pages

Understanding Bootstrapping in Statistics

The document discusses various data analysis techniques, including linear regression and KNN (K-Nearest Neighbors), highlighting their strengths and weaknesses. It emphasizes the importance of understanding model accuracy, overfitting, and underfitting in predictive modeling. Additionally, it touches on concepts like sampling distributions and the significance of sample size in statistical estimates.

Uploaded by

hejjdh
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

品slice map.

de
等[Link] sunmavi2epivet longerfacet [Link]
5ᵗʰmdnuafǎueilCdata
Nmayan

gǎnsáoin C
[Link]

In
[Link] id
Summarise Pivot


linearregressionmodel
knn 器简 i
mǎni spec linear_reg ws3 spe se
tsǎimidiiüǜiau

器 Rn
regression
classifiknn [Link]器货
䲜䲜 z cation z stencntncaaaii step
[Link]

[Link]
workflowC 1
管䲜器管品 regression1 nǎaünn
lit datas
pullc
workflow ieeipeih
ikllowistoretoobicanpute vlold
[Link]
[Link] is
ngispeccnujhruipecx
cucm_aimat.in data11s data train
all
add C
set_engine Steps spec
data predict111
v strata inode step at 1 ftctest datalbind

col metric ctuuthi estiinte. [Link]


Kun RMSPE 3 34 39 40
s
classification 35 36

0 Auurany 41
通识
寻找边养
KNN on becreated
dinner 器䜌望蒜 颜
和 ㄨㄨ 2 致
histogram nineǎǐin multipleohserves
banchant的
凝敷 nenohsc [Link]

Problems accuracy
performanceonly
_overall
真实 classili Sǚipeccǎni

cation
reua 䲜 ag 1 器
在 新则
workflow

difference
between
of between diffwarstandardized
me

Disadvantage
Narm to
classili an
standardizedmovemakesen
using Regression asdistances
undersame
remove missing scalenotaffected的
data [Link] Ueig
alex
longerso
ness my h
missing bar
notberandem Regte N'S9
cgpyzyunbalancechoosethe
sehe when km
onewithlangerament
impnte
mean oflabel eventhepatter
isdifferent
替 NA dis每个obs不同1 predictin affeited
布 P 1
42 20处 kunAccuracy could behigh
as the
function使 在 199.9 butmeaningless
allone
summary are
surroundings

y butcannot lind
之后因为
regression workflow
workflow 仅为estimate正确k uheregzis
使 summary重祈 跑 准
算data的rmse来找
it step_upscalecl
确 Ka 出 y summary

出最 rmse邻居然后使 7

validation111系时 通讼 了
cross 出 test data
在不 拿 model
前提下 estimate
的accuracy 值
[Link]
年不 寻找mincdilf.nl9
slice
adsyaucg
griduals
[Link] ii 的

RMSPE越 越 accurate
overfitting
_underfitting c 错误的 trend
both bad
line too smooth Cannot follow
under fitted the trend well
Over fitted toorough 点到 点 lose generali
有 minin 法展示niaaoma

代码
[Link] [Link]
dataI test.data1
17 17 17
17
in metrics
Summary Collectmetrics
set_mode C [Link]
Vlold [Link] name
loutput国表 data
pusE
D 17
[Link] 器

vloldivgloldnumbe
s1 lor [Link] regression certain
KNN
cdǎ[Link]'rec workflowHit y
1 neighbors tunec Strata
lüiinien 1 出
iiwctunea cinlit add
[Link]
addeiipeworkllonlit
tǎnā model 17 Cresamples
howmany 17

vlold kgoma_vlji.lt
resamples 17
try

collect_metrics 选
结 aauang
knn
Strengths
通 改 2 simple intuitive
容易理解
KNN vs linear
Kan can capture ndata requires few
about
non linear trend assumptions
hon data look
to interpret
workswell for
linear easy exceeded
evenfor range
non linear

prod Cons
nun
mt [Link]
trainingdata larger
maynotperfan
熊 RMPE
with large
number
predictors
we
not predict
vangeof
如呆beyond beyond
input in
[Link] range values data
interpret 会 都变成 个 training
值 直线 1
因为最近 的 只有 个
只有 个 y
neighbor
代码 2

[Link]
se [Link]
engine y
lm 1 L regression
17 P
group_by
Spector ft ly
replicate
cg
17
summarisec hypothesis
总㥪 2 Calculate test extract fit
sngǎismnl
proportion
sgnun lor parsnip
否 须
stpung_distrimionsinlpleslewuei [Link]

ti
taking igtiigc
sample pop
Samples
aescxi in Simulationmean
and Calc pc
samplep P Samplerep.sample_n
labs x
data size reps
y cont data_sample 1
summarize Sample n n
Sun new_type
怠偻
n

popsncnuntnlh
通识 了
What does k meansdo
k which is the numberof cluster
choose
to group observations
datato k clusters ran donn
initialize assign
Iterate each duster
calculateaveragecoordinates of
A
to find Center Closest center to
data
B reassign label
to

update Checkifuo
times
C After many
labelschanged terminate

what is agoodcluster
distance
书 0 25 26
the smalleravg
to Centroid Cws1
fronpointsbetter
the 䲜
Wss Diner case
What's WssD why mid decrease
characteristics ofwssis at of
cluster Eun
within distance alwaysdecrease f _kmieanswight.no
t
ofsquared each langer K thebestsolution
calculated f decrease
get
dusterthen up
add sloweddown
fat rateshouldbe
chosen
solve restart
kmeans
celbowl [Link]
通识 4
whole set population
population
ofinterest
subjects
distribution
parametersi mean
Var quantitative distribution
population property sǎmpe

n 个 absǎion
Sample
estimates Simulation
point
从 random sample _know parameters
中取 数值 estimates cnotrealistic
an unknown
population _calc sample mean
interest times to
parameters of manyClose to
is a random
get parameters
Sample mean
variable as
selectsome obsfirst sampling_distri
but
_randomly
sample mean shape_
calculate ametw bell shaped
baek data select
put men symmetric
sample new _centeredaround
true population

So we needtoknowdistribution proportion

to inferane replicate
distribution
sampling 第 个sample
samplemean
distribution of 避
通识 5
size on

Effect of sample Sampling_distribution


Sample side increase become
distribution
sampling 的 越
newer 抽取 sample range
nar more possible to get a estimate
to

Samplemean

bootstrap 93 52 重听
抽 sample
[Link] as population1
sample aet
simulation
点 population
ibqange
为什么 要 plans is
52 reliable idea
estimate
[Link]
给 how an
不会 gives
Plausible
if
leikegwearetoge [Link]

how
sample
another
代码 了

[Link]
[Link]
replane reps

sample in
taking
bootstrap

You might also like