0% found this document useful (0 votes)
9 views10 pages

Bootstrap Logistic Regression Analysis

The document describes using bootstrapping to analyze a dataset with over 1.7 million rows. It specifies stratified bootstrapping with 1000 samples, a 95% confidence level, and percentile confidence intervals. It then performs logistic regression on the dataset to identify predictor variables for a target variable, entering 13 independent variables using a significance level of 0.05 for inclusion.

Uploaded by

Preetham Karthik
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
9 views10 pages

Bootstrap Logistic Regression Analysis

The document describes using bootstrapping to analyze a dataset with over 1.7 million rows. It specifies stratified bootstrapping with 1000 samples, a 95% confidence level, and percentile confidence intervals. It then performs logistic regression on the dataset to identify predictor variables for a target variable, entering 13 independent variables using a significance level of 0.05 for inclusion.

Uploaded by

Preetham Karthik
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

GET

FILE='C:¥Users¥Arun Krishnan¥Downloads¥[Link]'.
DATASET NAME DataSet1 WINDOW=FRONT.
BOOTSTRAP
/SAMPLING METHOD=STRATIFIED(STRATA=V32 )
/VARIABLES TARGET=V32 INPUT=V2MALE1FEMALE0 V5 V7MO1MC2SC3 V8 V9 V10 V13
V14 V15 V18 V19 V22 V31
/CRITERIA CILEVEL=95 CITYPE=PERCENTILE NSAMPLES=1000
/MISSING USERMISSING=EXCLUDE.

Bootstrap

Notes

Output Created 15-NOV-2019 18:52:57


Comments
Input Data C:¥Users¥Arun
Krishnan¥Downloads¥[Link]
Active Dataset DataSet1
Filter <none>
Weight <none>
Split File <none>
Syntax BOOTSTRAP
/SAMPLING
METHOD=STRATIFIED(STRATA=V32
)
/VARIABLES TARGET=V32
INPUT=V2MALE1FEMALE0 V5
V7MO1MC2SC3 V8 V9 V10 V13 V14
V15 V18 V19 V22 V31
/CRITERIA CILEVEL=95
CITYPE=PERCENTILE
NSAMPLES=1000
/MISSING
USERMISSING=EXCLUDE.
Resources Processor Time 00:00:00.02

Elapsed Time 00:00:00.02

[DataSet1] C:¥Users¥Arun Krishnan¥Downloads¥[Link]


Bootstrap Specifications

Sampling Method Stratified


Number of Samples 1000
Confidence Interval Level 95.0%
Confidence Interval Type Percentile
Strata Variables Target variable ( 1: Responders / 0:
Non-Responders)

LOGISTIC REGRESSION VARIABLES V32


/METHOD=ENTER V2MALE1FEMALE0 V5 V7MO1MC2SC3 V8 V9 V10 V13 V14 V15 V18 V19
V22 V31
/CONTRAST (V2MALE1FEMALE0)=Indicator
/CONTRAST (V7MO1MC2SC3)=Indicator
/CLASSPLOT
/PRINT=GOODFIT
/CRITERIA=PIN(0.05) POUT(0.10) ITERATE(20) CUT(0.3).

Logistic Regression

Notes

Output Created 15-NOV-2019 18:52:57


Comments
Input Data C:¥Users¥Arun
Krishnan¥Downloads¥[Link]
Active Dataset DataSet1
Filter <none>
Weight <none>
Split File <none>
N of Rows in Working Data
179125847
File
Missing Value Handling Definition of Missing User-defined missing values are
treated as missing
Syntax LOGISTIC REGRESSION VARIABLES
V32
/METHOD=ENTER
V2MALE1FEMALE0 V5
V7MO1MC2SC3 V8 V9 V10 V13 V14
V15 V18 V19 V22 V31
/CONTRAST
(V2MALE1FEMALE0)=Indicator
/CONTRAST
(V7MO1MC2SC3)=Indicator
/CLASSPLOT
/PRINT=GOODFIT
/CRITERIA=PIN(0.05) POUT(0.10)
ITERATE(20) CUT(0.3).
Resources Processor Time 01:08:00.30

Elapsed Time 01:08:00.57

Case Processing Summary

Unweighted Casesa N Percent

Selected Cases Included in Analysis 282924 100.0

Missing Cases 0 .0

Total 282924 100.0


Unselected Cases 0 .0
Total 282924 100.0

a. If weight is in effect, see classification table for the total number of


cases.

Dependent Variable Encoding

Original Value Internal Value

0 0
1 1

Categorical Variables Codings

Parameter coding

Frequency (1) (2)


Product code of Two 1 81053 1.000 .000
wheeler (MC : Motorcycle , 2 92819 .000 1.000
MO : Moped, SC : Scooter) 3 109052 .000 .000
Gender 0 42078 1.000

1 240846 .000

Block 0: Beginning Block

Classification Tablea,b

Predicted

Target variable ( 1: Responders /


0: Non-Responders) Percentage
Observed 0 1 Correct

Step 0 Target variable ( 1: 0 262356 0 100.0


Responders / 0: Non- 1
20568 0 .0
Responders)

Overall Percentage 92.7

a. Constant is included in the model.


b. The cut value is .300

Variables in the Equation

B S.E. Wald df Sig. Exp(B)

Step 0a Constant -2.546 .007 123628.454 1 .000 .078

a. Variable(s) entered on step 1: V2MALE1FEMALE0, V5, V7MO1MC2SC3, V8, V9, V10, V13, V14,
V15, V18, V19, V22, V31.

Bootstrap for Variables in the Equation

Bootstrapa

95% Confidence Interval

B Bias Std. Error Sig. (2-tailed) Lower Upper

Step 0 Constant -2.546 .000 .000 .001 -2.546 -2.546

a. Unless otherwise noted, bootstrap results are based on 1000 stratified bootstrap samples
Variables not in the Equationa

Score df Sig.

Step 0 Variables V2MALE1FEMALE0(1) 524.113 1 .000

V5 184.812 1 .000

V7MO1MC2SC3 406.655 2 .000

V7MO1MC2SC3(1) 6.273 1 .012

V7MO1MC2SC3(2) 362.914 1 .000

V8 .623 1 .430

V9 34.879 1 .000

V10 12.551 1 .000

V13 40.036 1 .000

V14 524.411 1 .000

V15 68.528 1 .000

V18 332.079 1 .000

V19 642.087 1 .000

V22 52.066 1 .000

V31 4525.928 1 .000

a. Residual Chi-Squares are not computed because of redundancies.

Block 1: Method = Enter

Omnibus Tests of Model Coefficients

Chi-square df Sig.

Step 1 Step 7225.047 14 .000

Block 7225.047 14 .000

Model 7225.047 14 .000

Model Summary

Cox & Snell R Nagelkerke R


Step -2 Log likelihood Square Square

1 140213.721a .025 .062


a. Estimation terminated at iteration number 6 because
parameter estimates changed by less than .001 for split file
$bootstrap_split = 0.

Hosmer and Lemeshow Test

Step Chi-square df Sig.

1 102.956 8 .000

Contingency Table for Hosmer and Lemeshow Test

Target variable ( 1: Responders / Target variable ( 1: Responders /


0: Non-Responders) = 0 0: Non-Responders) = 1

Observed Expected Observed Expected Total

Step 1 1 27644 27502.308 648 789.692 28292

2 27217 27147.284 1075 1144.716 28292

3 27058 26936.984 1234 1355.016 28292

4 26822 26771.574 1470 1520.426 28292

5 26691 26616.551 1601 1675.449 28292

6 26450 26449.553 1842 1842.447 28292

7 26246 26253.128 2046 2038.872 28292

8 25783 25986.522 2509 2305.478 28292

9 25194 25493.231 3098 2798.769 28292

10 23251 23198.865 5045 5097.135 28296

Classification Tablea

Predicted

Target variable ( 1: Responders /


0: Non-Responders) Percentage
Observed 0 1 Correct

Step 1 Target variable ( 1: 0 260589 1767 99.3


Responders / 0: Non- 1
19770 798 3.9
Responders)

Overall Percentage 92.4

a. The cut value is .300


Variables in the Equation

B S.E. Wald df Sig.

Step 1a V2MALE1FEMALE0(
-.459 .027 298.540 1 .000
1)

V5 -.118 .004 940.342 1 .000

V7MO1MC2SC3 481.059 2 .000

V7MO1MC2SC3(1) -.157 .028 31.914 1 .000

V7MO1MC2SC3(2) .338 .019 303.459 1 .000

V8 -.083 .008 118.999 1 .000

V9 .000 .000 331.814 1 .000

V10 .000 .000 15.410 1 .000

V13 -.021 .002 82.427 1 .000

V14 .140 .004 1288.515 1 .000


V15 .028 .007 14.856 1 .000

V18 .022 .004 28.461 1 .000

V19 .026 .003 71.653 1 .000

V22 -.365 .012 933.880 1 .000

V31 .190 .003 3320.072 1 .000

Constant -2.244 .054 1700.885 1 .000

Variables in the Equation

Exp(B)

Step 1a V2MALE1FEMALE0(1) .632

V5 .889

V7MO1MC2SC3

V7MO1MC2SC3(1) .854

V7MO1MC2SC3(2) 1.402

V8 .920
V9 1.000

V10 1.000

V13 .979

V14 1.150

V15 1.028

V18 1.022

V19 1.026

V22 .694
V31 1.210

Constant .106
a. Variable(s) entered on step 1: V2MALE1FEMALE0, V5, V7MO1MC2SC3, V8, V9, V10, V13, V14, V15, V18,
V19, V22, V31.

Bootstrap for Variables in the Equation

Bootstrapa

95%
Confidenc
e Interval

B Bias Std. Error Sig. (2-tailed) Lower

Step 1 V2MALE1FEMALE0(
-.459 .001 .026 .001 -.511
1)

V5 -.118 .000 .004 .001 -.124

V7MO1MC2SC3(1) -.157 .000 .027 .001 -.209

V7MO1MC2SC3(2) .338 .001 .020 .001 .300

V8 -.083 .000 .008 .001 -.098

V9 .000 .000 .000 .001 .000

V10 .000 .000 .000 .001 .000

V13 -.021 .000 .002 .001 -.025

V14 .140 .000 .004 .001 .132

V15 .028 -.001 .007 .001 .012

V18 .022 .000 .004 .001 .014

V19 .026 .000 .003 .001 .020

V22 -.365 .000 .012 .001 -.390

V31 .190 .000 .003 .001 .184

Constant -2.244 -.001 .051 .001 -2.348

Bootstrap for Variables in the Equation

Bootstrap

95% Confidence Interval

Upper

Step 1 V2MALE1FEMALE0(1) -.406

V5 -.111

V7MO1MC2SC3(1) -.106

V7MO1MC2SC3(2) .378

V8 -.067

V9 .000
V10 .000
V13 -.016

V14 .147

V15 .041

V18 .029

V19 .031

V22 -.341

V31 .197

Constant -2.145

a. Unless otherwise noted, bootstrap results are based on 1000 stratified bootstrap samples

Step number: 1

Observed Groups and Predicted Probabilities

80000 +
+
I
I
I
I
F I
I
R 60000 +
+
E I
I
Q I 0
I
U I 01
I
E 40000 + 000
+
N I 0001
I
C I 0000
I
Y I 00000
I
20000 + 000000
+
I 0000001
I
I 000000000
I
I 00000000000001
I
Predicted ---------+---------+---------+---------+---------+---------+-----
----+---------+---------+----------
Prob:
0 .1 .2 .3 .4 .5 .6 .7
.8 .9 1
Group:
000000000000000000000000000000111111111111111111111111111111111111111111111
1111111111111111111111111

Predicted Probability is of Membership for 1


The Cut Value is .30
Symbols: 0 - 0
1 - 1
Each Symbol Represents 5000 Cases.

You might also like