0% found this document useful (0 votes)
21 views26 pages

False Negative in Screening Tests

The document discusses the importance of screening tests in identifying individuals at risk for diseases, emphasizing the concepts of sensitivity, specificity, and predictive accuracy. It outlines various study designs, including case control studies, cohort studies, and randomized control trials (RCTs), detailing their methodologies, advantages, and disadvantages. Additionally, it highlights the significance of minimizing bias through randomization, allocation concealment, and blinding in clinical trials.

Uploaded by

drakrana0905
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
21 views26 pages

False Negative in Screening Tests

The document discusses the importance of screening tests in identifying individuals at risk for diseases, emphasizing the concepts of sensitivity, specificity, and predictive accuracy. It outlines various study designs, including case control studies, cohort studies, and randomized control trials (RCTs), detailing their methodologies, advantages, and disadvantages. Additionally, it highlights the significance of minimizing bias through randomization, allocation concealment, and blinding in clinical trials.

Uploaded by

drakrana0905
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

BASIC

STATS FOR
SS DNB
SENSITIVITY VS SPECIFICITY
Screening tests are those tests done among apparently well people to identify those at an
increased risk of a disease or disorder. Those patients who are identified by a screening test are
then offered a subsequent diagnostic test or procedure, or in some instances, a treatment or
preventative medication.
A reliable screening test can improve health, but inappropriate screening harms healthy
individuals and squanders resources

Validity of a test means to what extent a screening test accurately measures what it is intended
to measure. It is the ability of a test to correctly differentiate those who have the disease from
those who don’t. Sensitivity, specificity and the predictive accuracy are the inherent properties
of a screening test. These terms are best explained by expressing the data is terms of a 2 x 2
table.

Screening test result by diagnosis

Disease
Test Total
Present Absent

Positive a b a+b
(TP) (FP)

Negative c d c+d
(FN) (TN)

Total a+c b+d N


a = True Positive, b = False positive, c = False negative, d = True negative
a + c = Total number of diseased patients
b + d = Total number of patients without the disease
a + b = Total number of patients with positive test results
c + d = Total number of patients with negative test results

Measure Formula
Sensitivity a / (a + c) x 100

Specificity d / (b + d) x 100
Predictive value of a positive test a / (a + b) x 100
Predictive value of a negative test d / (c + d) x 100
False positive b / (b + d) x 100
False negative c / (a + c) x 100
Sensitivity:
Sensitivity is the ability of a test to correctly identify all those who have the disease, that is
‘true positives’. It refers to the detection rate of a screening test.
It is a statistical index of diagnostic accuracy.
Eg: 90% sensitivity means that 90% of the diseased people who are screened with the test will
give a ‘true positive’ result, the rest 10% with the disease will give a ‘false negative’ result.

Specificity:
Specificity is the ability of a test to correctly identify those who do not have the disease, that is
‘true negatives’.
Eg: 90% specificity means that 90% of the non-diseased people will give a ‘true negative’
result, and 10% will give a false positive result, that is 10% people without the disease will be
wrongly classified as having the diseases
Sensitivity and specificity are inversely related, which means sensitivity can be increased at
the expense of specificity (or vice versa).

Predictive accuracy:
Predictive value reflects the diagnostic power of a test; it depends on both sensitivity and
specificity, and the disease prevalence.
Predictive value of a positive test indicates the probability that a patient with a positive test,
actually has the disease. Higher the prevalence of a disease, higher is the accuracy or the
predictive value of a positive test.
False negative means patients who actually have a disease test negative for the same. Higher
the sensitivity, lower is the false negative. It is like giving a false reassurance to a patient. The
patient may tend to ignore any signs and symptoms of the disease, as he may feel he has tested
negative to the screening test. This could be detrimental if the disease is a serious one.
False positive means patients who do not actually have disease, test positive on screening, and
are told that they have the disease. Higher the specificity, lower is the rate of patients testing
false positive. This may result in subjecting normal people to further diagnostic evaluation and
the associated anxiety & unnecessary expenses. False positives reduce the credibility of
screening programs, as the patients may feel they are being unnecessarily being subjected to
tests when they finally test negative with the diagnostic tests.
STUDY DESIGNS
Observational or Interventional
Observational – include case report, case series, case control study, cohort study
Interventional – Randomised control studies, non-randomised control studies

Case Control Study


 Observational/ Analytical study
 Generally retrospective = “backward looking” study design
 Effect to Cause
 Begins with the effect/ disease condition, and the investigator goes back in time to try
and identify any factor associated with the condition
 A similar group of patients with matched characteristics apart from the disease/effect
in question, is taken as the control = CONTROL GROUP
 Control group should be closely matched with the Case group

EXPOSURE PRESENT
DISEASE PRESENT

EXPOSURE ABSENT

STUDY BEGINS HERE

EXPOSURE PRESENT
DISEASE ABSENT

EXPOSURE ABSENT

Go back in time to see how many patients presently with the disease, have been exposed to the
factor
Once the values of the above are obtained, Odd’s ratio is obtained

2 x 2 contingency table in a case control study


Suspected risk factor Cases Controls
“Exposure” (Disease present) (Disease absent)
Present a b
Absent c d
a+c b +d
 Odd’s Ratio/ Cross product ratio = ad/bc
 Measures the strength of the association between risk factor and outcome
STEPS: Selection of cases & controls – Matching – Measurement of exposure – Analysis
& Interpretation
1. Selection of Case: “Case” is defined based on certain diagnostic criteria.
Source of Case – hospital or the community (general population)
Selection of Control: Must be free from the disease under study, but in other aspects,
as similar as possible to the cases. Eg, age, gender distribution, lab tests, habits, etc.
Source of Control - hospital or the community (general population)

2. Matching: Done to ensure comparability between cases and controls, and to try and
reduce the effect of ‘confounding factor’, a factor that can be associated with both the
cases and controls, and can independently affect the disease

3. Measurement of Exposure: Done as defined in the protocol, by interviews,


questionnaires, reviewing medical records

4. Analysis & Interpretation: exposure rate among cases and controls studied, Odd’s ratio
calculated
Odd’s Ratio: ad/bc

Advantages:
 Simple/easy/less expensive
 Less time consuming
 Useful in diseases with long latent period
 No drop out of cases/ attrition
 No loss to follow up
 Useful in research on RARE diseases
 No patient risk
 Minimal ethical problems
Disadvantages:
 Recall bias, selection bias, observation bias can be present
 Temporality of association cannot be established. (Temporality = cause precedes effect)
 Natural history of disease cannot be established
 Incidence of disease cannot be calculated
 Relative risk, Attributable risk cannot be calculated
 Selection of appropriate control group may be difficult
 Cannot distinguish between main cause and other associated factors
 Cannot evaluate prophylactic treatment of a disease
Cohort Study
 Observational/Analytical study
 Obtain evidence to support or refute existence of association between a cause and
disease
 Usually Prospective study/ Longitudinal study/ “Forward looking study”
 Follow up at different time points, while actively looking for the outcome
 Cause to effect
 Eg.: Role of exposure to dyes in development of Bladder cancer

DISEASE PRESENT

EXPOSURE PRESENT
DISEASE ABSENT

DISEASE PRESENT

EXPOSURE ABSENT

DISEASE ABSENT

Forward-looking study, with follow-ups at various time points


Points to remember:
 Cohorts identified before disease appearance
 Study groups observed over specified period of time, to determine disease incidence
Framework of a cohort study:
COHORT Disease present Disease absent Total

Exposure Present a b a+b


Exposure Absent c d c+d

a+c b +d
Incidence of disease among exposed = a/ a + b
Incidence of disease among non-exposed =c/ c + d
STEPS: Selection of study subjects – Obtaining data on exposure – Selection of
Comparison Group – Follow up - Analysis & Interpretation
1. Selection of study subjects: Can be general population if the exposure is frequent, or
can be special groups (professional, certain habits)
2. Obtaining data on exposure: directly from subjects, medical records, periodic medical
examination, surveys
3. Selection of Comparison Group: can be internal or external comparison
4. Follow up: medical check-up, reviewing medical and hospital records, death
surveillance
5. Analysis & Interpretation: Incidence rate among exposed and non-exposed calculated,
Estimation of Risk done
Estimation of Risk – Relative Risk
RR = Incidence of disease among exposed (a/a+b)
Incidence of disease among non-exposed (c/c+d)
Advantages:
 No recall bias
 Temporality of association can be established, as the study follows cause to effect
design
 Natural history of disease can be studied
 One exposure leading to multiple outcomes can be studied
 Incidence can be calculated (incidence = number of new cases)
 Relative risk, attributable risk can be calculated
Disadvantages:
 More time consuming
 More expensive/ resources
 Usually involves larger number of people
 Patient loss to follow can occur/ drop out/ attrition
 Rare diseases cannot be studied
Variants of Cohort Study
 “Nested case control” study, which is a type of cohort study, begins like a typical case
control study, but then follows up the patients over a period of time
 Retrospective / Historic cohort = involves studying hospital records, hence possible
only in situations where perfect patient records are maintained.
Eg.: 1) Diethyl stilbestrol was identified as a causative agent of vaginal adenocarcinoma
in young women, whose mothers had received the drug during pregnancy
2) High dose Oxygen as a cause of Retrolental fibroplasia in premature infants with
NICU admissions
Randomised control trial (RCT)
 Experimental, best study design
 Cross over of patients can be done – further increases the quality of the study
 Target population should be identified before conducting the study – refers to the
population on whom the results of the study will be generalised
 The study should include a part of this target population itself
 Ethics committee approval before the study - mandatory
 Written informed consent, in a language which patient can understand, taken before the
study – mandatory
Blinding:
 Blinding done to remove bias.
 Can be single blinded (patient does not know what group he/she is in), double blinded
(both patient and investigator does not know what group each patient is in), triple
blinded (patient, investigator, statistician unaware of the group distribution)
Control Group:
 These patients are patients with the same characteristics as the Test group, but they do
not receive the test medication
 Placebo controlled trials – placebo medication given to the patient. Being considered
unethical, as the patient may be denied treatment
 Current practice – the accepted standard of care presently available, given to patients
in Control group.
 Placebo given only in conditions where no other treatment options are available
Randomisation
 Process to ensure complete removal of any forms of bias
 Computer generated algorithms used to ensure fair distribution of patients across both
groups
Allocation Concealment
 Allocation concealment is the technique of ensuring that implementation of the
random allocation sequence occurs without knowledge of which patient will receive
which treatment, as knowledge of the next assignment could influence whether a
patient is included or excluded based on perceived prognosis.

CROSS OVER TRIAL- Improving study RCT design: Once both the groups receive the
drug, and once the washout period of the drugs is over, the groups can be interchanged.
Therefore, each patient becomes his own control, confounding factors are totally removed,
and both groups become totally comparable
Cross over not possible –
 If the drug cures the disease being studied
 If comparison being made between medical and surgical modalities of treatment
 Disproportionate drop outs occur in either of the groups
 In certain psychiatric drugs/ conditions

Study parts:
Protocol developed

Ethics committee approval

Patient recruitment – Written informed consent

Patient refuses to give IC – not included Patient agrees to give IC

Included in the study

RANDOMISED

Test Group (receives test drug) Control Group (receives control)


(Blinding can be done to remove bias)

Effective Not Effective Effective Not Effective

STATISTICAL ANALYSIS

Results
Results can be summarised as follows:

INTERVENTION GIVEN Disease cured/Drug Disease not cured/Drug not


effective effective

Test Group/New drug a b


Control Group/ Placebo/ c d
Current standard of care

a+c b +d

Advantages of RCT:
 Causal inference can be made, strongest evidence among all study designs
 Randomisation ensures removal of bias, minimisation of effect of confounding factors
 Study can be tailored to answer any specific question
Disadvantages:
 Expensive
 If not blinded, investigator bias may be present
 Loss to follow up may occur
 If a new drug is being tested, and is seen to be toxic, study may have to be
discontinued
METHODS TO ELIMINATE BIAS IN CLINICAL TRIALS:
 Randomization
 Allocation Concealment
 Blinding (Masking)

Bias:
Bias may be defined as a systematic error, or “difference between the true value and that
actually obtained due to all causes other than sampling variability.”
Bills “kills” the scientific value of the study

A. Randomization and allocation concealment


B. Actual assignment that can be followed by masking/blinding subjects as to their
assigned group
C. Prospective evaluation period
D. Outcome evaluation during which outcome assessors can be masked as to the subjects’
assigned group
RANDOMIZATION
The process by which each subject has the same chance of being assigned to either intervention
or control.
It begins with the sequence generation process, and extends until the subjects are assigned to
their groups. What follows sequence generation is allocation concealment, so that until the
subjects are assigned to their respective groups, this assignment remains a secret, and no
external influence must be made to alter the original assignment.
The successful implementation of randomization is dependent on allocation concealment.
Importance of randomisation:
 Reduces many types of Bias: Investigator related or patient related
 Adds Validity to Statistical Tests: differences between intervention and control groups
should behave like differences between two random samples from the population
 Minimizes Confounding: Randomization tends to produce groups that are similar in
terms of both known and unknown prognostic factors.
(Confounding factor- something that can independently affect the outcome in a trial,
unrelated to the intervention being given)
 Ethical aspects: Randomisation ensures a fair distribution of patients, and gives each
patient an equal chance of being either in the control or test group.
 Produces groups similar with respect to known or unknown risk and confounding
Types of Randomisation:
 Simple randomisation: Toss a coin, Random digit table
Easy to implement, but unequal distribution of subjects may occur
 Random permuted blocks: to equalize the number of subjects on each treatment
 Stratified Permuted Block randomisation: stratum is first defined, following which
block randomisation is performed within each stratum
 Play the winner design
 Two armed bandit design
 Adaptive minimisation
ALLOCATION CONCEALMENT IN CLINICAL TRIAL
Allocation concealment refers to preventing the next assignment in the clinical trial from being
known. It ensured correct implementation of the sequence generated by randomisation.
In an RCT, a system must be in place to ensure that subjects and investigators, do not know to
which group a subject will be allocated before that subject is entered into the study.
The use of Sequentially Numbered, Opaque, Sealed Envelopes (SNOSE) is an economical
and straightforward means of assuring allocation concealment
Consider the following scenarios:
Scenario 1: Consider the investigator knows the group to which the next patient will be
assigned; that investigator may try to influence the assignment, resulting in the randomisation
being broken, and the trial ends up being non-randomised.
Scenario 2: Consider the subject learns of the assignment into the placebo group; the subject
may refuse to take the medication or may even drop out of the trial.
Both the above scenarios result in loss of randomisation and introduction of bias.
Inadequate allocation concealment may result in deciphering of the randomisation code.
Examples of Methods of Deciphering Allocation Concealment:
 Holding translucent envelopes up to bright lights to reveal upcoming assignment (even
using the hot light in a radiology department for more opaque envelopes)
 Opening unsealed assignment envelopes, or well-sealed, opaque envelope in advance
of consent
 Opening unnumbered envelopes until desired allocation found
 Determining different weights of the assignment envelopes (eg, the heavier envelope
means intervention group)
 Asking a central randomization center for the next several assignments all at once
 Deciphering assignments to active drug or placebo based on appearance of drug
container labels

Allocation concealment vs Blinding (Masking)


Masking only refers to the prevention of knowledge of assigned groups and the allocation
sequence after actual allocation. The goal of masking is to prevent ascertainment bias.
Allocation concealment refers to the prevention of knowledge of upcoming assignment from
the time of generation of the randomized sequence up until actual allocation.
BLINDING

Blinding refers to the concealment of group allocation from one or more individuals involved
in a clinical research study, most commonly a randomized controlled trial (RCT). More the
number of parties ‘blinded’ in a clinical trial, the better it is.
Blinding in a surgical trial is more difficult, tricky and at times unethical, when compared to
medicines, which can easily be blinded with a placebo medication.
Blinding is completely different from Allocation Concealment- this is done to eliminate
selection bias during the process of recruitment and randomization

Why Blind?
 In order to reduce performance and ascertainment bias after randomization
 The purpose of blinding is to prevent differential treatment of the groups later in the
trial or the differential assessment of outcomes, which may result in biased estimates
of treatment effects.

Levels of blinding:
Single blinded trial:
Participant/subject does not know whether he is getting the test drug or control drug.
If participants are not blinded, knowledge of group assignment may affect their behaviour in
the trial and their responses to subjective outcome measures. This can be minimised by blinding
the participants.
Double blinded trial:
Both the participant and the investigator is unaware of the group assignment.
Avoids the clinicians showing differential attitude towards patients receiving either the test or
the control medication.
Triple blinded trial:
Participant, Investigator and the Data collectors/Statisticians are unaware of the group
assignment.
Ensures unbiased data collection and analysis, as there may be a subconscious effort to see a
positive result with a new drug. Blinding would reduce this.
When is blinding not possible?
 Surgical interventions being compared to medical management- ethical blinding may
not be possible, eg concealment of incisions/scars. Sham surgeries are done in animal
experiments, but may not be ethically possible in human participants.
 Data collectors may not be possible to be blinded in certain situations, eg research
involving highly contagious or infectious disease
If blinding is not possible
• Standardize the treatment of the groups (apart from the intervention)
• Consider an expertise-based trial design
• Use objective, reliable outcomes if possible
• Consider duplicate assessment
• Acknowledge the limitations

To conclude, Blinding is an important methodologic feature of RCTs to minimize bias and


maximize the validity of the results. Researchers should strive to blind as many parties
involved in a trial, as possible.
Basic Biostats concepts
HYPOTHESIS TESTING
Null hypothesis
When we begin the study, we begin with the assumption that there is no significant difference
between the two study groups. This is the null hypothesis (h0).
So if the study concludes that there is a difference between two groups, it means we have
‘disproved’ the null [Link] level of significance is set before the study begins.
A p value of less than 0.05 is considered statistically significant. It means the chance of type I
error having occurred is less than 5%. That is the study wrongly accepting a difference between
the two groups (thereby disproving the null hypo) is less than 5%.
Alternate hypothesis:
Begin the study with the assumption that there exists a difference between the two groups.

Based on the study result:


Null Hypothesis Based on study result, NH can be:
ACCEPTED REJECTED
TRUE (no Correct decision False positive error
difference) Type I
5%
FALSE False negative error Correct decision
(difference present) Type II
20%

Type I error:
When a study concludes that there is a difference between the analysed groups when, in fact,
there is no difference (False positive error).
Here we have rejected a null hypothesis which in reality is true.
By convention, type I is more dangerous than type II error.
Max permissible in any study is 5%.
(Therefore, Confidence interval should be more than 95%.)

Type II error (Beta error):


When a study concludes that there is no difference between the two groups, when in fact, there
is a difference (False negative error); i.e. the study has failed to demonstrate this difference.
We have accepted a null hypothesis, which in reality is false.
Max permissible is 20%.
Power of a study:
Power = 1 – beta error
Hence, minimum valid power is 80% (because beta error max is 20%)
“Capability of a test to detect the difference between two groups, if such a difference actually
exists”
Ability to reject a null hypothesis which in reality is false.

p Value
Used for probability of committing type I error/ false positive error.
Always expressed in decimals, never %
Since max type I error permissible is 5%, p value max is 0.05
It is not related to the power of the study, type II error/beta error
p value less than 0.05% implies that the probability of committing type I error is < 5%.

Significance of statistical “significance”:


*A statistically significant result need not always mean clinically significant. Ex. A drug may
reduce the systolic BP by 10 mm Hg, which may be statistically significant. But clinically this
small drop may not be significant.
*A statistically non-significant result, may not mean that the drug under investigation is not
effective. It may be because that particular study failed prove its significance (i.e fail to
disprove the null hypothesis)
Sample size:
To prove a statistical significance, the number of samples collected or the number of people
included in a study is very important.
If the number is too small, a rare disease may not be picked up. If the sample is too large, the
result may get diluted. Hence the optimal number should be used. This is referred to as sample
size.
It refers to the number of samples/patients included in the study, so that the result obtained is
as close as possible to the true positive.
DATA ANALYSIS
Data can broadly be divided into either “Quantitative” (that which can be measured
numerically), or “Qualitative” (numerically cannot be expressed)

Data

Qualitative (Non-numerical) Quantitative (Numerical)

Nominal Ordinal Discrete Continuous

Interval Ratio

Qualitative: cannot be quantified in numerical terms


*Nominal – cannot be expressed in any order, can be either one of the options only. Eg male
or female, alcoholic or non-alcoholic, yes or no
*Ordinal- can be expressed in an order, eg mild, moderate, severe
Quantitative: parameters such as BP, biochemical levels, Hb levels
*Discrete- number of males or females, number of patients with each type of blood group,
number of patients who underwent surgery
*Continuous- height, weight, temperature, PSA level
*Interval- temperature, zero-point is arbitrary
*Ratio- continuous, ordered and constant scale. BP, weight, height, age
STATISTICAL TESTS

Mathematical calculation or analysis of data. The result obtained refers to the statistical power
of the test.
The statistical power is defined as a measure of the probability that a statistical test rejects a
false null hypothesis. It is the probability of finding a significant result, if there is any difference
between the two treatment groups.
Higher the power of a statistical, more likely one is to find statistical significance, if the null
hypothesis is false.
*Type I error (alpha error): When a study concludes that there is a difference between the
analysed groups when, in fact, there is no difference (False positive)
*Type II error (beta error): When a study concludes that there is no difference between the two
groups, when in fact, there is a difference (False negative)
Max type I error allowed – 5%, type II is 20%
As we have different types of data, there are different types of stat tests.

Biostatistical tests

Parametric tests Non-parametric tests

ANOVA T Test Chi square test, Fisher’s exact test


Mann Whitney U test
Unpaired t test Paired t test Wilcoxon signed rank test
Friedman test, Spearman rank order test
Parametric Tests
Used for quantitative data, when data follows a ‘normal’ distribution.
Used to test the Null hypothesis that there is no difference between the ‘means’ of two samples.
More stat power than non-parametric tests
Criteria:
 Data should be numerical scale
 Distribution should be normal. If data is skewed, it is converted into log form
 Variance should be same
 Sample should be randomly drawn from the population
 Observations within a group are independent

STUDENT’S T TEST
Done for comparison of one or two means
Finally, the test concludes whether the null hypothesis is accepted or rejected.
Done when the sample size is less than 30.
It is usually applicable to measurement data (graded data), such as blood sugar level, body
weight, reaction time – “quantitative data”
‘t’ distribution curve is similar to that of the normal curve
Two independent variables are required for unpaired t test. Eg Control group and treatment
group
Single continuous variable for paired t test. Eg. Pre and post treatment for the same group of
patients

Unpaired t test Paired t test


Two independent groups are compared Sample group is compared, ex. Before and
after intervention
Intragroup variability is present No intragroup variability
More sample size required to increase the Comparative lesser sample size is required to
power of the study achieve the same power of the study
ANALYSIS OF VARIANCE (ANOVA)
Extension of t test, whereas t test can compare only two means, ANOVA can compare more
than two means, across various groups.
When there are >2 groups, a series of t tests would be needed to compare them, which can add
on to the experimental error (Type I error). Hence ANOVA is applied.
If two means are compared using ANOVA instead of t tests, we would get the same results.
Hence, t test can just be considered a type of ANOVA.
ANOVA can be one-way or multifactor.
One way ANOVA is used to compare the means of two or more levels of a single independent
variable. Multifactor ANOVA is used when there are two or more independent variables.
Another version is “MANOVA” (multiple analysis of variance). Used when two or more
dependent variables that are generally related in some way.

CHI-SQUARE TEST
Compares categorical responses between two or more groups
Used to examine differences between two groups consisting of categorical variables, ex.
Gender and marital status
Data should be nominal or ordinal, as this is a test of proportions
It summarises discrepancy between observed and expected frequencies. Smaller the
discrepancy between the observed and the expected scores, smaller is the value of the
chisquare.

Lung cancer Smoking profile


present absent
Yes
No
2 x 2 table is first made before applying the test.
If the value in any of the boxes is less than 5, then Fischer exact test is applied.
Odd’s ratio and relative risk or risk ratio are also calculated from the above table.
Formula:
CORRELATION
Relationship between two quantitative variables on a scatter plot.
Two variables can have no correlation, positive/ direct relationship, or indirect/ negative
relationship.
Positive- more the PSA, higher the chance of prostate cancer
Negative- more the HDL, lesser the chance of cardiovascular disease
Correlation coefficient gives the strength and direction of association

REGRESSION
Linear regression:
Quantitative data
Method of estimating or predicting a value on some dependent variable, given the values of
one or more independent variables. Like correlations, stat regression tells the association or
relationship between variables. The primary purpose of regression is prediction.
Ex. Health is predicted by body weight, medical history, disease status, marital status
Two types of regression- simple, multiple
Simple regression- attempt to predict the dependent variable with the help of a single
independent variable. Ex. Smoking and lung cancer
Multiple regression- many independent variables are used to predict one dependent variable.
Ex. Multiple factors to predict the health
Logistic regression:
Mix of quantitative and qualitative data
To predict dichotomous variables, such as the presence or absence of a specific outcome, based
on a specific set of independent or predictor variables.
Provides information about the strength and direction of the association between the variables.
Also, it can be used to estimate the Odd’s ratio for each independent variables.
This Odds ratio can tell us how likely a dichotomous outcome is to occur, given a particular
set of independent variables.
Application- to determine whether and to what degree a set of hypothesized risk fators might
predict the onset of a certain condition.
KAPLAN MEIER SURVIVAL ANALYSIS
Kaplan-Meier estimate is one of the best options to be used to measure the fraction of subjects
living for a certain amount of time after treatment.
In clinical trials, the effect of an intervention is assessed by measuring the number of subjects
survived or saved after that intervention over a period of time.
The time starting from a defined point to the occurrence of a given event, for example death is
called as survival time and the analysis of group data as survival analysis.
The Kaplan-Meier estimate is the simplest way of computing the survival over time in spite of
difficulties associated with subjects or situations like loss to follow up or drop out
(“censoring”).
Analysis technique:
The Kaplan-Meier survival curve is defined as the probability of surviving in a given length of
time while considering time in many small intervals.
Involves computing of probabilities of occurrence of event at a certain point of time, and then
creating a survival curve using these probabilities.
This can be calculated for two groups of subjects who have received two different modalities
of treatment. Accordingly, the statistical difference in the survivals between the groups can be
compared.
Analysis is based on three assumptions:
1. Firstly, we assume that at any time patients who are censored have the same survival
prospects as those who continue to be followed.
2. Secondly, we assume that the survival probabilities are the same for subjects recruited
early and late in the study.
3. Thirdly, we assume that the event happens at the time specified.
Based on these assumptions, the survival probability at any given point of time can be
calculated using the following formula:

The subjects who have died, dropped out or whose data is not available for analysis (i.e subjects
who are censored) are not a part of the denominator.
Benefits & Limitations:
Benefits- easy useful method of calculating survival analysis
Limitations: Based on certain assumptions, which may not always be accurate. Also, cannot
measure the effect of other co-variates which may independently affect the survival.
Kaplan Meier Plot:

In this curve, Time is marked on the X axis, Survival is marked on the Y axis.
Each step indicates (number of) the cumulative survival of patients at that given point of time.
If the steps are progressively lowered, it indicates that many subjects are being “censored”
(dead or drop outs).
Two or more curves corresponding to different treatment methods can be compared using this
analysis, as seen below:

The log-rank test can be used to


test whether the difference between
survival times between two groups
is statistically significant
Extra
Measures of Data/Measures of Central Tendency
Mean, Median, Mode
Mean: arithmetic average of any given data
Mean = ∑ x
N
Median: middle value of any given data, when data is arranged either in ascending or
descending order. If two central values are there, then the arithmetic mean of these two values
is taken.
Mode: The most frequently occurring value in any data set

Measures of Dispersion
Range, Variance, Standard deviation (SD), standard error of mean (SEM), Coefficient of
variation, difference between SD and SEM
Range: Highest value – Lowest value
Gives the value between the maximum and the minimum values in any data set
Ex: 2, 4, 6, 5, 3, 6, 8, 18, 14, 16.
Range = 18 – 2 = 16
Variance: dispersion of a group of data around their mean value
σ = ∑ (X - µ)2
n
where, X = individual score, mu = mean of total score, n = number of observations
Standard deviation: measure of dispersion or variability of a data from its mean value.
Calculated by the square root of variance.

Standard error of mean: measure of variability


Normal distribution curve: GAUSSIAN curve

In the above curve:


Mean ± 1SD ----- 68.2% of distribution
Mean ± 2SD -----95.4% of distribution
Mean ± 3SD -----99.7% of distribution

You might also like