Predicting High/Low User of Social Networking Siotes among Students
Social networking is the grouping of individuals into specific groups like small rural communities
or a neighbourhood subdivision. Although social networking is possible in person, especially in
the workplace, universities, and schools, it is most popular online. This is a because the internet
is filled with millions of individuals who are looking to meet other people, to gather and share
first-hand information and experiences about any number of topics-from golfing, gardening,
developing friendships to professional alliances.
When it comes to online social networking, websites are commonly used. These websites are
known as social networking sites they function like online communities of internet users.
Depending on the website in question, many of these online community members share common
interests in hobbies, religion, or politics. Once you are granted access to a social networking
website you can begin to socialize. This socialization may include reading the profile pages of
other members and possibly even contacting them.
Contrary to the widely held assumption that people fake themselves on social networking sites,
a new study has claimed that netizens use their profiles to communicate real personalities,
instead of an idealized virtual identity.
According to scientists at the University of Texas , Austin, online social networking profiles like
on Face book convey rather accurate images of the profile owners, either because people aren’t
trying to look good or because they are trying and failing to pull it off.
‘I was surprised by the finding because the widely held assumption is that people are using their
profiles to promote an enhanced impression of themselves.’ Said lead author Sam Gosling of the
research over 700 million people worldwide who have online profiles.
He said, ‘These findings suggest that online social networks are not so much about providing
positive spin for the profile owners but are instead just another medium for engaging in genuine
social interactions, much like the telephone’.
A brief survey of literature on social networking sites reveals that there has been an upsurge of
interest in the study of this relatively new domain in the past few years. Academic researchers
have started studying the use of social networking sites, with questions ranging from their role
in identify construction and expression(Bod and Heer,2006) to the building and maintenance of
social capital(Fllison, Steinfeld, and Lampe,2007) and concerns about privacy. Majority of these
studies generally use Facebook as the subject of study, reflecting the popularity and huge user
base of Facebook.
Hargittai(2007) ,conducted a study to look at the predictors of social networking sites usage
among a diverse group of mainly 18-and 19 –year –old college students studying in the University
of Illinois, Chicago. He found that a person’s gender, race and ethnicity, and parental educational
background are all associated with use, but in most cases only when the aggregate concept of
social networking sites is disaggregated by service. Additionally, people with more experience
and autonomy of use are more likely to be users of such of such sites.
Ellison, stenfield and Lampe (2007) stated that ‘our findings demonstrate a robust connection
between Face book usage and indicators of social capital, especially of the bridging type. Internet
use alone did not predict social capital accumulation, but intensive use of Facebook did.’ Stressing
the role of social networking sites in the formation of social capital, the study shows a strong
linkage between Facebook use and high school connections, and that social networking sites may
also facilitates connections when students graduate from college, with alumni keeping their
school email address and using Facebook to stay in touch with the college community. Such
connections could have strong payoffs in terms of jobs, internships, and other opportunities.
A study was conducted to identify the variables which distinguish between heavy/light users of
social networking sites among students. A questionnaire was designed for the purpose. The social
networking sites considered for the study were Facebook, Linked-In, Twitter, etc. the online
survey was conducted on a sample of 61 students in the age group of 20 to 30. The following
questions were asked of the respondent:
1. How much time do you spend daily on networking sites during weekdays (Monday to
Frida)? X1
(a) Less than 1 hour [1]
(b) 1 to less than 3 hours [2]
(c) 3 to less than 5 hours [3]
(d) More than 5 hours [4]
2. How much time do you spend on networking sites during weekends (Saturday and
Sunday)? (X2)
(a) Less than 2 hour [1]
(b) 2 to less than 4 hours [2]
(c) 4 to less than 6 hours [3]
(d) More than 6 hours [4]
3. Rate the uses of social networking on a scale of 1 to 5 ( 1 being least useful and 5 being
extremely useful) with respect to the following parameters:
(a) To link with professionals (X3A)
(b) Messaging/chatting (X3B)
(c) Networking with friends/relatives (X3C)
(d) To make new friends (X3D)
(e) To promote events/information (X3E)
(f) Blogging (X3F)
(g) News updates(X3G)
(h) Games (X3H)
(i) Educational (X3I)
(j) Photo-sharing (X3J)
(k) Job seeking (X3K)
(l) Online dating (X3L)
The data for the study is herewith attached in Excel sheet.
Questions:
1. Divide the sample into two groups- one that is using the social networking site for less
than 1 hour on weekdays (low users) and the second which is using the social networking
site for one or more hours (high users). Run a two-group Logistic regression analysis with
high/low user as a dependent variable and the variable X3A to X3L as independent
variables to:
(a) Estimate the Logistic regression model
(b) Compute the percentage of respondents that it is able to classify correctly
(c) Determine the statistical significance of the Logistic regression
(d) Test the goodness of fit of the model
(e) Identify which of the predictor variables are relatively better in classifying between
the two groups
Answer:
I do not have SPSS, so I have used EXCEL to solve the question. First of all, one needs to install
Real statistics resource pack as AddIn to Excel which has Logistic Regression solver.
a) Compute the percentage of respondents that it is able to classify correctly.
The result of the classification table inserted below shows us numbers of classifications classified
correctly and incorrectly.
Classification Table
Suc-Obs Fail-Obs SUM
Suc-Pred 19 11 30
Fail-Pred 10 21 31
SUM 29 32 61
The table has values categorized as follows:
Classification Table
Suc-Obs Fail-Obs
Suc-Pred TP FP
Fail-Pred FN TN
True Positives (TP) = the number of cases which were correctly classified to be positive, i.e. were
predicted to be a success and were actually observed to be a success
False Positives (FP) = the number of cases which were incorrectly classified as positive, i.e. were
predicted to be a success but were actually observed to be a failure
True Negatives (TN) = the number of cases which were correctly classified to be negative, i.e. were
predicted to be a failure and were actually observed to be a failure
False Negatives (FN) = the number of cases which were incorrectly classified as negative, i.e. were
predicted to be a negative but were actually observed to be a success.
Percentage of respondents that it is able to classify correctly = (TP+TN/Total)*100 = ((19+21)/61)*100
= 65.37%
b. Determine the statistical significance of the logistic function
LL0 -42.2082
LL1 -35.9327
Chi-Sq 12.55094
df 12
p-value 0.402506
alpha 0.05
sig no
From the analysis, we get the significance test values as given in the table. The coefficient of the
logistic regressions a = -42.2 and b = -35.9
The logistic regression model is given as
c. Identify which of the predictor variables are relatively better in classifying between the two groups.
coeff b s.e. Wald p-value exp(b) lower upper
Intercept -0.14963 1.820771 0.006753 0.934505 0.861029
X3A -0.0146 0.281978 0.002682 0.958695 0.985502 0.567071 1.712685
X3B 0.559262 0.498285 1.259722 0.261704 1.749381 0.658784 4.645424
X3C -0.1649 0.503224 0.107385 0.743141 0.847974 0.316255 2.27367
X3D -0.29475 0.296051 0.991247 0.319438 0.744716 0.416861 1.330423
X3E 0.146831 0.392836 0.139706 0.708574 1.158159 0.536272 2.501215
X3F 0.604464 0.446656 1.831451 0.175956 1.830272 0.762643 4.39248
X3G -0.36108 0.42638 0.717155 0.397079 0.696923 0.302169 1.607387
X3H 0.461482 0.338309 1.86073 0.172541 1.586424 0.817429 3.078848
X3I -0.07459 0.300851 0.061477 0.804177 0.92812 0.514659 1.673743
X3J -0.59604 0.401764 2.200962 0.137925 0.550987 0.250703 1.210944
X3K 0.221605 0.339166 0.426908 0.51351 1.248078 0.642012 2.426275
X3L -0.31477 0.27987 1.264982 0.26071 0.729954 0.421765 1.263342
None of the predictor is significantly above to differntitae the high and low users.(p value in the table
aobe)