0% found this document useful (0 votes)
12 views71 pages

Election Prediction Using Deep Learning

This systematic review examines the use of social media data, particularly from platforms like Twitter, to predict election outcomes, highlighting the evolution of research in this area. It identifies significant challenges such as data representativeness and the low success rates of traditional sentiment analysis methods, while suggesting that newer approaches like regression methods yield better results. The paper emphasizes the need for improved methodologies and the integration of advanced machine learning techniques for future research in electoral predictions.

Uploaded by

Siva Kumar
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
12 views71 pages

Election Prediction Using Deep Learning

This systematic review examines the use of social media data, particularly from platforms like Twitter, to predict election outcomes, highlighting the evolution of research in this area. It identifies significant challenges such as data representativeness and the low success rates of traditional sentiment analysis methods, while suggesting that newer approaches like regression methods yield better results. The paper emphasizes the need for improved methodologies and the integration of advanced machine learning techniques for future research in electoral predictions.

Uploaded by

Siva Kumar
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

SYSTEMATIC REVIEW OF PREDICTING ELECTION BASED ON

SOCIAL MEDIA DATA RESEARCH CHALLANGES AND FUTURES


USING DEEP LEARNING

ABSTRACT:

The way politicians communicate with the electorate and run electoral campaigns
was reshaped by the emergence and popularization of contemporary social media
(SM), such as Facebook, Twitter, and Instagram social networks (SNs). Due to the
inherent capabilities of SM, such as the large amount of available data accessed in
real time, a new research subject has emerged, focusing on using the SM data to
predict election outcomes. Despite many studies conducted in the last decade,
results are very controversial and many times challenged. In this context, this
article aims to investigate and summarize how research on predicting elections
based on the SM data has evolved since its beginning, to outline the state of both
the art and the practice, and to identify research opportunities within this field. In
terms of method, we performed a systematic literature review analyzing the
quantity and quality of publications, the electoral context of studies, the main
approaches to and characteristics of the successful studies, as well as their main
strengths and challenges and compared our results with previous reviews. We
identified and analyzed 83 relevant studies, and the challenges were identified in
many areas such as process, sampling, modeling, performance evaluation, and
scientific rigor. Main findings include the low success of the most-used approach,
namely volume and sentiment analysis on Twitter, and the better results with new
approaches, such as regression methods trained with traditional polls. Finally, a
vision of future research on integrating advances in process definitions, modeling,
and evaluation is also discussed, pointing out, among others, the need for better
investigating the application of state-of-the-art machine learning approaches.

CHAPTER-1

INTRODUCTION

1.1 OVERVIEW

Scholars have suggested that the social sciences are experiencing a momentous
shift from data scarcity to data abundance [1,2], noting, however, that full access to
data is likely to be reserved for the most powerful corporate, governmental, and
academic institutions [3]. Indeed, as a growing share of our social interactions
happen on digital platforms, our capacity to predict attitudes and behaviors should
increase as well, given the scale, richness, and temporality of such data. Still, there
is a growing debate in the research community about the usefulness of social media
signals in understanding public opinion. Social scientists have identified the
problems with using “repurposed” data afforded by social media platforms, which
are often incomplete, noisy, unstructured, unrepresentative, and algorithmically
confounded [4] compared to traditional surveys, which are carefully designed,
representative, and structured.
Many recent studies have utilized social media data as a “social sensor”
aiming to predict different economic, social, and political phenomena, ranging
from election results to the box office success of Hollywood movies—with some
achieving considerable success in this endeavor. Still, most of these studies have
been primarily data-driven and largely atheoretical. Furthermore, instead of
predicting the future, these studies have made predictions about the present [5] or
the past, and have generally not examined the validity of their predictions by
comparing them with more robust types of data, such as surveys and government
censuses. Researchers have also raised concerned about the lack of
representativeness and authenticity of data harvested from social media platforms,
noting the demographic biases, the curated nature of social media representations
[6], biases introduced by sampling methods [7] and pre-processing steps [8],
platform affordances [9], and false positives introduced by bot accounts [10]. Still,
when compared to the more traditional methods, social media analytics promise to
provide an entirely new level of insight into social phenomena, by allowing
researchers unprecedented access to people’s daily conversations and their
networks of friendship and influence, across time and geographic boundaries.
Thus far, only one meta-analysis assessing the predictive power of social
media data has been published, focusing solely on Twitter-based predictions of
elections [11]. This paper, reviewing only nine studies examining the predictive
utility of Twitter messages, concluded that although social media data provides
some insights regarding electoral outcomes, it is unlikely to replace survey-based
predictions in the near future. Given the lack of scientific consensus about the
merits of social media-based predictions and the growing number of studies using
data from diverse social media sources, a more comprehensive meta-review is
needed. We thus aim to compare the merits of the most commonly used
approaches in social media-based predictions and also examine the roles of
contextual variables, including those related to the level of democracy and type of
electoral system.

1.2. Survey vs. Social Media Approaches


Social media-based predictions have challenged the key underlying principles
of traditional, survey based research namely, probability-based sampling and the
structured, solicited-nature of participant data.
First, probability-based sampling, as one of the foundations of scientific
survey research, is based on the assumption that all opinions are worth the same
and should be weighted accordingly [12], thereby representing “a possible
compromise to measure the climate of public opinion” [13] (p. 45). However, this
method does not take into account the differences in the levels of interpersonal
influence that exist in the population and are highly important in the process of
public opinion formation [14,15]. Traditional polling can occasionally reveal
whether the people holding a given opinion can be thought of as constituting a
cohesive group; however, it does not provide much information about the opinions
of leaders who may have an important role in shaping opinions of many other
citizens. Ceron et al. noted that the predictive capacity of social media analysis
does not necessarily rely on user representativeness, suggesting that “if we assume
that the politically active internet users act like opinion-makers who are able to
influence (or to ‘anticipate’) the preference of a wider audience: consequently, it
would be found that the preferences expressed through social media today would
affect (predict) the opinion of the entire population tomorrow” [16] (p. 345). In
other words, simply discounting social media data as being invalid due to its
inability to represent a population misses important dynamics that might make the
data useful for opinion mining.

1.3 Problem Statement

The way politicians communicate with the electorate and run electoral campaigns
was reshaped by the emergence and popularization of contemporary socialmedia
(SM), such as Facebook, Twitter, and Instagram social networks (SNs). Due to the
inherent capabilities ofSM, such as the large amount of available data accessed in
real time a new research subject has emerged focusing on using the SM data to
predict election outcomes. Main findings include the low success of the most-used
approach, namely volume and sentiment analysis on Twitter, and the better results
with new approaches, such as regression methods trained with traditional polls.
Finally, a vision of future research on integrating advances in process definitions,
modeling, and evaluation is also discussed, pointing out, among others, the need
for better investigating the application of state-of the-art machine learning
approaches.
CHAPTER-2

LITERATURE SURVEY

[Link] and Public Opinion Forecasts with Social Media Data: A Meta-
Analysis

Author - Prof. Marko M. Skoric, Jing Liu

Year – 2020

In recent years, many studies have used social media data to make estimates of
electoral outcomes and public opinion. This paper reports the findings from a
meta-analysis examining the predictive power of social media data by focusing on
various sources of data and different methods of prediction; i.e., (1) sentiment
analysis, and (2) analysis of structural features. Our results, based on the data from
74 published studies, show significant variance in the accuracy of predictions,
which were on average behind the established benchmarks in traditional survey
research. In terms of the approaches used, the study shows that machine learning-
based estimates are generally superior to those derived from pre-existing lexica,
and that a combination of structural features and sentiment analyses provides the
most accurate predictions. Furthermore, our study shows some differences in the
predictive power of social media data across different levels of political democracy
and different electoral systems. We also note that since the accuracy of election
and public opinion forecasts varies depending on which statistical estimates are
used, the scientific community should aim to adopt a more standardized approach
to analyzing and reporting social media data-derived predictions in the future.

2. A Systematic Review of Predicting Elections Based on Social Media Data:


Research Challenges and Future Directions
Author- Kellyton dos Santos Brito, Member, IEEE, Rogério Luiz Cardoso Silva
Filho, and Paulo Jorge Leitão Adeodato

Year-2021

The way politicians communicate with the electorate and run electoral campaigns
was reshaped by the emergence and popularization of contemporary social media
(SM), such as Facebook, Twitter, and Instagram social networks (SN). Due to
inherent capabilities of SM, such as the large amount of available data accessed in
real time, a new research subject has emerged, focusing on using SM data to
predict election outcomes. Despite many studies conducted in the last decade,
results are very controversial, and many times challenged. In this context, this
work aims to investigate and summarize how research on predicting elections
based on SM data has evolved since its beginning, to outline the state of both the
art and the practice, and to identify research opportunities within this field. In
terms of method, we performed a systematic literature review analyzing the
quantity and quality of publications, the electoral context of studies, the main
approaches to and characteristics of the successful studies, as well as their main
strengths and challenges, and compared our results with previous reviews. We
identified and analyzed 83 relevant studies, and the challenges were identified in
many areas such as process, sampling, modeling, performance evaluation and
scientific rigor. Main findings include the low success of the most-used approach,
namely volume and sentiment analysis on Twitter, and the better results with new
approaches, such as regression methods trained with traditional polls. Finally, a
vision of future research on integrating advances on process definitions, modeling,
and evaluation is also discussed, pointing out, among others, the need for better
investigating the application of state-of-art machine learning approaches.
3. On the frontiers of Twitter data and sentiment analysis in election prediction: a
review

Author – Quratulain Alvi,1 Syed Farooq Ali,1 Sheikh Bilal Ahmed,1

Year-2023

Election prediction using sentiment analysis is a rapidly growing field that utilizes
natural language processing and machine learning techniques to predict the
outcome of political elections by analyzing the sentiment of online conversations
and news articles. Sentiment analysis, or opinion mining, involves using text
analysis to identify and extract subjective information from text data sources. In
the context of election prediction, sentiment analysis can be used to gauge public
opinion and predict the likely winner of an election. Significant progress has been
made in election prediction in the last two decades. Yet, it becomes easier to have
its comprehensive view if it has been appropriately classified approach-wise,
citation-wise, and technology-wise. The main objective of this article is to examine
and consolidate the progress made in research about election prediction using
Twitter data. The aim is to provide a comprehensive overview of the current state-
of-the-art practices in this field while identifying potential avenues for further
research and exploration.

4. A systematic review of social media-based sentiment analysis: Emerging


trends and challenges

Author- Qianwen Ariel Xu, Victor Chang

Year-2022

In the present information age, a wide and significant variety of social media
platforms have been developed and become an important part of modern life.
Massive amounts of user-generated data sourced from various social networking
platforms also provide new insights for businesses and governments. However, it
has become difficult to extract useful information from the vast amount of
information effectively. Sentiment analysis provides an automated method of
analyzing sentiment, emotion and opinion in written language to address this issue.
In the existing literature, a large number of scholars have worked on improving the
performance of various sentiment classifiers or applying them to various domains
using data from social networking platforms. This paper explores the challenges
that scholars have encountered and other potential problems in studying sentiment
analysis in social media. It gives insights into the goals of the sentiment analysis
task, the implementation process, and the ways in which it is utilized in various
application domains. It also provides a comparison of different studies and
highlights several challenges related to the datasets, text languages, analysis
methods and evaluation metrics. The paper contributes to the research on sentiment
analysis and can help practitioners select a suitable methodology for their
applications.

5. Design and analysis of tweet-based election models for the 2021 Mexican
legislative election

Author- Alejandro Vigna-G´omez1,2,3* , Javier Murillo2,4 , Manelik Ramirez5,4 ,


Alberto Borbolla4 , Ian M´arquez6,4 and Prasun K. Ray7

Year-2023

Modelling and forecasting real-life human behaviour using online social media is
an active endeavour of interest in politics, government, academia, and industry.
Since its creation in 2006, Twitter has been proposed as a potential laboratory that
could be used to gauge and predict social behaviour. During the last decade, the
user base of Twitter has been growing and becoming more representative of the
general population. Here we analyse this user base in the context of the 2021
Mexican Legislative Election. To do so, we use a dataset of 15 million election-
related tweets in the six months preceding election day. We explore different
election models that assign political preference to either the ruling parties or the
opposition. We find that models using data with geographical attributes determine
the results of the election with better precision and accuracy than conventional
polling methods. These results demonstrate that analysis of public online data can
outperform conventional polling methods, and that political analysis and general
forecasting would likely benefit from incorporating such data in the immediate
future. Moreover, the same Twitter dataset with geographical attributes is
positively correlated with results from official census data on population and
internet usage in Mexico. These findings suggest that we have reached a period in
time when online activity, appropriately curated, can provide an accurate
representation of offline behaviour
CHAPTER-3

PROJECT DECRIPTION

EXISTING SYSTEM

The way politicians communicate with the electorate and run electoral campaigns
was reshaped by the emergence and popularization of contemporary socialmedia
(SM), such as Facebook, Twitter, and Instagram social networks (SNs). Due to the
inherent capabilities ofSM, such as the large amount of available data accessed on
using the SM data to predict election outcomes. Main findings include the low
success of the most-used approach, namely volume and sentiment analysis on
Twitter, and the better results withnew approaches, such as regression methods
trained withtraditional polls. Finally, a vision of future research on integrating
advances in process definitions, modeling, and evaluation is also discussed,
pointing out, among others, the need for better investigating the application ofstate-
of the-art machine learning approaches.

Disadvantages

 Social media (SM) has played a central role in politics andelections


throughout this decade.
 We have entered a new era mediated by SM in which politicians conduct
permanent campaigns without geographic or time constraints, and additional
information about them can be obtained not only by the press but also
directly from their profiles on social networks(SNs) and through other
people sharing and amplifying their voices on SM. amplifying their voices
on SM.
 In this new scenario, SM isused extensively in electoral campaigns, and an
online campaign’s success can even decide elections.

PROPOSED SYSTEM

Social media platforms such as Facebook and Twitter carry a big load of people’s
opinions about politics and leaders, which makes them a good source of
information for researchers to exploit different tasks that include election
predictions. Objective. Identify, categorize, and present a comprehensive overview
of the approaches, techniques, and tools used in election predictions on
Twitter. Method. Conducted a systematic mapping study (SMS) on election
predictions on Twitter and provided empirical evidence for the work published
between January 2010 and January 2021. Results. This research identified 787
studies related to election predictions on Twitter. 98 primary studies were selected
after defining and implementing several inclusion/exclusion criteria. The results
show that most of the studies implemented sentiment analysis (SA) followed by
volume-based and social network analysis (SNA) approaches.

ADVANTAGES
 The majority of the studies employed supervised learning techniques,
subsequently, lexicon-based approach SA, volume-based, and unsupervised
learning. Besides this, 18 types of dictionaries were identified.
 Elections of 28 countries were analyzed, mainly USA (28%) and Indian
(25%) elections. Furthermore, the results revealed that 50% of the primary
studies used English tweets.
 The demographic data showed that academic organizations and conference
venues are the most active.
 The evolution of the work published in the past 11 years shows that most of
the studies employed SA. The implementation of SNA techniques is lower
as compared to SA. Appropriate political labelled datasets are not available,
especially in languages other than English. Deep learning needs to be
employed in this domain to get better predictions.

CHAPTER-3

SYSTEM CONFIGURATION

H/W SYSTEM CONFIGURATION:-


Processor – Intel core2 Duo
Speed - 2.93 Ghz
RAM – 2GB RAM
Hard Disk - 500 GB
Key Board - Standard Windows Keyboard
Mouse - Two or Three Button Mouse
Monitor – LED

S/W SYSTEM CONFIGURATION:-


Operating System: XP and windows 7
Python coding
s/w –google colabs,or spider

CHAPTER -4

SYSTEM IMPLEMENTATION

MODULES DESCRIPTION

Data Sources
A second research gap in the literature is that it is currently not known whether
one platform or data source yields more accurate predictions than others. Because
of the public nature of posts on Twitter, most studies have utilized Twitter data to
predict public opinion, followed by the use of Facebook, forums, blogs, and
YouTube. Given that each platform suffers from its own set of algorithmic
confounds, privacy constraints, and post restrictions, it is unknown whether
consolidating data from multiple platforms would have any advantage over
predictions from a single data source. Most studies have utilized data from a single
platform, while very few use data from multiple platforms.

Diversity in Political Contexts


The existing literature has not provided specific insights regarding the role of
political systems in influencing the predictive power of social media data mining.
The predictive power in a study conducted in semi-authoritarian Singapore was
significantly lower than in studies done in established democracies [30,60]. It can
be inferred that the context in which the elections take place also matters. Issues
like media freedom, competitiveness of the election, and idiosyncrasies of electoral
systems may lead to over- and under-estimations of voters’ preferences. Studies
suggest incorporating local context into election prediction, i.e., by controlling for
incumbency [61], because often the incumbents are better placed to build networks
and use media resources to sustain the campaign through to victory . To further
explore the predictive power of social media data in different political contexts, we
include the factors that are related to the ability of citizens to freely express
themselves and exercise their voting rights, such as the level of political
democracy.

Predictors
Social media predictors are first categorized into two types: sentiment vs.
structure. Sentiment predictors refer to the sentiment/preference extracted from the
users’ text-based social media posts. Sentiment scores are obtained by using (1)
lexicon-based approaches, which use dictionaries of weighted words, in which
accuracy is highly dependent on the quality and relevance of the lexical resources
to the domain it is being applied to; and (2) (supervised) machine learning-based
models, which predict the sentiment score of a piece of text based on the regression
weights that it learns from example data (a corpus) labeled by a human.

Predicted Election and Traditional Poll Results


The types of outcomes predicted are: (1) votes share that political candidates
or parties received during the election; (2) winning party or candidate in the
election; (3) seat share that political candidates or parties received in the election;
and (4) public support or approval towards certain political parties or candidates,
which are often obtained through polls. We have categorized the dependent
variables as either predicting (1) election results, or (2) a traditional poll result.

CHAPTER-5

SOFTWARE ENVIRONMENT
PYHTON SOFTWARE DEVELOPMENT

Python is an easy to learn, powerful programming language. It has efficient high-


level data structures and a simple but effective approach to object-oriented
programming. Python’s elegant syntax and dynamic typing, together with its
interpreted nature, make it an ideal language for scripting and rapid application
development in many areas on most platforms.
The Python interpreter and the extensive standard library are freely available in
source or binary form for all major platforms from the Python web
site, [Link] and may be freely distributed. The same site also
contains distributions of and pointers to many free third party Python modules,
programs and tools, and additional documentation.

The Python interpreter is easily extended with new functions and data types
implemented in C or C++ (or other languages callable from C). Python is also
suitable as an extension language for customizable applications.

This tutorial introduces the reader informally to the basic concepts and features of
the Python language and system. It helps to have a Python interpreter handy for
hands-on experience, but all examples are self-contained, so the tutorial can be
read off-line as well.

For a description of standard objects and modules, see The Python Standard
Library. The Python Language Reference gives a more formal definition of the
language. To write extensions in C or C++, read Extending and Embedding the
Python Interpreter and Python/C API Reference Manual. There are also several
books covering Python in depth.

What is Python

Python is a popular programming language. It was created by Guido van Rossum,


and released in 1991.

It is used for:

 web development (server-side),


 software development,
 mathematics,
 system scripting.

Python Install

Many PCs and Macs will have python already installed.

To check if you have python installed on a Windows PC, search in the start bar for
Python or run the following on the Command Line ([Link]):

C:\Users\Your Name>python --version

To check if you have python installed on a Linux or Mac, then on linux open the
command line or on Mac open the Terminal and type:

python --version

If you find that you do not have Python installed on your computer, then you can
download it for free from the following website: [Link]

Python Quickstart

Python is an interpreted programming language, this means that as a developer you


write Python (.py) files in a text editor and then put those files into the python
interpreter to be executed.

The way to run a python file is like this on the command line:

C:\Users\Your Name>python [Link]

Where "[Link]" is the name of your python file.

Let's write our first Python file, called [Link], which can be done in any
text editor.
The Python Command Line

To test a short amount of code in python sometimes it is quickest and easiest not to
write the code in a file. This is made possible because Python can be run as a
command line itself.

Type the following on the Windows, Mac or Linux command line:

C:\Users\Your Name>python

Or, if the "python" command did not work, you can try "py":

C:\Users\Your Name>py

From there you can write any python, including our hello world example from
earlier in the tutorial:

C:\Users\YourName>python

Argument Passing

When known to the interpreter, the script name and additional arguments thereafter
are turned into a list of strings and assigned to the argv variable in the sys module.
You can access this list by executing import sys. The length of the list is at least
one; when no script and no arguments are given, [Link][0] is an empty string.
When the script name is given as '-' (meaning standard input), [Link][0] is set
to '-'. When -c command is used, [Link][0] is set to '-c'. When -m module is
used, [Link][0] is set to the full name of the located module. Options found
after -c command or -m module are not consumed by the Python interpreter’s
option processing but left in [Link] for the command or module to handle.

Interactive Mode
When commands are read from a tty, the interpreter is said to be in interactive
mode. In this mode it prompts for the next command with the primary prompt,
usually three greater-than signs (>>>); for continuation lines it prompts with
the secondary prompt, by deFraud three dots (...). The interpreter prints a welcome
message stating its version number and a copyright notice before printing the first
prompt:

$ python3.10

Python 3.10 (deFraud , June 4 2019, 09:25:04)

[GCC 4.8.2] on linux

Type "help", "copyright", "credits" or "license" for more information.

Continuation lines are needed when entering a multi-line construct. As an example,


take a look at this if statement:

the_world_is_flat = True

if the_world_is_flat:

print("Be careful not to fall off!")

Be careful not to fall off!

For more on interactive mode, see Interactive Mode.

The Interpreter and Its Environment

By deFraud , Python source files are treated as encoded in UTF-8. In that


encoding, characters of most languages in the world can be used simultaneously in
string literals, identifiers and comments — although the standard library only uses
ASCII characters for identifiers, a convention that any portable code should follow.
To display all these characters properly, your editor must recognize that the file is
UTF-8, and it must use a font that supports all the characters in the file.

To declare an encoding other than the deFraud one, a special comment line should
be added as the first line of the file. The syntax is as follows:

# -*- coding: encoding -*-

where encoding is one of the valid codecs supported by Python.

For example, to declare that Windows-1252 encoding is to be used, the first line of
your source code file should be:

# -*- coding: cp1252 -*-

One exception to the first line rule is when the source code starts with a UNIX
“shebang” line. In this case, the encoding declaration should be added as the
second line of the file. For example:

#!/usr/bin/env python3

# -*- coding: cp1252 -*-

An Informal Introduction to Python

In the following examples, input and output are distinguished by the presence or
absence of prompts (>>> and …): to repeat the example, you must type everything
after the prompt, when the prompt appears; lines that do not begin with a prompt
are output from the interpreter. Note that a secondary prompt on a line by itself in
an example means you must type a blank line; this is used to end a multi-line
command.
You can toggle the display of prompts and output by clicking on >>> in the upper-
right corner of an example box. If you hide the prompts and output for an example,
then you can easily copy and paste the input lines into your interpreter.

Many of the examples in this manual, even those entered at the interactive prompt,
include comments. Comments in Python start with the hash character, #, and
extend to the end of the physical line. A comment may appear at the start of a line
or following whitespace or code, but not within a string literal. A hash character
within a string literal is just a hash character. Since comments are to clarify code
and are not interpreted by Python, they may be omitted when typing in examples.

Some examples:

# this is the first comment

spam = 1 # and this is the second comment

# ... and now a third!

text = "# This is not a comment because it's inside quotes."

Using Python as a Calculator

Let’s try some simple Python commands. Start the interpreter and wait for the
primary prompt, >>>. (It shouldn’t take long.)

Numbers

The interpreter acts as a simple calculator: you can type an expression at it and it
will write the value. Expression syntax is straightforward: the
operators +, -, * and / work just like in most other languages (for example, Pascal
or C); parentheses (()) can be used for grouping. For example:

>>>

>>> 2 + 2

>>> 50 - 5*6

20

>>> (50 - 5*6) / 4

5.0

>>> 8 / 5 # division always returns a floating point number

1.6

The integer numbers (e.g. 2, 4, 20) have type int, the ones with a fractional part
(e.g. 5.0, 1.6) have type float. We will see more about numeric types later in the
tutorial.

Division (/) always returns a float. To do floor division and get an integer result
(discarding any fractional result) you can use the // operator; to calculate the
remainder you can use %:

>>> 17 / 3 # classic division returns a float

5.666666666666667

>>> 17 // 3 # floor division discards the fractional part


5

>>> 17 % 3 # the % operator returns the remainder of the division

>>> 5 * 3 + 2 # floored quotient * divisor + remainder

17

With Python, it is possible to use the ** operator to calculate powers 1:

>>> 5 ** 2 # 5 squared

25

>>> 2 ** 7 # 2 to the power of 7

128

The equal sign (=) is used to assign a value to a variable. Afterwards, no result is
displayed before the next interactive prompt:

>>> width = 20

>>> height = 5 * 9

>>> width * height

900

If a variable is not “defined” (assigned a value), trying to use it will give you an
error:

>>> n # try to access an undefined variable

Traceback (most recent call last):


File "<stdin>", line 1, in <module>

NameError: name 'n' is not defined

There is full support for floating point; operators with mixed type operands convert
the integer operand to floating point:

>>> 4 * 3.75 - 1

14.0

In interactive mode, the last printed expression is assigned to the variable _. This
means that when you are using Python as a desk calculator, it is somewhat easier to
continue calculations, for example:

>>> tax = 12.5 / 100

>>> price = 100.50

>>> price * tax

12.5625

>>> price + _

113.0625

>>> round(_, 2)

113.06

This variable should be treated as read-only by the user. Don’t explicitly assign a
value to it — you would create an independent local variable with the same name
masking the built-in variable with its magic behavior.
In addition to int and float, Python supports other types of numbers, such
as Decimal and Fraction. Python also has built-in support for complex numbers,
and uses the j or J suffix to indicate the imaginary part (e.g. 3+5j).

Strings

Besides numbers, Python can also manipulate strings, which can be expressed in
several ways. They can be enclosed in single quotes ('...') or double quotes ("...")
with the same result 2. \ can be used to escape quotes:

>>>

>>> 'spam eggs' # single quotes

'spam eggs'

>>> 'doesn\'t' # use \' to escape the single quote...

"doesn't"

>>> "doesn't" # ...or use double quotes instead

"doesn't"

>>> '"Yes," they said.'

'"Yes," they said.'

>>> "\"Yes,\" they said."

'"Yes," they said.'

>>> '"Isn\'t," they said.'

'"Isn\'t," they said.'


In the interactive interpreter, the output string is enclosed in quotes and special
characters are escaped with backslashes. While this might sometimes look different
from the input (the enclosing quotes could change), the two strings are equivalent.
The string is enclosed in double quotes if the string contains a single quote and no
double quotes, otherwise it is enclosed in single quotes. The print() function
produces a more readable output, by omitting the enclosing quotes and by printing
escaped and special characters:

>>> '"Isn\'t," they said.'

'"Isn\'t," they said.'

>>> print('"Isn\'t," they said.')

"Isn't," they said.

>>> s = 'First line.\nSecond line.' # \n means newline

>>> s # without print(), \n is included in the output

'First line.\nSecond line.'

>>> print(s) # with print(), \n produces a new line

First line.

Second line.

If you don’t want characters prefaced by \ to be interpreted as special characters,


you can use raw strings by adding an r before the first quote:

>>>

>>> print('C:\some\name') # here \n means newline!


C:\some

ame

>>> print(r'C:\some\name') # note the r before the quote

C:\some\name

String literals can span multiple lines. One way is using triple-
quotes: """...""" or '''...'''. End of lines are automatically included in the string, but
it’s possible to prevent this by adding a \ at the end of the line. The following
example:

print("""\

Usage: thingy [OPTIONS]

-h Display this usage message

-H hostname Hostname to connect to

""")

produces the following output (note that the initial newline is not included):

Usage: thingy [OPTIONS]

-h Display this usage message

-H hostname Hostname to connect to

Strings can be concatenated (glued together) with the + operator, and repeated
with *:

>>>
>>> # 3 times 'un', followed by 'ium'

>>> 3 * 'un' + 'ium'

'unununium'

Two or more string literals (i.e. the ones enclosed between quotes) next to each
other are automatically concatenated.

>>>

>>> 'Py' 'thon'

'Python'

This feature is particularly useful when you want to break long strings:

>>>

>>> text = ('Put several strings within parentheses '

... 'to have them joined together.')

>>> text

'Put several strings within parentheses to have them joined together.'

This only works with two literals though, not with variables or expressions:

>>>

>>> prefix = 'Py'

>>> prefix 'thon' # can't concatenate a variable and a string literal

File "<stdin>", line 1


prefix 'thon'

SyntaxError: invalid syntax

>>> ('un' * 3) 'ium'

File "<stdin>", line 1

('un' * 3) 'ium'

SyntaxError: invalid syntax

If you want to concatenate variables or a variable and a literal, use +:

>>>

>>> prefix + 'thon'

'Python'

Strings can be indexed (subscripted), with the first character having index 0. There
is no separate character type; a character is simply a string of size one:

>>>

>>> word = 'Python'

>>> word[0] # character in position 0

'P'

>>> word[5] # character in position 5


'n'

Indices may also be negative numbers, to start counting from the right:

>>> word[-1] # last character

'n'

>>> word[-2] # second-last character

'o'

>>> word[-6]

'P'

Note that since -0 is the same as 0, negative indices start from -1.

In addition to indexing, slicing is also supported. While indexing is used to obtain


individual characters, slicing allows you to obtain substring:

>>>

>>> word[0:2] # characters from position 0 (included) to 2 (excluded)

'Py'

>>> word[2:5] # characters from position 2 (included) to 5 (excluded)

'tho'

Slice indices have useful deFraud s; an omitted first index deFraud s to zero, an
omitted second index deFraud s to the size of the string being sliced.

>>> word[:2] # character from the beginning to position 2 (excluded)

'Py'
>>> word[4:] # characters from position 4 (included) to the end

'on'

>>> word[-2:] # characters from the second-last (included) to the end

'on'

Note how the start is always included, and the end always excluded. This makes
sure that s[:i] + s[i:] is always equal to s:

>>> word[:2] + word[2:]

'Python'

>>> word[:4] + word[4:]

'Python'

One way to remember how slices work is to think of the indices as


pointing between characters, with the left edge of the first character numbered 0.
Then the right edge of the last character of a string of n characters has index n, for
example:

The first row of numbers gives the position of the indices 0…6 in the string; the
second row gives the corresponding negative indices. The slice from i to j consists
of all characters between the edges labeled i and j, respectively.

For non-negative indices, the length of a slice is the difference of the indices, if
both are within bounds. For example, the length of word.
CHAPTER -4

PROJECT DESCRIPTION

INPUT DESIGN

Twitter is the SN used in most of the studies (84%), and in many of them (75%), it
was the only SN used as input. However, Twitter is not a good sample, even
considering only SM users, due to its having very few active users (326 million),
relative to other SN, such as Facebook (2.6 billion) and Instagram (1,1 billion),
according to a 2020 report [36]. Despite these data, it is hard to find a discussion
about why studies focused on Twitter. After analyzing the API of these SNs [41],
[50], we hypothesized that Twitter was chosen because it is easier for researchers
to collect data on this platform. For example, starting on August 2018, the approval
process to gather data from Facebook and Instagram consisted of developing and
deploying a fully functional system, creating and publishing a privacy policy and
terms of use, recording a video showing all the functionalities related to Facebook
and Instagram data collection, creating test accounts allowing Facebook employees
to test the system and, in many cases, sending formal documentation of an
institution responsible for the system. By contrast, in August 2019, the Twitter
approval process only involved completing a form with information about the
system.

OBJECTIVES

In most studies, many data collection choices were arbitrary, such as the data
collection period, which usually varied from 3 days to 3 months before elections,
and the keywords used for open search on volume/sentiment approaches. This
created many problems, such as those presented by [30], in which the performance
was too unstable because it depended strongly on such parameterizations, and
unintentional data dredging could occur, due to post hoc analysis. Also, it
reinforces the argument presented by Jungherr [22] who, after replicating the
seminal study of Tumasjan [5], argued that “the results are contingent on arbitrary
choices of the authors,” and indicated that simply including one more party or day
of collection would greatly change the results.

OUTPUT DESIGN
Modelling and forecasting real-life human behaviour using online social media is
an active endeavour of interest in politics, government, academia, and industry.
Since its creation in 2006, Twitter has been proposed as a potential laboratory that
could be used to gauge and predict social behaviour. During the last decade, the
user base of Twitter has been growing and becoming more representative of the
general population. Here we analyse this user base in the context of the 2021
Mexican Legislative Election. To do so, we use a dataset of 15 million election-
related tweets in the six months preceding election day. We explore different
election models that assign political preference to either the ruling parties or the
opposition. We find that models using data with geographical attributes determine
the results of the election with better precision and accuracy than conventional
polling methods. These results demonstrate that analysis of public online data can
outperform conventional polling methods, and that political analysis and general
forecasting would likely benefit from incorporating such data in the immediate
future. Moreover, the same Twitter dataset with geographical attributes is
positively correlated with results from official census data on population and
internet usage in Mexico. These findings suggest that we have reached a period in
time when online activity, appropriately curated, can provide an accurate
representation of offline behaviour.
FEASIBILITY STUDY

The feasibility of the project is analyzed in this phase and business proposal is
put forth with a very general plan for the project and some cost estimates. During
system analysis the feasibility study of the proposed system is to be carried out.

This is to ensure that the proposed system is not a burden to the company. For
feasibility analysis, some understanding of the major requirements for the system
is essential.

Three key considerations involved in the feasibility analysis are

 ECONOMICAL FEASIBILITY
 TECHNICAL FEASIBILITY
 SOCIAL FEASIBILITY

ECONOMICAL FEASIBILITY

This study is carried out to check the economic impact that the system will
have on the organization. The amount of fund that the company can pour into the
research and development of the system is limited.

The expenditures must be justified. Thus the developed system as well


within the budget and this was achieved because most of the technologies used are
freely available. Only the customized products had to be purchased.

TECHNICAL FEASIBILITY
This study is carried out to check the technical feasibility, that is, the
technical requirements of the system. Any system developed must not have a high
demand on the available technical resources.
This will lead to high demands on the available technical resources. This
will lead to high demands being placed on the client. The developed system must
have a modest requirement, as only minimal or null changes are required for
implementing this system.

SOCIAL FEASIBILITY

The aspect of study is to check the level of acceptance of the system by the
user. This includes the process of training the user to use the system efficiently.

The user must not feel threatened by the system, instead must accept it as a
necessity.

The level of acceptance by the users solely depends on the methods that are
employed to educate the user about the system and to make him familiar with it.
CHAPTER-5

SYSTEM TESTING
The purpose of testing is to discover errors. Testing is the process of trying
to discover every conceivable fault or weakness in a work product. It provides a
way to check the functionality of components, sub assemblies, assemblies and/or a
finished product It is the process of exercising software with the intent of ensuring
that the Software system meets its requirements and user expectations and does not
fail in an unacceptable manner. There are various types of test. Each test type
addresses a specific testing requirement.

TYPES OF TESTS
Unit Testing
Unit testing involves the design of test cases that validate that the internal
program logic is functioning properly, and that program inputs produce valid
outputs. All decision branches and internal code flow should be validated. It is the
testing of individual software units of the application .it is done after the
completion of an individual unit before integration. This is a structural testing, that
relies on knowledge of its construction and is invasive. Unit tests perform basic
tests at component level and test a specific business process, application, and/or
system configuration. Unit tests ensure that each unique path of a business process
performs accurately to the documented specifications and contains clearly defined
inputs and expected results.

Integration Testing
Testing is event driven and is more concerned with the basic outcome of
screens or fields. Integration tests demonstrate that although the components were
individually satisfaction, as shown by successfully unit testing, the combination of
components is correct and consistent.

Functional Testing

Functional tests provide systematic demonstrations that functions tested are


available as specified by the business and technical requirements, system
documentation, and user manuals.

Functional testing is centered on the following items:

Valid Input : identified classes of valid input must be accepted.

Invalid Input : identified classes of invalid input must be


rejected.

Functions : identified functions must be exercised.

Output : identified classes of application outputs must be

exercised.

Systems/Procedures : interfacing systems or procedures must be


invoked.
Organization and preparation of functional tests is focused on requirements,
key functions, or special test cases. In addition, systematic coverage pertaining to
identify Business process flows; data fields, predefined processes, and successive
processes must be considered for testing.

Before functional testing is complete, additional tests are identified and the
effective value of current tests is determined.

System Testing
System testing ensures that the entire integrated software system meets
requirements. It tests a configuration to ensure known and predictable results. An
example of system testing is the configuration oriented system integration test.
System testing is based on process descriptions and flows, emphasizing pre-driven
process links and integration points.

White Box Testing


White Box Testing is a testing in which in which the software tester has
knowledge of the inner workings, structure and language of the software, or at least
its purpose. It is purpose. It is used to test areas that cannot be reached from a black
box level.

Black Box Testing


Black Box Testing is testing the software without any knowledge of the inner
workings, structure or language of the module being tested.

Black box tests, as most other kinds of tests, must be written from a
definitive source document, such as specification or requirements document, such
as specification or requirements document. It is a testing in which the software
under test is treated, as a black box .you cannot “see” into it. The test provides
inputs and responds to outputs without considering how the software works.
Unit Testing

Unit testing is usually conducted as part of a combined code and unit test
phase of the software lifecycle, although it is not uncommon for coding and unit
testing to be conducted as two distinct phases.

Test strategy and approach


Field testing will be performed manually and functional tests will be written
in detail.

Test objectives

 All field entries must work properly.


 Pages must be activated from the identified link.
 The entry screen, messages and responses must not be delayed.
Features to be tested

 Verify that the entries are of the correct format


 No duplicate entries should be allowed
 All links should take the user to the correct page.
Integration Testing
Software integration testing is the incremental integration testing of two or
more integrated software components on a single platform to produce failures
caused by interface defects.

Test Results: All the test cases mentioned above passed successfully. No defects
encountered.
Acceptance Testing
User Acceptance Testing is a critical phase of any project and requires
significant participation by the end user. It also ensures that the system meets the
functional requirements.

Conclusion

Digital traces of human behavior now enable the testing of social science theories
in new ways, and are also likely to pave the way for new theories in social science.
Both in terms of its nature and scale, user-generated social media data is
dramatically different from the existing market information regime data, which is
constructed to serve a specific purpose and typically relies on smaller samples of
unconnected citizens [25]. As new studies and new methods for social media-based
predictions proliferate, it is important to consider both how data are generated, and
how they are analyzed, and consolidate a validated framework for data analysis. As
Kuhn argued, there is a close, intimate link between the scientific puzzle and the
methodology designed to solve it [68]. Social media analytics will not replace
survey-based public opinion studies, but they do offer additional data sources and
tools for improving our understanding of public opinion and political behavior,
while at the same time changing the norms around voter engagement, mobilization,
and preference elicitation. Among other things, we see the potential of social
media data to provide deeper insights regarding the opinion dynamics and the role
of opinion leaders in the process of public opinion formation and change.
Furthermore, the problems that plague our digital ecologies today, such as filter
bubbles, disinformation, incivility, and hate speech, can all be better understood
using social media data and computational methods, than with survey-based
research. Still, we note a general lack of theoretically-informed work in the field—
most studies report predictions without paying sufficient attention to the
underlying mechanisms and processes. We are also concerned about the future
availability of social media data, as citizens switch to more private, encrypted
instant messaging apps and social media companies further restrict access to user
data because of legal, commercial, and privacy concerns. Thus, policymakers
should create legal and regulatory environments that will promote access to social
media data for the research community in ways that guarantee strong privacy
protection rights for the users.
SOURCE CODE

import face_recognition as fr

import cv2

import numpy as np

import os

path = "./train/"

known_names = []

known_name_encodings = []

images = [Link](path)

for _ in images:

image = fr.load_image_file(path + _)

image_path = path + _

encoding = fr.face_encodings(image)[0]

known_name_encodings.append(encoding)
known_names.append([Link]([Link](image_path))
[0].capitalize())

print(known_names)

test_image = "./test/[Link]"

image = [Link](test_image)

# image = [Link](image, cv2.COLOR_BGR2RGB)

face_locations = fr.face_locations(image)

face_encodings = fr.face_encodings(image, face_locations)

for (top, right, bottom, left), face_encoding in zip(face_locations, face_encodings):

matches = fr.compare_faces(known_name_encodings, face_encoding)

name = ""

face_distances = fr.face_distance(known_name_encodings, face_encoding)

best_match = [Link](face_distances)
if matches[best_match]:

name = known_names[best_match]

[Link](image, (left, top), (right, bottom), (0, 0, 255), 2)

[Link](image, (left, bottom - 15), (right, bottom), (0, 0, 255),


[Link])

font = cv2.FONT_HERSHEY_DUPLEX

[Link](image, name, (left + 6, bottom - 6), font, 1.0, (255, 255, 255), 1)

[Link]("Result", image)

[Link]("./[Link]", image)

[Link](0)

[Link]()

# this file is used to detect face

# and then store the data of the face

import cv2

import numpy as np
# import the file where data is

# stored in a csv file format

import npwriter

name = input("Enter your name: ")

# this is used to access the web-cam

# in order to capture frames

cap = [Link](0)

classifier = [Link](r"C:\Users\LENOVO\Desktop\FACE\dataset\
haarcascade_frontalface_default.xml")

# this is class used to detect the faces as provided

# with a haarcascade_frontalface_default.xml file as data

f_list = []

while True:

ret, frame = [Link]()


# converting the image into gray

# scale as it is easy for detection

gray = [Link](frame, cv2.COLOR_BGR2GRAY)

# detect multiscale, detects the face and its coordinates

faces = [Link](gray, 1.5, 5)

# this is used to detect the face which

# is closest to the web-cam on the first position

faces = sorted(faces, key = lambda x: x[2]*x[3],

reverse = True)

# only the first detected face is used

faces = faces[:1]

# len(faces) is the number of

# faces showing in a frame

if len(faces) == 1:

# this is removing from tuple format


face = faces[0]

# storing the coordinates of the

# face in different variables

x, y, w, h = face

# this is will show the face

# that is being detected

im_face = frame[y:y + h, x:x + w]

[Link]("face", im_face)

if not ret:

continue

[Link]("full", frame)

key = [Link](1)
# this will break the execution of the program

# on pressing 'q' and will click the frame on pressing 'c'

if key & 0xFF == ord('q'):

break

elif key & 0xFF == ord('c'):

if len(faces) == 1:

gray_face = [Link](im_face, cv2.COLOR_BGR2GRAY)

gray_face = [Link](gray_face, (100, 100))

print(len(f_list), type(gray_face), gray_face.shape)

# this will append the face's coordinates in f_list

f_list.append(gray_face.reshape(-1))

else:

print("face not found")

# this will store the data for detected

# face 10 times in order to increase accuracy

if len(f_list) == 10:

break
# declared in npwriter

[Link](name, [Link](f_list))

[Link]()

[Link]()

# import general libraries

import pandas as pd

import [Link] as plt

import math

import numpy as np

# import sklearn

from sklearn import preprocessing

from sklearn.model_selection import train_test_split, cross_val_score

from [Link] import DecisionTreeClassifier

from sklearn import metrics


from sklearn import tree

# import dataset

df = pd.read_csv('/content/county_census_and_election_result.csv')

[Link]()

# check dataset dimension

print(f"The dimension of the dataset is: {[Link]}")

# check data types of each columns

[Link]

"""## 1.1. Process Original Dataset"""

# select the rows in the dataset that have the voting labels (years 2008, 2012, 2016,
2020)

selected_df = [Link][df['year'].isin([2008, 2012, 2016, 2020])]

# reset index
selected_df = selected_df.reset_index(drop=True)

# view the dataset

selected_df

# check dataset dimension

print(f"The dimension of the dataset is: {[Link]}")

# Examine and replace missing values

print(selected_df.isnull().[Link]())

"""**Note:** As we consider each row as a data point to our ML model, we will


drop the categorical state names and abbreviation. We will also drop the all the
number of votes labels except for "winner" columns (0 for democrats and 1 for
republican). We don't concern other parties as they do not have a history of
winning anyway.

**Drop**:

- county_fips: as there is no correlation between fips and vote numbers.

- state_po: as there is no correlation between state appreviation and vote numbers.


- county_name: as there is no correlation between county names and vote numbers.

- democrat: as we do classification task.

- green: as we do classification task.

- liberitarian: as we do classification task.

- other: as we do classification task.

- republican: as we do classification task.

"""

# Feature selection

more_selected_df = selected_df.drop(columns=['county_fips', 'state_po',


'county_name', 'democrat',

'green', 'liberitarian', 'other', 'republican'])

more_selected_df.head()

# fill the NaN with both forward and backward fills

more_selected_df = more_selected_df.fillna(method='ffill')

more_selected_df = more_selected_df.fillna(method='bfill')
# check if there is any NaN values

print(more_selected_df.isnull().[Link]())

more_selected_df.head()

print(f"The dimension of the dataset is: {more_selected_df.shape}")

"""**Note:** divide the dataset into 2 datasets. Dataset 1 is within 2008, 2012,
2016 for machine learning modeling. Dataset 2 is of 2020 for testing the prediction
power of the trained and validated data

## 1.2. Process The Dataset Of 2008, 2012, 2016 Election Years

"""

# select rows with the years 2008, 2012, 2016 presidential election

df1 = more_selected_df.loc[more_selected_df['year'].isin([2008, 2012, 2016])]

# Model-dev 4: Feature selection

df1 = [Link](columns=['year'])
# reset index

df1 = df1.reset_index(drop=True)

# view the dataset

[Link]()

print(f"The dimension of the dataset is: {[Link]}")

# normalization + divide the dataset into input features and output labels

X_df1 = df1[list([Link][:-1])]

y_df1 = df1['winner'].to_numpy().reshape(-1, 1)

# data normalization

X_scaler = [Link](feature_range=(-1,1))

X_df1 = X_scaler.fit_transform(X_df1)

y_scaler = [Link](feature_range=(-1,1))

y_df1 = y_scaler.fit_transform(y_df1)
# partition into training/test/validation

X_train, X_test, y_train, y_test = train_test_split(X_df1, y_df1, test_size=0.3,


random_state=0)

"""## 1.3. Process The Dataset Of 2020 Election Year"""

# select rows with year 2020 presidential election

df2 = more_selected_df.loc[more_selected_df['year'].isin([2020])]

# Model-dev 4: Feature selection

df2 = [Link](columns=['year'])

# reset index

df2 = df2.reset_index(drop=True)

# view the dataset

[Link]()

print(f"The dimension of the dataset is: {[Link]}")


# normalization + divide the dataset into input features and output labels

X_df2 = df2[list([Link][:-1])]

y_df2 = df2['winner'].to_numpy().reshape(-1, 1)

# data normalization

X_scaler = [Link](feature_range=(-1,1))

X_df2 = X_scaler.fit_transform(X_df2)

y_scaler = [Link](feature_range=(-1,1))

y_df2 = y_scaler.fit_transform(y_df2)

"""## 2.1. Decision Tree With Gini Criterion"""

# create and train the model

clf1 = DecisionTreeClassifier(criterion='gini', splitter='best')

[Link](X_train, y_train)
# predict on validate set

y_test_pred = [Link](X_test)

# predict on train set

y_train_pred = [Link](X_train)

accuracy = metrics.accuracy_score(y_test, y_test_pred)

f1 = metrics.f1_score(y_test, y_test_pred)

prec = metrics.precision_score(y_test, y_test_pred)

recall = metrics.recall_score(y_test, y_test_pred)

roc_auc = metrics.roc_auc_score(y_test, y_test_pred)

print("Training MSE = %f" % metrics.mean_squared_error(y_train, y_train_pred))

print("Validating MSE = %f" % metrics.mean_squared_error(y_test, y_test_pred))

print('Accuracy = %f' % (accuracy))

print('F1 Score = %f' % (f1))

print('Precision Score = %f' % (prec))

print('Recall Score = %f' % (recall))

print('ROC-AUC Score = %f' % (roc_auc))


"""## 2.2. Decision Tree With Entropy Criterion"""

# create and train the model

clf2 = DecisionTreeClassifier(criterion='entropy', splitter='best')

[Link](X_train, y_train)

# predict on validate set

y_test_pred = [Link](X_test)

# predict on train set

y_train_pred = [Link](X_train)

accuracy = metrics.accuracy_score(y_test, y_test_pred)

f1 = metrics.f1_score(y_test, y_test_pred)

prec = metrics.precision_score(y_test, y_test_pred)

recall = metrics.recall_score(y_test, y_test_pred)

roc_auc = metrics.roc_auc_score(y_test, y_test_pred)


print("Training MSE = %f" % metrics.mean_squared_error(y_train, y_train_pred))

print("Validating MSE = %f" % metrics.mean_squared_error(y_test, y_test_pred))

print('Accuracy = %f' % (accuracy))

print('F1 Score = %f' % (f1))

print('Precision Score = %f' % (prec))

print('Recall Score = %f' % (recall))

print('ROC-AUC Score = %f' % (roc_auc))

"""# 3. Test Trained Models With The 2020 Dataset

## 3.1. Decision Tree With Gini Criterion

"""

clf1_pred_y_2020 = [Link](X_df2)

accuracy = metrics.accuracy_score(y_df2, clf1_pred_y_2020)

f1 = metrics.f1_score(y_df2, clf1_pred_y_2020)

prec = metrics.precision_score(y_df2, clf1_pred_y_2020)


recall = metrics.recall_score(y_df2, clf1_pred_y_2020)

roc_auc = metrics.roc_auc_score(y_df2, clf1_pred_y_2020)

print("Testing MSE = %f" % metrics.mean_squared_error(y_df2,


clf1_pred_y_2020))

print('Accuracy = %f' % (accuracy))

print('F1 Score = %f' % (f1))

print('Precision Score = %f' % (prec))

print('Recall Score = %f' % (recall))

print('ROC-AUC Score = %f' % (roc_auc))

# count the vote each counties

demo_count1 = repu_count1 = 0

for n in clf1_pred_y_2020:

if n == -1:

demo_count1 += 1

else:

repu_count1 += 1
"""## 3.2. Decision Tree With Entropy Criterion"""

clf2_pred_y_2020 = [Link](X_df2)

accuracy = metrics.accuracy_score(y_df2, clf2_pred_y_2020)

f1 = metrics.f1_score(y_df2, clf2_pred_y_2020)

prec = metrics.precision_score(y_df2, clf2_pred_y_2020)

recall = metrics.recall_score(y_df2, clf2_pred_y_2020)

roc_auc = metrics.roc_auc_score(y_df2, clf2_pred_y_2020)

print("Testing MSE = %f" % metrics.mean_squared_error(y_df2,


clf2_pred_y_2020))

print('Accuracy = %f' % (accuracy))

print('F1 Score = %f' % (f1))

print('Precision Score = %f' % (prec))

print('Recall Score = %f' % (recall))

print('ROC-AUC Score = %f' % (roc_auc))


# count the vote each counties

demo_count2 = repu_count2 = 0

for n in clf2_pred_y_2020:

if n == -1:

demo_count2 += 1

else:

repu_count2 += 1

"""# 4. Plot Performance Of 2 Decision Tree Models"""

votes = [demo_count1, repu_count1, demo_count2, repu_count2]

names = ['Democrats Prediction Model 1', 'Republican Prediction Model 1',

'Democrats Prediction Model 2', 'Republican Prediction Model 2']

[Link](figsize=(20, 5))

[Link](names, votes, color=['blue', 'red', 'blue', 'red'])

[Link]('Party Votes Counts in 2020 at Each County in the US - Decision Tree


Models Prediction Comparison')
OUTPUT SCREENCHAT
REFFERENCE

1. King, G. Ensuring the data-rich future of the social sciences. Science 2011, 331,
719–721. [Google Scholar] [CrossRef] [PubMed]
2. Lazer, D.; Pentland, A.S.; Adamic, L.; Aral, S.; Barabasi, A.L.; Brewer, D.;
Gutmann, M. Life in the network: The coming age of computational social
science. Science 2009, 323, 721–723. [Google Scholar] [CrossRef][Green Version]
3. boyd, D.; Crawford, K. Critical questions for big data: Provocations for a cultural,
technological, and scholarly phenomenon. Inf. Commun. Soc. 2012, 15, 662–679.
[Google Scholar] [CrossRef]
4. Salganik, M.J. Bit by Bit: Social Research in the Digital Age; Princeton University
Press: Princeton, NJ, USA, 2018. [Google Scholar]
5. Choi, H.; Varian, H. Predicting the present with Google Trends. Econ.
Rec. 2012, 88, 2–9. [Google Scholar] [CrossRef]
6. Manovich, L. Trending: The promises and the challenges of big social data.
In Debates in the Digital Humanities; Gold, M.K., Ed.; The University of
Minnesota Press: Minneapolis, MN, USA, 2011; pp. 460–475. [Google Scholar]
7. González-Bailón, S.; Wang, N.; Rivero, A.; Borge-Holthoefer, J.; Moreno, Y.
Assessing the bias in samples of large online networks. Soc. Netw. 2014, 38, 16–
27. [Google Scholar] [CrossRef]
8. Jungherr, A.; Jürgens, P.; Schoen, H. Why the Pirate party won the German
election of 2009 or the trouble with predictions: A response to Tumasjan, A.,
Sprenger, T.O., Sander, P.G., & Welpe, I.M. “Predicting elections with Twitter:
What 140 characters reveal about political sentiment.”. Soc. Sci. Comput.
Rev. 2012, 30, 229–234. [Google Scholar] [CrossRef]
9. Bucher, T.; Helmond, A. The Affordances of Social Media Platforms. In The
SAGE Handbook of Social Media; Burgess, J., Marwick, A., Poell, T., Eds.; Sage
Publications: New York, NY, USA, 2018; pp. 233–253. [Google Scholar]
10. Ratkiewicz, J.; Conover, M.; Meiss, M.R.; Gonçalves, B.; Flammini, A.; Menczer,
F. Detecting and Tracking Political Abuse in Social Media. In Proceedings of the
20th international conference companion on World Wide Web, Hyderabad, India,
28 March–1 April 2011. [Google Scholar]
11. Gayo-Avello, D. A meta-analysis of state-of-the-art electoral prediction from
Twitter data. Soc. Sci. Comput. Rev. 2013, 31, 649–679. [Google Scholar]
[CrossRef][Green Version]
12. Bourdieu, P. Public opinion does not exist. In Communication and Class Strugzle;
Mattelart, A., Siegelaub., S., Eds.; International General/IMMRC: New York, NY,
USA, 1979; pp. 124–310. [Google Scholar]
13. Lazarsfeld, P.F. Public opinion and the classical tradition. Public Opin.
Q. 1957, 21, 39–53. [Google Scholar] [CrossRef]
14. Katz, E. The two-step flow of communication: An up-to-date report on an
hypothesis. Public Opin. Q. 1957, 21, 61–78. [Google Scholar] [CrossRef]
15. Katz, E.; Lazarsfeld, P.F. Personal Influence: The Part Played by People in the
Flow of Mass Communications; Free Press: New York, NY, USA, 1955. [Google
Scholar]
16. Ceron, A.; Curini, L.; Iacus, S.M.; Porro, G. Every tweet counts? How sentiment
analysis of social media can improve our knowledge of citizens’ political
preferences with an application to Italy and France. New Media Soc. 2014, 16,
340–358. [Google Scholar] [CrossRef]
17. Espeland, W.N.; Sauder, M. Rankings and reactivity: How public measures
recreate social worlds. Am. J. Sociol. 2007, 113, 1–40. [Google Scholar]
[CrossRef][Green Version]
18. Berinsky, A.J. The two faces of public opinion. Am. J. Political Sci. 1999, 43,
1209–1230. [Google Scholar] [CrossRef]
19. Blumer, H. Public opinion and public opinion polling. Am. Sociol. Rev. 1948, 13,
542–549. [Google Scholar] [CrossRef]
20. Anstead, N.; O’Loughlin, B. Social media analysis and public opinion: The 2010
UK general election. J. Comput. Mediat. Commun. 2015, 37, 204–220. [Google
Scholar] [CrossRef][Green Version]

You might also like