Election Prediction Using Deep Learning
Election Prediction Using Deep Learning
ABSTRACT:
The way politicians communicate with the electorate and run electoral campaigns
was reshaped by the emergence and popularization of contemporary social media
(SM), such as Facebook, Twitter, and Instagram social networks (SNs). Due to the
inherent capabilities of SM, such as the large amount of available data accessed in
real time, a new research subject has emerged, focusing on using the SM data to
predict election outcomes. Despite many studies conducted in the last decade,
results are very controversial and many times challenged. In this context, this
article aims to investigate and summarize how research on predicting elections
based on the SM data has evolved since its beginning, to outline the state of both
the art and the practice, and to identify research opportunities within this field. In
terms of method, we performed a systematic literature review analyzing the
quantity and quality of publications, the electoral context of studies, the main
approaches to and characteristics of the successful studies, as well as their main
strengths and challenges and compared our results with previous reviews. We
identified and analyzed 83 relevant studies, and the challenges were identified in
many areas such as process, sampling, modeling, performance evaluation, and
scientific rigor. Main findings include the low success of the most-used approach,
namely volume and sentiment analysis on Twitter, and the better results with new
approaches, such as regression methods trained with traditional polls. Finally, a
vision of future research on integrating advances in process definitions, modeling,
and evaluation is also discussed, pointing out, among others, the need for better
investigating the application of state-of-the-art machine learning approaches.
CHAPTER-1
INTRODUCTION
1.1 OVERVIEW
Scholars have suggested that the social sciences are experiencing a momentous
shift from data scarcity to data abundance [1,2], noting, however, that full access to
data is likely to be reserved for the most powerful corporate, governmental, and
academic institutions [3]. Indeed, as a growing share of our social interactions
happen on digital platforms, our capacity to predict attitudes and behaviors should
increase as well, given the scale, richness, and temporality of such data. Still, there
is a growing debate in the research community about the usefulness of social media
signals in understanding public opinion. Social scientists have identified the
problems with using “repurposed” data afforded by social media platforms, which
are often incomplete, noisy, unstructured, unrepresentative, and algorithmically
confounded [4] compared to traditional surveys, which are carefully designed,
representative, and structured.
Many recent studies have utilized social media data as a “social sensor”
aiming to predict different economic, social, and political phenomena, ranging
from election results to the box office success of Hollywood movies—with some
achieving considerable success in this endeavor. Still, most of these studies have
been primarily data-driven and largely atheoretical. Furthermore, instead of
predicting the future, these studies have made predictions about the present [5] or
the past, and have generally not examined the validity of their predictions by
comparing them with more robust types of data, such as surveys and government
censuses. Researchers have also raised concerned about the lack of
representativeness and authenticity of data harvested from social media platforms,
noting the demographic biases, the curated nature of social media representations
[6], biases introduced by sampling methods [7] and pre-processing steps [8],
platform affordances [9], and false positives introduced by bot accounts [10]. Still,
when compared to the more traditional methods, social media analytics promise to
provide an entirely new level of insight into social phenomena, by allowing
researchers unprecedented access to people’s daily conversations and their
networks of friendship and influence, across time and geographic boundaries.
Thus far, only one meta-analysis assessing the predictive power of social
media data has been published, focusing solely on Twitter-based predictions of
elections [11]. This paper, reviewing only nine studies examining the predictive
utility of Twitter messages, concluded that although social media data provides
some insights regarding electoral outcomes, it is unlikely to replace survey-based
predictions in the near future. Given the lack of scientific consensus about the
merits of social media-based predictions and the growing number of studies using
data from diverse social media sources, a more comprehensive meta-review is
needed. We thus aim to compare the merits of the most commonly used
approaches in social media-based predictions and also examine the roles of
contextual variables, including those related to the level of democracy and type of
electoral system.
The way politicians communicate with the electorate and run electoral campaigns
was reshaped by the emergence and popularization of contemporary socialmedia
(SM), such as Facebook, Twitter, and Instagram social networks (SNs). Due to the
inherent capabilities ofSM, such as the large amount of available data accessed in
real time a new research subject has emerged focusing on using the SM data to
predict election outcomes. Main findings include the low success of the most-used
approach, namely volume and sentiment analysis on Twitter, and the better results
with new approaches, such as regression methods trained with traditional polls.
Finally, a vision of future research on integrating advances in process definitions,
modeling, and evaluation is also discussed, pointing out, among others, the need
for better investigating the application of state-of the-art machine learning
approaches.
CHAPTER-2
LITERATURE SURVEY
[Link] and Public Opinion Forecasts with Social Media Data: A Meta-
Analysis
Year – 2020
In recent years, many studies have used social media data to make estimates of
electoral outcomes and public opinion. This paper reports the findings from a
meta-analysis examining the predictive power of social media data by focusing on
various sources of data and different methods of prediction; i.e., (1) sentiment
analysis, and (2) analysis of structural features. Our results, based on the data from
74 published studies, show significant variance in the accuracy of predictions,
which were on average behind the established benchmarks in traditional survey
research. In terms of the approaches used, the study shows that machine learning-
based estimates are generally superior to those derived from pre-existing lexica,
and that a combination of structural features and sentiment analyses provides the
most accurate predictions. Furthermore, our study shows some differences in the
predictive power of social media data across different levels of political democracy
and different electoral systems. We also note that since the accuracy of election
and public opinion forecasts varies depending on which statistical estimates are
used, the scientific community should aim to adopt a more standardized approach
to analyzing and reporting social media data-derived predictions in the future.
Year-2021
The way politicians communicate with the electorate and run electoral campaigns
was reshaped by the emergence and popularization of contemporary social media
(SM), such as Facebook, Twitter, and Instagram social networks (SN). Due to
inherent capabilities of SM, such as the large amount of available data accessed in
real time, a new research subject has emerged, focusing on using SM data to
predict election outcomes. Despite many studies conducted in the last decade,
results are very controversial, and many times challenged. In this context, this
work aims to investigate and summarize how research on predicting elections
based on SM data has evolved since its beginning, to outline the state of both the
art and the practice, and to identify research opportunities within this field. In
terms of method, we performed a systematic literature review analyzing the
quantity and quality of publications, the electoral context of studies, the main
approaches to and characteristics of the successful studies, as well as their main
strengths and challenges, and compared our results with previous reviews. We
identified and analyzed 83 relevant studies, and the challenges were identified in
many areas such as process, sampling, modeling, performance evaluation and
scientific rigor. Main findings include the low success of the most-used approach,
namely volume and sentiment analysis on Twitter, and the better results with new
approaches, such as regression methods trained with traditional polls. Finally, a
vision of future research on integrating advances on process definitions, modeling,
and evaluation is also discussed, pointing out, among others, the need for better
investigating the application of state-of-art machine learning approaches.
3. On the frontiers of Twitter data and sentiment analysis in election prediction: a
review
Year-2023
Election prediction using sentiment analysis is a rapidly growing field that utilizes
natural language processing and machine learning techniques to predict the
outcome of political elections by analyzing the sentiment of online conversations
and news articles. Sentiment analysis, or opinion mining, involves using text
analysis to identify and extract subjective information from text data sources. In
the context of election prediction, sentiment analysis can be used to gauge public
opinion and predict the likely winner of an election. Significant progress has been
made in election prediction in the last two decades. Yet, it becomes easier to have
its comprehensive view if it has been appropriately classified approach-wise,
citation-wise, and technology-wise. The main objective of this article is to examine
and consolidate the progress made in research about election prediction using
Twitter data. The aim is to provide a comprehensive overview of the current state-
of-the-art practices in this field while identifying potential avenues for further
research and exploration.
Year-2022
In the present information age, a wide and significant variety of social media
platforms have been developed and become an important part of modern life.
Massive amounts of user-generated data sourced from various social networking
platforms also provide new insights for businesses and governments. However, it
has become difficult to extract useful information from the vast amount of
information effectively. Sentiment analysis provides an automated method of
analyzing sentiment, emotion and opinion in written language to address this issue.
In the existing literature, a large number of scholars have worked on improving the
performance of various sentiment classifiers or applying them to various domains
using data from social networking platforms. This paper explores the challenges
that scholars have encountered and other potential problems in studying sentiment
analysis in social media. It gives insights into the goals of the sentiment analysis
task, the implementation process, and the ways in which it is utilized in various
application domains. It also provides a comparison of different studies and
highlights several challenges related to the datasets, text languages, analysis
methods and evaluation metrics. The paper contributes to the research on sentiment
analysis and can help practitioners select a suitable methodology for their
applications.
5. Design and analysis of tweet-based election models for the 2021 Mexican
legislative election
Year-2023
Modelling and forecasting real-life human behaviour using online social media is
an active endeavour of interest in politics, government, academia, and industry.
Since its creation in 2006, Twitter has been proposed as a potential laboratory that
could be used to gauge and predict social behaviour. During the last decade, the
user base of Twitter has been growing and becoming more representative of the
general population. Here we analyse this user base in the context of the 2021
Mexican Legislative Election. To do so, we use a dataset of 15 million election-
related tweets in the six months preceding election day. We explore different
election models that assign political preference to either the ruling parties or the
opposition. We find that models using data with geographical attributes determine
the results of the election with better precision and accuracy than conventional
polling methods. These results demonstrate that analysis of public online data can
outperform conventional polling methods, and that political analysis and general
forecasting would likely benefit from incorporating such data in the immediate
future. Moreover, the same Twitter dataset with geographical attributes is
positively correlated with results from official census data on population and
internet usage in Mexico. These findings suggest that we have reached a period in
time when online activity, appropriately curated, can provide an accurate
representation of offline behaviour
CHAPTER-3
PROJECT DECRIPTION
EXISTING SYSTEM
The way politicians communicate with the electorate and run electoral campaigns
was reshaped by the emergence and popularization of contemporary socialmedia
(SM), such as Facebook, Twitter, and Instagram social networks (SNs). Due to the
inherent capabilities ofSM, such as the large amount of available data accessed on
using the SM data to predict election outcomes. Main findings include the low
success of the most-used approach, namely volume and sentiment analysis on
Twitter, and the better results withnew approaches, such as regression methods
trained withtraditional polls. Finally, a vision of future research on integrating
advances in process definitions, modeling, and evaluation is also discussed,
pointing out, among others, the need for better investigating the application ofstate-
of the-art machine learning approaches.
Disadvantages
PROPOSED SYSTEM
Social media platforms such as Facebook and Twitter carry a big load of people’s
opinions about politics and leaders, which makes them a good source of
information for researchers to exploit different tasks that include election
predictions. Objective. Identify, categorize, and present a comprehensive overview
of the approaches, techniques, and tools used in election predictions on
Twitter. Method. Conducted a systematic mapping study (SMS) on election
predictions on Twitter and provided empirical evidence for the work published
between January 2010 and January 2021. Results. This research identified 787
studies related to election predictions on Twitter. 98 primary studies were selected
after defining and implementing several inclusion/exclusion criteria. The results
show that most of the studies implemented sentiment analysis (SA) followed by
volume-based and social network analysis (SNA) approaches.
ADVANTAGES
The majority of the studies employed supervised learning techniques,
subsequently, lexicon-based approach SA, volume-based, and unsupervised
learning. Besides this, 18 types of dictionaries were identified.
Elections of 28 countries were analyzed, mainly USA (28%) and Indian
(25%) elections. Furthermore, the results revealed that 50% of the primary
studies used English tweets.
The demographic data showed that academic organizations and conference
venues are the most active.
The evolution of the work published in the past 11 years shows that most of
the studies employed SA. The implementation of SNA techniques is lower
as compared to SA. Appropriate political labelled datasets are not available,
especially in languages other than English. Deep learning needs to be
employed in this domain to get better predictions.
CHAPTER-3
SYSTEM CONFIGURATION
CHAPTER -4
SYSTEM IMPLEMENTATION
MODULES DESCRIPTION
Data Sources
A second research gap in the literature is that it is currently not known whether
one platform or data source yields more accurate predictions than others. Because
of the public nature of posts on Twitter, most studies have utilized Twitter data to
predict public opinion, followed by the use of Facebook, forums, blogs, and
YouTube. Given that each platform suffers from its own set of algorithmic
confounds, privacy constraints, and post restrictions, it is unknown whether
consolidating data from multiple platforms would have any advantage over
predictions from a single data source. Most studies have utilized data from a single
platform, while very few use data from multiple platforms.
Predictors
Social media predictors are first categorized into two types: sentiment vs.
structure. Sentiment predictors refer to the sentiment/preference extracted from the
users’ text-based social media posts. Sentiment scores are obtained by using (1)
lexicon-based approaches, which use dictionaries of weighted words, in which
accuracy is highly dependent on the quality and relevance of the lexical resources
to the domain it is being applied to; and (2) (supervised) machine learning-based
models, which predict the sentiment score of a piece of text based on the regression
weights that it learns from example data (a corpus) labeled by a human.
CHAPTER-5
SOFTWARE ENVIRONMENT
PYHTON SOFTWARE DEVELOPMENT
The Python interpreter is easily extended with new functions and data types
implemented in C or C++ (or other languages callable from C). Python is also
suitable as an extension language for customizable applications.
This tutorial introduces the reader informally to the basic concepts and features of
the Python language and system. It helps to have a Python interpreter handy for
hands-on experience, but all examples are self-contained, so the tutorial can be
read off-line as well.
For a description of standard objects and modules, see The Python Standard
Library. The Python Language Reference gives a more formal definition of the
language. To write extensions in C or C++, read Extending and Embedding the
Python Interpreter and Python/C API Reference Manual. There are also several
books covering Python in depth.
What is Python
It is used for:
Python Install
To check if you have python installed on a Windows PC, search in the start bar for
Python or run the following on the Command Line ([Link]):
To check if you have python installed on a Linux or Mac, then on linux open the
command line or on Mac open the Terminal and type:
python --version
If you find that you do not have Python installed on your computer, then you can
download it for free from the following website: [Link]
Python Quickstart
The way to run a python file is like this on the command line:
Let's write our first Python file, called [Link], which can be done in any
text editor.
The Python Command Line
To test a short amount of code in python sometimes it is quickest and easiest not to
write the code in a file. This is made possible because Python can be run as a
command line itself.
C:\Users\Your Name>python
Or, if the "python" command did not work, you can try "py":
C:\Users\Your Name>py
From there you can write any python, including our hello world example from
earlier in the tutorial:
C:\Users\YourName>python
Argument Passing
When known to the interpreter, the script name and additional arguments thereafter
are turned into a list of strings and assigned to the argv variable in the sys module.
You can access this list by executing import sys. The length of the list is at least
one; when no script and no arguments are given, [Link][0] is an empty string.
When the script name is given as '-' (meaning standard input), [Link][0] is set
to '-'. When -c command is used, [Link][0] is set to '-c'. When -m module is
used, [Link][0] is set to the full name of the located module. Options found
after -c command or -m module are not consumed by the Python interpreter’s
option processing but left in [Link] for the command or module to handle.
Interactive Mode
When commands are read from a tty, the interpreter is said to be in interactive
mode. In this mode it prompts for the next command with the primary prompt,
usually three greater-than signs (>>>); for continuation lines it prompts with
the secondary prompt, by deFraud three dots (...). The interpreter prints a welcome
message stating its version number and a copyright notice before printing the first
prompt:
$ python3.10
the_world_is_flat = True
if the_world_is_flat:
To declare an encoding other than the deFraud one, a special comment line should
be added as the first line of the file. The syntax is as follows:
For example, to declare that Windows-1252 encoding is to be used, the first line of
your source code file should be:
One exception to the first line rule is when the source code starts with a UNIX
“shebang” line. In this case, the encoding declaration should be added as the
second line of the file. For example:
#!/usr/bin/env python3
In the following examples, input and output are distinguished by the presence or
absence of prompts (>>> and …): to repeat the example, you must type everything
after the prompt, when the prompt appears; lines that do not begin with a prompt
are output from the interpreter. Note that a secondary prompt on a line by itself in
an example means you must type a blank line; this is used to end a multi-line
command.
You can toggle the display of prompts and output by clicking on >>> in the upper-
right corner of an example box. If you hide the prompts and output for an example,
then you can easily copy and paste the input lines into your interpreter.
Many of the examples in this manual, even those entered at the interactive prompt,
include comments. Comments in Python start with the hash character, #, and
extend to the end of the physical line. A comment may appear at the start of a line
or following whitespace or code, but not within a string literal. A hash character
within a string literal is just a hash character. Since comments are to clarify code
and are not interpreted by Python, they may be omitted when typing in examples.
Some examples:
Let’s try some simple Python commands. Start the interpreter and wait for the
primary prompt, >>>. (It shouldn’t take long.)
Numbers
The interpreter acts as a simple calculator: you can type an expression at it and it
will write the value. Expression syntax is straightforward: the
operators +, -, * and / work just like in most other languages (for example, Pascal
or C); parentheses (()) can be used for grouping. For example:
>>>
>>> 2 + 2
>>> 50 - 5*6
20
5.0
1.6
The integer numbers (e.g. 2, 4, 20) have type int, the ones with a fractional part
(e.g. 5.0, 1.6) have type float. We will see more about numeric types later in the
tutorial.
Division (/) always returns a float. To do floor division and get an integer result
(discarding any fractional result) you can use the // operator; to calculate the
remainder you can use %:
5.666666666666667
17
>>> 5 ** 2 # 5 squared
25
128
The equal sign (=) is used to assign a value to a variable. Afterwards, no result is
displayed before the next interactive prompt:
>>> width = 20
>>> height = 5 * 9
900
If a variable is not “defined” (assigned a value), trying to use it will give you an
error:
There is full support for floating point; operators with mixed type operands convert
the integer operand to floating point:
>>> 4 * 3.75 - 1
14.0
In interactive mode, the last printed expression is assigned to the variable _. This
means that when you are using Python as a desk calculator, it is somewhat easier to
continue calculations, for example:
12.5625
>>> price + _
113.0625
>>> round(_, 2)
113.06
This variable should be treated as read-only by the user. Don’t explicitly assign a
value to it — you would create an independent local variable with the same name
masking the built-in variable with its magic behavior.
In addition to int and float, Python supports other types of numbers, such
as Decimal and Fraction. Python also has built-in support for complex numbers,
and uses the j or J suffix to indicate the imaginary part (e.g. 3+5j).
Strings
Besides numbers, Python can also manipulate strings, which can be expressed in
several ways. They can be enclosed in single quotes ('...') or double quotes ("...")
with the same result 2. \ can be used to escape quotes:
>>>
'spam eggs'
"doesn't"
"doesn't"
First line.
Second line.
>>>
ame
C:\some\name
String literals can span multiple lines. One way is using triple-
quotes: """...""" or '''...'''. End of lines are automatically included in the string, but
it’s possible to prevent this by adding a \ at the end of the line. The following
example:
print("""\
""")
produces the following output (note that the initial newline is not included):
Strings can be concatenated (glued together) with the + operator, and repeated
with *:
>>>
>>> # 3 times 'un', followed by 'ium'
'unununium'
Two or more string literals (i.e. the ones enclosed between quotes) next to each
other are automatically concatenated.
>>>
'Python'
This feature is particularly useful when you want to break long strings:
>>>
>>> text
This only works with two literals though, not with variables or expressions:
>>>
('un' * 3) 'ium'
>>>
'Python'
Strings can be indexed (subscripted), with the first character having index 0. There
is no separate character type; a character is simply a string of size one:
>>>
'P'
Indices may also be negative numbers, to start counting from the right:
'n'
'o'
>>> word[-6]
'P'
Note that since -0 is the same as 0, negative indices start from -1.
>>>
'Py'
'tho'
Slice indices have useful deFraud s; an omitted first index deFraud s to zero, an
omitted second index deFraud s to the size of the string being sliced.
'Py'
>>> word[4:] # characters from position 4 (included) to the end
'on'
'on'
Note how the start is always included, and the end always excluded. This makes
sure that s[:i] + s[i:] is always equal to s:
'Python'
'Python'
The first row of numbers gives the position of the indices 0…6 in the string; the
second row gives the corresponding negative indices. The slice from i to j consists
of all characters between the edges labeled i and j, respectively.
For non-negative indices, the length of a slice is the difference of the indices, if
both are within bounds. For example, the length of word.
CHAPTER -4
PROJECT DESCRIPTION
INPUT DESIGN
Twitter is the SN used in most of the studies (84%), and in many of them (75%), it
was the only SN used as input. However, Twitter is not a good sample, even
considering only SM users, due to its having very few active users (326 million),
relative to other SN, such as Facebook (2.6 billion) and Instagram (1,1 billion),
according to a 2020 report [36]. Despite these data, it is hard to find a discussion
about why studies focused on Twitter. After analyzing the API of these SNs [41],
[50], we hypothesized that Twitter was chosen because it is easier for researchers
to collect data on this platform. For example, starting on August 2018, the approval
process to gather data from Facebook and Instagram consisted of developing and
deploying a fully functional system, creating and publishing a privacy policy and
terms of use, recording a video showing all the functionalities related to Facebook
and Instagram data collection, creating test accounts allowing Facebook employees
to test the system and, in many cases, sending formal documentation of an
institution responsible for the system. By contrast, in August 2019, the Twitter
approval process only involved completing a form with information about the
system.
OBJECTIVES
In most studies, many data collection choices were arbitrary, such as the data
collection period, which usually varied from 3 days to 3 months before elections,
and the keywords used for open search on volume/sentiment approaches. This
created many problems, such as those presented by [30], in which the performance
was too unstable because it depended strongly on such parameterizations, and
unintentional data dredging could occur, due to post hoc analysis. Also, it
reinforces the argument presented by Jungherr [22] who, after replicating the
seminal study of Tumasjan [5], argued that “the results are contingent on arbitrary
choices of the authors,” and indicated that simply including one more party or day
of collection would greatly change the results.
OUTPUT DESIGN
Modelling and forecasting real-life human behaviour using online social media is
an active endeavour of interest in politics, government, academia, and industry.
Since its creation in 2006, Twitter has been proposed as a potential laboratory that
could be used to gauge and predict social behaviour. During the last decade, the
user base of Twitter has been growing and becoming more representative of the
general population. Here we analyse this user base in the context of the 2021
Mexican Legislative Election. To do so, we use a dataset of 15 million election-
related tweets in the six months preceding election day. We explore different
election models that assign political preference to either the ruling parties or the
opposition. We find that models using data with geographical attributes determine
the results of the election with better precision and accuracy than conventional
polling methods. These results demonstrate that analysis of public online data can
outperform conventional polling methods, and that political analysis and general
forecasting would likely benefit from incorporating such data in the immediate
future. Moreover, the same Twitter dataset with geographical attributes is
positively correlated with results from official census data on population and
internet usage in Mexico. These findings suggest that we have reached a period in
time when online activity, appropriately curated, can provide an accurate
representation of offline behaviour.
FEASIBILITY STUDY
The feasibility of the project is analyzed in this phase and business proposal is
put forth with a very general plan for the project and some cost estimates. During
system analysis the feasibility study of the proposed system is to be carried out.
This is to ensure that the proposed system is not a burden to the company. For
feasibility analysis, some understanding of the major requirements for the system
is essential.
ECONOMICAL FEASIBILITY
TECHNICAL FEASIBILITY
SOCIAL FEASIBILITY
ECONOMICAL FEASIBILITY
This study is carried out to check the economic impact that the system will
have on the organization. The amount of fund that the company can pour into the
research and development of the system is limited.
TECHNICAL FEASIBILITY
This study is carried out to check the technical feasibility, that is, the
technical requirements of the system. Any system developed must not have a high
demand on the available technical resources.
This will lead to high demands on the available technical resources. This
will lead to high demands being placed on the client. The developed system must
have a modest requirement, as only minimal or null changes are required for
implementing this system.
SOCIAL FEASIBILITY
The aspect of study is to check the level of acceptance of the system by the
user. This includes the process of training the user to use the system efficiently.
The user must not feel threatened by the system, instead must accept it as a
necessity.
The level of acceptance by the users solely depends on the methods that are
employed to educate the user about the system and to make him familiar with it.
CHAPTER-5
SYSTEM TESTING
The purpose of testing is to discover errors. Testing is the process of trying
to discover every conceivable fault or weakness in a work product. It provides a
way to check the functionality of components, sub assemblies, assemblies and/or a
finished product It is the process of exercising software with the intent of ensuring
that the Software system meets its requirements and user expectations and does not
fail in an unacceptable manner. There are various types of test. Each test type
addresses a specific testing requirement.
TYPES OF TESTS
Unit Testing
Unit testing involves the design of test cases that validate that the internal
program logic is functioning properly, and that program inputs produce valid
outputs. All decision branches and internal code flow should be validated. It is the
testing of individual software units of the application .it is done after the
completion of an individual unit before integration. This is a structural testing, that
relies on knowledge of its construction and is invasive. Unit tests perform basic
tests at component level and test a specific business process, application, and/or
system configuration. Unit tests ensure that each unique path of a business process
performs accurately to the documented specifications and contains clearly defined
inputs and expected results.
Integration Testing
Testing is event driven and is more concerned with the basic outcome of
screens or fields. Integration tests demonstrate that although the components were
individually satisfaction, as shown by successfully unit testing, the combination of
components is correct and consistent.
Functional Testing
exercised.
Before functional testing is complete, additional tests are identified and the
effective value of current tests is determined.
System Testing
System testing ensures that the entire integrated software system meets
requirements. It tests a configuration to ensure known and predictable results. An
example of system testing is the configuration oriented system integration test.
System testing is based on process descriptions and flows, emphasizing pre-driven
process links and integration points.
Black box tests, as most other kinds of tests, must be written from a
definitive source document, such as specification or requirements document, such
as specification or requirements document. It is a testing in which the software
under test is treated, as a black box .you cannot “see” into it. The test provides
inputs and responds to outputs without considering how the software works.
Unit Testing
Unit testing is usually conducted as part of a combined code and unit test
phase of the software lifecycle, although it is not uncommon for coding and unit
testing to be conducted as two distinct phases.
Test objectives
Test Results: All the test cases mentioned above passed successfully. No defects
encountered.
Acceptance Testing
User Acceptance Testing is a critical phase of any project and requires
significant participation by the end user. It also ensures that the system meets the
functional requirements.
Conclusion
Digital traces of human behavior now enable the testing of social science theories
in new ways, and are also likely to pave the way for new theories in social science.
Both in terms of its nature and scale, user-generated social media data is
dramatically different from the existing market information regime data, which is
constructed to serve a specific purpose and typically relies on smaller samples of
unconnected citizens [25]. As new studies and new methods for social media-based
predictions proliferate, it is important to consider both how data are generated, and
how they are analyzed, and consolidate a validated framework for data analysis. As
Kuhn argued, there is a close, intimate link between the scientific puzzle and the
methodology designed to solve it [68]. Social media analytics will not replace
survey-based public opinion studies, but they do offer additional data sources and
tools for improving our understanding of public opinion and political behavior,
while at the same time changing the norms around voter engagement, mobilization,
and preference elicitation. Among other things, we see the potential of social
media data to provide deeper insights regarding the opinion dynamics and the role
of opinion leaders in the process of public opinion formation and change.
Furthermore, the problems that plague our digital ecologies today, such as filter
bubbles, disinformation, incivility, and hate speech, can all be better understood
using social media data and computational methods, than with survey-based
research. Still, we note a general lack of theoretically-informed work in the field—
most studies report predictions without paying sufficient attention to the
underlying mechanisms and processes. We are also concerned about the future
availability of social media data, as citizens switch to more private, encrypted
instant messaging apps and social media companies further restrict access to user
data because of legal, commercial, and privacy concerns. Thus, policymakers
should create legal and regulatory environments that will promote access to social
media data for the research community in ways that guarantee strong privacy
protection rights for the users.
SOURCE CODE
import face_recognition as fr
import cv2
import numpy as np
import os
path = "./train/"
known_names = []
known_name_encodings = []
images = [Link](path)
for _ in images:
image = fr.load_image_file(path + _)
image_path = path + _
encoding = fr.face_encodings(image)[0]
known_name_encodings.append(encoding)
known_names.append([Link]([Link](image_path))
[0].capitalize())
print(known_names)
test_image = "./test/[Link]"
image = [Link](test_image)
face_locations = fr.face_locations(image)
name = ""
best_match = [Link](face_distances)
if matches[best_match]:
name = known_names[best_match]
font = cv2.FONT_HERSHEY_DUPLEX
[Link](image, name, (left + 6, bottom - 6), font, 1.0, (255, 255, 255), 1)
[Link]("Result", image)
[Link]("./[Link]", image)
[Link](0)
[Link]()
import cv2
import numpy as np
# import the file where data is
import npwriter
cap = [Link](0)
classifier = [Link](r"C:\Users\LENOVO\Desktop\FACE\dataset\
haarcascade_frontalface_default.xml")
f_list = []
while True:
reverse = True)
faces = faces[:1]
if len(faces) == 1:
x, y, w, h = face
[Link]("face", im_face)
if not ret:
continue
[Link]("full", frame)
key = [Link](1)
# this will break the execution of the program
break
if len(faces) == 1:
f_list.append(gray_face.reshape(-1))
else:
if len(f_list) == 10:
break
# declared in npwriter
[Link](name, [Link](f_list))
[Link]()
[Link]()
import pandas as pd
import math
import numpy as np
# import sklearn
# import dataset
df = pd.read_csv('/content/county_census_and_election_result.csv')
[Link]()
[Link]
# select the rows in the dataset that have the voting labels (years 2008, 2012, 2016,
2020)
# reset index
selected_df = selected_df.reset_index(drop=True)
selected_df
print(selected_df.isnull().[Link]())
**Drop**:
"""
# Feature selection
more_selected_df.head()
more_selected_df = more_selected_df.fillna(method='ffill')
more_selected_df = more_selected_df.fillna(method='bfill')
# check if there is any NaN values
print(more_selected_df.isnull().[Link]())
more_selected_df.head()
"""**Note:** divide the dataset into 2 datasets. Dataset 1 is within 2008, 2012,
2016 for machine learning modeling. Dataset 2 is of 2020 for testing the prediction
power of the trained and validated data
"""
# select rows with the years 2008, 2012, 2016 presidential election
df1 = [Link](columns=['year'])
# reset index
df1 = df1.reset_index(drop=True)
[Link]()
# normalization + divide the dataset into input features and output labels
X_df1 = df1[list([Link][:-1])]
y_df1 = df1['winner'].to_numpy().reshape(-1, 1)
# data normalization
X_scaler = [Link](feature_range=(-1,1))
X_df1 = X_scaler.fit_transform(X_df1)
y_scaler = [Link](feature_range=(-1,1))
y_df1 = y_scaler.fit_transform(y_df1)
# partition into training/test/validation
df2 = more_selected_df.loc[more_selected_df['year'].isin([2020])]
df2 = [Link](columns=['year'])
# reset index
df2 = df2.reset_index(drop=True)
[Link]()
X_df2 = df2[list([Link][:-1])]
y_df2 = df2['winner'].to_numpy().reshape(-1, 1)
# data normalization
X_scaler = [Link](feature_range=(-1,1))
X_df2 = X_scaler.fit_transform(X_df2)
y_scaler = [Link](feature_range=(-1,1))
y_df2 = y_scaler.fit_transform(y_df2)
[Link](X_train, y_train)
# predict on validate set
y_test_pred = [Link](X_test)
y_train_pred = [Link](X_train)
f1 = metrics.f1_score(y_test, y_test_pred)
[Link](X_train, y_train)
y_test_pred = [Link](X_test)
y_train_pred = [Link](X_train)
f1 = metrics.f1_score(y_test, y_test_pred)
"""
clf1_pred_y_2020 = [Link](X_df2)
f1 = metrics.f1_score(y_df2, clf1_pred_y_2020)
demo_count1 = repu_count1 = 0
for n in clf1_pred_y_2020:
if n == -1:
demo_count1 += 1
else:
repu_count1 += 1
"""## 3.2. Decision Tree With Entropy Criterion"""
clf2_pred_y_2020 = [Link](X_df2)
f1 = metrics.f1_score(y_df2, clf2_pred_y_2020)
demo_count2 = repu_count2 = 0
for n in clf2_pred_y_2020:
if n == -1:
demo_count2 += 1
else:
repu_count2 += 1
[Link](figsize=(20, 5))
1. King, G. Ensuring the data-rich future of the social sciences. Science 2011, 331,
719–721. [Google Scholar] [CrossRef] [PubMed]
2. Lazer, D.; Pentland, A.S.; Adamic, L.; Aral, S.; Barabasi, A.L.; Brewer, D.;
Gutmann, M. Life in the network: The coming age of computational social
science. Science 2009, 323, 721–723. [Google Scholar] [CrossRef][Green Version]
3. boyd, D.; Crawford, K. Critical questions for big data: Provocations for a cultural,
technological, and scholarly phenomenon. Inf. Commun. Soc. 2012, 15, 662–679.
[Google Scholar] [CrossRef]
4. Salganik, M.J. Bit by Bit: Social Research in the Digital Age; Princeton University
Press: Princeton, NJ, USA, 2018. [Google Scholar]
5. Choi, H.; Varian, H. Predicting the present with Google Trends. Econ.
Rec. 2012, 88, 2–9. [Google Scholar] [CrossRef]
6. Manovich, L. Trending: The promises and the challenges of big social data.
In Debates in the Digital Humanities; Gold, M.K., Ed.; The University of
Minnesota Press: Minneapolis, MN, USA, 2011; pp. 460–475. [Google Scholar]
7. González-Bailón, S.; Wang, N.; Rivero, A.; Borge-Holthoefer, J.; Moreno, Y.
Assessing the bias in samples of large online networks. Soc. Netw. 2014, 38, 16–
27. [Google Scholar] [CrossRef]
8. Jungherr, A.; Jürgens, P.; Schoen, H. Why the Pirate party won the German
election of 2009 or the trouble with predictions: A response to Tumasjan, A.,
Sprenger, T.O., Sander, P.G., & Welpe, I.M. “Predicting elections with Twitter:
What 140 characters reveal about political sentiment.”. Soc. Sci. Comput.
Rev. 2012, 30, 229–234. [Google Scholar] [CrossRef]
9. Bucher, T.; Helmond, A. The Affordances of Social Media Platforms. In The
SAGE Handbook of Social Media; Burgess, J., Marwick, A., Poell, T., Eds.; Sage
Publications: New York, NY, USA, 2018; pp. 233–253. [Google Scholar]
10. Ratkiewicz, J.; Conover, M.; Meiss, M.R.; Gonçalves, B.; Flammini, A.; Menczer,
F. Detecting and Tracking Political Abuse in Social Media. In Proceedings of the
20th international conference companion on World Wide Web, Hyderabad, India,
28 March–1 April 2011. [Google Scholar]
11. Gayo-Avello, D. A meta-analysis of state-of-the-art electoral prediction from
Twitter data. Soc. Sci. Comput. Rev. 2013, 31, 649–679. [Google Scholar]
[CrossRef][Green Version]
12. Bourdieu, P. Public opinion does not exist. In Communication and Class Strugzle;
Mattelart, A., Siegelaub., S., Eds.; International General/IMMRC: New York, NY,
USA, 1979; pp. 124–310. [Google Scholar]
13. Lazarsfeld, P.F. Public opinion and the classical tradition. Public Opin.
Q. 1957, 21, 39–53. [Google Scholar] [CrossRef]
14. Katz, E. The two-step flow of communication: An up-to-date report on an
hypothesis. Public Opin. Q. 1957, 21, 61–78. [Google Scholar] [CrossRef]
15. Katz, E.; Lazarsfeld, P.F. Personal Influence: The Part Played by People in the
Flow of Mass Communications; Free Press: New York, NY, USA, 1955. [Google
Scholar]
16. Ceron, A.; Curini, L.; Iacus, S.M.; Porro, G. Every tweet counts? How sentiment
analysis of social media can improve our knowledge of citizens’ political
preferences with an application to Italy and France. New Media Soc. 2014, 16,
340–358. [Google Scholar] [CrossRef]
17. Espeland, W.N.; Sauder, M. Rankings and reactivity: How public measures
recreate social worlds. Am. J. Sociol. 2007, 113, 1–40. [Google Scholar]
[CrossRef][Green Version]
18. Berinsky, A.J. The two faces of public opinion. Am. J. Political Sci. 1999, 43,
1209–1230. [Google Scholar] [CrossRef]
19. Blumer, H. Public opinion and public opinion polling. Am. Sociol. Rev. 1948, 13,
542–549. [Google Scholar] [CrossRef]
20. Anstead, N.; O’Loughlin, B. Social media analysis and public opinion: The 2010
UK general election. J. Comput. Mediat. Commun. 2015, 37, 204–220. [Google
Scholar] [CrossRef][Green Version]