0% encontró este documento útil (0 votos)
85 vistas20 páginas

OSINT en Ciberseguridad: Análisis en Colombia

Este documento describe cómo la inteligencia de fuentes abiertas (OSINT) puede apoyar las operaciones de ciberseguridad mediante la obtención y análisis de información sobre posibles adversarios. Presenta diferentes herramientas de OSINT y cómo pueden usarse para realizar tareas de ciberinteligencia. También desarrolla modelos de aprendizaje automático para realizar análisis de sentimientos sobre la información recolectada de redes sociales con el fin de comprender la motivación de los adversarios y definir estrategias de ciberdefensa

Cargado por

Gulay Khanum
Derechos de autor
© All Rights Reserved
Nos tomamos en serio los derechos de los contenidos. Si sospechas que se trata de tu contenido, reclámalo aquí.
Formatos disponibles
Descarga como PDF, TXT o lee en línea desde Scribd
0% encontró este documento útil (0 votos)
85 vistas20 páginas

OSINT en Ciberseguridad: Análisis en Colombia

Este documento describe cómo la inteligencia de fuentes abiertas (OSINT) puede apoyar las operaciones de ciberseguridad mediante la obtención y análisis de información sobre posibles adversarios. Presenta diferentes herramientas de OSINT y cómo pueden usarse para realizar tareas de ciberinteligencia. También desarrolla modelos de aprendizaje automático para realizar análisis de sentimientos sobre la información recolectada de redes sociales con el fin de comprender la motivación de los adversarios y definir estrategias de ciberdefensa

Cargado por

Gulay Khanum
Derechos de autor
© All Rights Reserved
Nos tomamos en serio los derechos de los contenidos. Si sospechas que se trata de tu contenido, reclámalo aquí.
Formatos disponibles
Descarga como PDF, TXT o lee en línea desde Scribd

Revista Vínculos

[Link]

A+T Actualidad Tecnologica

Inteligencia de fuentes abierta (OSINT) para operaciones de


ciberseguridad.
“Aplicación de OSINT en un contexto colombiano y análisis de
sentimientos”
Open source intelligence (OSINT) as support of cybersecurity operations.
“Use of OSINT in a colombian context and sentiment Analysis”

Ricardo Andrés Pinto Rico1 Martin José Hernández Medina2 Cristian Camilo Pinzón Hernández3
Daniel Orlando Díaz López4 Juan Carlos Camilo García Ruíz5

Para citar este artículo: R. A. Pinto, M. J. Hernández, C. C. Pinzón, D. O. Díaz y J. C. C. García, “Inteligencia de
fuentes abierta (OSINT) para operaciones de ciberseguridad. “Aplicación de OSINT en un contexto colombiano
y análisis de sentimientos””. Revista Vínculos: Ciencia, Tecnología y Sociedad, vol 15, n° 2, julio-diciembre 2018,
195-214. DOI: [Link]

Recibido: 21-04-2018 / Aprobado: 03-06-2018

Resumen de sentimientos sobre dicha información, con el fin de


La Inteligencia de fuentes abiertas (OSINT) es una rama saber la posición del adversario respecto a determi-
de la ciber inteligencia usada para obtener y analizar nados temas y así entender la motivación que puede
información relacionada a posibles adversarios, para tener, lo cual permite definir estrategias de ciberdefensa
que esta pueda apoyar evaluaciones de riesgo y ayudar apropiadas. Finalmente, algunos desafíos relacionados
a prevenir afectaciones contra activos críticos. Este ar- a la aplicación de técnicas OSINT también son iden-
tículo presenta una investigación acerca de diferentes tificados y descritos al respecto de su aplicación por
tecnologías OSINT y como estas pueden ser usadas para agencias de seguridad del estado.
desarrollar tareas de ciber inteligencia de una nación.
Un conjunto de transformadas apropiadas para un Palabras clave: Análisis de sentimientos, aprendi-
contexto colombiano son presentadas y contribuidas a zaje automático, ciber inteligencia, ciencia de da-
la comunidad, permitiendo a organismos de seguridad tos, inteligencia de fuentes abiertas, perfilamiento
adelantar procesos de recolección de información de de adversarios.
fuentes abiertas colombianas. Sin embargo, el verda-
dero aprovechamiento de la información recolectada Abstract
se da mediante la implementación de tres modelos de Open source intelligence (OSINT) is a cyber-intelli-
aprendizaje automático usados para desarrollar análisis gence branch used to obtain and analyze information

1. Estudiante Ingeniería de Sistemas. Escuela Colombiana de Ingeniería Julio Garavito. Correo electrónico: [Link]@[Link]
2. Estudiante Ingeniería de Sistemas. Escuela Colombiana de Ingeniería Julio Garavito. Correo electrónico: [Link]@[Link]
3. Estudiante Ingeniería de Sistemas. Escuela Colombiana de Ingeniería Julio Garavito. Correo electrónico: [Link]@[Link]
4. Doctor en Informática; profesor asistente, Escuela Colombiana de Ingeniería Julio Garavito. Correo electrónico: [Link]@[Link]
5. Especialista en Seguridad Informática; jefe División de Ciberdefensa, Dirección de Cibernética Naval. Armada Nacional. Correo electrónico: juan.
garciaru@[Link]

[ 195 ]
Vínculos
ISSN 1794-211X • e-ISSN 2322-939X • Vol 15, N° 2 (julio-diciembre 2018). pp. 195-214. Universidad Distrital Francisco José de Caldas-Facultad Tecnológica.
Inteligencia de fuentes abierta (OSINT) para operaciones de ciberseguridad.

related to potential adversaries, so it can support risk The more high-quality information is acquired, the more
assessments and help to prevent damages against cri- conclusions can be established in different processes.
tical assets. This paper presents a research about diffe- Since there is a vast amount of information on internet
rent OSINT technologies and how these can be used which does not have a good quality or is wrong, many
to perform cyber intelligence tasks of a nation. A set of false positives can arise in OSINT. An example of a
transforms addressed to the Colombian context are pre- misguided data collection process can occur when the
sented, which were implemented and contributed to the collected data belongs to a namesake, i.e. someone
community allowing to the law enforcement agencies to with the same name, with the legitimate target.
develop information gathering process from Colombian Cyber intelligence should not be confused with
open sources. However, the real use of the information criminal analysis, the latter seeks to obtain the "in-
is given by the implementation of three machine learning formation necessary for the prosecution and repres-
models used to perform sentiment analysis over this in- sion of crimes (evidence)" i.e. a crime has already
formation, in order to know the opinion of the adversary occurred. Instead the cyber intelligence applies in
about certain topic and understand his motivation and, the pre-criminal scope of the threat and the risk [1].
in this way, define proper cyber defense strategies. Fina- OSINT is generally used by law enforcement agen-
lly, some challenges related to the application of OSINT cies, companies or organizations with high value
techniques are identified and described regarding its use assets and even cybercriminals. Law enforcement
by state security agencies. agencies can use OSINT for example to research
around the triangle of three aspects (Motive, Oppor-
Keywords: Sentiment analysis, machine learning, cyber tunity, Means) that must be established for a crime.
intelligence, data science, open source intelligence, ad- In this way it could be possible to find out about the
versary profiling. motives (the reason to develop an attack) behind an
adversary, the opportunities (how much the asset is
1. Introducción exposed or vulnerable) offered by the victim and the
Means (capabilities required to perform the attack)
Having information has meant ‘power’ from a long that the adversary holds.
time ago, in this way information from open sources So, this paper aims the following objectives:
has changed this paradigm because it is not initia-
lly restricted and can be accessed for many. So, 1. Identify different OSINT tools, most of them
the ‘power’ behind open source information is not open source, that can be useful in cyber inte-
related with the property (ownership) but with the lligence labors.
knowledge about how to use it. Information from 2. Implement data collectors (transforms) which
open sources can be recognized because is freely allows to use an OSINT tool in a Colombian
available in social networks, search engines, forums, context.
photographs, wikis, online libraries, conferences, 3. Develop models for the processing of informa-
metadata, etc. tion collected from social networks to realize a
This paper will address cyber intelligence research sentiment analysis.
conducted from open source intelligence, named
OSINT (Open Source Intelligence). Through OSINT The achievement of these objectives will allow to
is possible to collect and process all types of infor- build strategic, tactical and operational intelligence
mation, which can be used for tasks such as conduc- products. These products will be supported by all
ting security profiling, psychological studies, market the collected and processed information regarding
trend evaluations, security audits, review of digital a target, such as: full names, identification, address,
and online reputation of a target, among others. user names, emails, etc.

[ 196 ]
Vínculos
ISSN 1794-211X • e-ISSN 2322-939X • Vol 15, N° 2 (julio-diciembre 2018). pp. 195-214. Universidad Distrital Francisco José de Caldas-Facultad Tecnológica.
Ricardo Andrés Pinto Rico; Martin José Hernández Medina; Cristian Camilo Pinzón Hernández; Daniel Orlando Díaz López; Juan Carlos Camilo García Ruíz

This paper is composed as follows. Section 2 offers Another very useful tool for domain data collection is
a review of OSINT tools, explains the OSINT ar- ReconNG [3]. This tool has a module for the scanning
chitecture and presents the transforms developed of a domain that collect a large amount of information
for the Colombian context. Section 2 also shows about it. ReconNG can be used in a similar way to
at last a demonstration of use of transforms and The Harvester obtaining emails from an organization,
OSINT tools to develop cybersecurity labors. Then, but as added value, ReconNG obtains information
Section 3 presents three different machine learning about the domain location, physical location of the
models used to make sentiment analysis over co- domain server, name of the administrator and ano-
llected data. Next, Section 4 makes a reflection on ther “Who Is” data. This information can be used to
OSINT tools and techniques used Cyber security profile an organization and its members.
teams belonging to law enforcement agencies. When the target is a person, the information provi-
Finally, some conclusions and future works are ded by ReconNG could allow to identify coworkers
included. from the emails in the same domain. If the target is
an organization, the information provided by Re-
2. Gathering information from open sour- conNG could deliver data about the employees and
ces in a Colombian context then perform a new OSINT iteration, but this time
over one of the found employees.
The collection of public information is a process
that can be done in various ways using manual or Table 1. OSINT tools.
automatic processes. This section makes a review
Tool License Input data Platform
of different OSINT tools which can support am au-
tomatic collection process and list a set of elements Domain, username,
url, email, image, Linux, Windows,
Maltego MIT
(transforms) that can be used to collect information DNS, IP, Location, Mac
phrase, etc.
e.g. emails, documents, domain, etc. in the Colom-
Url and type of file
bian context. Metagoofil GNU 2.0 (extension), limit of Linux, Windows
results, etc.
2.1 OSINT tools overview Type of file,
The Foca GPL 3.0 domain, search Linux, Windows
engine, etc.
Cyber intelligence labors over open information Ip, country,
sources can be developed with several tools. Table port, keywords,
Shodan MIT Web
hostname, DNS,
1 shows the most representative open source tools protocol, url, etc.
that can help in the development of researching and Domain, number
The of desired results, Linux, Windows,
profiling. Most of these tools are complementary Harvester GPL 2.0 sources to search Mac
between them. (Google, Bing, etc.)
The first tool that will be mentioned is “The Harves- Domain, api-key,
domain, special
ter”[2], which is focused to find emails addressed Recon-NG GNU 2.0 Linux
modules for
from a domain name. The emails are searched in gathering, etc.
servers such as Google, Bing, LinkedIn, etc. Emails Domain, username, Linux, Windows
Spiderfoot GPL 2.0 files, url, email, etc.
can be useful as possible points of entry to the infor-
Personal
mation of an organization e.g. identifying an email information (name,
Intel phone number,
accounts with easy to guess passwords or validating Techniques N/A identification Web
the infection of accounts using cyberattacks that document, social
network profile)
have as starting point a malicious email (phishing,
trojans, spam, spear, whaling). Source: Own.

[ 197 ]
Vínculos
ISSN 1794-211X • e-ISSN 2322-939X • Vol 15, N° 2 (julio-diciembre 2018). pp. 195-214. Universidad Distrital Francisco José de Caldas-Facultad Tecnológica.
Inteligencia de fuentes abierta (OSINT) para operaciones de ciberseguridad.

FOCA [4] allows to find metadata in Open Office, Mi- Relationship between Entities are called "Transforms"
crosoft Office and PDF documents and correlate it to which develop the process of moving from one type
obtain relevant information. Metadata could indicate of data "X" to another type of data "Y". For example,
the file creation and modification date, in addition with a national ID number (Entity X) it is possible to
to the name of the user who made those changes. obtain the full name of a person (Entity Y), or with
Shodan [4] is a search engine used to find servers, an email address (Entity Z) it is possible to obtain
routers or any type of device reachable by internet the social network account (Entity A).
through protocols such as HTTP, SSH, Telnet, or Maltego offers 15 categories of transforms, being
others. Shodan uses a set of search filters like IP between the most representative a transform that
address, country name, TCP port, keywords, opera- make extraction of document metadata which recei-
ting system, among others. For example, a Shodan ves an input (domain or IP address) and generate an
search using the keyword "Password:" could list all output (author, publishing date, country and other
devices that have between their reachable files the metadata from documents found as related to the
keyword as part of the content. domain). Another popular transform obtains informa-
Finally, Metagoofil [5] helps the user to extract me- tion associated with an organization domain, which
tadata from Microsoft Office and PDF documents, receive an input (organization domain) and produce
amongst others. With Metagoofil can be possible an output (IPs, network range, server names, register
to make a search using a domain name and a type name, registered phone number). Transform results
of file as parameter, which will give all the public can be represented in graphs which helps to illustrate
found files from that domain. The use of this tool is relationships between entities as shown in Figure 1.
very similar to The Foca since it also takes benefit Maltego offers a very simple and easy to use in-
of the metadata of documents. terface. It introduces the concept of “machines”
Maltego, Spiderfoot and Intel Techniques are three which are groups of associated and preconfigured
powerful tools which allow to gather complementary transforms intended to obtain information regar-
data from an adversary. Due its technical features ding to a single type of entity. Maltego uses a TAS
they can support widely a cyber intelligence pro- (Transform Application Server) server that contains
cess. A detailed description of these tools will be the transforms executed by every request of the user.
presented in the following sections. It is also possible to have a private TAS server with
custom-made and not public transforms.[7]
2.1.1 Maltego A private TAS server would also avoid configuring
a development environment on each Maltego wor-
Maltego in one of the tools that facilitates to obtain kstation, since it could be installed in the TAS ser-
information from different open and public sources. ver. Additionally, a private TAS server maintain the
It has an enormous potential to find information privacy of transforms, the integrity of the source
about people, companies or organizations, suppor- code and support a centralized control of versions
ting the recognition which is the first stage of an transparent for Maltego users.
attack [6]. Maltego allows to search all information
around a single initial point (Entity) and then pivot 2.1.2 SpiderFoot
from it to make a new search and obtain more in-
formation. An entity is an abstract representation of SpiderFoot is an open source intelligence tool that
any type of information that is found in real life, such can be used offensively, as part of a black box pe-
as system user names, emails, full names, telephone netration test to gather information about a target.
numbers, addresses, social network accounts, IPs, Also, it can be used defensively by an organization
geographical locations, and even phrases [6]. to identify information that it freely provides and

[ 198 ]
Vínculos
ISSN 1794-211X • e-ISSN 2322-939X • Vol 15, N° 2 (julio-diciembre 2018). pp. 195-214. Universidad Distrital Francisco José de Caldas-Facultad Tecnológica.
Ricardo Andrés Pinto Rico; Martin José Hernández Medina; Cristian Camilo Pinzón Hernández; Daniel Orlando Díaz López; Juan Carlos Camilo García Ruíz

that could be used by attackers. SpiderFoot uses Spiderfoot allows to perform different modes of
more than fifty data sources such as search engines, search. “All mode” allows to get available target
web pages and public servers, to obtain the data. information from all Spiderfoot modules. “Footprint
Information recovered by Spiderfoot is represented mode” makes a fast identification of what informa-
as a graph of nodes as shown in Figure 2. tion the target exposes to internet. “Investigate mode”
develops different secretive collection tasks useful
when there is a suspect that the target is malicious.
“Passive mode” is the most noiseless mode used
to collect information and do not make the target
suspect that is being investigated.
SpiderFoot can even access information hosted on
the Deep Web using an "autonomous" TOR client
and enabling control connections so SpiderFoot
can manage it.
A module in Spiderfoot is “equivalent” to a “trans-
form” in Maltego. A module traduces one type of
data or information to another very different using
an existing relation. When a module discovers a
piece of data, it is transmitted to all other modules
Figure 1. Graphs of entities in Maltego version 4. that can be 'interested' who will use that data to
Source: Own develop its own data discovering processes. Data

Figura 2. Estructura de las unidades de investigación.


Fuente: elaboración propia.

[ 199 ]
Vínculos
ISSN 1794-211X • e-ISSN 2322-939X • Vol 15, N° 2 (julio-diciembre 2018). pp. 195-214. Universidad Distrital Francisco José de Caldas-Facultad Tecnológica.
Inteligencia de fuentes abierta (OSINT) para operaciones de ciberseguridad.

discovered will also feed other modules that can 2. Facebook: This module allows to perform different
be ‘interested’ and a new data discovering process searches using a Facebook username as input
starts, and so on successively. Table 2 summarizes parameter (Figure 3). These searches bring public
some important Spiderfoot modules. Facebook information related with the target.
3. Documents: This module supported by Google
Module Description
Hacking techniques can perform searches of
Look for associated accounts on almost 200
websites (Ebay, Slashdot, reddit and others) documents such as docx, xlsx, pdf and others.
Accounts Input: Email, domain name These documents can contain data about the
Output: Username, Account on External Site,
User Account on External Site
target such as organizations where had been
Check if a domain or IP is malicious affiliated, company where had worked, curri-
according to [Link] culum, documents created by the person, etc.
Input: Internet name, IP, Affiliate - Internet
[Link] Name, Affiliate - IP Address, Co-Hosted Site 4. Image Reversal: This module uses Google ar-
Output: Malicious IP, malicious internet tificial intelligence algorithms for facial recog-
name, malicious affiliate IP address,
malicious affiliate, malicious co-host site
nition. It receives as input parameter an image
Identify Base64-encoded strings in any URL and develop searches to find similar images
content and URLs, often revealing interesting published on internet.
Base64 hidden information
Input: Linked URL - Internal, web content
5. UserSherlock: This module allows to search ac-
Output: Base64-encoded Data counts in internet having a similar username than
Search Bing for hosts sharing the same IP the target. It has more services where is possible
Input: IP Address, Netblock Ownership
Bing
Output: Co-Hosted Site, Search Engines Web to have an account than Pipl database.
Content
Table 2. SpiderFoot modules.
Source: Own

2.1.3 Intel Techniques

Intel Techniques is a website created by Michael


Bazzel that provides a set of services oriented to
OSINT. Between the services offered by Intel Tech-
niques is possible to find modules that group a set
of queries toward common web services e.g. Pipl,
Facebook, among others, that can provide public
information. Intel Techniques also offers links to
books and conferences, access to forums, informa-
tive blogs and podcasts. Figure 3. Facebook tools offered by Intel Techniques.
The use of each one of the modules provided by In- Source: Own
tel Techniques depends on the available target data
(Input parameter). The most useful modules offered 2.1.4 OSINT Process
by Intel Techniques are:
The experiments developed in this paper follow the
1. Pipl: This module uses the target username as an activities shown in Figure 4 which are part of an
input parameter and develops a search to find OSINT process. This process is composed by three
accounts existing in popular websites that were important phases named gathering, processing and
created with the same username. taking advantage.

[ 200 ]
Vínculos
ISSN 1794-211X • e-ISSN 2322-939X • Vol 15, N° 2 (julio-diciembre 2018). pp. 195-214. Universidad Distrital Francisco José de Caldas-Facultad Tecnológica.
Ricardo Andrés Pinto Rico; Martin José Hernández Medina; Cristian Camilo Pinzón Hernández; Daniel Orlando Díaz López; Juan Carlos Camilo García Ruíz

The gathering phase is the collection of all public OSINT tools. It is also needed to discard no-rela-
information available around the target. This infor- ted information that could had been collected by a
mation may be in public servers, web pages, blogs, non-proper filter set in an OSINT tool. Detect the
social networks and, in the context of this paper, refers false positives is also needed given that some collec-
to Colombian open source websites, e.g. Open Data ted information is not really related to the target (e.g.
initiative web page1, National Civil Registry page2, namesake) due the existence of similarities between
amongst others. As part of the gathering, some data the target and another person. All of the activities
correlation process can be developed that allows to in this phase involve the analysis and correlation of
complement an adversary profile using different kind no structured data to enrich the target profile with
of information, e.g. personal and professional data. validated information. Another activity that can be
This can allow us to obtain basic information about the done in this phase is the analysis of sentiment of
person such as emails, names, addresses, aliases on the text produced by the target, such as comments
websites and also national identification documents, on social networks, blogs, etc., regarding a specific
helping us to know their possible location, motivations topic, making it possible to estimate his thought or
and behaviors. For this, a set of different Colombian position. This analysis will determine a feeling whe-
sources are consumed by the set of transforms for ther negative, positive or neutral, through different
the Colombian context mentioned in section 2.1.5. prediction models as presented in section 3.1.
This phase requires a broad use and understanding The taking advantage phase aims to use the collected
of OSINT tools such as those mentioned in Table 1. and processed information to support cyber intelli-
The processing phase refers to the actions applied gence objectives. One of the possible uses of this
over the collected data to make it usable. One of information is to predict an attacks or crime. Predic-
the first activities is to eliminate duplicate informa- tion could be possible through the determination of
tion that could have been collected by the different the triangle of three aspects (Motive, Opportunity,

Figure 4. Open source intelligence phases.


Source: Own

1. [Link]
2. [Link]

[ 201 ]
Vínculos
ISSN 1794-211X • e-ISSN 2322-939X • Vol 15, N° 2 (julio-diciembre 2018). pp. 195-214. Universidad Distrital Francisco José de Caldas-Facultad Tecnológica.
Inteligencia de fuentes abierta (OSINT) para operaciones de ciberseguridad.

Means) that must come together to develop a crime. Using the personal email as input parameter in Intel
With the collected information could be possible for Techniques was possible to obtain the curriculum
example to determine the means, e.g. capabilities and and then his national ID number. The national ID
skills of the adversary obtained from the curriculum, number was then used as input parameter for the
and the motives, e.g. reasons to perform a crime su- transforms shown in Table 3.
pported by the feelings expressed by the adversary. The transforms that were applied are following:

2.1.5 OSINT Transforms for the Colombian context • National Civil Registry Transform: This transform
uses the national ID number to consult citizen
The collected information in the open source intelli- information available in the National Civil Re-
gence phases can be useful to develop cyber security gistry website. The obtained information (Po-
labors. However most of the OSINT tools do not con- lling place and place of issue of the Colombian
sult Colombian open sources so there is a big amount identification card) becomes a Maltego entity.
of useful information that a Colombian law agency can • Colombia National Police Transform: This trans-
miss. In this paper a set of components (transforms) form accesses the judicial background check
are developed and proposed to allow the collection service available in the website of the Colombian
of information from Colombian open sources. National Police and through Selenium library in-
For this purpose, a set of Colombian open sources teract with the web page to enter the national ID
offering information useful in OSINT processes were number and submit the service request. Obtained
identified. Then, a set of transforms were developed information (Legal background of a citizen) is
using Python 2.7.x as programming language3. One used to construct an entity with the data recovered
of the main Python libraries used in the transforms from the Colombian National Police database.
construction was "Selenium" that allows to interact • National Army of Colombia Transform: This trans-
with web pages emulating a browser client to access form uses the national ID number as input pa-
information. Another’s important Python libraries rameter in the website of the National Army of
used in the transforms were "Pandas" and "Sodapy", Colombia. Using Selenium library makes the in-
the first one used to handle information in different teraction with the website and recover the citizen
modes (datagram, xml, html) and the second one data that is used to build two entities, one with the
to make connections to API, e.g. Open Data API. name of the person and another with the Military
Table 3 shows the identified main Colombian open status of the citizen which includes the military
sources, the service being consulted and the des- district where the military card was issued.
cription of the developed transform. • Open data transform: This transform generates
request to the API offered by the web services
2.1.6 OSINT process using transforms in a of Open Data. It uses the national ID number as
Colombian context input parameter and the response obtained is used
to build two entities. One entity contains politic
An intelligence exercise was done over a person with information, in this case the target is an elected
job in the Colombian state. The target name was re- councilor, so the recovered information refers to
named as Adam Goldstein and its personal collected the municipality that represents, the region where
data were changed to not expose the privacy of the is located and the name. The other entity has a
person. The input parameter was the personal email single property that refers to the Colombian po-
[Link]@[Link]. litical party where the target is affiliated.

3. [Link]

[ 202 ]
Vínculos
ISSN 1794-211X • e-ISSN 2322-939X • Vol 15, N° 2 (julio-diciembre 2018). pp. 195-214. Universidad Distrital Francisco José de Caldas-Facultad Tecnológica.
Ricardo Andrés Pinto Rico; Martin José Hernández Medina; Cristian Camilo Pinzón Hernández; Daniel Orlando Díaz López; Juan Carlos Camilo García Ruíz

Table 3. Colombian open sources and transforms.

Colombian open source Service Transform Input / Output

Policía Nacional de Colombia / Colombia Input → Colombian identification number


Judicial background check
National Police Output → Legal background of a citizen

Sistema de Identificación de Potenciales


Input → Colombian identification number
Beneficiarios de Programas Sociales
Affiliation and score check Output → Socio-economic stratum of a citizen and
(SISBEN) / Identification System for
SISBEN score
Potential Beneficiaries of Social Programs
Administradora de los Recursos
del Sistema General de Seguridad Social
General System of Social Input → Colombian identification number Output →
en Salud (ADRES) / Administrator of
Security affiliation Affiliation of a citizen to a health service provider
Resources of the General System of Social
Security in Health
Registro Único Nacional de Tránsito Input → Colombian identification number
Traffic infractions check
(RUNT) / National Registry of Traffic Output → Traffic infractions of a citizen

Ejército Nacional de Colombia / National Input → Colombian identification number


Military service check
Army of Colombia Output → Military status of a citizen
Instituto Colombiano de Crédito Educativo
Input → Colombian identification number
y Estudios Técnicos en el Exterior (ICETEX)
Affiliation check Output → Educational program where a citizen is
/ Colombian Institute of Educational Credit
enrolled
and Abroad Technical Studies

Information associated with Input → Colombian identification number


Registraduría Nacional del Estado Civil /
the Colombian identification Outputs → Polling place and place of issue of the
National Civil Registry
card Colombian identification card

Procuraduría General de la Nación / Input → Colombian identification number


Criminal background check
Attorney General Office Output → Criminal background of a citizen
Servicio Nacional Input → Colombian identification number
de Aprendizaje (SENA) / National Learning Affiliation Check Output → SENA educational program where a citizen
Service is enrolled
List of councilors of the municipality of San Andres,
Santander
Input → Colombian identification number, Output →
Job position, Political affiliation, Full names, Gender,
Municipality
Government secretaries of municipalities of Valle del
Cauca 2016-2019
Input → Full name, Email
Output → Phone number, Location
Consult open data published
Datos Abiertos / Open Data Mayors of Municipalities of Antioquia 2016-2019
by government organizations
Input → Email
Output → Phone number, Location, Full name, Period

Senate employees
Input → Phone number
Output → Email, Full name, Location
Delegates from the Ministry of Defense of Santa
Marta
Input → Email
Output → Full name, Location, Phone number

Source: Own.

[ 203 ]
Vínculos
ISSN 1794-211X • e-ISSN 2322-939X • Vol 15, N° 2 (julio-diciembre 2018). pp. 195-214. Universidad Distrital Francisco José de Caldas-Facultad Tecnológica.
Inteligencia de fuentes abierta (OSINT) para operaciones de ciberseguridad.

2.1.6 OSINT process using transforms in a interact with the web page to enter the national
Colombian context ID number and submit the service request. Obtai-
ned information (Legal background of a citizen) is
An intelligence exercise was done over a person with used to construct an entity with the data recovered
job in the Colombian state. The target name was re- from the Colombian National Police database.
named as Adam Goldstein and its personal collected • National Army of Colombia Transform: This trans-
data were changed to not expose the privacy of the form uses the national ID number as input pa-
person. The input parameter was the personal email rameter in the website of the National Army of
[Link]@[Link]. Colombia. Using Selenium library makes the in-
Using the personal email as input parameter in Intel teraction with the website and recover the citizen
Techniques was possible to obtain the curriculum data that is used to build two entities, one with the
and then his national ID number. The national ID name of the person and another with the Military
number was then used as input parameter for the status of the citizen which includes the military
transforms shown in Table 3. district where the military card was issued.
• Open data transform: This transform generates
The transforms that were applied are following: request to the API offered by the web services
of Open Data. It uses the national ID number as
• National Civil Registry Transform: This trans- input parameter and the response obtained is used
form uses the national ID number to consult to build two entities. One entity contains politic
citizen information available in the National information, in this case the target is an elected
Civil Registry website. The obtained informa- councilor, so the recovered information refers to
tion (Polling place and place of issue of the the municipality that represents, the region where
Colombian identification card) becomes a is located and the name. The other entity has a
Maltego entity. single property that refers to the Colombian po-
• Colombia National Police Transform: This trans- litical party where the target is affiliated.
form accesses the judicial background check
service available in the website of the Colombian All entities created by transforms are included in
National Police and through Selenium library Maltego workspace as shown in Figure 5.

Figure 5. Maltego workspace with entities derived by transforms execution.


Source: Own

[ 204 ]
Vínculos
ISSN 1794-211X • e-ISSN 2322-939X • Vol 15, N° 2 (julio-diciembre 2018). pp. 195-214. Universidad Distrital Francisco José de Caldas-Facultad Tecnológica.
Ricardo Andrés Pinto Rico; Martin José Hernández Medina; Cristian Camilo Pinzón Hernández; Daniel Orlando Díaz López; Juan Carlos Camilo García Ruíz

3. Developing models to understand co- Machine learning refers to the practice of teaching
llected information a computer how to detect patterns and make con-
nections by showing it a massive volume of data.
In this section we will talk about how machine Another definition of machine learning tells about
learning and one of the branches of data scien- the nontrivial extraction of implicit, previously unk-
ce, i.e. the sentiment analysis, can support cyber nown and potentially useful information from data.
intelligence labors. Three specific descriptive The large amount of data managed by fields such as
models will be presented which can be used to commerce, banking or healthcare and the power of
make sentiment analysis of phrases written in new computers to perform data processing opera-
spanish. tions give impetus to machine learning technologies.
The training set used to train the models was obtai- For the development of this paper, two models were
ned from a repository published as part of the Fee- mainly considered: descriptive and predictive.
ling Analysis Workshop of the Spanish Society for
the Processing of Natural Language (SEPLN)4. The 3.1.1 Descriptive models
training set contains approximately 60,000 tweets,
each of them labeled with a certain feeling (positi- Descriptive models look for interpretable patterns
ve or negative). Models were trained with the 70% to describe data. These models include: clustering,
of the training set and were tested using the 30% discovery of association rules and discovery of se-
residuary percentage. quential patterns [8].
Permiten establecer relevancia / irrelevancia de fac-
3.1 Machine learning and Sentiment analysis tores y si aquélla es positiva o negativa respecto a
otro factor o variable a estudiar.
Machine learning is an area of study derived from
Data Science that has been presented for many de- 3.1.2 Predictive models
cades but has acquired a high popularity in last
years for its multiples application possibilities in Predictive models employ some variables to pre-
many fields, including cybersecurity. dict future or unknown values of another variables.

Figure 6. Example of sentiment analysis for a twit. [8-11]

4. [Link]

[ 205 ]
Vínculos
ISSN 1794-211X • e-ISSN 2322-939X • Vol 15, N° 2 (julio-diciembre 2018). pp. 195-214. Universidad Distrital Francisco José de Caldas-Facultad Tecnológica.
Inteligencia de fuentes abierta (OSINT) para operaciones de ciberseguridad.

These models include: classification, regression and This technique takes advantage of machine learning
deviation detection [8]. elements such as latent semantics analysis, support
This paper uses predictive models to make sen- vector machines, bag of words, among others [14] .
timent analysis of information published by an
adversary in a social network (Twitter) around a • Concept level.
specific topic.
Conceptual level approaches employ elements for
3.1.3 Sentiment analysis knowledge representation, e.g. semantic networks.
Conceptual level detects semantics that are expres-
Sentiment analysis [9] is related with Opinion Mi- sed in a subtle manner, e.g. through the analysis
ning [10] and aims to: extract subjective informa- of concepts that do not explicitly convey relevant
tion from data, perform massive classification of information but which are implicitly linked to other
the positive, negative or neutral connotation of a concepts that do so [15].
text, and try to determine the attitude of a writer
regarding a topic. 3.2 Model 1: Bayes Naïve with a derivation of
Human supervision is required in sentiment analysis bag of words
exercises due these models have some limitations
that only a human can overcome, e.g. sentiment Model 1 was implemented using Bayes Naive with
analysis models cannot analyze historical trends. a variation of bag of words methodology [8]. The
An example of sentiment analysis can be seen in analysis of sentiment for twits written by an adversary
Figure 6 where a client is upset due the poor service regarding a specific topic must go through a series of
provided by a bank, and post a twit expressing its preprocessing steps to finally enter into Bayes Naive
nonconformity. The twit can be processed to de- classification model. Next, each of these steps will
termine positive or negative words and identify a be described which are necessary for the analysis
sentiment. of sentiments [16][17].
Sentiment analysis can be developed through one or
more of the following techniques for text processing 3.2.1 Tokenization
[11] : keyword location, lexical affinity, statistical
methods, concept level. This step includes deletion of characters, change of
uppercase to lowercase and separation of a phrase in
• Keywork location. list. Deletion of characters eliminates any character
that is not important and that does not change the
This technique classifies the text into effect catego- feeling of the phrase such as: @ {} []? ¡¨ *? ¡! "#! &%
ries based on the presence of unambiguous affection $ ¬ | | + * -. Phrases containing questions, irony or
words, such as happy, sad, frightened and bored exclamation were not taken in account at the mo-
[12] . ment and will be considered as future work.
Change of uppercase to lowercase pass all the words
• Lexical affinity. of the phrase that are in uppercase to lowercase to
avoid having repeated words that are considered as
This technique assigns to arbitrary words (birth, different only because of having capital letters. For
growth, speed, etc.) a probable "affinity" with par- example, words "bad" and "Bad" have the same feeling
ticular emotions [13]. so should not be set as different words. Separation of
a phrase in list separates the phrase into a list of words
• Statistical methods. for easy handling of the model and training phase.

[ 206 ]
Vínculos
ISSN 1794-211X • e-ISSN 2322-939X • Vol 15, N° 2 (julio-diciembre 2018). pp. 195-214. Universidad Distrital Francisco José de Caldas-Facultad Tecnológica.
Ricardo Andrés Pinto Rico; Martin José Hernández Medina; Cristian Camilo Pinzón Hernández; Daniel Orlando Díaz López; Juan Carlos Camilo García Ruíz

3.2.2 Context • count (di, C) is the number of occurrences of the


word 𝑑𝑖 for each class C (Positive or negative)
This step gives context to the phrase making an • Vc is the total number of words belonging to
analysis by word to determine if it "empower" ano- class C (Positive or negative) existing in the tra-
ther word. For example, "very" is a feeling empower ining set
word, so the words "very bad" have to be considered • n is the total number of words in the training set
together instead of considering them as individual
words. 3.2.5 Testing

3.2.3 Stop words As mentioned previously accuracy calculation (tes-


ting) was done using a testing set composed by the
This step eliminates from the list of words previously 30% of the training set. The trained model received
processed the "Stop words". Stop words are recog- the testing set and calculated the classification of
nized because they do not alter the feeling of the each twit. The area under the curve method was
phrase, such as: an, a, over, all, also, in addition, used using the Python "sklearn" library to compare
where, some, etc. the classification of each twit with its original and
correct classification. Bayes Naïve model obtained
3.2.4 Training a score of 0.59.

Analysis of sentiment is done using Bayes Naive 3.3 Model 2: Support Vector Machine
equation derived from the bag of word (Equation
1). Each word di of the phrase is selected and coun- Model 2 uses Support Vector Machines (SVM) [18]
t(di,C) counts how many times word di exists in for the analysis of sentiment of sentences written by
positive and negative sentences of the training set. an adversary. As indicated in [19] the phrases (twits)
Then, the probability associated with the word di is to be classified has to go through a series of steps to
calculated by dividing the number of occurrences of finally be entered into the SVM classification model.
the word di in the training set, (count(di,C)), on the Next, each of these steps will be described which
total number of words belonging to class C (Positive are necessary for the analysis of sentiments [20].
or negative) existing in the training set, V(Cn).
Probability of each word is cumulated in the total 3.3.1 Tokenization
probability of the phrase Finally, the classifica-
tion of the phrase will be determined by the accumu- This step includes deletion of characters, change of
lated probability for each class (positive, negative). uppercase to lowercase and separation of a phrase in
list. Deletion of characters eliminates any character
(1) that is not important and that does not change the
feeling of the phrase such as: @ {} []? ¡¨ *? ¡! "#! &%
$ ¬ | | + * -. Phrases containing questions, irony or
Where exclamation were not taken in account at the mo-
ment and will be considered as future work.
• P is the final probability of the phrase Change of uppercase to lowercase pass all the words
• C is the phrase class which can be positive or of the phrase that are in uppercase to lowercase to
negative avoid having repeated words that are considered
• p(C) is the probability that a word is part of a as different only because of having capital letters.
positive or negative phrase For example, words "bad" and "Bad" have the same

[ 207 ]
Vínculos
ISSN 1794-211X • e-ISSN 2322-939X • Vol 15, N° 2 (julio-diciembre 2018). pp. 195-214. Universidad Distrital Francisco José de Caldas-Facultad Tecnológica.
Inteligencia de fuentes abierta (OSINT) para operaciones de ciberseguridad.

feeling so should not be set as different words. Se-


paration of a phrase in list separates the phrase into
a list of words for easy handling of the model and
training phase. (2)

3.3.2 Stop words

This step eliminates from the list of words previously Where:


processed the "Stop words". Stop words are recogni-
zed because they do not alter the sentiment of the • C is the constant of capacity
phrase, such as: an, a, over, all, also, in addition, • Core is used to transform input data to the
where, some, etc. function space
• γ and ɗ represents parameters for the manage-
3.3.3 Stemming ment of non-separable data

This step reduces a word to its root or stem, so that 3.3.6 Testing
the model finds it easy to train and calculate the
sentiment of a word. This is due the sentiment can As mentioned previously accuracy calculation (tes-
be the same for a word in plural or singular. ting) was done using a testing set composed by the
30% of the training set. The trained model received
3.3.4 Vectorization the testing set and calculated the classification of
each twit. The area under the curve method was
This step changes the phrase representation toward used using the Python "sklearn" library to compare
a matrix where the columns are the words proces- the classification of each twit with its original and
sed by previous steps, and the row is the number of correct classification. SVM model obtained a score
occurrences that word appears in the phrase. of 0.823.

3.3.5 Training 3.4 Model 3 Bernoulli

Analysis of sentiment is doing using SVM model Model 3 uses Bernoulli model for the analysis of
which has been trained with multiple vectorized sentiment. It is important to understand that all the
and tagged phrase placed in a multi-plane. As part models presented in this paper for analysis of sen-
of the training phase, the SVM model define a sin- timent are oriented to calculate the polarity of a
gle plane that separate negative and positive phra- phrase and not just a word. These models serve as
ses. Unclassified phrases follow the preprocessing a starting point to develop other studies related to
steps defined previously and enter to the model to the disambiguation of polarity presented by phrases
be classified. or documents[21].
The single plane that separates positive and negative As mentioned in[22], there are preprocessing steps
phrases is chosen according to Equation 2 where that must be performed to prepare the phrase to
Xi and Xj represent points that define the plane entry to the Bernoulli model. Those steps are tokeni-
that classifies the different vectors in the multipla- zation and vectorization. Additionally, it is required
ne. At last, SVM model determines the sentiment to build a pipeline element that is used to serialize
of a phrase according to its spatial location in the the trained model to store it and streamline the clas-
multi-plane. sification of phrases.

[ 208 ]
Vínculos
ISSN 1794-211X • e-ISSN 2322-939X • Vol 15, N° 2 (julio-diciembre 2018). pp. 195-214. Universidad Distrital Francisco José de Caldas-Facultad Tecnológica.
Ricardo Andrés Pinto Rico; Martin José Hernández Medina; Cristian Camilo Pinzón Hernández; Daniel Orlando Díaz López; Juan Carlos Camilo García Ruíz

3.4.1 Tokenization classification. Equation 3 shows how to calculate


the probability.
This step includes deletion of characters Deletion of
characters eliminates any character that is not im- (3)
portant and that does not change the feeling of the
phrase such as: @ {} []? ¡¨ *? ¡! "#! &% $ ¬ | | + * -.
Phrases containing questions, irony or exclamation Where
were not taken in account at the moment and will
be considered as future work. is the number of occurrences of t in the training
set
3.4.2 Stop words is the number of elements that the class C has
within the training set
This step eliminates from the list of words previously
processed the "Stop words". Stop words are recogni- 3.4.5 Testing
zed because they do not alter the sentiment of the
phrase, such as: an, a, over, all, also, in addition, Different tests were carried out with phrases with a
where, some, etc. well mark polarity and the response of the model
for all these cases was correct. In addition, accura-
3.4.3 Stemming cy calculation (testing) was done using a testing set
composed by the 30% of the training set. The trained
This step reduces a word to its root or stem, so that model received the testing set and calculated the
the model finds it easy to train and calculate the classification of each twit. The area under the curve
sentiment of a word. This is due the sentiment can method was used using the Python "sklearn" library
be the same for a word in plural or singular. to compare the classification of each twit with its
original and correct classification. Bernoulli model
3.4.4 Count vectorizer obtained a score of 0.81.

Count vectorizer counts the occurrences of a word in 3.5 Comparison of results


positive and negative tweets belonging to the training
set. This count is used to determine the probability The SVM model was the one with the highest ac-
of the phrase of being positive or negative. curacy (0.82) in comparison with Bernoulli model
(0.81) and Bayes Naïve model (0.59). The reliabi-
• Training lity of the models depends on several aspects like
the preprocessing stage, which were applied in a
The Bernoulli model is of binomial type, which different way. For example, SVM and Bernoulli mo-
means that it mainly analyzes successes and fai- dels share the Vectorization/Count Vectorizer step,
lures, in a sentiment context phrases whose con- however Bayes Naive does not. Also, SVM and Bayes
notation are negative or positive are analyzed. It Naive models implement a stop words step, however
uses a formula where the occurrences of a certain Bernoulli does not.
word are counted within the training set and also The accuracy is probably higher for the SVM mo-
infer the phrase classification from the number of del because of a preprocessing stage that prepare
occurrences of the same in negative or positive phrases (twits) in the proper way with nor excess
sentences. Bernoulli model uses the word coun- in data cleaning neither scarcity. A scarcity of data
ter to calculate the probability related with the preprocessing can avoid that the model identifies

[ 209 ]
Vínculos
ISSN 1794-211X • e-ISSN 2322-939X • Vol 15, N° 2 (julio-diciembre 2018). pp. 195-214. Universidad Distrital Francisco José de Caldas-Facultad Tecnológica.
Inteligencia de fuentes abierta (OSINT) para operaciones de ciberseguridad.

really meaningful keywords, but an extreme cleaning or Iraq where we are at war, and somebody would be
can also produce that the phrase loses its original ahead in terms of intelligence and technology." [23].
meaning. Bernoulli model having just a few prepro- This indicates that one of the greatest concerns of na-
cessing steps has the lowest accuracy. tions is the cyberspace, called the fifth war domain,
Another possible reason that justifies the highest is generally included in the national security strategy.
accuracy of the SVM model is its linear statistical Thus, on July 14, 2011, the Colombian State, throu-
nature that does not apply the classic binomial and gh the National Department of Planning (DNP), set
multinomial paradigms applied by the other models. the guidelines for the Policy of Cybersecurity and
Instead, SVM opts for an alternative mathematical Cyberdefense of Colombia with the document of
modeling where the proximity between words is the National Council of Economic and Social Policy
measured using a hyperplane. (CONPES 3701) [24]. This document presents the
road map and defines the roles of each of the orga-
4. OSINT for Colombian Law Enforcement nism in charge of cybersecurity, cyberdefense and
Agencies cyberincidents management in Colombia. The aggru-
pation of these organisms is called the Intersectoral
In an interview with the former FBI director, Robert Commission (Figure 7) and is responsible for provide
Mueller, he was asked: “3000 lives were lost on 9/11, technical assistance, coordinate incident mana-
what is your worry from a cyber-perspective that would gement, offer emergency assistance, development
be catastrophic like that?” Muller answered: “The way operational capabilities, provide cyberintelligence
of looking at the power grids and our infrastructure, and information and advice and support cyberdefense.
financially because it would cripple us if there was a Today, 7 years after creation of CONPES 3701, the
substantial attack on wall street on the exchanges. It also Colombian state security agencies in charge of cy-
could lead to a loss of lives to the extent that we can bersecurity are in growing its cyber capacities and
have our command and control knocked in Afghanistan face the following challenges:

Figure 7. Intersectoral commission defined in CONPES 3701.


Source: Own.

[ 210 ]
Vínculos
ISSN 1794-211X • e-ISSN 2322-939X • Vol 15, N° 2 (julio-diciembre 2018). pp. 195-214. Universidad Distrital Francisco José de Caldas-Facultad Tecnológica.
Ricardo Andrés Pinto Rico; Martin José Hernández Medina; Cristian Camilo Pinzón Hernández; Daniel Orlando Díaz López; Juan Carlos Camilo García Ruíz

I. Cyber terrorism representing an asymmetric personal data, e.g. photos, biographical data,
warfare tool since it is cheap and has a high mobile numbers and emails, through mobile
destructive power[25], since the attacker e.g. applications and web services. This situation
terrorist, adversary state or guerrilla, does not becomes a challenge for law enforcement agen-
need to buy large arsenals but with few resources cies who must carry out investigations around
can develop a cybernetic weapon, transport it, a person, since they require proper tools and
e.g. using a USB and generate an impact over methodologies to support cyber intelligence la-
a cybernetic infrastructure. bors over big amount of available personal data.
II. Application of soft power [26] by foreign coun-
tries which use cultural, ideological and diplo- Many of these challenges are related to large flows
matic means to influence national politics or of information available on the Internet that must be
society, e.g. the use of social networks with fake analyzed by the state security agencies to support
news presidential campaigns in behalf of one cybersecurity and cyberdefense objectives defined in
candidate to obtain geostrategic advantage in a CONPES 3701. The collection, processing and analy-
foreign country. sis of all this public information can be supported by
III. Increase of Internet connections, mainly due open source intelligence tools and methodologies.
systems appear to be interconnected: securi- State security agencies in charge of cybersecurity and
ty, defense, commercial, energy, health, com- cyber defense uses OSINT to perform intelligence
munication, transportation, banking, librarians, which represents a strategic advantage that allows to
etc. These highly connected systems make in- anticipate a possible attack mitigating the risks and
ternet crucial and vital for the most advanced conduct an investigation after an attack or crime.
societies[27]. This paper presented in section 2 an OSINT architec-
IV. Growing technological dependency, which be- ture applicable to state security agencies supported
comes one of the biggest challenges since a cy- in a set of transforms applicable to the Colombian
berattack can produce chaos in state and society context. This contribution represents an advance
affecting delivery of essential services like water, in the prevention of attacks coming from internal
electricity and banking. and external agents allowing in many cases to de-
V. Secure Operation Technologies (TO) like indus- velop a strategic anticipation to the adversary. In
trial control systems designed to be functional addition, the transforms improve the focus that the
but not safe. TO systems can exist for example cyber analyst must have to deliver an intelligence
in energy companies with different connected product with accurate information, since there is
electrical substations which can be adversely contextualized. Additionally, this paper presents in
controlled to alter the energy supply. Stuxnet section 3 an application of machine learning mo-
[28] is other representative case where an ura- dels to make sentiment analysis, which allows cyber
nium enrichment plant was vulnerated leaving analyst to generate early warnings regarding illicit
it useless and delaying the Irani project to pro- acts and anticipate sabotage campaigns.
duce energy by means of radioactive elements. Contributions included in Section 3 and 4 can be
VI. Internet of Things technologies like Smart TVs, used not only for cybersecurity purposes, but it could
IP cameras, smart toys, among others, which be used to prevent suicides, human traffic, citizen
lately has been used by hackers to carry out security perception, citizen complaints expecting to
denial of services attacks against technological be attended, drug traffic, among others. Definition of
infrastructure. public policies could also be aided from information
VII. Large amount of people connected to different collected and analyzed by transforms and models,
networks voluntarily or involuntarily, sharing due this could help to build a society profile that

[ 211 ]
Vínculos
ISSN 1794-211X • e-ISSN 2322-939X • Vol 15, N° 2 (julio-diciembre 2018). pp. 195-214. Universidad Distrital Francisco José de Caldas-Facultad Tecnológica.
Inteligencia de fuentes abierta (OSINT) para operaciones de ciberseguridad.

allow to design policies accordingly. In this way is of information from a Colombian or Latin American
possible to analyze, identify, counteract, mitigate and context pose the necessity of develop own cyberinte-
prevent threats and even predict possible attacks. lligence solutions. For these purpose, it is important
These contributions allow state security agencies to to joint efforts between state, industry and acade-
prospectively anticipate the threats new wars defined mia around research, development and innovation
by George Friedman[29] in where he predicts that in cybersecurity solutions supported using open
the upcoming wars will be based on a combination source frameworks.
of computer sabotage and high-precision attacks
against targets strategic to generate chaos. 5. Conclusions
Among the most popular OSINT solutions are: Vo-
yager5 from Voyager Labs (Israel), iSIHT6 from Fireye Any blog, web page, online newspaper, social ne-
(USA), iNSIGTH7 from Checkpoint (Israel) and Flas- twork, forum and even free datasets can become a
hpoint8 from Flashpoint-intel (USA). The mentioned great source of information, which can be accessed
solutions are proprietary with high costs in licensing to collect data used in cyber intelligence labors.
and maintenance. Additionally, in many cases these Despite of this, privacy of data is a very serious
solutions are offered as Software as a Service (SaaS), issue, so it is also mandatory to know and recog-
which can be a disadvantage since highly confi- nize the difference between violating privacy and
dential information is hosted in the cloud. When a collecting information in reason of the protection
proprietary solution is contracted, some intelligence of critical assets.
labors can be developed by people not belonging The analysis of sentiments can represent a great tool
to state security agencies, which could decrease the for law agencies to face crime, because it allows to
development of internal capacity and expertise. In determine and analyze the position of a criminal
addition, data collected by foreign solutions gene- regarding a specific subject. It could allow to iden-
rally belongs to American and European sources but tify the reasons to carry out an attack on a person
not Latin American or Colombian, which difficult or organization. The identification of a possible ad-
the cyberintelligence tasks due Colombian security versaries serves to design and implement a cyber
agencies requires mainly data within a Colombian defense strategy that prevent future attacks.
context. One of the most valuable skill for cyber intelligence
When the aforementioned solutions are not availa- labors is to know how to look for and find infor-
ble, or these do not provide enough information, the mation regarding a target. Law agencies and cyber
cyber analyst proceeds to carry out search and co- intelligence organizations values this skill because
llection in a manual way according to his knowledge it represents an advantage against adversaries that
of the threat which could generate subjectivity and can be extremely useful to handle national security
vagueness. Additionally, manual activities increase incidents. A final suggestion for organizations or in-
the time required to develop a cyber intelligence dividuals is to be aware of all the information that
cycle, making difficult to carry out other cyber se- is shared or published on social networks or any
curity tasks in parallel. web page. This is due through OSINT is possible
The high ownership costs of proprietary OSINT so- that collect information which could be used by a
lutions, the confidentiality of data and the collection criminal to achieve an attack.

5. [Link]
6. [Link]
7. [Link]
8. [Link]

[ 212 ]
Vínculos
ISSN 1794-211X • e-ISSN 2322-939X • Vol 15, N° 2 (julio-diciembre 2018). pp. 195-214. Universidad Distrital Francisco José de Caldas-Facultad Tecnológica.
Ricardo Andrés Pinto Rico; Martin José Hernández Medina; Cristian Camilo Pinzón Hernández; Daniel Orlando Díaz López; Juan Carlos Camilo García Ruíz

6. Future Works [3] W. Alcorn, C. Frichot, and M. Orrù, “The Brow-


ser hacker’s handbook”, New Jersey: John Wiley
We state as future work the automation of captcha and Sons, 2014.
resolution present in the services consulted by the [4] M. Gregg, “Certified Ethical Hacker (CEH) Ver-
transforms presented in this paper, so the user does sion 9 Cert Guide” London: Pearson Education,
not have to solve it manually. Different techniques 2017.
can be explored like optical character recognition [5] P. Engebretson, “The basics of hacking and pe-
algorithms (OCR). Development of new transforms netration testing” Syngressr Publishing, 2013.
could also be considered as future work, which can [6] D. Bradbury, “In plain view: open source inte-
be able to develop advanced searches, like the ones lligence”, Computers in Human Behavior, no.
done by Intel Techniques. These new transforms 4, pp. 5–9, 2011.
could be integrated in an open source tool like Mal- [7] B. de S. G. Rodrigues, “Open-source intelligen-
tego looking for the integration of OSINT tools [30]. ce em sistemas SIEM” Lisboa: Universidade de
Transforms able to analyze unstructured information Lisboa, 2015.
like Microsoft documents could also be useful to de- [8] C. Pérez, “Minería de datos: técnicas y herra-
crease the manual review and achieve more efficient mientas” Paraninfo Cengage Learning, 2007.
cyber intelligence processes. Additionally, looking [9] G. Subramanian, “R Data analysis projects:
for information in the deep web, e.g. through TOR build end to end analytics systems to get deeper
network, could be useful. Finally, improve the accu- insights from your data”, Birmingham: Packt
racy of descriptive machine learning models should Publishing, 2017.
also be considered for example using better training [10] L. Zhang and B. Liu, “Sentiment Analysis and
data sets which can be customized for the specific Opinion Mining”. in Encyclopedia of Ma-
variations of a language (Colombian Spanish) or for chine Learning and Data Mining, Boston:
specific contexts (professional or informal). Springer, 2017, pp. 1152–1161, [Link]
org/10.1007/978-1-4899-7687-1_907
Acknowledgment [11] E. Cambria, B. Schuller, Y. Xia, and C. Havasi,
“New Avenues in Opinion Mining and Senti-
This work has been supported partially by the Colom- ment Analysis”, IEEE Intelligent Systems, vol. 28,
bian School of Engineering Julio Garavito (Colombia) no. 2, pp. 15–21, 2013, [Link]
through the project “Cyber Security Architecture MIS.2013.30
for Incident Management”, funded by the Internal [12] A. Ortony, G. L. Clore, and A. Collins, “The
Research Opening 2017. cognitive structure of emotions” Cambridge:
Cambridge University Press, 1988, [Link]
References org/10.1017/CBO9780511571299
[13] R. A. Stevenson, J. A. Mikels, and T. W. Ja-
[1] M. Glassman and M. J. Kang, “Intelligence in mes, “Characterization of the Affective Nor-
the internet age: The emergence and evolu- ms for English Words by discrete emotional
tion of Open Source Intelligence (OSINT)”, categories”, Behavior Research Methods, vol.
Computers in Human Behavior, vol. 28, no. 2, 39, no. 4, pp. 1020–1024, 2007, [Link]
pp. 673–682, 2012, [Link] org/10.3758/BF03192999
chb.2011.11.014 [14] P. D. Turney, “Thumbs Up or Thumbs Down?
[2] L. Brotherston and A. Berlin, “Defensive se- Semantic Orientation Applied to Unsupervised
curity handbook: best practices for securing Classification of Reviews”, In Proceedings of
infrastructure”. O’Reilly Media, 2017. the 40th Annual Meeting of the Association for

[ 213 ]
Vínculos
ISSN 1794-211X • e-ISSN 2322-939X • Vol 15, N° 2 (julio-diciembre 2018). pp. 195-214. Universidad Distrital Francisco José de Caldas-Facultad Tecnológica.
Inteligencia de fuentes abierta (OSINT) para operaciones de ciberseguridad.

Computational Linguistics (ACL), Philadelphia, [22] H. Wang, D. Can, A. Kazemzadeh, F. Bar and
july 2002, pp. 417-424. S. Narayanan, “A System for Real-time Twitter
[15] S. M. Kim and E. Hovy, “Identifying and Analyzing Sentiment Analysis of 2012 U.S. Presidential
Judgment Opinions”, Association for Computatio- Election Cycl,”. In 50th Annual Meeting of the
nal Linguistics Stroudsburg, pp. 200–207, 2006, Association for Computational Linguistics, Jeju
[Link] Island, july, 2012.
[16] Liangxiao Jiang, H. Zhang, and Zhihua Cai, “A [23]C-SPAN, “Robert Mueller on Cyberse-
Novel Bayes Model: Hidden Naive Bayes”, IEEE curity” [En línea] Disponible en: ht-
Transactions on Knowledge and Data Enginee- tps://[Link]/video/?319726-3/
ring, vol. 21, no. 10, pp. 1361–1371, 2009, robert-mueller-cybersecurity&start=1876
[Link] [24] Departamento Nacional de Planeación,
[17] Y. Yang and G. I. Webb, “A Comparative Study “CONPES 3701 - Lineamientos de Política para
of Discretization Methods for Naive-Bayes Clas- Ciberseguridad y Ciberdefensa. Colombia”.
sifiers”, J. Res., vol. 2, p. 267-324, 2007. Consejo Nacional de Política Económica y So-
[18] M. A. Hearst, S. T. Dumais, E. Osuna, J. Platt, cial, 2011.
and B. Scholkopf, “Support vector machines”, [25] R. Rodríguez, “Guerra Asimétrica”. [En línea].
IEEE Intelligent Systems and their Applications, Disponible en: [Link]
vol. 13, no. 4, pp. 18–28, 1998, [Link] carga/articulo/[Link]
org/10.1109/5254.708428 [26] J. Nye, “Bound to Lead: The Changing Nature
[19] F. Sebastiani, “Machine Learning in Automated of American Power” Hachette U. Basic Books,
Text Categorization”, ACM Computing Sur- 2016.
veys, vol. 34, no. 1, pp. 1–47, 1999, https:// [27] G. S. Medero, “Ciberespacio y el crimen or-
[Link]/10.1145/505282.505283 ganizado. Los nuevos desafíos del siglo XXI”,
[20] B. Pang and L. Lee, “A Sentimental Education: Revista Enfoques, vol.10, no. 16, pp. 71–87,
Sentiment Analysis Using Subjectivity Sum- 2012.
marization Based on Minimum Cuts”, Proce- [28] R. Langner, “Stuxnet: Dissecting a cyberwarfare
edings of ACL, pp. 271-278, 2004, [Link] weapon”, IEEE Security and Privacy, vol. 9, no.
org/10.3115/1218955.1218990 3, pp. 49–51, 2011, [Link]
[21] T. Wilson, J. Wiebe, and P. Hoffmann, “Recogni- MSP.2011.67
zing contextual polarity in phrase-level sentiment [29] G. Friedman, “The next 100 years: a forecast for
analysis”, Proceedings of the conference on Hu- the 21st century”, Knopf Doubleday Publishing
man Language Technology and Empirical Methods Group, 2009, pp. 193–212.
in Natural Language Processing, pp. 347–354, [30] R. Steele, “Handbook of Intelligence Studies”
2005, [Link] London: Routledge, 2007.

[ 214 ]
Vínculos
ISSN 1794-211X • e-ISSN 2322-939X • Vol 15, N° 2 (julio-diciembre 2018). pp. 195-214. Universidad Distrital Francisco José de Caldas-Facultad Tecnológica.

[ 195 ]
Vínculos
ISSN 1794-211X • e-ISSN 2322-939X • Vol 15, N° 2 (julio-diciembre 2018). pp. 195-214. Universidad Distrital
Inteligencia de fuentes abierta (OSINT) para operaciones de ciberseguridad.  
[ 196 ]
Vínculos
ISSN 1794-211X • e-ISSN 2322-9
Ricardo Andrés Pinto Rico; Martin José Hernández Medina; Cristian Camilo Pinzón Hernández; Daniel Orlando Díaz López; Juan Ca
Inteligencia de fuentes abierta (OSINT) para operaciones de ciberseguridad.  
[ 198 ]
Vínculos
ISSN 1794-211X • e-ISSN 2322-9
Ricardo Andrés Pinto Rico; Martin José Hernández Medina; Cristian Camilo Pinzón Hernández; Daniel Orlando Díaz López; Juan Ca
Inteligencia de fuentes abierta (OSINT) para operaciones de ciberseguridad.  
[ 200 ]
Vínculos
ISSN 1794-211X • e-ISSN 2322-9
Ricardo Andrés Pinto Rico; Martin José Hernández Medina; Cristian Camilo Pinzón Hernández; Daniel Orlando Díaz López; Juan Ca
Inteligencia de fuentes abierta (OSINT) para operaciones de ciberseguridad.  
[ 202 ]
Vínculos
ISSN 1794-211X • e-ISSN 2322-9
Ricardo Andrés Pinto Rico; Martin José Hernández Medina; Cristian Camilo Pinzón Hernández; Daniel Orlando Díaz López; Juan Ca
Inteligencia de fuentes abierta (OSINT) para operaciones de ciberseguridad.  
[ 204 ]
Vínculos
ISSN 1794-211X • e-ISSN 2322-9

También podría gustarte