0% found this document useful (0 votes)
18 views11 pages

Public Sector Data Analysis Insights

The documents provide correct answers to five questions about data analysis, data use by the public sector, open data, and the Access to Information Law. The questions were correctly answered with option "c" in all cases.

Translated by

ScribdTranslations
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
18 views11 pages

Public Sector Data Analysis Insights

The documents provide correct answers to five questions about data analysis, data use by the public sector, open data, and the Access to Information Law. The questions were correctly answered with option "c" in all cases.

Translated by

ScribdTranslations
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Question 1

Correct
Achieved 1.00 out of 1.00

Mark question

Text of the question

About data analysis in the public sector, indicate the correct alternative:

Choose an option:

a. The use of data by public bodies to base decision-making is still a


rarely used practice.

b. The data cannot assist public institutions in the decision-making process.


decisions, even if well worked, monitored, and analyzed.

c. The data may be used by public bodies to assist in detecting


anomalies in the monitoring of indicators or in the improvement of processes.

d. Although the importance of using data for transparency is recognized, there is no record of
its availability by public agencies in Brazil.

The data made available in panels by Brazilian public organs must be used
only for online consultation, meaning there is no possibility for a user to download the data.
from the panel to do your own analyses.
Feedback

Your answer is correct.

The letter 'A' is wrong. More and more public bodies are making decisions based on
in data, whether for anomaly detection, monitoring indicators, or improvement of
processes.

The letter 'B' is wrong. When well worked, monitored, and analyzed, the data serves
to assist institutions in many aspects.

The letter 'C' is correct. Public agencies are increasingly making decisions based on
in the data, whether for anomaly detection, monitoring indicators, or improvement of
processes.

The letter 'D' is wrong. In order to provide more transparency in their actions, the agencies
public entities have made government data available on their websites, following the
provided in the LAI.

The letter 'E' is wrong. Most public agencies make their data available on the internet,
except for those that fall under the hypotheses of confidentiality and personal information. Many
these data are available in panels on various topics, allowing them to be
downloaded for individual analysis.
Question 2
Correct
Reached 1.00 out of 1.00
Mark question

Text of the question

Regarding open data and the Access to Information Law, mark the correct alternative:

Choose an option:

a. The availability of open data cannot be considered a way of giving


transparency to citizens.

b. Only public servers or agents can access, use, modify, and share
open data for any purpose.

c. In Brazil, all government data can be freely accessed by anyone.


citizen.

Only developed countries provide databases via the internet.

We can define open data as that which anyone can freely access.
the, use them, modify them, and share them for any purpose, being subject to,
maximum, to the requirements aimed at preserving its provenance and its openness.
Feedback

Your answer is correct.

The letter 'A' is wrong. The availability of open data is a way to provide transparency.
to the citizens.

The letter 'B' is wrong. Anyone can freely access, use, modify and
share the open data for any purpose.

The letter 'C' is wrong. In Brazil, we have the Access to Information Law (LAI) that defines the
hypotheses of secrecy and personal information, that is, not all government data can
to be accessed.

The letter "D" is incorrect. Several countries provide databases online.


government data classified as open data, aiming to provide transparency to citizens,
whether developed countries or not.

The letter "E" is correct. This is the definition provided on the Brazilian Data Portal.
Open.
Question 3
Correct
Achieved 1.00 out of 1.00

Mark question
Text of the question

Select the option that presents the 4 Vs that characterize Big Data:

Choose an option:

a. Volume, veracity, volatility, and vulnerability.

b. Variety, volume, viability, and validity.

c. Variety, volatility, velocity, and veracity.

d. Volume, variety, velocity, and veracity.

e. Feasibility, speed, accuracy, and variety.


Feedback

Your answer is correct.

The correct alternative is option 'D'. The 4 Vs that characterize Big Data are volume, variety,
speed and truthfulness.
Question 4
Correct
Achieved 1.00 out of 1.00

Mark question

Text of the question

Select the alternative that presents the concept of Hadoop studied in this module:

Choose an option:

Hadoop is a programming language for statistical analysis and presentation.

[Link] is an ecosystem for data analysis that does not allow storage
distributed.

Hadoop is a relational database for storing structured data.

Hadoop is an open source framework for processing and managing small


data volumes.

Hadoop is an open source framework for processing and managing large


data volumes.
Feedback

Your answer is correct.

The letter 'A' is wrong. Hadoop is an open source framework for processing and
management of large volumes of data.

The letter 'B' is wrong. Hadoop is an ecosystem of tools and methods for
distributed storage and analysis of structured and unstructured data.
The letter 'C' is wrong. Hadoop is not a database.

The letter 'D' is incorrect. OHadoop is used for processing and managing large
data volumes and not small ones.

The letter 'E' is correct. Hadoop is an open source framework for processing and
management of large volumes of data.

Question 5
Correct
Reached 1.00 out of 1.00

Mark question

Text of the question

Regarding the definition of data science, mark the correct alternative:

Choose an option:

It is an ecosystem for data analysis and distributed storage.

It consists of developing and applying methods to collect, analyze, and interpret data.

It is the art of extracting knowledge from data to make better decisions,


make predictions and understand the past.

d. Area of study that seeks to give computers the ability to learn.

It is a technique used to group data based on similar characteristics.


Feedback

Your answer is correct.

The letter 'A' is wrong. This is the concept of Hadoop.

The letter 'B' is wrong. Statistics is the branch of science that consists of developing and applying
methods to collect, analyze and interpret data.

The letter "C" is correct. Data science or data science is the art of extracting knowledge from
through data to make better decisions, make forecasts, and understand the past.

The letter 'D' is wrong. This is the concept of machine learning.

The letter 'E' is wrong. This is the concept of unsupervised learning.

Question 1
Correct
Reached 1.00 out of 1.00
Mark question

Text of the question

Regarding machine learning and types of learning, mark the correct alternative:

Choose an option:

a. Machine learning algorithms are applied to a dataset with the aim of


identify existing relationships and generate a model from this data.

[Link] is an activity used in supervised learning algorithms to


group the data that have similar characteristics.

c. Unsupervised learning can be used to solve problems of


classification and regression.

d. Classification results in a numerical output. Regression results in


a categorical/discrete output.

e. Any algorithm can be used in machine learning regardless of the


problem to be solved.
Feedback

Your answer is correct.

The letter 'A' is correct. This is exactly the function of machine learning algorithms.

The letter 'B' is wrong. Clustering is an activity often used for grouping.
the data that have similar characteristics in learning algorithms do not
supervised.

The letter 'C' is wrong. The supervised learning that can be used to solve the
classification and regression problems.

The letter 'D' is wrong. The classification results in a categorical/discrete output. Already the
regression results in a numerical output. The alternative inverted the concepts.

The letter 'E' is wrong. There are many algorithms used in machine learning, for
Yes, it is important to choose the most suitable for the proposed problem.

Question 2
Correct
Achieved 1.00 out of 1.00

Mark question
Text of the question

Regarding supervised or unsupervised learning algorithms, indicate the


alternativacorreta

Choose an option:

a. Linear regression is a classification algorithm based on the nearest neighbors.


next.
b. The objective of the KNN algorithm is to predict the value of a continuous variable.

c. Decision tree is a structure that stores decision rules and has nodes, branches, and
leaves. The nodes represent the variables, the branches represent the possible values of each
the nodes and the leaves represent the final value of a node.

d. The goal of linear regression is to divide the data into groups based on similarity.
data (clusters), that is, we have data that are similar within a group, but different.
when compared to the data from other groups.

Clustering is an activity frequently used to group data that have


distinct characteristics.
Feedback

Your answer is correct.

The wrong letter "A". The classification algorithm that is based on the nearest neighbors is the
KNN.
The wrong letter 'B'. This is the objective of linear regression.

The correct letter is 'C'. This is the definition of a decision tree.

The wrong letter 'D'. This is the goal of the K-means algorithm.

The letter 'E' is wrong. Clustering is an activity frequently used to group data.
that have similar characteristics.

Question 3
Incorrect
Reached 0.00 out of 1.00

Mark question

Text of the question

Regarding the stages and techniques of constructing the machine learning model, indicate the
correct alternative

Choose an option:
Feature selection consists of the technique of creating variables from a dataset.
to improve the model's performance.

Feature engineering is a technique used to select the most relevant attributes that
will be used to train the model.

c. The normalization technique is used to train and validate a model with the same
dataset.

Cross-validation is a technique used in the data preprocessing phase.

e. Data splitting into training and testing is a technique used in the pre-processing phase of
dados.
Feedback

Your answer is incorrect.

The letter 'A' is wrong. The technique described is feature engineering.

The letter 'B' is wrong. The technique described is feature selection.

The letter 'C' is wrong. Cross-validation is the technique used to train and validate a model.
with the same dataset. Normalization is used to standardize the data that
contain variables on different scales.

The letter 'D' is wrong. Cross-validation is a technique used during the phase of
learning (model training).
The letter 'E' is correct. The division of data into training and testing is a technique used in the phase of
data preprocessing, as well as feature selection, feature engineering,
normalization and dimensionality reduction.

Question 4
Incorrect
Reached 0.00 out of 1.00

Mark question

Text of the question

Select the alternative that presents the sequence of steps used in building the model.
predictive

Choose an option:

a. Model evaluation, data preprocessing, learning, and prediction.

b. Data preprocessing, model evaluation, prediction, and learning.

c. Model evaluation, data preprocessing, prediction, and learning.

d. Data preprocessing, learning, model evaluation, and prediction.


e. Prediction, model evaluation, data preprocessing, and learning.
Feedback

Your answer is incorrect.


The correct alternative is letter "D". The steps used in the construction of the predictive model
correspond to the following sequence: data pre-processing, learning
(model construction), model evaluation, and prediction.

Question 1
Correct
Reached 1.00 out of 1.00

Mark question

Text of the question

Identify the option that presents the sequence of steps in the data analysis process.
in R language:

Choose an option:

a. Data acquisition, data exploration, transformation of the obtained data, definition of the
problem, visualization of results and evaluation of results.

b. Data collection, problem definition, transformation of the obtained data and


visualization of results.

c. Problem definition, data collection, transformation of the obtained data, exploration


two dice and visualization of the results.

d. Obtaining data, exploring data, defining the problem, transforming the


data obtained and visualization of the results.

e. Data collection, problem definition, data exploration, data transformation


obtained data, visualization of results, and evaluation of results.
Feedback

Your answer is correct.

The correct alternative is letter 'C'. The stages of the data analysis process in R language
are: problem definition, data collection, transformation of the obtained data, exploration
of the dice and visualization of the results.

Question2
Incorrect
Reached 0.00 out of 1.00
Mark question

Text of the question

Regarding the functions used in the stages of the data analysis process, mark the
correct alternative:

Choose an option:
a. To use the "[Link]()" function, it is necessary to insert the parameter "file", which is the directory of the
file that you wish to upload.
b. The parameter "sep" represents the separator for decimal places.

c. The "view()" function allows for better presentation in graph format.


With the function 'dim()', it is possible to retrieve some information from the dataset, such as value.
mínimo

The function 'summary()' checks the number of observations and columns in the dataset.
Feedback

Your answer is incorrect.

The letter 'A' is correct. To use the '[Link]()' function, it is necessary to insert the parameter
"file", which is the directory of the file to be loaded.

The letter 'B' is incorrect. The parameter 'dec()' represents the decimal separator.

The letter 'C' is wrong. The 'view()' function allows for a better presentation in format of
table.
The letter 'D' is wrong. The function 'dim()' checks the number of observations and columns.
dodataset.
The letter 'E' is wrong. The function 'summary()' retrieves information from the dataset, such as value.
mínimo

Question 3
Incorrect
Reached 0.00 out of 1.00

Mark question

Text of the question


Mark the option that presents the sequence of steps necessary for the construction of
predictive model:

Choose an option:

a. Definition of the problem, data preparation, exploratory analysis, model building,


visualization of results and evaluation of results.

b. Exploratory analysis, problem definition, data acquisition, model construction and


visualization of the results.

c. Data acquisition, data preparation, problem definition, model building


and exploratory analysis.

d. Problem definition, data collection, data preparation, exploratory analysis,


construction of the model and visualization of the results.

e. Data acquisition, problem definition, data preparation, exploratory analysis,


model building, result visualization and result evaluation.
Feedback

Your answer is incorrect.

The correct answer is letter 'D'. The necessary steps for building the predictive model.
are: problem definition, data acquisition, data preparation, exploratory analysis,
construction of the model and visualization of the results.

Question 4
Correct
Reached 1.00 out of 1.00

Mark question

Text of the question

Considering the application of the R language in data analysis for the predictive model,
mark the correct alternative:

Choose an option:
a. It is possible to verify if the division was done correctly with the function 'predict()'.

b. To train the construction of the predictive model, first, it is necessary to split the data into
training and testing. For this, the 'caTools' package should be used, which has the function '[Link]()'.

c. To split the dataset into training and testing sets, the function 'train_test_split()' from the package should be used.
caret

The function "predict()" is used to calculate the model's performance.

e. The parameter "bestTune" of the function "train()" is used to test different values of a
determined parameter.
Feedback
Your answer is correct.

The letter 'A' is wrong. The function 'predict()' is used to generate new predictions. For
to check if the division was done correctly, the function 'dim()' is used.

The letter 'B' is correct. The initial step in constructing the predictive model is the division of
training and test data. This operation is possible using the "caTools" package, in
What is the function '[Link]()'?

The letter 'C' is incorrect. Currently, there is no function 'train_test_split()' in the 'caret' package.

The letter 'D' is wrong. It is the function 'confusionMatrix()' that allows you to calculate performance of
model.

The letter 'E' is wrong. To test different values of a certain parameter, one uses the
parameter "tuneGrid" of the function "train()".

You might also like