MInESC and MECDad Master
MECDad03 Matemática Aplicada à Ciência dos Dados
Project 2024
Student Learning Outcomes
• The project group (2 students) must choose a project topic of their choice.
• For each project data are given.
• The project group will produce python program.
• The project group will analyze statistics data (with Excel and Python).
• The project group will write a short report for each data set summarizing the characteristics
of your sample in regards to your survey question (What did you find out?)
• The project group will examine the obtained results to interpret what the data implies.
• Project defense and presentation 15 and 16 may 2024 (10 min presentation + 5 min
questions).
Projects subjects
Project 1: Banana Quality analysis
Objective: The project group identify good bananas by their numerical characteristics?
Key Features: Dataset contains characteristics of bananas (size, weight, sweetness, softness, harvest
time, ripeness, acidity and quality)
Size Weight Sweetness Softness HarvestTime Ripeness Acidity Quality
-1.9249682 0.46807805 3.0778325 -1.4721768 0.2947986 2.4355695 0.27129033 Good
-2.4097514 0.48686993 0.34692144 -2.4950993 -0.8922133 2.0675488 0.30732512 Good
Project 2: Global Data on Sustainable Energy
Objective: Explore Global energy consumption patterns and indicators from all nations in this
comprehensive dataset.
Key Features:
• Entity: The name of the country or region for which the data is reported.
• Year: The year for which the data is reported, ranging from 2000 to 2020.
• Access to electricity (% of population): The percentage of population with access to
electricity.
• Access to clean fuels for cooking (% of population): The percentage of the population with
primary reliance on clean fuels.
1
• Renewable-electricity-generating-capacity-per-capita: Installed Renewable energy capacity
per person
• Financial flows to developing countries (US $): Aid and assistance from developed countries
for clean energy projects.
• Renewable energy share in total final energy consumption (%): Percentage of renewable
energy in final energy consumption.
• Electricity from fossil fuels (TWh): Electricity generated from fossil fuels (coal, oil, gas) in
terawatt-hours.
• Electricity from nuclear (TWh): Electricity generated from nuclear power in terawatt-hours.
• Electricity from renewables (TWh): Electricity generated from renewable sources (hydro,
solar, wind, etc.) in terawatt-hours.
• Low-carbon electricity (% electricity): Percentage of electricity from low-carbon sources
(nuclear and renewables).
• Primary energy consumption per capita (kWh/person): Energy consumption per person in
kilowatt-hours.
• Energy intensity level of primary energy (MJ/$2011 PPP GDP): Energy use per unit of GDP
at purchasing power parity.
• Value_CO2_emissions (metric tons per capita): Carbon dioxide emissions per person in
metric tons.
• Renewables (% equivalent primary energy): Equivalent primary energy that is derived from
renewable sources.
• GDP growth (annual %): Annual GDP growth rate based on constant local currency.
• GDP per capita: Gross domestic product per person.
• Density (P/Km2): Population density in persons per square kilometer.
• Land Area (Km2): Total land area in square kilometers.
• Latitude: Latitude of the country's centroid in decimal degrees.
• Longitude: Longitude of the country's centroid in decimal degrees.
Project 3: Air quality data
Objective: Analyze the amount of particulate matter in the air
Key Features: The dataset includes information on air pollution levels, including particulate matter
(PM2.5 and PM10) levels, nitrogen dioxide (NO2), sulfur dioxide (SO2), carbon dioxide (CO2),
ozone (O3), and other pollutants. The data is based on the amount of particulate matter in the air,
which is measured in micrograms per cubic meter (µg/m³). The particulate matter includes dust,
smoke, and other small particles.
2
Project 4: IoT Agriculture 2024
Objective: The Project involved the construction of a smart greenhouse equipped with advanced
technologies for monitoring and controlling environmental conditions.
Key Features:
date (datetime64): The date and time the measurements were recorded.
temperature (int64): The recorded temperature in degrees Celsius.
humidity (int64): The percentage of humidity in the environment.
water_level (int64): The water level as a percentage.
N (int64): The nitrogen level in the soil, scaled from 0 to 255.
P (int64): The phosphorus level in the soil, scaled from 0 to 255.
K (int64): The potassium level in the soil, scaled from 0 to 255.
Fan_actuator_OFF (float64): Indicator for the fan actuator if it is off (0 or 1).
Fan_actuator_ON (float64): Indicator for the fan actuator if it is on (0 or 1).
Watering_plant_pump_OFF (float64): Indicator for the plant watering pump if it is off (0 or 1).
Watering_plant_pump_ON (float64): Indicator for the plant watering pump if it is on (0 or 1).
Water_pump_actuator_OFF (float64): Indicator for the water pump actuator if it is off (0 or 1).
Water_pump_actuator_ON (float64): Indicator for the water pump actuator if it is on (0 or 1).
Watering Watering Water
Fan Fan
Water Plant Plant pump
temperature humidity N P K Actuator Actuator
level Pump Pump actuator
OFF ON
OFF ON OFF
41 63 100 255 255 255 0.0 1.0 1.0 0.0 1.0
41 59 100 255 255 255 0.0 1.0 1.0 0.0 1.0
Project 5: Bacteria Dataset analysis
Objective:
• Predictive modeling to understand factors influencing bacterial habitats and human health
implications.
• Clustering analyses to uncover patterns and relationships among bacterial families and
their characteristics.
• Data visualization projects to illustrate the diversity of bacterial life and its relevance to
ecosystems and health.
Key Features:
• Name: The scientific name of the bacterial species.
• Family: The taxonomic family to which the bacterium belongs.
• Where Found: Natural habitats or common environments where the bacterium is typically
found, including multiple locations if applicable.
• Harmful to Humans: Indicates whether the bacterium is known to have harmful effects on
human health ("Yes" or "No").
3
Project 6: Apple Quality
Objective: Explore the World of Fruits.
• Fruit Classification: Develop a classification model to categorize fruits based on their
features.
• Quality Prediction: Build a model to predict the quality rating of fruits using various
attributes.
This dataset contains information about various attributes of a set of fruits, providing. The dataset
includes details such as fruit ID, size, weight, sweetness, crunchiness, juiciness, ripeness, acidity,
and quality
Key Features:
• A_id: Unique identifier for each fruit
• Size: Size of the fruit
• Weight: Weight of the fruit
• Sweetness: Degree of sweetness of the fruit
• Crunchiness: Texture indicating the crunchiness of the fruit
• Juiciness: Level of juiciness of the fruit
• Ripeness: Stage of ripeness of the fruit
• Acidity: Acidity level of the fruit
• Quality: Overall quality of the fruit
Project 7: Quality Prediction in a Mining Process
Objective: Explore real industrial data and help manufacturing plants to be more efficient. The main
goal is to use this data to predict how much impurity is in the ore concentrate. As this impurity is
measured every hour, if we can predict how much silica (impurity) is in the ore concentrate, we can
help the engineers, giving them early information to take actions (empowering!). Hence, they will
be able to take corrective actions in advance (reduce impurity, if it is the case) and also help the
environment (reducing the amount of ore that goes to tailings as you reduce silica in the ore
concentrate).
Key Features:
• The first column shows time and date range.
• The second and third columns are quality measures of the iron ore pulp right before it is
fed into the flotation plant.
• Column 4 until column 8 are the most important variables that impact in the ore quality in
the end of the process.
• From column 9 until column 22, we can see process data (level and air flow inside the
flotation columns, which also impact in ore quality.
• The last two columns are the final iron ore pulp quality measurement from the lab.
Target is to predict the last column, which is the % of silica in the iron ore concentrate.
4
Project 8: Google Stock Prediction
Objective: Google stock price prediction and forecast
Key Features: This dataset contains 14 columns and 1257 Rows. Each columns are assigned to a
attribute and rows contains the values for that attribute.
The 14 columns are:
1. symbol : - Name of the company (in this case Google).
2. date :- year and date
3. close:- closing of stock value
4. high:- highest value of stock at that day
5. low:- lowest value of stock at that day
6. open:- opening value of stock at that day
7. volume
8. adjClose
9. adjHigh
10. adjLow
11. adjOpen
12. adjVolume
13. divCash
14. splitFactor
Project 9: Data for Admission in the University
Objective: Analysis the admission in the University for Higher Studies
Key Features: This dataset includes various information like
GRE score, TOEFL score, university rating, SOP (Statement of Purpose), LOR (Letter of
Recommendation), CGPA, research and chance of admit.
In this dataset, 400 entries are included.
GRE Scores ( out of 340 )
TOEFL Scores ( out of 120 )
University Rating ( out of 5 )
Statement of Purpose (SOP) and Letter of Recommendation (LOR) Strength ( out of 5 )
Undergraduate GPA ( out of 10 )
Research Experience ( either 0 or 1 )
Chance of Admit ( ranging from 0 to 1 ).
5
Project 10: White Wine Quality
Objective: The two datasets are related to red and white variants of the Portuguese "Vinho Verde"
wine. Due to privacy and logistic issues, only physicochemical (inputs) and sensory (the output)
variables are available (e.g. there is no data about grape types, wine brand, wine selling price, etc.).
These datasets can be viewed as classification or regression tasks. The classes are ordered and not
balanced (e.g. there are many more normal wines than excellent or poor ones). Outlier detection
algorithms could be used to detect the few excellent or poor wines. Also, we are not sure if all input
variables are relevant. So it could be interesting to test feature selection methods.
Key Features: Input variables based on physicochemical tests:
1 - fixed acidity
2 - volatile acidity
3 - citric acid
4 - residual sugar
5 - chlorides
6 - free sulfur dioxide
7 - total sulfur dioxide
8 - density
9 - pH
10 - sulphates
11 - alcohol
Output variable (based on sensory data):
12 - quality (score between 0 and 10)
Project 11: Time Series Room Temperature Data
Objective: Analyze the Dataset is generated with help of an IOT Device data represents room air
temperature values with respect time. In Time Series observations are function of time, each data
corresponds to instance of time, so there is relationship between different data points of dataset a
special case of time series is univariate time series where you have only one feature to deal with
Key Features: The Dataset represents **Univariate time series **values of temperature, which is
specific to time series analysis
1) Hourly _Temp contains mean Supply Air Temperature value in degree centigrade per hour,
2) Datetime shows date and Hour of data recording
6
Project 12: Average global IQ per country with other stats
Objective: 190 countries + small territories with nobel prices, HDI, GNI ...
This dataset contains information about the average IQ in countries around the world, with another
infos like Nobel Prices won collectively in that specific country. I also added more stats like GNI,
HDI and Mean Years of Schooling from another dataset of mine since it provides direct
correlation of why some people in a country are more prone to be more intelligent.
Key Features:
[Link] => Contains data from different measures to measure a country, like GNI,
HDI and Mean Years OF Schooling. Some studies suggest that there's a correlation between
overall quality of life and average iq per person in a country.
IQ_classification.csv => This table distinguishes an IQ score by classifications, for example,
someone might be a genius or a slightly gifted depending in how much IQ points he's got.
Project 13: Human Development Index
Objective: Analyze the human development and other measures concerning most countries in the
world.
Key Features:
• HDI rank (numerical) -> Measures the HDI position in 2021
• Country (character) -> Country name
• Human Development Index (HDI) - 2021 (numerical) -> Value of HDI
• Life expectancy at birth - 2021 (numerical) -> Expected years of living
• Expected years of schooling - 2021 (numerical) -> Years expected for kids/teens to be in
school
• Mean years of schooling - 2021 (numerical) -> Average years attending school
• Gross national income (GNI) per capita - 2021 (numerical) -> Total domestic and foreign
output claimed by residents of a country divided by the total population
2021_development.csv -> Contains columns that measures indicators about a country and its
development.
regions_development.csv -> Contains the mean average summarized by continents and
development groups.
Project 14: Sonar data Analysis
7
Objective: Analyzing sonar data. Sonar data, short for Sound Navigation and Ranging, is a
fascinating technology used primarily in underwater environments. It works by emitting sound
waves and analyzing the echoes that bounce back from objects, creating detailed maps or images
of the underwater terrain. Sonar is crucial for various applications, including underwater
navigation, marine exploration, fishing, and military purposes.
Key Features:
"sonar_data.csv" is likely a file containing data collected from sonar technology.
It could include information such as the depth of the water, the time it takes for sound waves to
travel and return (ping time), and the amplitude of the returned signal.
Each row in the CSV file likely represents a single measurement or observation, while the
columns represent different parameters or features.
The data in "sonar_data.csv" could be used for various purposes, such as analyzing underwater
terrain, detecting underwater objects like submarines or rocks, or studying marine life by
identifying the echoes produced by different species.
Project 15: Random Data on Skills
Objective: Randomly Generated Dataset: Insights into Artificial Intelligence with Randomize.
This dataset provides a collection of randomly generated data aimed at simulating attributes
related to artificial intelligence (AI). It offers a diverse set of attributes commonly associated with
AI, allowing researchers, practitioners, and enthusiasts to explore various aspects of artificial
intelligence through a simulated dataset. Machine Learning Models: Data scientists can utilize this
dataset to build and train machine learning models for tasks such as classification, regression, or
clustering. It serves as a synthetic resource for exploring various aspects of AI through simulated
data. Ideal for educational purposes, exploratory data analysis, and benchmarking AI algorithms.
Key Features:
1) ID: Unique identifier for each entry in the dataset.
2) Name: Randomly generated name for each entry, composed of five uppercase letters.
3) Age: Randomly generated age for each entry, ranging from 1 to 100.
4) Favorite_Color: Randomly assigned favorite color for each entry, chosen from a set of
common colors (Red, Blue, Green, Yellow).
5) Skill_Level: Randomly assigned skill level for each entry, representing proficiency in AI-
related tasks, ranging from 1 to 10.
6) Favorite_Topic: Randomly assigned favorite topic related to artificial intelligence for each
entry, chosen from a set of predefined topics (Artificial Intelligence, Machine Learning,
Natural Language Processing).
7) Favorite_Activity: Randomly assigned favorite activity related to artificial intelligence for
each entry, representing typical tasks or interests within the field (Chatting, Generating
Text, Learning).
8) Random_Fact: A constant value indicating that the data was randomly generated and is not
based on actual individuals. This field includes the statement: "I'm an AI language model
created by OpenAI."
8
Project 16: INTELLIGENT IRRIGATION SYSTEM
Objective: The main purpose is to provide an automatic water supply. The product experience
includes the farmers are satisfied by automatic water supply according to their crops. Product
functions include provide water according to requirement, activate relay motor automatically.
Product features are real-time sensing and control, self-controllable, complete elimination of men
power. Customer revalidation includes farmers who can automatically supply the water to
different crops on their requirements.
The main aim of this project is to generate an intelligent irrigation system that measures the
moisture of soil and helps to take the decision to turns on or off the water supply.
Key Features:
Project 17: Analyze Wine data
Objective: In a classification context, this is a well posed problem with "well behaved" class
structures. A good data set for first testing of a new classifier, but not very challenging.
Key Features:
These data are the results of a chemical analysis of wines grown in the same region in Italy but
derived from three different cultivars. The analysis determined the quantities of 13 constituents
found in each of the three types of wines.
1) Alcohol
2) Malic acid
3) Ash
4) Alcalinity of ash
5) Magnesium
6) Total phenols
7) Flavanoids
8) Nonflavanoid phenols
9) Proanthocyanins
10)Color intensity
11)Hue
9
12)OD280/OD315 of diluted wines
13)Proline
Project 18: Analyzing exam scores
Objective: To study the impact of different factors on test scores
Key Features:
• "gender" - male / female
• "race/ethnicity" - one of 5 combinations of race/ethnicity
• "parent education level" - highest education level of either parent
• "lunch" - whether the student receives free/reduced or standard lunch
• "test prep course" - whether the student took the test preparation course
• "math" - exam score in math
Project 19: Global Electricity Statistics (1980-2021)
Objective:
• This dataset contains the Yearly data from 1980 to 2021 on world electricity statistics.
• Time series forecasting: Time series forecasting to predict the electricity production and
consumption in future.
• Capacity: Find the current statistics of capacity of power plants and forecast the future
values.
• Reduce Electricity Losses: Analyzing patterns of distribution losses and finding methods to
reduce that.
Key Features:
The dataset has total of 4 features and details of each feature is given below (All the information is
in the billion kWh and million kW).
• Country: Name of the Country
• Region: Region of the Country
• Electricity Transaction: Different 7 types of transactions/activity, details of which
is given below.
• Years: Total 41 columns from year 1980 to 2021.
Electricity Activities/Transactions:
• Net Generation (billion kWh): Electricity generation/production
• Net Consumption (billion kWh): Electricity consumption
• Imports (billion kWh): Electricity imports
• Exports (billion kWh): Electricity exports
• Net Imports (billion kWh): Electricity net imports
10
• Installed Capacity (million kW): The maximum amount of electricity that a
generating station (also known as a power plant) can produce under specific
conditions designated by the manufacturer
• Distribution Losses (billion kWh): Transmission and distribution losses refer to the
losses that occur in transmission of electricity between the sources of supply and
points of distribution.
Project 20: Traffic, Driving Style and Road Surface Condition
Objective:
• Low-level parameters acquired by the car via OBD-II and through the
micro-devices embedded in the user smartphone, with the goal of
accurately characterizing the overall system composed by driver,
vehicle and environment
• Predicted attribute: road surface, traffic and driving style
• road surface condition class: SmoothCondition, FullOfHolesCondition, UnevenCondition;
• traffic congestion condition class: LowCongestionCondition, NormalCongestionCondition,
HighCongestionCondition;
• driving style class: EvenPaceStyle, AggressiveStyle.
•
Key Features: 14 numeric features and Data related to the following cars: Peugeot 207 1.4 HDi
(70 CV) and Opel Corsa 1.3 HDi (95 CV)
• altitude change, calculated over 10 seconds;
• current speed value; average speed in the last 60 seconds;
• speed variance in the last 60 seconds;
• speed variation for every second of detection;
• longitudinal acceleration, measured by the smartphone accelerometer and pre-processed
with a low-pass filter;
• engine load, expressed as a percentage;
• engine coolant temperatures in Celsius degree;
• Manifold Air Pressure (MAP), a parameter the internal combustion engine uses to compute
the optimal air/fuel ratio;
• Revolutions Per Minute (RPM) of the engine;
• Mass Air Flow (MAF) Rate measured in g/s, used by the engine to set fuel delivery and
spark timing;
• Intake Air Temperature (IAT) at the engine entrance;
• Vertical acceleration, measured by the smartphone accelerometer and pre-processed with a
low-pass filter;
• Average fuel consumption, calculated as needed liters per 100 km.
11