Tested AI Models:
[Link] Forests:
1.1. Importing Libraries:
import pandas as pd (Handles reading and processing the CSV file)
import numpy as np(Used for mathematical operations)
from [Link] import StandardScaler(Normalizes feature values to improve model
performance)
from sklearn.model_selection import train_test_split(Splits data into training and test sets)
from [Link] import RandomForestRegressor(Machine learning model used for prediction)
from [Link] import mean_squared_error, mean_absolute_error, r2_score(Used to evaluate
model accuracy)
import [Link] as plt(Used to visualize predictions vs. actual values)
1.2. Reading and Preparing Data:
df = pd.read_csv('vm_metrics.csv'): Reads the CSV file (vm_metrics.csv) into a pandas DataFrame
(df).
df['Timestamp'] = pd.to_datetime(df['Timestamp'], unit='s'): Converts the 'Timestamp' column
into a datetime format for time-based feature extraction
1.3. Feature Engineering:
. Extracts time-based features from the Timestamp column:
def create_features(data):
data['hour'] = data['Timestamp'].[Link]
data['day'] = data['Timestamp'].[Link]
data['month'] = data['Timestamp'].[Link]
data['dayofweek'] = data['Timestamp'].[Link]
1.4. Creating Lag Features(previous measurements):
data['lag_1'] = data['Value'].shift(1)
data['lag_2'] = data['Value'].shift(2)
data['rolling_mean'] = data['Value'].rolling(window=3).mean()
return [Link]()
lag_1 and lag_2 → Stores the previous values to help the model learn from past trends.
rolling_mean → Computes the average value over the last 3 records to smooth out fluctuations.
dropna() → Removes rows with NaN values (caused by lagging)
1.5. Training and Evaluating the Model:
def train_evaluate_model(df_metric, metric_name):
df_processed = create_features(df_metric)