0% found this document useful (0 votes)
1 views6 pages

Tested AI Model4

The document outlines the process of training and evaluating a Random Forest model for predicting various metrics such as CPU usage, memory usage, and network bytes using Python libraries. It details steps including data preparation, feature engineering, model training, and performance evaluation, highlighting results such as R² scores and potential overfitting issues. Additionally, it mentions the use of XGBoost for further model training, though details on that are not provided.

Uploaded by

anas
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
1 views6 pages

Tested AI Model4

The document outlines the process of training and evaluating a Random Forest model for predicting various metrics such as CPU usage, memory usage, and network bytes using Python libraries. It details steps including data preparation, feature engineering, model training, and performance evaluation, highlighting results such as R² scores and potential overfitting issues. Additionally, it mentions the use of XGBoost for further model training, though details on that are not provided.

Uploaded by

anas
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Tested AI Models:

[Link] Forests model:

1.1. Importing Libraries:

import pandas as pd (Handles reading and processing the CSV file)

import numpy as np(Used for mathematical operations)

from [Link] import StandardScaler(Normalizes feature values to improve model


performance)

from sklearn.model_selection import train_test_split(Splits data into training and test sets)

from [Link] import RandomForestRegressor(Machine learning model used for prediction)

from [Link] import mean_squared_error, mean_absolute_error, r2_score(Used to evaluate


model accuracy)

import [Link] as plt(Used to visualize predictions vs. actual values)

1.2. Reading and Preparing Data:

df = pd.read_csv('vm_metrics.csv'): Reads the CSV file (vm_metrics.csv) into a pandas DataFrame


(df).

df['Timestamp'] = pd.to_datetime(df['Timestamp'], unit='s'): Converts the 'Timestamp' column into


a datetime format for time-based feature extraction

1.3. Feature Engineering:

. Extracts time-based features from the Timestamp column:

def create_features(data):

data['hour'] = data['Timestamp'].[Link]

data['day'] = data['Timestamp'].[Link]

data['month'] = data['Timestamp'].[Link]

data['dayofweek'] = data['Timestamp'].[Link]

1.4. Creating Lag Features(previous measurements):

data['lag_1'] = data['Value'].shift(1)

data['lag_2'] = data['Value'].shift(2)

data['rolling_mean'] = data['Value'].rolling(window=3).mean()

return [Link]()

lag_1 and lag_2 → Stores the previous values to help the model learn from past trends.

rolling_mean → Computes the average value over the last 3 records to smooth out fluctuations.

dropna() → Removes rows with NaN values (caused by lagging)


1.5. Training and Evaluating the Model: Calls create_features() to add new time-based and lag
features to the data

def train_evaluate_model(df_metric, metric_name):

df_processed = create_features(df_metric)

[Link] Selection & Scaling:

X = df_processed[['hour', 'day', 'month', 'dayofweek', 'lag_1', 'lag_2', 'rolling_mean']]

y = df_processed['Value']

Defines X (features) and y (target variable):

X = Time-based features (hour, day, etc.) + Lag values.

y = The actual metric value we want to predict.

scaler = StandardScaler()

X_scaled = scaler.fit_transform(X)

StandardScaler() → Normalizes the values so that large numerical differences don’t affect model
performance.

1.5.2. Splitting Data into Training & Test Sets:

X_train, X_test, y_train, y_test = train_test_split(X_scaled, y, test_size=0.2, random_state=42)

.Splits data into:

• X_train (80%) → Used for training.

• X_test (20%) → Used for testing.

• y_train, y_test → Corresponding target values.

. random_state=42 → Ensures reproducibility (same split every time).

1.5.3. Training the Random Forest Model:

rf = RandomForestRegressor(n_estimators=100, random_state=42)

.Creates a RandomForestRegressor model:

• n_estimators=100 → Uses 100 decision trees for prediction.

• random_state=42 → Ensures reproducibility

[Link](X_train, y_train)

Trains the model (fit) on X_train and y_train.

1.5.4. Making Predictions:

y_pred = [Link](X_test): Uses the trained model to predict values for X_test

1.5.5. Evaluating Model Performance:


print(f'\nMetrics for {metric_name}:')

print(f'RMSE: {[Link](mean_squared_error(y_test, y_pred)):.2f}'): Measures how much the


predictions deviate from actual values.

print(f'MAE: {mean_absolute_error(y_test, y_pred):.2f}'): Measures the absolute difference


between actual and predicted values.

print(f'R2: {r2_score(y_test, y_pred):.2f}'): Measures how well the model explains the variability
(closer to 1.0 is better).

1.5.6. Plotting Actual vs. Predicted Values: Creates a line plot comparing actual vs. predicted values.

[Link](figsize=(12,6))

[Link](y_test.values, label='Actual')

[Link](y_pred, label='Predicted')

[Link](f'{metric_name} - Actual vs Predicted')

[Link]()

[Link]()

1.6. Looping Through Multiple Metrics: List of metrics that the script will process separately.

metrics = ['CPU Usage', 'Memory Used', 'Disk Usage', 'Network RX Bytes', 'Network TX Bytes', 'Device
Status']

1.6.1. Training & Evaluating for Each Metric:

for metric in metrics:

df_metric = df[df['Metric'] == metric].copy()

train_evaluate_model(df_metric, metric)

.Filters the dataset (df[df['Metric'] == metric]) to process each metric independently.

.Calls train_evaluate_model() for each metric.


[Link] of the model:

CPU:

The model performs well in predicting CPU usage, with an R² score of 0.92, meaning it explains 92%
of the variance in CPU usage. The RMSE (2.52) and MAE (1.05) indicate small errors, suggesting a
fairly accurate prediction.

Memory:

The R² score of 0.97 is excellent, meaning the model explains 97% of the variance in memory usage.
However, the RMSE and MAE values are very large, likely because the memory values are naturally in
the range of millions of bytes. If the dataset contains large values, these errors may still be
acceptable.

Disk:

The R² score of 1.00 suggests that the model fits the data almost perfectly, which is unusual. The
RMSE and MAE values are also extremely low, suggesting that disk usage is highly predictable and has
very little fluctuation in the data. However, an R² of 1.00 could indicate that the model might be
overfitting.

Network RX Bytes:
The model perfectly fits the data with an R² score of 1.00, meaning it captures all variance in the
network RX bytes. The RMSE and MAE values are large, but if the actual values are in the range of
millions of bytes, the error could be reasonable. The perfect R² might indicate the model has either
a very strong correlation in the data or possible overfitting.

Network TX Bytes:

Similar to Network RX Bytes, this model also achieves an R² of 1.00, meaning it explains 100% of the
variance. The RMSE and MAE are quite large, but their significance depends on the scale of actual
values.

Overfitting Reasons in Disk and Network Metrics(possible issues):

. disk usage and network metrics follow a repetitive or nearly constant pattern,and the dataset has
very little change over time the model learns these patterns perfectly, leading to an R² of 1.00.

. the dataset is too small or the test data is very similar to the training data, the model might
memorize instead of learning general trends.

. the train/test split is not random (e.g., if all test data comes from a period similar to training
data), the model performs well on the test set but may fail in real-world scenarios.

[Link] model:

You might also like