Final Doc (Star Classification)
Final Doc (Star Classification)
Submitted By
BACHELOR OF TECHNOLOGY
IN
COMPUTER SCIENCE AND ENGINEERING(AI&ML)
(2021-2025)
Website : [Link] Ph : 08622-212781
BONAFIDE CERTIFICATE
This is to certify that the project work entitled “AUTOMATED STAR TYPE
CLASSIFICATION WITH MACHINE LEARNING USING NASA DATA ” is a Bonafide
record done by [Link] (212U1A3302), [Link] (212U1A3307), [Link]
(212U1A3317), [Link] (212U1A3340) in the Department of Computer Science &
Engineering (AI&ML), Geethanjali Institute of Science and Technology, Nellore and is
submitted to Jawaharlal Nehru Technological University, Anantapur in the partial fulfilment for
the award of B. Tech degree in Computer Science & Engineering(AI&ML). This work has
been carried out under my supervision.
(2021-2025)
ACKNOWLEDGEMENTS
The satisfaction that accompanies the successful completion of the project would be
incomplete without the people who made it possible. Their constant guidance and encouragement
crowned the efforts with success.
We owe our gratitude to Dr. G. SUBBA RAO, [Link], Ph.D., MIE, LMISTE, MSAE,
DIRECTOR, Geethanjali Institute of Science and Technology, Nellore, for his consistent help
and valuable suggestions.
Our sincere special thanks to Dr. P. NAGENDRA KUMAR, Professor & Head of the
Department Computer Science & Engineering(AI&ML), Geethanjali Institute of Science and
Technology, Nellore, for his keen interest, critical, constructive and skillful guidance and
constant encouragement throughout the course and for successful completion of project.
It is indeed our proud privilege to express our deep sense of gratitude and indebtedness
to our guide, Dr. K. SUNDEEP KUMAR., Professor Computer Science & Engineering ,
Geethanjali Institute of Science and Technology, Nellore, for his keen interest, critical,
constructive and skillful guidance and constant encouragement throughout the course and for
successful completion of project.
During the entire course of dissertation work, we received valuable academic inputs as
well as moral support from other departments, general teaching and non-teaching faculty at
GEETHANJALI INSTITUTE OF SCIENCE AND TECHNOLOGY, Nellore. We were
motivated by the uphold and moral encouragement given to us by our beloved parents. Finally,
we wish to express our sincere thanks for all those who helped me directly or indirectly to
complete the work.
PROJECT ASSOCIATES
[Link] (212U1A3302)
[Link] (212U1A3307)
[Link] (212U1A3317)
[Link] (212U1A3340)
TABLE OF CONTENTS
ABSTRACT i
LIST OF FIGURES ii
LIST OF TABLES iii
LIST OF GRAPHS iv
LIST OF ABBREVIATIONS v
1.3 Classification 4
1.4 Star Classification 6
1.4.1 Key Applications 6
1.4.2 Traditional Star Classification 7
1.4.3 Star Classification using Machine Learning 8
1.4.4 Comparative Classification of Different Stars 8
2 LITERATURE SURVEY 10
2.1 Related Work 11
2.2 Research Gaps 18
2.3 Summary 19
3 SYSTEM REQUIREMENT SPECIFICATION 20
4 SYSTEM ANALYSIS 25
4.1 Existing Systems and Its Disadvantages 26
4.2 Proposed Systems and Its Advantages 27
4.3 Dataset 30
i
LIST OF FIGURES
[Link] FIGURE NO FIGURE NAME PAGE NO
1 1.1 Machine Learning 2
2 1.2 Classification Algorithms 5
3 1.3 Types of Stars 6
4 4.1 Random Forest Classifier 29
5 4.2 NASA Dataset 30
ii
LIST OF TABLES
iii
LIST OF GRAPHS
iv
LIST OF ABBREVIATIONS
1 AI Artificial Intelligence
2 ML Machine Learning
3 DL Deep Learning
4 LC Logistic Classifier
v
AUTOMATED STAR TYPE CLASSIFICATION WITH MACHINE LEARNING USING NASA DATA
CHAPTER-1
INTRODUCTION
1. INTRODUCTION
1.1 MACHINE LEARNING
Machine Learning (ML) is a subset of Artificial Intelligence (AI) that focuses on developing
algorithms capable of learning from and making decisions based on data. Unlike traditional
programming, where explicit instructions are provided, ML models learn from patterns within
datasets to improve performance over time. ML has become a powerful tool across numerous
domains, including healthcare, finance, marketing, cybersecurity, and space exploration.
ML systems rely on large amounts of data and computational power to build models that
generalize patterns and make accurate predictions. The primary goal of machine learning is to
minimize human intervention in complex problem-solving by allowing systems to adapt and
improve with experience.
spectral types:
O, B, A, F, G, K, M — from the hottest and most massive (O-type, blue) to the coolest
and smallest (M-type, red).
Each type is subdivided (e.g., G2, K5) for more precision.
Luminosity classes (I to V) indicate the star’s size and brightness, e.g., I for super
giants, V for main-sequence stars like our Sun.
1.4.3 Star Classification Using Machine Learning
Modern astrophysics leverages machine learning (ML) to automate and enhance the star
classification process. Unlike traditional methods that rely heavily on manual spectral
analysis, ML models can efficiently process large astronomical datasets, uncovering patterns
and features beyond human perception.
Machine learning systems are trained on labelled datasets containing stellar parameters,
enabling them to learn classification rules and make predictions on unseen data. These
parameters typically include:
Temperature
Luminosity
Radius
Absolute Magnitude
Color Index (B–V)
Spectral Type (used as labels in supervised learning)
1.4.4 Comparative Classification of Different Star Types
The classification of stars is based on key physical parameters such as temperature,
luminosity, radius, spectral type, color, and absolute magnitude. These parameters help
astronomers distinguish between various stages of stellar evolution and understand the
properties of celestial objects.
This comparative chart summarizes the primary characteristics of major star types:
Red Dwarfs: These are small, cool stars with long lifespans. They fall in the spectral
class M and emit low luminosity.
White Dwarfs: Representing the final stage of low-mass stars, they are dense stellar
remnants with no ongoing fusion, typically very hot but faint.
Main Sequence Stars: These include the Sun and form the majority of stars in the
universe. They range widely in temperature and luminosity, undergoing hydrogen fusion
in their cores.
Giant Stars: Evolved from main sequence stars, giants are larger and brighter, typically
found in spectral types K and M.
Supergiants: These stars are extremely massive and luminous but have shorter lifespans.
They exhibit a wide temperature range and span multiple spectral classes.
Brown Dwarfs: Substellar objects that do not sustain hydrogen fusion, often considered
failed stars. They have very low temperatures and luminosity.
CHAPTER-2
LITERATURE SURVEY
2. LITERATURE SURVEY
In this paper, they explore the application of machine learning methods for classifying
astronomical sources using photometric data, including normal and emission line galaxies
(ELGs; star forming, starburst, AGN, broad-line), quasars, and stars. They utilized samples from
Sloan Digital Sky Survey (SDSS) Data Release 17 (DR17) and the ALLWISE catalogue, which
contain spectroscopically labelled sources from SDSS. Our methodology comprises two parts.
First, they conducted experiments, including three-class, four-class, and seven-class
classifications, employing the Random Forest (RF) algorithm. This phase aimed to achieve
optimal performance with balanced data sets. In the second part, they trained various machine
learning methods, such as k-nearest neighbors (KNN), RF, XGBoost (XGB), voting, and
artificial neural network (ANN), using all available data based on promising results from the first
phase. They achieved an accuracy of 98.83% from XGBoost.
[2] Cody, S. E., Scher, S., McDonald, I., Zijlstra, A., Alexander, E., & Cox, N. L. J. (2024).
“Machine learning based stellar classification with highly sparse photometry data”
[3] Savyanavar, Amit Sadanand, Nikhil Mhala, and Shiv H. Sutar. “Star Galaxy Classification
Using Machine Learning Algorithms and Deep Learning”
In this paper, the star-galaxy dataset was classified into two categories—star and galaxy—
using machine learning algorithms, and their classification performance was compared. It was
observed that the random forest classifier achieved an accuracy of 78%, outperforming other ML
classifiers. To further improve accuracy, a Convolutional Neural Network (CNN) model was
proposed, achieving 92.44% accuracy. Since the CNN model automatically extracts features, it
demonstrated superior classification performance compared to traditional machine learning
approaches.
[4] Tamez Villarreal, J., & Barton, S. (2023). “ Stellar Classification based on Various Star
Characteristics using Machine Learning Algorithms”
In this paper, the authors discuss how the task of stellar classification can be tedious and
lengthy when done manually. They propose expediting stellar classification by creating an
artificial intelligence model to automate the process. As humanity continues to explore the
frontier of the observable universe, automating time-intensive tasks like stellar classification
becomes crucial. The current stellar classification model effectively categorizes stars for research
purposes, aiding in the study of their distribution across the universe. By automating this process,
researchers can allocate more time to exploring the limits of our understanding of space and the
universe, leading to more efficient and advanced astronomical discoveries. In this work they used
various ML Models like Decision Tree, Ridge Classifier and Random Forest and achieved an
accuracy of 94%.
In this paper, the authors discuss how the task of stellar classification can be tedious and
lengthy when done manually. They propose expediting stellar classification by creating an
artificial intelligence model to automate the process. As humanity continues to explore the
frontier of the observable universe, automating time-intensive tasks like stellar classification
becomes crucial. The current stellar classification model effectively categorizes stars for research
purposes, aiding in the study of their distribution across the universe. Popular classification
models including logistic classifier, decision tree, random forest, support vector machine (SVM),
multilayer perceptron, naive bayes, neural networks have proven to be efficient and accurate
applied to many industrial and scientific problems. They achieved an accuracy of 96.9% from
Logistic Regression, 99.3% from Neural Networks.
In this paper, the authors discuss how the task of stellar classification can be tedious and
lengthy when done manually. They propose expediting stellar classification by creating an
artificial intelligence model to automate the process. As humanity continues to explore the
frontier of the observable universe, automating time-intensive tasks like stellar classification
becomes crucial. The paper has used three algorithms (Decision Tree, Random Forest and
Support Vector Machine) to build prediction models. By automating this process, researchers
can allocate more time to exploring the limits of our understanding of space and the universe,
leading to more efficient and advanced astronomical discoveries. They achieved an accuracy of
98%.
[7] Zhao, Z., Wei, J., & Jiang, B. (2022). “Automated Stellar Spectra Classification with
Ensemble Convolutional Neural Network”
[8] Dafonte, C., Rodríguez, A., Manteiga, M., Gómez, Á., & Arcay, B. (2020). “A Blended
Artificial Intelligence Approach for Spectral Classification of Stars in Massive Astronomical
Surveys”
In this paper, the authors analyze and compare the sensitivity and suitability of various
artificial intelligence techniques applied to the Morgan–Keenan (MK) system for stellar
classification. The MK system, based on a sequence of spectral prototypes, classifies stars
according to their effective temperature and luminosity by examining their optical stellar spectra.
This study presents the methodology and results of different intelligent models developed in an
ongoing stellar classification project, including fuzzy knowledge-based systems,
backpropagation, radial basis function (RBF) networks, and Kohonen artificial neural networks.
They achieved the confidence level of 0.99.
[9] Kyritsis, E., Maravelias, G., Zezas, A., Bonfini, P., Kovlakas, K., & Reig, P. (2022). “ A new
automated tool for the spectral classification of OB stars”
In this paper, the authors address the necessity of an automated approach to spectral
classification as the availability of spectroscopic surveys continues to expand. Given the
significance of massive stars, identifying their phenomenological parameters, such as spectral
type, is crucial as these serve as proxies for their physical properties, including mass and
temperature. This study focuses on leveraging the Random Forest (RF) algorithm to develop a
tool for the automated spectral classification of OB-type stars based on their sub-types. They
achieved an accuracy of 70%.
[10] Kuntzer,T., Tewes, M., & Courbin, F. (2016). “Stellar classification from single-band
imaging using machine learning”
In this paper, the authors explore the classification of stars into spectral types using the
shape of their diffraction pattern in a single broad-band image, a method of particular interest for
space-based imaging surveys. They propose a supervised machine learning approach that
incorporates principal component analysis (PCA) for dimensionality reduction, followed by
artificial neural networks (ANNs) to estimate spectral types. The analysis is conducted using
image simulations designed to replicate observations from the Hubble Space Telescope (HST)
Advanced Camera for Surveys (ACS) in the F606W and F814W bands, as well as the Euclid VIS
imager. They achieved the success rate of 0.99 and F1 score of 0.78.
[11] Cody, Seán Enis, Sebastian Scher, Iain McDonald, Albert Zijlstra, Emma Alexander, and
Nick Cox. (2024) “Machine learning based stellar classification with highly sparse
photometry data.”
In this paper, the authors explore the use of machine learning for stellar classification,
addressing the need for automated methods in large-scale astronomical surveys. They employ
Decision Tree, Random Forest, and Support Vector Machine algorithms to classify stars,
galaxies, and quasars while also analyzing artificial intelligence techniques for the Morgan–
Keenan (MK) system. Additionally, they develop a Random Forest-based tool for OB-type stars
and utilize PCA with artificial neural networks for classification based on diffraction patterns.
Finally, XGBoost and PySSED are applied to classify stars using photometric data, addressing
challenges like data sparsity and class imbalance.
[12] Huichaqueo, M. O., & Orrego, R. M. (2022). “Automatic Spectral Classification of Stars
using Machine Learning: An Approach based on the use of Unbalanced Data”
In this paper, the authors present a machine learning-based method for the automatic
spectral classification of stars using data from the latest SDSS release. They propose a
combinatorial approach, integrating spectral data, derived stellar parameters, and calculated
values to create classification patterns. A Random Forest model is developed to classify stars into
six spectral types: A, F, G, K, M, and Carbon stars. To address data imbalance, the model is
trained using original, under-sampled, and over-sampled datasets. Experimental results indicate
that combining multiple data sources improves classification accuracy, with models trained on
augmented data performing best.
[13] Tamez Villarreal, J., & Barton, S. (2023). “Stellar Classification based on Various Star
Characteristics using Machine Learning Algorithms”
In this paper, the authors explore the application of machine learning in astronomy, where
it has been widely used for data processing and predictive modeling. The study employs three
algorithms—Decision Tree, Random Forest, and Support Vector Machine—to develop
classification models for distinguishing stars, galaxies, and quasars. A comparative analysis of
these models reveals that the Random Forest algorithm achieves the highest prediction accuracy
of approximately 98%, demonstrating superior performance in both accuracy and computational
efficiency.
[15] Kateryna, M., & Nataliia, F. (2024). “Peculiarities of methods for determining the class of
stars from photographs using neural networks and machine learning”
In this paper, the authors explore the development of a star classification algorithm using
artificial intelligence based on spectral analysis. The study investigates the application of neural
networks and machine learning methods for automated classification by analyzing stellar
characteristics. With the increasing volume of astronomical data, effective automated methods
are essential. Various machine learning models, including Random Forest, K-Nearest Neighbors,
Gaussian Naive Bayes, and Support Vector Machine, are utilized. Results indicate that K-Nearest
Neighbors and Support Vector Machine achieve the highest classification accuracy, confirming
their effectiveness.
[16] Zhang, J., Lian, J., Yi, Z., Yang, S., & Shan, Y. (2021). “High-Accuracy Guide Star
Catalogue Generation with a Machine Learning Classification Algorithm”
In this paper, the authors investigate methods to enhance the guide star catalogue (GSC)
for improving star identification in satellite attitude determination using star sensors (SSR).
Accurate attitude measurement is crucial for efficient laser link docking in gravitational wave
detection missions. The study examines the relationship between the number of stars in the field
of view, brightness, and sensor accuracy. Various machine learning models, including Decision
Trees, K-Nearest Neighbors, Support Vector Machine, and Neural Networks, are assessed for
GSC optimization. Results indicate that the K-Nearest Neighbors method generates the most
effective GSC, offering superior storage efficiency, uniformity, and completeness for high-
accuracy spacecraft applications.
[17] Kadam, S. S., Chaudhari, K. S., Chaudhari, J. V., Mahajan, T. N., & Kosamkar, P. (2024).
“Star Galaxy Classification Using Deep Learning”
In this paper, the authors present a deep learning-based approach for star and galaxy
classification using Convolutional Neural Networks (CNNs) and the VGG16 architecture. By
leveraging VGG16’s hierarchical feature extraction, the model captures intricate patterns in
astronomical images. A labeled dataset from Kaggle is used for training and evaluation, with
transfer learning applied to fine-tune the model. The VGG16 approach achieves 95.56%
accuracy, while CNN attains 94.15%, demonstrating the effectiveness of deep learning in
distinguishing stars from galaxies.
[18] Hosenie, Z., Lyon, R., Stappers, B., Mootoovaloo, A., & McBride, V. (2020).“ Imbalance
learning for variable star classification”
In this paper, the authors address the challenge of accurately classifying variable stars into
subtypes using machine learning, which is hindered by class imbalance. Building on previous
work with a hierarchical classifier that outperformed traditional methods on Catalina Real-Time
Survey (CRTS) data, they further enhance performance using data augmentation techniques.
Three methods—RASLE, GpFit, and SMOTE—were applied, improving classification accuracy
by 1–4%. The GpFit method showed the highest classification rate, highlighting the need for
better-labeled datasets and enhanced feature extraction.
[19] Hinners, T. A., Tat, K., & Thorp, R. (2018). “ Machine Learning Techniques for Stellar
Light Curve Classification”
In this paper, the authors apply machine learning techniques to predict and classify stellar
properties using noisy and sparse time-series data from Kepler light curves. Over 94 GB of data
from the Mikulski Archive for Space Telescopes (MAST) is preprocessed to classify stars based
on 10 physical properties using both representation learning and feature engineering. While long
short-term memory networks yielded no successful predictions, feature engineering proved
effective, achieving low error (∼2%-4%) for stellar properties and ∼75% accuracy in transit
classification, demonstrating the potential for future astrophysical analysis.
[20] Hoffman, D. I., Harrison, T. E., & McNamara, B. J. (2009). “AUTOMATED VARIABLE
STAR CLASSIFICATION USING THE NORTHERN SKY VARIABILITY SURVEY”
This paper presents the classification of 4,659 variable objects from the Northern Sky
Variability Survey into five distinct variable star classes: Algol/β Lyr systems, W Ursae Majoris
variables, long-period variables, RR Lyr pulsating variables, and short-period δ Scuti stars.
Classification was performed using Fourier coefficients, period analysis, and light-curve
properties, followed by manual verification. The study provides coordinates, periods, infrared
colors, amplitude variations, and prior classifications, excluding 548 previously identified
Algols.
REF
AUTHORS(YEAR) WORK DONE OBSERVATIONS
NO
[1] Zera atgari, et al.(2024) KNN, RF, XGBoost, ANN Accuracy = 98.93%
[4] Tamez Villarreal, et al. Decision Tree, RF, Ridge Accuracy = 94%
(2023) Classifier and SVM
F1 Score of 0.78
2.3 SUMMARY
From the literature survey, several drawbacks had been identified in existing star
classification models. Many studies had struggled with issues such as data imbalance, sensitivity
to noisy observations, and high dependency on handcrafted features. Some models had exhibited
limited generalizability when applied to diverse datasets, while others had faced computational
inefficiencies due to the complexity of deep learning architectures. Additionally, the need for
extensive preprocessing and feature engineering had posed challenges in automating the
classification process effectively.
To overcome these limitations, this work aimed to implement advanced techniques such
as hyperparameter tuning, feature selection optimization, and data augmentation methods. By
fine-tuning model parameters and employing more robust machine learning architectures,
classification accuracy and efficiency could be improved. Additionally, leveraging deep learning
models with automated feature extraction capabilities would reduce dependency on manual
feature engineering. The proposed approach sought to enhance adaptability across various
astronomical datasets while minimizing computational overhead, ultimately improving the
reliability of automated star classification.
CHAPTER-3
SYSTEM
REQUIREMENT
SPECIFICATION
allowing for clear and concise code with minimal complexity. Additionally, Python provides an
extensive standard library, reducing external dependencies and enhancing development
efficiency. In this project, key libraries such as NumPy and pandas facilitate numerical
computation and data manipulation, while matplotlib, seaborn, and Plotly enable effective data
visualization. Machine learning tasks, including classification and data balancing, are handled
using Sklearn and imblearn. Moreover, Joblib and pickle simplify model serialization, ensuring
efficient storage and retrieval. The tqdm library enhances user experience by providing progress
tracking for iterative operations.
pd.read_csv()
The pd.read_csv() function in Pandas is used to read data from a CSV (Comma-Separated
Values) file into a Data Frame. In our project, we use it to load datasets containing stellar
characteristics, such as temperature, luminosity, radius, and spectral classification, for further
analysis and model training.
head() and tail()
The head() function displays the first few rows of the Data Frame, while tail() returns the
last few rows. These functions are useful for quickly inspecting the dataset structure and verifying
its integrity after loading.
shape
The shape attribute provides the number of rows and columns in the dataset. This helps in
understanding the dataset size and ensuring that all necessary features are available for model
training.
isnull()
The isnull() function identifies missing values in the dataset. Handling missing values is
crucial in our project since incomplete data can negatively impact the accuracy of our
classification model.
Matplotlib and Seaborn
Matplotlib and Seaborn are used for data visualization, enabling us to explore relationships
between stellar features. Matplotlib allows us to create line plots, scatter plots, histograms, and
bar charts. Seaborn enhances visualization aesthetics and provides built-in statistical plotting
functions.
bar()
Matplotlib's bar() function is used to plot the frequency distribution of different star types,
helping us understand class imbalances before model training.
Scikit-learn
Scikit-learn provides essential machine learning tools for preprocessing data, training
models, and evaluating performance.
StandardScaler
The StandardScaler class standardizes numerical features by removing the mean and
scaling to unit variance. This ensures that our model treats all features equally, improving
accuracy.
train_test_split()
This function splits the dataset into training and testing sets, allowing us to evaluate model
performance effectively.
accuracy_score()
The accuracy_score() function evaluates how well our model classifies stars based on their
attributes.
Keras and TensorFlow
Keras and TensorFlow power our deep learning models for star classification.
Sequential()
We use the Sequential() function to construct a neural network with multiple layers,
allowing for efficient classification of stars based on their features.
compile()
The compile() function is used to configure the model’s optimizer, loss function, and
evaluation metrics.
add()
The add() function in Keras is used to insert layers such as Dense layers, which help the
network learn complex patterns in stellar data.
fit()
The fit() function trains the model using our dataset, adjusting weights to minimize
classification errors.
CHAPTER-4
SYSTEM ANALYSIS
[Link] ANALYSIS
algorithms—Logistic Classifier, KNN, and Naïve Bayes—with our proposed model, Random Forest
Classifier. The Random Forest enhances performance by capturing complex patterns, handling high-
dimensional data, and reducing overfitting through its ensemble of decision trees. This approach
ensures more accurate and reliable classification, aiding both stellar research and the broader
understanding of cosmic evolution.
Random Forest Classifier
The Random Forest Classifier is an ensemble learning method that builds multiple decision
trees during training and merges their outputs to improve prediction accuracy and control overfitting.
It is robust, handles large datasets with higher dimensionality well, and works efficiently for both
classification and regression tasks.
Algorithm:
Step 1: Data Sampling
Randomly select samples (with replacement) from the original training dataset to form
multiple subsets—this technique is known as bootstrapping.
Step 2: Tree Construction
For each subset, build an individual decision tree. At each node, a random subset of features
is selected to find the best split, ensuring diversity among the trees.
Step 3: Tree Growth
Each tree is grown to its maximum depth without pruning, capturing complex patterns from
its respective sample.
Step 4: Voting Mechanism
Once all trees are trained, predictions for a new input are made by aggregating the outputs
of all trees. For classification, the final class is determined by majority voting.
Step 5: Final Prediction
The model outputs the class with the highest number of votes from all the decision trees,
ensuring a more accurate and stable prediction than any individual tree.
4.3 DATASET
Our project focuses on analysing the NASA Star Dataset from Kaggle, which contains
structured stellar characteristics essential for automated star classification. The dataset includes
multiple attributes that define the physical and spectral properties of stars, enabling machine
learning models to classify them accurately.
The dataset comprises stellar observations with key attributes such as temperature,
luminosity, radius, colour, spectral type, and absolute magnitude. These features help categorize
stars into distinct types, including Main Sequence, White Dwarfs, Giants, and Super giants. To
enhance classification accuracy, we apply multiple machines learning algorithms, comparing
existing methods such as Logistic Classifier, K-Nearest Neighbors (KNN), and Naïve Bayes with
our proposed Random Forest Classifier.
During the training phase, the dataset is utilized to identify patterns and relationships
between these stellar attributes, allowing models to predict star types effectively. After
classification, the dataset is enriched with predicted labels, indicating the assigned star type for
each instance. The implementation of the Random Forest algorithm improves classification
performance by leveraging multiple decision trees, enhancing accuracy and generalization.
By applying machine learning to star classification, our project streamlines the process,
reducing manual effort and increasing reliability. The ability to classify stars automatically aids
in advancing astronomical research and contributes to a better understanding of stellar evolution,
making the system valuable for future astronomical studies.
dataset used for training, NASA Star Dataset from Kaggle, is publicly available, eliminating
the need for expensive data acquisition. Since machine learning models such as Logistic
Classifier, K-Nearest Neighbors (KNN), Naïve Bayes, and Random Forest are
computationally efficient, the system does not require high-end computational resources for
training and classification. Additionally, cloud platforms such as Google Colab provide cost-
effective solutions for large-scale data processing, making the system highly economically
viable.
4.4.2 TECHNICAL FEASIBILITY
The project is technically feasible due to the availability of robust data processing
tools and machine learning frameworks. The NASA Star Dataset is processed using Python-
based libraries, including Pandas for data manipulation, Matplotlib and Seaborn for
visualization, and Scikit-learn for machine learning algorithms. The implementation of
Random Forest as the proposed classification model ensures improved accuracy compared to
existing models like Logistic Classifier, KNN, and Naïve Bayes. Furthermore, the abundance
of astronomical datasets and research papers enhances the accessibility of necessary resources
for developing and refining the classification system. The technical requirements align with
modern machine learning and data science standards, ensuring smooth implementation.
4.4.3 OPERATIONAL FEASIBILITY
The project is operationally feasible due to the availability of structured astronomical
datasets and the ease of implementing machine learning models. The essential skills required
for this project—Python programming, data preprocessing, feature engineering, and model
training—are well-supported by extensive online resources. The automated nature of the
classification process reduces human intervention, ensuring scalability and efficiency.
Additionally, the project follows a structured workflow, including data collection,
preprocessing, model training, evaluation, and deployment, making it adaptable for further
research and practical applications in astronomy and astrophysics.
CHAPTER-5
SYSTEM DESIGN
5. SYSTEM DESIGN
This workflow outlines the structured approach for analysing NASA star data using
machine learning techniques. The process begins with Collecting NASA Star Data, which
involves acquiring datasets containing information on various star characteristics. Next, the
Preprocessing Data step ensures data consistency by handling missing values, scaling numerical
features, and encoding categorical variables. This is followed by Exploratory Data Analysis
(EDA), where statistical and visual techniques are employed to uncover trends, correlations, and
distributions within the dataset.
After gaining insights from EDA, the Train Model step involves selecting and training an
appropriate machine learning model to learn the patterns in the data. Once trained, the model's
Performance is Evaluated using metrics such as accuracy, precision, recall, and F1-score to ensure
reliability. With a well-trained and evaluated model, it proceeds to Make Predictions on unseen
data, classifying new star observations based on learned features. Finally, the Interpret and Save
Results step ensures the model's outputs are documented, analysed, and stored for further scientific
research and applications.
5.2 ARCHITECTURAL DESIGN
Once the model is finalized, it proceeds to make predictions on new or unseen data. The
results of these predictions are then saved and exported, ensuring they are available for further
analysis or deployment. The workflow concludes with the end state, marking the completion of
the machine learning pipeline.
CHAPTER-6
TESTING
6. TESTING
TESTING
Testing is a systematic process crucial for evaluating the functionality and reliability of the
system to ensure it aligns with specified requirements. It involves executing different components
under controlled conditions to identify defects and maintain quality standards. The testing
process for our project is categorized into three main types: Unit Testing, Integration Testing,
and System Testing.
Unit Testing
Unit testing ensures that individual modules or functions in our system perform correctly
in isolation. Each function is tested independently to verify expected outputs and detect early-
stage bugs. For our project, unit tests focus on ensuring correct validation of user inputs, such as
login and signup, testing API endpoints for expected responses, and verifying database
interactions like CRUD operations on stored records. By catching issues early, unit testing
enhances code maintainability and software reliability.
Integration Testing
Integration testing verifies the interaction between different components of the system. It
ensures smooth communication between various modules and services. For our project,
integration testing ensures seamless interaction between the frontend (React) and backend
(Django or [Link]), proper database interactions through API requests, and the smooth
functioning of third-party integrations like email reminders and payment gateways. This type of
testing helps identify inconsistencies or misconfigurations between interconnected modules.
System Testing
System testing evaluates the entire project holistically to ensure it meets all functional and
non-functional requirements. It validates the overall system performance and readiness for
deployment. System testing includes Functional Testing, which checks the correctness of core
features like user authentication, data storage, and retrieval. It also includes Performance Testing,
which measures response times and system behaviour under different loads, and Usability
Testing, which ensures that the interface is user-friendly and accessible.
CHAPTER-7
RESULTS AND
SCREENSHOTS
Where:
Number of Correct Predictions refers to instances where the model correctly classifies the
given data.
Total Number of Predictions is the sum of all instances tested by the model.
Precision
Precision is a metric used to measure the proportion of correctly predicted positive cases out of all
predicted positive cases. It helps in evaluating how many of the predicted positive instances are
actually relevant. The formula for precision is:
Where:
True Positives (TP) are cases where the model correctly predicts a positive outcome.
False Positives (FP) are cases where the model incorrectly predicts a positive outcome.
Recall
Recall, also known as Sensitivity, measures the proportion of actual positive cases that were correctly
identified by the model. It is useful in scenarios where missing a positive instance is critical. The
formula for recall is:
Where:
True Positives (TP) are correctly predicted positive cases.
False Negatives (FN) are actual positive cases that the model incorrectly classified as negative.
F1 Score
F1 Score is a harmonic mean of precision and recall, providing a balanced evaluation when the
dataset is imbalanced. It is calculated as:
Where:
Precision measures the relevance of positive predictions.
Recall measures the completeness of positive predictions.
Confusion Matrix
A confusion matrix is a table used to evaluate the performance of a classification model by showing
the distribution of predictions across different classes. It consists of four components:
Where:
True Positives (TP): Correctly classified positive instances.
False Positives (FP): Incorrectly classified negative instances as positive.
True Negatives (TN): Correctly classified negative instances.
False Negatives (FN): Incorrectly classified positive instances as negative.
Classification Report
A classification report provides a summary of different evaluation metrics such as precision, recall,
F1 score, and support for each class. It gives a comprehensive overview of model performance across
all categories.
The classification report includes:
Precision for each class
Recall for each class
F1 Score for each class
Support (number of instances for each class)
Table 7.2. presents that the Logistic Classifier model performed exceptionally well with an
overall accuracy of 98%, showing perfect precision, recall, and F1-score for Brown Dwarfs, Main
Sequence stars, Super Giants, and Hyper Giants. However, minor misclassifications occurred with
Red Dwarfs (precision: 0.88) and White Dwarfs (recall: 0.88), which slightly reduced the model's
overall performance. The high recall value for Red Dwarfs (1.00) indicates that the model correctly
identified all Red Dwarfs but had a few false positives in classifying other star types as Red Dwarfs.
Despite this, the model demonstrates strong classification performance across all categories, making
it a reliable method for stellar classification.
Table 7.3. presents that the Naive Bayes model achieved an overall accuracy of 96%, slightly
lower than Logistic Classifier. The model perfectly classified Brown Dwarfs and Hyper Giants
(precision & recall = 1.00), showing its effectiveness in differentiating these star types. However,
Red Dwarfs had a recall of only 0.86, meaning some Red Dwarfs were incorrectly classified as other
star types. Additionally, Main Sequence and Super Giants had lower precision (0.89), which suggests
that the model sometimes misclassified other stars as these types. Although the performance is strong,
the slightly lower precision and recall values for Red Dwarfs, Main Sequence, and Super Giants
indicate that Naive Bayes may struggle with feature overlap among similar stellar categories.
Table [Link] that K-Nearest Neighbors (KNN) model performed slightly lower than
Logistic Classifier and Naive Bayes, achieving an accuracy of 94%. The model performed well in
predicting Brown Dwarfs and Hyper Giants (precision & recall = 1.00), but Super Giants and White
Dwarfs had lower recall values (0.88), indicating misclassification of some samples. The lower recall
for Brown Dwarfs (0.91) suggests that some Brown Dwarfs were incorrectly classified into other
categories. Despite these limitations, KNN still demonstrates a high overall classification
performance, though it struggles slightly when dealing with overlapping feature distributions in the
dataset.
Table 7.5. presents that the Random Forest model delivered a perfect performance across all
metrics, achieving 100% accuracy, precision, recall, and F1-score for every star type. Unlike the
other models, Random Forest did not misclassify any star, making it the most reliable and effective
model for stellar classification. The perfect classification indicates that the model successfully
captures intricate patterns within the dataset and differentiates between Red Dwarfs, Brown Dwarfs,
White Dwarfs, Main Sequence, Super Giants, and Hyper Giants with absolute certainty. This
outstanding performance solidifies Random Forest as the best choice for automated star classification
using machine learning.
Figure 7.1 presents the confusion matrices for four different machine learning models—
Logistic Classifier, Naïve Bayes, K-Nearest Neighbors (KNN), and Random Forest—to analyse their
classification performance in identifying star types. Each confusion matrix visually represents the
true class labels on the y-axis and the predicted class labels on the x-axis, showing the model’s ability
to correctly classify each star type while also highlighting misclassifications. The Logistic Classifier
model (Fig. 7.1a) performs well, but it misclassifies one White Dwarf as a Red Dwarf and one Red
Dwarf as a White Dwarf, indicating a slight overlap between these categories. Similarly, Naïve Bayes
(Fig. 7.1b) exhibits a similar trend, with one Red Dwarf misclassified as a White Dwarf and one
White Dwarf misclassified as a Main Sequence star. These errors suggest that both Logistic Classifier
and Naïve Bayes may struggle with distinguishing overlapping features in some star classes.
The K-Nearest Neighbors (Fig. 7.1c) model shows slightly more misclassification errors
compared to the previous two models, misclassifying a Brown Dwarf as a White Dwarf and a Red
Dwarf as a White Dwarf. Additionally, one Super Giant was misclassified as a Main Sequence star,
further reducing its reliability. In contrast, the Random Forest model (Fig. 7.1d) achieves perfect
classification, correctly predicting all instances without any misclassification errors. This indicates
that Random Forest effectively learns complex relationships in the dataset and provides the best
classification performance. The perfect diagonal alignment in the confusion matrix of Random Forest
confirms its 100% accuracy, making it the most suitable model for stellar classification tasks.
(a) (b)
(c) (d)
Fig 7.1. Confusion Matrices. (a) Logistic Classifier. (b) Naïve Bayes.
(c) K-Nearest Neighbor. (d) Proposed Random Forest.
Table 7.6. presents the True Positives (TP), True Negatives (TN), False Positives (FP), and
False Negatives (FN) for different machine learning models used in star classification. Logistic
Classifier achieved 47 TP and 192 TN, with only 1 FP and 1 FN, indicating a highly reliable
classification with minimal errors. Naïve Bayes performed slightly worse, with 46 TP, 191 TN, 2 FP,
and 1 FN, suggesting a marginal increase in misclassification. K-Nearest Neighbors (KNN) exhibited
the highest error rate among the models, with 44 TP, 189 TN, 4 FP, and 2 FN, highlighting its
struggles in distinguishing between certain star types. In contrast, the Random Forest model achieved
perfect classification, with 48 TP, 192 TN, 0 FP, and 0 FN, demonstrating its superior learning
capability and robustness in identifying stellar types. This analysis confirms Random Forest as the
most effective model for star classification, followed by Logistic Classifier and Naïve Bayes, with
KNN being the least reliable due to its higher misclassification rate.
Model TP TN FP FN
7.3 SCREENSHOTS
CHAPTER-8
CONCLUSION
8. CONCLUSION
This Star Type Classification System leverages the power of the Random Forest algorithm to
automate the categorization of stars, outperforming traditional models such as Logistic Classifier,
Naïve Bayes, and K-Nearest Neighbors (KNN). Unlike these models, Random Forest efficiently
handles complex, non-linear relationships in astronomical data, reducing overfitting and improving
overall classification accuracy. By processing vast datasets with precision, it enhances not only star
classification but also contributes to exoplanet detection, stellar evolution studies, and broader
cosmic research. The integration of machine learning in this domain minimizes human errors,
optimizes computational efficiency, and enables large-scale analysis, paving the way for future
advancements in automated celestial studies. This project underscores the transformative potential
of AI in astronomy, offering a scalable and robust solution for astronomical data analysis.
CHAPTER-9
FUTURE SCOPE
9. FUTURE SCOPE
The project can be significantly improved by integrating deep learning techniques such as
Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs) to enhance
classification accuracy and model robustness. Expanding the dataset by incorporating high-precision
astronomical surveys like Gaia and the Sloan Digital Sky Survey (SDSS) will provide richer and
more diverse data, leading to better generalization of the model. Additionally, the system can be
extended to support exoplanet detection by analysing stellar light curves and transit events, aiding in
the discovery of habitable planets beyond our solar system. Implementing automated anomaly
detection will allow the identification of rare celestial objects and unknown star types, contributing
to new astronomical discoveries. Collaboration with space agencies such as NASA and ESA can
further enhance the project by leveraging high-resolution datasets and computational resources for
large-scale star mapping, deep-space research, and real-time astronomical event monitoring. This
integration will not only refine the accuracy and efficiency of star classification but also open new
frontiers in astrophysics and space exploration.
CHAPTER-10
BIBLIOGRAPHY
[Link]
[1] Zeraatgari, Fatemeh Zahra, Fatemeh Hafezianzadeh, Yanxia Zhang, Liquan Mei, Ashraf
Ayubinia, Amin Mosallanezhad, and Jingyi Zhang. "Machine learning-based photometric
classification of galaxies, quasars, emission-line galaxies, and stars." Monthly Notices of the
Royal Astronomical Society 527, no. 3 (2024): 4677-4689.
[2] Cody, S. E., Scher, S., McDonald, I., Zijlstra, A., Alexander, E., & Cox, N. L. J. (2024).
Machine learning based stellar classification with highly sparse photometry data. arXiv.
[Link]
[3] Savyanavar, Amit Sadanand, Nikhil Mhala, and Shiv H. Sutar. "Star Galaxy
Classification Using Machine Learning Algorithms and Deep Learning." International
Journal on Information Technologies & Security 15, no. 2 (2023)..
[4] Tamez Villarreal, J., & Barton, S. (2023). Stellar Classification based on Various Star
Characteristics using Machine Learning Algorithms. In Journal of Student Research (Vol.
12, Issue 1). [Link]
[6] Qi, Z. (2022). Stellar Classification by Machine Learning. In A. Luqman, Q. Zhang, &
W. Liu (Eds.), SHS Web of Conferences (Vol. 144, p. 03006). EDP Sciences.
[Link]
[7] Zhao, Z., Wei, J., & Jiang, B. (2022). Automated Stellar Spectra Classification with
Ensemble Convolutional Neural Network. In K. Yakut (Ed.), Advances in Astronomy
(Vol. 2022, pp. 1–7). Hindawi Limited. [Link]
[8] Dafonte, C., Rodríguez, A., Manteiga, M., Gómez, Á., & Arcay, B. (2020). A Blended
Artificial Intelligence Approach for Spectral Classification of Stars in Massive
Astronomical Surveys. In Entropy (Vol. 22, Issue 5, p. 518). MDPI AG.
[Link]
[9] Kyritsis, E., Maravelias, G., Zezas, A., Bonfini, P., Kovlakas, K., & Reig, P. (2022). A
new automated tool for the spectral classification of OB stars. In Astronomy &
for variable star classification. In Monthly Notices of the Royal Astronomical Society (Vol. 493,
Issue 4, pp. 6050–6059). Oxford University Press (OUP). [Link]
[19] Hinners, T. A., Tat, K., & Thorp, R. (2018). Machine Learning Techniques for Stellar Light Curve
Classification. In The Astronomical Journal (Vol. 156, Issue 1, p. 7). American Astronomical
Society. [Link]
[20] Hoffman, D. I., Harrison, T. E., & McNamara, B. J. (2009). AUTOMATED VARIABLE STAR
CLASSIFICATION USING THE NORTHERN SKY VARIABILITY SURVEY. In The
Astronomical Journal (Vol. 138, Issue 2, pp. 466–477). American Astronomical Society.
[Link]