0% found this document useful (0 votes)
3 views14 pages

Machine Learning for Autism Detection

The document discusses a research project focused on detecting Autism Spectrum Disorder (ASD) using machine learning techniques, particularly Support Vector Machines (SVM) and Convolutional Neural Networks (CNN). It highlights the limitations of traditional diagnostic methods and proposes a systematic approach that includes data collection, preprocessing, model development, and evaluation to enhance early detection of ASD in children. The goal is to create a low-cost diagnostic tool that can be used in areas lacking expert resources, ultimately improving treatment outcomes for affected children.

Uploaded by

l7272551
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views14 pages

Machine Learning for Autism Detection

The document discusses a research project focused on detecting Autism Spectrum Disorder (ASD) using machine learning techniques, particularly Support Vector Machines (SVM) and Convolutional Neural Networks (CNN). It highlights the limitations of traditional diagnostic methods and proposes a systematic approach that includes data collection, preprocessing, model development, and evaluation to enhance early detection of ASD in children. The goal is to create a low-cost diagnostic tool that can be used in areas lacking expert resources, ultimately improving treatment outcomes for affected children.

Uploaded by

l7272551
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Page 1 of 14 - Cover Page Submission ID trn:oid:::1:3398019136

Arun K H
Autism Spectrum Disorder Detection Using Machine Learning
Articles

Acharya LRC

Acharya Institute of Technology, Bengaluru

Document Details

Submission ID

trn:oid:::1:3398019136 8 Pages

Submission Date 4,492 Words

Nov 4, 2025, 12:50 PM GMT+5:30


26,811 Characters

Download Date

Nov 4, 2025, 12:51 PM GMT+5:30

File Name

ASD_final_paper3_-_ARYA_TADAS_AIT22BEIS142.docx

File Size

478.6 KB

Page 1 of 14 - Cover Page Submission ID trn:oid:::1:3398019136


Page 2 of 14 - Integrity Overview Submission ID trn:oid:::1:3398019136

16% Overall Similarity


The combined total of all matches, including overlapping sources, for each database.

Filtered from the Report


Bibliography

Quoted Text

Match Groups Top Sources

69 Not Cited or Quoted 16% 10% Internet sources


Matches with neither in-text citation nor quotation marks
14% Publications
1 Missing Quotations 0% 4% Submitted works (Student Papers)
Matches that are still very similar to source material

0 Missing Citation 0%
Matches that have quotation marks, but no in-text citation

0 Cited and Quoted 0%


Matches with in-text citation present, but no quotation marks

Integrity Flags
0 Integrity Flags for Review
Our system's algorithms look deeply at a document for any inconsistencies that
No suspicious text manipulations found. would set it apart from a normal submission. If we notice something strange, we flag
it for you to review.

A Flag is not necessarily an indicator of a problem. However, we'd recommend you


focus your attention there for further review.

Page 2 of 14 - Integrity Overview Submission ID trn:oid:::1:3398019136


Page 3 of 14 - Integrity Overview Submission ID trn:oid:::1:3398019136

Match Groups Top Sources

69 Not Cited or Quoted 16% 10% Internet sources


Matches with neither in-text citation nor quotation marks
14% Publications
1 Missing Quotations 0% 4% Submitted works (Student Papers)
Matches that are still very similar to source material

0 Missing Citation 0%
Matches that have quotation marks, but no in-text citation

0 Cited and Quoted 0%


Matches with in-text citation present, but no quotation marks

Top Sources
The sources with the highest number of matches within the submission. Overlapping sources will not be displayed.

1 Publication

Sondekola Rudra Swamy, Luminita-Ioana Cotîrlă. "A New Pseudo-Type κ-Fold Sym… 1%

2 Publication

S.P. Jani, M. Adam Khan. "Applications of AI in Smart Technologies and Manufactu… 1%

3 Publication

Pushpa Choudhary, Sambit Satpathy, Arvind Dagur, Dhirendra Kumar Shukla. "Re… <1%

4 Publication

"Applications of Computational Intelligence in Management and Mathematics I", … <1%

5 Internet

[Link] <1%

6 Publication

Mohamed Rochdi Keffala, Safoua Mahjoub. "Forecasting Spanish Olive Oil Prices … <1%

7 Internet

[Link] <1%

8 Publication

"Data Science and Big Data Analytics", Springer Science and Business Media LLC, … <1%

9 Publication

Shankar Babu, Mahesh Babu Kota. "Synergies in Smart and Virtual Systems using… <1%

10 Internet

[Link] <1%

Page 3 of 14 - Integrity Overview Submission ID trn:oid:::1:3398019136


Page 4 of 14 - Integrity Overview Submission ID trn:oid:::1:3398019136

11 Internet

[Link] <1%

12 Internet

[Link] <1%

13 Publication

H L Gururaj, Francesco Flammini, V Ravi Kumar, N S Prema. "Recent Trends in He… <1%

14 Publication

Khlefat, Hamza. "Recognizing Counterfeit Brand Logo Using Machine Learning", … <1%

15 Internet

[Link] <1%

16 Internet

[Link] <1%

17 Internet

[Link] <1%

18 Internet

[Link] <1%

19 Student papers

Technische Hochschule Deggendorf <1%

20 Publication

Ashok Kumar, Geeta Sharma, Anil Sharma, Pooja Chopra, Punam Rattan. "Advanc… <1%

21 Publication

Puneet Bawa, Virender Kadyan, Archana Mantri, Harsh Vardhan. "Investigating M… <1%

22 Publication

Nazmul Siddique, Mohammad Shamsul Arefin, K. M. Azharul Hasan, M. Shamim K… <1%

23 Student papers

Manchester Metropolitan University <1%

24 Internet

[Link] <1%

Page 4 of 14 - Integrity Overview Submission ID trn:oid:::1:3398019136


Page 5 of 14 - Integrity Overview Submission ID trn:oid:::1:3398019136

25 Internet

[Link] <1%

26 Internet

[Link] <1%

27 Internet

[Link] <1%

28 Internet

[Link] <1%

29 Internet

[Link] <1%

30 Internet

[Link] <1%

31 Internet

[Link] <1%

32 Internet

[Link] <1%

33 Internet

[Link] <1%

34 Internet

[Link] <1%

35 Publication

Ibtissam Essadik, Anass Nouri, Raja Touahni, Romain Bourcier, Florent Autrussea… <1%

36 Publication

Rafael Ferreira, George D.C. Cavalcanti, Fred Freitas, Rafael Dueire Lins, Steven J. … <1%

37 Publication

Wisdom Richard Mgomezulu, Paul Thangata, Bertha Mkandawire, Nana Amoah. "… <1%

38 Internet

[Link] <1%

Page 5 of 14 - Integrity Overview Submission ID trn:oid:::1:3398019136


Page 6 of 14 - Integrity Overview Submission ID trn:oid:::1:3398019136

39 Internet

[Link] <1%

40 Internet

[Link] <1%

41 Publication

Fahima Hajjej, Sarra Ayouni, Manal Abdullah Alohali, Mohamed Maddeh. "Novel F… <1%

42 Internet

[Link] <1%

Page 6 of 14 - Integrity Overview Submission ID trn:oid:::1:3398019136


Page 7 of 14 - Integrity Submission Submission ID trn:oid:::1:3398019136
13
Autism Spectrum Disorder Detection Using
Machine Learning
Prof. Arun K H
15 Department of Information Science and Aastha
Engineering Department of Information Science and
1
Acharya Institute of Technology Engineering
Bengaluru, India Acharya Institute of Technology
arun2976@[Link] Bengaluru, India
[Link]@[Link]
Namrata R N
Department of Information Science and
Engineering Arya Tadas
1 Acharya Institute of Technology Department of Information Science and
1
Bengaluru, India Engineering
[Link]@[Link] Acharya Institute of Technology
Bengaluru, India
Himalaya K S [Link]@[Link]
1 Department of Information Science and
Engineering
Acharya Institute of Technology
Bengaluru, India
[Link]@[Link]

31 Abstract—Autism Spectrum Disorder(ASD), also I. INTRODUCTION


known as a neurodevelopmental disorder,
impacts how a person communicates, behaves, Autism Spectrum Disorder Detection (ASD), also known as
and interacts socially. The detection of autism neurodevelopmental disorder impacts how a person
13
11 communicates, behaves and interacts socially. It is called a
should be done at an early stage. Traditional “spectrum” disorder because the symptoms and their severity
diagnosis methods for autism detection are time- vary greatly from person to person. The facial features of
consuming and need expert evaluation, but using children with ASD can often be distinguished more easily
modern techniques like machine learning and than those of typically developing children. Detecting ASD
deep learning, autism detection can be done faster at an early stage is crucial, as it allows children and their
families to receive timely support, therapy and treatment that
and can help doctors and parents in making better can greatly enhance their quality of life. Traditional methods
decisions. The project on autism prediction for diagnosing ASD involve behavioral assessments,
focuses on using machine learning methods for questionnaires, and clinical observations conducted by
Behavioral analysis and convolutional neural trained professionals. However, these methods can be time-
networks (CNNs) to analyze image data for early consuming and are not always available in rural or under-
resourced areas. As a result, many children may remain
autism diagnosis. This research intends to build a undiagnosed or be diagnosed later, leading to delays in
model that can detect autism using computer- receiving proper treatment and support.
based methods. The datasets includes Behavioral
and demographic information of children with ASD is distinguished by challenges in three core areas:
and without ASD, such as age, gender, language,
and social interaction scores. Using machine Social Interaction: Children with ASD usually experience
learning, the system analyzes data and predicts difficulties in establishing and maintaining social
relationships. Between the ages of 0-11years, this may
4 whether a person may have autism. Our goal is to present as limited interest in playing with peers, reduced eye
create a simple, low-cost tool that makes autism contact, lack of response to social cues, or difficulty
detection easier, especially in places where expert understanding other’s emotions. Such challenges can make
doctors may not be available. group participation in preschool or primary school settings
more complex.
34 Keywords— Autism detection, Support Vector Machine
(SVM), Convolutional Neural Network (CNN), image recognition, Communication: Communication difficulties are especially
10 machine learning, deep learning.
prominent in younger children with ASD. These may include
delayed speech development, limited vocabulary, repetitive
27 language, or challenges in using non-verbal communication

Page 7 of 14 - Integrity Submission Submission ID trn:oid:::1:3398019136


Page 8 of 14 - Integrity Submission Submission ID trn:oid:::1:3398019136

27 such as gestures, facial expressions, and tone of voice. Some 2. Data Preprocessing: After collecting the datasets, the next
18 children find it difficult to initiate or sustain conversations, step involved data preprocessing, which plays a vital role in
making it harder for them to express their needs and feelings converting raw data into a format suitable for machine
effectively. learning models. The effectiveness of preprocessing directly
influences the model's performance, as clean, structured, and
behavioral Patterns and Interests: Children within the 0-11 relevant data that leads to more reliable and accurate
age group often display restricted or repetitive patterns or predictions. The dataset, stored within the project’s model
behaviour. This may involve repetitive movements such as folder contains 12 behavioral questions with responses
28
hand-flapping, rocking, intense focus on specific interests, marked as “Yes” or “No”, along with demographic attributes
2 insistence on routine, or heightened sensitivity to sensory such as age, gender, ethnicity. The structured preprocessing
2 stimuli such as sounds, lights, or textures. Such behaviors can step ensures that the datasets is machine readable and ready
make it difficult for children to manage their daily for effective model training and evaluation.
functioning at home or in school environments.
35 TABLE 1. Feature Description of ASD Datasets
II. PROBLEM STATEMENT
21 Autism Spectrum Disorder (ASD) is a developmental
condition that deeply affects a child’s communication, social
interaction, and behavioral patterns in children. The growing
number of ASD cases around the world shows the urgent
need for early detection, especially in children younger than
11, as timely intervention can greatly improve their
development and social growth. However, the current
diagnostic process is often lengthy, time-consuming and
dependent on expert evaluations, clinical observations, and
behavioral assessments, that are not always accessible in rural
or under-resourced areas. These challenges frequently result
in delayed or missed diagnosis. Table 1. Feature description of ASD datasets

3 To overcome these limitations, this research explores how 3. Model Development: In this project, four different models
machine learning methods can be used to support early and were built for ASD detection in children. The Support Vector
8 efficient ASD detection. Specifically, Support Vector Machine model was utilized to classify questionnaire and
Machine (SVM) algorithm is applied to classify demographic data, while a Convolutional Neural Network
3 questionnaire-based data, while Convolutional Neural was employed for image-based recognition tasks. The
Networks (CNN) are employed for image-based recognition Decision tree model was chosen for its easy interpretation in
of behavioral and facial features. This dual-approach system identifying important behavioral features, and Random
aims to improve accuracy, reduce reliance on expert-only Forest was employed as an ensemble method to improve
evaluations, and provide an interpretable and scalable accuracy and reduce over fitting. All the models were
solution. By leveraging these models, the study seeks to developed using Python with sci-kit-learn for SVM, Decision
3
create a supportive diagnostic tool that enhances early Tree and Random Forest and TensorFlow / Keras for CNN.
identification of ASD in children aged 0-11 years, thereby 4. Model Evaluation: Each model’s performance was
5
facilitating timely treatment and better long-term outcomes. assessed using k-fold cross-validation to ensure
24 III. PROPOSED METHODOLOGY generalization. Metrics such as accuracy, precision, recall,
F1-score were computed to compare classifier effectiveness.
5. System Design and Integration: The best-performance
9 Methodology—This research follows a systematic approach,
model was integrated into a prototype decision-support
combining data collection. Preprocessing, model
system designed to assist doctors and caregivers in making
development, and evaluation to detect Autism Spectrum
quick, data-driven screening decision.
Disorder effectively.
2 6. Validation and Testing: The system was tested on unseen
data to evaluate real-world applicability and robustness,
1. Data Collection: In this study, Autism Spectrum Disorder followed by expert feedback for clinical relevance.
(ASD) related datasets were collected from open-access Performance metrics such as accuracy, precision, recall, and
22
repositories and verified clinical sources. The datasets F1-score were used to check the reliability of predictions. The
includes both behavioral assessment data and demographic models demonstrated strong generalization capability,
38 details, that play a key role in identifying early signs of ASD minimizing over-fitting issues during validation.
in children. The behavioral data is based on 12 screening Additionally, the results showed that machine learning can
8
questions, where responses are recorded in a binary format as serve as a supportive tool for early ASD detection,
“Yes” or “No”. These questions aim to identify behavioral complementing traditional diagnostic approaches.
patterns linked with ASD in children. Along with behavioral
responses, the datasets also contains demographic attributes
2 such as age, gender, and ethnicity, which provide context for
the screening results.

Page 8 of 14 - Integrity Submission Submission ID trn:oid:::1:3398019136


Page 9 of 14 - Integrity Submission Submission ID trn:oid:::1:3398019136

IV. SYSTEM DESIGN

A. System Architecture

Fig 4.1: Architecture diagram

40 The diagram represents the workflow of an Autism Spectrum


Disorder (ASD) detection system using machine learning. It
starts with an adult autism screening datasets, followed by
data preprocessing and feature selection. The processed data
is then fed into classifiers like Random Forest and Decision
Tree. Based on the analysis, the system predicts whether a
person has autism (Yes/No), and finally, performance metrics
and graphs are generated to evaluate the model's accuracy.
B. Class Diagram
The diagram represents a simplified overview of the input-
4 output process used in Autism Spectrum Disorder (ASD)
detection system using machine learning. The process begins
on the input side with datasets acquisition, where a structured
datasets—typically collected from public sources like the
UCI Machine Learning Repository—is loaded. This datasets
contains various features including behavioral, demographic,
and personal attributes relevant to autism screening.
Fig 4.2: Class diagram
The figure defines the key components and their interactions,
starting user input to final prediction. The User class enables Following datasets acquisition, data preprocessing is
actions such as login, model selection and data input. Input is performed. This stage includes cleaning the data, handling
represented through the input data class, which is sub-classed missing or inconsistent values, normalizing or scaling
into Questionnaire for behavioral data and Image File for numerical features, and encoding categorical data.
13 facial or MRI image inputs. All inputs are stored in the dataset Preprocessing ensures the datasets is suitable for training
class for Training and testing. The Model class implements accurate and reliable machine learning models. On the output
4 algorithms such as SVM, Random Forest, Logical Regression, side, the system moves onto feature extraction and selection,
KNN, and CNN to train and predict outcomes. The Prediction where the most informative attributes are identified to
class generates the Final output, providing diagnostic results simplify the data and improve model performance. These
(Autistic/Non-Autistic), confidence scores and explanation chosen features are then passed into classification models,
17
through an explainable AI component. This design ensures like Random Forest Classifier and Decision Tree Classifier,
modularity, scalability and interpretability of the ASD which are trained to detect patterns indicative of ASD. Finally,
detection system. the trained models are used to classify new input data. The
predicted result is displayed to the user, indicating whether
The modular architecture of the ASD detection system the individual is likely to have autism (Yes or No).
enables integration with the healthcare and educational
platforms. The flexible design ensures compatibility with V. MODEL SELECTION
tele-health systems, cloud databases and learning tools, 1) Support Vector Machine: Support Vector Machine (SVM)
12
making the solution scalable and accessible across diverse is a supervised machine learning algorithm designed for both
real-world applications classification and regression tasks. In ASD detection, it is
applied to classify children into ASD or Non-ASD groups

Page 9 of 14 - Integrity Submission Submission ID trn:oid:::1:3398019136


Page 10 of 14 - Integrity Submission Submission ID trn:oid:::1:3398019136

using behavioral and demographic features. It was selected as  Dependent variable: A categorical outcome, such as
one of the main models because of its strong capability in “high vs. low symptom severity” or “presence of
handling classification tasks, especially when dealing with internalizing behaviours” (Yes/No).
high-dimensional datasets such as behavioral questionnaires  Independent variables: Scores from standardized
and demographic information. SVM functions by identifying assessments like the Autism Diagnostic Observation
the best possible boundary that separates the different classes Schedule (ADOS) or adaptive behaviour scales.
with maximum margin, which helps improve the model’s  Analysis: The model could show which behavioral
generalization ability. behavioral and demographic datasets indicators such as challenges with social reciprocity or
often contain a limited number of samples. SVM performs repetitive behaviours are stronger predictors of a
well even when the datasets is not very large, making it specific symptom profile in older children.
suitable for ASD screening data. Each questionnaire response  Comparing age groups: A single model could
and demographic attribute can be considered a feature, incorporate an “age group” variable to compare how
leading to high-dimensional feature space. SVM efficiently behavioral predictors differ across the two ranges.
handles such spaces while reducing the risk of over fitting.  Dependent variables: A binary outcome, such as “severe
SVM Formula: If we have data points (xi, yj) where, communication difficulties”.
 Independent variables: behavioral scores and a
categorical variable for age group.
 Analysis: The model’s coefficients would reveal how
the odds of severe communication difficulties vary
across age group, while accounting for the specific
19 Then the equation of the hyperplane is: behaviours.
w. x + b = 0
where: Algorithm
 w = weight vector Logistic regression uses an optimistic algorithm, most
 b = bias term commonly gradient descent, to find the optimal coefficients
that minimize the model’s error. The process is as follows:
25 2) Convolutional Neural Network: CNN is a deep learning 1. Initialize Coefficients: The model begins with initial
technique used for image-based recognition tasks, coefficient values (weights), which are typically set to zero.
particularly analyzing facial features and behavioral cues that 2. Calculate Probabilities: For every child in the training
17 may indicate Autism Spectrum Disorder (ASD). Unlike datasets, the model calculates a predicted probability (p) of
traditional machine learning models, CNN's automatically the outcome using the current coefficients.
2
extract layered features from images, making them highly 3. Evaluate Loss: The model’s performance is measured
efficient for tasks related to visual and pattern recognition. using a loss function (cross entropy), which penalizes large
CNN was chosen for its superior ability to handle image- errors.
based data, automatically extract discriminative features, and 4. Compute Gradient: The algorithm calculates how the loss
provide accurate classification. Its use in this project changes, to determine the coefficients should move in to
enhances the reliability of ASD detection by complementing reduce errors.
5 questionnaire based machine learning models like SVM, 5. Update Coefficients: The coefficients are updated by
Decision Tree and Random Forest. taking a small step in the opposite direction of the gradient.
6. Repeat: Steps 2-5 are repeated until the model’s
3) Logistic Regression: To use logistic regression for the coefficients converge to a stable solution.
behavioral analysis in autistic children, the method creates a
predictive model that categorizes children (e.g., presence or Formula
absence of a certain behaviour) based on multiple factors. Linear Combination (z): z = β0 + β1x1 + β2x2 + ... + βnxn
This is particularly useful for comparing the behavioral
differences between children aged 0-3years and 4-11 years. Sigmoid Function (p): p = 1 / (1 + e^-z)
i) In the 0-3 age group: The model helps in early
identification and prediction of an ASD diagnosis. Cross-Entropy Loss Function (L):
16
 Dependent variable: A binary outcome, such as “later L = - (1/m) * Σ [ yi log(pi) + (1-yi) log(1-pi)]
ASD diagnosis” (Yes/No). where yi represents the actual label and pi denotes predicted
 Independent variables: Early behavioral markers, such probability.
as delayed language, lack of eye contact, or repetitive
6
movements. 4) Random Forest Classifier: The Random Forest algorithm
 Analysis: The model could identify which early is an ensemble learning approach for both classification and
behaviours are most predictive of a diagnosis later in regression tasks. It works by building several decision trees
childhood. For instance, a model could reveal that “lack during training and combining their results to produce the
of pointing” at 18 months has high odds ratio for a later final prediction, which enhances accuracy and reduces over
14
diagnosis. fitting. Each tree in the forest is trained using a random
ii) In the 4-11 age group: The model can investigate how sample of the data, and final decision is made through
different behavioral traits relate to different degrees of majority voting (for classification) or averaging (for
20
symptom severity or specific symptom profiles. regression). Random Forest can recognize and understand
complex patterns in the data. In our project, Random Forest

Page 10 of 14 - Integrity Submission Submission ID trn:oid:::1:3398019136


Page 11 of 14 - Integrity Submission Submission ID trn:oid:::1:3398019136

37 Classifier will analyze extracted features (e.g., behavioral and 6. Calculate the distance: The algorithm calculates the
cognitive data) to classify children as having ASD or not, distance between the new child’s data point and every other
32 leveraging its high accuracy for binary classification tasks. child’s data point in the training set using a distance metric,
Here’s how the Random Forest algorithm works: such as the Euclidean distance.
 Random portions of data are selected from the original 7. Identify the nearest neighbours: It then finds the k closest
training datasets. to the new child based on these distances.
 A separate decision tree is built for each of these subsets. 8. Assign a class via majority vote: The new child is assigned
 The number of decision trees to be generated is to the most common diagnostic class (e.g., ASD or Non-ASD)
determined by selecting a predefined value of ‘N’. among their k nearest neighbours.
 Steps 1 and 2 are repeated until every tree in the forest
has been built. Formula
7  For each test sample, find the predictions of each The Euclidean distance is the metric most often used to
decision tree, and assign the test sample a class value measure similarity in the KNN algorithm. For a new data
based on majority voting. point x with features (x1,x2,…..,xn) and an existing data
23 Algorithm point y with features (y1,y2,….,yn), the euclidean distance is
The Random Forest algorithm combines bootstrapping and calculated as follows:
random feature selection to create an ensemble of decision d(x,y) = sqrt (Σ (xi-yi)2 )
9 trees. This formula essentially measures the straight-line distance
29 1. Bootstrap Samples: Create multiple subsets of the original between the two points in the multi-dimensional behavioral
training data by randomly sampling with replacement. space.
2. Random Feature Selection: For each tree, and at each split
point, consider only a random subset of available features.
3. Train Decisions Tress: Grow a decision tree on each VI. EXPERIMENTAL RESULTS AND ANALYSIS
20
bootstrap sample using randomly selected features.
4. Aggregate Predictions: For a new child, feed their data The designed ASD detection model was tested across four
through every tree. Each tree predicts, and the final different approaches, each focusing on different input data
classification is decided based on majority vote. and age groups. The first two models were designed for
behavioral and developmental data of children aged 0-3 and
Formula (Gini Impurity) 4-11 years respectively, while the other two models i.e. facial
Gini Impurity for a Node (G): recognition and brain MRI image analysis utilized image-
10 G = 1 - Σ pi^2 based data. The machine learning algorithms implemented
33 where pi is the probability of a data point belonging to class includes Support Vector Machine (SVM), Logistic
i. Splitting Decision: Choose the split that results in the Regression, and Random Forest for behavioral datasets, and
8 greatest decrease in Gini Impurity. a Convolutional Neural Network (CNN) for the image-based
30 Decrease in Gini = G_parent - (Weighted G_children) models. Performance was assessed using accuracy, precision ,
recall, and F1-score as evaluation. The results showed that the
5 5) KNN Algorithm: Instead of learning a single predictive CNN-based models achieved better performance compared
41 equation, the KNN algorithm makes predictions by finding to traditional machine learning algorithms, particularly in
the ‘k’ most similar children from the training data and using facial and MRI image recognition tasks, because of their
2 their diagnosis to classify the new case. For example, if a new strong feature extraction capabilities. Among the behavioral
5-year-old child is being evaluated for autism, the algorithm models, Random forest achieved the best performance,
would look at the behavioral profiles of the children in the effectively capturing complex patterns within the data.
datasets most similar to them.
Example Steps: A. Age group (0-3 & 4-11 years
1. Select behavioral features: Create a datasets with variables
for a child’s age group (0-3, 4-11) and specific behavioral
26 scores from a standardized test, such as Social
Responsiveness Scale (SRS) or the Autism Diagnostic
Observation Schedule (ADOS).
2. Define a known diagnostic label: Each child in the training
datasets must have a known classification (e.g., ASD or Non-
ASD).
3. Represent the children as data points: Each child’s
42 behavioral profile is plotted as a point in a multi-dimensional
‘behavioral space’. Fig 6.1 Result snapshot for behavioral and demographic questionnaire for
4. Introduce a new child for classification: When a new child age group between 0-3 & 4-11 years
with an unknown diagnosis is evaluated, the KNN algorithm
is used to classify them. The proposed autism detection system was evaluated using
5. Choose the number neighbours(k): The user must select behavioral questionnaire data for early-age subjects. Fig X
the number of neighbours (k) to examine. Choosing an odd illustrates one of the model’s output screens generated after
number for k is a common practice to avoid ties in processing user responses. The system collects user details,
classification. including name and email, and provides a diagnostic status

Page 11 of 14 - Integrity Submission Submission ID trn:oid:::1:3398019136


Page 12 of 14 - Integrity Submission Submission ID trn:oid:::1:3398019136

along with an explainable reasoning section. In this test case, The developed system was tested using MRI scan inputs to
the model predicted the user as “Non- Autistic”. The decision evaluate its performance in disgusting autistic from non-
was primarily influenced by positive behavioral indicators autistic brain patterns. Fig shows a sample output generated
such as the subject’s ability to look when called, point to by the system for one of the test subjects. The user
share interest, follow gaze and use simple gestures, all of details(name and email) are recorded along with the
3 which align with typical social developmental patterns. diagnostic status predicted by the Convolutional Neural
Additionally, the absence of responses indicating lack of Network (CNN) model. In the particular case, the model
comfort, repetitive staring, or social withdrawal contributed analysed the MRI features and produced a probability of
to the non-autistic classification. Demographic and biological 0.0016 for autistic and 0.9984 for non autistic. Since the
3 parameters such as age, gender, ethnicity, and family history probability of being non-autistic is significantly higher, the
were also processed by the trained machine learning model, final diagnosis is reported as normal. The “Explained Why”
which collectively supported the “Non-Autistic” outcome. section provides interpretability by displaying the reasoning
The “Explained Why” module improves the system’s clarity behind the prediction, thereby improving the model’s
2 by offering a transparent explaination of how each prediction transparency. This result clearly demonstrates the CNN
2 is made. These results validate the model’s ability to integrate model’s ability to accurately classify MRI scans and provide
behavioral data effectively and provide a reliable autism risk confidence values for each prediction. Similar evaluations
assessment with clear interpretive feedback to end users. conducted across multiple test cases confirmed the robustness
and reliability of the proposed system in detecting autism
B. Facial Image Recognition related brain patterns.

VII. CONCLUSION

In this research, we designed an integrated framework for


Autism Spectrum Disorder detection that combines
behavioral, facial and MRI-based analysis using both
36 Machine Learning (ML) and Convolutional Neural Network
(CNN) techniques. The framework was designed to identify
autism indicators across different age groups, providing a
unified approach to early diagnosis. behavioral data for
toddlers and children were processed using ML classifiers,
Fig 6.2 Result snapshot of Facial Image Recognition while CNN models analysed facial expressions and MRI
4 scans to enhance diagnostic precision. The experimental
In this phase of experimentation, the CNN based facial results confirmed that the proposed hybrid framework
analysis model was tested to predict autism using facial effectively distinguishes between autistic and non-autistic
features and emotional cues. As shown in Fig 6.2, the system subjects with strong interpretability.
analysed the input image and classified the subject as Non-
Autistic. The explanation module indicated that the detected Furthermore, the inclusion of an explainable AI module
facial patterns reflected normal emotional in interpreting provides transparency by highlighting the factors influencing
facial patterns reflected normal emotional responses and each prediction, which enhances clinical trust and usability.
social engagement. This demonstrated the model’s This approach supports healthcare professionals in
effectiveness in interpreting facial cues and its capability to understanding key behavioral and neurological indicators
provide reliable, explainable predictions for autism detection. linked to ASD, enabling more accurate and timely
Additionally, the interpretability feature enhances user trust interventions. Although the current model’s performance is
by clearly explaining how the decision was made based on limited by data-size and diversity, it demonstrates significant
observable facial characteristics. potential as a decision-support tool in clinical and educational
settings. Upcoming research will focus on expanding the
C. Brain MRI image analysis datasets, refining deep learning architectures, and
implementing real-time web-based deployment for broader
accessibility and scalability in early ASD detection.

VIII. FUTURE ENHANCEMENTS


3 In the future, this project can be enhanced by expanding the
datasets to include larger and more diverse populations,
39 improving the models ability to generalize across different
age groups and cultural contexts. Additional data models like
speech motor activity, genetic markers, and longitudinal
behavioral tracking can be incorporated to enable more
Fig: 6.3 Result snapshot of Brain MRI image analysis comprehensive early risk assessment.

The explainable AI component can be further redefined to


provide deeper insights into the reasoning behind each

Page 12 of 14 - Integrity Submission Submission ID trn:oid:::1:3398019136


Page 13 of 14 - Integrity Submission Submission ID trn:oid:::1:3398019136

prediction, helping clinicians and caregivers better [16] V Kumaravel, K. Helen Prabha, “Analysis of Autism
understanding and to act on the results. Adaptive learning Spectrum Disorder Prediction using various Machine
Learning Models”, IEEE, 2024.
mechanisms can be integrated to personalize diagnostic
suggestions and intervention of pathways for individual
profiles.

Furthermore, embedding the system into tele-health


platforms, wearable devices, and electronic health record
systems can increases accessibility, scalability, and real- time
monitoring.

Overall, these future enhancements aim to make this project


a more accurate, interpretable, and practical tool for
supporting early diagnosis in ASD.

IX. REFERENCES

[1] S.M. Mahedy Hasan, MD Palash Uddin, MD AL


Mamun, Muhammad Imran Sharif, Anwaar Ulhaq, and
Govind Krishnamoorty, “A Machine Learning
Framework for Early-Stage Detection of Autism
Spectrum Disorders”, IEEE, 2023.
[2] Leslie Mertz, “ Using AI and ML to Predict Autism
Spectrum Disorder”, IEEE, 2024.
[3] Khushbu Garg, Dr. Nripendra Narayan Das, Dr. Gaurav
Agarwal, “Autism Spectrum Disorder Detection by
Machine Learning Using Small Video”, IEEE, 2023.
[4] Ganesh Zambre, Harshad Albhar, Toukir Masli, Sarita
Patil, “Detection and Analysis of Autism Spectrum
Disorder Using Random Forest Classifier”, IEEE, 2024.
[5] Oumaima Ben Mohamed, Olfa Souki, “Early Detection
of Autism Spectrum Disorder in Toddlers Using
behavioral Indicators”, IEEE, 2024.
[6] Nutan Hemant Deshmukh, Haridas Gadade, “Autism
Spectrum Disorder: A Global Review of Analysis,
Intervention and Support Approaches”, IEEE, 2025.
[7] YingTong Ai, “An Children Autism Spectrum Disorder
Diagnosis using Convolution Neural Network and
Recurrent Neural Network”, IEEE, 2024.
[8] Praveena R S, Pradeep R, “Brain Wave Based Autism
Spectrum Disorder Detection - A Machine Learning
Approach”, IEEE, 2025.
[9] T. Arunprasath, M. Niranjana, Sharumathi M, Buvanesh
pandian V, M Pallikonda Rajasekaran, Kottaimalai
Ramaraj, “Prediction of Autism Spectrum Disorder
Using Machine Learning”, IEEE, 2024.
[10] V. Kavitha, R. Siva, “Classification of Toddler, Child,
Adolescent and Adult for Autism Spectrum Disorder
Using Machine Learning Algorithm”, IEEE, 2023.
[11] Ambika Rani Subhash, Ashwin Kumar UM, “Ensemble-
based Machine Learning Classification for Early
Detection of Autism Spectrum Disorder”, IEEE, 2025.
[12] Anas D. Sallibi, Khattab [Link] Alheeti, “Detection of
Autism Spectrum Disorder by Using Common Machine
Learning Algorithms”, IEEE, 2023.
[13] Ruhi Patankar, Shreyas Vedpathak, Vaidehi Thakre,
Prajol Sethi, Sejal Sawarkar, “AutiScan: Screening of
Autism Spectrum Disorder Specific to Indian Region”,
IEEE, 2022.
[14] Pushpmala Nawghare, Jayashree Rajesh Prasad, “Early
Detection of Autism Spectrum Disorder Using AI and
Machine Learning Models: A Systematic Review for
Effective Intervention”, IEEE, 2024.
[15] D Harish Sree Prasanna Kuamr, Julia Punitha Maraldhas,
I. Devapriya, “Machine Learning Approaches for
Autism Spectrum Disorder Prediction”, IEEE, 2024.

Page 13 of 14 - Integrity Submission Submission ID trn:oid:::1:3398019136


Page 14 of 14 - Integrity Submission Submission ID trn:oid:::1:3398019136

IEEE conference templates contain guidance text for


mission to the conference. Fai

Page 14 of 14 - Integrity Submission Submission ID trn:oid:::1:3398019136

Common questions

Powered by AI

Combining traditional diagnostic methods with machine learning techniques enhances overall ASD detection effectiveness by leveraging the strengths of both approaches. Traditional methods provide a foundational understanding through expert evaluation, while machine learning provides data-driven insights and faster processing of complex datasets. This hybrid approach enables comprehensive diagnostics, facilitating early intervention and personalized treatment plans, and improving diagnosis accuracy and patient outcomes by utilizing detailed behavioral and imaging data .

Challenges in using machine learning models for ASD detection include issues such as data quality, feature selection, and overfitting. Limited availability of annotated datasets can hinder model training, while poor data preprocessing may lead to inaccurate predictions. These challenges can be mitigated through rigorous data preprocessing, feature selection to enhance model performance, and techniques like k-fold cross-validation to ensure generalization. Continuous refinement and validation of models with diverse datasets can also improve their robustness and applicability across different demographics .

The integration of explainable AI components enhances transparency and user trust in ASD detection systems by providing clear, understandable reasoning for each prediction made by the model. This feature allows clinicians and caregivers to gain insights into the decision-making process, facilitating better validation of outcomes and increasing acceptance of AI-driven diagnostics. Explainability ensures users can comprehend how input data, such as behavioral questionnaires, leads to specific diagnostic results, fostering confidence in the technology's reliability and helping integrate AI into clinical settings .

Machine learning techniques offer benefits in early ASD detection by reducing the time-consuming nature of traditional methods that require expert evaluation. They allow for faster diagnosis, enabling healthcare providers and parents to make quicker decisions. These techniques enhance accuracy and can manage the high-dimensional data often associated with ASD assessments. The use of models such as Support Vector Machines and Convolutional Neural Networks in processing diverse data types such as questionnaires and imaging data improves detection rates and provides a complementary tool to traditional approaches .

Demographic attributes such as age, gender, and ethnicity significantly enhance machine learning models' predictions for ASD detection by providing contextual information that can influence autism risk. These features allow models to adjust predictions based on patterns observed in specific demographics, improving accuracy and relevance. Incorporating demographic data helps in tailoring assessments to individual cases, thereby supporting personalized diagnostics and interventions in ASD detection systems .

Convolutional Neural Networks (CNNs) play a critical role in ASD detection through their application in image-based tasks such as facial and MRI image analysis. They are particularly effective because they excel at feature extraction from raw pixel data, making them suitable for interpreting complex image patterns that might indicate ASD traits. CNNs demonstrated superior performance over traditional machine learning models in tasks that require precise image recognition, thereby enhancing ASD detection and providing insightful data for early diagnosis .

Modularity and scalability in ASD detection systems are crucial for their integration with healthcare and educational platforms, allowing for adaptation to various technological infrastructures. Modularity ensures that each component of the system can be independently updated or replaced, facilitating maintenance and feature enhancement. Scalability allows for handling increased data volumes or user demands without performance loss. These can be achieved through a flexible architecture using standardized APIs, cloud-based resources, and a layered design that separates data processing from model training and prediction components, ensuring seamless extensions or integrations .

K-fold cross-validation is essential in evaluating ASD detection models because it systematically partitions the dataset into k equal subsets, ensuring that each data point gets to be part of both training and testing sets across different iterations. This process contributes to the model's generalization capability by providing a more comprehensive view of how the model performs on unseen data, thus minimizing the risk of overfitting and improving the reliability and robustness of the model's predictions .

Support Vector Machines (SVMs) are employed in diagnosing Autism Spectrum Disorder by classifying data into ASD or non-ASD groups based on behavioral and demographic features. Their suitability for ASD screening comes from their capability to handle high-dimensional datasets and provide robust classification even with a limited number of samples. SVMs identify an optimal boundary that maximizes the margin between classes, which in turn improves model accuracy and generalization .

Data preprocessing and feature selection enhance machine learning models' effectiveness in diagnosing ASD by ensuring that datasets are clean, complete, and machine-readable. Preprocessing deals with inconsistencies and normalizes or scales numerical features, while feature selection identifies the most informative attributes. This simplification reduces noise and dimensionality, which improves model performance and accuracy in predicting ASD outcomes. Effective preprocessing ensures robust and reliable classifiers, which is crucial for clinical applications .

You might also like