0% found this document useful (0 votes)
14 views67 pages

Project Report

The project report titled 'Milk Quality Prediction' by Diwakar focuses on developing a conversational AI agent to provide mental health support using natural language processing and emotion recognition. It aims to address the growing need for accessible mental health care by offering real-time assistance and personalized coping strategies. The report includes sections on system design, implementation, and literature review related to AI in mental health applications.

Uploaded by

gatizmourya353
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
14 views67 pages

Project Report

The project report titled 'Milk Quality Prediction' by Diwakar focuses on developing a conversational AI agent to provide mental health support using natural language processing and emotion recognition. It aims to address the growing need for accessible mental health care by offering real-time assistance and personalized coping strategies. The report includes sections on system design, implementation, and literature review related to AI in mental health applications.

Uploaded by

gatizmourya353
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

A PROJECT REPORT

ON
“Milk Quality Prediction”
Submitted in Fulfillment of the Requirement for the Degree of
BACHELOR OF TECHNOLOGY
IN
Artificial Intelligence and Machine Learning

Submitted by
Diwakar
213022008

Under the Supervision of


Mrs. Supriya Shukla
Assistant Professor

DEPARTMENT OF ARTIFICIAL INTELLIGENCE AND


MACHINE LEARNING
COER University, Roorkee
7th, KM Haridwar, National Highway Vardhmanpuram, Roorkee,
Rehmadpur, Uttarakhand, 247667
Session: 2024 - 2025
COER UNIVERSITY

Department of Artificial Intelligence and Machine Learning

Project Progress Report

Student ID: U213022008

Student Name: Diwakar

Program: [Link] Semester/Section: 8th / A

Session: Even Semester (2024-2025)

Project Mentor Name: Mrs. Supriya Shukla

Project Title: Milk Quality Prediction

Details of Visit to Mentor:


S. No Date Time Remarks Signature

Project In-Charge Signature

Diwakar (213022008) I
CANDIDATE’S DECLARATION

I hereby declare that this project report titled "Milk Quality Prediction" is an
original work done by me under the supervision of Mrs. Supriya Shukla. It has not
been submitted previously for the award of any degree.

Diwakar
213022008
01/05/2025

Diwakar (213022008) II
CERTIFICATE

This is to certify that the project titled "Milk Quality Prediction " submitted by
Diwakar, Roll No. 213022008, has been carried out under my guidance and is
approved for submission.

Mrs. Supriya Shukla Signature


Assistant Professor (External Examiner)
Department of AI/ML Name:
COER University, Roorkee Designation:

Date:

Diwakar (213022008) III


ACKNOWLEDGEMENT

I sincerely express my gratitude to Mrs. Supriya Shukla, my project guide, for their
valuable guidance, encouragement, and support throughout this project. I also extend
my thanks to my department faculty, family, and friends for their cooperation.

Diwakar
[Submission Date]

Diwakar (213022008) IV
TABLE OF CONTENTS

Sr. Title Page No.


No.

1. Progress Report I

2. Candidate’s Declaration II

3. Certificate III
4. Acknowledgements IV

5. Table of Contents V

6. Abstract 1

7. Introduction 2

8. Literature Review 3-6

9. System Analysis 7-9


9.1 Existing System
9.2 Proposed System

System Design
10. 10-11
10.1 Architecture Diagram
10.2 Data Flow Diagram (DFD)

11. Implementation
12-41
11.1 Technologies Used
11.2 Coding & Modules
11.3 Testing & Validation

12. Results & Discussion 42-45

46-49
13. Conclusions & Future Work
52
14. References

Diwakar (213022008) V
LIST OF FIGURES

Figure No. Name of the Figure Page No.


3.1 Architecture of AI Mental Health Chatbot 3
3.2 Multimodal Emotion Recognition System 5
5.1 System Architecture Diagram 10
5.2 Data Flow Diagram 11
7.1 3D Avatar interacting during a live session 42
7.2 Voice output playing in real-time during a user query 42
7.3 Gemini-powered chatbot response to a user expressing 43
stress
7.4 Flask server logs showing active session and chatbot 43
requests
7.5 Final integrated Milk Quality Prediction interface 44

Diwakar (213022008) VI
LIST OF TABLES

Table Name of the Table Page


No. No.
3.1 Popular Mental Health Chatbots and Their 4
Features
3.2 Techniques for Emotion Detection in AI Systems 5
3.3 Comparison of AI Models for Dialogue 5
Generation
3.4 Voice-Based Emotion Recognition Methods 6
4.1 Comparison of Existing Systems vs Proposed 9
System
7.1 Unit Testing Test Cases 44
7.2 Test Cases 45

Diwakar (213022008) VII


CHAPTER- 1
ABSTRACT

The rising prevalence of mental health issues globally highlights the urgent need for
accessible and scalable psychological support systems. This project presents the
development of an Milk Quality Prediction—a conversational agent designed to
provide empathetic, context-aware mental health assistance using artificial
intelligence and natural language processing (NLP). The system incorporates real-
time emotion and sentiment analysis to personalize responses, offering coping
strategies, mental wellness guidance, and crisis intervention support. Key features
include voice interaction via text-to-speech and speech-to-text integration, emotion
recognition through facial analysis, conversation logging, and a secure user
authentication system. Built using open-source tools and frameworks such as Flask,
React, Eleven Labs, and Ready Player Me avatars, the chatbot simulates human-like
interactions and fosters a comforting environment. This project aims to bridge the
gap in mental healthcare accessibility by delivering timely, supportive, and
intelligent mental health assistance, especially for individuals lacking immediate
access to professional therapists.

Keywords: Artificial Intelligence (AI), Mental Health, Chatbot, Natural


Language Processing (NLP), Sentiment Analysis, Emotion Detection, Virtual
Therapist, Conversational AI, Real-time Interaction, Speech-to-Text (STT), Text-
to-Speech (TTS), Eleven Labs Integration, Ready Player Me Avatar, Flask
Framework, React Frontend, Empathetic Response Generation, Crisis
Intervention, Mental Health Support System, User Authentication, AI-driven
Therapy

Diwakar (213022008) 8
CHAPTER- 2
INTRODUCTION
2.1 Milk Quality Prediction
Mental health is an essential aspect of overall human well-being, yet it is often
neglected due to societal stigma, lack of awareness, and insufficient availability of
professional care. With the growing incidence of mental health disorders such as
anxiety, depression, PTSD, and stress-related conditions, there is a need for
accessible, scalable, and supportive technological solutions. The Milk Quality
Prediction is a conversational agent developed to provide emotional support, mental
wellness guidance, and empathetic interaction using artificial intelligence, natural
language processing, and emotion recognition technologies.

2.1.1 Real-Time Applications of Milk Quality Prediction


The Milk Quality Prediction can be applied in various real-time scenarios:
 Providing immediate support to individuals experiencing emotional distress.

 Assisting students with academic stress, social anxiety, or exam-related


pressure.

 Offering 24/7 assistance to employees dealing with workplace stress.

 Supporting isolated individuals or those in rural areas without access to


therapists.

 Acting as a first-level responder before professional intervention in mental


health crises.

2.2 Categorization of Milk Quality Prediction


Milk Quality Predictions can be categorized based on their interaction modes,
features, and data analysis capabilities:
 Text-based vs. Voice-based chatbots

 Emotion-aware vs. rule-based systems

 Static response vs. real-time NLP-generated dialogue

 Basic vs. avatar-enhanced visual interaction

Diwakar (213022008) 1
2.2.1 AI Mental Health Therapy Process
The process involves:
1. User interaction via voice or text.

2. Emotion and sentiment analysis through NLP and facial expression


recognition.

3. Response generation using a pretrained NLP model like GPT-3.5 or


DialoGPT.

4. Voice synthesis and avatar animation for realistic interaction.

5. Data logging and recommendation of wellness strategies or referral to


professional help if necessary.

2.2.2 Sensor-Based Emotion Detection


Sensor-based systems (e.g., wearable devices) can monitor physiological indicators
such as heart rate, skin temperature, and galvanic skin response. Although not used
directly in this version, they are relevant for future expansion into HAR (Human
Activity Recognition) integration for enhanced emotional profiling.
2.2.3 Vision-Based Emotion Detection
Facial expression analysis via webcam allows for real-time emotion detection using
computer vision models. Emotions such as sadness, anger, fear, happiness, and
neutrality are classified to tailor chatbot responses empathetically.
2.3 Machine Learning and Deep Learning Algorithms for Milk Quality
Prediction
 Natural Language Processing (NLP): Transformer-based models like GPT-3.5
and DialoGPT enable context-aware, human-like conversations.

 Sentiment Analysis Models: Tools like VADER, TextBlob, or custom BERT-


based classifiers detect emotional tones in textual input.

 Facial Emotion Recognition: CNN-based models (e.g., FER2013, OpenCV


DNN modules) analyze facial expressions from live webcam input.

 Voice Analysis: Emotion inference from vocal tone is a future enhancement


using deep learning audio classifiers.

Diwakar (213022008) 2
2.4 Research Gaps
 Lack of real-time, multi-modal emotional context understanding in current
chatbots.

 Limited integration of 3D avatars and emotion-aware speech synthesis in


mental health tools.

 Absence of regional language support for inclusivity.

 Inadequate personalization of coping strategies based on user history.

 Limited support for emergency escalation and professional help linkage.

2.5 Objectives
 To develop a real-time Milk Quality Prediction that can engage in empathetic,
emotionally-aware conversations.

 To implement speech-to-text and text-to-speech capabilities for voice


interaction.

 To integrate emotion recognition using facial expressions and textual


sentiment analysis.

 To simulate human-like presence with a 3D avatar using Ready Player Me.

 To provide privacy-focused user authentication and secure logging of


conversations.

 To offer contextual coping suggestions, mindfulness strategies, and referral


guidance for professional mental health support.

Diwakar (213022008) 3
CHAPTER- 3
LITERATURE REVIEW
3.1 Introduction
Mental health technologies have evolved significantly in the last decade due to
advancements in artificial intelligence (AI), natural language processing (NLP), and
human-computer interaction. Various studies and models have contributed to the
development of AI-based mental health applications. This chapter provides a
comprehensive review of the literature relevant to the Milk Quality Prediction project,
focusing on key domains such as chatbot technologies, emotion detection, therapeutic
interactions, and AI-based mental health systems.

3.2 Conversational AI and Mental Health Chatbots


Conversational agents, commonly referred to as chatbots, are AI-powered systems
designed to simulate human-like conversations using NLP and machine learning. In
recent years, these systems have been increasingly applied in the healthcare sector,
particularly in the domain of mental health, where they serve as accessible, scalable
tools for providing emotional support, guidance, and therapy.
One of the earliest and most well-documented applications is Woebot, developed by
Stanford researchers and introduced by Fitzpatrick et al. (2017). Woebot uses
principles of Cognitive Behavioral Therapy (CBT) to engage users in daily mood
tracking and therapeutic conversations. Clinical trials demonstrated a reduction in
depressive symptoms among users after just two weeks of usage, highlighting the
potential of AI-driven mental health interventions in improving psychological well-
being.
Other prominent tools include Wysa, which combines AI with human coaching, and
Replika, a GPT-based chatbot designed for emotional companionship. Wysa guides
users through evidence-based practices such as mindfulness, journaling, and CBT
techniques. It also offers escalation to human therapists when required. Replika, on
the other hand, emphasizes emotionally supportive conversations and learns from
interactions to personalize its responses.

Diwakar (213022008) 4
Figure 3.1: Architecture of AI Mental Health Chatbot

Table 3.1: Popular Mental Health Chatbots and Their Features

Chatbot Technology Used Features Limitation

Woebot CBT + NLP Mood tracking, daily No voice or emotion


check-ins recognition

Wysa AI + Pre-scripted Mental wellness tools, Limited context


flows journaling awareness

Replika GPT-based AI Personalized Less clinically


conversations, empathy validated

Explanation: This table highlights the comparative strengths and weaknesses of


widely-used mental health chatbots. It shows that while the systems have core
therapeutic features, their lack of multimodal interaction and emotional intelligence
presents limitations.

3.3 Emotion Recognition in Mental Health AI


Emotion recognition is a critical component in developing empathetic AI systems.
Research by Ko (2018) emphasizes the role of facial emotion recognition using CNNs
trained on datasets like FER2013 and CK+ for accurate detection of emotions such as

Diwakar (213022008) 5
sadness, happiness, and anger.
Text-based sentiment analysis using models like VADER, TextBlob, or fine-tuned
BERT models have shown efficacy in detecting emotional cues in user input.
Additionally, multimodal approaches integrating text, speech, and facial expressions
are gaining popularity due to their higher accuracy in understanding user sentiment.
Figure 3.2: Multimodal Emotion Recognition System

Table 3.2: Techniques for Emotion Detection in AI Systems

Modality Techniques Used Strengths Challenges

Text VADER, BERT, Simple, fast, Lacks tonal/emotional


TextBlob adaptable cues

Audio MFCC + Captures vocal tones Sensitive to noise


RNN/CNN

Visual CNNs trained on Effective for facial Requires camera access


FER datasets cues and lighting

Multimodal Fusion of all above High accuracy and Complex

Diwakar (213022008) 6
context awareness implementation

Explanation: The table shows how emotion detection varies across different input
modalities, highlighting that multimodal approaches are more robust, though
technically demanding.

3.4 Natural Language Processing and Dialogue Generation


The advent of transformer-based models like GPT-2, GPT-3, DialoGPT, and BERT
has revolutionized dialogue generation. These models enable chatbots to generate
human-like, context-aware responses. Zhang et al. (2020) demonstrated the use of
DialoGPT in open-domain dialogue generation with emotional relevance.
In mental health applications, generating supportive and non-judgmental responses is
critical. Literature indicates that models fine-tuned on mental health-related
conversations improve therapeutic quality. However, careful filtering of training data
and response validation are necessary to avoid generating insensitive or incorrect
advice.
Table 3.3: Comparison of AI Models for Dialogue Generation

Model Strengths Weaknesses Use in Therapy


Chatbots

GPT-3 Highly fluent, May generate unsafe High


contextual responses

BERT Great for classification, Not generative Moderate (intent


Q&A detection)

DialoGP Tuned for dialogue Lacks emotional High


T generation grounding

T5 Multi-task capabilities Slower inference Moderate

3.5 3D Avatar-Based Interaction and Realism


Human-like avatars enhance user engagement in digital therapy. Research by
Bickmore and Picard (2005) highlights that embodied conversational agents (ECAs)
with realistic avatars build user trust and increase retention in long-term therapeutic

Diwakar (213022008) 7
interventions.
Platforms like Ready Player Me offer customizable avatars that can be integrated into
3D web applications. Studies suggest that animated avatars synchronized with audio
(lip-syncing) and emotional expressions significantly improve the realism and
therapeutic effectiveness of AI therapists.

3.6 Voice Interaction and Speech Emotion Recognition


Speech emotion recognition (SER) uses audio features such as pitch, tone, and speech
rate to detect user emotions. Studies by Schuller et al. (2011) explored the use of deep
learning techniques such as recurrent neural networks (RNNs) for classifying
emotions from audio input.
Table 3.4: Voice-Based Emotion Recognition Methods

Methodology Features Pros Cons


Used

MFCC + RNN Pitch, energy, Effective in clean Noise-sensitive,


timbre environments language limitations

Spectrogram + Frequency High accuracy with Computationally


CNN patterns enough data expensive

Hybrid Models MFCC + Robust multi- Requires GPU and large


Deep CNN emotion detection dataset

Explanation: This table outlines the different methods used in SER, highlighting
trade-offs between accuracy, noise robustness, and computational requirements.

3.7 Security, Privacy, and Ethical Considerations


Mental health data is highly sensitive. Literature emphasizes the importance of
encryption, anonymization, and ethical AI usage. Studies by Price and Cohen (2019)
discuss GDPR-compliant data handling practices and informed consent mechanisms
that should be embedded into AI systems.
AI therapists must also be designed to detect crisis scenarios (e.g., suicidal ideation)
and escalate them appropriately, either through alerts or referrals to professional

Diwakar (213022008) 8
services.

3.8 Summary of Literature Gaps


 Many current mental health chatbots lack real-time multimodal emotion
detection.

 Absence of integrated 3D avatars in open-source mental health solutions.

 Limited personalization in therapy recommendations based on user history.

 Inadequate handling of crisis situations by rule-based systems.

 Scarce regional language and dialect support for inclusivity.

Diwakar (213022008) 9
CHAPTER- 4
SYSTEM ANALYSIS

4.1 Existing System

In the realm of digital mental health support, several AI-driven chatbots have
emerged to provide users with conversational therapy and emotional assistance.
Notable examples include:

 Woebot: A chatbot based on Cognitive Behavioral Therapy (CBT) principles.


It provides pre-scripted responses aimed at helping users manage negative
thought patterns.

 Wysa: A mental health app using AI and professional therapist guidance. It


focuses on journaling and guided self-help using evidence-based tools.

 Replika: A general-purpose AI companion designed for open-ended


conversations, often used for emotional support.

 Youper: Combines AI with psychological techniques to track emotions and


provide brief therapeutic conversations.

These systems primarily operate through text-based interfaces on mobile or web


platforms. Most offer pre-defined therapeutic interventions, and some provide
limited AI-generated responses.

Drawbacks of the Existing Systems:

While current solutions have positively impacted mental health accessibility, they
suffer from the following limitations:

 Lack of Visual Engagement: Most systems rely solely on text-based


interfaces without any visual avatars or embodied AI, making them less
immersive and emotionally engaging.

Diwakar (213022008) 10
 No Real-Time Voice Interaction: Few systems provide dynamic text-to-
speech capabilities or natural voice outputs, reducing the sense of presence and
realism.
 Limited Flexibility in Conversations: Bots like Woebot follow a structured
CBT model with minimal freedom for general conversation or natural dialogue
branching.
 No Web-Based 3D Interfacing: There is a noticeable absence of web-based
mental health assistants that combine 3D avatars, voice response, and chatbot
intelligence in one unified platform.

4.2 Proposed System

The proposed project introduces a web-based Milk Quality Prediction that integrates
multiple cutting-edge technologies to overcome the shortcomings of existing systems.
The system features:

 React Frontend with 3D Avatar: Utilizes Ready Player Me avatars rendered


in a React environment to create a friendly, approachable virtual therapist.

 Voice Output using [Link]: Provides real-time text-to-speech responses,


offering a more human-like interaction.

 Backend Integration with Flask: A lightweight Flask server facilitates media


handling, routing, and potential extensions like session tracking or logging.

 Gemini API for AI Conversation: Powers the natural language


understanding and response generation, allowing the chatbot to converse
meaningfully on mental health topics.

 Modular Architecture: Designed with separation of concerns, making it


easier to integrate emotion detection, journaling, video analysis, or database
logging in the future.

Advantages of the Proposed System:

 Enhanced User Engagement: The inclusion of a 3D avatar and voice output


creates a more immersive and empathetic environment.

 Web-Based Accessibility: Runs directly in modern browsers, removing the

Diwakar (213022008) 11
need for installation and increasing platform independence.

 Customizability and Extensibility: The modular architecture supports future


upgrades such as user accounts, emotion detection, or personalized coping
strategies.

 Natural Conversations: The use of Gemini API provides more fluid and
contextual AI responses compared to pre-scripted bots.

 Lightweight and Scalable Backend: Flask provides a scalable backend suited


for both small deployments and future expansions.

 Privacy-Friendly: As a client-server system, it allows for local deployment if


needed, enhancing data privacy for sensitive mental health interactions.

Table 4.1: Comparison of Existing Systems vs Proposed System

Feature / Criteria Woebot Wysa Replika Proposed System


(AI Therapist)

Interface Type Text-based Text-based Text-based Web-based with 3D


avatar

Voice Output ❌ ❌ Limited ✅ ([Link] integration)

CBT Integration ✅ ✅ ❌ Partially (contextual


AI via Gemini)

Flexibility in Low Medium High High (via Gemini


Conversation (Scripted) API)

Emotion ❌ Limited ❌ ❌ (future-ready for


Detection integration)

Avatar ❌ ❌ ❌ ✅ (Ready Player Me


Integration in React)

Scalable Backend ❌ (App- ❌ (App- ❌ (App- ✅ (Flask-powered

Diwakar (213022008) 12
based) based) based) modular backend)

Platform ❌ (Mobile) ❌ (Mobile) ✅ ✅ (Browser-based)


Independence

Diwakar (213022008) 13
CHAPTER- 5
SYSTEM DESIGN
5.1 Architecture Diagram
The architecture of the Milk Quality Prediction system follows a modular, decoupled
structure combining a ReactJS frontend with a Flask-based backend, along with
external APIs and voice libraries.
Components:
1. Frontend (ReactJS):
o Renders the 3D AI Avatar using Ready Player Me.
o Integrates JavaScript say library for real-time text-to-speech output.
o Handles user input, UI interactions, and displays responses.
2. Backend (Flask):
o Receives user input via API calls.
o Interfaces with the Gemini API for generating chatbot responses.
o Manages routing, media handling, and response delivery to the
frontend.
3. Gemini API:
o Processes mental health queries and returns contextual, emotionally
aware responses.
o Acts as the conversational brain of the system.
4. Voice Engine ([Link]):
o Converts generated text into speech using browser-compatible TTS.

Diwakar (213022008) 14
Figure 5.1: System Architecture Diagram

5.2 Data Flow Diagram (DFD)


The Data Flow Diagram represents how data moves through the system in response to
user actions.
DFD Levels:
 Level 0 (Context Diagram):
o Shows the user interacting with the system.
o Basic flow: User Input → System → Response Output.
 Level 1:
1. User submits a mental health query through the chatbot interface.
2. Query is sent to Flask backend.
3. Backend calls Gemini API and receives response.
4. Response is processed and returned to the frontend.
5. [Link] generates voice output and displays avatar animation.

Diwakar (213022008) 15
Figure 5.2: Data Flow Diagram

Diwakar (213022008) 16
CHAPTER- 6
IMPLEMENTATION
6.1 Overview
The Milk Quality Prediction application was implemented as a full-stack web solution
combining conversational AI, avatar interaction, voice synthesis, and a web-based
interface. The primary goal of this implementation was to create a scalable, interactive
system that enables users to engage in supportive, human-like mental health
conversations.

6.2 Technologies used

Table 6.1: Technologies used

Component Technology Purpose

Frontend [Link] Building the dynamic UI and


rendering the 3D avatar

Avatar System Ready Player Me + Generating and displaying the


[Link] animated 3D character

Voice Synthesis [Link] Converting AI text responses to


natural-sounding speech

AI Chat Logic Google Gemini API Generating empathetic, intelligent


replies

Backend Server Flask (Python) Serving media, routing data, and


handling API calls

Communicatio RESTful API / Fetch Connecting frontend and backend via


n asynchronous requests

Deployment Localhost / GitHub Testing and showcasing the


Pages + Flask application

Diwakar (213022008) 17
6.3 Coding & Modules

The entire system is divided into functional modules for separation of concerns:

1. User Interaction Module ([Link])

 Handles user input from chat UI

 Sends the input to the backend via fetch API

 Displays the chatbot’s response in text

 Plays the response using [Link]

Key File: [Link]

import { createContext, useContext, useEffect, useState, useRef } from "react";

const backendUrl = [Link].VITE_API_URL || "[Link]

const ChatContext = createContext();

export const ChatProvider = ({ children }) => {

// State for bot message queue (audio playback)

const [messages, setMessages] = useState([]);

const [message, setMessage] = useState(null); // Current message for audio playback

// State for general loading and UI

const [loading, setLoading] = useState(false);

const [cameraZoomed, setCameraZoomed] = useState(true);

// State for conversation history

const [chatHistory, setChatHistory] = useState([]);

// State for speech recognition

const [isRecording, setIsRecording] = useState(false);

const [transcribedText, setTranscribedText] = useState(""); // Stores the *final*

Diwakar (213022008) 18
transcript

const [isSpeechRecognitionSupported, setIsSpeechRecognitionSupported] =


useState(false);

const recognitionRef = useRef(null);

const finalTranscriptRef = useRef(""); // Ref to store final transcript reliably

// --- Speech Recognition Setup ---

useEffect(() => {

const SpeechRecognition = [Link] ||


[Link];

if (SpeechRecognition) {

setIsSpeechRecognitionSupported(true);

[Link] = new SpeechRecognition();

[Link] = false; // Stop listening after pause

[Link] = 'en-US';

[Link] = false; // We only want final results

[Link] = () => {

[Link]('Voice recognition started.');

[Link] = ""; // Clear previous final transcript

setTranscribedText(""); // Clear state for UI feedback if needed

setIsRecording(true);

};

[Link] = (event) => {

// Get the final transcript from the last result segment

const transcript = [Link][[Link] - 1][0].[Link]();

Diwakar (213022008) 19
[Link]('Speech recognized (final):', transcript);

[Link] = transcript; // Store final transcript in ref

// Optionally update state if you need live feedback, but ref is safer for triggering
chat

// setTranscribedText(transcript);

};

[Link] = (event) => {

[Link]('Speech recognition error:', [Link]);

// Handle specific errors if needed, e.g., 'no-speech', 'audio-capture'

setIsRecording(false); // Reset recording state

[Link] = ""; // Clear any partial transcript on error

};

[Link] = () => {

[Link]('Voice recognition ended.');

setIsRecording(false);

// --- Trigger chat *after* recognition ends and if we have a final transcript ---

const finalTranscript = [Link];

if (finalTranscript) {

// Set the state here just before calling chat, if UI needs it

setTranscribedText(finalTranscript);

// Call chat with the final transcript

chat(finalTranscript); // Pass the final transcript directly

[Link] = ""; // Clear the ref after processing

Diwakar (213022008) 20
};

} else {

[Link]('Web Speech API is not supported in this browser.');

setIsSpeechRecognitionSupported(false);

// Cleanup

return () => {

if ([Link]) {

[Link](); // Stop recognition if component unmounts

};

}, []); // Run once on mount

// --- Recording Control ---

const startRecording = () => {

if ([Link] && !isRecording && !loading) {

// Clear any stale transcript from previous attempts

[Link] = "";

setTranscribedText("");

[Link]();

};

const stopRecording = () => {

// Manually stopping should trigger 'onend' which handles the logic

if ([Link] && isRecording) {

Diwakar (213022008) 21
[Link]();

};

// --- Chat Function (Handles both text input and transcribed audio) ---

const chat = async (messageText) => {

if (!messageText || ![Link]()) {

[Link]("Chat attempt with empty message blocked.");

return; // Don't proceed if the message is empty

[Link]("Sending message to chat:", messageText);

setLoading(true);

// Add user message to history *first*

setChatHistory((prev) => [...prev, { sender: 'user', text: messageText }]);

// Clear the transcribed text state used for UI feedback *after* adding to history

// If chat was called via text input, messageText is not from transcription,

// so clearing transcribedText state here is fine.

setTranscribedText("");

try {

const response = await fetch(`${backendUrl}/chat`, {

method: "POST",

headers: {

"Content-Type": "application/json",

},

Diwakar (213022008) 22
body: [Link]({ message: messageText }), // Send the actual message text

});

if (![Link]) {

// Log specific error from backend if available

const errorBody = await [Link]();

[Link]("API Error Response:", errorBody);

throw new Error(`HTTP error! status: ${[Link]}`);

const data = await [Link]();

const botResponses = [Link] || []; // Expecting an array like [{ text: "...",


audio: "..." }, ...]

if ([Link] > 0) {

// Add bot responses to history

const newHistoryMessages = [Link](msg => ({

sender: 'bot',

text: [Link] || "(Bot response has no text)", // Ensure text exists

}));

setChatHistory((prev) => [...prev, ...newHistoryMessages]);

// Add bot responses to the playback queue

setMessages((prev) => [...prev, ...botResponses]);

} else {

[Link]("Received empty 'messages' array from bot.");

// Add a placeholder if desired

Diwakar (213022008) 23
setChatHistory((prev) => [...prev, { sender: 'bot', text: "(No specific
response)" }]);

} catch (error) {

[Link]("Failed to fetch chat response:", error);

// Add error message to history

setChatHistory((prev) => [...prev, { sender: 'bot', text: "Sorry, I couldn't get a


response. Please try again." }]);

} finally {

setLoading(false); // Ensure loading is turned off

};

// --- Audio Playback Handling ---

const onMessagePlayed = () => {

// Remove the message that just finished playing from the playback queue

setMessages((prev) => [Link](1));

};

// Update the current message for playback when the queue changes

useEffect(() => {

setMessage([Link] > 0 ? messages[0] : null);

}, [messages]);

// --- Context Provider ---

// Remove the useEffect that previously watched isRecording and transcribedText,

// as chat is now called directly from the recognition 'onend' event.

Diwakar (213022008) 24
return (

<[Link]

value={{

loading,

cameraZoomed,

setCameraZoomed,

message, // Current message for audio playback

chatHistory, // Expose history

isRecording, // Expose recording status

isSpeechRecognitionSupported, // Expose support status

startRecording, // Expose function

stopRecording, // Expose function

chat, // Expose chat function (for text input)

onMessagePlayed, // Expose playback callback

// messages queue is internal, message object is exposed for playback component

// transcribedText state is mostly internal now, expose if needed for UI indication

}}

>

{children}

</[Link]>

);

};

// --- Custom Hook ---

export const useChat = () => {

Diwakar (213022008) 25
const context = useContext(ChatContext);

if (!context) {

throw new Error("useChat must be used within a ChatProvider");

return context;

};

2. AI Chat Processing Module (Flask + Gemini API)

 Receives message via POST request

 Sends the message to the Gemini API

 Returns the generated response in JSON

Key File: [Link]

@[Link]('/chat', methods=['POST'])

def chat():

user_input = [Link]['message']

response = model.generate_content(user_input)

return jsonify({'reply': [Link]})

3. Avatar Rendering Module (React + [Link])

 Loads and displays 3D Ready Player Me avatar

 Supports animations and interactivity

Key File: [Link]

import React, { useEffect, useRef, useState } from "react";

// Import helpers like useGLTF and useAnimations from drei

import { useGLTF, useAnimations } from '@react-three/drei';

// Import core hooks like useGraph and useFrame from fiber

Diwakar (213022008) 26
import { useGraph, useFrame } from '@react-three/fiber';

import { SkeletonUtils } from 'three-stdlib';

import * as THREE from "three";

import { useChat } from "../hooks/useChat"; // Assuming this path is correct

import { button, useControls } from "leva";

const facialExpressions = {

default: {},

smile: {

browInnerUp: 0.17,

eyeSquintLeft: 0.4,

eyeSquintRight: 0.44,

noseSneerLeft: 0.1700000727403593,

noseSneerRight: 0.14000002836874015,

mouthPressLeft: 0.61,

mouthPressRight: 0.41000000000000003,

},

funnyFace: {

jawLeft: 0.63,

mouthPucker: 0.53,

noseSneerLeft: 1,

noseSneerRight: 0.39,

mouthLeft: 1,

eyeLookUpLeft: 1,

Diwakar (213022008) 27
eyeLookUpRight: 1,

cheekPuff: 0.9999924982764238,

mouthDimpleLeft: 0.414743888682652,

mouthRollLower: 0.32,

mouthSmileLeft: 0.35499733688813034,

mouthSmileRight: 0.35499733688813034,

},

sad: {

mouthFrownLeft: 1,

mouthFrownRight: 1,

mouthShrugLower: 0.78341,

browInnerUp: 0.452,

eyeSquintLeft: 0.72,

eyeSquintRight: 0.75,

eyeLookDownLeft: 0.5,

eyeLookDownRight: 0.5,

jawForward: 1,

},

surprised: {

eyeWideLeft: 0.5,

eyeWideRight: 0.5,

jawOpen: 0.351,

mouthFunnel: 1,

browInnerUp: 1,

Diwakar (213022008) 28
},

angry: {

browDownLeft: 1,

browDownRight: 1,

eyeSquintLeft: 1,

eyeSquintRight: 1,

jawForward: 1,

jawLeft: 1,

mouthShrugLower: 1,

noseSneerLeft: 1,

noseSneerRight: 0.42,

eyeLookDownLeft: 0.16,

eyeLookDownRight: 0.16,

cheekSquintLeft: 1,

cheekSquintRight: 1,

mouthClose: 0.23,

mouthFunnel: 0.63,

mouthDimpleRight: 1,

},

crazy: {

browInnerUp: 0.9,

jawForward: 1,

noseSneerLeft: 0.5700000000000001,

noseSneerRight: 0.51,

Diwakar (213022008) 29
eyeLookDownLeft: 0.39435766259644545,

eyeLookUpRight: 0.4039761421719682,

eyeLookInLeft: 0.9618479575523053,

eyeLookInRight: 0.9618479575523053,

jawOpen: 0.9618479575523053,

mouthDimpleLeft: 0.9618479575523053,

mouthDimpleRight: 0.9618479575523053,

mouthStretchLeft: 0.27893590769016857,

mouthStretchRight: 0.2885543872656917,

mouthSmileLeft: 0.5578718153803371,

mouthSmileRight: 0.38473918302092225,

tongueOut: 0.9618479575523053,

},

};

const corresponding = {

A: "viseme_PP",

B: "viseme_kk",

C: "viseme_I",

D: "viseme_AA",

E: "viseme_O",

F: "viseme_U",

G: "viseme_FF",

H: "viseme_TH",

X: "viseme_PP",

Diwakar (213022008) 30
};

let setupMode = false;

export function Avatar(props) {

// Model loading setup remains the same, useGLTF is now imported correctly

const { scene: modelScene } = useGLTF('/models/[Link]');

// Animation loading setup remains the same, useGLTF is now imported correctly

const { animations } = useGLTF("/models/[Link]");

// Clone the scene for independent use

const clone = [Link](() => [Link](modelScene),


[modelScene]);

const { nodes, materials } = useGraph(clone); // useGraph is correctly imported from


fiber

const { message, onMessagePlayed, chat } = useChat();

const [lipsync, setLipsync] = useState();

const [facialExpression, setFacialExpression] = useState("");

const [audio, setAudio] = useState();

const [blink, setBlink] = useState(false);

const [winkLeft, setWinkLeft] = useState(false);

const [winkRight, setWinkRight] = useState(false);

// Animation setup

const group = useRef();

// useAnimations is now imported correctly from drei

const { actions, mixer } = useAnimations(animations, group);

const [animation, setAnimation] = useState(

Diwakar (213022008) 31
[Link]((a) => [Link] === "Idle") ? "Idle" : animations[0]?.name // Safer
access to name

);

// --- Resource Disposal Effect ---

useEffect(() => {

// This cleanup function runs when the component unmounts

return () => {

[Link]("Disposing of avatar resources...");

if (clone) {

[Link]((child) => {

if ([Link]) {

// Dispose of geometry

if ([Link]) {

[Link]();

// Dispose of materials

if ([Link]) {

// If material is an array, dispose each material

if ([Link]([Link])) {

[Link]((material) => [Link]());

} else {

// Dispose single material

[Link]();

Diwakar (213022008) 32
}

});

// Note: Textures are typically disposed along with materials if they have maps.

// If you have custom textures not part of materials, dispose them separately here.

// Optional: Dispose of the original loaded scene if it's not managed by useGLTF
caching

// [Link](... dispose logic ...); // Use with caution if useGLTF


handles caching

[Link]("Avatar resources disposed.");

};

}, [clone]); // Dependency on clone ensures cleanup runs when clone changes


(though useMemo makes it stable)

// Effect for handling incoming messages

useEffect(() => {

// [Link](message); // Keep for debugging if needed

if (!message) {

setAnimation("Idle");

// Optionally stop existing audio if message becomes null

if (audio) {

[Link](); // Use pause for HTMLAudioElement

[Link] = 0;

setAudio(undefined);

Diwakar (213022008) 33
}

return;

setAnimation([Link] || "Idle"); // Default to Idle if not specified

setFacialExpression([Link] || "default"); // Default expression

setLipsync([Link]);

// Stop previous audio before playing new one

if (audio) {

[Link](); // Use pause

[Link] = 0;

// Assuming audio is Base64 encoded WAV from backend

const newAudio = new Audio("data:audio/wav;base64," + [Link]);

[Link]().catch(e => [Link]("Audio play failed:", e)); // Add catch for


autoplay policy issues

setAudio(newAudio);

[Link] = () => {

onMessagePlayed(); // Call the callback from useChat hook

setAudio(undefined); // Clear audio state when ended

setAnimation("Idle"); // Optionally return to Idle after speaking

};

// Cleanup function to stop audio if component unmounts or message changes


quickly

return () => {

Diwakar (213022008) 34
if (newAudio && ![Link]) {

[Link](); // Use pause instead of stop if stop is not available

[Link] = 0;

};

}, [message, onMessagePlayed]); // Added onMessagePlayed to dependency array if


it can change

// Effect for playing the current animation

useEffect(() => {

// Ensure actions and the specific animation exist

if (actions && actions[animation]) {

const currentAction = actions[animation];

// Fade in the new animation

currentAction

.reset()

.fadeIn([Link] === 0 ? 0 : 0.5)

.play();

// Fade out the previous animation on cleanup

return () => {

if (currentAction) { // Check if action still exists on cleanup

[Link](0.5);

};

Diwakar (213022008) 35
// Cleanup in case the animation name doesn't exist in actions

return () => {

// Optional: Stop all animations if the desired one isn't found?

// Or fade out the currently playing one if different?

// [Link](); // Could be too abrupt

};

}, [animation, actions, mixer]); // Dependencies are correct

// Leva Controls state - needed for lerpMorphTarget's set call

const [, set] = useControls("MorphTarget", () =>

[Link](

{},

// Check if [Link] exists before trying to access its properties

...(nodes?.Wolf3D_Head?.morphTargetDictionary ?
[Link](nodes.Wolf3D_Head.morphTargetDictionary).map((key) => {

// Ensure influences array exists and has an entry for the dictionary key

const influenceValue = nodes.Wolf3D_Head.morphTargetInfluences &&


nodes.Wolf3D_Head.morphTargetDictionary[key] !== undefined

?
nodes.Wolf3D_Head.morphTargetInfluences[nodes.Wolf3D_Head.morphTargetDicti
onary[key]]

: 0; // Default to 0 if not found

return {

[key]: {

label: key,

value: influenceValue, // Use the actual current influence as initial value

Diwakar (213022008) 36
min: 0, // Morph target range is typically 0 to 1

max: 1,

onChange: (val) => {

if (setupMode) {

// Pass the scene directly to lerpMorphTarget if needed,

// or ensure it's accessible in that scope.

// Using `clone` from useGraph's scope directly is fine here.

lerpMorphTarget(clone, key, val, 1); // Pass scene, target, value, speed

},

},

};

}) : []) // Return empty array if Wolf3D_Head or dictionary doesn't exist

);

// Helper function to smoothly change morph targets

const lerpMorphTarget = (objectScene, target, value, speed = 0.1) => {

// Check if objectScene is valid

if (!objectScene) return;

[Link]((child) => {

if ([Link] && [Link]) {

const index = [Link][target];

// Check if index is valid and influences array exists

if (

Diwakar (213022008) 37
index !== undefined &&

[Link] &&

[Link][index] !== undefined

){

[Link][index] = [Link](

[Link][index],

value,

speed

);

// Update Leva control state only if NOT in setup mode and 'set' function is
available

// Note: This direct update might fight with lerping.

// Consider only setting initial values or using a different approach

// for reflecting state back to Leva outside setup mode.

// if (!setupMode && set) {

// try {

// set({ [target]: [Link][index] }); // Reflect lerped


value

// } catch (e) {

// // [Link](`Leva control for ${target} not found or error setting:`, e);

// }

// }

Diwakar (213022008) 38
});

};

// Frame loop for continuous updates (expressions, blinking, lip-sync)

useFrame((state, delta) => { // delta can be used for speed adjustments if needed

// Ensure nodes are available before proceeding

if (!nodes?.Wolf3D_Head?.morphTargetDictionary) return; // Use Wolf3D_Head as


it has most morph targets

const currentFacialExpression = facialExpressions[facialExpression] ||


[Link];

// Update facial expression morph targets (excluding blink/wink and visemes)

if (!setupMode) {

[Link](nodes.Wolf3D_Head.morphTargetDictionary).forEach((key) => {

// Exclude blink, wink, and visemes from direct expression control

if (key === "eyeBlinkLeft" || key === "eyeBlinkRight" ||


[Link](corresponding).includes(key)) {

return; // Handled separately

const targetValue = currentFacialExpression[key] !== undefined ?


currentFacialExpression[key] : 0;

lerpMorphTarget(clone, key, targetValue, 0.1); // Use the clone scene

});

// Update blink/wink targets

lerpMorphTarget(clone, "eyeBlinkLeft", blink || winkLeft ? 1 : 0, 0.5);

lerpMorphTarget(clone, "eyeBlinkRight", blink || winkRight ? 1 : 0, 0.5);

Diwakar (213022008) 39
// Lip Sync Logic

if (!setupMode && message && lipsync && audio && [Link] >= 2) { //
Check audio readyState

const currentAudioTime = [Link];

let appliedMorphTarget = null; // Track the single active viseme

for (let i = 0; i < [Link]; i++) {

const mouthCue = [Link][i];

if (

currentAudioTime >= [Link] &&

currentAudioTime <= [Link]

){

const visemeName = corresponding[[Link]];

if (visemeName) {

lerpMorphTarget(clone, visemeName, 1, 0.2);

appliedMorphTarget = visemeName;

break; // Found the active cue

[Link](corresponding).forEach((visemeName) => {

if (visemeName !== appliedMorphTarget) {

lerpMorphTarget(clone, visemeName, 0, 0.1);

});

Diwakar (213022008) 40
} else if (!setupMode) {

[Link](corresponding).forEach((visemeName) => {

lerpMorphTarget(clone, visemeName, 0, 0.1);

});

});

useControls("FacialExpressions", {

chat: button(() => chat()),

winkLeft: button(() => {

setWinkLeft(true);

setTimeout(() => setWinkLeft(false), 300);

}),

winkRight: button(() => {

setWinkRight(true);

setTimeout(() => setWinkRight(false), 300);

}),

animation: {

value: animation || "Idle", // Use current animation state

options: [Link]((a) => [Link]),

onChange: (value) => setAnimation(value),

},

facialExpression: {

value: facialExpression || "default", // Use current expression state

options: [Link](facialExpressions),

Diwakar (213022008) 41
onChange: (value) => setFacialExpression(value),

},

enableSetupMode: button(() => {

[Link]("Setup Mode Enabled");

setupMode = true

disableSetupMode: button(() => {

[Link]("Setup Mode Disabled");

setupMode = false;

}),

logMorphTargetValues: button(() => {

if(!nodes?.Wolf3D_Head?.morphTargetInfluences|| !
nodes?.Wolf3D_Head?.morphTargetDictionary) {

[Link]("Morph target data not available.");

return;

const emotionValues = {};

[Link](nodes.Wolf3D_Head.morphTargetDictionary).forEach((key) => {

// Exclude blink, wink, and visemes from the logged emotion values

if (key === "eyeBlinkLeft" || key === "eyeBlinkRight" ||


[Link](corresponding).includes(key)) {

return;

const index = nodes.Wolf3D_Head.morphTargetDictionary[key];

if (index !== undefined &&

Diwakar (213022008) 42
nodes.Wolf3D_Head.morphTargetInfluences[index] !== undefined) {

const value = nodes.Wolf3D_Head.morphTargetInfluences[index];

if (value > 0.01) { // Only log significant values

emotionValues[key] = parseFloat([Link](4)); // Clean up the number

});

[Link]([Link](emotionValues, null, 2));

}),

}, [animation, facialExpression, animations, chat, nodes, clone]); // Added clone to


dependencies

// Effect for automatic blinking

useEffect(() => {

let blinkTimeout;

const nextBlink = () => {

blinkTimeout = setTimeout(() => {

setBlink(true);

setTimeout(() => {

setBlink(false);

nextBlink(); // Schedule the next blink

}, 150); // Shorter blink duration

}, [Link](2000, 6000)); // Random interval between blinks

};

nextBlink(); // Start the blinking loop

Diwakar (213022008) 43
return () => clearTimeout(blinkTimeout); // Cleanup on unmount

}, []); // Empty dependency array ensures this runs only once

// Render the avatar

return (

// Spread any additional props onto the group

<group {...props} dispose={null} ref={group}>

{/* Ensure nodes and materials are loaded before rendering meshes */}

{nodes && materials && (

<>

<primitive object={[Link]} />

{/* Render skinned meshes, passing morph target props where they exist */}

{nodes.Wolf3D_Hair && <skinnedMesh name="Wolf3D_Hair"


geometry={nodes.Wolf3D_Hair.geometry} material={materials.Wolf3D_Hair}
skeleton={nodes.Wolf3D_Hair.skeleton} />}

{nodes.Wolf3D_Body && <skinnedMesh name="Wolf3D_Body"


geometry={nodes.Wolf3D_Body.geometry} material={materials.Wolf3D_Body}
skeleton={nodes.Wolf3D_Body.skeleton} />}

{nodes.Wolf3D_Outfit_Bottom && <skinnedMesh


name="Wolf3D_Outfit_Bottom"
geometry={nodes.Wolf3D_Outfit_Bottom.geometry}
material={materials.Wolf3D_Outfit_Bottom}
skeleton={nodes.Wolf3D_Outfit_Bottom.skeleton} />}

{nodes.Wolf3D_Outfit_Footwear && <skinnedMesh


name="Wolf3D_Outfit_Footwear.geometry"
geometry={nodes.Wolf3D_Outfit_Footwear.geometry}
material={materials.Wolf3D_Outfit_Footwear}

Diwakar (213022008) 44
skeleton={nodes.Wolf3D_Outfit_Footwear.skeleton} />}

{nodes.Wolf3D_Outfit_Top && <skinnedMesh name="Wolf3D_Outfit_Top"


geometry={nodes.Wolf3D_Outfit_Top.geometry}
material={materials.Wolf3D_Outfit_Top}
skeleton={nodes.Wolf3D_Outfit_Top.skeleton} />}

{[Link] && <skinnedMesh

name="EyeLeft"

geometry={[Link]}

material={materials.Wolf3D_Eye}

skeleton={[Link]}

morphTargetDictionary={[Link]}

morphTargetInfluences={[Link]}

/>}

{[Link] && <skinnedMesh

name="EyeRight"

geometry={[Link]}

material={materials.Wolf3D_Eye}

skeleton={[Link]}

morphTargetDictionary={[Link]}

morphTargetInfluences={[Link]}

/>}

{nodes.Wolf3D_Head && <skinnedMesh

name="Wolf3D_Head"

geometry={nodes.Wolf3D_Head.geometry}

Diwakar (213022008) 45
material={materials.Wolf3D_Skin}

skeleton={nodes.Wolf3D_Head.skeleton}

morphTargetDictionary={nodes.Wolf3D_Head.morphTargetDictionary}

morphTargetInfluences={nodes.Wolf3D_Head.morphTargetInfluences}

/>}

{nodes.Wolf3D_Teeth && <skinnedMesh

name="Wolf3D_Teeth"

geometry={nodes.Wolf3D_Teeth.geometry}

material={materials.Wolf3D_Teeth}

skeleton={nodes.Wolf3D_Teeth.skeleton}

morphTargetDictionary={nodes.Wolf3D_Teeth.morphTargetDictionary}

morphTargetInfluences={nodes.Wolf3D_Teeth.morphTargetInfluences}

/>}

</>

)}

</group>

);

// Preload models and animations

[Link]('/models/[Link]');

[Link]("/models/[Link]");

4. Voice Synthesis Module ([Link])

 Converts chatbot responses into speech

 Enhances user engagement and realism

Diwakar (213022008) 46
import say from 'say';

const speak = (text) => {

[Link](text); // Play chatbot response via system voice };


CHAPTER- 7
TESTING & VALIDATION
1. Unit Testing
Tested individual components like the Gemini API response handler, speech synthesis
(say), and avatar rendering. Face Recognition Model:
1. Frontend (React JS)
 Avatar Renderer
o Test: Does the avatar render correctly on load?
o Tool: Jest + React Testing Library
o

Figure 7.1: 3D Avatar interacting during a live session

 Voice Output Function ([Link]())


o Test: Is the correct text converted to voice?
o Mock the say function and check that it is called with the correct
arguments.

Figure 7.2: Voice output playing in real-time during a user query

Diwakar (213022008) 47
 Input Field and Chat Display
o Test: When a user types a message, does it appear in the chat window?

Figure 7.3: Gemini-powered chatbot response to a user expressing stress

2. Backend (Flask)
 Streaming Route
o Test: Does the video/audio route return expected content type?
o Check if Flask correctly handles requests/responses.

Figure 7.4: Flask server logs showing active session and chatbot requests

Diwakar (213022008) 48
Table 7.1: Unit Testing Test Cases
Test Componen Description Input Expecte Result
Case ID t d Output
UT01 Chat API Check if "I'm Valid Pass
Endpoint chatbot sad" JSON
returns with
response response
text
UT02 Voice Ensure "Hello Voice Pass
Output correct text- there" played
Function to-speech without
output errors
UT03 Avatar Avatar Page Avatar Pass
Render loads load appears
correctly on in scene
page load

2. Integration Testing
Integration Testing was performed to validate the coordination between the frontend

Diwakar (213022008) 49
([Link]), backend (Flask), and external libraries like Gemini API and [Link]. Multiple
end-to-end scenarios were tested to ensure the system operates reliably as an
interconnected unit. The results showed seamless interaction among modules, with no
major integration failures.

Figure 7.5: Final integrated Milk Quality Prediction interface

Table 7.2 Test Cases


Test Component Input Expected Output Result
Case
ID
TC01 Chat API "I'm feeling Empathetic response from Pass
low" Gemini API
TC02 Voice Output Any chatbot Natural voice playback Pass
response using [Link]()
TC03 Avatar Load Open app 3D avatar renders correctly Pass
with default animation
TC04 API Failure No internet App shows fallback Pass
message or error handling
is invoked
TC05 UI Resize window Layout adjusts correctly Pass
Responsiveness / mobile view without breaking
components

Diwakar (213022008) 50
7.3 Validation
 Chat Validation: Responses from the Gemini API were reviewed for
empathy, contextual relevance, and appropriateness. Results were found
acceptable for non-clinical settings.
 User Feedback (Optional): If any users interacted with the system during
testing, include general feedback (e.g., “Easy to use”, “Feels interactive”).
 Output Consistency: The system consistently produced voice and text outputs
in response to different mental health inputs, indicating functional reliability.

Diwakar (213022008) 51
CHAPTER- 8
RESULTS & DISCUSSION

8.1 Results

The results of the testing phase reflect the system's performance in key areas such as
functional testing, voice output, avatar interaction, and system integration. The
following sections summarize the outcomes from each of these areas:

8.1.1 Functional Testing Results

The Milk Quality Prediction system, powered by the Gemini API, demonstrated
strong performance in handling a variety of mental health-related queries. During
testing, the chatbot accurately generated contextually relevant and emotionally
sensitive responses to topics such as stress, anxiety, and sadness. The chatbot was able
to respond effectively to open-ended questions and provided responses that aligned
with the user's expressed needs. This indicates that the Gemini-based conversational
model is capable of simulating basic therapeutic engagement and offering valuable
mental health support.

8.1.2 Voice Feedback

The integration of the [Link] JavaScript library allowed for seamless conversion of
text-based responses into natural-sounding speech. Testing confirmed that the voice
output was clear, articulate, and properly synchronized with the chatbot's responses.
The real-time voice feedback added a human-like dimension to the interactions,
improving user engagement. The quality of the voice output was consistent throughout
various test cases, providing a more immersive and natural experience for users.

8.1.3 Avatar Interaction

The 3D avatar, created using Ready Player Me and rendered via React Three Fiber,
was an essential part of the user interface. During testing, the avatar maintained a
conversational posture and exhibited realistic body language. It was effectively
synchronized with the voice output and interacted in real time with users. The visual
representation of the avatar enhanced the user's sense of engagement, making the
system feel more lifelike and approachable. This feature significantly contributed to
the overall user experience by offering a visual and emotional interface to the

Diwakar (213022008) 52
conversation.

8.1.4 System Integration Results

The decoupled architecture, with React handling the frontend and Flask managing the
backend, worked as expected, ensuring smooth data flow between components. The
Flask server efficiently managed routing and media transmission, enabling real-time
communication between the chatbot and the user. Integration testing revealed no
major issues in data exchange between the frontend and backend components. This
seamless integration allowed the system to function smoothly, without lag or delays,
during user interactions.

8.2 Discussion

The following discussion interprets the results in the context of the project’s goals and
compares the performance of the proposed system with existing chatbot platforms.

8.2.1 Achievement of Objectives

The system successfully met its primary objective of simulating empathetic, human-
like conversations using a 3D avatar, voice feedback, and AI-driven responses. The
chatbot was able to engage users in meaningful conversations on mental health topics,
providing personalized and relevant responses. By integrating real-time voice and
avatar interaction, the system was able to enhance the user experience, making the
therapy session feel more authentic and human-like.

The use of Gemini API for generating AI responses allowed for flexible and dynamic
conversations, compared to more rigid, pre-scripted chatbot models. This flexibility
makes the system adaptable to a wide range of user inquiries, contributing to its
broader application in mental health support.

8.2.2 System Strengths

 Interactive Interface: The combination of avatar, voice feedback, and real-


time responses significantly enhanced user engagement and provided a more
immersive experience than traditional text-based chatbots. The avatar created a
sense of presence, which is essential in mental health support systems, as it
helps users feel less isolated during conversations.

Diwakar (213022008) 53
 AI-Driven Conversations: The Gemini API-based chatbot facilitated more
adaptive and context-sensitive conversations, making the system flexible in
handling various mental health issues. Unlike traditional chatbots with scripted
responses, the AI model provided a more personalized interaction.

 Platform Independence: Built on web technologies, the system can be


accessed through any modern browser, eliminating the need for specific
installations or compatibility requirements. This enhances accessibility,
allowing users to interact with the therapist from anywhere with an internet
connection.

 Scalable Backend: The use of Flask as the backend framework provides a


lightweight, scalable solution. This makes it easier to expand the system in the
future, such as by adding features like user authentication, storing conversation
histories, or implementing a more advanced AI model for deeper interactions.

8.2.3 System Limitations

Despite the success of the system, several limitations were identified during testing:

 No Real Emotion Detection: Currently, the system does not analyze user
facial expressions or voice tone to detect their emotional state. As a result, the
system’s ability to fully empathize with the user is limited. Adding emotion
detection would significantly improve the personalization and emotional
intelligence of the system.

 Crisis Response Handling: The system does not yet incorporate crisis
management features, such as detecting suicide risk or emergency escalation
protocols. This is a critical aspect for any mental health application, as timely
intervention can save lives. Future iterations of the system should include such
protocols to address high-risk situations effectively.

 Pre-trained Model Dependency: The Gemini API responses are dependent


on the dataset it was trained on. In some cases, especially in highly sensitive
conversations, the chatbot’s responses might lack the necessary clinical
relevance. While the system is designed to be supportive, human intervention
or supervision may be necessary in more complex cases.

Diwakar (213022008) 54
8.2.4 Comparison with Existing Systems

When compared to existing mental health chatbot platforms like Woebot, Replika, and
Wysa, the proposed system offers several unique advantages:

 3D Avatar and Voice Feedback: While systems like Woebot and Replika
provide text-based conversations, they do not incorporate visual elements like
a 3D avatar or voice feedback. The avatar in the proposed system significantly
enhances user interaction, providing a more lifelike experience.

 Flexible Conversations: Unlike Woebot, which is primarily focused on


Cognitive Behavioral Therapy (CBT) methods, the AI-driven nature of this
system allows for broader, more generalized conversations. This makes the
system more adaptable to users' diverse emotional needs.

 Modular Architecture: The system’s architecture (React for frontend and


Flask for backend) is modular and scalable. This makes it easier to extend the
system with new features, such as emotion detection, journaling, or video-
based consultations, which are not supported by all existing systems.

8.2.5 Potential Improvements

 Emotion Detection: Adding emotion recognition capabilities, such as


analyzing facial expressions or voice tone, would significantly enhance the
emotional intelligence of the system. This would allow the chatbot to better
understand the user’s emotional state and adjust its responses accordingly.

 Crisis Response Handling: Incorporating emergency escalation protocols,


such as detecting suicide risk and providing resources for urgent care, is a
crucial feature for any mental health support system. Future versions should
include these safety measures.

 Advanced NLP Models: The use of more advanced NLP models could
improve the chatbot’s ability to understand complex emotional states and
provide more tailored, clinically relevant responses.

Diwakar (213022008) 55
CHAPTER- 9
CONCLUSIONS & FUTURE WORK

9.1 Conclusion
The development of the Milk Quality Prediction marks a significant step towards
leveraging conversational AI and immersive technologies for mental well-being
support. This project successfully combines advanced frontend and backend
technologies to deliver a user-centric virtual therapy assistant that is both interactive
and accessible.
Through the integration of a 3D avatar using [Link] and Ready Player Me, the
application creates a visually engaging interface that simulates human-like presence.
The voice output using the say JavaScript library further enhances the naturalness of
the interaction, making the therapy sessions feel more empathetic and less mechanical.
On the backend, Flask serves as a robust and lightweight framework that efficiently
manages server-side operations, ensuring seamless data routing, real-time
communication, and system scalability.
The system’s use of the Gemini API to power conversations allows it to respond
contextually to a variety of mental health-related queries, from stress and anxiety to
emotional support needs. Unlike traditional CBT-based bots, the flexibility of the
Gemini model provides more generalized and adaptive responses, helping users feel
heard and understood.
While this application does not yet support emotion detection or crisis intervention, its
modular architecture allows easy extension in future iterations. The project also opens
opportunities for integrating features like journaling, user authentication, session
history, and mental health analytics.
In conclusion, this Milk Quality Prediction offers a compelling proof-of-concept for
the role of conversational agents in mental health care. It provides an accessible, non-
judgmental, and engaging platform for users seeking emotional support. As
technology continues to advance, such systems hold promise in supplementing
traditional therapy and reaching underserved populations with limited access to mental
health professionals

Diwakar (213022008) 56
9.2 Future Works
While the current implementation of the Milk Quality Prediction provides a solid
foundation, there are several directions in which the system can be enhanced:
1. Emotion Detection Integration
Future versions can include real-time facial emotion recognition or sentiment
analysis through webcam or voice inputs to personalize responses based on the
user's emotional state. This would allow the system to adapt tone, language,
and suggestions accordingly.
2. User Authentication and Session Logging
Introducing user registration and login functionality would enable session
management, storing past conversations and offering personalized mental
health progress tracking over time.
3. Crisis Detection and Escalation
Incorporating logic to identify signs of crisis or high-risk phrases can enable
the system to recommend immediate helpline support or notify emergency
contacts, making it safer for vulnerable users.
4. Multilingual Support
Expanding the system’s capabilities to support regional and global languages
would increase accessibility and inclusivity for diverse user groups.
5. Mobile Optimization and PWA (Progressive Web App)
Optimizing the platform for mobile and developing it as a PWA would allow
users to access therapy support on-the-go without needing to install native
applications.
6. Integration with Mental Health Resources
The system can link users to credible mental health articles, mindfulness
exercises, or even schedule tele-therapy appointments with human
professionals for more comprehensive care.
7. Therapeutic Journaling & Mood Tracking
Enabling users to record their daily thoughts, track their mood, and visualize
trends over time could help in self-reflection and emotional regulation.

Diwakar (213022008) 57
By expanding the system in these directions, the Milk Quality Prediction can evolve
from a conversational companion to a holistic mental health support platform that
actively promotes emotional well-being and resilience.
CHAPTER- 10
REFERENCES

 [Link]
 [Link]
 [Link]
Time+Sentiment+Analysis+for+Customer+Support
 [Link]
 [Link]
 [Link]
 [Link]
 [Link]
 [Link]
 [Link]
 [Link]
 [Link]

Diwakar (213022008) 58

Common questions

Powered by AI

The proposed system employs a modular architecture using React for the frontend and Flask for the backend, which supports future upgrades like emotion detection and journaling . It uses the Gemini API for generating flexible, context-sensitive AI dialogues, contrasting with the rigid, pre-scripted responses of current systems . Moreover, it integrates a 3D avatar and real-time voice response to address user engagement inadequacies .

The use of a 3D avatar enhances the user experience by providing visual engagement, which contributes to a more lifelike interaction and the feeling of presence during therapy sessions . This visual element makes the sessions feel more authentic and can reduce user isolation by creating a sense of companionship, which is crucial in mental health contexts .

The proposed system enhances user engagement by integrating a 3D avatar and voice output, creating a more immersive and empathetic environment that suggests a more human-like presence . It also runs on modern browsers without requiring installation, which increases its accessibility . The inclusion of these visual and auditory elements significantly improves user interaction compared to current text-only systems like Woebot and Replika .

To improve effectiveness and emotional intelligence, the proposed system could integrate emotion detection through facial expression and voice tone analysis, enhancing its ability to empathize with users . Incorporating crisis response management, such as detecting suicide risk and implementing emergency protocols, would make the system more comprehensive. Furthermore, leveraging advanced NLP models could improve the chatbot’s ability to provide clinically relevant and tailored responses .

Current AI-driven mental health chatbots primarily operate through text-based interfaces, which lack visual engagement, such as the presence of 3D avatars or embodied AI. This makes these systems less immersive and emotionally engaging for users . Additionally, most systems follow structured models like Cognitive Behavioral Therapy (CBT) with pre-defined interventions, limiting flexibility in conversations and preventing natural dialogue branching .

Platform independence is achieved through the use of web-based technologies, allowing the system to run directly in modern browsers without the need for specific software installations . This enhances accessibility, enabling users to engage with the system from any location with internet access and ensures broader reach and ease of use .

The Gemini API allows the chatbot to conduct flexible and dynamic conversations by enabling it to generate contextually-sensitive and empathetic replies, unlike the rigid models of traditional chatbots . This adaptability makes the system capable of covering a broader range of mental health topics, enhancing its applicability to diverse user needs .

The proposed system offers several advantages over existing ones like Woebot or Replika, such as the inclusion of a 3D avatar and voice feedback for richer user interaction . It allows more flexible conversation dynamics due to its AI-driven nature, which accommodates a broader range of topics beyond the structured CBT models used by systems like Woebot . The modular architecture provides a scalable and extensible framework for future enhancements .

To make the virtual therapy assistant more supportive, implementing emotion detection can allow for more personalized interactions by understanding user emotions through voice and facial expressions . Incorporating advanced NLP models can improve the ability to deliver clinically relevant responses. Adding journaling and session tracking can support ongoing therapeutic engagement, while crisis response mechanisms can address high-risk situations effectively .

The proposed system faces challenges such as the lack of real emotion detection capable of analyzing facial expressions or voice tone, which limits its empathetic capability . Additionally, it lacks crisis management features like emergency escalation protocols, which are vital for handling high-risk situations . Dependency on pre-trained models like Gemini API could also lead to inadequacies in clinically relevant responses, suggesting the need for human oversight in complex scenarios .

You might also like