An Interactive Question Answer Based System
on Alzheimer’s Disease Using Retrieval
Augmented Generation
Sujoy Sen1 , Samay Sarkar1 , Partha Ghosh2 , Takaaki Goto3 , and Soumya Sen1(B)
1 University of Calcutta, Kolkata, India
iamsoumyasen@[Link]
2 Academy of Technology, Adisaptagram, West Bengal, India
3 Toyo University, Saitama, Japan
tg@[Link]
Abstract. Alzheimer’s Disease (AD) presents profound challenges to healthcare
systems worldwide, necessitating efficient access to accurate information for opti-
mal care delivery. Effective management of AD requires timely access to accurate
information spanning disease etiology, diagnosis, treatment options, and caregiv-
ing strategies. However, the vast and constantly evolving body of AD-related liter-
ature poses a considerable barrier to efficient information retrieval, particularly for
healthcare professionals operating in time-constrained environments. This paper
outlines the objectives of a specialized Retrieval-Augmented Generation (RAG)
system designed for answering questions related to Alzheimer’s Disease (AD),
employing prompt engineering and utilizing Pinecone as the vector database.
With the goal of enhancing accessibility and comprehension of AD-related infor-
mation, the system aims to efficiently retrieve relevant data from diverse sources
and generate contextually relevant answers tailored to user queries. By leverag-
ing advanced techniques in prompt engineering and vector similarity search, the
RAG system empowers healthcare professionals, patients, and caregivers with
timely access to accurate and comprehensive information, ultimately facilitating
informed decision-making and improving patient outcomes in AD management.
Keywords: Alzheimer’s Disease · Question Answering · Retrieval-Augmented
Generation · Prompt Engineering · Pinecone · Vector Database
1 Introduction
Alzheimer’s Disease (AD) presents an escalating challenge in modern healthcare, char-
acterized by its progressive neurodegenerative effects, cognitive decline, and profound
societal impact. With an aging global population, the prevalence of AD continues to
surge, projecting a doubling of cases by 2050, thereby amplifying the strain on healthcare
systems worldwide.
Alzheimer’s disease is a progressive neurodegenerative disorder that predominantly
impacts memory, cognition, and behaviour. It stands as the leading cause of dementia
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2025
G. Hu et al. (Eds.): CAINE 2024, CCIS 2242, pp. 30–40, 2025.
[Link]
An Interactive Question Answer Based System 31
in older adults, responsible for 60–80% of cases. The disease is marked by the build-up
of amyloid plaques and tau tangles in the brain, which results in the death of nerve
cells and subsequent cognitive decline. Initial symptoms often include minor memory
lapses, especially concerning short-term memory, and gradually escalate to significant
difficulties in language, reasoning, and spatial navigation. As the disease advances,
individuals may find it challenging to perform basic daily activities and eventually may
require full-time care.
The exact cause of Alzheimer’s disease is not yet fully understood, but several risk
factors have been identified, including age, genetics, and lifestyle choices. Age is the most
significant risk factor, with the likelihood of developing Alzheimer’s increasing substan-
tially in individuals over the age of 65. While there is no cure for Alzheimer’s, treatments
are available to manage symptoms and improve quality of life. These include medica-
tions that help regulate neurotransmitter activity and behavioural therapy to address
mood and behaviour changes. Researchers continue to investigate potential treatments
and preventative measures, aiming to better understand the disease and develop more
effective interventions.
Amidst this burgeoning crisis, the application of cutting-edge methodologies
becomes imperative. Retrieval-Augmented Generation (RAG), a paradigm that inte-
grates retrieval-based techniques with generative models, emerges as a promising avenue
for revolutionizing Alzheimer’s Disease management. By harnessing the power of large-
scale data repositories and advanced natural language processing (NLP) models, RAG
offers a novel framework for synthesizing information, facilitating decision-making, and
optimizing patient care pathways in the realm of AD.
At its essence, the RAG framework for Alzheimer’s Disease management revolves
around the synergistic interplay of two fundamental components:
Retrieval: The retrieval aspect of RAG involves the systematic aggregation and extrac-
tion of pertinent data from diverse sources, including electronic health records (EHRs),
medical literature, imaging databases, and patient registries. Leveraging sophisticated
information retrieval (IR) techniques, healthcare practitioners can access a wealth of
structured and unstructured data, encompassing clinical assessments, biomarker pro-
files, genetic predispositions, and treatment histories. This comprehensive data retrieval
lays the groundwork for informed decision-making, personalized risk stratification, and
targeted intervention planning in Alzheimer’s Disease management.
Generation: Complementing the retrieval phase, the generation component of RAG
encompasses the utilization of generative models, such as language models and deep
learning architectures, to synthesize contextually relevant insights, recommendations,
and prognostications. Through advanced natural language generation (NLG) techniques,
RAG facilitates the creation of tailored care plans, patient education materials, and care-
giver support resources, thereby empowering stakeholders with actionable information
and enhancing communication within the healthcare ecosystem.
The integration of retrieval-augmented generation in Alzheimer’s Disease manage-
ment is inherently data-centric, leveraging large-scale datasets, ontologies, and knowl-
edge graphs to fuel model training and inference. By capitalizing on state-of-the-art AI
32 S. Sen et al.
technologies, including transformers, attention mechanisms, and reinforcement learn-
ing algorithms, RAG enables the seamless integration of clinical expertise, empirical
evidence, and patient preferences into a unified decision-support framework.
In this documentation, we embark on a comprehensive exploration of the RAG
paradigm in the context of Alzheimer’s Disease management. Drawing upon interdis-
ciplinary insights from computational linguistics, biomedical informatics, and clini-
cal neuroscience, we elucidate the theoretical foundations, methodological intricacies,
and practical applications of RAG in optimizing patient outcomes, enhancing clinical
workflows, and driving innovation in AD care delivery.
As we navigate the evolving landscape of Alzheimer’s Disease management in an
era of data-driven healthcare transformation, the adoption of RAG represents a paradigm
shift towards precision medicine, personalized care delivery, and collaborative decision-
making. By embracing the synergistic potential of retrieval-augmented generation, we
endeavour to mitigate the burden of Alzheimer’s Disease, empower individuals and fam-
ilies affected by the condition, and pave the way towards a brighter future for dementia
care.
The organization of the paper is outlined as follows: Sect. 2 covers the related studies,
followed by the research objectives in Sect. 3. Sect. 4 explains the relevant terminologies.
The proposed methodology is detailed in Sect. 5. In Sect. 6, we conduct a case study to
evaluate the efficiency of the proposed model. Lastly, Sect. 7 presents the conclusion.
2 Related Work
Alzheimer’s disease is one of the threats to the human life as more and more people are
being affected over the time. Modern lifestyle, food habits, psychological issues, genetic
issues are believed to be the different reasons behind this disease. The exact reasons or
trigger points are not known and there is no treatment that can cure Alzheimer’s totally.
Treatments are done to manage symptoms and to slowdown the advancement of the
disease. Use of modern day technology is helping the treatment process of Alzheimer’s.
Early detection of the disease is one of the key issues of the treatment process.
[1] presents a comprehensive survey on different deep learning based methods for
Alzheimer’s disease detection. A simple neural network is used to detect the Alzheimer’s
disease analysing cerebral MRI [2]. Classification of Alzheimer’s disease [3] and its
prodormal stage, MCI (Mild Cognitive Impairment) helps to prevent progress of mem-
ory impairment hence improving the life quality of Alzheimer’s patients. These research
articles are helpful for those actively involved in the research associated with the detec-
tion, prevention and cure of the Alzheimer’s disease. The patient or the family of the
patient are not interested about the research article rather they search the documents
that give the instant result about the cure or treatment of the disease. The research
article described in [4] form LLM and Knowledge graph to answer few questions on
Alzheimer’s disease. The advancement in the area of Generative AI changes the way
users interact with the system. The use of Large Language Model (LLM) [5] helps us
to get curated answer from any GenAI based system. However LLM suffers from many
problems such has hallucination, security issues. Henceforth Retrieval Augmented Gen-
eration (RAG) [6] is evolved to avoid the previously mentioned problems. Using the
An Interactive Question Answer Based System 33
RAG at backend AI chatbots are developed that use NLP [7] and Text mining [8] for
human like conversation. These chatbots are improved with cognitive skills such that
it able to engage clinicians and patients to discuss about patients’ health conditions to
come up with the dimension towards diagnosis and treatment. However concerns are
there about the accuracy of the results that we get from the chatbot. The article in [9]
discusses about the associated benefits, limitation, and risks of GPT-4 as an AI Chatbot
for Medicine. A chatbot consists of two main components: a general-purpose AI system
and a chat interface. This article [9] used GPT-4 (Generative Pretrained Transformer 4)
with a chat interface. In the area of medical science the scope to have an error should be
minimized as any error could be as fatal as death of a person. Hence precautions need
to be taken to make sure that the result is correct so that based on that right answers
are given to the user. These answers are the decisive factors about the treatment. Any
research in the area of medical science based on Generative AI must focus on accurate
answer retrieval. Moreover prompt engineering [10] can be incorporated to fine tune the
answer.
3 Objectives
This research work aims to build a Retrieval Augmented Generation System for AD
question-answering aims to provide accurate, contextually relevant, and user-friendly
answers to queries related to Alzheimer’s Disease. By leveraging cutting-edge gener-
ative AI and prompt engineering, the system will empower healthcare professionals,
caregivers, and patients with timely access to valuable information, ultimately improv-
ing the understanding and management of AD. This system will enhance user experience
and facilitate exploration of AD-related topics through the generation of related queries.
The objective of this research work is pointed below:
i. Providing accurate and up-to-date information using RAG.
ii. Reducing hallucination using RAG and prompt engineering.
iii. Storing the text data in the form of vector database by converting the text to vector
and then applying chunking and embedding.
iv. Related question generation by understanding the pattern of the questions from
the user using prompt engineering.
4 Related Terminologies
4.a. Large Language Models(LLMs): Large language models (LLMs) are a type of
foundational model trained on vast datasets, enabling them to understand and generate
natural language as well as other content types to perform various tasks. These models
utilize deep learning techniques and extensive textual data. Typically based on a trans-
former architecture, such as the generative pre-trained transformer (GPT), they excel
at processing sequential data like text input. LLMs consist of multiple neural network
layers, each with parameters that can be adjusted during training. An additional layer,
known as the attention mechanism, enhances their capability by focusing on specific
parts of the data.
34 S. Sen et al.
During training, LLMs learn to predict the next word in a sentence by considering
the context provided by the preceding words. They achieve this by assigning a proba-
bility score to the recurrence of tokenized words—these tokens are smaller sequences
of characters. These tokens are then converted into embedding, which are numerical
representations of the context. This training involves using an extensive corpus of text,
spanning billions of pages, enabling the model to learn grammar, semantics, and concep-
tual relationships through zero-shot and self-supervised learning. Once trained, LLMs
can generate text by predicting the next word based on the input they receive, utilizing
the patterns and knowledge they have acquired. This results in coherent and contextually
relevant language generation, useful for various natural language understanding (NLU)
and content creation tasks.
Model performance can be enhanced through techniques like prompt engineering,
prompt-tuning, and fine-tuning. Additionally, reinforcement learning with human feed-
back (RLHF) is used to mitigate biases, hateful speech, and “hallucinations” (factually
incorrect answers) that can arise from training on vast amounts of unstructured data.
Ensuring that enterprise-grade LLMs are reliable and safe for use is crucial to avoid
potential liabilities and protect organizational reputations.
4.b. Vector Database: A vector database is designed to store, manage and index massive
quantities of high-dimensional vector data efficiently. In contrast to traditional databases
that handle tabular or document-based data, vector databases are optimized for managing
spatial information represented as geometric shapes such as lines and polygons. Vector
databases organize spatial data using a vector data model, which represents geometric
objects as collections of vertices and edges. Each object is defined by its geometry (shape
and location) and may also include additional attributes such as metadata or descriptive
information.
Vector databases use various indexing techniques to enable faster searching. Vector
indexing along with distance-calculating algorithms such as nearest neighbour search,
are particularly helpful with searching for relevant results across millions if not billions
of data points, with optimized performance.
Vector databases find applications in various domains, including urban planning,
transportation management, environmental modelling, and asset tracking. They are par-
ticularly well-suited for scenarios that require complex spatial queries and advanced
spatial analysis capabilities.
4.c. Prompt Engineering: Prompt engineering refers to the process of designing and
crafting prompts or inputs to language models in order to elicit desired responses or
behaviours. This technique has gained significant attention in the field of natural lan-
guage processing (NLP), particularly with the rise of large language models (LLMs)
such as OpenAI’s GPT (Generative Pre-trained Transformer) series and Google’s BERT
(Bidirectional Encoder Representations from Transformers).
At its core, prompt engineering aims to shape the behaviour of language models
by providing them with relevant context and instructions. This context could include
keywords, phrases, or specific cues tailored to the task at hand. By carefully designing
the input, prompt engineers aim to elicit responses from the language model that align
with the desired task objectives.
An Interactive Question Answer Based System 35
Moreover, prompt engineering often involves iterative experimentation and evalua-
tion. Engineers may test different variations of prompts, fine-tune language models on
task-specific data, and analyse the quality of generated outputs. This iterative approach
allows prompt engineers to refine their prompts, improve the performance of language
models, and address any shortcomings or biases.
4.d. Retrieval Augmented Generation (RAG): Retrieval-Augmented Generation
(RAG) is a novel approach in natural language processing (NLP) that integrates both
retrieval-based and generation-based models. Unlike traditional generative models, RAG
incorporates information retrieval techniques to retrieve relevant context from a large cor-
pus of text data before generating a response. This retrieved context serves as input to the
generation component, which then produces a response informed by the retrieved infor-
mation. RAG models are trained using supervised learning and reinforcement learning
techniques to optimize both the retrieval and generation components. They find applica-
tions in various NLP tasks such as conversational agents, question answering, summa-
rization, and content generation, offering the advantage of producing more informative
and contextually relevant responses compared to purely generative models. However,
RAG also faces challenges such as scalability of the retrieval component, integration of
retrieved context with the generation process, and potential biases in the retrieved data.
Despite these challenges, RAG represents a promising approach in NLP and is expected
to play a significant role in advancing the field.
5 Proposed Methodology
This research work utilizes a conversational AI system that combines advanced tech-
nologies to create an engaging and effective user experience. At the heart of the system
is a sophisticated language model capable of understanding user inputs and generating
contextually relevant responses. This allows the chatbot to hold natural and fluent con-
versations with users, making interactions feel human-like. The system uses memory
management techniques to keep track of conversation history, which helps the chatbot
provide more personalized and consistent answers over time.
To enhance the chatbot’s performance and efficiency, the system leverages a special-
ized database designed for fast searching and data retrieval. This database stores data in
a way that enables the system to quickly find the most similar information when a user
asks a question. By efficiently matching user queries with relevant data, the system can
deliver comprehensive and accurate responses to the user.
In this research work, the process of generating a question-answer (Q&A) model
revolves around effectively managing and searching through a collection of text doc-
uments that have been split into smaller chunks and stored as embedding in a vec-
tor database. This system can be thought of as an advanced information retrieval and
question-answering framework that leverages natural language processing and machine
learning techniques. First, the text documents are processed and divided into manageable
chunks. These chunks are then embedded using a pre-trained model, which transforms
the text into high-dimensional vectors. By storing these vectors in a vector database we
can efficiently retrieve and search for relevant information in response to user queries.
36 S. Sen et al.
When a user submits a query, the query is refined to enhance its relevance and
accuracy. This refinement may include pre-processing steps such as removing stop words,
stemming, or lemmatization. Once the query is ready, it is embedded using the same
model used for the text chunks. The query vector is then used to search the vector database
to find the most similar or relevant text chunks. Once relevant chunks are identified, they
are sent back to the user, along with the original query context, to provide a coherent
and helpful response. This process can enhance the quality and speed of information
retrieval, making it easier for users to access the precise information they need from a
large collection of text data Fig. 1.
Fig. 1. Process Diagram of this RAG system
5.1 Customizing Prompt Engineering
In this research work we have used prompt engineering in three different places, i.e.
Query Refiner, Related Questions Generator and System message template.
An Interactive Question Answer Based System 37
5.1. A) Query Refiner:
The Query Refiner method is designed to enhance the user’s query by refining it for
clarity, grammar, and spelling. The goal is to transform the initial user query into a well-
formed question that can effectively retrieve relevant information from the knowledge
base. The prompt for the query refiner is given below.
Prompt: “Given the following user query and conversation log, formulate a
question that would be the most relevant to provide the user with an answer
from a knowledge base.\n\nCONVERSATION LOG: \n{conversation}\n\nQuery:
{query}\n\nRefined Query:”
The provided prompt guides this process, instructing the model to generate a refined
question based on the user’s original query and chat history. By ensuring that the refined
question is free of grammatical and spelling errors, the prompt aims to improve the
precision of the information retrieval process. The inclusion of an example in the prompt
serves as a helpful guide for the language model, demonstrating the expected output
format and style for the refined question. Overall, this approach enhances the quality
and relevance of the responses provided to the user.
5.1. B) Related Questions Generator:
The Related Questions Generator is an essential part of providing additional context
and depth to the user’s query. This method generates three questions related to the user’s
initial query, offering a broader understanding of the topic and potentially enriching the
user’s experience with the knowledge base. The prompt for the related question generator
is given below.
Prompt: “Provide three diverse, complete one-line questions related to ‘{query}’
that delve into various aspects of the topic, ensuring they are straightforward and directly
related. Avoid repeating the same content as the query. You need not to give any
numbering of the question”.
The prompt guides the model to generate three alternative questions based on the
user’s original query. These questions are intended to be straightforward, accessible, and
directly related to the query. By offering these alternative questions, the model provides
the user with a broader perspective on the topic and encourages further exploration. This
enhances the overall interaction with the knowledge base and supports a richer, more
informative experience for the user.
System Message Template:
The System Message Template in this work outlines how the AI model should handle
user queries and provide responses based on the available context. This method sets
expectations for the model’s behaviour, ensuring it answers questions as accurately as
possible using the provided information from the vector database. The prompt for the
system message template is given below.
Prompt: “Answer the question as truthfully as possible using the provided context
in a minimum of 200 words. Ensure the response is well-structured with proper spacing,
and highlight important words in bold. If the answer includes multiple points, present
the response in a clear point-by-point format with proper spacing between points. If the
answer is not contained within the text below, say ‘I do not know, because it is irrelevant
to our context’”.
38 S. Sen et al.
The prompt for this method instructs the model to respond truthfully and transpar-
ently, indicating when it lacks sufficient information to provide a reliable answer. By
explicitly stating “I DO NOT KNOW, because the data is not present in our Vector DB”
when the necessary data is unavailable, the model fosters trust and clarity in the conver-
sation. This approach helps manage user expectations and establishes a reliable standard
for the AI’s responses, enhancing the overall user experience and communication quality.
6 Case Study
• On the ChatUI page, the user inputs a prompt into the chatbot (Fig. 2). The prompt
is refined (Fig. 2) and passed to the Pinecone vector database.
• The refined prompt generates a response and displays it (Fig. 2).
• Related questions based on the user’s prompt are also generated (Fig. 3).
• If the query is not based on Alzheimer’s, then our system will display “I do not know,
because it is irrelevant to our context.”(Fig. 4)
Fig. 2. Response of the query
Fig. 3. Related questions of the given prompt
An Interactive Question Answer Based System 39
Fig. 4. If the query is irrelevant, the system replies this
7 Conclusion
This research work develops an interactive RAG based question answering system for
Alzheimer’s disease for doctor, patient, caregiver as well as the researchers in this
domain. It uses the advanced techniques such vector database, vector embedding, text
chunking, prompt engineering etc. Prompt engineer is used to perform query refinement,
domain knowledge enhancement, response format. Moreover it is used to generate related
question that help the user of the system to have better interaction. The system also able
to reduces the hallucination using RAG and prompt engineering.
This system could be replicated to other domains by appropriately forming the RAG
from the underlying domain. Moreover as new data are being gathered by the system
through the interaction on Alzheimer’s disease, this text corpus could be analysed further
by using sentiment analysis [11] on the patient and well-wishers of patient for more
effective information generation.
References
1. Saikia, P., Kalita, S.K.: Alzheimer disease detection using MRI: deep learning review. SN
Comput. Sci. 5, 507 (2024). [Link]
2. Alhyane, R., Kassimi El Bakkali, A., Bouroumi, A., Rémy, F., El Boustani, A.: “Detection of
Alzheimer’s Disease using a convolutional neural network”. In: International Conference on
Advanced Intelligent Systems for Sustainable Development. AI2SD (2022). [Link]
10.1007/978-3-031-35248-5_66
3. Li, F., Tran, L., Thung, K.H., Ji, S., Shen, D., Li, J.: A robust deep model for improved
classification of AD/MCI Patients. IEEE J. Biomed. Health Inf. 19(5) 1610−1616 (2015).
[Link]
4. Li, D., et al.: “DALK: Dynamic Co-Augmentation of LLMs and KG to answer Alzheimer’s
Disease Questions with Scientific Literature” (2024). [[Link]
5. Wu, Y.: Large language model and text generation. In: Xu, H., Demner Fushman, D. (eds)
Natural Language Processing in Biomedicine. Cognitive Informatics in Biomedicine and
Healthcare (2024)
40 S. Sen et al.
6. Wang, C., et al.: Potential for GPT technology to optimize future clinical decision-making
using retrieval-augmented generation. Ann. Biomed. Eng. 52, 1115−1118 (2024). [Link]
org/10.1007/s10439-023-03327-6
7. Masoumi, S., et al.: Natural language processing (NLP) to facilitate abstract review in medical
research: the application of BioBERT to exploring the 20-year use of NLP in medical research.
Syst. Rev. 13, 107 (2024). [Link]
8. Roy, S., Cortesi, A., Sen, S.: Context-aware OLAP for textual data warehouses. Int. J. Inf.
Manage. Data Insights 2(2), 100129 (2022)
9. Lee, P., Bubeck, S., Petro, J.: ”Benefits, limits, and risks of GPT-4 as an AI Chatbot for
Medicine” The New England Journal of Medicine Vol. 388(13) 1233−1239 (2023) https://
[Link]/10.1056/NEJMsr2214184v
10. Kansal, A.: “Prompt engineering techniques. In: Building Generative AI-Powered Apps.”
(2024). [Link]/[Link]
11. Ghosh, P., Samanta, O., Goto, T., Sen, S.: Sales forecasting of overrated products: fine tun-
ing of customer’s rating by integrating sentiment analysis. IEEE Access 12, 69578−69592
(2024)[Link]