0% found this document useful (0 votes)
6 views21 pages

Hallucination Detection in Ai Responses: B. Tech Project On

The document outlines a B.Tech project titled 'Hallucination Detection in AI Responses' by students Krishan, Amit Kumar, and Ravi at Netaji Subhas University of Technology. It presents a dynamic framework for verifying claims made by AI models to detect factual inaccuracies, detailing a seven-layer architecture for processing and evaluating AI-generated text. The project aims to enhance trust in AI systems by providing explainable risk assessments of generated content.

Uploaded by

RAVI
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views21 pages

Hallucination Detection in Ai Responses: B. Tech Project On

The document outlines a B.Tech project titled 'Hallucination Detection in AI Responses' by students Krishan, Amit Kumar, and Ravi at Netaji Subhas University of Technology. It presents a dynamic framework for verifying claims made by AI models to detect factual inaccuracies, detailing a seven-layer architecture for processing and evaluating AI-generated text. The project aims to enhance trust in AI systems by providing explainable risk assessments of generated content.

Uploaded by

RAVI
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

B.

TECH PROJECT
on
HALLUCINATION DETECTION
IN AI RESPONSES

REPORT SUBMITTED AS PARTIAL REQUIREMENT FOR THE AWARD


OF BACHELOR OF TECHNOLOGY DEGREE

IN

COMPUTER SCIENCE WITH ARTIFICIAL


INTELLIGENCE ENGINEERING

Name Roll Number


Krishan 2022UCA1847
Amit Kumar 2022UCA1877
Ravi 2022UCA8065

Under the Supervision of


Dr. Gaurav Singal

COMPUTER SCIENCE ENGINEERING DEPARTMENT


Netaji Subhas University of Technology, Dwarka, New Delhi.

Academic Session: 2025-2026


Submission Stage: 8th Semester Mid-Semester Evaluation
CERTIFICATE

DEPARTMENT OF COMPUTER SCIENCE & ENGINEERING

This is to certify that the work embodied in the project thesis titled ”Hallucination
Detection in AI Responses” by Krishan (2022UCA1847), Amit Kumar (2022UCA1877),
and Ravi (2022UCA8065) is the bonafide work of the group submitted to Netaji Subhas
University of Technology for consideration in 8th Semester [Link] Project Mid-Semester
Evaluation. The research and implementation work has been carried out by the team
under my guidance and supervision during the academic year 2025-2026. This work has
not been submitted for any other diploma or degree of any university.
On the basis of declaration made by the group, I recommend the project report for
evaluation.

Dr. Gaurav Singal


(Assistant Professor)
Department of Computer Science & Engineering
Netaji Subhas University of Technology

1
CANDIDATE(S) DECLARATION

DEPARTMENT OF COMPUTER SCIENCE & ENGINEERING

We, Krishan (2022UCA1847), Amit Kumar (2022UCA1877), and Ravi (2022UCA8065)


of [Link], Department of Computer Science & Engineering (Artificial Intelligence), hereby
declare that the project thesis titled ”Hallucination Detection in AI Responses” submit-
ted to the Department of Computer Science & Engineering, Netaji Subhas University of
Technology (NSUT) Dwarka, New Delhi, in partial fulfillment of the requirements for the
award of the degree of Bachelor of Technology, is our original work.
We further declare that this manuscript has not previously formed the basis of award
of any degree and all referenced material has been duly cited.

Place: Netaji Subhas University of Technology


Date: 01 March 2026

Krishan (2022UCA1847)
Amit Kumar (2022UCA1877)
Ravi (2022UCA8065)

2
ACKNOWLEDGEMENT

We would like to express our sincere gratitude to our project supervisor Dr. Gaurav
Singal for his guidance, continuous encouragement, and valuable feedback throughout
this work. His insights helped us shape the project from idea to implementation for the
8th semester mid-semester milestone.
We are also thankful to the faculty and staff of the Department of Computer Science &
Engineering, NSUT, for providing infrastructure and support. We acknowledge our peers
and batchmates for constructive discussions during model design, evaluation planning,
and report preparation.

Krishan (2022UCA1847)
Amit Kumar (2022UCA1877)
Ravi (2022UCA8065)

3
Contents

CERTIFICATE 1

CANDIDATE(S) DECLARATION 2

ACKNOWLEDGEMENT 3

1 Abstract 6

2 Introduction 7
2.1 Background . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
2.2 Hallucination in LLMs . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7
2.3 Need for Post-Hoc Verification . . . . . . . . . . . . . . . . . . . . . . . . 7
2.4 Project Direction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7

3 Motivation 8

4 Problem Statement 9

5 Literature Survey 10
5.1 Hallucination Surveys . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10
5.2 Retrieval-Augmented Verification . . . . . . . . . . . . . . . . . . . . . . 10
5.3 NLI for Factual Consistency . . . . . . . . . . . . . . . . . . . . . . . . . 10
5.4 Gap Identified . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10

6 Objectives 11

7 Methodology (Dynamic 7-Layer Architecture) 12


7.1 Layer 1: Target Generation Layer . . . . . . . . . . . . . . . . . . . . . . 12
7.2 Layer 2: Claim Extraction and Normalization . . . . . . . . . . . . . . . 12
7.3 Layer 3: Dynamic Web Retrieval . . . . . . . . . . . . . . . . . . . . . . 13
7.4 Layer 4: Ephemeral Vector Database . . . . . . . . . . . . . . . . . . . . 13

4
7.5 Layer 5: Semantic Filtering . . . . . . . . . . . . . . . . . . . . . . . . . 13
7.6 Layer 6: NLI Verification Engine . . . . . . . . . . . . . . . . . . . . . . 13
7.7 Layer 7: Risk Scoring and Explainable UI . . . . . . . . . . . . . . . . . 14

8 Simulation Platform and Requirements 15


8.1 Software Stack . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
8.2 Hardware/Runtime . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
8.3 Dataset for Evaluation . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
8.4 Mid-Sem Deliverables . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15

9 Results and Discussion (Mid-Sem Progress) 16


9.1 Functional Validation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16
9.2 Evaluation Pipeline . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16
9.3 Discussion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16

10 Future Scope of Work 18

11 Conclusion 19

References 20

5
Chapter 1

Abstract

Large Language Models (LLMs) have transformed how users consume information, yet
they can generate fluent but factually incorrect statements, commonly known as hallu-
cinations. This report presents our project ”Hallucination Detection in AI Responses,”
an end-to-end dynamic framework that verifies model claims using live open-domain ev-
idence and logical inference.
Our architecture follows seven layers: (1) response generation using cloud-hosted
LLMS, (2) claim extraction with normalization and coreference resolution, (3) live web re-
trieval, (4) in-memory ephemeral vector indexing, (5) semantic evidence selection through
cosine similarity, (6) contradiction/entailment scoring via NLI, and (7) risk calibration
with explainable visualization.
A key design principle is stateless in-memory processing for intermediate artifacts.
Claims, embeddings, and temporary indexes are retained only in RAM during execution
and discarded after inference. For quantitative evaluation, we use HaluEval and report
confusion matrix, precision, and recall, along with graphical analysis such as confusion-
matrix heatmaps and precision-recall curves.
This 8th-semester mid-semester submission documents architecture finalization, pipeline
implementation, and evaluation setup for robust factual risk detection in AI-generated
text.

6
Chapter 2

Introduction

2.1 Background

Large Language Models are autoregressive sequence models that predict the next token
based on prior context. Their generative strength enables coherent responses across QA,
summarization, and dialogue tasks. However, linguistic fluency does not guarantee factual
reliability.

2.2 Hallucination in LLMs

Hallucination refers to generation of unsupported, fabricated, or contradictory content


relative to ground truth. These errors are high-risk in education, legal finance, and
healthcare settings where factual correctness is mandatory.

2.3 Need for Post-Hoc Verification

Model-size scaling alone does not eliminate hallucinations. Therefore, a practical solution
is post-hoc claim verification: decompose response into claims, retrieve evidence, and run
logical consistency checks to estimate hallucination risk.

2.4 Project Direction

Our project builds a dynamic open-domain checker that does not depend on static knowl-
edge snapshots. Instead, it fetches live knowledge, reasons over claim-evidence pairs, and
produces explainable risk labels for each sentence.

7
Chapter 3

Motivation

1) Trust Gap in AI Systems: Users often cannot distinguish between true and fab-
ricated statements when output is grammatically correct and confident. This creates an
urgent trust deficit.
2) Need for Explainability: Binary labels alone are insufficient. Users require
evidence-backed explanations: which claim is risky, why it is risky, and what source
contradicts/supports it.
3) Dynamic Knowledge Requirement: Static training corpora become outdated.
A robust detector must use live retrieval to evaluate claims against current information.
4) Practical Deployment Value: An efficient pipeline with in-memory processing
and modular architecture supports real-time deployment and low storage overhead.

8
Chapter 4

Problem Statement

Title: Dynamic Open-Domain Hallucination Detection in AI Responses


Problem: Given a user prompt and an LLM-generated response, identify which
response lines are potentially hallucinated by verifying claim-level factual consistency
against retrieved evidence from live web sources.
Technical Challenges:

• 1) Atomic claim extraction from free-form generated text.


• 2) Coreference normalization (”it”, ”they”, ”this mission”) into explicit entities.
• 3) Retrieval noise filtering from web sources.
• 4) Efficient temporary indexing for fast semantic lookup.
• 5) Robust logical judgment beyond keyword matching.
• 6) Transparent risk visualization for end users.

9
Chapter 5

Literature Survey

5.1 Hallucination Surveys

Ji et al. (2023) provide foundational taxonomy and evaluation directions for hallucination
across NLG systems. Zhang et al. (2023) discuss why modern LLMs still hallucinate
despite scaling.

5.2 Retrieval-Augmented Verification

Lewis et al. (2020) introduced Retrieval-Augmented Generation (RAG), showing exter-


nal knowledge retrieval can improve factual grounding. Reimers and Gurevych (2019)
introduced Sentence-BERT, enabling efficient semantic similarity in embedding space.

5.3 NLI for Factual Consistency

Natural Language Inference models classify entailment, neutral, or contradiction. Cross-


encoders provide stronger claim-evidence reasoning than lexical overlap.

5.4 Gap Identified

Most systems either depend on static corpora or provide weak interpretability. Our
framework combines live retrieval, in-memory vector search, and NLI scoring to deliver
dynamic, explainable, claim-level risk estimation.

10
Chapter 6

Objectives

Primary Objective: Develop an explainable pipeline that flags hallucinated lines in AI


responses using retrieval + semantic filtering + NLI contradiction reasoning.
Specific Objectives:

• 1) Integrate cloud generation models (Groq/HF) for fast response generation.


• 2) Extract and normalize factual claims using spaCy + fastcoref.
• 3) Retrieve live evidence using Wikipedia and DuckDuckGo APIs.
• 4) Build ephemeral vector DB in memory with ChromaDB + MiniLM embeddings.
• 5) Select best evidence by cosine similarity for each claim.
• 6) Classify claim-evidence relation using DeBERTa NLI cross-encoder.
• 7) Produce risk labels and explainable Streamlit dashboard output.
• 8) Evaluate with HaluEval using confusion matrix, precision, and recall.

11
Chapter 7

Methodology (Dynamic 7-Layer


Architecture)

System Flow: User Prompt → LLM Generation → Claim Extraction → Live Retrieval
→ Ephemeral Vector DB → Semantic Filtering → NLI Engine → Risk Scoring → Ex-
plainable Dashboard

7.1 Layer 1: Target Generation Layer

Role: Generate the initial response (which may contain hallucinations).


Tech: Llama-3 (8B) / Mistral-7B via Groq API or Hugging Face Serverless.
Input: User prompt.
Output: Raw response paragraph.

7.2 Layer 2: Claim Extraction and Normalization

Role: Split response into testable claims and resolve pronouns.


Tech: spaCy sentence segmentation + fastcoref coreference resolution.
Operational Note: This processing is done locally in memory. No claim text is sent to
external services at this stage.
Output Example:

• ”Chandrayaan-3 is an Indian lunar mission.”


• ”Chandrayaan-3 was launched on August 1, 2022.”
• ”Chandrayaan-3 successfully landed near the lunar south pole.”

12
7.3 Layer 3: Dynamic Web Retrieval

Role: Fetch fresh evidence for each normalized claim.


Tech: wikipedia and duckduckgo-search (fallback).
Behavior: Top relevant pages/snippets are downloaded into RAM for immediate pro-
cessing. No mandatory persistent storage is required.

7.4 Layer 4: Ephemeral Vector Database

Role: Create temporary semantic index from retrieved text.


Tech: ChromaDB EphemeralClient + all-MiniLM-L6-v2 embeddings.
Process:

• 1) Split retrieved content into sentence chunks.


• 2) Embed chunks using MiniLM.
• 3) Insert vectors into in-memory Chroma collection.

7.5 Layer 5: Semantic Filtering

Role: Select the best evidence sentence for each claim.


Tech: Chroma cosine similarity query.
Output: Top evidence sentence with similarity score and source metadata.

7.6 Layer 6: NLI Verification Engine

Role: Judge logical relation between claim and evidence.


Tech: cross-encoder/nli-deberta-v3-small.
Output: Three probabilities per claim: Entailment, Neutral, Contradiction.

13
7.7 Layer 7: Risk Scoring and Explainable UI

Role: Convert NLI probabilities into actionable risk labels.


Rule (Default):

• If Contradiction ≥ 0.60 → HIGH RISK (hallucinated)


• Else if Entailment ≥ 0.60 → LOW RISK (supported)
• Else → MEDIUM RISK (uncertain)

UI: Streamlit dashboard highlights hallucinated lines in RED and shows: claim text, top
evidence sentence, contradiction/entailment/neutral probabilities, source title and URL.
Data Residency and Security Notes:

• 1) Layers 2, 4, and 5 are in-memory and ephemeral.


• 2) External communication is primarily Layer 1 (generation API) and Layer 3
(retrieval API/web search).
• 3) The architecture minimizes disk footprint and supports stateless execution.

14
Chapter 8

Simulation Platform and


Requirements

8.1 Software Stack

Python 3.10+, streamlit, spacy, fastcoref, wikipedia, duckduckgo-search, chromadb, sentence-


transformers, transformers, torch, scikit-learn, matplotlib, pandas, datasets, python-
dotenv, requests.

8.2 Hardware/Runtime

CPU-only execution is supported for prototype mode. Better latency with GPU for NLI
and embeddings. Internet connectivity required for generation and live retrieval.

8.3 Dataset for Evaluation

HaluEval benchmark for hallucination detection evaluation.

8.4 Mid-Sem Deliverables

End-to-end pipeline modules implemented. Streamlit explainable dashboard integrated.


HaluEval batch runner and metrics scripts completed.

15
Chapter 9

Results and Discussion (Mid-Sem


Progress)

At mid-semester stage, we completed architecture implementation and functional valida-


tion with end-to-end runs.

9.1 Functional Validation

Case Prompt: ”Tell me about the Chandrayaan-3 mission.”


Observed Behavior: Claim extraction identifies mission facts as separate claims. Re-
trieval fetches current evidence from web sources. NLI assigns high contradiction to
incorrect launch date claim. Dashboard marks hallucinated line in RED with source-
backed explanation.

9.2 Evaluation Pipeline

Our evaluation scripts produce: Confusion Matrix, Precision, Recall, Confusion Matrix
Heatmap, Precision-Recall Curve.

9.3 Discussion

Strengths:

• 1) Modular and interpretable architecture.


• 2) Dynamic evidence retrieval instead of static internal memory.
• 3) Claim-level diagnostics improve trust and usability.

16
Current Limitations:

• 1) Retrieval quality can vary by query wording.


• 2) NLI confidence calibration can be domain-dependent.
• 3) Runtime depends on API/network latency.

17
Chapter 10

Future Scope of Work

1) Hybrid Retrieval: Combine Wikipedia, trusted news APIs, and domain-specific


knowledge bases.
2) Better Claim Decomposition: Introduce relation extraction and temporal normal-
ization to improve complex multi-hop fact checks.
3) Calibration and Robustness: Calibrate NLI outputs with temperature scaling and
evaluate domain transfer.
4) Human-in-the-Loop Verification: Allow analyst feedback to refine thresholds and
source quality weighting.
5) Multilingual Extension: Extend extraction, retrieval, and NLI for non-English
content.
6) Production Readiness: Add caching policies, asynchronous retrieval, and observ-
ability dashboards.

18
Chapter 11

Conclusion

This project report presents a complete and practical architecture for hallucination de-
tection in AI responses. Unlike purely confidence-based methods, our system performs
evidence-grounded verification through claim extraction, live retrieval, semantic search,
and NLI reasoning.
The model architecture is aligned with explainable AI principles: every flagged line is
backed by explicit evidence and contradiction probability. The in-memory ephemeral de-
sign supports speed, privacy-aware processing of intermediate data, and minimal storage
overhead.
As an 8th semester mid-semester submission, this report documents a strong and
functional foundation. The next phase focuses on deeper benchmarking, calibration, and
robustness improvements for final deployment-grade performance.

19
References

1. Ji, Z., Lee, N., Frieske, R., et al. (2023). Survey of Hallucination in Natural
Language Generation. ACM Computing Surveys.

2. Zhang, Y., Li, Y., Cui, L., et al. (2023). Siren’s Song in the AI Ocean: A Survey
on Hallucination in Large Language Models. arXiv preprint.

3. Lewis, P., Perez, E., Piktus, A., et al. (2020). Retrieval-Augmented Generation for
Knowledge-Intensive NLP Tasks. NeurIPS.

4. Reimers, N., and Gurevych, I. (2019). Sentence-BERT: Sentence Embeddings using


Siamese BERT-Networks. EMNLP.

5. Sanh, V., Debut, L., Chaumond, J., and Wolf, T. (2019). DistilBERT, a distilled
version of BERT: smaller, faster, cheaper and lighter.

6. Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019). BERT: Pre-training
of Deep Bidirectional Transformers for Language Understanding.

7. Karpukhin, V., Oguz, B., Min, S., et al. (2020). Dense Passage Retrieval for
Open-Domain Question Answering. EMNLP.

8. Thorne, J., Vlachos, A., Christodoulopoulos, C., and Mittal, A. (2018). FEVER: a
large-scale dataset for fact extraction and verification.

9. HaluEval Dataset Card, Hugging Face Datasets (accessed 2026).

10. Groq API Documentation and Hugging Face Inference API Documentation (ac-
cessed 2026).

20

You might also like