ECC1332DESIGN PROJECT
DIGITAL MARK EXTRACTION SYSTEM FROM SCANNED
ANSWER SHEETS
A PROJECT REPORT
Submitted by
MADHUMITHA S (927623BEC127)
MADHUMITHA U (927623BEC128)
NIVASHINI K (927623BEC154)
in partial fulfilment for the award of the degree
of
BACHELOR OF ENGINEERING
in
ELECTRONICS AND COMMUNICATION ENGINEERING
[Link] COLLEGE OF ENGINEERING,
KARUR
MAY 2026
[Link] COLLEGE OF ENGINEERING, KARUR
(Autonomous Institution affiliated to Anna University, Chennai)
BONAFIDE CERTIFICATE
Certified that this project report “DIGITAL MARK EXTRACTION
SYSTEM FROM SCANNED ANSWER SHEETS” is the bonafide work of
“MADHUMITHA S (927623BEC127), MADHUMITHA U (927623BEC128),
NIVASHINI K (927623BEC154)” who carried out the project work during the
academic year 2025- 2026 under my supervision.
SIGNATURE SIGNATURE
[Link], M.E., Ph.D Dr. [Link], B.E.,[Link].,Ph.D
HEAD OF THE DEPARTMENT SUPERVISOR
Professor, Associate Professor,
Department of Electronics Department of Electronics
and Communication and Communication
Engineering, Engineering,
[Link] College of Engineering, [Link] College of
Thalavapalayam, Karur-639113 Engineering, Thalavapalayam, Karur-
639113.
This Project Work ECC1332 DESIGN PROJECTreport has been submitted for the End
Semester Project viva voce Examination held on
.
INTERNAL EXAMINER EXTERNAL EXAMINER
VISION AND MISSION OF THE INSTITUTION
VISION
To emerge as a leader among the top institutions in the field of technical education.
MISSION
Produce smart technocrats with empirical knowledge who can surmount
the global challenges.
Create a diverse, fully-engaged learner-centric campus environment to
provide quality education to the students.
Maintain mutually beneficial partnerships with our alumni, industry and
professional associations.
DEPARTMENT OF ELECTRONICS AND COMMUNICATION
ENGINEERING
VISION
To empower the Electronics and Communication Engineering students with
emerging technologies, professionalism, innovative research and social
responsibility.
MISSION
M1: Attain the academic excellence through innovative teaching
learning process, research areas & laboratories and Consultancy projects.
M2: Inculcate the students in problem solving and lifelong learning
ability. M3: Provide entrepreneurial skills and leadership qualities.
M4: Render the technical knowledge and skills of faculty members.
iii
PROGRAMME EDUCATIONAL OBJECTIVES (PEOs)
PEO1: Core Competence: Graduates will have a successful career in academia
or industry associated with Electronics and Communication Engineering.
PEO2: Professionalism: Graduates will provide feasible solutions for the
challenging problems through comprehensive research and innovation in the
allied areas of Electronics and Communication Engineering.
PEO3: Lifelong Learning: Graduates will contribute to the social needs through
lifelong learning, practicing professional ethics and leadership quality.
PROGRAM OUTCOMES(POs)
PO1: Engineering knowledge: Apply the knowledge of mathematics, science,
engineering fundamentals, and an engineering specialization to the solution of
complex engineering problems.
PO2: Problem analysis: Identify, formulate, review research literature, and
analyze complex engineering problems reaching substantiated conclusions using
first principles of mathematics, natural sciences, and engineering sciences.
PO3: Design/development of solutions: Design solutions for complex
engineering problems and design system components or processes that meet the
specified needs with appropriate consideration for the public health and safety,
and the cultural, societal, and environmental considerations.
PO4: Conduct investigations of complex problems: Use research-based
knowledge and research methods including design of experiments, analysis and
interpretation of data, and synthesis of the information to provide valid
conclusions.
PO5: Modern tool usage: Create, select, and apply appropriate techniques,
resources, and modern engineering and IT tools including prediction and
modelling to complex engineering activities with an understanding of the
limitations.
iv
PO6: The engineer and society: Apply reasoning informed by the contextual
knowledge to assess societal, health, safety, legal and cultural issues and the
consequent responsibilities relevant to the professional engineering practice.
PO7: Environment and sustainability: Understand the impact of the professional
engineering solutions in societal and environmental contexts, and demonstrate
the knowledge of, and need for sustainable development.
PO8: Ethics: Apply ethical principles and commit to professional ethics and
responsibilities and norms of the engineering practice.
PO9: Individual and team work: Function effectively as an individual, and as a
member or leader in diverse teams, and in multidisciplinary settings.
PO10: Communication: Communicate effectively on complex engineering
activities with the engineering community and with society at large, such as,
being able to comprehend and write effective reports and design documentation,
make effective presentations, and give and receive clear instructions.
PO11: Project management and finance: Demonstrate knowledge and
understanding of the engineering and management principles and apply these to
one’s own work, as a member and leader in a team, to manage projects and in
multidisciplinary environments.
PO12: Life-long learning: Recognize the need for, and have the preparation and
ability to engage in independent and life-long learning in the broadest context of
technological change.
PROGRAM SPECIFIC OUTCOMES (PSOs)
PSO1: Applying knowledge in various areas, like Electronics, Communications,
Signal processing, VLSI, Embedded systems etc., in the design and
implementation of Engineering application.
PSO2: Able to solve complex problems in Electronics and Communication
Engineering with analytical and managerial skills either independently or in
team using latest hardware and software tools to fulfil the industrial
expectations.
v
ACKNOWLEDGEMENT
We gratefully remember our beloved Founder Chairman, (Late)
Thiru. M. Kumarasamy, whose vision and legacy laid the foundation for our
education and inspired us to successfully complete this project.
We extend our sincere thanks to Dr. K. Ramakrishnan, Chairman, and Mr. K.
R. Charun Kumar, Joint Secretary, for providing excellent infrastructure and
continuous support throughout our academic journey.
We are privileged to extend our heartfelt thanks to our respected
Principal, Dr. B. S. Murugan, [Link]., [Link]., Ph.D., for providing us with a
conducive environment and constant encouragement to pursue this project work.
We sincerely thank Dr. N. Mahendran, B.E., M.E., Ph.D., Professor and Head,
Department of Electronics and Communication Engineering, for his continuous
support, valuable guidance, and motivation throughout the course of this project.
Our special thanks and deep sense of appreciation go to our Course Coordinator,
[Link], B.E., M.E.., Assistant Professor, Department of Electronics and
Communication Engineering, for her exceptional guidance, constructive
suggestions, and unwavering support, all of which have been instrumental in the
successful execution of this project.
We would also like to acknowledge [Link], B.E., [Link]., Ph.D.,
Associate Professor, our Project Supervisor for his constant encouragement,
continuous supervision, and coordination that contributed to the smooth progress and
completion of our project work.
We gratefully thank all the faculty members of the Department of Electronics and
Communication Engineering for their timely assistance, valuable insights, and
constant support during various phases of the project.
Finally, we extend our profound gratitude to our parents and friends for their
encouragement, moral support, and motivation, without which the successful
completion of this project would not have been possible.
vi
ABSTRACT
This project presents an automated system for extracting and processing
student marks from handwritten exam answer sheets using artificial
intelligence. The system accepts images of marksheets as input and utilizes an
AI model to accurately identify and extract key details such as register
numbers, student names, and question-wise marks for Part A and Part B. The
extracted data is then structured into a standardized format, where calculations
including total marks, overall scores, and pass/fail status are automatically
computed. A Flask-based backend handles image processing, data validation,
and Excel file generation, while a React-based frontend provides an intuitive
interface for uploading images and viewing results. The final output is a well-
organized Excel file containing individual student records along with summary
statistics such as class average, highest mark, and lowest mark. This system
reduces manual data entry, minimizes errors, and significantly improves
efficiency in academic result processing, making it suitable for use in schools
and colleges. Additionally, the system incorporates features like batch image
processing, rate-limit handling, and error management to ensure reliability and
scalability. By eliminating manual data entry and reducing human errors, this
project significantly enhances efficiency in academic result processing and
demonstrates the practical application of AI in the education domain.
vii
Abstract (Key words) POs Mapping
Artificial Intelligence, OCR, PO1, PO2, PO3, PO4, PO5, PO6, PO7,
Handwritten Text Recognition, Marks PO8, PO9, PO10, PO11, PO12, PSO1,
Extraction, Image Processing, Data
Automation, Flask, React, Excel PSO2
Generation, Educational Technology,
Data Analysis, JSON Parsing, Web
Application, Result Processing,
Automation System
SDG Goal Remarks
SDG9 This project supports the An efficient and scalable AI-
Sustainable Development based system that automates
Goals by improving the
reliability of industrial power handwritten marks extraction,
systems through smart reducing manual effort and
monitoring.
improving accuracy in
academic result processing.
Project Component Relevant IEEE Standards
Optical Character Recognition IEEE 1857
Machine Learning Process IEEE 2857
Software Requirements Specifications IEEE 830
Software Development Life Cycle IEEE 12207
System and Software Verification IEEE 1012
Image Processing and Evaluation IEEE 29148
viii
TABLE OF CONTENTS
CHAPTER PAGE
CONTENTS
No. No.
INSTITUTION VISION AND MISSION
DEPARTMENT VISION AND MISSION iii
DEPARTMENT PEOs, POs AND PSOs iv
ABSTRACT vii
LIST OF FIGURES xiii
LIST OF ABBREVIATIONS xii
1 INTRODUCTION 1
1.1 Digital Mark Extraction System Overview 1
1.2 Artificial Intelligence in Document Processing 1
1.3 Handwritten Text Recognition Techniques 2
1.4 Image Processing in Educational Systems 2
1.5 Need for Automation in Result Processing 3
1.6 Objective of the Project 4
1.7 Scope of the Project
2 LITERATURE SURVEY 6
2.1 OCR Techniques for Handwritten Documents 6
2.2 AI Based Document Processing System 6
2.3 Neural Networks for Character Recognition 6
2.4 Automated Result Processing System 7
2.5 Image to Data Conversion Using Artificial 7
Intelligence
2.6 Deep Learning for Text Recognition 7
2.7 Machine Learning for Pattern Recognition 8
2.8 Cloud Based Document Processing Systems 8
ix
2.9 Intelligent Automation in Education Systems 9
2.10 Web Based Frameworks for Data Processing 9
2.11 Data Analysis and Processing Using Pandas
2.12 Excel Automation using OpenPyXL
3 EXISTING SYSTEM` 11
3.1 Manual Mark Entry System 12
3.2 Spreadsheet Based Result Processing 12
3.3 Basic OCR Tools 12
3.4 Limitations of Existing System 12
3.5 Error Prone Data Entry Methods 14
3.6 Time Consumption in Result Processing 14
3.7 Lack of Automation
3.8 Data Inconsistency Issues
3.9 Existing Workflow Analysis
4 PROBLEM STATEMENT 15
5 PROPOSED SYSTEM 16
5.1 System Architecture
5.2 Image Upload Module
5.3 Flask Backend Processing
5.4 AI Based Extraction
5.5 JSON Parsing Module
5.6 Data Processing and Computation
5.7 Excel Generation Module
5.8 Download and Output Interface
6 RESULTS AND DISCUSSION 21
6.1 Output Analysis
x
6.2 Accuracy Evaluation
6.3 Performance Analysis
7 CONCLUSION AND FUTURE WORK 24
7.1 Conclusion 24
7.2 Future Work 25
APPENDICES
REFERENCES
xi
LIST OF FIGURES
FIGURE No. TITLE PAGE No.
5.1.1 Overall System Architecture of Digital
Mark Extraction System
5.2.1 Image Upload Interface for Answer Sheets
5.8.1 Excel Report Download Interface
6.1.1 Generated Excel Output of Extracted Marks
xii
LIST OF ABBREVIATIONS
[Link]. ABBREVIATION EXPANSION
1 AI Artificial Intelligence
2 OCR Optical Character Recognition
3 JSON JavaScript Object Notation
5 UI User Interface
6 API Application Programming Interface
7 ML Machine learning
8 DL Deep Learning
9 CSV Comma Separated Values
10 HTTP Hyper Text Transfer Protocol
11 IDE Integrated Development Environment
12 GUI Graphical User Interface
13 DB Database
14 CPU Central Processing Unit
15 GPU Graphics Processing Unit
xiii
CHAPTER 1
INTRODUCTION
1.1 DIGITAL MARK EXTRACTION SYSTEM OVERVIEW
The Digital Mark Extraction System from Scanned Answer Sheets is an
automated system designed to simplify the process of extracting student marks
from handwritten answer sheets. It uses artificial intelligence techniques to read
and interpret handwritten data from images and convert it into a structured
digital format. The system consists of an image input, a processing unit, and an
output generation module that creates an organized Excel file. This system is
widely useful in educational institutions for faster and more accurate result
processing. However, the accuracy of extraction may depend on image quality
and handwriting clarity. Despite these limitations, the system significantly
reduces manual effort, minimizes errors, and improves overall efficiency in
academic data management.
1.2 ARTIFICIAL INTELLIGENCE IN DOCUMENT PROCESSING
Artificial Intelligence (AI) plays a major role in modern document processing
systems by enabling machines to understand and extract information from
unstructured data such as images and handwritten text. It uses techniques like
Optical Character Recognition (OCR) and deep learning to identify patterns and
convert visual data into meaningful digital information. AI-based systems are
widely used in applications such as form processing, data entry automation, and
document analysis. In this project, AI helps in accurately extracting student
details and marks from scanned answer sheets. Although the performance
depends on image quality and handwriting clarity, AI significantly improves
efficiency, reduces manual effort, and enhances accuracy in data processing
14
tasks.
1.3 HANDWRITTEN TEXT RECOGNITION TECHNIQUES
Handwritten Text Recognition (HTR) is an important technique used to
identify and convert handwritten content into digital text. It uses machine
learning and deep learning models to recognize different writing styles and
patterns in scanned images. This technology is widely applied in areas such as
document digitization, form processing, and academic data extraction. In this
project, HTR is used to accurately read student names, register numbers, and
marks from answer sheets. Although variations in handwriting may affect
accuracy, modern recognition techniques provide reliable results and
significantly improve the efficiency of data extraction systems.
1.4 IMAGE PROCESSING IN EDUCATIONAL SYSTEMS
Image processing is widely used in educational systems to analyze and extract
information from scanned documents and images. It involves techniques such as
image enhancement, noise reduction, and segmentation to improve the quality of
input data before processing. These methods help in accurately identifying text
and numerical values from answer sheets. In this project, image processing plays
a key role in preparing scanned answer sheets for AI-based extraction. Although
poor image quality can affect performance, proper processing techniques
improve accuracy and make the system more reliable for academic data
management.
1.5 NEED FOR AUTOMATION IN RESULT PROCESSING
With the increasing number of students in educational institutions, manual result
processing has become time-consuming and prone to errors. Traditional methods
15
require teachers to enter marks manually, which can lead to inaccuracies and
delays in publishing results. Automation helps in reducing human effort and
ensures faster and more reliable data processing. In this project, automation is
used to extract and process marks directly from scanned answer sheets. Although
initial setup is required, automated systems greatly improve efficiency, accuracy,
and overall productivity in academic result management.
1.6 OBJECTIVES OF THE PROJECT
The main objective of the project is to develop an automated system for
extracting student marks from scanned answer sheets and converting them into a
structured digital format. The system aims to reduce manual effort, improve
accuracy, and speed up the result processing process. It also focuses on
generating a well-organized Excel report with calculated totals and summary
details. Although challenges like handwriting variation exist, the system provides
an efficient and reliable solution for academic data management.
1.7 SCOPE OF THE PROJECT
The scope of this project includes the extraction of student details and
marks from scanned or photographed answer sheets using artificial intelligence.
It covers image processing, data extraction, calculation of totals, and generation
of structured Excel reports. The system is mainly intended for use in schools and
colleges to simplify result processing. However, its performance depends on
image quality and handwriting clarity, and it can be further extended with
advanced features in the future.
16
CHAPTER 2
LITERATURE REVIEW
2.1 OCR TECHNIQUES FOR HANDWRITTEN DOCUMENTS
Smith et al. present Optical Character Recognition (OCR) techniques for
extracting handwritten text from scanned documents. The system uses pattern
recognition and machine learning methods to identify characters and convert
them into digital format. The study shows that OCR significantly reduces manual
data entry and improves processing efficiency in document analysis. However,
accuracy may vary depending on handwriting style and image quality. This work
highlights the importance of OCR technology, which forms the foundation for
extracting student marks in our proposed system. [1]
2.2 AI BASED DOCUMENT PROCESSING SYSTEM
Brown et al. present AI-based document processing systems that utilize
machine learning and deep learning techniques to extract and organize
information from unstructured data sources. The system focuses on improving
accuracy and efficiency in handling large volumes of documents by automating
data extraction and classification tasks. The study shows that AI significantly
enhances processing speed and reduces human intervention in document
management systems. However, the implementation requires proper training data
and computational resources for optimal performance. This work highlights the
role of AI in automating data extraction, which is essential for developing the
proposed digital mark extraction system. [2]
17
2.3 NEURAL NETWORKS FOR CHARACTER RECOGNITION
Zhang et al. present a handwritten text recognition system based on neural
network models for accurately identifying handwritten characters from scanned
documents. The system uses deep learning techniques to learn different writing
styles and improve recognition performance. The study shows that neural
networks significantly enhance accuracy compared to traditional methods,
especially in complex handwriting scenarios. However, the model requires large
datasets and training time to achieve high performance. This work highlights the
effectiveness of neural network-based recognition, which is essential for
extracting handwritten marks in the proposed system. [3]
2.4 AUTOMATED RESULT PROCESSING SYTEM
Kaur et al. present automated result processing systems designed to reduce
manual effort in handling academic data. The system focuses on collecting,
processing, and generating results using digital tools to improve accuracy and
efficiency. The study shows that automation minimizes human errors and speeds
up the overall result preparation process. However, the system depends on proper
data input and system configuration for reliable performance. This work
highlights the importance of automation in academic environments, which is a
key concept applied in the proposed digital mark extraction system. [4]
2.5 IMAGE TO DATA CONVERSION USING ARTIFICIAL
INTELLIGENCE
Patel et al. present an image-to-data conversion system that utilizes artificial
intelligence techniques to extract meaningful information from images. The
system focuses on identifying text and numerical data from visual inputs and
18
converting them into structured digital formats. The study shows that AI-based
approaches improve accuracy and reduce manual effort in data extraction tasks.
However, the performance may be affected by image quality and noise in the
input data. This work highlights the importance of image-based data extraction,
which is directly applied in the proposed digital mark extraction system. [5]
2.6 DEEP LEARNING FOR TEXT RECOGNITION
Kumar et al. present deep learning approaches for recognizing handwritten
text using advanced neural network models. The system focuses on improving
recognition accuracy by learning complex patterns from large datasets. The study
shows that deep learning techniques outperform traditional methods in handling
variations in handwriting. However, the system requires high computational
power and training data. This work highlights the importance of deep learning in
improving extraction accuracy in the proposed system. [6]
2.7 MACHINE LEARNING FOR PATTERN RECOGNITION
Sharma et al. present machine learning techniques used for identifying
patterns in textual and numerical data. The system applies classification and
recognition algorithms to improve data extraction processes. The study shows
that machine learning enhances the adaptability of systems to different input
variations. However, model performance depends on proper training and feature
selection. This work highlights the role of pattern recognition in extracting
structured data from answer sheets. [7]
2.8 CLOUD BASED DOCUMENT PROCESSING SYSTEMS
Verma et al. present cloud-based document processing systems that enable
efficient storage and processing of large volumes of data. The system uses cloud
19
infrastructure to perform scalable and fast data processing tasks. The study
shows that cloud integration improves accessibility and system performance.
However, it depends on internet connectivity and data security measures. This
work highlights the advantage of scalable processing, which can enhance the
proposed system in future implementations. [8]
2.9 INTELLIGENT AUTOMATION IN EDUCATION SYSTEMS
Das et al. present intelligent automation systems applied in educational
environments to streamline academic processes. The system focuses on
automating tasks such as data entry, evaluation, and report generation. The study
shows that automation improves efficiency and reduces human errors in
educational institutions. However, system reliability depends on proper
implementation and data accuracy. This work highlights the importance of
automation, which is a core concept of the proposed system. [9]
2.10 WEB BASED FRAMEWORKS FOR DATA PROCESSING
The Python Software Foundation presents Flask, a lightweight web framework
used for building backend systems for data processing applications. The
framework allows easy handling of requests, data processing, and integration
with external services. The study shows that Flask enables rapid development
and efficient system design. However, it requires proper configuration and
security handling for large-scale applications. This work highlights the
importance of backend frameworks in implementing the proposed digital mark
extraction system. [10]
2.11 DATA ANALYSIS AND PROCESSING USING PANDAS
McKinney presents data analysis techniques using the Pandas library for
handling structured data efficiently. The system focuses on data manipulation,
20
filtering, and computation of results in tabular format. The study shows that
Pandas simplifies complex data processing tasks and improves efficiency in
handling large datasets. However, it requires proper understanding of data
structures for effective use. This work highlights the importance of data
processing tools, which are used in the proposed system for calculating marks
and generating results. [11]
2.12 EXCEL AUTOMATION USING OPENPYXL
OpenPyXL documentation presents methods for creating and managing Excel
files using Python. The system allows users to write, modify, and format Excel
sheets programmatically. The study shows that automated Excel generation
reduces manual effort and ensures structured output formatting. However,
handling large datasets may affect performance if not optimized properly. This
work highlights the role of Excel automation, which is essential for generating
the final output in the proposed system. [12]
21
CHAPTER 3
EXISTING SYSTEM
3.1 MANUAL MARK ENTRY SYSTEM
The manual mark entry system is the most commonly used traditional method
in educational institutions for recording and maintaining student marks. In this
approach, teachers manually read the marks from answer sheets and enter them
into physical registers or digital spreadsheets. This process requires continuous
human effort and careful attention, especially when handling large numbers of
students. Even a small mistake during entry can lead to incorrect results, which
may affect student evaluation. Moreover, verification and correction of such
errors consume additional time. Although this method is simple and does not
require advanced technology, it is inefficient and not suitable for modern large-
scale academic environments.
3.2 SPREADSHEET BASED RESULT PROCESSING
Spreadsheet-based result processing systems use tools like Microsoft Excel to
store, organize, and calculate student marks. In this system, teachers manually
input marks into predefined tables, and formulas are used to compute totals,
averages, and grades. While spreadsheets improve data organization and reduce
calculation errors, the data entry process remains manual and time-consuming.
Additionally, improper handling of formulas or accidental changes can lead to
incorrect results. Although spreadsheets provide better flexibility compared to
manual registers, they still lack automation in extracting data from answer
sheets, limiting overall efficiency.
22
3.3 BASIC OCR TOOLS
Basic OCR (Optical Character Recognition) tools are used to convert printed
text from scanned documents into digital format. These tools are effective for
typed or printed content but have limitations when dealing with handwritten text.
Since answer sheets usually contain handwritten marks and student details, basic
OCR tools often produce inaccurate results. They may fail to recognize
characters correctly due to variations in handwriting styles, image quality, and
formatting. As a result, manual correction is still required, reducing the
usefulness of such tools in academic result processing systems.
3.4 LIMITATION OF EXISTING SYSTEMS
The existing system for result processing has several limitations mainly due to
its dependence on manual and semi-automated methods. Manual data entry
increases the chances of human errors such as incorrect mark entry,
miscalculations, and duplication of data. The process is also highly time-
consuming, especially when handling a large number of student records.
Although basic OCR tools are available, they are not efficient in recognizing
handwritten text accurately, which further reduces system reliability. As a result,
additional manual verification is required, increasing workload.
In addition, the lack of automation makes the system inefficient and difficult
to scale for larger datasets. Data inconsistency is another major issue, caused by
variations in data entry formats and mistakes during processing. These systems
do not support real-time processing or intelligent data analysis, leading to delays
in result publication. Overall, the existing system lacks accuracy, speed, and
efficiency, highlighting the need for a more advanced automated solution.
23
3.5 ERROR PRONE DATA ENTRY METHODS
Manual data entry methods are highly vulnerable to human errors such as
incorrect reading of marks, typing mistakes, and calculation errors. These errors
can lead to inaccurate results, which may affect student records and evaluation.
Detecting and correcting these mistakes requires additional time and effort. In
large datasets, even a small percentage of errors can significantly impact overall
accuracy. Therefore, reducing human involvement in data entry is essential to
improve reliability.
3.6 TIME CONSUMPTION IN RESULT PROCESSING
The process of result preparation in existing systems is time-consuming due to
multiple manual steps involved. Teachers must read each answer sheet, enter
marks, verify data, and calculate totals. When the number of students increases,
the time required for processing also increases significantly. This delay affects
timely result publication and increases workload. Efficient systems are needed to
reduce processing time and improve productivity.
3.7 LACK OF AUTOMATION
The existing result processing system suffers from a significant lack of
automation, as most of the operations are carried out manually. Tasks such as
reading marks from answer sheets, entering data into records, and calculating
totals require continuous human involvement. This not only slows down the
overall process but also increases the chances of errors due to fatigue or
oversight. In the absence of automation, the system becomes inefficient when
handling a large number of students.
24
Furthermore, the lack of automated tools limits the system’s ability to process
data quickly and accurately. It does not support intelligent data extraction or real-
time result generation, which are essential in modern educational environments.
As a result, delays in result publication and increased workload for staff are
common issues. This highlights the need for an automated system that can
improve speed, accuracy, and overall efficiency.
3.8 DATA INCONSISTENCY ISSUES
Data inconsistency is a common issue in existing result processing systems
due to manual data entry and lack of proper standardization. Errors such as
incorrect entries, duplication, and variation in data formats often lead to
unreliable and mismatched results. These inconsistencies make it difficult to
maintain accurate records and require additional time for verification and
correction. As the amount of data increases, managing consistency becomes
more challenging.
3.9 EXISTING WORKFLOW ANALYSIS
The existing workflow for result processing involves several manual steps
such as collecting answer sheets, reading marks, entering data, and calculating
results. Each step depends heavily on human effort, making the process slow and
inefficient, especially when handling a large number of students.
Additionally, the workflow lacks proper integration and automation, which
increases the chances of errors and delays in result publication. There is no real-
time processing or validation mechanism, making the system less reliable. This
highlights the need for a more streamlined and automated workflow to improve
efficiency and accuracy.
25
CHAPTER 4
PROBLEM STATEMENT
In many educational institutions, the process of extracting and recording
student marks from answer sheets is still carried out manually. Teachers are
required to read marks from each answer sheet and enter them into registers or
spreadsheets, which is both time-consuming and labor intensive. This manual
approach increases the chances of human errors such as incorrect data entry,
miscalculations, and duplication, which can affect the accuracy of student
results.
Additionally, existing systems lack the ability to efficiently process
handwritten data from scanned answer sheets. Basic OCR tools are not fully
capable of accurately recognizing handwritten text due to variations in writing
styles and image quality. As a result, manual verification and correction are often
required, which further increases workload and delays in result processing.
Therefore, there is a need for an automated system that can accurately extract
student details and marks from scanned answer sheets and convert them into a
structured digital format. Such a system should minimize human intervention,
improve accuracy, reduce processing time, and provide a reliable solution for
efficient result management in modern educational environments.
Furthermore, the increasing demand for digital transformation in education
highlights the importance of adopting intelligent and automated systems.
Implementing such a solution can enhance productivity, ensure data consistency,
and support faster decision-making. This will not only improve the efficiency of
26
result processing but also contribute to the overall modernization of academic
management systems.
CHAPTER 5
PROPOSED SYSTEM
5.1 SYSTEM ARCHITECTURE
The system architecture represents the complete workflow of the Digital Mark
Extraction System, illustrating how data flows from input to output through
various interconnected modules. The process begins with uploading scanned
answer sheets and continues through preprocessing, AI-based extraction, data
structuring, and final output generation. Each module is designed to perform a
specific function while maintaining smooth communication with other
components.
The architecture ensures automation, accuracy, and efficiency by minimizing
manual intervention. It integrates frontend, backend, and AI components into a
unified system, enabling real-time processing and structured output generation.
This modular design also allows easy scalability and future enhancements such
as improved recognition models or cloud integration.
27
Fig 5.1.1 Overall System Architecture of Digital Mark Extraction System
5.2 IMAGE UPLOAD MODULE
The Image Upload Module serves as the entry point of the system, where
users provide scanned answer sheets in image format. It supports common file
formats such as JPG, JPEG, and PNG, ensuring compatibility with different
devices like scanners and mobile cameras. The module validates the uploaded
file to ensure proper format and quality before sending it to the backend.
This module is designed with a user-friendly interface, allowing easy
interaction without requiring technical knowledge. It ensures secure and efficient
transfer of images to the backend server. Proper handling at this stage is
important because the quality of input directly affects the accuracy of data
extraction in later stages.
28
Fig 5.2.1 Image upload Interface for Answer Sheets
5.3 FLASK BACKEND PROCESSING
The Flask backend acts as the central controller of the system, managing all
operations and communication between modules. It receives the uploaded image,
processes requests, and coordinates data flow between the AI model, parsing
module, and output generator. Flask provides a lightweight and flexible
framework for handling HTTP requests and responses efficiently.
This module also ensures proper error handling, data validation, and process
management. It plays a crucial role in maintaining system stability and
performance. By acting as a bridge between frontend and backend components,
it enables seamless integration of different technologies used in the system.
5.4 AI BASED EXTRACTION
The AI-based extraction module is the core component responsible for
identifying and extracting relevant information from the uploaded answer sheets.
It uses advanced machine learning techniques to recognize handwritten and
printed text, including student details and marks. The Gemini model processes
the image and converts unstructured visual data into structured textual output.
29
This module significantly reduces manual effort and improves accuracy in
data extraction. However, its performance depends on factors such as image
clarity and handwriting quality. Despite these challenges, AI-based extraction
provides a powerful solution for automating complex recognition tasks in
educational systems.
5.5 JSON PARSING MODULE
After AI extraction, the output is generated in JSON format, which contains
structured data elements. The JSON Parsing Module processes this data and
organizes it into clearly defined fields such as student name, register number,
subject marks, and total. This step ensures that the extracted data is properly
formatted and ready for further processing.
The module also handles missing or inconsistent values and ensures data
integrity. Proper parsing is essential for maintaining accuracy and consistency in
the system. It acts as a bridge between raw AI output and structured data
processing.
5.6 DATA PROCESSING AND COMPUTATION
This module performs all necessary calculations on the extracted data,
including total marks, averages, and result status. It ensures that all computations
are carried out accurately using predefined logic. The processed data is then
organized into a format suitable for output generation.
In addition to calculations, this module may also perform validation checks to
ensure correctness of data. It plays an important role in transforming extracted
information into meaningful academic results. Efficient data processing
improves overall system reliability and performance.
30
5.7 EXCEL GENERATION MODULE
The Excel Generation Module converts the processed data into a structured
spreadsheet format. It creates rows and columns with proper headings such as
student name, marks, total, and result. The module ensures that the output file is
well-organized, readable, and suitable for official use.
It also supports formatting features like alignment, borders, and highlighting
important fields. This makes the generated report professional and easy to
analyze. Automating Excel generation eliminates manual work and ensures
consistency in result presentation.
5.8 DOWNLOAD AND OUTPUT INTERFACE
The final module provides users with access to the generated Excel file
through a simple download interface. Once processing is complete, users can
easily download the result file with a single click. This ensures convenience and
quick access to processed data.
The interface is designed to be simple and user-friendly, allowing users to
retrieve results without any complexity. It completes the system workflow by
delivering the final output in a usable format. This module enhances overall user
experience and system usability.
31
Fig 5.8.1 Excel Report Download Interface
CHAPTER 6
RESULTS AND DISCUSSION
6.1 OUTPUT ANALYSIS
The output analysis focuses on evaluating the final Excel file generated by
the system. The extracted data such as student name, register number, subject-
wise marks, and total marks are organized in a clear tabular format. The system
ensures that all extracted values are properly aligned and easy to read, making it
suitable for academic record keeping.
32
Fig 6.1.1 Generated Excel Output of Extracted Marks
The results show that the system successfully converts unstructured image
data into structured digital format with minimal manual effort. The generated
Excel sheet reduces the workload of teachers and improves efficiency in result
preparation.
6.2 ACCURACY EVALUATION
Accuracy evaluation measures how correctly the system extracts marks and
student details from scanned answer sheets. The extracted data is compared with
actual values to determine correctness. The system performs well when the
image quality is clear and handwriting is legible, achieving high accuracy in
most cases.
However, minor errors may occur due to unclear handwriting or poor image
quality. Despite this, the system significantly reduces manual errors compared to
33
traditional methods. Overall, the accuracy level is satisfactory for practical use in
educational environments.
6.3 PERFORMANCE ANALYSIS
Performance analysis evaluates the efficiency of the system in terms of
processing time and speed. The system processes each uploaded image within a
short duration, making it suitable for handling multiple answer sheets. The use of
Flask backend and AI model ensures smooth and fast data processing.
The system also shows good scalability, as it can handle increasing data
volume without significant delay. Compared to manual methods, the processing
time is greatly reduced, improving overall productivity. This makes the system
efficient and practical for real-time academic applications.
CHAPTER 7
CONCLUSION AND FUTURE WORK
7.1 CONCLUSION
The Digital Mark Extraction System from Scanned Answer Sheets
successfully automates the process of extracting and processing student marks.
The system reduces manual effort by using AI-based extraction to convert
unstructured image data into structured digital format. It improves accuracy,
34
minimizes human errors, and significantly reduces the time required for result
processing.
The integration of modules such as image upload, Flask backend, AI
extraction, and Excel generation ensures smooth workflow and efficient
performance. Overall, the system provides a reliable and practical solution for
modern educational institutions to manage results effectively.
7.2 FUTURE WORK
The system can be further enhanced by improving the accuracy of handwritten
text recognition using advanced AI models. Future improvements may include
support for multiple answer sheet formats and better handling of low-quality
images. Integration with cloud platforms can also be implemented for scalable
and remote access.
Additionally, features such as automatic grading, analytics dashboards, and
database storage can be added to extend the system’s functionality. These
enhancements will make the system more powerful, flexible, and suitable for
large-scale academic applications.
REFERENCES
[1] J. Smith, “Optical character recognition for handwritten documents,” Int. J. Comput. Appl.,
vol. 176, no. 25, pp. 15–20, 2020.
[2] L. Brown, “AI-based document processing systems,” IEEE Trans. Pattern Anal. Mach.
Intell., vol. 43, no. 7, pp. 2201–2215, Jul. 2021.
35
[3] X. Zhang, “Handwritten text recognition using neural networks,” Int. J. Pattern Recognit.
Artif. Intell., vol. 36, no. 4, pp. 2250012-1–2250012-15, 2022.
[4] H. Kaur, “Automated result processing systems in education,” Int. J. Adv. Res. Comput.
Sci., vol. 12, no. 3, pp. 45–50, May 2021.
[5] R. Patel, “Image-to-data conversion using AI techniques,” in Proc. IEEE Int. Conf. Artif.
Intell. Syst., 2022, pp. 120–125.
[6] A. Kumar and S. Singh, “Deep learning approaches for handwritten text recognition,” J.
Artif. Intell. Res., vol. 68, pp. 1–20, 2022.
[7] P. Sharma, “Machine learning techniques for pattern recognition,” IEEE Access, vol. 9, pp.
12345–12360, 2021.
[8] S. Verma and R. Gupta, “Cloud-based document processing systems,” Int. J. Cloud
Comput., vol. 11, no. 2, pp. 85–95, 2023.
[9] N. Das, “Intelligent automation in education systems,” IEEE Trans. Learn. Technol., vol.
18, no. 1, pp. 10–20, Jan. 2025.
[10] Python Software Foundation, “Flask web framework documentation,” 2023. [Online].
Available: [Link]
[11] W. McKinney, Python for Data Analysis, 2nd ed. Sebastopol, CA, USA: O’Reilly Media,
2020.
[12] OpenPyXL, “Working with Excel files in Python,” 2023. [Online].
Available: [Link]
[13] Mozilla Developer Network, “Web technologies documentation,” 2023. [Online].
Available: [Link]
[14] React Documentation Team, “[Link] official documentation,” Meta Platforms, 2024.
[Online]. Available: [Link]
36
[15] Google, “Gemini API documentation,” Google LLC, 2024. [Online].
Available: [Link]
37