0% found this document useful (0 votes)
14 views2 pages

Enhancing Summarization Accuracy in ILSUM

The research presented by Rishika Jha and her team focuses on enhancing summarization accuracy for Indian languages, particularly Hindi and Gujarati, through the ILSUM shared task. They employed the T5-small model for generating summaries and classified factual inaccuracies using machine learning classifiers, achieving notable rankings in evaluation metrics. The study emphasizes the need for addressing factual errors in machine-generated summaries and outlines plans to explore advanced models for future improvements.

Uploaded by

rishikajhashiny
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
14 views2 pages

Enhancing Summarization Accuracy in ILSUM

The research presented by Rishika Jha and her team focuses on enhancing summarization accuracy for Indian languages, particularly Hindi and Gujarati, through the ILSUM shared task. They employed the T5-small model for generating summaries and classified factual inaccuracies using machine learning classifiers, achieving notable rankings in evaluation metrics. The study emphasizes the need for addressing factual errors in machine-generated summaries and outlines plans to explore advanced models for future improvements.

Uploaded by

rishikajhashiny
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Good morning, everyone.

I am Rishika jha on behalf of my team and co-authors


rakshit ahuja, sainik kumar mahata ,monalisa dey and dipankar das . Today, I
will be presenting our research on enhancing accuracy in summarization for
Indian languages. As we know, text summarization is a vital area in Natural
Language Processing, and with India's rich linguistic diversity, developing
effective summarization systems is crucial for improving knowledge
accessibility.
We participated in the ILSUM shared task, which focuses on summarization
across multiple languages, including English, Hindi, Gujarati, and Bengali as
well as Our goal was to improve the reliability of these summaries by
identifying and classifying factual errors., particularly in Hindi and Gujarati
language.
we utilized the ILSUM 2024 dataset. In Task 1, we worked with an English
dataset containing 9,500 training articles and 2,500 test articles. In Task 2, we
focused on Hindi and Gujarati language, analyzing the test articles to evaluate
the factual inaccuracies of machine-generated summaries. This diverse dataset
allowed us to train and test our models effectively.
For the Methodology , We approached our research in two main tasks:
1. Task 1: We employed the T5-small model, a pre-trained sequence-to-
sequence model, for generating summaries. We used the T5 tokenizer
for text processing and developed a summarization function that
adheres to specific length constraints.
2. Task 2: We classified factual inaccuracies in machine-generated
summaries into four categories: misrepresentation, inaccurate
quantities, false attribution, and fabrication. We utilized various machine
learning classifiers, including Logistic Regression and Random Forest, to
detect these errors.
Slide 5: Results In our evaluation of machine-generated summaries using
ROUGE metrics, we ranked 4th in the English language category. Here are the
scores:
We were also assessed using BERT scores, which measure semantic similarity.
Our results were as follows:
We ranked 7th in this category,.
Task 2: For the task of detecting inaccuracies in machine-generated
summaries, we achieved
 Gujarati F1 Score: 0.2127 (Rank 4)
 Hindi F1 Score: 0.2132 (Rank 6)
While these scores reflect the challenges in detecting factual inaccuracies, they
also highlight areas for future improvement.

Slide 6: Conclusion and Future Work In conclusion, our research highlights the
importance of addressing factual inaccuracies in machine-generated
summaries. We successfully classified multi-label errors and improved
summarization accuracy. Moving forward, we plan to explore advanced
models, such as GPT-4, and techniques like few-shot learning to further
enhance the reliability of machine-generated summaries.
Thank you for your attention. I am happy to take any questions you may have!

Common questions

Powered by AI

The key findings in terms of ranking and evaluation metrics showed that the researchers' summarization models ranked 4th in the English language category according to ROUGE metrics. This ranking was indicative of how well the summaries approximated reference texts in terms of frequent text segments. Additionally, using BERT scores for measuring semantic similarity, they ranked 7th. These evaluations highlighted the models’ relative strengths and areas for improvement in generating and assessing semantic content accuracy in machine-generated summaries .

In Task 1, the research focused on generating accurate summaries using an English dataset. They employed the T5-small model for pre-trained sequence-to-sequence summarization. The goal was to enhance the summarization quality compared to existing methods. Task 2 shifted focus to detecting factual inaccuracies in summaries specifically for Hindi and Gujarati. The research aimed to classify inaccuracies into categories such as misrepresentation and false attribution and utilized machine learning classifiers for detection. Each task had a specific goal: Task 1 aimed at summary generation, while Task 2 focused on error detection .

The authors considered employing techniques like few-shot learning in future summarization projects to address the limitations associated with training data scarcity, especially for less-resourced languages. Few-shot learning allows models to leverage limited examples to learn and apply generalization capabilities, which could significantly enhance the performance and adaptability of summarization models in linguistically diverse settings.<br>This approach is anticipated to enable efficient model training, improve the handling of varied language nuances, and reduce the likelihood of errors due to limited data .

The ILSUM 2024 dataset was crucial to the research as it provided a diverse linguistic base, essential for training and evaluating summarization models across multiple languages, including English, Hindi, Gujarati, and Bengali. This diversity allowed the researchers to develop and test models that could handle the complexities and nuances of different languages, ensuring broader applicability and improved accuracy of summarization systems. It facilitated a more comprehensive assessment of the models' capability to handle diverse linguistic features and errors, highlighting areas needing improvement .

The choice of the T5-small model impacted summarization outcomes by providing a pre-trained sequence-to-sequence architecture that aids in generating accurate summaries within specified length constraints. The use of the ROUGE metrics for evaluation allowed the researchers to quantitatively assess the quality of the summaries by comparing them to reference texts, where they ranked 4th in the English language category. This combination helped to ensure that summaries not only adhered to structural constraints but also maintained a level of accuracy in conveying the content .

The use of machine learning classifiers, including Logistic Regression and Random Forest, was key in detecting factual inaccuracies within machine-generated summaries. By classifying errors into categories like misrepresentation and fabrication, these classifiers allowed the researchers to systematically evaluate the presence of inaccuracies in the summaries. The effectiveness of these classifiers was quantitatively assessed using F1 Scores, where they achieved a score of 0.2127 for Gujarati and 0.2132 for Hindi, ranking them 4th and 6th respectively in this task .

The researchers faced the challenge of improving summarization accuracy across multiple Indian languages, which is difficult due to the linguistic diversity in India. Their methodology addressed these challenges by utilizing the ILSUM 2024 dataset to provide training and evaluation data for multiple languages like English, Hindi, and Gujarati. They employed a T5-small model to generate summaries and developed specific methodologies for assessing factual inaccuracies in machine-generated summaries by classifying errors into categories such as misrepresentation and false attribution. They also utilized machine learning classifiers like Logistic Regression and Random Forest to detect inaccuracies .

Addressing factual inaccuracies is critical in summarization due to the significant linguistic diversity in India, where each language presents unique semantic and syntactic challenges. Factual inaccuracies can mislead readers and diminish trust in automated summarization systems. Accurately capturing the essence and facts of the source material is essential for maintaining information credibility and reliability across different languages, particularly in a multilingual context where each error may propagate distinct misunderstandings. Properly addressing these inaccuracies ensures that summarization systems are useful and trustworthy for varied linguistic audiences .

Classifying factual inaccuracies into distinct categories such as misrepresentation, inaccurate quantities, false attribution, and fabrication allowed for a systematic evaluation and targeted improvement of machine-generated summaries. By identifying specific types of errors, the researchers could address them more effectively with tailored methodologies using machine learning classifiers. This approach potentially enhances the reliability of summaries by ensuring that different error types are detected and corrected, leading to more trustworthy representations of the original texts .

The authors identified potential future directions such as exploring advanced models like GPT-4 and techniques like few-shot learning. These approaches aim to improve the reliability and accuracy of machine-generated summaries by leveraging more sophisticated algorithms and learning paradigms that require fewer examples to understand patterns. Such advancements may address current limitations related to linguistic diversity and error detection by enhancing the models' ability to generalize across different languages and accurately classify multiple types of factual inaccuracies .

You might also like