0% found this document useful (0 votes)
24 views5 pages

Bangla Dialect Classification with ML

This conference paper discusses the classification of Bangla language dialects, specifically Chatgaiya and Pabna, using machine learning techniques. The authors collected a dataset of approximately 5,000 annotated texts from local sources and applied various machine learning algorithms, achieving a maximum accuracy of 96%. The paper outlines the methodology, including exploratory data analysis and feature extraction, and presents performance evaluations of different models used in the classification task.

Uploaded by

Arup Barua
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
24 views5 pages

Bangla Dialect Classification with ML

This conference paper discusses the classification of Bangla language dialects, specifically Chatgaiya and Pabna, using machine learning techniques. The authors collected a dataset of approximately 5,000 annotated texts from local sources and applied various machine learning algorithms, achieving a maximum accuracy of 96%. The paper outlines the methodology, including exploratory data analysis and feature extraction, and presents performance evaluations of different models used in the classification task.

Uploaded by

Arup Barua
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

See discussions, stats, and author profiles for this publication at: [Link]

net/publication/370613797

Bangla Language Dialect Classification using Machine Learning

Conference Paper · December 2022


DOI: 10.1109/ICECTE57896.2022.10114552

CITATIONS READS

2 369

4 authors, including:

Md Raihanul Islam Tomal Abdul Kadar Muhammad Masum


Universiti Malaysia Pahang Al-Sultan Abdullah Daffodil International University
2 PUBLICATIONS 3 CITATIONS 84 PUBLICATIONS 1,445 CITATIONS

SEE PROFILE SEE PROFILE

All content following this page was uploaded by Md Raihanul Islam Tomal on 28 October 2024.

The user has requested enhancement of the downloaded file.


2022 4th International Conference on Electrical, Computer & Telecommunication Engineering (ICECTE)
29-31 December 2022, Rajshahi-6204, Bangladesh

Bangla Language Dialect


Classification using Machine Learning
Md Raihanul Islam Tomal Tanveer Kader Md. Kalim Amzad Chy
Department of Computer Science and Department of Computer Science and Department of Computer Science and
Engineering Engineering Engineering
International Islamic University International Islamic University International Islamic University
Chittagong, Bangladesh Chittagong, Bangladesh Chittagong, Bangladesh
raihanultomal@[Link] tanveerkaderedu@[Link] [Link]@[Link]

Abdul Kadar Muhammad Masum


Department of Computer Science and
Engineering
International Islamic University
Chittagong, Bangladesh
akmmasum@[Link]

Abstract—Dialect classification of a language is a complex and take appropriate action against them. Rescue operations
work as it is the variation of the same language. This paper like firefighters, police, ambulances, etc. often struggle to
classifies dialect based on local Bengali text. The classification identify emergencies.
becomes harder when it is about a language that is not very
much available in written format or stored in any other way This study is concerned with classifying the dialects from
except spoken among local people. For natural language texts. To train and test the models in this, a significant
processing (NLP) a good amount of data is essential to get the amount of data is needed. About 5,000 data including local
job done. It focuses on generating an enriched dataset of the languages from Chattogram and Pabna districts in
local Bangla language. The dataset introduces two popular Bangladesh, were collected by researchers. These dialects are
dialects which are Chatgaiya and Pabna which are spoken by a referred to as Chatgaiya and Pabna. Researchers used a
large number of people. It comprises about 5000 data variety of social media platforms and online news portals to
regarding these local languages which are annotated with their develop the dataset. In addition to collecting the data, the
respective dialects. A five-step Exploratory Data Analysis annotation process is carried out. These two dialects have
(EDA) is carried out. Feature extraction is conducted using about 5,000 annotations. Once a five-step EDA has been
three different techniques like CountVectorizer, Term completed on the data feature extraction is carried out to
Frequency-Inverse Document Frequency (TF-IDF) and train the machine learning models. The results section
Word2vec. With this huge amount of data, it worked on
includes a performance analysis.
classifying Bangla language dialect using machine learning
algorithms such as Support Vector Machine (SVM), The rest of the paper is organized as follows. The
Multinomial Naïve Bayes (MNB), Logistic Regression (LR), literature review is included in Section 2, and the dataset
Random Forest (RF), Decision Tree (DT), K Nearest Neighbor development process is discussed in Section 3. Section 4
(KNN). This study obtained the highest 96% accuracy. describes the working process, which includes EDA, feature
extraction techniques, machine learning algorithms, and.
Index Terms—bangla dialect classificaiton, machine Section 5 analyzes the results, while Section 6 provides a
learning, countvectorizer, tf-idf, word2vec
conclusion and suggestions for future scopes.
I. INTRODUCTION II. LITERATURE REVIEW
Language is the most convenient element of Md. Ashraful Haider Chowdhury et al. [1] used a deep
communication. There are more than 6500 languages all over learning model to examine the features of Bengali sentences
the world. Bangla is the fifth most native language and which are compatibility, proximity and expectancy. This
seventh most spoken language in this world. Gradually the paper proposed a neural network model consisting of one hot
number of Bangla-speaking people is increasing. They are encoding, word embedding and Long Short-Term Memory
making contributions in a number of areas, including the (LSTM) network. They tested their model on 75000 Simple
economy and education. Bangladesh has recently gained Bengali sentences. Susmoy Chakraborty et al. [2] worked on
popularity as a tourist destination due to its stunning natural the readability of Bengali texts, classifying them as simple or
surroundings and distinctive culture. Most tourist attractions complex sentences using deep learning models. They
are located in rural and remote areas. People communicate in combined character length and consonant conjunct with
various dialects, which makes it challenging for outsiders to Bidirectional LSTM (BiLSTM) which provided significant
understand. Criminal actions including cyberbullying, results. Tanzina Akter Tani et al. [3] applied machine
threats, illegal trade, etc. are also carried out by locals from learning algorithms to detect hateful text form twitter.
various regions. Due to their poor grip on the local tongue,
law enforcement officials find it challenging to identify them

979-8-3503-2054-1/22/$31.00 ©2022 IEEE


A. Dataset Generation Related Researches
A dataset of 3000 Bengali texts was provided by Utpal
Rudra et al. [4] that contains data on 5 types of personality
traits such as conscientiousness, agreeableness, extroversion,
openness and neuroticism. They demonstrated that for this
multiclass classification task, deep learning-based
approaches outperformed other methods. Avishek Das et al.
[5] developed a dataset consisting of about 6243 Bangla text
corpora for their classification process. This paper introduced
a transformer-based technique to classify six basic emotions
which are anger, disgust, fear, joy, surprise and sadness.
Prashengit Dhar et al. [6] collected data from popular Fig. 1. Bivariate Analysis
Bengali news portals in three categories like health,
technology and sports. The dataset dimension is 100x3. Tree- Management is the most important matter for any task.
based pipeline optimization tool is used to improve the ml Corpus management is not the exceptional. Each local
classifiers. language has 25-30% pure language. That means we have
some common words in different types of dialects. That’s
B. Dialect Classification Related Researches why we tried to collect pure local sentences and paragraphs
There have been some studies on dialect classification in as much as possible. Figure 1 shows that the dataset has
various languages. Omar F. Zaidan et al. [7] introduced a more than 3000 Chatgaiya data and 2000 Pabna data. We
huge monolingual data set with plenty of dialectal Arabic have to maintain the ratio carefully because if one data is
content named as Arabic On-line Commentary (AOC) extremely available in the dataset and other is struggling then
Dataset. They annotated more than 100000 Arabic dialects our dataset will not be considered as a handsome dataset due
by crowdsourcing. Their classification approach provides to lack of maintaining a proper ratio. We also keep identical
near-human accuracy. Mohamed Elaraby et al. [8] conducted entries in our dataset. Though it’s a collection of more than
additional research on the AOC dataset by applying deep 5000 data but it contains unique data which improves its
learning models and benchmarking the dataset. Another level.
work on identifying Arabic dialect is conducted by
Mohammed Ali [9] who proposes three models with the IV. METHODOLOGY
same architecture with the exception of the first layer by one We have accounted in several steps to classify the corpus
hot encoding, embedding layer, and gated recurrent unit we have collected. At the very beginning we did the EDA to
(GRU) recurrent layer. The best result was obtained from the our dataset to make the dataset useable for our work There
third model with the convolutional neural network. The same are some loadings, cleaning, elimination steps which are
architecture was followed for German dialect identification nicely accomplished. After that our data is ready for further
where the third model performed best [10]. Chiyu Zhang et work.
al. [11] came up with two deep learning models GRU and
Bidirectional Encoder Representations from Transformers
(BERT) for the identification of the Arabic dialects based on
21 country-level dialects from the MADAR twitter corpus
where the BERT-based semi-supervised model produced the
best results.
III. DATA
We have collected data from local news portals, local
Bangla dramas and local content in social media. Table 1
shows the dialect difference in Bengali language.

TABLE I. REPRESENTATION OF DATA


Standard Bangla Chatgaiya Dialect Pabna Dialect
Language Fig. 2. Methodology
কমন আেছন কন আছন অনরা? িকরাম আেছন
আপনারা? আপেনরা? Now it is time to make the corpora understandable to our
algorithms as the algorithms cannot understand the raw text
A lot of dramas and video contents are available in Pabna without any processing. There are various ways to do this job
and Chittagong local language. We have collected those which are known as text feature extraction and word
videos and converted the speeches to our corpus. We embedding techniques. We have used TF-IDF,
collected data in sentences and paragraphs with their CountVectorizer for text feature extraction and Word2vec as
corresponding dialect. We have collected more than 5000 our word embedding technique.
entries of paragraphs and sentences manually. A. Exploratory Data Analysis (EDA)
EDA is an essential process for data scientists. It is the
analysis of the data that is held in several steps. This process
removes any kind of ambiguity such as garbage values, null
values or any other unnecessary features from the dataset.
C. Splitting
We split our dataset for training and testing. We used
70% of our data for training models and remaining 30% data
used for testing.
D. Applied Machine Learning Algorithms
We have applied several machine learning algorithms on
our corpus. We have used Multinomial Naïve Bayes,
Fig. 3. EDA Logistic Regression, Support Vector Machine, Random
Forest, Decision Tree and K Nearest Neighbor. A brief
Figure 3 shows all the steps of EDA. Dataset loading overview of their working principle is discussed in this
means to read the dataset which is created for this specific section.
task. Corrupted file type will send error. We can view the
features/columns/fields, as well as the data type and the The Naive Bayes classifier works on the principle of
number of nulls. It is a must to check the information of data conditional probability, as given by the Bayes theorem. SVM
to ensure about all the features are working correctly. There works by mapping data to a high-dimensional feature space
were columns that contains standard Bengali language, time so that data points can be categorized. A decision tree is a
and date of the data entry which are not necessary for the graphical representation of all possible solutions to a
classification task. So, these columns were removed. There decision based on certain conditions. On each step or node of
were a lot of garbage data or null values in the dataset All a decision tree, used for classification, we try to form a
existing null values were removed. Bivariate analysis is condition on the features to separate all the labels or classes
already shown in figure 1. in the dataset. Random Forest grows multiple decision trees
which are merged together for a more accurate prediction.
B. Text Feature Extraction Logistic regression is a statistical analysis method to predict
To process natural language text and extract usable a binary outcome, such as yes or no, based on prior
information from a given word or sentence using machine observations of a data set. KNN works by finding the
learning techniques, the text must be translated into a set of distances between a query and all the examples in the data,
real integers which is known as vectors. The process of selecting the specified number of examples (K) closest to the
converting words into vectors are called vectorization. This query, then voting for the most frequent label. A comparative
paper uses three most popular vectorization techniques study of the performance is shown in the result section.
which are CountVectorizer, TF-IDF, Word2vec. E. Model Evaluation
CountVectorizer creates a matrix with documents and Model assessment is the process that analyze the
token counts where each sentence is a document and words performance of a machine learning model. It uses several
in the sentence are tokens. It is concerned with the frequency evaluation criteria to find out a machine learning model’s
of vocabulary words in a given document. As a result, strength and weakness. Model efficiency determination and
articles, prepositions, and conjunctions which don’t monitoring a model are done with the use of model
contribute a lot to the meaning get as much importance as evaluation. We can use a variety of evaluation metrics to see
meaningful words. TF-IDF scales down the impact of tokens if the models are doing well with provided data. For
that occur very frequently in a given corpus and that are less evaluating classification performance some of the commonly
informative than features that only appear in a tiny portion of used matrices are accuracy, confusion matrix, recall, f1-
the training corpus. It measures the frequency of a word in a score, precision. All this evaluation matrices are done for
text against its overall frequency in the corpus. In our model every models. Only the best results evaluation is showed in
the training and testing data shapes are (3502, 2120) and the result section.
(1502, 2120) where the vector size is 2120 for both
CountVectorizer and TF-IDF. In above methods every word V. RESULT
was treated as an individual entity, and semantics were We have used three different kinds of models with
completely ignored. Word embeddings give us a way to use various machine learning algorithms for the dialect
an efficient, dense representation in which similar words classification task. Performance of each model is discussed
have a similar encoding. An embedding is a dense vector of in this section. Also attached various figures and output
floating point values. Instead of specifying the values for the results with each model.
embedding manually, they are trainable parameters.
Word2vector represent the semantic and syntactic properties TABLE II. DIALECT CLASSIFICATION ACCURACY
of words through word embeddings. Skip Gram (SG) and
Classifiers CountVectorizer (%) TF-IDF (%) Word2vec
Continuous Bag of Words (CBOW) are two main
(%)
implementations of Word2Vec. Both train a shallow neural
MNB 96 96 -
network to represent words as feature vectors of variable
length. These vectors are dense, meaning that they consist of LR 96* 95 60
mostly floating point values, rather than zeros. In our model
CBOW is used with a vector size of 100 and window size of SVM 94 96* 60
5. A distance metric, usually cosine similarity, is used to RF 94 95 89*
determine semantic textual similarity. Cosine similarity
measures the cosine angle between two vectors projected in DT 90 92 83
an n-dimensional vector space. KNN 89 91 87
From table II it is clear that Multinomial Naïve Bayes research. In future we will work on it and will try to achieve
and Logistic Regression results are very close. The Logistic better result.
regression performs best on CountVectorizer with 96%
accuracy. Multinomial Naïve Bayes and Support Vector References
Machine performed best. The Support Vector Machine [1] M. A. H. Chowdhury, N. Mumenin, M. Taus, and M. A. Yousuf,
performs slightly better according to macro avg score on TF- "Detection of Compatibility, Proximity and Expectancy of Bengali
IDF vectorizer with 96% accuracy. Random Forest performs Sentences using Long Short Term Memory," in 2021 2nd
International Conference on Robotics, Electrical and Signal
best on Word2Vec with 89% accuracy. Processing Techniques (ICREST), 2021: IEEE, pp. 233-237.
[2] S. Chakraborty, M. T. Nayeem, and W. U. Ahmad, "Simple or
TABLE III. RESULT OF BEST MODELS complex? learning to predict readability of bengali texts," in
Proceedings of the AAAI Conference on Artificial Intelligence, 2021,
Precision (%)

F1-Score (%)
vol. 35, no. 14, pp. 12621-12629.
Recall (%)

Classifier

Accuracy
Dialect

Model
[3] T. A. Tani, T. Islam, S. A. Newaz, and N. Sultana, "Systematic

(%)
Analysis of Hateful Text Detection Using Machine Learning
Classifiers," in 2021 13th International Conference on Information &
Communication Technology and System (ICTS), 2021: IEEE, pp.
330-335.
Chatgaiya 96 98 97 [4] U. Rudra, A. N. Chy, and M. H. Seddiqui, "Personality traits
TF-IDF SVM 96* detection in bangla: A benchmark dataset with comparative
Pabna 97 93 95 performance analysis of state-of-the-art methods," in 2020 23rd
International Conference on Computer and Information Technology
Chatgaiya 95 98 97 Count- (ICCIT), 2020: IEEE, pp. 1-6.
LR 96
Pabna 97 91 94 Vectorizer [5] A. Das, O. Sharif, M. M. Hoque, and I. H. Sarker, "Emotion
classification in a resource constrained language using transformer-
Chatgaiya 90 82 86 based approach," arXiv preprint arXiv:2104.08613, 2021.
Word2vec RF 89 [6] P. Dhar and M. Abedin, "Bengali News Headline Categorization
Pabna 89 94 91 Using Optimized Machine Learning Pipeline," International Journal
of Information Engineering & Electronic Business, vol. 13, no. 1,
2021.
Table III shows that SVM on TF-IDF performed best [7] O. F. Zaidan and C. Callison-Burch, "Arabic dialect identification,"
with 96% accuracy by analyzing precision, recall and f1 Computational Linguistics, vol. 40, no. 1, pp. 171-202, 2014.
score. [8] M. Elaraby and M. Abdul-Mageed, "Deep models for arabic dialect
identification on benchmarked data," in Proceedings of the Fifth
TABLE IV. CONFUSION MATRIX Workshop on NLP for Similar Languages, Varieties and Dialects
(VarDial 2018), 2018, pp. 263-274.
Predicted Level [9] M. Ali, "Character level convolutional neural network for Arabic
Chattogram Pabna dialect identification," in Proceedings of the Fifth Workshop on NLP
True Level Chattogram 909 17 for Similar Languages, Varieties and Dialects (VarDial 2018), 2018,
Pabna 39 537 pp. 122-127.
[10] M. Ali, "Character level convolutional neural network for German
dialect identification," in Proceedings of the Fifth Workshop on NLP
Table IV shows the confusion matrix of SVM where for Similar Languages, Varieties and Dialects (VarDial 2018), 2018,
maximum number of predictions are correct hence the score pp. 172-177.
is highest. [11] C. Zhang and M. Abdul-Mageed, "No army, no navy: Bert semi-
supervised learning of arabic dialects," in Proceedings of the Fourth
VI. FUTURE WORK AND CONCLUSION Arabic Natural Language Processing Workshop, 2019, pp. 279-284.
This thesis is about dialect classification on local
language and gender detection. We have classified local
language of two different area ‘Pabna’ and ‘Chittagong’. We
have got 96% accuracy in dialect classification. So, our main
findings to improve accuracy result and extend our dataset.
We mainly implement CountVectorizer, Tf-idf, and
Word2vec. So, we need to implement more models and
algorithm on our dataset such as LSTM, BiLSTM and more
deep learning and neural network-based algorithm. So, we
can consider it as our finding. In future we will try to
implement and recover those findings. We stated that there is
a vast area of research on Bangla local language. We just
introduce our local language and implement some basic
algorithm in this paper. But this field is so vast. In future we
can work on ‘natural language generation automatically’. We
also will work on ‘voice detection of a local people’ for
dialect classify.
We have applied several algorithms and we have
different results. We provide several plots and screen shots
of our work to understand our thesis result easily. We also
mentioned our plan for future work and finding of this

View publication stats

Common questions

Powered by AI

The dataset collection process ensures diversity and accuracy by including over 5000 data entries that feature sentences and paragraphs annotated with the respective dialects of Chatgaiya and Pabna. These data were manually collected from various sources like local news, dramas, and social media, which ensures a wide range of context and usage scenarios. Care was taken to maintain a balanced ratio between the dialects, with over 3000 entries for Chatgaiya and 2000 for Pabna, to enable effective classification without skewing towards one dialect . Unique data entries were emphasized to improve the dataset's level and diversity .

Word embeddings contributed to the dialect classification process by providing dense vector representations of words, capturing semantic and syntactic similarities. The study employed the Word2vec method, specifically using the Continuous Bag of Words (CBOW) approach, with a vector size of 100 and a window size of 5. This method allowed the models to learn a low-dimensional representation of words in the dataset, facilitating better pattern recognition and classification accuracy in processing dialectal text .

The authors proposed exploring future research opportunities in deep learning models and neural network-based algorithms to further improve dialect classification accuracy. They suggested implementing advanced models such as LSTM and BiLSTM and working on natural language generation automatically. Additionally, they aimed to expand the dataset and explore the voice detection of local people for dialect classification. These steps are intended to delve deeper into the vast area of Bangla local language research, which has been relatively underexplored .

Exploratory Data Analysis (EDA) plays a critical role in preparing data for dialect classification by identifying and rectifying issues such as missing values, redundant features, and inconsistencies. In this study, EDA involved loading, cleaning, and processing data to remove irrelevant columns and null values and to ensure all features worked correctly. This process helped in creating a clean and consistent dataset, essential for achieving high accuracy in machine learning models. EDA enabled the identification of meaningful patterns and relationships within the data, facilitating more effective feature extraction and model training .

The study employed three feature extraction techniques: CountVectorizer, TF-IDF, and Word2vec. CountVectorizer creates a matrix of token counts, focusing on word frequency, while TF-IDF scales down the influence of frequently occurring terms providing a more informative feature set. Word2vec provides dense vector representations of words, capturing semantic similarities. These techniques translate text into numerical vectors, enabling machine learning algorithms to process and classify dialect data effectively .

Using multiple feature extraction techniques, such as CountVectorizer, TF-IDF, and Word2vec, impacts dialect classification models by providing varied textual representations that cater to different aspects of feature analysis. CountVectorizer focuses on frequency, TF-IDF balances term frequency and document relevance, and Word2vec captures semantic similarities. This diversity in feature extraction allows models to leverage the strengths of each technique, leading to improved accuracy and robustness in dialect classification, as evidenced by the high performance of models like SVM with TF-IDF .

The key evaluation metrics used to assess the performance of dialect classification models included accuracy, precision, recall, F1-score, and confusion matrix. These metrics are important as they provide a comprehensive view of a model's effectiveness. Accuracy indicates overall correctness of classifications, precision measures the ratio of correctly predicted positive observations to total predicted positives, recall evaluates the model's ability to identify all relevant instances, and the F1-score balances precision and recall. The confusion matrix visually presents the performance of the classifier, highlighting correct and incorrect predictions .

Support Vector Machine (SVM) demonstrated comparative advantages by achieving the highest accuracy of 96% when combined with the TF-IDF feature extraction method. It outperformed other algorithms like Multinomial Naïve Bayes and Logistic Regression, which achieved similar accuracy but were more reliant on the feature extraction technique. SVM's ability to map data to a high-dimensional space allows it to find an optimal hyperplane for classification, which contributes to its superior performance in this context .

Challenges in Bangla dialect classification include the language's variance and lack of written documentation for dialects like Chatgaiya and Pabna. Overcoming these challenges involved creating a robust dataset from spoken sources like local dramas, news, and social media, which ensured a diverse representation of dialects. Exploratory Data Analysis (EDA) was used to eliminate noise and enhance data quality. Feature extraction methods, such as TF-IDF and Word2vec, captured meaningful linguistic patterns. Various machine learning algorithms were tested to maximize accuracy, with SVM showing a high performance of 96% accuracy .

The study utilized several machine learning algorithms, including Multinomial Naïve Bayes, Logistic Regression, Support Vector Machine, Random Forest, Decision Tree, and K Nearest Neighbor. The Support Vector Machine achieved the highest accuracy of 96% when used with the TF-IDF feature extraction method, making it the most effective algorithm for Bangla dialect classification in this study .

You might also like