Arabic Text Categorization Algorithm
Using Vector Space Model
Essam Hanandeh and Mohamed Shajahan
Abstract Text categorization is one of the most important jobs in knowledge
retrieval and data mining. This study seeks to investigate several vector space model
(VSMs) variants using the K-Nearest Neighbor algorithm (k-NN) method. 242
abstract Arabic documents that they used was utilized in this paper. The compar-
ison is often based on well-known text assessment metrics, recall calculation, preci-
sion measurement, and measurement F1. Cosine outperformed in tests conducted on
Saudi data sets, according to the results.
Keywords Arabic data sets · Data mining · Text categorization · Term
weighting · Vector space model
1 Introduction
One of the most important topics in knowledge retrieval and data mining is text cate-
gorization (TC) [1]. Natural language text is vital, so a significant amount of text is
stored online alongside easily accessible knowledge libraries and document collec-
tions. Furthermore, the relevance of TC grows as it relates to the encoding and classi-
fication of natural language text utilizing a variety of methods and approaches, which
facilitates retrieval and other text manipulation activities. On the basis of text simi-
larity, TC typically VSM, when combined with k-NN, produces the greatest results
in terms of F1, precision, and recall. So far, the author has not made comparisons with
The Saudi Newspapers (SNP) using VSM. The article is structured as follows: The
second part presents a synthesis of the literature on the subject. Section 3 contains
an outline of the TC problem. In sections four and five, conclusions and future
E. Hanandeh
Computer Science, Zarqa University, Zarqa, Jordan
e-mail: hanandeh@[Link]
M. Shajahan (B)
Software Engineering, University of Business and Technology, Jeddah, Saudi Arabia
e-mail: shajahan@[Link]
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2023 41
A. Hannoon and A. Mahmood (eds.), Artificial Intelligence, Internet of Things,
and Society 5.0, Studies in Computational Intelligence 1113,
[Link]
42 E. Hanandeh and M. Shajahan
research, the experiment results are ultimately examined and explained. The goal of
this paper is to demonstrate and contrast k-NN algorithm findings against Arabic text
sets. The Arabic data sets we take into consideration are subjected to three distinct
experimental runs of similarity computation. One of TC’s acknowledged ideas is
term weighting. It is a factor that is given to a term to reflect performs two steps:
categorization assignment and the significance of that phrase, according to one defi-
nition. In this study, we compare several VSM iterations to the KNN [2] approach,
which makes use of inverse document frequency (IDF). There are three different
word weighting techniques: IDF Weighted (WIDF), Inverse Term Frequency (ITF),
and ITF [3]. The researcher’s comparison of the various KNN implementations is
based on the F1, Recall, and Precision measurements. To put it another way, we’re
trying to decide which the KNN Algorithm employing three different VSMs.
2 Related Works
Along with others, [4] offers Superior Multi-Kernel CNN Model for Arabic Text
Categorization that has been improved with n-gram word embedding for classifying
Arabic news documents (SATCDM). When compared to current work on Arabic
text categorization, the proposed technique achieves extraordinarily high accuracy
using 15 publicly accessible datasets. The model’s performance, which varies from
97.58 to 99.90%, is greater than that of research on the categorization of Arabic
papers that are equivalent to this one. In terms of the accuracy of the results, the
proposed model performs better than previous Arabic studies and can be applied to
any Arabic-language textual material, without regard to the normalization, stemming,
or pre-processing techniques or methods. We believe that SATCDM will definitely
be of great use to many scientists working in the field of information retrieval.
The aim of this research was to design a way of categorizing the Arabic text on the
basis of the weighting of the terms and the reduction concept belonging to the rough
set theory [5] for reducing the number of words needed to construct the classification
principles which make up the classifier. Research has proposed the minimum multiple
reduction extraction method as a way to improve the quick reduction algorithm.
Several reductions are utilized to create a set of classification rules that act as an
approximate sentence classifier. A corpus of 2700 Arabic documents divided into
nine categories was used to evaluate the proposed method. In the experiment, the
researcher compared the results of the proposed strategy using multiple and single
minimal reductions. The findings revealed that the suggested strategy had a 94%
accuracy rate when using the accuracy of 86%. The outcomes of the tests further
demonstrated that the suggested strategy.
In their study, a brand-new vector assessment technique for categorizing Arabic
text is proposed [6]. The recommended technique makes use of a corpus of classified
Arabic texts before calculating the word weights in the tested document to determine
its keywords, which will be compared to the keywords of the corpus categorizations
to choose the optimum category for the tested document. The experimental findings
Arabic Text Categorization Algorithm Using Vector Space Model 43
demonstrated that, after evaluating the proposal’s contents, the suggested algorithm
prepared the documents to ensure a better selection of the majority of papers, with
98% in one category and 93% in another. The provided technique exhibits the ability
to accurately and quickly categorize Arabic text documents into the appropriate
groups.
This essay suggests a fresh approach to reading Arabic text. Reviewing the relevant
papers in the field demonstrated that numerous strategies for classifying Arabic
text have been suggested by researchers. Examines the use of N-Gram Frequency
Statistics to classify Arabic text sources [7]. This approach uses Manhattan distance
and Dice similarity measurements to categorize an Arabic corpus. Several online
Arabic newspapers provided the Arabic corpus for this study. Four categories are
linked to the data [8]. The “Sport” category beat the other categories in terms of the
recall evaluation measure after numerous iterations of pre-processing the data and
experimentation were conducted. “Economy” had the lowest recall, at about 40%.
The N-gram Dice similarity index generally.
In general, the Manhattan distance similarity measure performed worse than the
N-gram Dice similarity measure. The vector space model was examined using the
k-NN approach by the authors of [9]. Term weighting methods for the Cosine, Dice,
and Jaccard moduli. roll the dice TF. IDF and TF are based on Jaccard. According to
the mean F1 results from the six Arabian data series, they were better than the others.
Using cosine, IDF performed better than TF. IDF, Dice-based WIDF, Dice-based
ITF, Dice-based log (1 + tf), Dice-based WIDF, Jaccard-based ITF and logs (1 + tf),
and Jaccard-based WIDF, ITF and log(1 + tf) on the cosine + tf). Arabic papers were
automatically classified using active collaboration trees [10]. The corpus’s documents
were compiled from Arab Web pages. The corpus contains 6825 entries, which are
separated into the following seven categories: Economics, law, athletics, medicine,
and religion are all related to politics. In order to compare the active collaboration
of trees with SVM, we used Naive Bayes seeding. It was discovered that the naive
Bayes seed and SVM algorithms performed the best in terms of accuracy. This is as
a result of the batch-strengthening technique that improves C4.5’s performance.
The decision-tree method has a downside, though. The association rule algorithm-
based SVM classification algorithm and the naive Bayesian approach, are used in
research to determine Arabic. 5121 Arabic papers of varying lengths from the seven
categories make up the data set. The experimental findings revealed that the asso-
ciation rule-based classification algorithm was superior to the SVM and the naïve
Bayesian technique in terms of f1, precession, retrieval, and measurements.
Alsaleem [11] compiled a big dataset of 5121 documents from various classifica-
tions and used NB and SVM to automatically classify Arabic language documents.
The effects of making the classes longer than 7 classes will be interesting to observe.
The outcomes demonstrate that SVM outperforms NB.
44 E. Hanandeh and M. Shajahan
3 Text Categorization Problem
Text categorization is the procedure of collecting texts into one of several prede-
termined categories based on their content. Newspaper articles could fall under the
following categories, for example: economics, politics, sports, etc., if the texts (study
data) are newspaper articles. Applications for this task include the automatic clas-
sification of emails and online pages. In today’s information-driven society, those
applications are becoming more and more significant. The text classification issue,
according to [12, 13], is as follows: Training datasets and testing datasets were created
from the documents. Let’s assume that the documents d1, d2, and dg make up the
training data set. The classifier uses these g documents as examples, therefore each
of the pertinent categories has to have a significant number of strong examples. The
performance of the classifier was tested using the test data set dg + 1, dg + 2…Dn.
The matrix in Table 1 presents the data divided into training and testing. If Cky = 1
and if Cky = 0, the document dy is considered a positive example of Ck.
In a TC task, evaluation, text categorization, and data pre-processing are the three
main processes. The data preprocessing phase entails preparing the text documents
for classifier training. The text classifier is then developed and improved utilizing a
text learning technique using a training data set. The text classifier is then assessed
using recall, precision, and other metrics. The next two subsections discuss the main
steps of the CT problem regarding the data we used in this article.
A. Arabic data pre-processing
The data for our trials was provided by the Saudi Newspapers (SNP) [14], which
included 5121 Arabic publications in 7 different categories with variable lengths.
Table 2 shows the number of works for every group. The classifications fall under
the headings of (Culture, Economics, General, Information Technology, Politics,
Social and Sport).
English text and Arabic text are not the same. To put it another way, Arabic is
an extremely inflectional and derivational language, making it difficult to analyze it
monophonically. In the Arabic alphabet, some vowels are also marked with diacritics,
which usually do not appear in the text. Furthermore, Arabic capitalizes proper nouns,
which obscures the meaning [15]. Each document file in the Arabic dataset we used
was stored in a disconnected file in the appropriate category directory.
Table 1 Text classification issue represented
Category Training data set Testing data set
d1 … dj dj + 1 … Dn
C1 C11 … C1j C1(j + 1) … C1n
… … … … … … …
Cm Cm1 … Cmg Cm(j + 1) … Cmn
Arabic Text Categorization Algorithm Using Vector Space Model 45
Table 2 Number of
Economics 739
documents in each category
Culture 738
Social 731
Sport 731
General 728
Information technology 728
Politics 726
Total 5121
Also, we made the Arabic dataset so that the classification algorithm grasps it.
At this stage we have processed the Arabic documents according to the procedures
described in [7, 16, 17] for data formatting.
• The Arabic data set’s articles are treated to get rid of the numbers and punctuation.
• We have complied with certain standards while normalizing a number of Arabic
letters, in all of its forms [18].
• All texts that were not in Arabic were filtered.
• Function terms, which were Arabic, were removed. Function words (stop words)
in Arabic are worthless words in IR systems.
B. Assignment of Classification
SVM [19], neural networks [16], and k-NN are only a few techniques for classifying
incoming text [11]. The k-NN, often referred to as text-to-text comparison (TTC), was
used in our investigation [20]. For more than 40 years, pattern recognition researchers
have studied the statistical classification method known as k-NN. As shown in [17, 21,
22], k-NN has been successfully applied to the CT problem and has shown promising
results compared to other statistical processes such as Bayes-Based Network.
4 K-Nearest Neighbors’ Algorithm (k-NN)
Do the outcomes of the supervised categorization of textual material by the algo-
rithm show its efficacy? The storing of the labeled instances is the learning step. By
measuring the separation between each data set instance that has been saved and the
vector representing [10] the new text, new texts are classified. A document gets a
majority class after choosing the next instances. We employed the three-similarity
metrics to conduct a comparative study and because they are crucial to the technique.
Vector Space Model
In the vector space approach, index words of documents and queries are given non-
binary weights [23, 24]. In this case, a suitable/irrelevant match is suggested as a
solution. The non-binary weights assigned to both queries and documents ultimately
46 E. Hanandeh and M. Shajahan
determine how similar each document stored in the system is to a user’s query. In
this way, the vector model also examines documents that only partially match the
search terms. In the vector approach, the document and the query are represented by
vectors of dimension t. Here are the t-dimensional representations of the document
(dj) and query (q), where j is the document number:
The following is the q representation:
q→ = (w1, q, w1.q, . . . wt.q)
The DJ for the document is as follows:
Where wi, q≥ 0, and t are the total number of index terms in the system.
To determine how similar document dj is to query q, the vector model recommends
computing the interconnection between vectors dj and q. This relationship can be
measured with the cosine of the angle, for example.
→
dj = (w1.j, w2.j, . . . wt.j)
∑t
Between these two vectors [23], That is, Sim ( i=1 wi jwi, q).
As shown in Table 3, a vector model can employ various similarity indices besides
cosine similarity [25].
The most common methods for calculating index term weights are as follows
[26]: Weights for binary terms.
The formula for calculating term weights using term frequency inverse to docu-
ment frequency (tf-idf) is as follows: Idf Wij = tf. The two values N and ni are
intended to indicate the total number of documents in the system and separately
the total number of documents containing the index word ki. In the dj-document,
let freqij replace the raw frequency of the term (that is, how many times the term
Table 3 Similarity measures
Similarity measure Evaluation for binary term vector Evaluation for weighted term
vector
|d∩q|
Cosine sim(d, q) = 2 1 1 sim(d j , q) =
|d| 2 ·|q| 2 ∑t
i−1 wi,/ j ×wi,q
/∑ ∑t
t
i=1 wi.2 j × j=1
2
wi,q
|d∩q| ∑t
Dice sim(d, q) = 2 |d|+|q| sim(d j , q) =
2 i=1
∑ 2 i,∑
w j ×wi,q
wi, j + wi,q
2
|d∩r |
Jaccard sim(d, q) = |d|+|q|−|d∩q| sim(d j , q) =
∑t
w j ×wi,q
∑ ∑i=12 i,∑ t
wi,2 j + wi,q − i=1 wi, j ×wi,q
Inner |di ∩ qk | ∑
t
sim = (dik · qk )
k=1
Note The phrase in document d is indicated by the number |d|
Arabic Text Categorization Algorithm Using Vector Space Model 47
ki occurs in the text of the dj-document). Next, the normalized frequency of the ki
phrase in the dj document is presented:
f r eqi j
fij =
max f r eqi j
The two values N and ni are intended to indicate the total number of documents
in the system and separately the total number of documents containing the index
word ki. In the dj-document, let freqij replace the raw frequency of the term (that is,
how many times the term ki occurs in the text of the dj-document). The normalized
frequency of the Ki sentence in the dj document is then.
Where the maximum is calculated using all phrases specified in the document’s
content. If the word “ki” is absent from the document “dj,” then fij equals zero. Let’s
define idfi, the counter document frequency for ki, as follows:
id f i = logni
N
The most common methods of weighting terms use weights provided by wi, j =
f i, j × log 10 ni
N
. These term weighting techniques are called tf-idf systems. For the
weighting of the keyword search, this suggests [25].
( )
0.5 f r eqi, j N
wi, j = 0.5 × × log
max l f r eqi, j ni
where freqi,j is the number of occurrences of the term ki in the document, and max.
Freqi,j is calculated based on all terms listed in the document.
5 Experiment Results
Arabic text differs from English text because it is a highly inflectional and deriva-
tional language, making monophonic analysis difficult. Additionally, some vowels
in Arabic script are represented by diacritics, which are typically omitted from the
text, and proper nouns are capitalized, which causes text ambiguity [27].
Three TC techniques (Cosine, Jaccard, and Dice) based on vector model similarity
have been investigated in terms of the F1 measure (1). These techniques classify
incoming text using the K-NN technique. There are numerous ways to build a text
categorization algorithm, thus we contrasted them using the IDF term weighting
method [28]. On three Pentium IV PCs with 1 GB of Memory, all of the experiments
were carried out using Java.
The subsequent equation is utilized to calculate the F1 measurement:
F1 = 2 ∗ Pr ecision ∗ Recall
Recall + Pr ecision (1)
48 E. Hanandeh and M. Shajahan
Accuracy and recall are common metrics in IR and ML;
See Table 4 for more information.
TP
Pr ecision = (2)
T P + FP
TP
Recall = (3)
T P + FN
The Cosine classification performed better than the Dice and Jaccard algorithms
on all criteria, according to Table 5 study (F1, Precision, and recall).
On 6, 5 data sets, respectively, Cosine outperforms end Dice and Jaccard in terms
of F1 results.
Additionally, recall findings show that on 5, 6 data sets, the cosine outscored the
dice and the Jaccard. And the exact numbers show that Cosinus beat Dice and Jaccard
with 6, 6, and respectively.
Dominant cosine Dice and Jaccard are used to average three measurements using
seven Arabic data sets.
Table 4 Arabic text classification results for precision, recall, and F1
Name of Cosine Dice Jaccard
category
Evaluation Precision Recall F1 Precision Recall F1 Precision Recall F1
measures
Economics 0.931 0.927 0.928 0.852 0.887 0.869 0.843 0.83 0.836
Culture 0.756 0.737 0.746 0.733 0.711 0.722 0.721 0.731 0.726
Social 0.651 0.623 0.636 0.513 0.52 0.516 0.591 0.542 0.565
Sport 0.917 0.979 0.947 0.96 0.934 0.946 0.911 0.94 0.925
General 0.497 0.532 0.514 0.451 0.442 0.446 0.339 0.392 0.364
Information 0.917 0.887 0.902 0.842 0.891 0.865 0.952 0.952 0.952
technology
Politics 0.835 0.873 0.854 0.832 0.921 0.874 0.885 0.842 0.863
Average 0.786 0.794 0.789 0.740 0.758 0.748 0.748 0.747 0.747
Table 5 Confusion matrix
Category Expected actual category Expected other
Totally true Positive true (PT) Negative false (NF)
Other category Positive false (PF) Negative true (NT)
Arabic Text Categorization Algorithm Using Vector Space Model 49
6 Conclusion
This study aimed to improve an Arabic text classifier that can sort Arabic text into
different categories. Using the K-NN method, we investigated several VSM differ-
ences. These variations are represented by the cosine, dice, and jaccard coefficients.
The IDF term weighting approach was also utilized. The computed average of the
three metrics in comparison to seven According to Arabic records, cosine is more
common than dice and Jaccard.
References
1. Musleh Al-Sartawi, A.M.A. (eds.): Artificial intelligence for sustainable finance and sustainable
technology. In: ICGER 2021. Lecture Notes in Networks and Systems, vol. 423. Springer, Cham
(2022)
2. Sawaf, H., Zaplo, J., Ney, H.: Statistical classification methods for Arabic news articles. In:
Arabic Natural Language Processing, Workshop on the ACL’2001. Toulouse, France, July
(2001)
3. Tokunaga, T., Iwayama, M.: Text categorisation based on weighted inverse document frequency.
Department of Computer Science, Tokyo Institute of Technology: Tokyo, Japan (1994)
4. Alhawarat, M., Aseeri, A.O.: A superior Arabic text categorization deep model (SATCDM).
IEEE Access 8, 24653–24661 (2020)
5. Al-Radaideh, Q.A., Al-Abrat, M.A.: An Arabic text categorization approach using term
weighting and multiple reducts. Soft Comput 23, 5849–5863 (2019)
6. Odeh, A., et al.: Arabic text categorization algorithm using vector evaluation method. arXiv
preprint arXiv:1501.01318 (2015)
7. Khreisat, L.: Arabic text classification using n-gram frequency statistics a comparative study.
DMIN 2006, 78–82 (2006)
8. Mohamed, B., Mounir, Z.: Text mining approaches for dependent bug report assembly and
severity prediction. Int. Arab J. Inform. Technol. (2022)
9. Thabtah, F., Hadi, W., Al-Shammare, G.: VSMs with K-nearest neighbour to categorise Arabic
text data. In: The World Congress on Engineering and Computer Science 2008, pp. 778–781,
22–44. San Francisco, USA (2008)
10. Guo, G., Wang, H., Bell, D., Bi, Y., Greer, K.: An kNN model-based approach and its application
in text categorization. In: Proceedings of 5th International Conference on Intelligent Text
Processing and Computational Linguistic, CICLing, LNCS 2945, Springer-Verlag, pp. 559–570
(2004)
11. Alsaleem, S.: Automated Arabic text categorization using SVM and NB. Int. Arab J.
e-Technology 2(2) (2011)
12. Sebastiani, F.: Text categorization. In: Alessandro, Z. (ed.), Text mining and its applications.
WIT Press, Southampton, UK, pp. 109–129 (2005)
13. Syiam, M.M., Fayed, Z.T., Habib, M.B.: An intelligent system for Arabic text categorization.
IJICIS 6(1) (2006)
14. Al-Harbi, S.: Automatic Arabic text classification. JADT 08:9esJournées internationalesd
Analysestatistique des Données Textuelles, pp. 77–83 (2008)
15. Hammo, B., Abu-Salem, H., Lytinen, S., Evens, M.: QARAB: a question answering system
to support the Arabic language. In: Workshop on Computational Approaches to Semitic
Languages. ACL, Philadelphia, PA, July, pp. 55–65 (2002)
16. Benkhalifa, M., Mouradi, A., Bouyakhf, H.: Integrating WordNet knowledge to supplement
training data in semi- supervised agglomerative hierarchical clustering for text categorization.
Int. J. Intel Syst 16(8), 929–947 (2001)
50 E. Hanandeh and M. Shajahan
17. Guo, Y., Shao, Z., Hua, N.: Automatic text categorization based on content analysis with
cognitive situation models. Inf. Sci. 180, 613–630 (2010)
18. Samir, A., Ata, W., Darwish, N.: A new technique for automatic text categorization for Arabic
documents. In: 5th IBIMA Conference (The internet & information technology in modern
organizations), Cairo, Egypt
19. Joachims, T.: Text categorisation with support vector machines: learning with many relevant
features. In: Proceedings of the European Conference on Machine Learning (ECML), pp. 173–
142. Berlin (1998)
20. El-Halees, A.: Mining Arabic association rules for text classification. In: The Proceedings of
the First International Conference on Mathematical Sciences. Al-Azhar University of Gaza,
Palestine, 15–17 (2006)
21. Junker, M., Hoch, R., Dengel, A.: On the evaluation of document analysis components by recall,
precision, and accuracy. In: Proceedings of the Fifth International Conference on Document
Analysis and Recognition (1999)
22. Yang, Y.: An evaluation of statistical approaches to text categorization. J. Inf. Retrieval 1(1/2),
67–88 (1999)
23. Salton, G.: Automatic information organization and retrieval (1968).
24. Hanandeh, E.S., Awwad, A.A., Khassawneh, Y.: Classify Arabic text using vector space models.
In: 2021 22nd International Arab Conference on Information Technology (ACIT). IEEE (2021)
25. Salton, G., Buckley, C.: Parallel text search methods. Commun. ACM 31(2), 202–215 (1988)
26. Salton, G., McGill, M.J.: Introduction to modern information retrieval. Mcgraw-Hill (1983)
27. El-Halees A.: Arabic text classification using maximum entropy the Islamic university journal
(Series of Natural Studies and Engineering) 15(1), 157–167 (2007)
28. El-Kourdi, M., Bensaid, A., Rachidi, T.: Automatic Arabic document categorisation based on
the Naïve Bayes algorithm. In: 20th International Conference on Computational Linguistics,
August 28th, Geneva (2004)