Improving Arabic Text Classification Accuracy Using Lightweight NLP Techniques
DOI:
https://doi.org/10.71285/icpt.v3i2.43Keywords:
Arabic natural language processing, Arabic text classification, TF-IDF method, Chi-Square feature selection, linear SVM classifier, Logistic Regression, BiLSTM, AraBERT, efficient NLP models, accuracy and efficiency trade-offAbstract
Arabic text classification is a challenging task because of the complex morphology of the language, the existence of different writing forms and a multitude of dialects, which can result in sparser common text representations. While transformer models such as AraBERT have obtained superior results on many Arabic NLP tasks, their high computational requirements make them difficult to deploy in environments with limited hardware resources. In some cases this can also make the model less practical for researchers working with basic computer systems.
This study focuses on a more practical issue: how much accuracy a simple classifier may lose when the amount of required computation is reduced. We use a combined TF-IDF representation based on both words and characters, then reduce the number of features using Chi-Square selection. The selected features are finally used with a linear SVM to create a model that is faster and more efficient, while still keeping good predictive performance. However, the method do not always provide the same level of accuracy as more complex models, especially with difficult text.
The proposed lightweight method is tested against several other approaches, including word-based TF-IDF, character-based TF-IDF, Naive Bayes, Logistic Regression, BiLSTM, and AraBERT. All methods are tested using the same data distribution to make the comparison fair. The dataset used is the arbml/arabic_100k_reviews corpus, which originally contains 99,999 balanced reviews from three sentiment categories. After cleaning the data and removing very short reviews, the dataset was reduced to 99,759 samples. The training, validation, and testing sets includes 69,831, 9,976, and 19,952 reviews, respectively. This setup allows the different models to be compared under similar conditions and with the same evaluation process.
After applying Chi-Square feature selection, the hybrid feature set was reduced from 250,000 to 50,000 dimensions, which represents an 80% reduction. That proposed method achieved 67.90% accuracy and a Macro-F1 score of 67.75%. Its training time was around 41 seconds, while the complete test set was processed in about 0.03 seconds. Among all the evaluated models, AraBERT achieved the best overall performance, reaching 74.08% accuracy and 74.25% Macro-F1. However, its computational requirements were much higher, with nearly 1,760 seconds needed for training and about 47 seconds for inference on the same test set. Depending on the processing stage, this makes AraBERT approximately 40 to 1,500 times slower than the proposed lightweight approach. The proposed method is therefore not presented as a replacement for AraBERT in terms of accuracy, since its accuracy is lower. Instead, its main advantage is the considerable and measurable reduction in computational cost. This trade-off can be useful in situations where GPU availability, memory capacity, or response time are limited.
References
R. Bensoltane and T. Zaki, “Towards Arabic aspect-based sentiment analysis: A transfer learning-based approach,” Social Network Analysis and Mining, vol. 12, no. 1, pp. 1–16, 2022, doi: 10.1007/s13278-021-00794-4.
R. Bensoltane and T. Zaki, “Aspect-based sentiment analysis: An overview in the use of Arabic language,” Artificial Intelligence Review, vol. 56, no. 3, pp. 2325–2363, 2023, doi: 10.1007/s10462-022-10215-3.
A. Al-Hassan and H. Al-Dossari, “Detection of hate speech in Arabic tweets using deep learning,” Multimedia Systems, vol. 28, no. 6, pp. 1963–1974, 2022, doi: 10.1007/s00530-020-00742-w.
A. S. Fadel, M. E. Saleh, and O. A. Abulnaja, “Arabic aspect extraction based on stacked contextualized embedding with deep learning,” IEEE Access, vol. 10, pp. 30526–30535, 2022, doi: 10.1109/ACCESS.2022.3159252.
N. Elhassan, G. Varone, R. Ahmed, M. Gogate, K. Dashtipour, H. Almoamari, M. A. El-Affendi, B. N. Al-Tamimi, F. Albalwy, and A. Hussain, “Arabic sentiment analysis based on word embeddings and deep learning,” Computers, vol. 12, no. 6, Art. no. 126, 2023, doi: 10.3390/computers12060126.
M. Masadeh, A. Moustapha, B. Sharada, J. Hanumanthappa, K. Hemachandran, C. Chola, and A. Y. Muaad, “Investigating the impact of preprocessing techniques and representation models on Arabic text classification using machine learning,” International Journal of Advanced Computer Science and Applications, vol. 15, no. 1, 2024, doi: 10.14569/IJACSA.2024.01501110.
S. M. Alzanin, A. Gumaei, M. A. Haque, and A. Y. Muaad, “An optimized Arabic multilabel text classification approach using genetic algorithm and ensemble learning,” Applied Sciences, vol. 13, no. 18, Art. no. 10264, 2023, doi: 10.3390/app131810264.
M. Hadni and H. Hjiaj, “A new metaheuristic optimization technique for solving feature selection and classification problems for Arabic text,” in Arabic Language Processing: From Theory to Practice, Commun. Comput. Inf. Sci., vol. 2340, pp. 221–235, Springer, 2024, doi: 10.1007/978-3-031-80438-0_17.
M. Hadni and H. Hjiaj, “New model of feature selection based chaotic firefly algorithm for Arabic text categorization,” The International Arab Journal of Information Technology, vol. 20, no. 3A, pp. 461–468, 2023, doi: 10.34028/iajit/20/3A/3.
T. Sabri, S. Bahassine, O. El Beggar, and M. Kissi, “An improved Arabic text classification method using word embedding,” International Journal of Electrical and Computer Engineering, vol. 14, no. 1, pp. 721–731, 2024, doi: 10.11591/ijece.v14i1.pp721-731.
H. Alangari and N. Algethami, “Exploring the effects of pre-processing techniques on topic modeling of an Arabic news article data set,” Applied Sciences, vol. 14, no. 23, Art. no. 11350, 2024, doi: 10.3390/app142311350.
S. Albahli, “An advanced natural language processing framework for Arabic named entity recognition: A novel approach to handling morphological richness and nested entities,” Applied Sciences, vol. 15, no. 6, Art. no. 3073, 2025, doi: 10.3390/app15063073.
T. El Moussaoui and C. Loqman, “Advancements in Arabic named entity recognition: A comprehensive review,” IEEE Access, vol. 12, pp. 180238–180266, 2024, doi: 10.1109/ACCESS.2024.3491897.
M. Lichouri, K. Lounnas, Z. N. Zahaf, and A. Rabiai, “dzNLP at NADI 2024 shared task: Multi-classifier ensemble with weighted voting and TF-IDF features,” in Proc. 2nd Arabic Natural Language Processing Conf. (ArabicNLP), 2024, pp. 754–757, doi: 10.18653/v1/2024.arabicnlp-1.84.
S. E. Bekhouche, A. Z. Sellam, H. Telli, C. Distante, and A. Hadid, “CVPD at QIAS 2025 shared task: An efficient encoder-based approach for Islamic inheritance reasoning,” arXiv preprint arXiv:2509.00457, 2025.
ARBML, “Arabic 100k reviews,” Hugging Face, Dataset. Online Available: https://huggingface.co/datasets/arbml/arabic_100k_reviews
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Rasha Majid Hassoon, Safana Mohammed Qasem, Sarah Hikmat Khaled, Reiam Abd Al-Kareim Abd, Rusul Hussein Hasan, Liqaa Mohammad Shoohi

This work is licensed under a Creative Commons Attribution 4.0 International License.





