Arabic Plagiarism Detection Using Word2Vec-Based Semantic Features and Random Forest Classification on the ExAraPlagDet Dataset

Abstract

Plagiarism detection has become a critical challenge in the digital age, particularly for languages with complex structures such as Arabic. Traditional methods relying on string matching and basic lexical analysis, are insufficient for detecting more sophisticated forms of plagiarism like paraphrasing and synonym substitution in Arabic texts. This research addresses this gap by proposing a novel approach that employs that employs word embedding and semantic similarity measures within a machine learning framework. Specifically, we utilize models such as Support Vector Machines (SVM) and neural networks to capture the nuanced semantic relationships between words, enabling more effective detection of subtle semantic similarities in text. Our methodology encompasses the development and evaluation of machine learning models, specifically tailored to the unique characteristics of the Arabic language [1]. We conducted extensive experiments on a diverse dataset of Arabic texts, consisting of over 50,000 documents from various sources, including academic publications, online articles, and literary works, to demonstrate the effectiveness of our approach. Our results demonstrate substantial improvements in both accuracy and robustness, surpassing traditional plagiarism detection techniques. The Random Forest classifier achieved the best performance with precision, recall, and F1-score all reaching 0.96, significantly outperforming Decision Tree, Logistic Regression, and SVM. These results confirm the superiority of the Random Forest approach for the given classification problem. This study contributes to the field by developing a more effective tool for plagiarism detection, which is crucial for maintaining academic integrity and protecting intellectual property in Arabic-speaking communities. The findings underscore the potential of advanced NLP techniques to overcome language-specific challenges, providing a foundation for future research in multilingual plagiarism detection and enhancing the development of tools for other languages facing similar challenges.

Author Biographies

  • Hanan Mohammed Fawzy, Department of Computer Science, Faculty of Computers and Informatics, Zagazig University, Egypt
    BIO
  • Ahmad Salah, Department of Computer Science, Faculty of Computers and Informatics, Zagazig University, Egypt
    BIO
  • Heba El-Fiqi, School of Systems and Computing, University of New South Wales Canberra, Australia
    BIO
  • Mahmoud Mahdi, Department of Computer Science, Faculty of Computers and Informatics, Zagazig University, Egypt
    BIO

References

Authors

  • Hanan Mohammed Fawzy Department of Computer Science, Faculty of Computers and Informatics, Zagazig University, Egypt
  • Ahmad Salah Department of Computer Science, Faculty of Computers and Informatics, Zagazig University, Egypt
  • Heba El-Fiqi School of Systems and Computing, University of New South Wales Canberra, Australia
  • Mahmoud Mahdi Department of Computer Science, Faculty of Computers and Informatics, Zagazig University, Egypt

DOI:

https://doi.org/10.31449/inf.v50i2.11936

Downloads

Published

08/04/2026

Issue

Section

Regular papers

How to Cite

Fawzy, H. M., Salah, A., El-Fiqi, H., & Mahdi, M. (2026). Arabic Plagiarism Detection Using Word2Vec-Based Semantic Features and Random Forest Classification on the ExAraPlagDet Dataset. Informatica, 50(2). https://doi.org/10.31449/inf.v50i2.11936