Sentiment Analysis Using Various Machine Learning Techniques on Depression Review Data
Dosyalar
Tarih
Dergi Başlığı
Dergi ISSN
Cilt Başlığı
Yayıncı
Erişim Hakkı
Özet
In this work, we investigate the effectiveness of different machine learning (ML) models combined with different text representation techniques on the 'Depression Corpus of Arabic Tweets' dataset. This dataset is crucial for sentiment analysis (SA), an important task in natural language processing (NLP). First, we perform data cleaning and preprocessing to remove noise and irrelevant information from the tweets, including non-Arabic characters, extra spaces, and stop words, in order to focus on meaningful content. Next, we investigate two popular text representation methods: Term Frequency Inverse Document Frequency (TF-IDF) and Bag-of-Words (BoW). With these methods, we transform text into numerical features that can be used by ML models, and build models using Naïve Bayes (NB), Decision Tree (DT), and Support Vector Machines (SVM) methods. We use standard evaluation metrics such as Precision (P), Recall (R), F1-Score (F1), and Accuracy (Acc) to evaluate the performance of these models. In this study, we obtained an accuracy of 0.9585 with the TF-IDF and SVM model, which is competitive with the literature. In the future, advanced techniques with deep learning and pre-trained models with newly developed word representations will be used to further improve various NLP tasks on Arabic texts. © 2024 IEEE.










