Metin temsili ve model seçiminin sınıflandırma performansına etkisi: Covid19-FNIR veri seti üzerinde TF-IDF, BoW ve Transformatör tabanlı yöntemlerin kapsamlı bir karşılaştırması

dc.contributor.authorBaşarslan, Muhammet Sinan
dc.contributor.authorBal, Fatih
dc.date.accessioned2025-11-16T19:21:40Z
dc.date.issued2025
dc.departmentİstanbul Medeniyet Üniversitesi
dc.description.abstractBu çalışmada, Terim Frekansı-Ters Doküman Frekansı (TF-IDF) ve Bag of Words (BoW) metin vektörleştirmesi kullanılarak %80 eğitim ve %20 teste ayrılmış bir veri kümesi üzerinde çeşitli makine öğrenimi (ML) modellerinin performansı değerlendirilmiştir. DistilBERT, RoBERTa ve alBERT gibi dönüştürücü tabanlı modeller, klasik makine öğrenimi algoritmaları ve Stacking, Hard Voting ve Soft Voting gibi topluluk yöntemleriyle entegre edilmiştir. Yığınlama her iki yöntemle de en yüksek performansı elde etmiştir- TF-IDF ile %92.62 Doğruluk ve %92.51 F1, BoW ile %92.29 Doğruluk ve %92.41 F1. BoW ile Hard Voting en yüksek geri çağırmayı (%95,23) vermiştir. Lojistik Regresyon ve DVM gibi klasik modeller BoW ile daha iyi performans göstererek sırasıyla %90.98 ve %90.51 Doğruluğa ulaşmıştır. Genel olarak, TF-IDF dengeli sonuçlar üretirken, BoW belirli durumlarda daha yüksek geri çağırma ve kesinlik sunmuştur. Bu sonuçlar, optimum sınıflandırma performansına ulaşmada hem model hem de metin temsili seçimlerinin önemini vurgulamaktadır.
dc.description.abstractThis study evaluates the performance of various machine learning (ML) models on a dataset split into 80% training and 20% testing using Term Frequency-Inverse Document Frequency (TF-IDF) and Bag of Words (BoW) text vectorization. Transformer-based models like DistilBERT, RoBERTa, and alBERT were integrated with classical ML algorithms and ensemble methods such as Stacking, Hard Voting, and Soft Voting. Stacking achieved the highest performance with both methods—92.62% Accuracy (Acc) and 92.51% F1-score (F1) with TF-IDF, and 92.29% Acc and 92.41% F1 with BoW. Hard Voting with BoW yielded the highest Recall (95.23%). Classical models like Logistic Regression (LR) and Support Vector Machine (SVM) performed better with BoW, reaching 90.98% and 90.51% Acc, respectively. Overall, TF-IDF produced balanced outcomes, while BoW offered higher Recall and Precision in specific cases. These results highlight the significance of both model and text representation choices in achieving optimal classification performance.
dc.identifier.doi10.28948/ngumuh.1694988
dc.identifier.endpage1461
dc.identifier.issn2564-6605
dc.identifier.issue4
dc.identifier.startpage1447
dc.identifier.urihttps://doi.org/10.28948/ngumuh.1694988
dc.identifier.urihttps://hdl.handle.net/20.500.14730/14456
dc.identifier.volume14
dc.language.isoen
dc.publisherNigde Omer Halisdemir University
dc.relation.ispartofNigde Omer Halisdemir University Journal of Engineering Sciences
dc.relation.publicationcategoryMakale - Ulusal Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_DergiPark_20251116
dc.subjectMachine Vision
dc.subjectYapay Görme
dc.subjectNatural Language Processing
dc.subjectDoğal Dil İşleme
dc.titleMetin temsili ve model seçiminin sınıflandırma performansına etkisi: Covid19-FNIR veri seti üzerinde TF-IDF, BoW ve Transformatör tabanlı yöntemlerin kapsamlı bir karşılaştırması
dc.title.alternativeThe effect of text representation and model selection on classification performance: A comprehensive comparison of TF-IDF, Bow and Transformer-based methods on the Covid19-FNIR dataset
dc.typeArticle

Dosyalar