Comparison of Spelling Error Detection in Turkish Texts With Machine Learning and Transformer-Based Approaches

dc.contributor.authorErdagi, Erturk
dc.date.accessioned2025-11-16T19:34:11Z
dc.date.issued2025
dc.departmentİstanbul Medeniyet Üniversitesi
dc.description.abstractThis study aims to compare the performance of five different models for spelling error detection, a crucial task in natural language processing. In this study, the performance of the Logistic Regression model, one of the traditional machine learning techniques, was investigated in both default and optimized hyperparameter forms. The LSTM structure, which provides good results on sequential data, the BERT model, which has a high ability to represent contextual language, and the Soft-Masked BERT model, designed specifically for the task of spelling error detection, were examined. All models were evaluated on a balanced Turkish dataset containing at least one spelling error based on the sentences in the dataset, and comparisons were made using basic classification metrics such as accuracy, precision, recall, F1-score, and ROC AUC. The findings show that the Logistic Regression model running with the default parameters failed to distinguish between classes with 0.366 accuracy and F1-score, while the best hyperparameter version of this model provided a significant improvement with 0.675 accuracy and F1-score. The LSTM model achieved partial success in learning sequential structure with 0.713 accuracy and 0.7 F1-score, outperforming the Logistic Regression model. The BERT model achieved 0.886 accuracy and 0.885 F1-score, outperforming both the Logistic Regression and LSTM models. The Soft-Masked BERT model achieved the highest success with 0.897 accuracy and F1-score. These results demonstrate that transformer-based models perform better on tasks involving both the morphological and contextual structure of the language. This study compares and demonstrates the effectiveness of various model architectures in detecting spelling errors in Turkish texts and highlights the contribution of context sensitivity to classification success.
dc.identifier.doi10.1109/ACCESS.2025.3619865
dc.identifier.endpage175675
dc.identifier.issn2169-3536
dc.identifier.scopus2-s2.0-105018720827
dc.identifier.scopusqualityQ1
dc.identifier.startpage175662
dc.identifier.urihttps://doi.org/10.1109/ACCESS.2025.3619865
dc.identifier.urihttps://hdl.handle.net/20.500.14730/15269
dc.identifier.volume13
dc.identifier.wosWOS:001594883400003
dc.identifier.wosqualityN/A
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherIeee-Inst Electrical Electronics Engineers Inc
dc.relation.ispartofIeee Access
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_WOS_20250302
dc.subjectAccuracy
dc.subjectContext modeling
dc.subjectLogistic regression
dc.subjectNatural language processing
dc.subjectData models
dc.subjectMerging
dc.subjectKeyboards
dc.subjectEncoding
dc.subjectDictionaries
dc.subjectTransformers
dc.subjectSpelling error detection
dc.subjecttext classification
dc.subjectlogistic regression
dc.subjectLSTM
dc.subjectBERT
dc.subjectsoft-masked BERT
dc.subjectTurkish language processing
dc.titleComparison of Spelling Error Detection in Turkish Texts With Machine Learning and Transformer-Based Approaches
dc.typeArticle

Dosyalar