NeXtBrain: Combining local and global feature learning for brain tumor classification

dc.contributor.authorPacal, Ishak
dc.contributor.authorAkhan, Ozan
dc.contributor.authorDeveci, Rumeysa Tuna
dc.contributor.authorDeveci, Muhammet
dc.date.accessioned2025-11-16T19:33:38Z
dc.date.issued2025
dc.departmentİstanbul Medeniyet Üniversitesi
dc.description.abstractThe accurate and timely diagnosis of brain tumors is of paramount clinical significance for effective treatment planning and improved patient outcomes. While deep learning has advanced medical image analysis, concurrently achieving high classification accuracy, robust generalization, and computational efficiency remains a formidable challenge. This is often due to the difficulty in optimally capturing both fine-grained local tumor features and their broader global contextual cues without incurring substantial computational costs. This paper introduces NeXtBrain, a novel hybrid architecture meticulously designed to overcome these limitations. NeXt-Brain's core innovations, the NeXt Convolutional Block (NCB) and the NeXt Transformer Block (NTB), synergistically enhance feature learning: NCB leverages Multi-Head Convolutional Attention and a SwiGLU-based MLP to precisely extract subtle local tumor morphologies and detailed textures, while NTB integrates self-attention with convolutional attention and a SwiGLU MLP to effectively model long-range spatial dependencies and global contextual relationships, crucial for differentiating complex tumor characteristics. Evaluated on two publicly available benchmark datasets, Figshare and Kaggle, NeXtBrain was rigorously compared against 17 state-of-the-art (SOTA) models. On Figshare, it achieved 99.78 % accuracy and a 99.77 % F1-score. On Kaggle, it attained 99.78 % accuracy and a 99.81 % F1-score, surpassing leading SOTA ViT, CNN, and hybrid models. Critically, NeXtBrain demonstrates exceptional computational efficiency, achieving these SOTA results with only 23.91 million parameters, requiring just 10.32 GFLOPs, and exhibiting a rapid inference time of 0.007 ms. This efficiency allows it to outperform significantly larger models such as DeiT3-Base with 85.82 M parameters, Swin-Base with 86.75 M parameters in both accuracy and computational demand.
dc.identifier.doi10.1016/j.brainres.2025.149762
dc.identifier.issn0006-8993
dc.identifier.issn1872-6240
dc.identifier.pmid40490088
dc.identifier.scopus2-s2.0-105007671415
dc.identifier.scopusqualityQ2
dc.identifier.urihttps://doi.org/10.1016/j.brainres.2025.149762
dc.identifier.urihttps://hdl.handle.net/20.500.14730/15107
dc.identifier.volume1863
dc.identifier.wosWOS:001510828300001
dc.identifier.wosqualityN/A
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.indekslendigikaynakPubMed
dc.language.isoen
dc.publisherElsevier
dc.relation.ispartofBrain Research
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_WOS_20250302
dc.subjectHealth
dc.subjectMedical imaging
dc.subjectDeep learning
dc.subjectBrain tumor detection
dc.subjectVision transformers
dc.titleNeXtBrain: Combining local and global feature learning for brain tumor classification
dc.typeArticle

Dosyalar