Mikrobiyota Verileri İçin Boyut İndirgemede Yeni Bir Yaklaşım

dc.contributor.authorAnkaralı, Handan
dc.contributor.authorYıldırım, Süleyman
dc.contributor.authorBulut, Nurgül
dc.date.accessioned2025-05-10T11:34:05Z
dc.date.issued2021
dc.departmentİMÜ, Fakülteler, Temel Tıp Bilimleri Bölümü
dc.description.abstractİnsan derisi, nazofaringeal ve ağız boşlukları, vajinal sistem ve gastrointestinal sistem ile ilişkili mikroorganizmalar insan mikrobiyotasını oluşturur. Fizyolojik, metabolik ve immun sistem üzerinde oldukça etkilidir ve birçok hastalık ile ilişkisi gösterilmiştir. DNA dizileme teknolojisindeki son gelişmeler, bakteriler için 16S rRNA, 18s rRNA veya ITS gibi marker genlerinin amplikonlarının yüksek verim dizilimi yoluyla, mikrobiyal toplulukların profillenmesi kolaylaşmıştır. Elde edilen veriler, çok büyük sayılarda mikrobiyota türlerine ait frekans değerlerinden oluşur ve bol miktarda sıfır değeri içerir. Mikrobiyota verileri gibi büyük boyutlu verilerin çeşitli istatistik modellerle analiz edilebilmesi için ön işleme aşamasında, sonuca anlamlı katkısı bulunmayan türlerin veri analizinden çıkarılması gerekmektedir. İstatistik literatüründe bu işlem, boyut indirgeme veya değişken eleme olarak adlandırılmaktadır. Bu çalışmada, çok sayıda sıfır değeri içeren frekans tipi büyük boyutlu veri setlerinde, boyut indirgeme amacıyla kullanılabilecek yeni bir yaklaşım önerildi. Bu amaçla, tek değişkenli testler, sıfır etkili negatif binomiyal model, sınıflama ve regresyon ağaçları ve değişken seçimi algoritması kullanıldı. Önerilen yaklaşım, Parkinson hastaları, erken demans ve kontrol bireylerinden elde edilen mikrobiyota cinsleri üzerinde denendi. Değişken seçimi sonucunda 199 bakteri cinsi içinden seçilen 19 adet aday cinsin, klinik açıdan da birçok çalışmada vurgulanan bakteri cinsleri olduğu görüldü. Aday olarak seçilen cinslerin hastalık tanısındaki başarısını değerlendirmek için kurulan multiple logistic regresyon modelinde yeniden stepwise değişken eleme yöntemi kullanıldı ve bu model sonucunda birkaç bakteri cinsi ile başarılı bir şekilde hasta ve kontrol gruplarının ayrımı yapıldı. Bu çalışma ile önerilen yeni hibrit yaklaşım, birden çok yöntemin ortak kararı neticesinde belirlenen değişkenleri veri analizine alma imkanı sunmaktadır. Benzeri yaklaşımlar farklı yöntemlerle denenerek farklı veri tipleri üzerinde kullanılabilir.
dc.description.abstractMicroorganisms associated with human skin, nasopharyngeal and oral cavities, vaginal tract, and gastrointestinal system make up the human microbiota. It is highly effective on the physiological, metabolic and immune system and has been shown to be associated with many diseases. Recent advances in DNA sequencing technology have facilitated profiling of these microbial communities through high throughput sequencing of amplicons of the marker genes such as 16S rRNA for bacteria, 18S rRNA or ITS. Data generated from such sequencing efforts are preprocessed into composition or relative abundance that are often presented in species abundance (OTU/ASV) tables. The data obtained consists of the frequency of microbiota species in very large numbers and it contains a large amount of zero values. Nonetheless, the high dimensional data in such tables must be treated with dimension reduction techniques to draw sensible conclusions from the data. In the statistical literature, this process is called dimension reduction or variable selection. The aim in this study is to propose a novel approach to reduce dimensions in high dimensional and inherently zero inflated and frequency character microbiota data. For this purpose, univariate tests, a zero-inflated negative binomial model, classification and regression trees, and a feature selection and variable screening algorithm were used. Using these four methods enabled us to select most important features of the microbiota dataset for the subsequent downstream analyses. We tested the above approach on our recent microbiota dataset we generated from stool samples of Parkinson’s disease patients cohort. Of 199 bacteria genera our approach enabled us to select 19 candidate biomarker genera, which are often implicated in serving critical metabolic activities in human body such as production of short-chain fatty acids. To assess the potential of these candidate biomarkers in differentiating disease and healthy states we developed a multiple logistic regression model and further selected their biomarker potential in a stepwise variable screening. Big data analysis necessarily entails use of increasingly more sophisticated and combinatorial modalities. Here we successfully demonstrated that hitherto untested combinatorial use of feature selection methods enables more useful predictive models. Similar approaches can be tried with different methods and used on different data types.
dc.identifier.endpage30
dc.identifier.issn2667-582X
dc.identifier.issue1
dc.identifier.startpage23
dc.identifier.urihttps://dergipark.org.tr/tr/pub/veri/issue/59505/768569
dc.identifier.urihttps://hdl.handle.net/20.500.14730/3217
dc.identifier.volume4
dc.language.isotr
dc.publisherMurat GÖK
dc.relation.ispartofVeri Bilimi
dc.relation.publicationcategoryMakale - Ulusal Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/openAccess
dc.snmzKA_DergiPark_20250302
dc.subjectSıfır etkili modeller
dc.subjectFrekans verisi
dc.subjectSınıflama ve Regresyon ağaçları
dc.subjectDeğişken seçim algoritmaları
dc.subjectMikrobiyota
dc.subjectParkinson hastalığı
dc.subjectZero-inflated models
dc.subjectFrequency data
dc.subjectClassification and Regression tree
dc.subjectVariable Screening algorithm
dc.subjectMicrobiota
dc.subjectParkinson’s disease
dc.titleMikrobiyota Verileri İçin Boyut İndirgemede Yeni Bir Yaklaşım
dc.titleA New Approach to Dimension Reduction for Microbiota Data
dc.typeArticle

Dosyalar

Orijinal paket

Listeleniyor 1 - 1 / 1
Yükleniyor...
Küçük Resim
İsim:
3217.pdf
Boyut:
467.45 KB
Biçim:
Adobe Portable Document Format