Online anomaly detection for multi-source VMware using a distributed streaming framework

dc.authorid0000-0002-7323-3695
dc.authorid0000-0002-9300-1576
dc.contributor.authorSolaimani, Mohiuddin
dc.contributor.authorIftekhar, Mohammed
dc.contributor.authorKhan, Latifur
dc.contributor.authorThuraisingham, Bhavani
dc.contributor.authorIngram, Joe
dc.contributor.authorSeker, Sadi Evren
dc.date.accessioned2025-05-10T19:54:03Z
dc.date.issued2016
dc.departmentİstanbul Medeniyet Üniversitesi
dc.description.abstractAnomaly detection refers to the identification of patterns in a dataset that do not conform to expected patterns. Such non-conformant patterns typically correspond to samples of interest and are assigned to different labels in different domains, such as outliers, anomalies, exceptions, and malware. A daunting challenge is to detect anomalies in rapid voluminous streams of data. This paper presents a novel, generic real-time distributed anomaly detection framework for multi-source stream data. As a case study, we investigate anomaly detection for a multi-source VMware-based cloud data center, which maintains a large number of virtual machines (VMs). This framework continuously monitors VMware performance stream data related to CPU statistics (e.g., load and usage). It collects data simultaneously from all of the VMs connected to the network and notifies the resource manager to reschedule its CPU resources dynamically when it identifies any abnormal behavior from its collected data. A semi-supervised clustering technique is used to build a model from benign training data only. During testing, if a data instance deviates significantly from the model, then it is flagged as an anomaly. Effective anomaly detection in this case demands a distributed framework with high throughput and low latency. Distributed streaming frameworks like Apache Storm, Apache Spark, S4, and others are designed for a lower data processing time and a higher throughput than standard centralized frameworks. We have experimentally compared the average processing latency of a tuple during clustering and prediction in both Spark and Storm and demonstrated that Spark processes a tuple much quicker than storm on average. Copyright (c) 2016 John Wiley & Sons, Ltd.
dc.description.sponsorshipLaboratory Directed Research and Development program at Sandia National Laboratories; National Science Foundation (NSF); US Department of Energy's National Nuclear Security Administration [DE-AC04-94AL85000]; NSF [CNS 1229652, DUE 1129435]; Direct For Computer & Info Scie & Enginr; Division Of Computer and Network Systems [1229652] Funding Source: National Science Foundation
dc.description.sponsorshipFunding for this work was partially supported by the Laboratory Directed Research and Development program at Sandia National Laboratories and The National Science Foundation (NSF). Sandia National Laboratories is a multi-program laboratory managed and operated by Sandia Corporation, a wholly owned subsidiary of Lockheed Martin Corporation, for the US Department of Energy's National Nuclear Security Administration under contract DE-AC04-94AL85000. NSF grant is contracted under NSF award No. CNS 1229652 and NSF Award No. DUE 1129435.
dc.identifier.doi10.1002/spe.2390
dc.identifier.endpage1497
dc.identifier.issn0038-0644
dc.identifier.issn1097-024X
dc.identifier.issue11
dc.identifier.scopus2-s2.0-84990184652
dc.identifier.scopusqualityQ1
dc.identifier.startpage1479
dc.identifier.urihttps://doi.org/10.1002/spe.2390
dc.identifier.urihttps://hdl.handle.net/20.500.14730/12917
dc.identifier.volume46
dc.identifier.wosWOS:000386150000003
dc.identifier.wosqualityQ2
dc.indekslendigikaynakWeb of Science
dc.indekslendigikaynakScopus
dc.language.isoen
dc.publisherWiley
dc.relation.ispartofSoftware-Practice & Experience
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/closedAccess
dc.snmzKA_WOS_20250302
dc.subjectreal-time anomaly detection
dc.subjectincremental clustering
dc.subjectresource scheduling
dc.subjectdata center
dc.subjectApache Spark
dc.subjectApache Storm
dc.titleOnline anomaly detection for multi-source VMware using a distributed streaming framework
dc.typeArticle

Dosyalar

Orijinal paket

Listeleniyor 1 - 1 / 1
Yükleniyor...
Küçük Resim
İsim:
12917.pdf
Boyut:
1.67 MB
Biçim:
Adobe Portable Document Format