Cross-domain One-shot Video Object Detection
| dc.authorid | 0000-0003-3628-3316 | |
| dc.authorid | 0000-0002-9857-3012 | |
| dc.contributor.author | Hanoğlu, Yusuf Kağan | |
| dc.contributor.author | Günsel, Bilge | |
| dc.contributor.author | Gürkan, Filiz | |
| dc.date.accessioned | 2026-09-09T10:40:40Z | |
| dc.date.issued | 2025 | |
| dc.department | İMÜ, Fakülteler, Mühendislik ve Doğa Bilimleri Fakültesi, Elektrik-Elektronik Mühendisliği Bölümü | |
| dc.description.abstract | One-shot-object detection (OSOD) aims to detect novel object classes using a single example of an unseen class. Cross-domain OSOD is a more challenging problem since the seen and unseen objects are sampled from the entirely disjoint datasets. The majority of the existing CD-OSOD methods focus on image datasets where the video domain remains largely unaddressed. To tackle this problem, we introduce a one-shot cross-domain video object detection (CD-OSVOD) model enabling adaptation from the still image to the video. Specifically the novel target object is designated as the query shot and a target driven cross-domain finetuning (FT) scheme is integrated with a baseline object detector. To address the requirements of the long term video object detection, the FT scheme is augmented with a novel Online Target Update (OTU) mechanism, enabling the detector to handle challenges such as appearance changes and occlusions. The OTU is controlled by a temporal aggregation module (TAM) which leverages temporal information in video and triggers update of the one-shot query when the temporal consistency is disrupted. The proposed CD-OSVOD utilizes base models trained on COCO and VOC still image datasets and successfully adapts to the video domain for novel object classes. Performance evaluations on challenging VOT-LT benchmarking video dataset demonstrate significant improvement in AP50 and mAP scores, highlighting the effectiveness of the proposed domain adaptation approach. | |
| dc.identifier.citation | Hanoğlu, Y. K., Günsel, B., & Gürkan, F. (2025). Cross-Domain One-Shot Video Object Detection. In 2025 33rd European Signal Processing Conference (EUSIPCO) (pp. 641-645). IEEE. | |
| dc.identifier.doi | 10.23919/EUSIPCO63237.2025.11226448 | |
| dc.identifier.endpage | 645 | |
| dc.identifier.scopus | 2-s2.0-105029830695 | |
| dc.identifier.scopusquality | Q3 | |
| dc.identifier.startpage | 641 | |
| dc.identifier.uri | https://doi.org/10.23919/EUSIPCO63237.2025.11226448 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.14730/15646 | |
| dc.identifier.wos | WOS:001831889700129 | |
| dc.identifier.wosquality | N/A | |
| dc.indekslendigikaynak | Web of Science | |
| dc.indekslendigikaynak | Scopus | |
| dc.language.iso | en | |
| dc.publisher | IEEE | |
| dc.relation.ispartof | 33rd European Signal Processing Conference, EUSIPCO | |
| dc.relation.publicationcategory | Konferans Öğesi - Uluslararası - Kurum Öğretim Elemanı | |
| dc.rights | info:eu-repo/semantics/openAccess | |
| dc.subject | Cross-domain learning | |
| dc.subject | Video object detection | |
| dc.title | Cross-domain One-shot Video Object Detection | |
| dc.type | Conference Object |










