English
Related papers

Related papers: Oceanship: A Large-Scale Dataset for Underwater Au…

200 papers

The demand for realistic virtual immersive audio continues to grow, with Head-Related Transfer Functions (HRTFs) playing a key role. HRTFs capture how sound reaches our ears, reflecting unique anatomical features and enhancing spatial…

Sound · Computer Science 2026-01-26 Xuyi Hu , Jian Li , Lorenzo Picinali , Aidan O. T. Hogg

In this paper, we present a novel Amplitude-Modulated Stochastic Perturbation and Vortex Convolutional Network, AMSP-UOD, designed for underwater object detection. AMSP-UOD specifically addresses the impact of non-ideal imaging factors on…

Computer Vision and Pattern Recognition · Computer Science 2024-01-19 Jingchun Zhou , Zongxin He , Kin-Man Lam , Yudong Wang , Weishi Zhang , ChunLe Guo , Chongyi Li

Devices capable of detecting and categorizing acoustic scenes have numerous applications such as providing context-aware user experiences. In this paper, we address the task of characterizing acoustic scenes in a workplace setting from…

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-12 Arindam Jati , Amrutha Nadarajan , Karel Mundnich , Shrikanth Narayanan

In this paper we aim to automatically discover high quality frame-level speech features and acoustic tokens directly from unlabeled speech data. A Multi-granular Acoustic Tokenizer (MAT) was proposed for automatic discovery of multiple sets…

Computation and Language · Computer Science 2017-07-19 Cheng-Tao Chung , Cheng-Yu Tsai , Chia-Hsiang Liu , Lin-Shan Lee

Fine-grained recognition of marine organisms is important for ecological research, biodiversity monitoring, habitat conservation, and evidence-based policy-making. However, many existing approaches primarily rely on object- or ROI-centered…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Donghwan Lee , Byeongjin Kim , Geunhee Kim , Hyukjin Kwon , Nahyeon Maeng , Wooju Kim

Ship detection has been an active and vital topic in the field of remote sensing for a decade, but it is still a challenging problem due to the large scale variations, the high aspect ratios, the intensive arrangement, and the background…

Computer Vision and Pattern Recognition · Computer Science 2020-07-27 Lingyi Liu , Yunpeng Bai , Ying Li

Vision-based semantic segmentation of waterbodies and nearby related objects provides important information for managing water resources and handling flooding emergency. However, the lack of large-scale labeled training and testing datasets…

Computer Vision and Pattern Recognition · Computer Science 2021-11-24 Seyed Mohammad Hassan Erfani , Zhenyao Wu , Xinyi Wu , Song Wang , Erfan Goharian

We introduce EPIC-SOUNDS, a large-scale dataset of audio annotations capturing temporal extents and class labels within the audio stream of the egocentric videos. We propose an annotation pipeline where annotators temporally label…

Sound · Computer Science 2025-07-17 Jaesung Huh , Jacob Chalk , Evangelos Kazakos , Dima Damen , Andrew Zisserman

In this research, we present an innovative, parameter-efficient model that utilizes the attention U-Net architecture for the automatic detection and eradication of non-speech vocal sounds, specifically breath sounds, in vocal recordings.…

Sound · Computer Science 2024-09-10 Nidula Elgiriyewithana , N. D. Kodikara

A methodology based on deep recurrent models for maritime surveillance, over publicly available Automatic Identification System (AIS) data, is presented in this paper. The setup employs a deep Recurrent Neural Network (RNN)-based model, for…

Machine Learning · Computer Science 2024-06-17 Constantine Maganaris , Eftychios Protopapadakis , Nikolaos Doulamis

Background: Underwater images, in general, suffer from low contrast and high color distortions due to the non-uniform attenuation of the light as it propagates through the water. In addition, the degree of attenuation varies with the…

Image and Video Processing · Electrical Eng. & Systems 2022-01-20 Prasen Kumar Sharma , Ira Bisht , Arijit Sur

Adapting pre-trained deep learning models to new and unknown environments remains a major challenge in underwater acoustic localization. We show that although the performance of pre-trained models suffers from mismatch between the training…

Sound · Computer Science 2025-10-14 Dariush Kari , Hari Vishnu , Andrew C. Singer

Closed-Set speaker identification aims to assign a speech utterance to one of a predefined set of enrolled speakers and requires robust modeling of speaker-specific characteristics across multiple temporal scales. While recent deep learning…

Sound · Computer Science 2026-05-11 Yassin Terraf , Youssef Iraqi

The use of conventional neutrino telescope methods and technology for detecting neutrinos with energies above 1 EeV from astrophysical sources would be prohibitively expensive and may turn out to be technically not feasible. Acoustic…

Instrumentation and Methods for Astrophysics · Physics 2019-08-13 Robert Lahmann

Environmental sound scene and sound event recognition is important for the recognition of suspicious events in indoor and outdoor environments (such as nurseries, smart homes, nursing homes, etc.) and is a fundamental task involved in many…

Sound · Computer Science 2023-08-31 Nan Che , Chenrui Liu , Fei Yu

Offloading computationally heavy tasks from an unmanned aerial vehicle (UAV) to a remote server helps improve the battery life and can help reduce resource requirements. Deep learning based state-of-the-art computer vision tasks, such as…

Computer Vision and Pattern Recognition · Computer Science 2023-02-07 Sedat Ozer , Enes Ilhan , Mehmet Akif Ozkanoglu , Hakan Ali Cirpan

Recognizing sounds is a key aspect of computational audio scene analysis and machine perception. In this paper, we advocate that sound recognition is inherently a multi-modal audiovisual task in that it is easier to differentiate sounds…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-03 Haytham M. Fayek , Anurag Kumar

Despite significant advances in recent years, the existing Computer-Assisted Pronunciation Training (CAPT) methods detect pronunciation errors with a relatively low accuracy (precision of 60% at 40%-80% recall). This Ph.D. work proposes…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-15 Daniel Korzekwa

Building a video retrieval system that is robust and reliable, especially for the marine environment, is a challenging task due to several factors such as dealing with massive amounts of dense and repetitive data, occlusion, blurriness, low…

Computer Vision and Pattern Recognition · Computer Science 2023-06-08 Tan-Sang Ha , Hai Nguyen-Truong , Tuan-Anh Vu , Sai-Kit Yeung

Speech activity detection (or endpointing) is an important processing step for applications such as speech recognition, language identification and speaker diarization. Both audio- and vision-based approaches have been used for this task in…

‹ Prev 1 8 9 10 Next ›