English
Related papers

Related papers: DDSC: Dynamic Dual-Signal Curriculum for Data-Effi…

200 papers

The Detection and Classification of Acoustic Scenes and Events (DCASE) 2019 challenge focuses on audio tagging, sound event detection and spatial localisation. DCASE 2019 consists of five tasks: 1) acoustic scene classification, 2) audio…

Sound · Computer Science 2019-06-11 Qiuqiang Kong , Yin Cao , Turab Iqbal , Yong Xu , Wenwu Wang , Mark D. Plumbley

Sound event detection (SED) is essential for recognizing specific sounds and their temporal locations within acoustic signals. This becomes challenging particularly for on-device applications, where computational resources are limited. To…

Sound · Computer Science 2024-02-07 Yang Xiao , Rohan Kumar Das

In this paper, we propose a new wireless video communication scheme to achieve high-efficiency video transmission over noisy channels. It exploits the idea of model division multiple access (MDMA) and extracts common semantic features…

Multimedia · Computer Science 2023-05-26 Zhicheng Bao , Haotai Liang , Chen Dong , Xiaodong Xu , Geng Liu

Batch selection is crucial for improving both training efficiency and predictive performance in deep multi-label classification (MLC). Existing batch selection methods typically rely on a single metric to assess instance importance and use…

Machine Learning · Computer Science 2026-05-12 Bin Liu , Haoyu Peng , Zhijia Wei , Jiajing Zhang , Grigorios Tsoumakas

Both visual and auditory information are valuable to determine the salient regions in videos. Deep convolution neural networks (CNN) showcase strong capacity in coping with the audio-visual saliency prediction task. Due to various factors…

Computer Vision and Pattern Recognition · Computer Science 2022-08-17 Yingzi Fan , Longfei Han , Yue Zhang , Lechao Cheng , Chen Xia , Di Hu

Humans do not understand individual events in isolation; rather, they generalize concepts within classes and compare them to others. Existing audio-video pre-training paradigms only focus on the alignment of the overall audio-video…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Kaixuan Cong , Yifan Wang , Rongkun Xue , Yuyang Jiang , Yiming Feng , Jing Yang

Most research in synthetic speech detection (SSD) focuses on improving performance on standard noise-free datasets. However, in actual situations, noise interference is usually present, causing significant performance degradation in SSD…

Sound · Computer Science 2024-04-17 Cunhang Fan , Mingming Ding , Jianhua Tao , Ruibo Fu , Jiangyan Yi , Zhengqi Wen , Zhao Lv

Most generative models of audio directly generate samples in one of two domains: time or frequency. While sufficient to express any signal, these representations are inefficient, as they do not utilize existing knowledge of how sound is…

Machine Learning · Computer Science 2020-01-15 Jesse Engel , Lamtharn Hantrakul , Chenjie Gu , Adam Roberts

Unsupervised domain adaptation (UDA) for semantic segmentation aims to adapt a segmentation model trained on the labeled source domain to the unlabeled target domain. Existing methods try to learn domain invariant features while suffering…

Computer Vision and Pattern Recognition · Computer Science 2021-07-28 Li Gao , Jing Zhang , Lefei Zhang , Dacheng Tao

Semi-supervised learning and domain adaptation techniques have drawn increasing attention in the field of domestic sound event detection thanks to the availability of large amounts of unlabeled data and the relative ease to generate…

Sound · Computer Science 2022-08-18 Fang-Ching Chen , Kuan-Dar Chen , Yi-Wen Liu

Addressing the rising concerns of privacy and security, domain adaptation in the dark aims to adapt a black-box source trained model to an unlabeled target domain without access to any source data or source model parameters. The need for…

Computer Vision and Pattern Recognition · Computer Science 2023-08-03 Chowdhury Sadman Jahan , Andreas Savakis

Image classification datasets exhibit a non-negligible fraction of mislabeled examples, often due to human error when one class superficially resembles another. This issue poses challenges in supervised contrastive learning (SCL), where the…

Computer Vision and Pattern Recognition · Computer Science 2023-11-29 Zijun Long , George Killick , Lipeng Zhuang , Richard McCreadie , Gerardo Aragon Camarasa , Paul Henderson

Most existing scene text detectors require large-scale training data which cannot scale well due to two major factors: 1) scene text images often have domain-specific distributions; 2) collecting large-scale annotated scene text images is…

Computer Vision and Pattern Recognition · Computer Science 2022-12-02 Zichen Tian , Chuhui Xue , Jingyi Zhang , Shijian Lu

Cross-domain sequential recommendation (CDSR) aims to uncover and transfer users' sequential preferences across multiple recommendation domains. While significant endeavors have been made, they primarily concentrated on developing advanced…

Information Retrieval · Computer Science 2024-08-22 Mingjia Yin , Hao Wang , Wei Guo , Yong Liu , Zhi Li , Sirui Zhao , Zhen Wang , Defu Lian , Enhong Chen

Convolutional neural networks (CNNs) have achieved exciting performance in joint segmentation of optic disc and optic cup on single-institution datasets. However, their clinical translation is hindered by two major challenges: limited…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Yusong Xiao , Yuxuan Wu , Li Xiao , Gang Qu , Haiye Huo , Yu-Ping Wang

Acoustic scene classification (ASC) has been approached in the last years using deep learning techniques such as convolutional neural networks or recurrent neural networks. Many state-of-the-art solutions are based on image classification…

Purpose: Segmentation of surgical instruments in endoscopic videos is essential for automated surgical scene understanding and process modeling. However, relying on fully supervised deep learning for this task is challenging because manual…

Computer Vision and Pattern Recognition · Computer Science 2021-03-03 Manish Sahu , Anirban Mukhopadhyay , Stefan Zachow

Deep Joint Source-Channel Coding (Deep-JSCC) has emerged as a promising semantic communication approach for wireless image transmission by jointly optimizing source and channel coding using deep learning techniques. However, traditional…

Networking and Internet Architecture · Computer Science 2025-07-29 Avi Deb Raha , Apurba Adhikary , Mrityunjoy Gain , Yumin Park , Walid Saad , Choong Seon Hong

Like humans, deep networks have been shown to learn better when samples are organized and introduced in a meaningful order or curriculum. Conventional curriculum learning schemes introduce samples in their order of difficulty. This forces…

Computer Vision and Pattern Recognition · Computer Science 2020-08-14 Madan Ravi Ganesh , Jason J. Corso

This thesis focuses on dealing with the task of acoustic scene classification (ASC), and then applied the techniques developed for ASC to a real-life application of detecting respiratory disease. To deal with ASC challenges, this thesis…

Sound · Computer Science 2021-07-21 Lam Pham