English
Related papers

Related papers: Masked Contrastive Pre-Training Improves Music Aud…

200 papers

People exploit the predictability of lexical structures during text comprehension. Though predictable structure is also present in speech, the degree to which prosody, e.g. intonation, tempo, and loudness, contributes to such structure…

Computation and Language · Computer Science 2025-06-04 Sarenne Wallbridge , Christoph Minixhofer , Catherine Lai , Peter Bell

Speaker verification system trained on one domain usually suffers performance degradation when applied to another domain. To address this challenge, researchers commonly use feature distribution matching-based methods in unsupervised domain…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-23 Wen Huang , Bing Han , Zhengyang Chen , Shuai Wang , Yanmin Qian

Recent methods in self-supervised learning have demonstrated that masking-based pretext tasks extend beyond NLP, serving as useful pretraining objectives in computer vision. However, existing approaches apply random or ad hoc masking…

Computer Vision and Pattern Recognition · Computer Science 2022-12-19 Dylan Sam , Min Bai , Tristan McKinney , Li Erran Li

Multi-pitch estimation is a decades-long research problem involving the detection of pitch activity associated with concurrent musical events within multi-instrument mixtures. Supervised learning techniques have demonstrated solid…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-27 Frank Cwitkowitz , Zhiyao Duan

Audio-based music structure analysis (MSA) is an essential task in Music Information Retrieval that remains challenging due to the complexity and variability of musical form. Recent advances highlight the potential of fine-tuning…

Sound · Computer Science 2025-07-21 Yixiao Zhang , Haonan Chen , Ju-Chiang Wang , Jitong Chen

Contrastive learning (CL) for Vision Transformers (ViTs) in image domains has achieved performance comparable to CL for traditional convolutional backbones. However, in 3D point cloud pretraining with ViTs, masked autoencoder (MAE) modeling…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Bin Ren , Guofeng Mei , Danda Pani Paudel , Weijie Wang , Yawei Li , Mengyuan Liu , Rita Cucchiara , Luc Van Gool , Nicu Sebe

Video Question Answering (Video QA) requires fine-grained understanding of both video and language modalities to answer the given questions. In this paper, we propose novel training schemes for multiple-choice video question answering with…

Computation and Language · Computer Science 2020-12-15 Seonhoon Kim , Seohyeong Jeong , Eunbyul Kim , Inho Kang , Nojun Kwak

Annotating musical beats is a very long and tedious process. In order to combat this problem, we present a new self-supervised learning pretext task for beat tracking and downbeat estimation. This task makes use of Spleeter, an audio source…

Sound · Computer Science 2023-07-18 Dorian Desblancs

Masked Autoencoders (MAE) based on a reconstruction task have risen to be a promising paradigm for self-supervised learning (SSL) and achieve state-of-the-art performance across different benchmark datasets. However, despite its impressive…

Machine Learning · Computer Science 2023-03-28 Qi Zhang , Yifei Wang , Yisen Wang

Content creators often use music to enhance their videos, from soundtracks in movies to background music in video blogs and social media content. However, identifying the best music for a video can be a difficult and time-consuming task. To…

Multimedia · Computer Science 2024-12-24 Shanti Stewart , Gouthaman KV , Lie Lu , Andrea Fanelli

A great challenge in speaker representation learning using deep models is to design learning objectives that can enhance the discrimination of unseen speakers under unseen domains. This work proposes a supervised contrastive learning…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-18 Zhe Li , Man-Wai Mak

Music auto-tagging is crucial for enhancing music discovery and recommendation. Existing models in Music Information Retrieval (MIR) struggle with real-world noise such as environmental and speech sounds in multimedia content. This study…

Sound · Computer Science 2024-01-30 Haesun Joung , Kyogu Lee

We study a novel neural architecture and its training strategies of speaker encoder for speaker recognition without using any identity labels. The speaker encoder is trained to extract a fixed-size speaker embedding from a spoken utterance…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-28 Ruijie Tao , Kong Aik Lee , Rohan Kumar Das , Ville Hautamäki , Haizhou Li

In this study, for the first time, we extensively investigate whether music foundation models (MFMs) or speech foundation models (SFMs) work better for singing voice deepfake detection (SVDD), which has recently attracted attention in the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-24 Orchid Chetia Phukan , Sarthak Jain , Swarup Ranjan Behera , Arun Balaji Buduru , Rajesh Sharma , S. R Mahadeva Prasanna

Depression detection research has increased over the last few decades, one major bottleneck of which is the limited data availability and representation learning. Recently, self-supervised learning has seen success in pretraining text…

Human-Computer Interaction · Computer Science 2021-10-29 Pingyue Zhang , Mengyue Wu , Heinrich Dinkel , Kai Yu

Self-supervised pretraining has been shown to yield powerful representations for transfer learning. These performance gains come at a large computational cost however, with state-of-the-art methods requiring an order of magnitude more…

Computer Vision and Pattern Recognition · Computer Science 2021-08-06 Olivier J. Hénaff , Skanda Koppula , Jean-Baptiste Alayrac , Aaron van den Oord , Oriol Vinyals , João Carreira

Learning-based synthetic multi-contrast MRI commonly involves deep models trained using high-quality images of source and target contrasts, regardless of whether source and target domain samples are paired or unpaired. This results in…

Image and Video Processing · Electrical Eng. & Systems 2021-05-13 Mahmut Yurt , Salman Ul Hassan Dar , Muzaffer Özbey , Berk Tınaz , Kader Karlı Oğuz , Tolga Çukur

Audio self-supervised learning (SSL) pre-training, which aims to learn good representations from unlabeled audio, has made remarkable progress. However, the extensive computational demands during pre-training pose a significant barrier to…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-09 Wenxi Chen , Yuzhe Liang , Ziyang Ma , Zhisheng Zheng , Xie Chen

Music structure analysis (MSA) methods traditionally search for musically meaningful patterns in audio: homogeneity, repetition, novelty, and segment-length regularity. Hand-crafted audio features such as MFCCs or chromagrams are often used…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-03 Ju-Chiang Wang , Jordan B. L. Smith , Wei-Tsung Lu , Xuchen Song

Machine learning has revolutionized the modeling of clinical timeseries data. Using machine learning, a Deep Neural Network (DNN) can be automatically trained to learn a complex mapping of its input features for a desired task. This is…

Machine Learning · Computer Science 2024-10-15 Ryan King , Shivesh Kodali , Conrad Krueger , Tianbao Yang , Bobak J. Mortazavi
‹ Prev 1 3 4 5 6 7 10 Next ›