English
Related papers

Related papers: A Multimodal Prototypical Approach for Unsupervise…

200 papers

Text-based voice editing (TBVE) uses synthetic output from text-to-speech (TTS) systems to replace words in an original recording. Recent work has used neural models to produce edited speech that is similar to the original speech in terms…

Sound · Computer Science 2022-10-31 Jason Fong , Yun Wang , Prabhav Agrawal , Vimal Manohar , Jilong Wu , Thilo Köhler , Qing He

Tag-based music retrieval is crucial to browse large-scale music libraries efficiently. Hence, automatic music tagging has been actively explored, mostly as a classification task, which has an inherent limitation: a fixed vocabulary. On the…

Information Retrieval · Computer Science 2020-11-02 Minz Won , Sergio Oramas , Oriol Nieto , Fabien Gouyon , Xavier Serra

Zero-shot learning (ZSL) makes object recognition in images possible in absence of visual training data for a part of the classes from a dataset. When the number of classes is large, classes are usually represented by semantic class…

Computer Vision and Pattern Recognition · Computer Science 2020-08-10 Yannick Le Cacheux , Adrian Popescu , Hervé Le Borgne

Low-light image enhancement is a crucial visual task, and many unsupervised methods tend to overlook the degradation of visible information in low-light scenes, which adversely affects the fusion of complementary information and hinders the…

Computer Vision and Pattern Recognition · Computer Science 2024-02-05 Xiaofeng Zhang , Zishan Xu , Hao Tang , Chaochen Gu , Wei Chen , Shanying Zhu , Xinping Guan

Recent advances in large pretrained language models have increased attention to zero-shot text classification. In particular, models finetuned on natural language inference datasets have been widely adopted as zero-shot classifiers due to…

Computation and Language · Computer Science 2022-11-01 Ariel Gera , Alon Halfon , Eyal Shnarch , Yotam Perlitz , Liat Ein-Dor , Noam Slonim

Recent work has demonstrated that pre-trained language models (PLMs) are zero-shot learners. However, most existing zero-shot methods involve heavy human engineering or complicated self-training pipelines, hindering their application to new…

Computation and Language · Computer Science 2022-11-24 Yu Fei , Ping Nie , Zhao Meng , Roger Wattenhofer , Mrinmaya Sachan

Text and vision foundation models can perform many tasks in a zero-shot setting, a desirable property that enables these systems to be applied in general and low-resource settings. There has been far less work, however, on the zero-shot…

Computation and Language · Computer Science 2024-03-29 Rao Ma , Adian Liusie , Mark J. F. Gales , Kate M. Knill

In this paper, we propose a novel end-to-end user-defined keyword spotting method that utilizes linguistically corresponding patterns between speech and text sequences. Unlike previous approaches requiring speech keyword enrollment, our…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-04 Hyeon-Kyeong Shin , Hyewon Han , Doyeon Kim , Soo-Whan Chung , Hong-Goo Kang

Eliminating the negative effect of non-stationary environmental noise is a long-standing research topic for automatic speech recognition that stills remains an important challenge. Data-driven supervised approaches, including ones based on…

This paper tackles automatically discovering phone-like acoustic units (AUD) from unlabeled speech data. Past studies usually proposed single-step approaches. We propose a two-stage approach: the first stage learns a subword-discriminative…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-08 Siyuan Feng , Piotr Żelasko , Laureano Moro-Velázquez , Odette Scharenborg

One of the key factors of enabling machine learning models to comprehend and solve real-world tasks is to leverage multimodal data. Unfortunately, annotation of multimodal data is challenging and expensive. Recently, self-supervised…

Computer Vision and Pattern Recognition · Computer Science 2020-12-11 Elad Amrani , Rami Ben-Ari , Daniel Rotman , Alex Bronstein

Text classification of unseen classes is a challenging Natural Language Processing task and is mainly attempted using two different types of approaches. Similarity-based approaches attempt to classify instances based on similarities between…

Computation and Language · Computer Science 2023-07-25 Tim Schopf , Daniel Braun , Florian Matthes

This paper proposes a new architecture for speaker adaptation of multi-speaker neural-network speech synthesis systems, in which an unseen speaker's voice can be built using a relatively small amount of speech data without transcriptions.…

Audio and Speech Processing · Electrical Eng. & Systems 2018-08-21 Hieu-Thi Luong , Junichi Yamagishi

Approximately 1.2% of the world's population has impaired voice production. As a result, automatic dysphonic voice detection has attracted considerable academic and clinical interest. However, existing methods for automated voice assessment…

Sound · Computer Science 2023-01-27 Jianwei Zhang , Julie Liss , Suren Jayasuriya , Visar Berisha

There has been a recent spike in interest in multi-modal Language and Vision problems. On the language side, most of these models primarily focus on English since most multi-modal datasets are monolingual. We try to bridge this gap with a…

Computation and Language · Computer Science 2020-12-10 Pranav Aggarwal , Ajinkya Kale

Human perception has the unique ability to focus on specific events in a mixture of signals--a challenging task for existing non-intrusive assessment methods. In this work, we introduce semi-intrusive assessment that emulates human…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-23 Jozef Coldenhoff , Milos Cernak

We explore self-supervised models that can be potentially deployed on mobile devices to learn general purpose audio representations. Specifically, we propose methods that exploit the temporal context in the spectrogram domain. One method…

Audio and Speech Processing · Electrical Eng. & Systems 2019-05-29 Marco Tagliasacchi , Beat Gfeller , Félix de Chaumont Quitry , Dominik Roblek

Decoding visual semantic representations from human brain activity is a significant challenge. While recent zero-shot decoding approaches have improved performance by leveraging aligned image-text datasets, they overlook a fundamental…

Neurons and Cognition · Quantitative Biology 2026-01-21 Zhengdi Zhang , Hao Zhang , Wenjun Xia

Multimodal learning allows us to leverage information from multiple sources (visual, acoustic and text), similar to our experience of the real world. However, it is currently unclear to what extent auxiliary modalities improve performance…

Computation and Language · Computer Science 2020-01-01 Tejas Srinivasan , Ramon Sanabria , Florian Metze

Unsupervised discovery of acoustic tokens from audio corpora without annotation and learning vector representations for these tokens have been widely studied. Although these techniques have been shown successful in some applications such as…

Computation and Language · Computer Science 2018-04-03 Da-Rong Liu , Kuan-Yu Chen , Hung-Yi Lee , Lin-shan Lee