中文
相关论文

相关论文: On feature representations for marmoset vocal comm…

200 篇论文

Self-supervised learning (SSL) models use only the intrinsic structure of a given signal, independent of its acoustic domain, to extract essential information from the input to an embedding space. This implies that the utility of such…

机器学习 · 计算机科学 2023-06-09 Eklavya Sarkar , Mathew Magimai. -Doss

Understanding evolution of vocal communication in social animals is an important research problem. In that context, beyond humans, there is an interest in analyzing vocalizations of other social animals such as, meerkats, marmosets, apes.…

音频与语音处理 · 电气工程与系统科学 2024-08-29 Imen Ben Mahmoud , Eklavya Sarkar , Marta Manser , Mathew Magimai. -Doss

Marmoset monkeys encode vital information in their calls and serve as a surrogate model for neuro-biologists to understand the evolutionary origins of human vocal communication. Traditionally analyzed with signal processing-based features,…

声音 · 计算机科学 2024-07-25 Eklavya Sarkar , Mathew Magimai. -Doss

Enhancing explainability in speech self-supervised learning (SSL) is important for developing reliable SSL-based speech processing systems. This study probes how speech SSL models encode speaker-specific information via a large-scale…

音频与语音处理 · 电气工程与系统科学 2026-03-06 Aemon Yat Fei Chiu , Kei Ching Fung , Roger Tsz Yeung Li , Jingyu Li , Tan Lee

Self-supervised language models are very effective at predicting high-level cortical responses during language comprehension. However, the best current models of lower-level auditory processing in the human brain rely on either…

计算与语言 · 计算机科学 2022-05-31 Aditya R. Vaidya , Shailee Jain , Alexander G. Huth

Automatic singing voice understanding tasks, such as singer identification, singing voice transcription, and singing technique classification, benefit from data-driven approaches that utilize deep learning techniques. These approaches work…

声音 · 计算机科学 2023-09-06 Yuya Yamamoto

Self-supervised learning (SSL) based models have been shown to generate powerful representations that can be used to improve the performance of downstream speech tasks. Several state-of-the-art SSL models are available, and each of these…

计算与语言 · 计算机科学 2023-02-21 A Arunkumar , Vrunda N Sukhadia , S. Umesh

In our previous work, we derived the acoustic features, that contribute to the perception of warmth and competence in synthetic speech. As an extension, in our current work, we investigate the impact of the derived vocal features in the…

声音 · 计算机科学 2022-04-05 Sai Sirisha Rallabandi , Sebastian Möller

The marmoset, a highly vocal primate, is a key model for studying social-communicative behavior. Unlike human speech, marmoset vocalizations are less structured, highly variable, and recorded in noisy, low-resource conditions. Learning…

声音 · 计算机科学 2025-08-13 Bin Wu , Shinnosuke Takamichi , Sakriani Sakti , Satoshi Nakamura

In this work, we study the features extracted by English self-supervised learning (SSL) models in cross-lingual contexts and propose a new metric to predict the quality of feature representations. Using automatic speech recognition (ASR) as…

计算与语言 · 计算机科学 2023-11-28 Shuyue Stella Li , Beining Xu , Xiangyu Zhang , Hexin Liu , Wenhan Chao , Leibny Paola Garcia

Articulatory distinctive features, as well as phonetic transcription, play important role in speech-related tasks: computer-assisted pronunciation training, text-to-speech conversion (TTS), studying speech production mechanisms, speech…

音频与语音处理 · 电气工程与系统科学 2019-07-04 Ievgen Karaulov , Dmytro Tkanov

Marmoset monkeys exhibit complex vocal communication, challenging the view that nonhuman primates vocal communication is entirely innate, and show similar features of human speech, such as vocal labeling of others and turn-taking. Studying…

计算与语言 · 计算机科学 2025-09-16 Talia Sternberg , Michael London , David Omer , Yossi Adi

The goal of our research is to automatically retrieve the satisfaction and the frustration in real-life call-center conversations. This study focuses an industrial application in which the customer satisfaction is continuously tracked down…

音频与语音处理 · 电气工程与系统科学 2023-10-10 Manon Macary , Marie Tahon , Yannick Estève , Daniel Luzzati

In this work, we propose a novel method for modeling numerous speakers, which enables expressing the overall characteristics of speakers in detail like a trained multi-speaker model without additional training on the target speaker's…

声音 · 计算机科学 2024-06-03 Jungil Kong , Junmo Lee , Jeongmin Kim , Beomjeong Kim , Jihoon Park , Dohee Kong , Changheon Lee , Sangjin Kim

Transformer based end-to-end modelling approaches with multiple stream inputs have been achieved great success in various automatic speech recognition (ASR) tasks. An important issue associated with such approaches is that the intermediate…

音频与语音处理 · 电气工程与系统科学 2022-07-11 Jin Li , Rongfeng Su , Xurong Xie , Nan Yan , Lan Wang

Masked latent prediction has emerged as a leading paradigm in self-supervised learning (SSL), especially for general audio and music representation learning. While recent methods have demonstrated strong performance, the role of the…

声音 · 计算机科学 2025-08-19 Aurian Quelennec , Pierre Chouteau , Geoffroy Peeters , Slim Essid

Audio deepfake model attribution aims to mitigate the misuse of synthetic speech by identifying the source model responsible for generating a given audio sample, enabling accountability and informing vendors. The task is challenging, but…

音频与语音处理 · 电气工程与系统科学 2026-03-17 Gabriel Pîrlogeanu , Adriana Stan , Horia Cucu

While FastSpeech2 aims to integrate aspects of speech such as pitch, energy, and duration as conditional inputs, it still leaves scope for richer representations. As a part of this work, we leverage representations from various…

计算与语言 · 计算机科学 2023-08-03 Ramanan Sivaguru , Vasista Sai Lodagala , S Umesh

To understand why self-supervised learning (SSL) models have empirically achieved strong performances on several speech-processing downstream tasks, numerous studies have focused on analyzing the encoded information of the SSL layer…

音频与语音处理 · 电气工程与系统科学 2024-06-07 Jialu Li , Mark Hasegawa-Johnson , Nancy L. McElwain

While supervised quality predictors for synthesized speech have demonstrated strong correlations with human ratings, their requirement for in-domain labeled training data hinders their generalization ability to new domains. Unsupervised…

音频与语音处理 · 电气工程与系统科学 2025-09-08 Erica Cooper , Takuma Okamoto , Yamato Ohtani , Tomoki Toda , Hisashi Kawai
‹ 上一页 1 2 3 10 下一页 ›