中文
相关论文

相关论文: Exploring Meta Information for Audio-based Zero-sh…

200 篇论文

Audio-visual zero-shot learning methods commonly build on features extracted from pre-trained models, e.g. video or audio classification models. However, existing benchmarks predate the popularization of large multi-modal models, such as…

计算机视觉与模式识别 · 计算机科学 2024-04-10 David Kurzendörfer , Otniel-Bogdan Mercea , A. Sophia Koepke , Zeynep Akata

Large audio-language models (LALMs) exhibit strong zero-shot capabilities in multiple downstream tasks, such as audio question answering (AQA) and abstract reasoning; however, these models still lag behind specialized models for certain…

声音 · 计算机科学 2026-03-24 Videet Mehta , Liming Wang , Hilde Kuehne , Rogerio Feris , James R. Glass , M. Jehanzeb Mirza

Recently, end-to-end speech translation (ST) has gained significant attention as it avoids error propagation. However, the approach suffers from data scarcity. It heavily depends on direct ST data and is less efficient in making use of…

计算与语言 · 计算机科学 2022-05-17 Tu Anh Dinh , Danni Liu , Jan Niehues

Accurate estimation of aircraft operations, such as takeoffs and landings, is critical for effective airport management, yet remains challenging, especially at non-towered facilities lacking dedicated surveillance infrastructure. This paper…

声音 · 计算机科学 2025-09-15 Abdullah All Tanvir , Chenyu Huang , Moe Alahmad , Chuyang Yang , Xin Zhong

Language-enabled robots have been widely studied over the past years to enable natural human-robot interaction and teaming in various real-world applications. Language-enabled robots must be able to comprehend referring expressions to…

机器人学 · 计算机科学 2023-12-22 Peng Gao , Ahmed Jaafar , Brian Reily , Christopher Reardon , Hao Zhang

Large audio-video language models can generate descriptions for both video and audio. However, they sometimes ignore audio content, producing audio descriptions solely reliant on visual information. This paper refers to this as audio…

多媒体 · 计算机科学 2024-01-19 Taichi Nishimura , Shota Nakada , Masayoshi Kondo

Anomaly detection has many important applications, such as monitoring industrial equipment. Despite recent advances in anomaly detection with deep-learning methods, it is unclear how existing solutions would perform under…

声音 · 计算机科学 2022-04-06 Bingqing Chen , Luca Bondi , Samarjit Das

Audio classification can distinguish different kinds of sounds, which is helpful for intelligent applications in daily life. However, it remains a challenging task since the sound events in an audio clip is probably multiple, even…

音频与语音处理 · 电气工程与系统科学 2019-11-22 Jiaxu Chen , Jing Hao , Kai Chen , Di Xie , Shicai Yang , Shiliang Pu

Research in bioacoustics, neuroscience, and linguistics often uses birdsong as a proxy to acquire knowledge across diverse areas. This requires audio models to annotate and parse the birdsong. Developing such models requires precise,…

机器学习 · 计算机科学 2026-05-20 Houtan Ghaffari , Lukas Rauch , Paul Devos

Automatic phonemic transcription tools are useful for low-resource language documentation. However, due to the lack of training sets, only a tiny fraction of languages have phonemic transcription tools. Fortunately, multilingual acoustic…

计算与语言 · 计算机科学 2020-02-28 Xinjian Li , Siddharth Dalmia , David R. Mortensen , Juncheng Li , Alan W Black , Florian Metze

The automatic classification of animal sounds presents an enduring challenge in bioacoustics, owing to the diverse statistical properties of sound signals, variations in recording equipment, and prevalent low Signal-to-Noise Ratio (SNR)…

声音 · 计算机科学 2024-07-08 Qiang Yang , Xiuying Chen , Changsheng Ma , Carlos M. Duarte , Xiangliang Zhang

Medical professionals frequently work in a data constrained setting to provide insights across a unique demographic. A few medical observations, for instance, informs the diagnosis and treatment of a patient. This suggests a unique setting…

计算与语言 · 计算机科学 2022-12-06 Pankaj Sharma , Imran Qureshi , Minh Tran

The massive growth of self-supervised learning (SSL) has been witnessed in language, vision, speech, and audio domains over the past few years. While discrete label prediction is widely adopted for other modalities, the state-of-the-art…

音频与语音处理 · 电气工程与系统科学 2022-12-20 Sanyuan Chen , Yu Wu , Chengyi Wang , Shujie Liu , Daniel Tompkins , Zhuo Chen , Furu Wei

This research introduces a novel approach, MBO-NB, that leverages Migrating Birds Optimization (MBO) coupled with Naive Bayes as an internal classifier to address feature selection challenges in text classification having large number of…

神经与进化计算 · 计算机科学 2024-01-22 Cem Kaya , Zeynep Hilal Kilimci , Mitat Uysal , Murat Kaya

In this paper, ensembles of classifiers that exploit several data augmentation techniques and four signal representations for training Convolutional Neural Networks (CNNs) for audio classification are presented and tested on three freely…

音频与语音处理 · 电气工程与系统科学 2021-11-18 Loris Nanni , Gianluca Maguolo , Sheryl Brahnam , Michelangelo Paci

Recently, sound recognition has been used to identify sounds, such as car and river. However, sounds have nuances that may be better described by adjective-noun pairs such as slow car, and verb-noun pairs such as flying insects, which are…

声音 · 计算机科学 2018-01-10 Sebastian Sager , Benjamin Elizalde , Damian Borth , Christian Schulze , Bhiksha Raj , Ian Lane

Political scientists often grapple with data scarcity in text classification. Recently, fine-tuned BERT models and their variants have gained traction as effective solutions to address this issue. In this study, we investigate the potential…

计算与语言 · 计算机科学 2024-11-11 Yu Wang , Wen Qu , Xin Ye

We propose a metadata-aware self-supervised learning~(SSL)~framework useful for fine-grained classification and ecological mapping of bird species around the world. Our framework unifies two SSL strategies: Contrastive Learning~(CL) and…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Srikumar Sastry , Subash Khanal , Aayush Dhakal , Di Huang , Nathan Jacobs

Many applications of speech technology require more and more audio data. Automatic assessment of the quality of the collected recordings is important to ensure they meet the requirements of the related applications. However, effective and…

音频与语音处理 · 电气工程与系统科学 2020-05-19 Qiang Huang , Thomas Hain

This paper introduces a zero-shot sound event classification (ZS-SEC) method to identify sound events that have never occurred in training data. In our previous work, we proposed a ZS-SEC method using sound attribute vectors (SAVs), where a…

声音 · 计算机科学 2023-03-21 Yi-Han Lin , Xunquan Chen , Ryoichi Takashima , Tetsuya Takiguchi