中文
相关论文

相关论文: Soundscape Captioning using Sound Affective Qualit…

200 篇论文

Zero-shot audio captioning aims at automatically generating descriptive textual captions for audio content without prior training for this task. Different from speech recognition which translates audio content that contains spoken language…

音频与语音处理 · 电气工程与系统科学 2023-11-15 Leonard Salewski , Stefan Fauth , A. Sophia Koepke , Zeynep Akata

Large audio-language models (LALMs) have achieved near-human performance in sentence-level transcription and emotion recognition. However, existing evaluations focus mainly on surface-level perception, leaving the capacity of models for…

计算与语言 · 计算机科学 2025-08-05 Wanqi Yang , Yanda Li , Yunchao Wei , Meng Fang , Ling Chen

Aesthetic image captioning (AIC) refers to the multi-modal task of generating critical textual feedbacks for photographs. While in natural image captioning (NIC), deep models are trained in an end-to-end manner using large curated datasets…

计算机视觉与模式识别 · 计算机科学 2019-08-30 Koustav Ghosal , Aakanksha Rana , Aljosa Smolic

Recent literature uses language to build foundation models for audio. These Audio-Language Models (ALMs) are trained on a vast number of audio-text pairs and show remarkable performance in tasks including Text-to-Audio Retrieval,…

In this paper, we propose SemanticAC, a semantics-assisted framework for Audio Classification to better leverage the semantic information. Unlike conventional audio classification methods that treat class labels as discrete vectors, we…

声音 · 计算机科学 2023-02-14 Yicheng Xiao , Yue Ma , Shuyan Li , Hantao Zhou , Ran Liao , Xiu Li

Speech language models have recently demonstrated great potential as universal speech processing systems. Such models have the ability to model the rich acoustic information existing in audio signals, beyond spoken content, such as emotion,…

声音 · 计算机科学 2025-01-16 Gallil Maimon , Amit Roth , Yossi Adi

State-of-The-Art (SoTA) image captioning models are often trained on the MicroSoft Common Objects in Context (MS-COCO) dataset, which contains human-annotated captions with an average length of approximately ten tokens. Although effective…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Luigi Celona , Simone Bianco , Marco Donzella , Paolo Napoletano

Automated Audio Captioning (AAC) involves generating natural language descriptions of audio content, using encoder-decoder architectures. An audio encoder produces audio embeddings fed to a decoder, usually a Transformer decoder, for…

声音 · 计算机科学 2023-09-04 Étienne Labbé , Thomas Pellegrini , Julien Pinquier

Video captioning, i.e. the task of generating captions from video sequences creates a bridge between the Natural Language Processing and Computer Vision domains of computer science. The task of generating a semantically accurate description…

计算机视觉与模式识别 · 计算机科学 2023-10-17 Md. Mushfiqur Rahman , Thasin Abedin , Khondokar S. S. Prottoy , Ayana Moshruba , Fazlul Hasan Siddiqui

Emotion recognition from human speech is a critical enabler for socially aware conversational AI. However, while most prior work frames emotion recognition as a categorical classification problem, real-world affective states are often…

声音 · 计算机科学 2026-02-05 Hong Jia , Weibin Li , Jingyao Wu , Xiaofeng Yu , Yan Gao , Jintao Cheng , Xiaoyu Tang , Feng Xia , Ting Dang

Acoustic Scene Classification (ASC) identifies an environment based on an audio signal. This paper explores ASC in low-resource conditions and proposes a novel model, DS-FlexiNet, which combines depthwise separable convolutions from…

音频与语音处理 · 电气工程与系统科学 2025-04-29 Zhi Chen , Yun-Fei Shao , Yong Ma , Mingsheng Wei , Le Zhang , Wei-Qiang Zhang

Image captioning is a significant field across computer vision and natural language processing. We propose and present AIC-AB NET, a novel Attribute-Information-Combined Attention-Based Network that combines spatial attention architecture…

计算机视觉与模式识别 · 计算机科学 2023-07-17 Guoyun Tu , Ying Liu , Vladimir Vlassov

In real-world environments, background noise significantly degrades the intelligibility and clarity of human speech. Audio-visual speech enhancement (AVSE) attempts to restore speech quality, but existing methods often fall short,…

音频与语音处理 · 电气工程与系统科学 2024-02-27 Tassadaq Hussain , Kia Dashtipour , Yu Tsao , Amir Hussain

In the last decade, soundscapes have become one of the most active topics in Acoustics, providing a holistic approach to the acoustic environment, which involves human perception and context. Soundscapes-elicited emotions are central and…

声音 · 计算机科学 2022-07-27 R. San Millán-Castillo , L. Martino , E. Morgado , F. Llorente

In this paper, we presents a low-complexity deep learning frameworks for acoustic scene classification (ASC). The proposed framework can be separated into three main steps: Front-end spectrogram extraction, back-end classification, and late…

声音 · 计算机科学 2021-06-17 Lam Pham , Hieu Tang , Anahid Jalali , Alexander Schindler , Ross King

Context-aware processing mechanisms have increasingly become a critical area of exploration for improving the semantic and contextual capabilities of language generation models. The Context-Aware Semantic Recomposition Mechanism (CASRM) was…

计算与语言 · 计算机科学 2025-03-27 Richard Katrix , Quentin Carroway , Rowan Hawkesbury , Matthias Heathfield

We present an iVector based Acoustic Scene Classification (ASC) system suited for real life settings where active foreground speech can be present. In the proposed system, each recording is represented by a fixed-length iVector that models…

音频与语音处理 · 电气工程与系统科学 2021-08-03 Siyuan Song , Brecht Desplanques , Celest De Moor , Kris Demuynck , Nilesh Madhu

Multimodal Large Language Models (MLLMs) have been widely applied in speech and music. This tendency has led to a focus on audio tokenization for Large Models (LMs). Unlike semantic-only text tokens, audio tokens must both capture global…

声音 · 计算机科学 2025-09-05 Lu Wang , Hao Chen , Siyu Wu , Zhiyue Wu , Hao Zhou , Chengfeng Zhang , Ting Wang , Haodi Zhang

Cross-lingual aspect-based sentiment analysis (ABSA) involves detailed sentiment analysis in a target language by transferring knowledge from a source language with available annotated data. Most existing methods depend heavily on often…

计算与语言 · 计算机科学 2025-08-14 Jakub Šmíd , Pavel Přibáň , Pavel Král

Real-world sound scenes consist of time-varying collections of sound sources, each generating characteristic sound events that are mixed together in audio recordings. The association of these constituent sound events with their mixture and…