中文
相关论文

相关论文: Deep Learning-Based Single-Ended Objective Quality…

200 篇论文

This paper studies the generalization of the targeted minimum loss-based estimation (TMLE) framework to estimation of effects of time-varying interventions in settings where both interventions, covariates, and outcome can happen at…

统计理论 · 数学 2021-05-06 Helene C. Rytgaard , Thomas A. Gerds , Mark J. van der Laan

We introduce DeSTA2.5-Audio, a general-purpose Large Audio Language Model (LALM) designed for robust auditory perception and instruction-following. Recent LALMs augment Large Language Models (LLMs) with auditory capabilities by training on…

Distance estimation from audio plays a crucial role in various applications, such as acoustic scene analysis, sound source localization, and room modeling. Most studies predominantly center on employing a classification approach, where…

音频与语音处理 · 电气工程与系统科学 2024-03-27 Michael Neri , Archontis Politis , Daniel Krause , Marco Carli , Tuomas Virtanen

Majority of the recent approaches for text-independent speaker recognition apply attention or similar techniques for aggregation of frame-level feature descriptors generated by a deep neural network (DNN) front-end. In this paper, we…

声音 · 计算机科学 2019-10-22 Sarthak Yadav , Atul Rai

The goal of this paper is to enhance Text-to-Audio generation at inference, focusing on generating realistic audio that precisely aligns with text prompts. Despite the rapid advancements, existing models often fail to achieve a reliable…

音频与语音处理 · 电气工程与系统科学 2025-09-25 Jaemin Jung , Jaehun Kim , Inkyu Shin , Joon Son Chung

The analysis of the structure of musical pieces is a task that remains a challenge for Artificial Intelligence, especially in the field of Deep Learning. It requires prior identification of structural boundaries of the music pieces. This…

音频与语音处理 · 电气工程与系统科学 2021-12-02 Carlos Hernandez-Olivan , Jose R. Beltran , David Diaz-Guerra

Language-queried audio source separation (LASS) aims to separate an audio source guided by a text query, with the signal-to-distortion ratio (SDR)-based metrics being commonly used to objectively measure the quality of the separated audio.…

声音 · 计算机科学 2025-01-07 Feiyang Xiao , Jian Guan , Qiaoxi Zhu , Xubo Liu , Wenbo Wang , Shuhan Qi , Kejia Zhang , Jianyuan Sun , Wenwu Wang

Environmental sound detection is a challenging application of machine learning because of the noisy nature of the signal, and the small amount of (labeled) data that is typically available. This work thus presents a comparison of several…

声音 · 计算机科学 2017-03-22 Juncheng Li , Wei Dai , Florian Metze , Shuhui Qu , Samarjit Das

Audio tagging aims to predict one or several labels in an audio clip. Many previous works use weakly labelled data (WLD) for audio tagging, where only presence or absence of sound events is known, but the order of sound events is unknown.…

声音 · 计算机科学 2018-08-07 Yuanbo Hou , Qiuqiang Kong , Shengchen Li

Accurate audio quality estimation is essential for developing and evaluating audio generation, retrieval, and enhancement systems. Existing non-intrusive assessment models predict a single Mean Opinion Score (MOS) for speech, merging…

音频与语音处理 · 电气工程与系统科学 2026-01-13 Yi-Cheng Lin , Jia-Hung Chen , Hung-yi Lee

Speech audio quality is subject to degradation caused by an acoustic environment and isotropic ambient and point noises. The environment can lead to decreased speech intelligibility and loss of focus and attention by the listener. Basic…

音频与语音处理 · 电气工程与系统科学 2022-04-05 Paula Sánchez López , Paul Callens , Milos Cernak

Sound matching algorithms seek to approximate a target waveform by parametric audio synthesis. Deep neural networks have achieved promising results in matching sustained harmonic tones. However, the task is more challenging when targets are…

声音 · 计算机科学 2023-03-14 Han Han , Vincent Lostanlen , Mathieu Lagrange

Deep learning has achieved impressive prediction performance in the field of sequence learning recently. Dissolved oxygen prediction, as a kind of time-series forecasting, is suitable for this technique. Although many researchers have…

信号处理 · 电气工程与系统科学 2019-11-22 Hongqian Qin

Neural network models for audio tasks, such as automatic speech recognition (ASR) and acoustic scene classification (ASC), are susceptible to noise contamination for real-life applications. To improve audio quality, an enhancement module,…

Currently, high-quality, synchronized audio is synthesized from video and optional text inputs using various multi-modal joint learning frameworks. However, the precise alignment between the visual and generated audio domains remains far…

声音 · 计算机科学 2025-03-31 Yunming Liang , Zihao Chen , Chaofan Ding , Xinhan Di

Remixing separated audio sources trades off interferer attenuation against the amount of audible deteriorations. This paper proposes a non-intrusive audio quality estimation method for controlling this trade-off in a signal-adaptive manner.…

音频与语音处理 · 电气工程与系统科学 2023-03-24 Matteo Torcoli , Jouni Paulus , Thorsten Kastner , Christian Uhle

Particle beam microscopy (PBM) performs nanoscale imaging by pixelwise capture of scalar values representing noisy measurements of the response from secondary electrons (SEs) integrated over a dwell time. Extended to metrology, goals…

数据分析、统计与概率 · 物理学 2023-08-15 Akshay Agarwal , Minxu Peng , Vivek K. Goyal

Human beings can perceive a target sound type from a multi-source mixture signal by the selective auditory attention, however, such functionality was hardly ever explored in machine hearing. This paper addresses the target sound detection…

声音 · 计算机科学 2022-07-08 Dongchao Yang , Helin Wang , Yuexian Zou , Fan Cui , Yujun Wang

Discrete audio tokens are compact representations that aim to preserve perceptual quality, phonetic content, and speaker characteristics while enabling efficient storage and inference, as well as competitive performance across diverse…

Audio quality assessment is critical for assessing the perceptual realism of sounds. However, the time and expense of obtaining ''gold standard'' human judgments limit the availability of such data. For AR&VR, good perceived sound quality…

音频与语音处理 · 电气工程与系统科学 2022-06-27 Pranay Manocha , Anurag Kumar , Buye Xu , Anjali Menon , Israel D. Gebru , Vamsi K. Ithapu , Paul Calamia