中文
相关论文

相关论文: NISQA: A Deep CNN-Self-Attention Model for Multidi…

200 篇论文

Data-driven speech enhancement employing deep neural networks (DNNs) can provide state-of-the-art performance even in the presence of non-stationary noise. During the training process, most of the speech enhancement neural networks are…

音频与语音处理 · 电气工程与系统科学 2021-04-01 Ziyi Xu , Maximilian Strake , Tim Fingscheidt

No-Reference Image Quality Assessment (NR-IQA) remains a challenging task due to the diversity of distortions and the lack of large annotated datasets. Many studies have attempted to tackle these challenges by developing more accurate…

计算机视觉与模式识别 · 计算机科学 2024-09-09 Nasim Jamshidi Avanaki , Abhijay Ghildyal , Nabajeet Barman , Saman Zadtootaghaj

Recent findings raise concerns about whether the evaluation of Multiple-Choice Question Answering (MCQA) accurately reflects the comprehension abilities of large language models. This paper explores the concept of choice sensitivity, which…

计算与语言 · 计算机科学 2026-01-13 Gyeongje Cho , Yeonkyoung So , Jaejin Lee

In spoken question answering, QA systems are designed to answer questions from contiguous text spans within the related speech transcripts. However, the most natural way that human seek or test their knowledge is via human conversations.…

计算与语言 · 计算机科学 2020-10-20 Chenyu You , Nuo Chen , Fenglin Liu , Dongchao Yang , Yuexian Zou

Current disfluency detection models focus on individual utterances each from a single speaker. However, numerous discontinuity phenomena in spoken conversational transcripts occur across multiple turns, hampering human readability and the…

计算与语言 · 计算机科学 2023-10-30 Hua Shen , Vicky Zayats , Johann C. Rocholl , Daniel D. Walker , Dirk Padfield

We propose a deep bilinear model for blind image quality assessment (BIQA) that handles both synthetic and authentic distortions. Our model consists of two convolutional neural networks (CNN), each of which specializes in one distortion…

图像与视频处理 · 电气工程与系统科学 2019-07-08 Weixia Zhang , Kede Ma , Jia Yan , Dexiang Deng , Zhou Wang

The Open Dataset of Audio Quality (ODAQ) was recently introduced to address the scarcity of openly available audio datasets with corresponding subjective quality scores. The dataset, released under permissive licenses, comprises audio…

音频与语音处理 · 电气工程与系统科学 2025-04-02 Sascha Dick , Christoph Thompson , Chih-Wei Wu , Matteo Torcoli , Pablo Delgado , Phillip A. Williams , Emanuel Habets

Classic public switched telephone networks (PSTN) are often a black box for VoIP network providers, as they have no access to performance indicators, such as delay or packet loss. Only the degraded output speech signal can be used to…

音频与语音处理 · 电气工程与系统科学 2020-07-30 Gabriel Mittag , Ross Cutler , Yasaman Hosseinkashi , Michael Revow , Sriram Srinivasan , Naglakshmi Chande , Robert Aichner

The conventional speaker recognition frameworks (e.g., the i-vector and CNN-based approach) have been successfully applied to various tasks when the channel of the enrolment dataset is similar to that of the test dataset. However, in…

音频与语音处理 · 电气工程与系统科学 2019-02-26 Xin Fang , Liang Zou , Jin Li , Lei Sun , Zhen-Hua Ling

Recently, increasing interest has been drawn in exploiting deep convolutional neural networks (DCNNs) for no-reference image quality assessment (NR-IQA). Despite of the notable success achieved, there is a broad consensus that training…

图像与视频处理 · 电气工程与系统科学 2020-04-14 Hancheng Zhu , Leida Li , Jinjian Wu , Weisheng Dong , Guangming Shi

Learning-based image quality assessment (IQA) has made remarkable progress in the past decade, but nearly all consider the two key components -- model and data -- in isolation. Specifically, model-centric IQA focuses on developing…

计算机视觉与模式识别 · 计算机科学 2023-12-11 Peibei Cao , Dingquan Li , Kede Ma

Deep learning based image quality assessment (IQA) models usually learn to predict image quality from a single dataset, leading the model to overfit specific scenes. To account for this, mixed datasets training can be an effective way to…

计算机视觉与模式识别 · 计算机科学 2022-11-15 Zhaopeng Feng , Keyang Zhang , Shuyue Jia , Baoliang Chen , Shiqi Wang

Deep neural networks have shown recent promise in many language-related tasks such as the modeling of conversations. We extend RNN-based sequence to sequence models to capture the long range discourse across many turns of conversation. We…

计算与语言 · 计算机科学 2016-07-18 John M. Pierre , Mark Butler , Jacob Portnoff , Luis Aguilar

Neural network based approaches to speech enhancement have shown to be particularly powerful, being able to leverage a data-driven approach to result in a significant performance gain versus other approaches. Such approaches are reliant on…

声音 · 计算机科学 2023-12-15 George Close , William Ravenscroft , Thomas Hain , Stefan Goetze

Image Quality Assessment (IQA) models aim to predict perceptual image quality in alignment with human judgments. No-Reference (NR) IQA remains particularly challenging due to the absence of a reference image. While deep learning has…

图像与视频处理 · 电气工程与系统科学 2025-07-18 Rajesh Sureddi , Saman Zadtootaghaj , Nabajeet Barman , Alan C. Bovik

This paper presents an improved deep embedding learning method based on convolutional neural network (CNN) for text-independent speaker verification. Two improvements are proposed for x-vector embedding learning: (1) Multi-scale convolution…

音频与语音处理 · 电气工程与系统科学 2020-01-15 Bin Gu , Wu Guo

Most state-of-the-art Deep Learning (DL) approaches for speaker recognition work on a short utterance level. Given the speech signal, these algorithms extract a sequence of speaker embeddings from short segments and those are averaged to…

声音 · 计算机科学 2019-07-03 Miquel India , Pooyan Safari , Javier Hernando

Large-scale pretraining datasets drive the success of large language models (LLMs). However, these web-scale corpora inevitably contain large amounts of noisy data due to unregulated web content or randomness inherent in data. Although LLM…

机器学习 · 计算机科学 2026-02-03 Qizhen Zhang , Ankush Garg , Jakob Foerster , Niladri Chatterji , Kshitiz Malik , Mike Lewis

Generative models for image restoration, enhancement, and generation have significantly improved the quality of the generated images. Surprisingly, these models produce more pleasant images to the human eye than other methods, yet, they may…

图像与视频处理 · 电气工程与系统科学 2022-04-28 Marcos V. Conde , Maxime Burchi , Radu Timofte

With the development of deep learning, speech enhancement has been greatly optimized in terms of speech quality. Previous methods typically focus on the discriminative supervised learning or generative modeling, which tends to introduce…

音频与语音处理 · 电气工程与系统科学 2025-10-31 Nan Xu , Zhaolong Huang , Xiaonan Zhi