English
Related papers

Related papers: ViSQOL v3: An Open Source Production Ready Objecti…

200 papers

Predicting audio quality in voice synthesis and conversion systems is a critical yet challenging task, especially when traditional methods like Mean Opinion Scores (MOS) are cumbersome to collect at scale. This paper addresses the gap in…

Sound · Computer Science 2023-12-27 Aditya Ravuri , Erica Cooper , Junichi Yamagishi

We introduce LRS-VoxMM, an in-the-wild benchmark for audio-visual speech recognition (AVSR). The benchmark is derived from VoxMM, a dataset of diverse real-world spoken conversations with human-annotated transcriptions. We select…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-01 Doyeop Kwak , Jeongsoo Choi , Suyeon Lee , Joon Son Chung

Neural codecs have become crucial to recent speech and audio generation research. In addition to signal compression capabilities, discrete codecs have also been found to enhance downstream training efficiency and compatibility with…

Open-vocabulary segmentation aims to achieve segmentation of arbitrary categories given unlimited text inputs as guidance. To achieve this, recent works have focused on developing various technical routes to exploit the potential of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Yong Liu , SongLi Wu , Sule Bai , Jiahao Wang , Yitong Wang , Yansong Tang

Non-intrusive assessment of speech quality and intelligibility is essential when clean reference signals are unavailable. In this work, we propose a multimodal framework that integrates audio features and visual cues to predict PESQ and…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-12 Shafique Ahmed , Ryandhimas E. Zezario , Nasir Saleem , Amir Hussain , Hsin-Min Wang , Yu Tsao

Even without directly hearing sounds, humans can effortlessly reason about auditory properties, such as pitch, loudness, or sound-source associations, drawing on auditory commonsense. In contrast, language models often lack this capability,…

Computation and Language · Computer Science 2026-01-29 Hyunjong Ok , Suho Yoo , Hyeonjun Kim , Jaeho Lee

Discrete audio tokens are compact representations that aim to preserve perceptual quality, phonetic content, and speaker characteristics while enabling efficient storage and inference, as well as competitive performance across diverse…

We introduce C3LLM (Conditioned-on-Three-Modalities Large Language Models), a novel framework combining three tasks of video-to-audio, audio-to-text, and text-to-audio together. C3LLM adapts the Large Language Model (LLM) structure as a…

Artificial Intelligence · Computer Science 2024-05-28 Zixuan Wang , Qinkai Duan , Yu-Wing Tai , Chi-Keung Tang

The quantification of audio aesthetics remains a complex challenge in audio processing, primarily due to its subjective nature, which is influenced by human perception and cultural context. Traditional methods often depend on human…

When video is shot in noisy environment, the voice of a speaker seen in the video can be enhanced using the visible mouth movements, reducing background noise. While most existing methods use audio-only inputs, improved performance is…

Computer Vision and Pattern Recognition · Computer Science 2018-06-14 Aviv Gabbay , Asaph Shamir , Shmuel Peleg

Modern software relies on a multitude of automated testing and quality assurance tools to prevent errors, bugs and potential vulnerabilities. This study sets out to provide a head-to-head, quantitative and qualitative evaluation of six…

Software Engineering · Computer Science 2025-08-07 Damian Gnieciak , Tomasz Szandala

In the development of spatial audio technologies, reliable and shared methods for evaluating audio quality are essential. Listening tests are currently the standard but remain costly in terms of time and resources. Several models predicting…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-21 Adrien Llave , Emma Granier , Grégory Pallone

Recent advances in text-to-speech (TTS) have been driven by large, multi-domain speech corpora, yet the expressive potential of audiobook data remains underexamined. We argue that human-narrated audiobooks, particularly fictional works,…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-22 Gaspard Michel , Elena V. Epure , Christophe Cerisara

Although Multimodal Large Language Models (MLLMs) demonstrate strong omni-modal perception, their ability to forecast future events from audio-visual cues remains largely unexplored, as existing benchmarks focus mainly on retrospective…

Computation and Language · Computer Science 2026-01-21 Qian Chen , Jinlan Fu , Changsong Li , See-Kiong Ng , Xipeng Qiu

Human speech processing is inherently multimodal, where visual cues (lip movements) help to better understand the speech in noise. Lip-reading driven speech enhancement significantly outperforms benchmark audio-only approaches at low…

Computer Vision and Pattern Recognition · Computer Science 2019-09-24 Ahsan Adeel , Mandar Gogate , Amir Hussain

Human perceptual studies are the gold standard for the evaluation of many research tasks in machine learning, linguistics, and psychology. However, these studies require significant time and cost to perform. As a result, many researchers…

Human-Computer Interaction · Computer Science 2022-03-10 Max Morrison , Brian Tang , Gefei Tan , Bryan Pardo

Streaming Speech-to-Text Translation (StreamST) requires producing translations concurrently with incoming speech, imposing strict latency constraints and demanding models that balance partial-information decision-making with high…

Computation and Language · Computer Science 2025-12-22 Marco Gaido , Sara Papi , Mauro Cettolo , Matteo Negri , Luisa Bentivogli

In this paper, we present an open-source software for developing a nonparallel voice conversion (VC) system named crank. Although we have released an open-source VC software based on the Gaussian mixture model named sprocket in the last VC…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-05 Kazuhiro Kobayashi , Wen-Chin Huang , Yi-Chiao Wu , Patrick Lumban Tobing , Tomoki Hayashi , Tomoki Toda

Video-conditioned audio generation, including Video-to-Sound (V2S) and Visual Text-to-Speech (VisualTTS), has traditionally been treated as distinct tasks, leaving the potential for a unified generative framework largely underexplored. In…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-23 Xin Cheng , Yuyue Wang , Xihua Wang , Yihan Wu , Kaisi Guan , Yijing Chen , Peng Zhang , Xiaojiang Liu , Meng Cao , Ruihua Song

This study has proposed an E-textbook platform, MetaCQ, which integrates ITS and OLM to enable users to monitor their study progress. The platform adopts a chatbot to generate MCQs and manage learners' study data and their learning model.…

Human-Computer Interaction · Computer Science 2025-12-02 Beier Wang , Xueting Huang
‹ Prev 1 4 5 6 7 8 10 Next ›