English
Related papers

Related papers: AnimeScore: A Preference-Based Dataset and Framewo…

200 papers

Generating high-quality cartoon animations multimodal control is challenging due to the complexity of non-human characters, stylistically diverse motions and fine-grained emotions. There is a huge domain gap between real-world videos and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Shuolin Xu , Bingyuan Wang , Zeyu Cai , Fangteng Fu , Yue Ma , Tongyi Lee , Hongchuan Yu , Zeyu Wang

Persona-based dialogue generation is an important milestone towards building conversational artificial intelligence. Despite the ever-improving capabilities of large language models (LLMs), effectively integrating persona fidelity in…

Computation and Language · Computer Science 2025-08-12 Arpita Saggar , Jonathan C. Darling , Vania Dimitrova , Duygu Sarikaya , David C. Hogg

Evaluating AI generated dubbed content is inherently multi-dimensional, shaped by synchronization, intelligibility, speaker consistency, emotional alignment, and semantic context. Human Mean Opinion Scores (MOS) remain the gold standard but…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-27 Ashwini Dasare , Nirmesh Shah , Ashishkumar Gudmalwar , Pankaj Wasnik

Spoken dialogue models currently lack the ability for fine-grained speech style control, a critical capability for human-like interaction that is often overlooked in favor of purely functional capabilities like reasoning and question…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-28 Wenming Tu , Guanrou Yang , Ruiqi Yan , Wenxi Chen , Ziyang Ma , Yipeng Kang , Kai Yu , Xie Chen , Zilong Zheng

Progress in speech processing has been facilitated by shared datasets and benchmarks. Historically these have focused on automatic speech recognition (ASR), speaker identification, or other lower-level tasks. Interest has been growing in…

Computation and Language · Computer Science 2022-08-01 Suwon Shon , Ankita Pasad , Felix Wu , Pablo Brusco , Yoav Artzi , Karen Livescu , Kyu J. Han

Objective estimators of multimedia quality are often judged by comparing estimates with subjective "truth data," most often via Pearson correlation coefficient (PCC) or mean-squared error (MSE). But subjective test results contain noise, so…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-16 Jaden Pieper , Stephen D. Voran

We present Kimi-Audio, an open-source audio foundation model that excels in audio understanding, generation, and conversation. We detail the practices in building Kimi-Audio, including model architecture, data curation, training recipe,…

MOS (Mean Opinion Score) is a subjective method used for the evaluation of a system's quality. Telecommunications (for voice and video), and speech synthesis systems (for generated speech) are a few of the many applications of the method.…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-26 Bálint Gyires-Tóth , Csaba Zainkó

Automatic Speech Recognition (ASR) in medical contexts has the potential to save time, cut costs, increase report accuracy, and reduce physician burnout. However, the healthcare industry has been slower to adopt this technology, in part due…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-14 Joel Shor , Ruyue Agnes Bi , Subhashini Venugopalan , Steven Ibara , Roman Goldenberg , Ehud Rivlin

Prompt-based text-to-speech (TTS) aims to generate speech that adheres to fine-grained style cues provided in a text prompt. However, most prior works depend on neither plausible nor faithful measures to evaluate prompt adherence. That is,…

Sound · Computer Science 2026-01-12 Chanhee Cho , Nayeon Kim , Bugeun Kim

Human feedback has become the de facto standard for evaluating the performance of Large Language Models, and is increasingly being used as a training objective. However, it is not clear which properties of a generated output this single…

Computation and Language · Computer Science 2024-01-17 Tom Hosking , Phil Blunsom , Max Bartolo

Speech severity evaluation is becoming increasingly important as the economic burden of speech disorders grows. Current speech severity models often struggle with generalization, learning dataset-specific acoustic cues rather than…

Sound · Computer Science 2025-10-02 Bence Mark Halpern , Tomoki Toda

Kawaii is the Japanese concept of cute++, a global export with local characteristics. Recent work has explored kawaii as a feature of user experience (UX) with social robots, virtual characters, and voice assistants, i.e., kawaii vocalics.…

Human-Computer Interaction · Computer Science 2023-10-10 Katie Seaborn , Katja Rogers , Somang Name , Miu Kojima

Word error rate (WER) and character error rate (CER) are standard metrics in Speech Recognition (ASR), but one problem has always been alternative spellings: If one's system transcribes adviser whereas the ground truth has advisor, this…

Computation and Language · Computer Science 2023-06-08 Shigeki Karita , Richard Sproat , Haruko Ishikawa

Large Language Models and commercial speech synthesis systems now enable highly realistic AI-generated voice scams (vishing), raising urgent concerns about deception at scale. Yet it remains unclear whether individuals can reliably…

Cryptography and Security · Computer Science 2026-03-27 Zoha Hayat Bhatti , Bakhtawar Ahtisham , Seemal Tausif , Niklas George , Nida ul Habib Bajwa , Mobin Javed

The growing capabilities of large language models and multimodal systems have spurred interest in voice-first AI assistants, yet existing benchmarks are inadequate for evaluating the full range of these systems' capabilities. We introduce…

Computation and Language · Computer Science 2025-09-29 Ke Wang , Houxing Ren , Zimu Lu , Mingjie Zhan , Hongsheng Li

Speech inherently contains rich acoustic information that extends far beyond the textual language. In real-world spoken language understanding, effective interpretation often requires integrating semantic meaning (e.g., content),…

Computation and Language · Computer Science 2026-03-17 Dingdong Wang , Junan Li , Jincenzi Wu , Dongchao Yang , Xueyuan Chen , Tianhua Zhang , Helen Meng

Language model benchmarks are pervasive and computationally-efficient proxies for real-world performance. However, many recent works find that benchmarks often fail to predict real utility. Towards bridging this gap, we introduce benchmark…

Artificial Intelligence · Computer Science 2026-05-28 Marco Gutierrez , Xinyi Leng , Hannah Cyberey , Jonathan Richard Schwarz , Ahmed Alaa , Thomas Hartvigsen

Large audio-language models (LALMs) have achieved near-human performance in sentence-level transcription and emotion recognition. However, existing evaluations focus mainly on surface-level perception, leaving the capacity of models for…

Computation and Language · Computer Science 2025-08-05 Wanqi Yang , Yanda Li , Yunchao Wei , Meng Fang , Ling Chen

Audio-to-score alignment aims at generating an accurate mapping between a performance audio and the score of a given piece. Standard alignment methods are based on Dynamic Time Warping (DTW) and employ handcrafted features, which cannot be…

Sound · Computer Science 2020-11-17 Ruchit Agrawal , Simon Dixon
‹ Prev 1 4 5 6 7 8 10 Next ›