English
Related papers

Related papers: Qwen2-Audio Technical Report

200 papers

Vision Language Action (VLA) models promise an open-vocabulary interface that can translate perceptual ambiguity into semantically grounded driving decisions, yet they still treat language as a static prior fixed at inference time. As a…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-30 Ziang Guo , Feng Yang , Xuefeng Zhang , Jiaqi Guo , Kun Zhao , Yixiao Zhou , Peng Lu , Sifa Zheng , Zufeng Zhang

Text-to-audio (TTA) generation is advancing rapidly, but evaluation remains challenging because human listening studies are expensive and existing automatic metrics capture only limited aspects of perceptual quality. We introduce AudioEval,…

Sound · Computer Science 2026-01-30 Hui Wang , Jinghua Zhao , Junyang Cheng , Cheng Liu , Yuhang Jia , Haoqin Sun , Jiaming Zhou , Yong Qin

Dialogue state tracking plays a crucial role in extracting information in task-oriented dialogue systems. However, preceding research are limited to textual modalities, primarily due to the shortage of authentic human audio datasets. We…

Sound · Computer Science 2023-12-05 Jihyun Lee , Yejin Jeon , Wonjun Lee , Yunsu Kim , Gary Geunbae Lee

Audio-driven 3D facial animation has several virtual humans applications for content creation and editing. While several existing methods provide solutions for speech-driven animation, precise control over content (what) and style (how) of…

Sound · Computer Science 2024-08-15 Qingju Liu , Hyeongwoo Kim , Gaurav Bharaj

Multimodal question answering (QA) often requires identifying which video, audio, or sensor tokens are relevant to the question. Yet modality disagreements are common: off-camera speech, background noise, or motion outside the field of view…

Computation and Language · Computer Science 2025-09-08 Subrata Biswas , Mohammad Nur Hossain Khan , Bashima Islam

Representation learning from unlabeled data has been of major interest in artificial intelligence research. While self-supervised speech representation learning has been popular in the speech research community, very few works have…

This paper introduces an automated framework WSW2.0 for analyzing vocal interactions in preschool classrooms, enhancing both accuracy and scalability through the integration of wav2vec2-based speaker classification and Whisper (large-v2 and…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-27 Anchen Sun , Tiantian Feng , Gabriela Gutierrez , Juan J Londono , Anfeng Xu , Batya Elbaum , Shrikanth Narayanan , Lynn K Perry , Daniel S Messinger

Analyses of self-supervised speech models have begun to reveal where and how they represent different types of information. However, almost all analyses have focused on English. Here, we examine how wav2vec2 models trained on four different…

Computation and Language · Computer Science 2025-06-13 Michele Gubian , Ioana Krehan , Oli Liu , James Kirby , Sharon Goldwater

We present a framework for generating appropriate facial responses from a listener in dyadic social interactions based on the speaker's words. Given an input transcription of the speaker's words with their timestamps, our approach…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Evonne Ng , Sanjay Subramanian , Dan Klein , Angjoo Kanazawa , Trevor Darrell , Shiry Ginosar

A voice AI agent that blends seamlessly into daily life would interact with humans in an autonomous, real-time, and emotionally expressive manner. Rather than merely reacting to commands, it would continuously listen, reason, and respond…

Artificial Intelligence · Computer Science 2025-05-06 Yemin Shi , Yu Shu , Siwei Dong , Guangyi Liu , Jaward Sesay , Jingwen Li , Zhiting Hu

In human dialogue, nonverbal information such as nodding and facial expressions is as crucial as verbal information, and spoken dialogue systems are also expected to express such nonverbal behaviors. We focus on nodding, which is critical…

Human-Computer Interaction · Computer Science 2025-08-05 Kazushi Kato , Koji Inoue , Divesh Lala , Keiko Ochi , Tatsuya Kawahara

Recent advances in diffusion-based video generation have enabled photo-realistic short clips, but current methods still struggle to achieve multi-modal consistency when jointly generating whole-body motion and natural speech. Current…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Xinhan Di , Kristin Qi , Pengqian Yu

Recent advances in large language models (LLMs) have significantly improved text-to-speech (TTS) systems, enhancing control over speech style, naturalness, and emotional expression, which brings TTS Systems closer to human-level…

Voice-based AI development faces unique challenges in processing both linguistic and paralinguistic information. This study compares how large audio-language models (LALMs) and humans integrate speaker characteristics during speech…

Computation and Language · Computer Science 2025-10-28 Hanlin Wu , Xufeng Duan , Zhenguang Cai

The task of predicting dialog acts (DA) based on conversational dialog is a key component in the development of conversational agents. Accurately predicting DAs requires a precise modeling of both the conversation and the global tag…

Computation and Language · Computer Science 2020-02-27 Pierre Colombo , Emile Chapuis , Matteo Manica , Emmanuel Vignon , Giovanna Varni , Chloe Clavel

This work investigates two strategies for zero-shot non-intrusive speech assessment leveraging large language models. First, we explore the audio analysis capabilities of GPT-4o. Second, we propose GPT-Whisper, which uses Whisper as an…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-22 Ryandhimas E. Zezario , Sabato M. Siniscalchi , Hsin-Min Wang , Yu Tsao

Wav2vec 2.0 is a recently proposed self-supervised framework for speech representation learning. It follows a two-stage training process of pre-training and fine-tuning, and performs well in speech recognition tasks especially ultra-low…

Sound · Computer Science 2021-01-15 Zhiyun Fan , Meng Li , Shiyu Zhou , Bo Xu

Spoken conversational systems require more than accurate speech generation to have human-like conversations: to feel natural and engaging, they must produce conversational behaviour that adapts dynamically to the context. Current spoken…

Computation and Language · Computer Science 2026-04-16 Maike Züfle , Ondrej Klejch , Nicholas Sanders , Jan Niehues , Alexandra Birch , Tsz Kin Lam

In recent years, audio-driven 3D facial animation has gained significant attention, particularly in applications such as virtual reality, gaming, and video conferencing. However, accurately modeling the intricate and subtle dynamics of…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Guinan Su , Yanwu Yang , Zhifeng Li

In this paper, we present TalkingMachines -- an efficient framework that transforms pretrained video generation models into real-time, audio-driven character animators. TalkingMachines enables natural conversational experiences by…

Sound · Computer Science 2025-06-04 Chetwin Low , Weimin Wang
‹ Prev 1 4 5 6 7 8 10 Next ›