中文
相关论文

相关论文: AVR: Synergizing Foundation Models for Audio-Visua…

200 篇论文

Automatic Speech Recognition (ASR) has achieved remarkable success with deep learning, driving advancements in conversational artificial intelligence, media transcription, and assistive technologies. However, ASR systems still struggle in…

声音 · 计算机科学 2026-03-17 Haoyuan Yang , Yue Zhang , Liqiang Jing , John H. L. Hansen

This paper deals with Audio-Visual Speech Recognition (AVSR) under multimodal input corruption situations where audio inputs and visual inputs are both corrupted, which is not well addressed in previous research directions. Previous studies…

多媒体 · 计算机科学 2023-03-21 Joanna Hong , Minsu Kim , Jeongsoo Choi , Yong Man Ro

This paper presents a novel optimization framework for automatic speech recognition (ASR) with the aim of reducing hallucinations produced by an ASR model. The proposed framework optimizes the ASR model to maximize an expected factual…

音频与语音处理 · 电气工程与系统科学 2023-02-27 Naoyuki Kanda , Takuya Yoshioka , Yang Liu

Despite being a critical communication skill, grasping humor is challenging -- a successful use of humor requires a mixture of both engaging content build-up and an appropriate vocal delivery (e.g., pause). Prior studies on computational…

计算与语言 · 计算机科学 2021-07-20 Xingbo Wang , Yao Ming , Tongshuang Wu , Haipeng Zeng , Yong Wang , Huamin Qu

Voice, as input, has progressively become popular on mobiles and seems to transcend almost entirely text input. Through voice, the voice search (VS) system can provide a more natural way to meet user's information needs. However, errors…

信息检索 · 计算机科学 2023-09-06 Yi-Cheng Wang , Tzu-Ting Yang , Hsin-Wei Wang , Bi-Cheng Yan , Berlin Chen

Audiovisual Automatic Speech Recognition (AV-ASR) aims to improve speech recognition accuracy by leveraging visual signals. It is particularly challenging in unconstrained real-world scenarios across various domains due to noisy acoustic…

音频与语音处理 · 电气工程与系统科学 2024-12-30 Yihan Wu , Yichen Lu , Yifan Peng , Xihua Wang , Ruihua Song , Shinji Watanabe

Employing pre-trained language models (LM) to extract contextualized word representations has achieved state-of-the-art performance on various NLP tasks. However, applying this technique to noisy transcripts generated by automatic speech…

计算与语言 · 计算机科学 2020-11-03 Chao-Wei Huang , Yun-Nung Chen

Automatic speech recognition (ASR) system is becoming a ubiquitous technology. Although its accuracy is closing the gap with that of human level under certain settings, one area that can further improve is to incorporate user-specific…

计算与语言 · 计算机科学 2020-05-05 Young Mo Kang , Yingbo Zhou

This paper presents a self-supervised method for visual detection of the active speaker in a multi-person spoken interaction scenario. Active speaker detection is a fundamental prerequisite for any artificial cognitive system attempting to…

计算机视觉与模式识别 · 计算机科学 2019-07-19 Kalin Stefanov , Jonas Beskow , Giampiero Salvi

Dialog systems, such as voice assistants, are expected to engage with users in complex, evolving conversations. Unfortunately, traditional automatic speech recognition (ASR) systems deployed in such applications are usually trained to…

音频与语音处理 · 电气工程与系统科学 2024-09-17 Hitesh Tulsiani , David M. Chan , Shalini Ghosh , Garima Lalwani , Prabhat Pandey , Ankish Bansal , Sri Garimella , Ariya Rastrow , Björn Hoffmeister

Voice activity detection is an essential pre-processing component for speech-related tasks such as automatic speech recognition (ASR). Traditional supervised VAD systems obtain frame-level labels from an ASR pipeline by using, e.g., a…

声音 · 计算机科学 2021-05-11 Heinrich Dinkel , Shuai Wang , Xuenan Xu , Mengyue Wu , Kai Yu

Audio-Visual Speech Recognition (AVSR) has gained significant attention recently due to its robustness against noise, which often challenges conventional speech recognition systems that rely solely on audio features. Despite this advantage,…

计算与语言 · 计算机科学 2025-06-06 Thai-Binh Nguyen , Thi Van Nguyen , Quoc Truong Do , Chi Mai Luong

Automatic Speech Recognition (ASR) systems in real-world settings need to handle imperfect audio, often degraded by hardware limitations or environmental noise, while accommodating diverse user groups. In human-robot interaction (HRI),…

机器人学 · 计算机科学 2025-08-26 Theresa Pekarek Rosin , Julia Gachot , Henri-Leon Kordt , Matthias Kerzel , Stefan Wermter

Audio-visual automatic speech recognition is a promising approach to robust ASR under noisy conditions. However, up until recently it had been traditionally studied in isolation assuming the video of a single speaking face matches the…

音频与语音处理 · 电气工程与系统科学 2022-05-13 Otavio Braga , Olivier Siohan

This paper proposes an efficient attempt to noisy speech emotion recognition (NSER). Conventional NSER approaches have proven effective in mitigating the impact of artificial noise sources, such as white Gaussian noise, but are limited to…

声音 · 计算机科学 2026-01-13 Xiaohan Shi , Jiajun He , Xingfeng Li , Tomoki Toda

In Speech Emotion Recognition (SER), textual data is often used alongside audio signals to address their inherent variability. However, the reliance on human annotated text in most research hinders the development of practical SER systems.…

音频与语音处理 · 电气工程与系统科学 2023-05-30 Yuanchao Li , Zeyu Zhao , Ondrej Klejch , Peter Bell , Catherine Lai

The usage of automatic speech recognition (ASR) systems are becoming omnipresent ranging from personal assistant to chatbots, home, and industrial automation systems, etc. Modern robots are also equipped with ASR capabilities for…

音频与语音处理 · 电气工程与系统科学 2022-10-25 Pradip Pramanick , Chayan Sarkar

Visual speech recognition (VSR), which decodes spoken words from video data, offers significant benefits, particularly when audio is unavailable. However, the high dimensionality of video data leads to prohibitive computational costs that…

计算机视觉与模式识别 · 计算机科学 2025-02-10 Iason Ioannis Panagos , Giorgos Sfikas , Christophoros Nikou

Visual signals can enhance audiovisual speech recognition accuracy by providing additional contextual information. Given the complexity of visual signals, an audiovisual speech recognition model requires robust generalization capabilities…

音频与语音处理 · 电气工程与系统科学 2024-09-20 Yihan Wu , Yifan Peng , Yichen Lu , Xuankai Chang , Ruihua Song , Shinji Watanabe

Visual Speech Recognition (VSR) stands at the intersection of computer vision and speech recognition, aiming to interpret spoken content from visual cues. A prominent challenge in VSR is the presence of homophenes-visually similar lip…

人工智能 · 计算机科学 2024-06-19 Young Jin Ahn , Jungwoo Park , Sangha Park , Jonghyun Choi , Kee-Eung Kim