中文
相关论文

相关论文: XLAVS-R: Cross-Lingual Audio-Visual Speech Represe…

200 篇论文

Audio-visual speech enhancement (AV-SE) is the task of improving speech quality and intelligibility in a noisy environment using audio and visual information from a talker. Recently, deep learning techniques have been adopted to solve the…

音频与语音处理 · 电气工程与系统科学 2019-11-05 Daniel Michelsanti , Zheng-Hua Tan , Sigurdur Sigurdsson , Jesper Jensen

Audio-Visual Speech Recognition (AVSR) offers a robust solution for speech recognition in challenging environments, such as cocktail-party scenarios, where relying solely on audio proves insufficient. However, current AVSR models are often…

声音 · 计算机科学 2025-06-04 Thai-Binh Nguyen , Ngoc-Quan Pham , Alexander Waibel

In this work, we present the Textless Vision-Language Transformer (TVLT), where homogeneous transformer blocks take raw visual and audio inputs for vision-and-language representation learning with minimal modality-specific design, and do…

计算机视觉与模式识别 · 计算机科学 2022-11-03 Zineng Tang , Jaemin Cho , Yixin Nie , Mohit Bansal

This report presents VibeVoice-ASR, a general-purpose speech understanding framework built upon VibeVoice, designed to address the persistent challenges of context fragmentation and multi-speaker complexity in long-form audio (e.g.,…

We present an Audio-Visual Language Model (AVLM) for expressive speech generation by integrating full-face visual cues into a pre-trained expressive speech model. We explore multiple visual encoders and multimodal fusion strategies during…

计算与语言 · 计算机科学 2025-08-29 Weiting Tan , Jiachen Lian , Hirofumi Inaguma , Paden Tomasello , Philipp Koehn , Xutai Ma

Voice activity detection is an essential pre-processing component for speech-related tasks such as automatic speech recognition (ASR). Traditional supervised VAD systems obtain frame-level labels from an ASR pipeline by using, e.g., a…

声音 · 计算机科学 2021-05-11 Heinrich Dinkel , Shuai Wang , Xuenan Xu , Mengyue Wu , Kai Yu

In this paper, we propose a weakly supervised multilingual representation learning framework, called cross-lingual self-training (XLST). XLST is able to utilize a small amount of annotated data from high-resource languages to improve the…

音频与语音处理 · 电气工程与系统科学 2021-03-16 Zi-Qiang Zhang , Yan Song , Ming-Hui Wu , Xin Fang , Li-Rong Dai

Compared with automatic speech recognition (ASR), the human auditory system is more adept at handling noise-adverse situations, including environmental noise and channel distortion. To mimic this adeptness, auditory models have been widely…

计算与语言 · 计算机科学 2016-09-16 Peng Dai , Xue Teng , Frank Rudzicz , Ing Yann Soon

Novel view acoustic synthesis (NVAS) aims to render binaural audio at any target viewpoint, given a mono audio emitted by a sound source at a 3D scene. Existing methods have proposed NeRF-based implicit models to exploit visual cues as a…

声音 · 计算机科学 2025-03-18 Swapnil Bhosale , Haosen Yang , Diptesh Kanojia , Jiankang Deng , Xiatian Zhu

Recent advances in Audio-Visual Speech Recognition (AVSR) have led to unprecedented achievements in the field, improving the robustness of this type of system in adverse, noisy environments. In most cases, this task has been addressed…

计算机视觉与模式识别 · 计算机科学 2025-05-07 David Gimeno-Gómez , Carlos-D. Martínez-Hinarejos

In this work, we propose a training algorithm for an audio-visual automatic speech recognition (AV-ASR) system using deep recurrent neural network (RNN).First, we train a deep RNN acoustic model with a Connectionist Temporal Classification…

计算机视觉与模式识别 · 计算机科学 2016-11-10 Abhinav Thanda , Shankar M Venkatesan

Multilingual automatic speech recognition (ASR) systems mostly benefit low resource languages but suffer degradation in performance across several languages relative to their monolingual counterparts. Limited studies have focused on…

计算与语言 · 计算机科学 2022-07-08 Muhammad Umar Farooq , Thomas Hain

Representation learning from unlabeled data has been of major interest in artificial intelligence research. While self-supervised speech representation learning has been popular in the speech research community, very few works have…

Automatic Speech Recognition (ASR) is an active field of research due to its large number of applications and the proliferation of interfaces or computing devices that can support speech processing. However, the bulk of applications are…

人工智能 · 计算机科学 2022-03-03 Jean Louis K. E. Fendji , Diane C. M. Tala , Blaise O. Yenke , Marcellin Atemkeng

Visual and linguistic pre-training aims to learn vision and language representations together, which can be transferred to visual-linguistic downstream tasks. However, there exists semantic confusion between language and vision during the…

计算机视觉与模式识别 · 计算机科学 2023-04-11 Shentong Mo , Jingfei Xia , Ihor Markevych

The development of Neural Radiance Fields (NeRFs) has provided a potent representation for encapsulating the geometric and appearance characteristics of 3D scenes. Enhancing the capabilities of NeRFs in open-vocabulary 3D semantic…

计算机视觉与模式识别 · 计算机科学 2024-09-24 Guibiao Liao , Kaichen Zhou , Zhenyu Bao , Kanglin Liu , Qing Li

Large language models (LLMs) have demonstrated potential in handling spoken inputs for high-resource languages, reaching state-of-the-art performance in various tasks. However, their applicability is still less explored in low-resource…

音频与语音处理 · 电气工程与系统科学 2025-08-08 Seraphina Fong , Marco Matassoni , Alessio Brutti

Conventional audio-visual methods for speaker verification rely on large amounts of labeled data and separate modality-specific architectures, which is computationally expensive, limiting their scalability. To address these problems, we…

计算机视觉与模式识别 · 计算机科学 2025-06-25 Gnana Praveen Rajasekhar , Jahangir Alam

When video is shot in noisy environment, the voice of a speaker seen in the video can be enhanced using the visible mouth movements, reducing background noise. While most existing methods use audio-only inputs, improved performance is…

计算机视觉与模式识别 · 计算机科学 2018-06-14 Aviv Gabbay , Asaph Shamir , Shmuel Peleg

Vision-language models (VLMs) have achieved impressive results on single-view vision tasks, but lack the multi-view spatial reasoning capabilities essential for embodied AI systems to understand 3D environments and manipulate objects across…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Suchae Jeong , Jaehwi Song , Haeone Lee , Hanna Kim , Jian Kim , Dongjun Lee , Dong Kyu Shin , Changyeon Kim , Dongyoon Hahm , Woogyeol Jin , Juheon Choi , Kimin Lee
‹ 上一页 1 8 9 10 下一页 ›