中文
相关论文

相关论文: Cross modal video representations for weakly super…

200 篇论文

Weakly supervised Audio-Visual Video Parsing (AVVP) aims to recognize and temporally localize audio, visual, and audio-visual events in videos using only coarse-grained labels. Faced with the challenging task settings, existing research…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Huilai Li , Xiaomeng Di , Ying Xing , Yonghao Dang , Yiming Wang , Jianqin Yin

Audio-visual speaker diarization aims at detecting "who spoke when" using both auditory and visual signals. Existing audio-visual diarization datasets are mainly focused on indoor environments like meeting rooms or news studios, which are…

计算机视觉与模式识别 · 计算机科学 2022-07-19 Eric Zhongcong Xu , Zeyang Song , Satoshi Tsutsui , Chao Feng , Mang Ye , Mike Zheng Shou

Audio-visual speech recognition (AVSR) aims to transcribe human speech using both audio and video modalities. In practical environments with noise-corrupted audio, the role of video information becomes crucial. However, prior works have…

音频与语音处理 · 电气工程与系统科学 2024-10-15 Sungnyun Kim , Kangwook Jang , Sangmin Bae , Hoirin Kim , Se-Young Yun

Cross-lingual self-supervised learning has been a growing research topic in the last few years. However, current works only explored the use of audio signals to create representations. In this work, we study cross-lingual self-supervised…

计算与语言 · 计算机科学 2023-03-17 Andreas Zinonos , Alexandros Haliassos , Pingchuan Ma , Stavros Petridis , Maja Pantic

Self-supervision has shown great potential for audio-visual speech recognition by vastly reducing the amount of labeled data required to build good systems. However, existing methods are either not entirely end-to-end or do not train joint…

音频与语音处理 · 电气工程与系统科学 2024-01-23 Jiachen Lian , Alexei Baevski , Wei-Ning Hsu , Michael Auli

Self-supervised learning has attracted plenty of recent research interest. However, most works for self-supervision in speech are typically unimodal and there has been limited work that studies the interaction between audio and visual…

音频与语音处理 · 电气工程与系统科学 2021-03-19 Abhinav Shukla , Stavros Petridis , Maja Pantic

We address the problem of active speaker detection through a new framework, called SPELL, that learns long-range multimodal graphs to encode the inter-modal relationship between audio and visual data. We cast active speaker detection as a…

计算机视觉与模式识别 · 计算机科学 2021-12-07 Sourya Roy , Kyle Min , Subarna Tripathi , Tanaya Guha , Somdeb Majumdar

Active speaker detection is a challenging task in audio-visual scenario understanding, which aims to detect who is speaking in one or more speakers scenarios. This task has received extensive attention as it is crucial in applications such…

计算机视觉与模式识别 · 计算机科学 2023-03-09 Junhua Liao , Haihan Duan , Kanghui Feng , Wanbing Zhao , Yanbing Yang , Liangyin Chen

Voice activity detection is an essential pre-processing component for speech-related tasks such as automatic speech recognition (ASR). Traditional supervised VAD systems obtain frame-level labels from an ASR pipeline by using, e.g., a…

声音 · 计算机科学 2021-05-11 Heinrich Dinkel , Shuai Wang , Xuenan Xu , Mengyue Wu , Kai Yu

We present a universal framework to model contextualized sentence representations with visual awareness that is motivated to overcome the shortcomings of the multimodal parallel data with manual annotations. For each sentence, we first…

计算与语言 · 计算机科学 2019-11-12 Zhuosheng Zhang , Rui Wang , Kehai Chen , Masao Utiyama , Eiichiro Sumita , Hai Zhao

Current methods for active speak er detection focus on modeling short-term audiovisual information from a single speaker. Although this strategy can be enough for addressing single-speaker scenarios, it prevents accurate detection when the…

计算机视觉与模式识别 · 计算机科学 2020-05-21 Juan Leon Alcazar , Fabian Caba Heilbron , Long Mai , Federico Perazzi , Joon-Young Lee , Pablo Arbelaez , Bernard Ghanem

In this work, we present a novel audio-visual dataset for active speaker detection in the wild. A speaker is considered active when his or her face is visible and the voice is audible simultaneously. Although active speaker detection is a…

计算机视觉与模式识别 · 计算机科学 2021-08-18 You Jin Kim , Hee-Soo Heo , Soyeon Choe , Soo-Whan Chung , Yoohwan Kwon , Bong-Jin Lee , Youngki Kwon , Joon Son Chung

This paper strives for activity recognition under domain shift, for example caused by change of scenery or camera viewpoint. The leading approaches reduce the shift in activity appearance by adversarial training and self-supervised…

计算机视觉与模式识别 · 计算机科学 2022-03-30 Yunhua Zhang , Hazel Doughty , Ling Shao , Cees G. M. Snoek

Speaker verification has been widely explored using speech signals, which has shown significant improvement using deep models. Recently, there has been a surge in exploring faces and voices as they can offer more complementary and…

声音 · 计算机科学 2023-09-29 R. Gnana Praveen , Jahangir Alam

Dense video captioning is a task of localizing interesting events from an untrimmed video and producing textual description (captions) for each localized event. Most of the previous works in dense video captioning are solely based on visual…

计算机视觉与模式识别 · 计算机科学 2020-05-07 Vladimir Iashin , Esa Rahtu

Active speaker detection (ASD) and virtual cinematography (VC) can significantly improve the remote user experience of a video conference by automatically panning, tilting and zooming of a video conferencing camera: users subjectively rate…

音频与语音处理 · 电气工程与系统科学 2022-05-26 Ross Cutler , Ramin Mehran , Sam Johnson , Cha Zhang , Adam Kirk , Oliver Whyte , Adarsh Kowdle

Audio-visual segmentation is a challenging task that aims to predict pixel-level masks for sound sources in a video. Previous work applied a comprehensive manually designed architecture with countless pixel-wise accurate masks as…

计算机视觉与模式识别 · 计算机科学 2023-11-28 Shentong Mo , Bhiksha Raj

In this paper, we study the associations between human faces and voices. Audiovisual integration, specifically the integration of facial and vocal information is a well-researched area in neuroscience. It is shown that the overlapping…

计算机视觉与模式识别 · 计算机科学 2018-11-05 Changil Kim , Hijung Valentina Shin , Tae-Hyun Oh , Alexandre Kaspar , Mohamed Elgharib , Wojciech Matusik

Augmented reality devices have the potential to enhance human perception and enable other assistive functionalities in complex conversational environments. Effectively capturing the audio-visual context necessary for understanding these…

计算机视觉与模式识别 · 计算机科学 2022-01-07 Hao Jiang , Calvin Murdock , Vamsi Krishna Ithapu

Humans are able to localize objects in the environment using both visual and auditory cues, integrating information from multiple modalities into a common reference frame. We introduce a system that can leverage unlabeled audio-visual data…

计算机视觉与模式识别 · 计算机科学 2019-10-28 Chuang Gan , Hang Zhao , Peihao Chen , David Cox , Antonio Torralba