中文
相关论文

相关论文: Action Dubber: Timing Audible Actions via Inflecti…

200 篇论文

Audio tagging aims to perform multi-label classification on audio chunks and it is a newly proposed task in the Detection and Classification of Acoustic Scenes and Events 2016 (DCASE 2016) challenge. This task encourages research efforts to…

声音 · 计算机科学 2017-03-20 Yong Xu , Qiuqiang Kong , Qiang Huang , Wenwu Wang , Mark D. Plumbley

Understanding the structure of complex activities in untrimmed videos is a challenging task in the area of action recognition. One problem here is that this task usually requires a large amount of hand-annotated minute- or even hour-long…

计算机视觉与模式识别 · 计算机科学 2020-10-01 Rosaura G. VidalMata , Walter J. Scheirer , Anna Kukleva , David Cox , Hilde Kuehne

Large audio language models are increasingly used for complex audio understanding tasks, but they struggle with temporal tasks that require precise temporal grounding, such as word alignment and speaker diarization. The standard approach,…

机器学习 · 计算机科学 2026-02-12 Joesph An , Phillip Keung , Jiaqi Wang , Orevaoghene Ahia , Noah A. Smith

Text-to-audio (TTA) systems have recently demonstrated strong performance in synthesizing monaural audio from text. However, the task of generating binaural spatial audio from text, which provides a more immersive auditory experience by…

音频与语音处理 · 电气工程与系统科学 2025-02-18 Linfeng Feng , Lei Zhao , Boyu Zhu , Xiao-Lei Zhang , Xuelong Li

Autonomous driving systems require huge amounts of data to train. Manual annotation of this data is time-consuming and prohibitively expensive since it involves human resources. Therefore, active learning emerged as an alternative to ease…

Explaining the decision of a multi-modal decision-maker requires to determine the evidence from both modalities. Recent advances in XAI provide explanations for models trained on still images. However, when it comes to modeling multiple…

计算机视觉与模式识别 · 计算机科学 2021-05-05 Yanbei Chen , Thomas Hummel , A. Sophia Koepke , Zeynep Akata

Accurately localizing audible objects based on audio-visual cues is the core objective of audio-visual segmentation. Most previous methods emphasize spatial or temporal multi-modal modeling, yet overlook challenges from ambiguous…

声音 · 计算机科学 2025-03-18 Chen Liu , Peike Li , Liying Yang , Dadong Wang , Lincheng Li , Xin Yu

Movie dubbing seeks to synthesize speech from a given script using a specific voice, while ensuring accurate lip synchronization and emotion-prosody alignment with the character's visual performance. However, existing alignment approaches…

声音 · 计算机科学 2025-12-22 Zhedong Zhang , Liang Li , Gaoxiang Cong , Chunshan Liu , Yuhan Gao , Xiaowan Wang , Tao Gu , Yuankai Qi

This paper presents a new task, the grounding of spatio-temporal identifying descriptions in videos. Previous work suggests potential bias in existing datasets and emphasizes the need for a new data creation schema to better model…

计算机视觉与模式识别 · 计算机科学 2019-04-09 Peratham Wiriyathammabhum , Abhinav Shrivastava , Vlad I. Morariu , Larry S. Davis

We consider the task of temporal human action localization in lifestyle vlogs. We introduce a novel dataset consisting of manual annotations of temporal localization for 13,000 narrated actions in 1,200 video clips. We present an extensive…

计算机视觉与模式识别 · 计算机科学 2022-02-22 Oana Ignat , Santiago Castro , Yuhang Zhou , Jiajun Bao , Dandan Shan , Rada Mihalcea

The perception and recognition of the surroundings is one of the essential tasks for a robot. With preliminary knowledge about a target object, it can perform various manipulation tasks such as rolling motion, palpation, and force control.…

The task of spatial-temporal action detection has attracted increasing attention among researchers. Existing dominant methods solve this problem by relying on short-term information and dense serial-wise detection on each individual frames…

计算机视觉与模式识别 · 计算机科学 2020-09-01 Yuxi Li , Weiyao Lin , Tao Wang , John See , Rui Qian , Ning Xu , Limin Wang , Shugong Xu

Active speaker detection requires a solid integration of multi-modal cues. While individual modalities can approximate a solution, accurate predictions can only be achieved by explicitly fusing the audio and visual features and modeling…

计算机视觉与模式识别 · 计算机科学 2021-10-06 Juan León-Alcázar , Fabian Caba Heilbron , Ali Thabet , Bernard Ghanem

Temporal Action Localization (TAL) aims to detect the start and end timestamps of actions in a video. However, the training of TAL models requires a substantial amount of manually annotated data. Data programming is an efficient method to…

人机交互 · 计算机科学 2025-05-26 Yuchen He , Jianbing Lv , Liqi Cheng , Lingyu Meng , Dazhen Deng , Yingcai Wu

The automatic movie dubbing model generates vivid speech from given scripts, replicating a speaker's timbre from a brief timbre prompt while ensuring lip-sync with the silent video. Existing approaches simulate a simplified workflow where…

计算与语言 · 计算机科学 2025-11-19 Rui Liu , Yuan Zhao , Zhenqi Jia

Toward the goal of automatic production for sports broadcasts, a paramount task consists in understanding the high-level semantic information of the game in play. For instance, recognizing and localizing the main actions of the game would…

计算机视觉与模式识别 · 计算机科学 2021-04-15 Silvio Giancola , Bernard Ghanem

Current methods for spatiotemporal action tube detection often extend a bounding box proposal at a given keyframe into a 3D temporal cuboid and pool features from nearby frames. However, such pooling fails to accumulate meaningful…

计算机视觉与模式识别 · 计算机科学 2022-10-26 Gurkirt Singh , Vasileios Choutas , Suman Saha , Fisher Yu , Luc Van Gool

Text-to-video editing aims to edit the visual appearance of a source video conditional on textual prompts. A major challenge in this task is to ensure that all frames in the edited video are visually consistent. Most recent works apply…

计算机视觉与模式识别 · 计算机科学 2024-03-04 Yuren Cong , Mengmeng Xu , Christian Simon , Shoufa Chen , Jiawei Ren , Yanping Xie , Juan-Manuel Perez-Rua , Bodo Rosenhahn , Tao Xiang , Sen He

Audio-visual speech recognition (AVSR) aims to transcribe human speech using both audio and video modalities. In practical environments with noise-corrupted audio, the role of video information becomes crucial. However, prior works have…

音频与语音处理 · 电气工程与系统科学 2024-10-15 Sungnyun Kim , Kangwook Jang , Sangmin Bae , Hoirin Kim , Se-Young Yun

Asynchronous inference has emerged as a prevalent paradigm in robotic manipulation, achieving significant progress in ensuring trajectory smoothness and efficiency. However, a systemic challenge remains unresolved, as inherent latency…

机器人学 · 计算机科学 2026-04-14 Haoyu Wei , Xiuwei Xu , Ziyang Cheng , Hang Yin , Angyuan Ma , Bingyao Yu , Jie Zhou , Jiwen Lu