English
Related papers

Related papers: Lumi\`ereNet: Lecture Video Synthesis from Audio

200 papers

Joint understanding of video and language is an active research area with many applications. Prior work in this domain typically relies on learning text-video embeddings. One difficulty with this approach, however, is the lack of…

Computer Vision and Pattern Recognition · Computer Science 2020-01-17 Antoine Miech , Ivan Laptev , Josef Sivic

Audio-visual speaker extraction has attracted increasing attention, as it removes the need for pre-registered speech and leverages the visual modality as a complement to audio. Although existing methods have achieved impressive performance,…

Multimedia · Computer Science 2026-03-03 Jiadong Wang , Ke Zhang , Xinyuan Qian , Ruijie Tao , Haizhou Li , Björn Schuller

The recent development of Video-based Large Language Models (VideoLLMs), has significantly advanced video summarization by aligning video features and, in some cases, audio features with Large Language Models (LLMs). Each of these VideoLLMs…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Kuan-Chen Mu , Zhi-Yi Chin , Wei-Chen Chiu

In this demo, we present VirtualConductor, a system that can generate conducting video from any given music and a single user's image. First, a large-scale conductor motion dataset is collected and constructed. Then, we propose Audio Motion…

Computer Vision and Pattern Recognition · Computer Science 2021-08-11 Delong Chen , Fan Liu , Zewen Li , Feng Xu

Lipreading has a lot of potential applications such as in the domain of surveillance and video conferencing. Despite this, most of the work in building lipreading systems has been limited to classifying silent videos into classes…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-03 Yaman Kumar , Rohit Jain , Khwaja Mohd. Salik , Rajiv Ratn Shah , Yifang yin , Roger Zimmermann

We propose to synthesize high-quality and synchronized audio, given video and optional text conditions, using a novel multimodal joint training framework MMAudio. In contrast to single-modality training conditioned on (limited) video data…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 Ho Kei Cheng , Masato Ishii , Akio Hayakawa , Takashi Shibuya , Alexander Schwing , Yuki Mitsufuji

Video-Text pre-training aims at learning transferable representations from large-scale video-text pairs via aligning the semantics between visual and textual information. State-of-the-art approaches extract visual features from raw pixels…

Computer Vision and Pattern Recognition · Computer Science 2021-12-07 Rui Yan , Mike Zheng Shou , Yixiao Ge , Alex Jinpeng Wang , Xudong Lin , Guanyu Cai , Jinhui Tang

Synthesizing realistic videos according to a given speech is still an open challenge. Previous works have been plagued by issues such as inaccurate lip shape generation and poor image quality. The key reason is that only motions and…

Computer Vision and Pattern Recognition · Computer Science 2023-09-12 Xiuzhe Wu , Pengfei Hu , Yang Wu , Xiaoyang Lyu , Yan-Pei Cao , Ying Shan , Wenming Yang , Zhongqian Sun , Xiaojuan Qi

With the growing demand for real-time video enhancement in live applications, existing methods often struggle to balance speed and effective exposure control, particularly under uneven lighting. We introduce RRNet (Rendering Relighting…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Wenlong Yang , Canran Jin , Weihang Yuan , Chao Wang , Lifeng Sun

We present LumiX, a structured diffusion framework for coherent text-to-intrinsic generation. Conditioned on text prompts, LumiX jointly generates a comprehensive set of intrinsic maps (e.g., albedo, irradiance, normal, depth, and final…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Xu Han , Biao Zhang , Xiangjun Tang , Xianzhi Li , Peter Wonka

When watching videos, the occurrence of a visual event is often accompanied by an audio event, e.g., the voice of lip motion, the music of playing instruments. There is an underlying correlation between audio and visual events, which can be…

Multimedia · Computer Science 2020-08-19 Ying Cheng , Ruize Wang , Zhihao Pan , Rui Feng , Yuejie Zhang

Recent advances have shown that large-scale video diffusion models can be repurposed as neural renderers by first decomposing videos into intrinsic scene representations and then performing forward rendering under novel illumination. While…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Weiqing Xiao , Hong Li , Xiuyu Yang , Houyuan Chen , Wenyi Li , Tianqi Liu , Shaocong Xu , Chongjie Ye , Hao Zhao , Beibei Wang

We present a technique for synthesizing a motion blurred image from a pair of unblurred images captured in succession. To build this system we motivate and design a differentiable "line prediction" layer to be used as part of a neural…

Computer Vision and Pattern Recognition · Computer Science 2019-06-21 Tim Brooks , Jonathan T. Barron

This paper presents Audio-Visual LLM, a Multimodal Large Language Model that takes both visual and auditory inputs for holistic video understanding. A key design is the modality-augmented training, which involves the integration of…

Computer Vision and Pattern Recognition · Computer Science 2023-12-15 Fangxun Shu , Lei Zhang , Hao Jiang , Cihang Xie

Assistive technology is a prerequisite for making a high-quality lecture video. It is therefore imperative to edit the lecture video after recording. In this study, we aim to reduce the cumbersome task of lecture video editing by developing…

Human-Computer Interaction · Computer Science 2021-10-13 Yuma Ito , Masato Kikuchi , Tadachika Ozono , Toramatsu Shintani

Deep networks have recently enjoyed enormous success when applied to recognition and classification problems in computer vision, but their use in graphics problems has been limited. In this work, we present a novel deep architecture that…

Computer Vision and Pattern Recognition · Computer Science 2015-06-24 John Flynn , Ivan Neulander , James Philbin , Noah Snavely

Event classification is inherently sequential and multimodal. Therefore, deep neural models need to dynamically focus on the most relevant time window and/or modality of a video. In this study, we propose the Multi-level Attention Fusion…

Computer Vision and Pattern Recognition · Computer Science 2021-06-15 Mathilde Brousmiche , Jean Rouat , Stéphane Dupont

Audio is the main form for the visually impaired to obtain information. In reality, all kinds of visual data always exist, but audio data does not exist in many cases. In order to help the visually impaired people to better perceive the…

Sound · Computer Science 2021-03-19 Hailong Ning , Xiangtao Zheng , Yuan Yuan , Xiaoqiang Lu

We introduce a neural relighting algorithm for captured indoors scenes, that allows interactive free-viewpoint navigation. Our method allows illumination to be changed synthetically, while coherently rendering cast shadows and complex…

Graphics · Computer Science 2021-06-28 Julien Philip , Sébastien Morgenthaler , Michaël Gharbi , George Drettakis

In this paper, we present a method for reprogramming pre-trained audio-driven talking face synthesis models to operate in a text-driven manner. Consequently, we can easily generate face videos that articulate the provided textual sentences,…

Graphics · Computer Science 2024-01-19 Jeongsoo Choi , Minsu Kim , Se Jin Park , Yong Man Ro