中文
相关论文

相关论文: Lumi\`ereNet: Lecture Video Synthesis from Audio

200 篇论文

Video description entails automatically generating coherent natural language sentences that narrate the content of a given video. We introduce CLearViD, a transformer-based model for video description generation that leverages curriculum…

计算机视觉与模式识别 · 计算机科学 2023-11-09 Cheng-Yu Chuang , Pooyan Fazli

With the emergence of large-scale open online courses and online academic conferences, it has become increasingly feasible and convenient to access online educational resources. However, it is time consuming and challenging to effectively…

人机交互 · 计算机科学 2022-01-28 Jiaohao Weng , Chao Zhang , Xi Yang , Haoran Xie

This work proposes an industry-level omni-modal large language model (LLM) pipeline that integrates auditory, visual, and linguistic modalities to overcome challenges such as limited tri-modal datasets, high computational costs, and complex…

The light transport (LT) of a scene describes how it appears under different lighting and viewing directions, and complete knowledge of a scene's LT enables the synthesis of novel views under arbitrary lighting. In this paper, we focus on…

Low-light image enhancement (LLIE) aims at improving the perception or interpretability of an image captured in an environment with poor illumination. Recent advances in this area are dominated by deep learning-based solutions, where many…

计算机视觉与模式识别 · 计算机科学 2021-11-08 Chongyi Li , Chunle Guo , Linghao Han , Jun Jiang , Ming-Ming Cheng , Jinwei Gu , Chen Change Loy

Ultrasound (US) is widely used for its advantages of real-time imaging, radiation-free and portability. In clinical practice, analysis and diagnosis often rely on US sequences rather than a single image to obtain dynamic anatomical…

计算机视觉与模式识别 · 计算机科学 2022-07-04 Jiamin Liang , Xin Yang , Yuhao Huang , Kai Liu , Xinrui Zhou , Xindi Hu , Zehui Lin , Huanjia Luo , Yuanji Zhang , Yi Xiong , Dong Ni

Commercially available light field cameras have difficulty in capturing 5D (4D + time) light field videos. They can only capture still light filed images or are excessively expensive for normal users to capture the light field video. To…

计算机视觉与模式识别 · 计算机科学 2019-12-24 Kyuho Bae , Andre Ivan , Hajime Nagahara , In Kyu Park

While multi-modal learning has advanced significantly, current approaches often create inconsistencies in representation and reasoning of different modalities. We propose UMaT, a theoretically-grounded framework that unifies visual and…

计算机视觉与模式识别 · 计算机科学 2025-06-11 Xiaowei Bi , Zheyuan Xu

We introduce MoNet, a novel functionally modular network for self-supervised and interpretable end-to-end learning. By leveraging its functional modularity with a latent-guided contrastive loss function, MoNet efficiently learns…

机器学习 · 计算机科学 2024-06-06 Hyunki Seong , David Hyunchul Shim

Generating detailed descriptions from multiple cameras and viewpoints is challenging due to the complex and inconsistent nature of visual data. In this paper, we introduce PerspectiveNet, a lightweight yet efficient model for generating…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Vinh Nguyen

In the last couple of years, weakly labeled learning has turned out to be an exciting approach for audio event detection. In this work, we introduce webly labeled learning for sound events which aims to remove human supervision altogether…

声音 · 计算机科学 2019-07-16 Anurag Kumar , Ankit Shah , Bhiksha Raj , Alex Hauptmann

The goal of this work is to synchronise audio and video of a talking face using deep neural network models. Existing works have trained networks on proxy tasks such as cross-modal similarity learning, and then computed similarities between…

计算机视觉与模式识别 · 计算机科学 2021-03-22 You Jin Kim , Hee Soo Heo , Soo-Whan Chung , Bong-Jin Lee

We present a novel Relightable Neural Renderer (RNR) for simultaneous view synthesis and relighting using multi-view image inputs. Existing neural rendering (NR) does not explicitly model the physical rendering process and hence has limited…

计算机视觉与模式识别 · 计算机科学 2020-06-16 Zhang Chen , Anpei Chen , Guli Zhang , Chengyuan Wang , Yu Ji , Kiriakos N. Kutulakos , Jingyi Yu

Generating realistic audio effects for movies and other media is a challenging task that is accomplished today primarily through physical techniques known as Foley art. Foley artists create sounds with common objects (e.g., boxing gloves,…

声音 · 计算机科学 2023-08-25 Matthew Martel , Jackson Wagner

Lecture slide element detection and retrieval are key problems in slide understanding. Training effective models for these tasks often depends on extensive manual annotation. However, annotating large volumes of lecture slides for…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Suyash Maniyar , Vishvesh Trivedi , Ajoy Mondal , Anand Mishra , C. V. Jawahar

Artists and video game designers often construct 2D animations using libraries of sprites -- textured patches of objects and characters. We propose a deep learning approach that decomposes sprite-based video animations into a disentangled…

计算机视觉与模式识别 · 计算机科学 2021-10-22 Dmitriy Smirnov , Michael Gharbi , Matthew Fisher , Vitor Guizilini , Alexei A. Efros , Justin Solomon

YouTube users looking for instructions for a specific task may spend a long time browsing content trying to find the right video that matches their needs. Creating a visual summary (abridged version of a video) provides viewers with a quick…

计算机视觉与模式识别 · 计算机科学 2022-08-16 Medhini Narasimhan , Arsha Nagrani , Chen Sun , Michael Rubinstein , Trevor Darrell , Anna Rohrbach , Cordelia Schmid

Unsupervised video object learning seeks to decompose video scenes into structural object representations without any supervision from depth, optical flow, or segmentation. We present VONet, an innovative approach that is inspired by MONet.…

计算机视觉与模式识别 · 计算机科学 2024-01-23 Haonan Yu , Wei Xu

Automatically generating a natural language sentence to describe the content of an input video is a very challenging problem. It is an essential multimodal task in which auditory and visual contents are equally important. Although audio…

计算机视觉与模式识别 · 计算机科学 2018-12-10 Yapeng Tian , Chenxiao Guan , Justin Goodman , Marc Moore , Chenliang Xu

Binaural audio gives the listener an immersive experience and can enhance augmented and virtual reality. However, recording binaural audio requires specialized setup with a dummy human head having microphones in left and right ears. Such a…

计算机视觉与模式识别 · 计算机科学 2021-11-17 Kranti Kumar Parida , Siddharth Srivastava , Gaurav Sharma