中文
相关论文

相关论文: EasyVolcap: Accelerating Neural Volumetric Video R…

200 篇论文

Diverse presentation formats play a pivotal role in effectively conveying code and analytical processes during data analysis. One increasingly popular format is tutorial videos, particularly those based on Jupyter notebooks, which offer an…

人机交互 · 计算机科学 2024-08-05 Yang Ouyang , Leixian Shen , Yun Wang , Quan Li

Volumetric video is emerging as a key medium for digitizing the dynamic physical world, creating the virtual environments with six degrees of freedom to deliver immersive user experiences. However, robustly modeling general dynamic scenes,…

The video visual relation detection (VidVRD) task is to identify objects and their relationships in videos, which is challenging due to the dynamic content, high annotation costs, and long-tailed distribution of relations. Visual language…

计算机视觉与模式识别 · 计算机科学 2025-03-13 Qi Liu , Weiying Xue , Yuxiao Wang , Zhenao Wei

We present the first real-time human performance capture approach that reconstructs dense, space-time coherent deforming geometry of entire humans in general everyday clothing from just a single RGB video. We propose a novel two-stage…

计算机视觉与模式识别 · 计算机科学 2019-01-28 Marc Habermann , Weipeng Xu , Michael Zollhoefer , Gerard Pons-Moll , Christian Theobalt

We are interested in enabling visual planning for complex long-horizon tasks in the space of generated videos and language, leveraging recent advances in large generative models pretrained on Internet-scale data. To this end, we present…

We present a method to capture temporally coherent dynamic clothing deformation from a monocular RGB video input. In contrast to the existing literature, our method does not require a pre-scanned personalized mesh template, and thus can be…

计算机视觉与模式识别 · 计算机科学 2020-11-24 Donglai Xiang , Fabian Prada , Chenglei Wu , Jessica Hodgins

4D reconstruction and rendering of human activities is critical for immersive VR/AR experience.Recent advances still fail to recover fine geometry and texture results with the level of detail present in the input images from sparse…

计算机视觉与模式识别 · 计算机科学 2021-03-16 Xin Suo , Yuheng Jiang , Pei Lin , Yingliang Zhang , Kaiwen Guo , Minye Wu , Lan Xu

High-quality 4D reconstruction of human performance with complex interactions to various objects is essential in real-world scenarios, which enables numerous immersive VR/AR applications. However, recent advances still fail to provide…

计算机视觉与模式识别 · 计算机科学 2021-05-03 Zhuo Su , Lan Xu , Dawei Zhong , Zhong Li , Fan Deng , Shuxue Quan , Lu Fang

Vision-language retrieval (VLR) has attracted significant attention in both academia and industry, which involves using text (or images) as queries to retrieve corresponding images (or text). However, existing methods often neglect the rich…

计算机视觉与模式识别 · 计算机科学 2025-05-27 GuangHao Meng , Sunan He , Jinpeng Wang , Tao Dai , Letian Zhang , Jieming Zhu , Qing Li , Gang Wang , Rui Zhang , Yong Jiang

In this work, we enhance a professional end-to-end volumetric video production pipeline to achieve high-fidelity human body reconstruction using only passive cameras. While current volumetric video approaches estimate depth maps using…

计算机视觉与模式识别 · 计算机科学 2022-03-01 Decai Chen , Markus Worchel , Ingo Feldmann , Oliver Schreer , Peter Eisert

We propose VisFusion, a visibility-aware online 3D scene reconstruction approach from posed monocular videos. In particular, we aim to reconstruct the scene from volumetric features. Unlike previous reconstruction methods which aggregate…

计算机视觉与模式识别 · 计算机科学 2023-04-24 Huiyu Gao , Wei Mao , Miaomiao Liu

Neuromorphic event cameras possess superior temporal resolution, power efficiency, and dynamic range compared to traditional cameras. However, their asynchronous and sparse data format poses a significant challenge for conventional deep…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Wei Fang , Priyadarshini Panda

Human daily activities can be concisely narrated as sequences of routine events (e.g., turning off an alarm) in video streams, forming an event vocabulary. Motivated by this, we introduce VLog, a novel video understanding framework that…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Kevin Qinghong Lin , Mike Zheng Shou

Constructing supervised machine learning models for real-world video analysis require substantial labeled data, which is costly to acquire due to scarce domain expertise and laborious manual inspection. While data programming shows promise…

计算机视觉与模式识别 · 计算机科学 2023-11-02 Jianben He , Xingbo Wang , Kam Kwai Wong , Xijie Huang , Changjian Chen , Zixin Chen , Fengjie Wang , Min Zhu , Huamin Qu

Event-based vision has drawn increasing attention due to its unique characteristics, such as high temporal resolution and high dynamic range. It has been used in video super-resolution (VSR) recently to enhance the flow estimation and…

计算机视觉与模式识别 · 计算机科学 2024-06-21 Dachun Kai , Jiayao Lu , Yueyi Zhang , Xiaoyan Sun

Active speaker detection (ASD) and virtual cinematography (VC) can significantly improve the remote user experience of a video conference by automatically panning, tilting and zooming of a video conferencing camera: users subjectively rate…

音频与语音处理 · 电气工程与系统科学 2022-05-26 Ross Cutler , Ramin Mehran , Sam Johnson , Cha Zhang , Adam Kirk , Oliver Whyte , Adarsh Kowdle

The widespread adoption of digital technology has ushered in a new era of digital transformation across all aspects of our lives. Online learning, social, and work activities, such as distance education, videoconferencing, interviews, and…

多媒体 · 计算机科学 2025-08-06 Baoquan Zhao , Xiaofan Ma , Qianshi Pang , Ruomei Wang , Fan Zhou , Shujin Lin

Direct volume rendering (DVR) aims to help users identify and examine regions of interest (ROIs) within volumetric data, and feature representations that support effective ROI classification and clustering play a fundamental role in volume…

图形学 · 计算机科学 2026-04-14 Haill An , Suhyeon Kim , Donghyuk Choo , Younhyun Jung

Volumetric models have become a popular representation for 3D scenes in recent years. One of the breakthroughs leading to their popularity was KinectFusion, where the focus is on 3D reconstruction using RGB-D sensors. However, monocular…

计算机视觉与模式识别 · 计算机科学 2014-10-27 Victor Adrian Prisacariu , Olaf Kähler , Ming Ming Cheng , Carl Yuheng Ren , Julien Valentin , Philip H. S. Torr , Ian D. Reid , David W. Murray

Visual Dialog is a vision-language task that requires an AI agent to engage in a conversation with humans grounded in an image. It remains a challenging task since it requires the agent to fully understand a given question before making an…

计算与语言 · 计算机科学 2019-12-19 Feilong Chen , Fandong Meng , Jiaming Xu , Peng Li , Bo Xu , Jie Zhou