English
Related papers

Related papers: 3DSPA: A 3D Semantic Point Autoencoder for Evaluat…

200 papers

Visual search in 3D environments requires embodied agents to actively explore their surroundings and acquire task-relevant evidence. However, existing visual search and embodied AI benchmarks, including EQA, typically rely on static…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Koya Sakamoto , Taiki Miyanishi , Daichi Azuma , Shuhei Kurita , Shu Morikuni , Naoya Chiba , Motoaki Kawanabe , Yusuke Iwasawa , Yutaka Matsuo

Enabling intelligent agents to comprehend and interact with 3D environments through natural language is crucial for advancing robotics and human-computer interaction. A fundamental task in this field is ego-centric 3D visual grounding,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-09 Henry Zheng , Hao Shi , Qihang Peng , Yong Xien Chng , Rui Huang , Yepeng Weng , Zhongchao Shi , Gao Huang

Video world models aim to simulate dynamic, real-world environments, yet existing methods struggle to provide unified and precise control over camera and multi-object motion, as videos inherently capture dynamics in the projected 2D image…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Sixiao Zheng , Minghao Yin , Wenbo Hu , Xiaoyu Li , Ying Shan , Yanwei Fu

An interesting problem in many video-based applications is the generation of short synopses by selecting the most informative frames, a procedure which is known as video summarization. For sign language videos the benefits of using the…

Computer Vision and Pattern Recognition · Computer Science 2024-10-28 Evangelos G. Sartinas , Emmanouil Z. Psarakis , Dimitrios I. Kosmopoulos

3D Gaussian Splatting has recently enabled fast and photorealistic reconstruction of static 3D scenes. However, dynamic editing of such scenes remains a significant challenge. We introduce a novel framework, Physics-Guided Score…

Graphics · Computer Science 2026-03-26 Gal Fiebelman , Hadar Averbuch-Elor , Sagie Benaim

Driving simulation plays a crucial role in developing reliable driving agents by providing controlled, evaluative environments. To enable meaningful assessments, a high-quality driving simulator must satisfy several key requirements:…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Junzhe Jiang , Nan Song , Jingyu Li , Xiatian Zhu , Li Zhang

Spatial audio is crucial for immersive 360-degree video experiences, yet most 360-degree videos lack it due to the difficulty of capturing spatial audio during recording. Automatically generating spatial audio such as first-order ambisonics…

Sound · Computer Science 2026-04-13 Ziyu Luo , Lin Chen , Qiang Qu , Xiaoming Chen , Yiran Shen

Spatial audio is crucial for immersive 360-degree video experiences, yet most 360-degree videos lack it due to the difficulty of capturing spatial audio during recording. Automatically generating spatial audio such as first-order ambisonics…

Sound · Computer Science 2026-05-05 Ziyu Luo , Lin Chen , Qiang Qu , Xiaoming Chen , Yiran Shen

Video generative models can be regarded as world simulators due to their ability to capture dynamic, continuous changes inherent in real-world environments. These models integrate high-dimensional information across visual, temporal,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Hengyuan Cao , Yutong Feng , Biao Gong , Yijing Tian , Yunhong Lu , Chuang Liu , Bin Wang

The growing capabilities of AI in generating video content have brought forward significant challenges in effectively evaluating these videos. Unlike static images or text, video content involves complex spatial and temporal dynamics which…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Xiao Liu , Xinhao Xiang , Zizhong Li , Yongheng Wang , Zhuoheng Li , Zhuosheng Liu , Weidi Zhang , Weiqi Ye , Jiawei Zhang

This paper proposes a novel 3D speech-to-animation (STA) generation framework designed to address the shortcomings of existing models in producing diverse and emotionally resonant animations. Current STA models often generate animations…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Xulong Zhang , Xiaoyang Qu , Haoxiang Shi , Chunguang Xiao , Jianzong Wang

Driving simulators play a large role in developing and testing new intelligent vehicle systems. The visual fidelity of the simulation is critical for building vision-based algorithms and conducting human driver experiments. Low visual…

Computer Vision and Pattern Recognition · Computer Science 2022-07-22 Ekim Yurtsever , Dongfang Yang , Ibrahim Mert Koc , Keith A. Redmill

Generating 3D worlds from text is a highly anticipated goal in computer vision. Existing works are limited by the degree of exploration they allow inside of a scene, i.e., produce streched-out and noisy artifacts when moving beyond central…

Computer Vision and Pattern Recognition · Computer Science 2025-09-17 Manuel-Andreas Schneider , Lukas Höllein , Matthias Nießner

Video summarization is a crucial research area that aims to efficiently browse and retrieve relevant information from the vast amount of video content available today. With the exponential growth of multimedia data, the ability to extract…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Hai-Dang Huynh-Lam , Ngoc-Phuong Ho-Thi , Minh-Triet Tran , Trung-Nghia Le

Current 3D scene understanding methods are limited by offline-collected multi-view data or pre-constructed 3D geometry. In this paper, we present ExtractAnything3D (EA3D), a unified online framework for open-world 3D object extraction that…

Computer Vision and Pattern Recognition · Computer Science 2025-10-30 Xiaoyu Zhou , Jingqi Wang , Yuang Jia , Yongtao Wang , Deqing Sun , Ming-Hsuan Yang

Managing chronic wounds remains a major healthcare challenge, with clinical assessment often relying on subjective and time-consuming manual documentation methods. Although 2D digital videometry frameworks aided the measurement process,…

The ability to generate virtual environments is crucial for applications ranging from gaming to physical AI domains such as robotics, autonomous driving, and industrial AI. Current learning-based 3D reconstruction methods rely on the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Sherwin Bahmani , Tianchang Shen , Jiawei Ren , Jiahui Huang , Yifeng Jiang , Haithem Turki , Andrea Tagliasacchi , David B. Lindell , Zan Gojcic , Sanja Fidler , Huan Ling , Jun Gao , Xuanchi Ren

With the development of deep neural networks, the demand for a significant amount of annotated training data becomes the performance bottlenecks in many fields of research and applications. Image synthesis can generate annotated images…

Computer Vision and Pattern Recognition · Computer Science 2020-12-24 Minghui Liao , Boyu Song , Shangbang Long , Minghang He , Cong Yao , Xiang Bai

Capturing and reconstructing a human actor's motion is important for filmmaking and gaming. Currently, motion capture systems with static cameras are used for pixel-level high-fidelity reconstructions. Such setups are costly, require…

Robotics · Computer Science 2023-08-02 Qingyuan Jiang , Volkan Isler

Video representation is a long-standing problem that is crucial for various down-stream tasks, such as tracking,depth prediction,segmentation,view synthesis,and editing. However, current methods either struggle to model complex motions due…

Computer Vision and Pattern Recognition · Computer Science 2024-06-27 Yang-Tian Sun , Yi-Hua Huang , Lin Ma , Xiaoyang Lyu , Yan-Pei Cao , Xiaojuan Qi
‹ Prev 1 8 9 10 Next ›