中文
相关论文

相关论文: SpatialVID: A Large-Scale Video Dataset with Spati…

200 篇论文

Stereoscopic video has long been the subject of research due to its capacity to deliver immersive three-dimensional content across a wide range of applications, from virtual and augmented reality to advanced human-computer interaction. The…

Understanding relations between objects is crucial for understanding the semantics of a visual scene. It is also an essential step in order to bridge visual and language models. However, current state-of-the-art computer vision models still…

计算机视觉与模式识别 · 计算机科学 2025-02-28 Palaash Agrawal , Haidi Azaman , Cheston Tan

We present the Moments in Time Dataset, a large-scale human-annotated collection of one million short videos corresponding to dynamic events unfolding within three seconds. Modeling the spatial-audio-temporal dynamics even for actions…

计算机视觉与模式识别 · 计算机科学 2019-02-19 Mathew Monfort , Alex Andonian , Bolei Zhou , Kandan Ramakrishnan , Sarah Adel Bargal , Tom Yan , Lisa Brown , Quanfu Fan , Dan Gutfruend , Carl Vondrick , Aude Oliva

Video depth estimation has long been hindered by the scarcity of consistent and scalable ground truth data, leading to inconsistent and unreliable results. In this paper, we introduce Depth Any Video, a model that tackles the challenge…

计算机视觉与模式识别 · 计算机科学 2025-03-13 Honghui Yang , Di Huang , Wei Yin , Chunhua Shen , Haifeng Liu , Xiaofei He , Binbin Lin , Wanli Ouyang , Tong He

Accurate 3D geometric perception is an important prerequisite for a wide range of spatial AI systems. While state-of-the-art methods depend on large-scale training data, acquiring consistent and precise 3D annotations from in-the-wild…

Spatial intelligence is emerging as a transformative frontier in AI, yet it remains constrained by the scarcity of large-scale 3D datasets. Unlike the abundant 2D imagery, acquiring 3D data typically requires specialized sensors and…

计算机视觉与模式识别 · 计算机科学 2025-07-28 Xingyu Miao , Haoran Duan , Quanhao Qian , Jiuniu Wang , Yang Long , Ling Shao , Deli Zhao , Ran Xu , Gongjie Zhang

We propose a method for annotating videos of complex multi-object scenes with a globally-consistent 3D representation of the objects. We annotate each object with a CAD model from a database, and place it in the 3D coordinate frame of the…

计算机视觉与模式识别 · 计算机科学 2023-08-15 Kevis-Kokitsi Maninis , Stefan Popov , Matthias Nießner , Vittorio Ferrari

Despite impressive high-level video comprehension, multimodal language models struggle with spatial reasoning across time and space. While current spatial training approaches rely on real-world video data, obtaining diverse footage with…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Ellis Brown , Arijit Ray , Ranjay Krishna , Ross Girshick , Rob Fergus , Saining Xie

Video action detection requires dense spatio-temporal annotations, which are both challenging and expensive to obtain. However, real-world videos often vary in difficulty and may not require the same level of annotation. This paper analyzes…

计算机视觉与模式识别 · 计算机科学 2025-08-20 Aayush Rana , Akash Kumar , Vibhav Vineet , Yogesh S Rawat

Recent advances in camera-controllable video generation have been constrained by the reliance on static-scene datasets with relative-scale camera annotations, such as RealEstate10K. While these datasets enable basic viewpoint control, they…

计算机视觉与模式识别 · 计算机科学 2025-04-14 Guangcong Zheng , Teng Li , Xianpan Zhou , Xi Li

We explore spatiotemporal data augmentation using video foundation models to diversify both camera viewpoints and scene dynamics. Unlike existing approaches based on simple geometric transforms or appearance perturbations, our method…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Jinfan Zhou , Lixin Luo , Sungmin Eum , Heesung Kwon , Jeong Joon Park

We present a dataset of 998 3D models of everyday tabletop objects along with their 847,000 real world RGB and depth images. Accurate annotations of camera poses and object poses for each image are performed in a semi-automated fashion to…

计算机视觉与模式识别 · 计算机科学 2022-08-10 Rakesh Shrestha , Siqi Hu , Minghao Gou , Ziyuan Liu , Ping Tan

Human image animation involves generating videos from a character photo, allowing user control and unlocking the potential for video and movie production. While recent approaches yield impressive results using high-quality training data,…

计算机视觉与模式识别 · 计算机科学 2024-11-22 Zhenzhi Wang , Yixuan Li , Yanhong Zeng , Youqing Fang , Yuwei Guo , Wenran Liu , Jing Tan , Kai Chen , Tianfan Xue , Bo Dai , Dahua Lin

Videos that are shot using commodity hardware such as phones and surveillance cameras record various metadata such as time and location. We encounter such geospatial videos on a daily basis and such videos have been growing in volume…

数据库 · 计算机科学 2024-07-16 Chanwut Kittivorawong , Yongming Ge , Yousef Helal , Alvin Cheung

In this work, we introduce a dataset of video annotated with high quality natural language phrases describing the visual content in a given segment of time. Our dataset is based on the Descriptive Video Service (DVS) that is now encoded on…

计算机视觉与模式识别 · 计算机科学 2015-03-04 Atousa Torabi , Christopher Pal , Hugo Larochelle , Aaron Courville

The large abundance of perspective camera datasets facilitated the emergence of novel learning-based strategies for various tasks, such as camera localization, single image depth estimation, or view synthesis. However, panoramic or…

计算机视觉与模式识别 · 计算机科学 2024-07-08 Kibaek Park , Francois Rameau , Jaesik Park , In So Kweon

Violence Detection (VD) has become an increasingly vital area of research. Existing automated VD efforts are hindered by the limited availability of diverse, well-annotated databases. Existing databases suffer from coarse video-level…

计算机视觉与模式识别 · 计算机科学 2025-06-09 Dimitrios Kollias , Damith C. Senadeera , Jianian Zheng , Kaushal K. K. Yadav , Greg Slabaugh , Muhammad Awais , Xiaoyun Yang

The recent and increasing interest in video-language research has driven the development of large-scale datasets that enable data-intensive machine learning techniques. In comparison, limited effort has been made at assessing the fitness of…

计算机视觉与模式识别 · 计算机科学 2022-03-29 Mattia Soldan , Alejandro Pardo , Juan León Alcázar , Fabian Caba Heilbron , Chen Zhao , Silvio Giancola , Bernard Ghanem

The application of deep learning to nursing procedure activity understanding has the potential to greatly enhance the quality and safety of nurse-patient interactions. By utilizing the technique, we can facilitate training and education,…

计算机视觉与模式识别 · 计算机科学 2023-10-23 Ming Hu , Lin Wang , Siyuan Yan , Don Ma , Qingli Ren , Peng Xia , Wei Feng , Peibo Duan , Lie Ju , Zongyuan Ge

We present the HANDAL dataset for category-level object pose estimation and affordance prediction. Unlike previous datasets, ours is focused on robotics-ready manipulable objects that are of the proper size and shape for functional grasping…

机器人学 · 计算机科学 2023-08-04 Andrew Guo , Bowen Wen , Jianhe Yuan , Jonathan Tremblay , Stephen Tyree , Jeffrey Smith , Stan Birchfield
‹ 上一页 1 2 3 10 下一页 ›