中文
相关论文

相关论文: Sekai: A Video Dataset towards World Exploration

200 篇论文

The human hand is our primary interface to the physical world, yet egocentric perception rarely knows when, where, or how forcefully it makes contact. Robust wearable tactile sensors are scarce, and no existing in-the-wild datasets align…

This paper presents DriveTrack, a new benchmark and data generation framework for long-range keypoint tracking in real-world videos. DriveTrack is motivated by the observation that the accuracy of state-of-the-art trackers depends strongly…

计算机视觉与模式识别 · 计算机科学 2023-12-18 Arjun Balasingam , Joseph Chandler , Chenning Li , Zhoutong Zhang , Hari Balakrishnan

Recent advances in world models have demonstrated strong capabilities in simulating physical reality, making them an increasingly important foundation for embodied intelligence. For UAV agents in particular, accurate prediction of complex…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Zile Guo , Zhan Chen , Enze Zhu , Kan Wei , Yongkang Zou , Xiaoxuan Liu , Lei Wang

World engines aim to synthesize long, 3D-consistent videos that support interactive exploration of a scene under user-controlled camera motion. However, existing systems struggle under aggressive 6-DoF trajectories and complex outdoor…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Yu-Cheng Chou , Xingrui Wang , Yitong Li , Jiahao Wang , Hanting Liu , Cihang Xie , Alan Yuille , Junfei Xiao

Cutting-edge robot learning techniques including foundation models and imitation learning from humans all pose huge demands on large-scale and high-quality datasets which constitute one of the bottleneck in the general intelligent robot…

机器人学 · 计算机科学 2026-04-27 Shuo Jiang , Haonan Li , Ruochen Ren , Yanmin Zhou , Zhipeng Wang , Bin He

The large abundance of perspective camera datasets facilitated the emergence of novel learning-based strategies for various tasks, such as camera localization, single image depth estimation, or view synthesis. However, panoramic or…

计算机视觉与模式识别 · 计算机科学 2024-07-08 Kibaek Park , Francois Rameau , Jaesik Park , In So Kweon

Multi-object tracking is a classic field in computer vision. Among them, pedestrian tracking has extremely high application value and has become the most popular research category. Existing methods mainly use motion or appearance…

计算机视觉与模式识别 · 计算机科学 2025-07-04 Teng Fu , Yuwen Chen , Zhuofan Chen , Mengyang Zhao , Bin Li , Xiangyang Xue

In the era of deep learning, data is the critical determining factor in the performance of neural network models. Generating large datasets suffers from various difficulties such as scalability, cost efficiency and photorealism. To avoid…

计算机视觉与模式识别 · 计算机科学 2022-10-04 Chahat Deep Singh , Riya Kumari , Cornelia Fermüller , Nitin J. Sanket , Yiannis Aloimonos

Recording animal behaviour is an important step in evaluating the well-being of animals and further understanding the natural world. Current methods for documenting animal behaviour within a zoo setting, such as scan sampling, require…

Annotated datasets are critical for training neural networks for object detection, yet their manual creation is time- and labour-intensive, subjective to human error, and often limited in diversity. This challenge is particularly pronounced…

机器人学 · 计算机科学 2025-06-06 Aneesh Deogan , Wout Beks , Peter Teurlings , Koen de Vos , Mark van den Brand , Rene van de Molengraft

We introduce the WorldScore benchmark, the first unified benchmark for world generation. We decompose world generation into a sequence of next-scene generation tasks with explicit camera trajectory-based layout specifications, enabling…

图形学 · 计算机科学 2025-12-02 Haoyi Duan , Hong-Xing Yu , Sirui Chen , Li Fei-Fei , Jiajun Wu

Object recognition and object pose estimation in robotic grasping continue to be significant challenges, since building a labelled dataset can be time consuming and financially costly in terms of data collection and annotation. In this…

计算机视觉与模式识别 · 计算机科学 2024-01-25 Dongmyoung Lee , Wei Chen , Nicolas Rojas

Hand-drawn cartoon animation employs sketches and flat-color segments to create the illusion of motion. While recent advancements like CLIP, SVD, and Sora show impressive results in understanding and generating natural video by scaling…

计算机视觉与模式识别 · 计算机科学 2024-05-24 Zhenglin Pan

In this paper we introduce a new dataset for 360-degree video summarization: the transformation of 360-degree video content to concise 2D-video summaries that can be consumed via traditional devices, such as TV sets and smartphones. The…

计算机视觉与模式识别 · 计算机科学 2024-06-06 Ioannis Kontostathis , Evlampios Apostolidis , Vasileios Mezaris

Dynamical systems theory and reinforcement learning view world evolution as latent-state dynamics driven by actions, with visual observations providing partial information about the state. Recent video world models attempt to learn this…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Zhen Li , Zian Meng , Shuwei Shi , Wenshuo Peng , Yuwei Wu , Bo Zheng , Chuanhao Li , Kaipeng Zhang

Educational recommenders have received much less attention in comparison to e-commerce and entertainment-related recommenders, even though efficient intelligent tutors have great potential to improve learning gains. One of the main…

信息检索 · 计算机科学 2021-09-15 Sahan Bulathwela , Maria Perez-Ortiz , Erik Novak , Emine Yilmaz , John Shawe-Taylor

We introduce DAiSEE, the first multi-label video classification dataset comprising of 9068 video snippets captured from 112 users for recognizing the user affective states of boredom, confusion, engagement, and frustration in the wild. The…

计算机视觉与模式识别 · 计算机科学 2022-07-08 Abhay Gupta , Arjun D'Cunha , Kamal Awasthi , Vineeth Balasubramanian

Walking has always been a primary mode of transportation and is recognized as an essential activity for maintaining good health. Despite the need for safe walking conditions in urban environments, sidewalks are frequently obstructed by…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Marios Thoma , Zenonas Theodosiou , Harris Partaourides , Vassilis Vassiliades , Loizos Michael , Andreas Lanitis

Egocentric "walking tour" videos provide a rich source of image data to develop rich and diverse visual models of environments around the world. However, the significant presence of humans in frames of these videos due to crowds and…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Yujin Ham , Junho Kim , Vivek Boominathan , Guha Balakrishnan

Physics-aware driving world model is essential for drive planning, out-of-distribution data synthesis, and closed-loop evaluation. However, existing methods often rely on a single diffusion model to directly map driving actions to videos,…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Zhenya Yang , Zhe Liu , Yuxiang Lu , Liping Hou , Chenxuan Miao , Siyi Peng , Bailan Feng , Xiang Bai , Hengshuang Zhao