English
Related papers

Related papers: CaptainCook4D: A Dataset for Understanding Errors …

200 papers

Wearable cameras allow to acquire images and videos from the user's perspective. These data can be processed to understand humans behavior. Despite human behavior analysis has been thoroughly investigated in third person vision, it is still…

Computer Vision and Pattern Recognition · Computer Science 2023-07-06 Francesco Ragusa , Antonino Furnari , Giovanni Maria Farinella

Deepfakes represent a growing concern across domains such as disinformation, fraud, and non-consensual media. In particular, the rise of video conference and identity-driven attacks in high-stakes scenarios--such as impostor hiring--demands…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Sarah Barrington , Maty Bohacek , Hany Farid

Food image segmentation is a critical task for dietary analysis, enabling accurate estimation of food volume and nutrients. However, current methods suffer from limited multi-view data and poor generalization to new viewpoints. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Ahmad AlMughrabi , Guillermo Rivo , Carlos Jiménez-Farfán , Umair Haroon , Farid Al-Areqi , Hyunjun Jung , Benjamin Busam , Ricardo Marques , Petia Radeva

Scene understanding is essential in determining how intelligent robotic grasping and manipulation could get. It is a problem that can be approached using different techniques: seen object segmentation, unseen object segmentation, or 6D pose…

Robotics · Computer Science 2022-11-29 Anas Gouda , Abraham Ghanem , Christopher Reining

In recent years, automatic video caption generation has attracted considerable attention. This paper focuses on the generation of Japanese captions for describing human actions. While most currently available video caption datasets have…

Computation and Language · Computer Science 2020-03-11 Yutaro Shigeto , Yuya Yoshikawa , Jiaqing Lin , Akikazu Takeuchi

Maintaining a healthy lifestyle has become increasingly challenging in today's sedentary society marked by poor eating habits. To address this issue, both national and international organisations have made numerous efforts to promote…

In the field of robotic manipulation, deep imitation learning is recognized as a promising approach for acquiring manipulation skills. Additionally, learning from diverse robot datasets is considered a viable method to achieve versatility…

Robotics · Computer Science 2024-03-20 Heecheol Kim , Yoshiyuki Ohmura , Yasuo Kuniyoshi

A person's movement or relative positioning can be effectively captured by different types of sensors and corresponding sensor output can be utilized in various manipulative techniques for the classification of different human activities.…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Utsab Saha , Sawradip Saha , Tahmid Kabir , Shaikh Anowarul Fattah , Mohammad Saquib

Generating long form narratives such as stories and procedures from multiple modalities has been a long standing dream for artificial intelligence. In this regard, there is often crucial subtext that is derived from the surrounding…

Computation and Language · Computer Science 2020-10-28 Khyathi Raghavi Chandu , Ruo-Ping Dong , Alan Black

Daily Activity Recordings for Artificial Intelligence (DARai, pronounced "Dahr-ree") is a multimodal, hierarchically annotated dataset constructed to understand human activities in real-world settings. DARai consists of continuous scripted…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Ghazal Kaviani , Yavuz Yarici , Seulgi Kim , Mohit Prabhushankar , Ghassan AlRegib , Mashhour Solh , Ameya Patil

The development and evaluation of machine vision in underwater environments remains challenging, often relying on trial-and-error-based testing tailored to specific applications. This is partly due to the lack of controlled, ground-truthed…

Computer Vision and Pattern Recognition · Computer Science 2025-09-25 Patricia Schöntag , David Nakath , Judith Fischer , Rüdiger Röttgers , Kevin Köser

In this work, we take aim towards increasing the effectiveness of surgical assistant robots. We intended to make assistant robots safer by making them aware about the actions of surgeon, so it can take appropriate assisting actions. In…

Current large-scale video datasets focus on general human activity, but lack depth of coverage on fine-grained activities needed to address physical skill learning. We introduce SportSkills, the first large-scale sports dataset geared…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Kumar Ashutosh , Chi Hsuan Wu , Kristen Grauman

Data-driven navigation algorithms are critically dependent on large-scale, high-quality real-world data collection for successful training and robust performance in realistic and uncontrolled conditions. To enhance the growing family of…

Complex activities in real-world audio unfold over extended durations and exhibit hierarchical structure, yet most prior work focuses on short clips and isolated events. To bridge this gap, we introduce MultiAct, a new dataset and benchmark…

Sound · Computer Science 2026-02-09 Peng Zhang , Qingyu Luo , Philip J. B. Jackson , Wenwu Wang

Progress in embodied intelligence increasingly depends on scalable data infrastructure. While vision and language have scaled with internet corpora, learning physical interaction remains constrained by the lack of large, diverse, and richly…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Yufan Deng , Daquan Zhou

The kinetics and dynamics of drug-protein binding and dissociation are crucial to understanding drug absorption and metabolism. Despite advances in artificial intelligence (AI) tools for drug-protein interaction studies, existing training…

Computational Physics · Physics 2026-02-17 Maodong Li , Jiying Zhang , Zhe Wang , Bin Feng , Wenqi Zeng , Dechin Chen , Zhijun Pan , Yu Li , Zijing Liu , Yi Isaac Yang

This paper addresses the long-standing challenge of reconstructing 3D structures from videos with dynamic content. Current approaches to this problem were not designed to operate on casual videos recorded by standard cameras or require a…

Computer Vision and Pattern Recognition · Computer Science 2024-06-28 Yoni Kasten , Wuyue Lu , Haggai Maron

Segmenting long videos into chapters enables users to quickly navigate to the information of their interest. This important topic has been understudied due to the lack of publicly released datasets. To address this issue, we present…

Computer Vision and Pattern Recognition · Computer Science 2023-09-26 Antoine Yang , Arsha Nagrani , Ivan Laptev , Josef Sivic , Cordelia Schmid

Spatio-temporal action detection is an important and challenging problem in video understanding. The existing action detection benchmarks are limited in aspects of small numbers of instances in a trimmed video or low-level atomic actions.…

Computer Vision and Pattern Recognition · Computer Science 2021-08-19 Yixuan Li , Lei Chen , Runyu He , Zhenzhi Wang , Gangshan Wu , Limin Wang
‹ Prev 1 8 9 10 Next ›