English
Related papers

Related papers: PhysLab: A Benchmark Dataset for Multi-Granularity…

200 papers

The paper develops datasets and methods to assess student participation in real-life collaborative learning environments. In collaborative learning environments, students are organized into small groups where they are free to interact…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Wenjing Shi , Phuong Tran , Sylvia Celedón-Pattichis , Marios S. Pattichis

Despite progress in video large language models (Video-LLMs), research on instructional video understanding, crucial for enhancing access to instructional content, remains insufficient. To address this, we introduce InstructionBench, an…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Haiwan Wei , Yitian Yuan , Xiaohan Lan , Wei Ke , Lin Ma

Spatio-temporal action detection is an important and challenging problem in video understanding. The existing action detection benchmarks are limited in aspects of small numbers of instances in a trimmed video or low-level atomic actions.…

Computer Vision and Pattern Recognition · Computer Science 2021-08-19 Yixuan Li , Lei Chen , Runyu He , Zhenzhi Wang , Gangshan Wu , Limin Wang

Understanding fine-grained human hand motion is fundamental to visual perception, embodied intelligence, and multimodal communication. In this work, we propose Fine-grained Finger-level Hand Motion Captioning (FingerCap), which aims to…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Xin Shen , Rui Zhu , Lei Shen , Xinyu Wang , Kaihao Zhang , Tianqing Zhu , Shuchen Wu , Chenxi Miao , Weikang Li , Yang Li , Deguo Xia , Jizhou Huang , Xin Yu

Modeling and rendering photorealistic avatars is of crucial importance in many applications. Existing methods that build a 3D avatar from visual observations, however, struggle to reconstruct clothed humans. We introduce PhysAvatar, a novel…

Measuring the perception of visual content is a long-standing problem in computer vision. Many mathematical models have been developed to evaluate the look or quality of an image. Despite the effectiveness of such tools in quantifying…

Computer Vision and Pattern Recognition · Computer Science 2022-11-24 Jianyi Wang , Kelvin C. K. Chan , Chen Change Loy

The application of deep learning to nursing procedure activity understanding has the potential to greatly enhance the quality and safety of nurse-patient interactions. By utilizing the technique, we can facilitate training and education,…

Computer Vision and Pattern Recognition · Computer Science 2023-10-23 Ming Hu , Lin Wang , Siyuan Yan , Don Ma , Qingli Ren , Peng Xia , Wei Feng , Peibo Duan , Lie Ju , Zongyuan Ge

Recent progress in imitation learning from human demonstrations has shown promising results in teaching robots manipulation skills. To further scale up training datasets, recent works start to use portable data collection devices without…

Robotics · Computer Science 2024-10-14 Sirui Chen , Chen Wang , Kaden Nguyen , Li Fei-Fei , C. Karen Liu

Flowcharts are graphical tools for representing complex concepts in concise visual representations. This paper introduces the FlowLearn dataset, a resource tailored to enhance the understanding of flowcharts. FlowLearn contains complex…

Computer Vision and Pattern Recognition · Computer Science 2024-07-11 Huitong Pan , Qi Zhang , Cornelia Caragea , Eduard Dragut , Longin Jan Latecki

Academic emotion analysis plays a crucial role in evaluating students' engagement and cognitive states during the learning process. This paper addresses the challenge of automatically recognizing academic emotions through facial expressions…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Luming Zhao , Jingwen Xuan , Jiamin Lou , Yonghui Yu , Wenwu Yang

Despite an exciting new wave of multimodal machine learning models, current approaches still struggle to interpret the complex contextual relationships between the different modalities present in videos. Going beyond existing methods that…

Computer Vision and Pattern Recognition · Computer Science 2023-09-20 Laura Hanu , Anita L. Verő , James Thewlis

Understanding the dynamic physical world, characterized by its evolving 3D structure, real-world motion, and semantic content with textual descriptions, is crucial for human-agent interaction and enables embodied agents to perceive and act…

While automatic monitoring and coaching of exercises are showing encouraging results in non-medical applications, they still have limitations such as errors and limited use contexts. To allow the development and assessment of physical…

Machine Learning · Computer Science 2025-01-14 Sao Mai Nguyen , Maxime Devanne , Olivier Remy-Neris , Mathieu Lempereur , André Thepaut

Body-worn first-person vision (FPV) camera enables to extract a rich source of information on the environment from the subject's viewpoint. However, the research progress in wearable camera-based egocentric office activity understanding is…

Computer Vision and Pattern Recognition · Computer Science 2022-09-13 Girmaw Abebe Tadesse , Oliver Bent , Komminist Weldemariam , Md. Abrar Istiak , Taufiq Hasan , Andrea Cavallaro

Multisensory object-centric perception, reasoning, and interaction have been a key research topic in recent years. However, the progress in these directions is limited by the small set of objects available -- synthetic objects are not…

Robotics · Computer Science 2021-11-09 Ruohan Gao , Yen-Yu Chang , Shivani Mall , Li Fei-Fei , Jiajun Wu

Point tracking models often struggle to generalize to real-world videos because large-scale training data is predominantly synthetic$\unicode{x2014}$the only source currently feasible to produce at scale. Collecting real-world annotations,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Inès Hyeonsu Kim , Seokju Cho , Jahyeok Koo , Junghyun Park , Jiahui Huang , Honglak Lee , Joon-Young Lee , Seungryong Kim

Understanding objects in videos in terms of fine-grained localization masks and detailed semantic properties is a fundamental task in video understanding. In this paper, we propose VoCap, a flexible video model that consumes a video and a…

Computer Vision and Pattern Recognition · Computer Science 2025-09-01 Jasper Uijlings , Xingyi Zhou , Xiuye Gu , Arsha Nagrani , Anurag Arnab , Alireza Fathi , David Ross , Cordelia Schmid

Low-light videos often exhibit spatiotemporal incoherent noise, leading to poor visibility and compromised performance across various computer vision applications. One significant challenge in enhancing such content using modern…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Nantheera Anantrasirichai , Ruirui Lin , Alexandra Malyugina , David Bull

Distilling knowledge from human demonstrations is a promising way for robots to learn and act. Existing methods, which often rely on coarsely-aligned video pairs, are typically constrained to learning global or task-level features. As a…

Robotics · Computer Science 2025-11-18 Sicheng Xie , Haidong Cao , Zejia Weng , Zhen Xing , Haoran Chen , Shiwei Shen , Jiaqi Leng , Zuxuan Wu , Yu-Gang Jiang

Data is the foundation for the development of computer vision, and the establishment of datasets plays an important role in advancing the techniques of fine-grained visual categorization~(FGVC). In the existing FGVC datasets used in…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Shuo Ye , Shiming Chen , Ruxin Wang , Tianxu Wu , Jiamiao Xu , Salman Khan , Fahad Shahbaz Khan , Ling Shao