中文
相关论文

相关论文: VISEM-Tracking, a human spermatozoa tracking datas…

200 篇论文

Inferring physical actions from visual observations is a fundamental capability for advancing machine intelligence in the physical world. Achieving this requires large-scale, open-vocabulary video action datasets that span broad domains. We…

计算机视觉与模式识别 · 计算机科学 2026-01-16 Delong Chen , Tejaswi Kasarla , Yejin Bang , Mustafa Shukor , Willy Chung , Jade Yu , Allen Bolourchi , Theo Moutakanni , Pascale Fung

Extracting physical dynamical system parameters from recorded observations is key in natural science. Current methods for automatic parameter estimation from video train supervised deep networks on large datasets. Such datasets require…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Alejandro Castañeda Garcia , Jan van Gemert , Daan Brinks , Nergis Tömen

Academic integrity continues to face the persistent challenge of examination cheating. Traditional invigilation relies on human observation, which is inefficient, costly, and prone to errors at scale. Although some existing AI-powered…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Van-Truong Le , Le-Khanh Nguyen , Trong-Doanh Nguyen

Image-based salient object detection (SOD) has been extensively studied in the past decades. However, video-based SOD is much less explored since there lack large-scale video datasets within which salient objects are unambiguously defined…

计算机视觉与模式识别 · 计算机科学 2017-05-10 Jia Li , Changqun Xia , Xiaowu Chen

The proper enforcement of motorcycle helmet regulations is crucial for ensuring the safety of motorbike passengers and riders, as roadway cyclists and passengers are not likely to abide by these regulations if no proper enforcement systems…

计算机视觉与模式识别 · 计算机科学 2023-04-20 Geoffery Agorku , Divine Agbobli , Vuban Chowdhury , Kwadwo Amankwah-Nkyi , Adedolapo Ogungbire , Portia Ankamah Lartey , Armstrong Aboah

Despite the significant impact of visual events on human cognition, understanding events in videos remains a challenging task for AI due to their complex structures, semantic hierarchies, and dynamic evolution. To address this, we propose…

计算机视觉与模式识别 · 计算机科学 2025-06-04 Baoyu Liang , Qile Su , Shoutai Zhu , Yuchen Liang , Chao Tong

The development of video large multimodal models (LMMs) has been hindered by the difficulty of curating large amounts of high-quality raw data from the web. To address this, we propose an alternative approach by creating a high-quality…

计算机视觉与模式识别 · 计算机科学 2025-08-04 Yuanhan Zhang , Jinming Wu , Wei Li , Bo Li , Zejun Ma , Ziwei Liu , Chunyuan Li

Autonomous driving, particularly navigating complex and unanticipated scenarios, demands sophisticated reasoning and planning capabilities. While Multi-modal Large Language Models (MLLMs) offer a promising avenue for this, their use has…

计算机视觉与模式识别 · 计算机科学 2025-10-15 Hidehisa Arai , Keita Miwa , Kento Sasaki , Yu Yamaguchi , Kohei Watanabe , Shunsuke Aoki , Issei Yamamoto

Driver drowsiness detection has been the subject of many researches in the past few decades and various methods have been developed to detect it. In this study, as an image-based approach with adequate accuracy, along with the expedite…

计算机视觉与模式识别 · 计算机科学 2021-04-02 Farnoosh Faraji , Faraz Lotfi , Javad Khorramdel , Ali Najafi , Ali Ghaffari

The goal of this study was to develop new reliable open surgery suturing simulation system for training medical students in situation where resources are limited or in the domestic setup. Namely, we developed an algorithm for tools and…

计算机视觉与模式识别 · 计算机科学 2022-01-11 Adam Goldbraikh , Anne-Lise D'Angelo , Carla M. Pugh , Shlomi Laufer

Procedural video understanding is gaining attention in the vision and language community. Deep learning-based video analysis requires extensive data. Consequently, existing works often use web videos as training resources, making it…

计算机视觉与模式识别 · 计算机科学 2024-08-06 Koki Maeda , Tosho Hirasawa , Atsushi Hashimoto , Jun Harashima , Leszek Rybicki , Yusuke Fukasawa , Yoshitaka Ushiku

Eye-tracking applications that utilize the human gaze in video understanding tasks have become increasingly important. To effectively automate the process of video analysis based on eye-tracking data, it is important to accurately replicate…

计算机视觉与模式识别 · 计算机科学 2024-04-15 Suleyman Ozdel , Yao Rong , Berat Mert Albaba , Yen-Ling Kuo , Xi Wang , Enkelejda Kasneci

Vision-based sign language recognition aims at helping deaf people to communicate with others. However, most existing sign language datasets are limited to a small number of words. Due to the limited vocabulary size, models learned from…

计算机视觉与模式识别 · 计算机科学 2020-01-22 Dongxu Li , Cristian Rodriguez Opazo , Xin Yu , Hongdong Li

Within medical imaging, manual curation of sufficient well-labeled samples is cost, time and scale-prohibitive. To improve the representativeness of the training dataset, for the first time, we present an approach to utilize large amounts…

计算机视觉与模式识别 · 计算机科学 2019-06-03 Fernando Navarro , Sailesh Conjeti , Federico Tombari , Nassir Navab

We study joint video and language (VL) pre-training to enable cross-modality learning and benefit plentiful downstream VL tasks. Existing works either extract low-quality video features or learn limited text embedding, while neglecting that…

计算机视觉与模式识别 · 计算机科学 2022-07-11 Hongwei Xue , Tiankai Hang , Yanhong Zeng , Yuchong Sun , Bei Liu , Huan Yang , Jianlong Fu , Baining Guo

Long-term video understanding requires interpreting complex temporal events and reasoning over procedural activities. While instructional video corpora, like HowTo100M, offer rich resources for model training, they present significant…

计算机视觉与模式识别 · 计算机科学 2026-04-30 Mingji Ge , Qirui Chen , Zeqian Li , Weidi Xie

Ball trajectory data are one of the most fundamental and useful information in the evaluation of players' performance and analysis of game strategies. Although vision-based object tracking techniques have been developed to analyze sport…

机器学习 · 计算机科学 2019-07-09 Yu-Chuan Huang , I-No Liao , Ching-Hsuan Chen , Tsì-Uí İk , Wen-Chih Peng

Computer-use agents can operate computers and automate laborious tasks, but despite recent rapid progress, they still lag behind human users, especially when tasks require domain-specific procedural knowledge about particular applications,…

计算机视觉与模式识别 · 计算机科学 2025-11-07 Yujian Liu , Ze Wang , Hao Chen , Ximeng Sun , Xiaodong Yu , Jialian Wu , Jiang Liu , Emad Barsoum , Zicheng Liu , Shiyu Chang

In this work, the novel task of detecting and classifying table tennis strokes solely using the ball trajectory has been explored. A single camera setup positioned in the umpire's view has been employed to procure a dataset consisting of…

计算机视觉与模式识别 · 计算机科学 2023-02-21 Kaustubh Milind Kulkarni , Rohan S Jamadagni , Jeffrey Aaron Paul , Sucheth Shenoy

Spatio-temporal action detection is an important and challenging problem in video understanding. The existing action detection benchmarks are limited in aspects of small numbers of instances in a trimmed video or low-level atomic actions.…

计算机视觉与模式识别 · 计算机科学 2021-08-19 Yixuan Li , Lei Chen , Runyu He , Zhenzhi Wang , Gangshan Wu , Limin Wang