中文
相关论文

相关论文: ProSkill: Segment-Level Skill Assessment in Proced…

200 篇论文

Five billion people in the world lack access to quality surgical care. Surgeon skill varies dramatically, and many surgical patients suffer complications and avoidable harm. Improving surgical training and feedback would help to reduce the…

计算机视觉与模式识别 · 计算机科学 2018-07-25 Amy Jin , Serena Yeung , Jeffrey Jopling , Jonathan Krause , Dan Azagury , Arnold Milstein , Li Fei-Fei

Video is a promising source of knowledge for embodied agents to learn models of the world's dynamics. Large deep networks have become increasingly effective at modeling complex video data in a self-supervised manner, as evaluated by metrics…

计算机视觉与模式识别 · 计算机科学 2023-04-27 Stephen Tian , Chelsea Finn , Jiajun Wu

Instructional videos provide a convenient modality to learn new tasks (ex. cooking a recipe, or assembling furniture). A viewer will want to find a corresponding video that reflects both the overall task they are interested in as well as…

计算机视觉与模式识别 · 计算机科学 2025-02-25 Karan Samel , Nitish Sontakke , Irfan Essa

We propose a novel multimodal video benchmark - the Perception Test - to evaluate the perception and reasoning skills of pre-trained multimodal models (e.g. Flamingo, SeViLA, or GPT-4). Compared to existing benchmarks that focus on…

Point-level supervised temporal action localization (PTAL) aims at recognizing and localizing actions in untrimmed videos where only a single point (frame) within every action instance is annotated in training data. Without temporal…

计算机视觉与模式识别 · 计算机科学 2023-10-10 Yuan Yin , Yifei Huang , Ryosuke Furuta , Yoichi Sato

Action recognition is an important and challenging problem in video analysis. Although the past decade has witnessed progress in action recognition with the development of deep learning, such process has been slow in competitive sports…

计算机视觉与模式识别 · 计算机科学 2020-02-11 Shenlan Liu , Xiang Liu , Gao Huang , Lin Feng , Lianyu Hu , Dong Jiang , Aibin Zhang , Yang Liu , Hong Qiao

Watching instructional videos are often used to learn about procedures. Video captioning is one way of automatically collecting such knowledge. However, it provides only an indirect, overall evaluation of multimodal models with no…

计算与语言 · 计算机科学 2020-10-12 Frank F. Xu , Lei Ji , Botian Shi , Junyi Du , Graham Neubig , Yonatan Bisk , Nan Duan

Temporal Action Localization (TAL) aims to predict both action category and temporal boundary of action instances in untrimmed videos, i.e., start and end time. Fully-supervised solutions are usually adopted in most existing works, and…

计算机视觉与模式识别 · 计算机科学 2022-12-02 Ding Li , Xuebing Yang , Yongqiang Tang , Chenyang Zhang , Wensheng Zhang

Video processing is generally divided into two main categories: processing of the entire video, which typically yields optimal classification outcomes, and real-time processing, where the objective is to make a decision as promptly as…

计算机视觉与模式识别 · 计算机科学 2025-08-14 Magdalena Trędowicz , Marcin Mazur , Szymon Janusz , Arkadiusz Lewicki , Jacek Tabor , Łukasz Struski

Whole slide image (WSI) classification is a critical task in computational pathology, requiring the processing of gigapixel-sized images, which is challenging for current deep-learning methods. Current state of the art methods are based on…

计算机视觉与模式识别 · 计算机科学 2023-10-06 Jingwei Zhang , Saarthak Kapse , Ke Ma , Prateek Prasanna , Joel Saltz , Maria Vakalopoulou , Dimitris Samaras

Video summarization is a technique to create a short skim of the original video while preserving the main stories/content. There exists a substantial interest in automatizing this process due to the rapid growth of the available material.…

计算机视觉与模式识别 · 计算机科学 2019-04-12 Mayu Otani , Yuta Nakashima , Esa Rahtu , Janne Heikkilä

There are substantial instructional videos on the Internet, which provide us tutorials for completing various tasks. Existing instructional video datasets only focus on specific steps at the video level, lacking experiential guidelines at…

计算机视觉与模式识别 · 计算机科学 2024-06-27 Jiafeng Liang , Shixin Jiang , Zekun Wang , Haojie Pan , Zerui Chen , Zheng Chu , Ming Liu , Ruiji Fu , Zhongyuan Wang , Bing Qin

While Large Language Models (LLMs) have evolved into tool-using agents, they remain brittle in long-horizon interactions. Unlike mathematical reasoning where errors are often rectifiable via backtracking, tool-use failures frequently induce…

Human poses and motions are important cues for analysis of videos with people and there is strong evidence that representations based on body pose are highly effective for a variety of tasks such as activity recognition, content retrieval…

计算机视觉与模式识别 · 计算机科学 2018-04-12 Mykhaylo Andriluka , Umar Iqbal , Eldar Insafutdinov , Leonid Pishchulin , Anton Milan , Juergen Gall , Bernt Schiele

Predictive process monitoring is a subfield of process mining that aims to estimate case or event features for running process instances. Such predictions are of significant interest to the process stakeholders. However, most of the…

Skills play a central role in the job market and many human resources (HR) processes. In the wake of other digital experiences, today's online job market has candidates expecting to see the right opportunities based on their skill set.…

计算与语言 · 计算机科学 2022-09-14 Jens-Joris Decorte , Jeroen Van Hautte , Johannes Deleu , Chris Develder , Thomas Demeester

We summarise popular methods used for skill rating in competitive sports, along with their inferential paradigms and introduce new approaches based on sequential Monte Carlo and discrete hidden Markov models. We advocate for a state-space…

应用统计 · 统计学 2024-08-20 Samuel Duffield , Samuel Power , Lorenzo Rimella

Proficiency in microanastomosis is a critical surgical skill in neurosurgery, where the ability to precisely manipulate fine instruments is crucial to successful outcomes. These procedures require sustained attention, coordinated hand…

计算机视觉与模式识别 · 计算机科学 2026-01-01 Yan Meng , Daniel Donoho , Marcelle Altshuler , Omar Arnaout

Procedural texts help AI enhance reasoning about context and action sequences. Transforming these into Semantic Role Labeling (SRL) improves understanding of individual steps by identifying predicate-argument structure like…

计算与语言 · 计算机科学 2025-05-28 Anil Batra , Laura Sevilla-Lara , Marcus Rohrbach , Frank Keller

Despite recent advances in AI, the development of systems capable of executing complex, multi-step reasoning tasks involving multiple tools remains a significant challenge. Current benchmarks fall short in capturing the real-world…

计算与语言 · 计算机科学 2025-01-03 Vaskar Nath , Pranav Raja , Claire Yoon , Sean Hendryx