中文
相关论文

相关论文: Learning Skill-Attributes for Transferable Assessm…

200 篇论文

Recent advances in computer vision have made it possible to automatically assess from videos the manipulation skills of humans in performing a task, which breeds many important applications in domains such as health rehabilitation and…

计算机视觉与模式识别 · 计算机科学 2019-04-11 Zhenqiang Li , Yifei Huang , Minjie Cai , Yoichi Sato

Assessing human skill levels in complex activities is a challenging problem with applications in sports, rehabilitation, and training. In this work, we present SkillFormer, a parameter-efficient architecture for unified multi-view…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Edoardo Bianchi , Antonio Liotta

Current large-scale video datasets focus on general human activity, but lack depth of coverage on fine-grained activities needed to address physical skill learning. We introduce SportSkills, the first large-scale sports dataset geared…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Kumar Ashutosh , Chi Hsuan Wu , Kristen Grauman

Learning manipulation skills from human demonstration videos presents a promising yet challenging problem, primarily due to the significant embodiment gap between human body and robot manipulators. Existing methods rely on paired datasets…

机器人学 · 计算机科学 2025-10-10 YuHang Tang , Yixuan Lou , Pengfei Han , Haoming Song , Xinyi Ye , Dong Wang , Bin Zhao

Human demonstration videos are a widely available data source for robot learning and an intuitive user interface for expressing desired behavior. However, directly extracting reusable robot manipulation skills from unstructured human videos…

机器人学 · 计算机科学 2023-10-02 Mengda Xu , Zhenjia Xu , Cheng Chi , Manuela Veloso , Shuran Song

Pre-trained language models are still far from human performance in tasks that need understanding of properties (e.g. appearance, measurable quantity) and affordances of everyday objects in the real world since the text lacks such…

计算与语言 · 计算机科学 2022-03-18 Woojeong Jin , Dong-Ho Lee , Chenguang Zhu , Jay Pujara , Xiang Ren

Skill assessment in procedural videos is crucial for the objective evaluation of human performance in settings such as manufacturing and procedural daily tasks. Current research on skill assessment has predominantly focused on sports and…

计算机视觉与模式识别 · 计算机科学 2026-01-29 Michele Mazzamuto , Daniele Di Mauro , Gianpiero Francesca , Giovanni Maria Farinella , Antonino Furnari

The objective of this work is to learn an object-centric video representation, with the aim of improving transferability to novel tasks, i.e., tasks different from the pre-training task of action classification. To this end, we introduce a…

计算机视觉与模式识别 · 计算机科学 2022-10-11 Chuhan Zhang , Ankush Gupta , Andrew Zisserman

We present Skill Transformer, an approach for solving long-horizon robotic tasks by combining conditional sequence modeling and skill modularity. Conditioned on egocentric and proprioceptive observations of a robot, Skill Transformer is…

机器学习 · 计算机科学 2023-08-22 Xiaoyu Huang , Dhruv Batra , Akshara Rai , Andrew Szot

We address the task of text translation on the How2 dataset using a state of the art transformer-based multimodal approach. The question we ask ourselves is whether visual features can support the translation process, in particular, given…

计算与语言 · 计算机科学 2019-08-20 Zixiu Wu , Julia Ive , Josiah Wang , Pranava Madhyastha , Lucia Specia

Recent end-to-end robotic manipulation research increasingly adopts architectures inspired by large language models to enable robust manipulation. However, a critical challenge arises from severe distribution shifts between robotic action…

机器人学 · 计算机科学 2025-12-10 Yuchi Zhang , Churui Sun , Shiqi Liang , Diyuan Liu , Chao Ji , Wei-Nan Zhang , Ting Liu

Human action recognition and analysis have great demand and important application significance in video surveillance, video retrieval, and human-computer interaction. The task of human action quality evaluation requires the intelligent…

计算机视觉与模式识别 · 计算机科学 2022-04-21 Shunli Wang , Dingkang Yang , Peng Zhai , Qing Yu , Tao Suo , Zhan Sun , Ka Li , Lihua Zhang

Action recognition models demonstrate strong generalization, but can they effectively transfer high-level motion concepts across diverse contexts, even within similar distributions? For example, can a model recognize the broad action…

计算机视觉与模式识别 · 计算机科学 2025-08-04 Raiyaan Abdullah , Jared Claypoole , Michael Cogswell , Ajay Divakaran , Yogesh Rawat

We present a method for assessing skill from video, applicable to a variety of tasks, ranging from surgery to drawing and rolling pizza dough. We formulate the problem as pairwise (who's better?) and overall (who's best?) ranking of video…

计算机视觉与模式识别 · 计算机科学 2018-03-30 Hazel Doughty , Dima Damen , Walterio Mayol-Cuevas

Unsupervised image-to-image translation is a recently proposed task of translating an image to a different style or domain given only unpaired image examples at training time. In this paper, we formulate a new task of unsupervised…

计算机视觉与模式识别 · 计算机科学 2018-06-12 Dina Bashkirova , Ben Usman , Kate Saenko

Transformer models have shown great success handling long-range interactions, making them a promising tool for modeling video. However, they lack inductive biases and scale quadratically with input length. These limitations are further…

计算机视觉与模式识别 · 计算机科学 2023-02-14 Javier Selva , Anders S. Johansen , Sergio Escalera , Kamal Nasrollahi , Thomas B. Moeslund , Albert Clapés

A more robust and holistic language-video representation is the key to pushing video understanding forward. Despite the improvement in training strategies, the quality of the language-video dataset is less attention to. The current plain…

多媒体 · 计算机科学 2024-06-21 Yuchen Yang , Yingxuan Duan

Mastering the technical skills required to perform surgery is an extremely challenging task. Video-based assessment allows surgeons to receive feedback on their technical skills to facilitate learning and development. Currently, this…

计算机视觉与模式识别 · 计算机科学 2022-07-07 Mona Fathollahi , Mohammad Hasan Sarhan , Ramon Pena , Lela DiMonte , Anshu Gupta , Aishani Ataliwala , Jocelyn Barker

We propose a novel multimodal video benchmark - the Perception Test - to evaluate the perception and reasoning skills of pre-trained multimodal models (e.g. Flamingo, SeViLA, or GPT-4). Compared to existing benchmarks that focus on…

This paper addresses the challenge of transferring the behavior expressivity style of a virtual agent to another one while preserving behaviors shape as they carry communicative meaning. Behavior expressivity style is viewed here as the…

多媒体 · 计算机科学 2023-08-22 Mireille Fares , Catherine Pelachaud , Nicolas Obin
‹ 上一页 1 2 3 10 下一页 ›