中文
相关论文

相关论文: Aligning Step-by-Step Instructional Diagrams to Vi…

200 篇论文

Assistants on assembly tasks show great potential to benefit humans ranging from helping with everyday tasks to interacting in industrial settings. However, evaluation resources in assembly activities are underexplored. To foster system…

Robust frame-wise embeddings are essential to perform video analysis and understanding tasks. We present a self-supervised method for representation learning based on aligning temporal video sequences. Our framework uses a transformer-based…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Keyne Oei , Amr Gomaa , Anna Maria Feit , João Belo

The extraction of text information in videos serves as a critical step towards semantic understanding of videos. It usually involved in two steps: (1) text recognition and (2) text classification. To localize texts in videos, we can resort…

计算机视觉与模式识别 · 计算机科学 2022-06-07 Ye Liu , Changchong Lu , Chen Lin , Di Yin , Bo Ren

We propose to learn legged robot locomotion skills by watching thousands of wild animal videos from the internet, such as those featured in nature documentaries. Indeed, such videos offer a rich and diverse collection of plausible motion…

机器人学 · 计算机科学 2024-12-06 Elliot Chane-Sane , Constant Roux , Olivier Stasse , Nicolas Mansard

Multimodal large language models have fueled progress in image captioning. These models, fine-tuned on vast image datasets, exhibit a deep understanding of semantic concepts. In this work, we show that this ability can be re-purposed for…

音频与语音处理 · 电气工程与系统科学 2024-10-10 Hugo Malard , Michel Olvera , Stéphane Lathuiliere , Slim Essid

Agents that can learn to imitate given video observation -- \emph{without direct access to state or action information} are more applicable to learning in the natural world. However, formulating a reinforcement learning (RL) agent that…

机器学习 · 计算机科学 2023-07-14 Glen Berseth , Florian Golemo , Christopher Pal

A seamless integration of robots into human environments requires robots to learn how to use existing human tools. Current approaches for learning tool manipulation skills mostly rely on expert demonstrations provided in the target robot…

机器人学 · 计算机科学 2021-11-08 Kateryna Zorina , Justin Carpentier , Josef Sivic , Vladimír Petrík

Traditional multimodal learning approaches require expensive alignment pre-training to bridge vision and language modalities, typically projecting visual features into discrete text token spaces. We challenge both fundamental assumptions…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Xuhui Zhan , Tyler Derr

This paper presents a framework for learning visual representations from unlabeled video demonstrations captured from multiple viewpoints. We show that these representations are applicable for imitating several robotic tasks, including pick…

计算机视觉与模式识别 · 计算机科学 2023-01-30 André Correia , Luís A. Alexandre

Assembling objects from parts requires understanding multimodal instructions, linking them to 3D components, and predicting physically plausible 6-DoF motions for each assembly step. Existing datasets focus on simplified scenarios,…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Danrui Li , Jiahao Zhang , Bernhard Egger , Moitreya Chatterjee , Suhas Lohit , Tim K. Marks , Anoop Cherian

Imagine a robot is shown new concepts visually together with spoken tags, e.g. "milk", "eggs", "butter". After seeing one paired audio-visual example per class, it is shown a new set of unseen instances of these objects, and asked to pick…

计算与语言 · 计算机科学 2019-04-16 Ryan Eloff , Herman A. Engelbrecht , Herman Kamper

Scaling up robot learning is hindered by the scarcity of robotic demonstrations, whereas human videos offer a vast, untapped source of interaction data. However, bridging the embodiment gap between human hands and robot arms remains a…

机器人学 · 计算机科学 2026-04-14 Yifu Xu , Bokai Lin , Xinyu Zhan , Hongjie Fang , Yong-Lu Li , Cewu Lu , Lixin Yang

Visual instruction tuning has made considerable strides in enhancing the capabilities of Large Multimodal Models (LMMs). However, existing open LMMs largely focus on single-image tasks, their applications to multi-image scenarios remains…

计算机视觉与模式识别 · 计算机科学 2024-07-30 Feng Li , Renrui Zhang , Hao Zhang , Yuanhan Zhang , Bo Li , Wei Li , Zejun Ma , Chunyuan Li

We study multi-modal summarization for instructional videos, whose goal is to provide users an efficient way to learn skills in the form of text instructions and key video frames. We observe that existing benchmarks focus on generic…

计算机视觉与模式识别 · 计算机科学 2025-06-17 Yuan Zang , Hao Tan , Seunghyun Yoon , Franck Dernoncourt , Jiuxiang Gu , Kushal Kafle , Chen Sun , Trung Bui

We present IMU2CLIP, a novel pre-training approach to align Inertial Measurement Unit (IMU) motion sensor recordings with video and text, by projecting them into the joint representation space of Contrastive Language-Image Pre-training…

计算机视觉与模式识别 · 计算机科学 2022-10-27 Seungwhan Moon , Andrea Madotto , Zhaojiang Lin , Alireza Dirafzoon , Aparajita Saraf , Amy Bearman , Babak Damavandi

Description-based person re-identification (Re-id) is an important task in video surveillance that requires discriminative cross-modal representations to distinguish different people. It is difficult to directly measure the similarity…

计算机视觉与模式识别 · 计算机科学 2019-06-25 Kai Niu , Yan Huang , Wanli Ouyang , Liang Wang

A key solution to temporal sentence grounding (TSG) exists in how to learn effective alignment between vision and language features extracted from an untrimmed video and a sentence description. Existing methods mainly leverage vanilla soft…

计算机视觉与模式识别 · 计算机科学 2021-09-15 Daizong Liu , Xiaoye Qu , Pan Zhou

True understanding of videos comes from a joint analysis of all its modalities: the video frames, the audio track, and any accompanying text such as closed captions. We present a way to learn a compact multimodal feature representation that…

计算机视觉与模式识别 · 计算机科学 2020-04-07 Vivek Sharma , Makarand Tapaswi , Rainer Stiefelhagen

The rapid advancement of Large Multi-modal Foundation Models (LMM) has paved the way for the possible Explainable Image Quality Assessment (EIQA) with instruction tuning from two perspectives: overall quality explanation, and attribute-wise…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Yiting Lu , Xin Li , Haoning Wu , Bingchen Li , Weisi Lin , Zhibo Chen

Detecting and interpreting operator actions, engagement, and object interactions in dynamic industrial workflows remains a significant challenge in human-robot collaboration research, especially within complex, real-world environments.…

计算机视觉与模式识别 · 计算机科学 2025-01-13 Naval Kishore Mehta , Arvind , Himanshu Kumar , Abeer Banerjee , Sumeet Saurav , Sanjay Singh