中文
相关论文

相关论文: Exploring the GLIDE model for Human Action-effect …

200 篇论文

Recent conditional image generation methods produce images of remarkable diversity, fidelity and realism. However, the majority of these methods allow conditioning only on labels or text prompts, which limits their level of control over the…

计算机视觉与模式识别 · 计算机科学 2023-02-14 Dina Bashkirova , Jose Lezama , Kihyuk Sohn , Kate Saenko , Irfan Essa

Automated vehicles operating in urban environments have to reliably interact with other traffic participants. Planning algorithms often utilize separate prediction modules forecasting probabilistic, multi-modal, and interactive behaviors of…

机器人学 · 计算机科学 2024-10-28 Sascha Rosbach , Stefan M. Leupold , Simon Großjohann , Stefan Roth

Understanding the decision process underlying gaze control is an important question in cognitive neuroscience with applications in diverse fields ranging from psychology to computer vision. The decision for choosing an upcoming saccade…

神经元与认知 · 定量生物学 2021-01-27 Noa Malem-Shinitski , Manfred Opper , Sebastian Reich , Lisa Schwetlick , Stefan A. Seelig , Ralf Engbert

For a given scene, humans can easily reason for the locations and pose to place objects. Designing a computational model to reason about these affordances poses a significant challenge, mirroring the intuitive reasoning abilities of humans.…

计算机视觉与模式识别 · 计算机科学 2024-07-23 Rishubh Parihar , Harsh Gupta , Sachidanand VS , R. Venkatesh Babu

Artificial Intelligence (AI) is transforming all fields of knowledge and production. From surgery, autonomous driving, to image and video creation, AI seems to make possible hitherto unimaginable processes of automation and efficient…

人工智能 · 计算机科学 2022-12-20 Frederic Guerrero-Sole

Image-generation diffusion models have been fine-tuned to unlock new capabilities such as image-editing and novel view synthesis. Can we similarly unlock image-generation models for visuomotor control? We present GENIMA, a behavior-cloning…

机器人学 · 计算机科学 2024-10-10 Mohit Shridhar , Yat Long Lo , Stephen James

The self-supervised Masked Image Modeling (MIM) schema, following "mask-and-reconstruct" pipeline of recovering contents from masked image, has recently captured the increasing interest in the multimedia community, owing to the excellent…

计算机视觉与模式识别 · 计算机科学 2022-04-19 Hao Liu , Xinghua Jiang , Xin Li , Antai Guo , Deqiang Jiang , Bo Ren

We address the challenging task of anticipating human-object interaction in first person videos. Most existing methods ignore how the camera wearer interacts with the objects, or simply consider body motion as a separate modality. In…

计算机视觉与模式识别 · 计算机科学 2020-07-21 Miao Liu , Siyu Tang , Yin Li , James Rehg

Scene graph generation (SGG) aims to automatically map an image into a semantic structural graph for better scene understanding. It has attracted significant attention for its ability to provide object and relation information, enabling…

计算机视觉与模式识别 · 计算机科学 2023-10-03 Xinyu Zhou , Zihan Ji , Anna Zhu

To interact with humans in collaborative environments, machines need to be able to predict (i.e., anticipate) future events, and execute actions in a timely manner. However, the observation of the human limb movements may not be sufficient…

机器人学 · 计算机科学 2020-06-19 Clebeson Canuto , Plinio Moreno , Jorge Samatelo , Raquel Vassallo , José Santos-Victor

Person image synthesis, e.g., pose transfer, is a challenging problem due to large variation and occlusion. Existing methods have difficulties predicting reasonable invisible regions and fail to decouple the shape and style of clothing,…

计算机视觉与模式识别 · 计算机科学 2021-04-01 Jinsong Zhang , Kun Li , Yu-Kun Lai , Jingyu Yang

We introduce the task of automatic human action co-occurrence identification, i.e., determine whether two human actions can co-occur in the same interval of time. We create and make publicly available the ACE (Action Co-occurrencE) dataset,…

计算机视觉与模式识别 · 计算机科学 2024-06-21 Oana Ignat , Santiago Castro , Weiji Li , Rada Mihalcea

This paper studies the problem of modeling multi-agent dynamical systems, where agents could interact mutually to influence their behaviors. Recent research predominantly uses geometric graphs to depict these mutual interactions, which are…

机器学习 · 计算机科学 2024-06-28 Xiao Luo , Yiyang Gu , Huiyu Jiang , Hang Zhou , Jinsheng Huang , Wei Ju , Zhiping Xiao , Ming Zhang , Yizhou Sun

Apparent personality analysis from short videos poses significant chal-lenges due to the complex interplay of visual, auditory, and textual cues. In this paper, we propose GAME, a Graph-Augmented Multimodal Encoder designed to robustly…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Kangsheng Wang , Yuhang Li , Chengwei Ye , Yufei Lin , Huanzhen Zhang , Bohan Hu , Linuo Xu , Shuyan Liu

Egocentric perception has grown rapidly with the advent of immersive computing devices. Human gaze prediction is an important problem in analyzing egocentric videos and has primarily been tackled through either saliency-based modeling or…

计算机视觉与模式识别 · 计算机科学 2021-04-30 Sathyanarayanan N. Aakur , Arunkumar Bagavathi

The attention mechanisms in deep neural networks are inspired by human's attention that sequentially focuses on the most relevant parts of the information over time to generate prediction output. The attention parameters in those models are…

计算机视觉与模式识别 · 计算机科学 2017-07-20 Youngjae Yu , Jongwook Choi , Yeonhwa Kim , Kyung Yoo , Sang-Hun Lee , Gunhee Kim

We propose Video Gaussian Masked Autoencoders (Video-GMAE), a self-supervised approach for representation learning that encodes a sequence of images into a set of Gaussian splats moving over time. Representing a video as a set of Gaussians…

计算机视觉与模式识别 · 计算机科学 2026-01-01 Tanish Baranwal , Himanshu Gaurav Singh , Jathushan Rajasegaran , Jitendra Malik

Generative models have made it possible to synthesize highly realistic images, potentially providing an abundant data source for training machine learning models. Despite the advantages of these synthesizable data sources, the…

计算机视觉与模式识别 · 计算机科学 2026-02-18 Shentong Mo , Sukmin Yun

Looking at a person's hands one often can tell what the person is going to do next, how his/her hands are moving and where they will be, because an actor's intentions shape his/her movement kinematics during action execution. Similarly,…

计算机视觉与模式识别 · 计算机科学 2016-10-05 Cornelia Fermüller , Fang Wang , Yezhou Yang , Konstantinos Zampogiannis , Yi Zhang , Francisco Barranco , Michael Pfeiffer

Image captioning is a fundamental task in vision-language understanding, where the model predicts a textual informative caption to a given input image. In this paper, we present a simple approach to address this task. We use CLIP encoding…

计算机视觉与模式识别 · 计算机科学 2021-11-19 Ron Mokady , Amir Hertz , Amit H. Bermano
‹ 上一页 1 8 9 10 下一页 ›