English
Related papers

Related papers: Collaborative Transformers for Grounded Situation …

200 papers

Estimating human pose, classifying actions, and predicting movement progress are essential for human-robot interaction. While vision-based methods suffer from occlusion and privacy concerns in realistic environments, tactile sensing avoids…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Isaac Han , Seoyoung Lee , Sangyeon Park , Ecehan Akan , Yiyue Luo , Joseph DelPreto , Kyung-Joong Kim

We propose a new task, dataset and model for grounded video caption generation. This task unifies captioning and object grounding in video, where the objects in the caption are grounded in the video via temporally consistent bounding boxes.…

Computer Vision and Pattern Recognition · Computer Science 2024-11-13 Evangelos Kazakos , Cordelia Schmid , Josef Sivic

During recent years transformers architectures have been growing in popularity. Modulated Detection Transformer (MDETR) is an end-to-end multi-modal understanding model that performs tasks such as phase grounding, referring expression…

Computer Vision and Pattern Recognition · Computer Science 2022-09-22 Tomás Crisol , Joel Ermantraut , Adrián Rostagno , Santiago L. Aggio , Javier Iparraguirre

We introduce the GANformer2 model, an iterative object-oriented transformer, explored for the task of generative modeling. The network incorporates strong and explicit structural priors, to reflect the compositional nature of visual scenes,…

Computer Vision and Pattern Recognition · Computer Science 2021-11-18 Drew A. Hudson , C. Lawrence Zitnick

In this work, we propose a generally applicable transformation unit for visual recognition with deep convolutional neural networks. This transformation explicitly models channel relationships with explainable control variables. These…

Computer Vision and Pattern Recognition · Computer Science 2020-03-30 Zongxin Yang , Linchao Zhu , Yu Wu , Yi Yang

Predicting human mobility holds significant practical value, with applications ranging from enhancing disaster risk planning to simulating epidemic spread. In this paper, we present the GeoFormer, a decoder-only transformer model adapted…

Machine Learning · Computer Science 2023-11-10 Aivin V. Solatorio

Capturing the dependencies between joints is critical in skeleton-based action recognition task. Transformer shows great potential to model the correlation of important joints. However, the existing Transformer-based methods cannot capture…

Computer Vision and Pattern Recognition · Computer Science 2022-11-04 Helei Qiu , Biao Hou , Bo Ren , Xiaohua Zhang

This research presents the idea of activity fusion into existing Pose Estimation architectures to enhance their predictive ability. This is motivated by the rise in higher level concepts found in modern machine learning architectures, and…

Computer Vision and Pattern Recognition · Computer Science 2022-10-27 David Poulton , Richard Klein

Generic Event Boundary Detection (GEBD) aims to detect moments where humans naturally perceive as event boundaries. In this paper, we present Structured Context Transformer (or SC-Transformer) to solve the GEBD task, which can be trained in…

Computer Vision and Pattern Recognition · Computer Science 2022-06-08 Congcong Li , Xinyao Wang , Dexiang Hong , Yufei Wang , Libo Zhang , Tiejian Luo , Longyin Wen

Video-based behavior recognition is essential in fields such as public safety, intelligent surveillance, and human-computer interaction. Traditional 3D Convolutional Neural Network (3D CNN) effectively capture local spatiotemporal features…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Xiuliang Zhang , Tadiwa Elisha Nyamasvisva , Chuntao Liu

Understanding and predicting pedestrian crossing behavioral intention is crucial for the driving safety of autonomous vehicles. Nonetheless, challenges emerge when using promising images or environmental context masks to extract various…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Chen Xie , Ciyun Lin , Xiaoyu Zheng , Bowen Gong , Antonio M. López

The ability to perceive and reason about social interactions in the context of physical environments is core to human social intelligence and human-machine cooperation. However, no prior dataset or benchmark has systematically evaluated…

Artificial Intelligence · Computer Science 2021-03-23 Aviv Netanyahu , Tianmin Shu , Boris Katz , Andrei Barbu , Joshua B. Tenenbaum

Collaborative localization is an essential capability for a team of robots such as connected vehicles to collaboratively estimate object locations from multiple perspectives with reliant cooperation. To enable collaborative localization,…

Robotics · Computer Science 2021-11-09 Peng Gao , Brian Reily , Rui Guo , Hongsheng Lu , Qingzhao Zhu , Hao Zhang

Human pose forecasting is a challenging problem involving complex human body motion and posture dynamics. In cases that there are multiple people in the environment, one's motion may also be influenced by the motion and dynamic movements of…

Computer Vision and Pattern Recognition · Computer Science 2022-08-31 Edward Vendrow , Satyajit Kumar , Ehsan Adeli , Hamid Rezatofighi

Transformer, which originates from machine translation, is particularly powerful at modeling long-range dependencies. Currently, the transformer is making revolutionary progress in various vision tasks, leading to significant performance…

Computer Vision and Pattern Recognition · Computer Science 2023-01-02 Yuxin Mao , Jing Zhang , Zhexiong Wan , Yuchao Dai , Aixuan Li , Yunqiu Lv , Xinyu Tian , Deng-Ping Fan , Nick Barnes

Grasping in cluttered scenes has always been a great challenge for robots, due to the requirement of the ability to well understand the scene and object information. Previous works usually assume that the geometry information of the objects…

Robotics · Computer Science 2021-09-28 Yiming Li , Tao Kong , Ruihang Chu , Yifeng Li , Peng Wang , Lei Li

Transformers have proved to be very effective for visual recognition tasks. In particular, vision transformers construct compressed global representations through self-attention and learnable class tokens. Multi-resolution transformers have…

Computer Vision and Pattern Recognition · Computer Science 2022-12-16 Loic Themyr , Clement Rambour , Nicolas Thome , Toby Collins , Alexandre Hostettler

Existing Transformer-based RGBT tracking methods either use cross-attention to fuse the two modalities, or use self-attention and cross-attention to model both modality-specific and modality-sharing information. However, the significant…

Computer Vision and Pattern Recognition · Computer Science 2023-04-25 Yabin Zhu , Chenglong Li , Xiao Wang , Jin Tang , Zhixiang Huang

Autonomous vehicles operating in complex real-world environments require accurate predictions of interactive behaviors between traffic participants. This paper tackles the interaction prediction problem by formulating it with hierarchical…

Robotics · Computer Science 2023-08-15 Zhiyu Huang , Haochen Liu , Chen Lv

In this paper we describe two new computational operators, called complex entropic form (CEF) and generalized complex entropic form (GEF), for pattern characterization of spatially extended systems. Besides of being a measure of regularity,…

Condensed Matter · Physics 2009-10-31 Fernando M. Ramos , Reinaldo R. Rosa , Camilo Rodrigues Neto , Ademilson Zanandrea