English
Related papers

Related papers: Collaborative Transformers for Grounded Situation …

200 papers

Predicting human gaze is important in Human-Computer Interaction (HCI). However, to practically serve HCI applications, gaze prediction models must be scalable, fast, and accurate in their spatial and temporal gaze predictions. Recent…

Computer Vision and Pattern Recognition · Computer Science 2023-07-04 Sounak Mondal , Zhibo Yang , Seoyoung Ahn , Dimitris Samaras , Gregory Zelinsky , Minh Hoai

This paper introduces a novel approach to Social Group Activity Recognition (SoGAR) using Self-supervised Transformers network that can effectively utilize unlabeled video data. To extract spatio-temporal information, we created local and…

Computer Vision and Pattern Recognition · Computer Science 2024-11-20 Naga VS Raviteja Chappa , Pha Nguyen , Alexander H Nelson , Han-Seok Seo , Xin Li , Page Daniel Dobbs , Khoa Luu

We present a novel unified framework that concurrently tackles recognition and future prediction for human hand pose and action modeling. Previous works generally provide isolated solutions for either recognition or prediction, which not…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Yilin Wen , Hao Pan , Takehiko Ohkawa , Lei Yang , Jia Pan , Yoichi Sato , Taku Komura , Wenping Wang

Understanding complex animal behaviors hinges on deciphering the neural activity patterns within brain circuits, making the ability to forecast neural activity crucial for developing predictive models of brain dynamics. This capability…

Graph convolutional networks (GCNs) have been widely used and achieved remarkable results in skeleton-based action recognition. We think the key to skeleton-based action recognition is a skeleton hanging in frames, so we focus on how the…

Computer Vision and Pattern Recognition · Computer Science 2023-12-07 Nguyen Huu Bao Long

We introduce CARMA, a system for situational grounding in human-robot group interactions. Effective collaboration in such group settings requires situational awareness based on a consistent representation of present persons and objects…

Human activity intensity prediction is crucial to many location-based services. Despite tremendous progress in modeling dynamics of human activity, most existing methods overlook physical constraints of spatial interaction, leading to…

Machine Learning · Computer Science 2025-10-27 Yi Wang , Zhenghong Wang , Fan Zhang , Chaogui Kang , Sijie Ruan , Di Zhu , Chengling Tang , Zhongfu Ma , Weiyu Zhang , Yu Zheng , Philip S. Yu , Yu Liu

The recent development of multimodal single-cell technology has made the possibility of acquiring multiple omics data from individual cells, thereby enabling a deeper understanding of cellular states and dynamics. Nevertheless, the…

Genomics · Quantitative Biology 2023-10-16 Wenzhuo Tang , Hongzhi Wen , Renming Liu , Jiayuan Ding , Wei Jin , Yuying Xie , Hui Liu , Jiliang Tang

Graph Transformers have garnered significant attention for learning graph-structured data, thanks to their superb ability to capture long-range dependencies among nodes. However, the quadratic space and time complexity hinders the…

Information Retrieval · Computer Science 2024-05-08 Huiyuan Chen , Zhe Xu , Chin-Chia Michael Yeh , Vivian Lai , Yan Zheng , Minghua Xu , Hanghang Tong

We introduce a new dynamic model with the capability of recognizing both activities that an individual is performing as well as where that ndividual is located. Our model is novel in that it utilizes a dynamic graphical model to jointly…

Artificial Intelligence · Computer Science 2012-07-02 Amarnag Subramanya , Alvin Raj , Jeff A. Bilmes , Dieter Fox

Grounded Multimodal Named Entity Recognition (GMNER) task aims to identify named entities, entity types and their corresponding visual regions. GMNER task exhibits two challenging attributes: 1) The tenuous correlation between images and…

Multimedia · Computer Science 2025-09-03 Jinyuan Li , Ziyan Li , Han Li , Jianfei Yu , Rui Xia , Di Sun , Gang Pan

Egocentric action anticipation is a challenging task that aims to make advanced predictions of future actions from current and historical observations in the first-person view. Most existing methods focus on improving the model architecture…

Computer Vision and Pattern Recognition · Computer Science 2023-07-11 Congqi Cao , Ze Sun , Qinyi Lv , Lingtong Min , Yanning Zhang

For embodied agents to effectively understand and interact within the world around them, they require a nuanced comprehension of human actions grounded in physical space. Current action recognition models, often relying on RGB video, learn…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Nicholas Babey , Tiffany Gu , Yiheng Li , Cristian Meo , Kevin Zhu

We propose a novel attention-based 2D-to-3D pose estimation network for graph-structured data, named KOG-Transformer, and a 3D pose-to-shape estimation network for hand data, named GASE-Net. Previous 3D pose estimation methods have focused…

Computer Vision and Pattern Recognition · Computer Science 2022-09-27 Weixi Zhao , Weiqiang Wang

Assistance in collaborative manipulation is often initiated by user instructions, making high-level reasoning request-driven. In fluent human teamwork, however, partners often infer the next helpful step from the observed outcome of an…

Robotics · Computer Science 2026-03-26 Fengkai Liu , Hao Su , Haozhuang Chi , Rui Geng , Congzhi Ren , Xuqing Liu , Yucheng Xu , Yuichi Ohsita , Liyun Zhang

Time series forecasting at scale presents significant challenges for modern prediction systems, particularly when dealing with large sets of synchronized series, such as in a global payment network. In such systems, three key challenges…

Predicting the motion of multiple agents is necessary for planning in dynamic environments. This task is challenging for autonomous driving since agents (e.g. vehicles and pedestrians) and their associated behaviors may be diverse and…

Heatmap-based anatomical landmark detection is still facing two unresolved challenges: 1) inability to accurately evaluate the distribution of heatmap; 2) inability to effectively exploit global spatial structure information. To address the…

Computer Vision and Pattern Recognition · Computer Science 2023-05-22 Qikui Zhu , Yihui Bi , Danxin Wang , Xiangpeng Chu , Jie Chen , Yanqing Wang

In this technical report, we introduce our solution to human-centric spatio-temporal video grounding task. We propose a concise and effective framework named STVGFormer, which models spatiotemporal visual-linguistic dependencies with a…

Computer Vision and Pattern Recognition · Computer Science 2022-07-07 Zihang Lin , Chaolei Tan , Jian-Fang Hu , Zhi Jin , Tiancai Ye , Wei-Shi Zheng

Predicting pedestrian behavior is a crucial task for intelligent driving systems. Accurate predictions require a deep understanding of various contextual elements that potentially impact the way pedestrians behave. To address this…

Computer Vision and Pattern Recognition · Computer Science 2022-10-17 Amir Rasouli , Iuliia Kotseruba