中文
相关论文

相关论文: Collaborative Transformers for Grounded Situation …

200 篇论文

In this paper, we present the first transformer-based model to address the challenging problem of egocentric gaze estimation. We observe that the connection between the global scene context and local visual information is vital for…

计算机视觉与模式识别 · 计算机科学 2024-10-17 Bolin Lai , Miao Liu , Fiona Ryan , James M. Rehg

Centralized multimodal learning commonly compresses language, acoustic, and visual signals into a single fused representation for prediction. While effective, this paradigm suffers from two limitations: modality dominance, where…

We address the task of identifying distracted driving by analyzing in-car videos using efficient transformers. Although transformer models have achieved outstanding performance in human action recognition tasks, their high computational…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Ricardo Pizarro , Roberto Valle , Rafael Barea , Jose M. Buenaposada , Luis Baumela , Luis Miguel Bergasa

We propose a grounded dialogue state encoder which addresses a foundational issue on how to integrate visual grounding with dialogue system components. As a test-bed, we focus on the GuessWhat?! game, a two-player game where the goal is to…

Human Activity Recognition (HAR) with wearable sensors is challenged by limited interpretability, which significantly impacts cross-dataset generalization. To address this challenge, we propose Motion-Primitive Transformer (MoPFormer), a…

计算机视觉与模式识别 · 计算机科学 2025-10-29 Hao Zhang , Zhan Zhuang , Xuehao Wang , Xiaodong Yang , Yu Zhang

Multi-agent collaborative perception enhances each agent perceptual capabilities by sharing sensing information to cooperatively perform robot perception tasks. This approach has proven effective in addressing challenges such as sensor…

机器学习 · 计算机科学 2025-07-02 Rujia Wang , Xiangbo Gao , Hao Xiang , Runsheng Xu , Zhengzhong Tu

Simulation-based inference (SBI) with neural networks has accelerated and transformed cognitive modeling workflows. SBI enables modelers to fit complex models that were previously difficult or impossible to estimate, while also allowing…

机器学习 · 统计学 2026-03-24 Jerry M. Huang , Lukas Schumacher , Niek Stevenson , Stefan T. Radev

From a visual perception perspective, modern graphical user interfaces (GUIs) comprise a complex graphics-rich two-dimensional visuospatial arrangement of text, images, and interactive objects such as buttons and menus. While existing…

计算机视觉与模式识别 · 计算机科学 2024-04-23 Yue Jiang , Zixin Guo , Hamed Rezazadegan Tavakoli , Luis A. Leiva , Antti Oulasvirta

Despite rapid progress, embodied agents still struggle with long-horizon manipulation that requires maintaining spatial consistency, causal dependencies, and goal constraints. A key limitation of existing approaches is that task reasoning…

机器人学 · 计算机科学 2026-02-04 Kewei Hu , Michael Zhang , Wei Ying , Tianhao Liu , Guoqiang Hao , Zimeng Li , Wanchan Yu , Jiajian Jing , Fangwen Chen , Hanwen Kang

Group detection, especially for large-scale scenes, has many potential applications for public safety and smart cities. Existing methods fail to cope with frequent occlusions in large-scale scenes with multiple people, and are difficult to…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Jinsong Zhang , Lingfeng Gu , Yu-Kun Lai , Xueyang Wang , Kun Li

Student engagement is crucial for improving learning outcomes in group activities. Highly engaged students perform better both individually and contribute to overall group success. However, most existing automated engagement recognition…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Saniah Kayenat Chowdhury , Muhammad E. H. Chowdhury

In this paper, we propose a transformer based approach for visual grounding. Unlike previous proposal-and-rank frameworks that rely heavily on pretrained object detectors or proposal-free frameworks that upgrade an off-the-shelf one-stage…

计算机视觉与模式识别 · 计算机科学 2022-03-15 Ye Du , Zehua Fu , Qingjie Liu , Yunhong Wang

Convolution neural networks and Transformers have their own advantages and both have been widely used for dense prediction in multi-task learning (MTL). Existing studies typically employ either CNNs (effectively capture local spatial…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Yangyang Xu , Yibo Yang , Bernard Ghanem , Lefei Zhang , Bo Du , Jun Zhu

We present TRACE, a novel system for live *common ground* tracking in situated collaborative tasks. With a focus on fast, real-time performance, TRACE tracks the speech, actions, gestures, and visual attention of participants, uses these…

Spatio-temporal traffic forecasting is challenging due to complex temporal patterns, dynamic spatial structures, and diverse input formats. Although Transformer-based models offer strong global modeling, they often struggle with rigid…

人工智能 · 计算机科学 2025-08-20 Jiayu Fang , Zhiqi Shao , S T Boris Choy , Junbin Gao

In this study, we propose GITSR, an effective framework for Graph Interaction Transformer-based Scene Representation for multi-vehicle collaborative decision-making in intelligent transportation system. In the context of mixed traffic where…

机器学习 · 计算机科学 2024-11-05 Xingyu Hu , Lijun Zhang , Dejian Meng , Ye Han , Lisha Yuan

Remarkable advancements have been made recently in point cloud analysis through the exploration of transformer architecture, but it remains challenging to effectively learn local and global structures within point clouds. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2023-11-01 Haibo Qiu , Baosheng Yu , Dacheng Tao

Pedestrian trajectory prediction, vital for selfdriving cars and socially-aware robots, is complicated due to intricate interactions between pedestrians, their environment, and other Vulnerable Road Users. This paper presents GSGFormer, an…

计算机视觉与模式识别 · 计算机科学 2023-12-08 Zhongchang Luo , Marion Robin , Pavan Vasishta

Grounded Multimodal Named Entity Recognition (GMNER) identifies named entities, including their spans and types, in natural language text and grounds them to the corresponding regions in associated images. Most existing approaches split…

计算机视觉与模式识别 · 计算机科学 2026-03-20 Hongbing Li , Jiamin Liu , Shuo Zhang , Bo Xiao

The COVID-19 pandemic and the internet's availability have recently boosted online learning. However, monitoring engagement in online learning is a difficult task for teachers. In this context, timely automatic student engagement…

计算机视觉与模式识别 · 计算机科学 2025-02-18 Sandeep Mandia , Kuldeep Singh , Rajendra Mitharwal , Faisel Mushtaq , Dimpal Janu