中文
相关论文

相关论文: MSGL-Transformer: A Multi-Scale Global-Local Trans…

200 篇论文

This paper introduces a modeling approach that employs multi-level global processing, encompassing both short-term frame-level and long-term sample-level feature scales. In the initial stage of shallow feature extraction, various scales are…

声音 · 计算机科学 2024-11-07 Chunyan Zeng , Yuhao Zhao , Zhifeng Wang

Grasp detection in a cluttered environment is still a great challenge for robots. Currently, the Transformer mechanism has been successfully applied to visual tasks, and its excellent ability of global context information extraction…

机器人学 · 计算机科学 2022-05-31 Mingshuai Dong , Xiuli Yu

This paper presents the first-rank solution for the Multi-Modal Action Recognition Challenge, part of the Multi-Modal Visual Pattern Recognition Workshop at the \acl{ICPR} 2024. The competition aimed to recognize human actions using a…

计算机视觉与模式识别 · 计算机科学 2025-02-11 Anh-Kiet Duong , Petra Gomez-Krämer

Automated social behaviour analysis of mice has become an increasingly popular research area in behavioural neuroscience. Recently, pose information (i.e., locations of keypoints or skeleton) has been used to interpret social behaviours of…

计算机视觉与模式识别 · 计算机科学 2025-01-09 Feixiang Zhou , Xinyu Yang , Fang Chen , Long Chen , Zheheng Jiang , Hui Zhu , Reiko Heckel , Haikuan Wang , Minrui Fei , Huiyu Zhou

Place recognition is an important capability for autonomously navigating vehicles operating in complex environments and under changing conditions. It is a key component for tasks such as loop closing in SLAM or global localization. In this…

机器人学 · 计算机科学 2023-04-21 Junyi Ma , Jun Zhang , Jintao Xu , Rui Ai , Weihao Gu , Xieyuanli Chen

Sign language is the preferred method of communication of deaf or mute people, but similar to any language, it is difficult to learn and represents a significant barrier for those who are hard of hearing or unable to speak. A person's…

计算机视觉与模式识别 · 计算机科学 2022-12-26 Neil Song , Yu Xiang

Leveraging sensing modalities across diverse spatial and temporal resolutions can improve performance of robotic manipulation tasks. Multi-spatial resolution sensing provides hierarchical information captured at different spatial scales and…

机器人学 · 计算机科学 2024-01-29 Saumya Saxena , Mohit Sharma , Oliver Kroemer

ChatGPT, a widely-recognized large language model (LLM), has recently gained substantial attention for its performance scaling, attributed to the billions of web-sourced natural language sentences used for training. Its underlying…

计算与语言 · 计算机科学 2023-11-21 Rui Fukushima , Jun Tani

Nonverbal behaviors, particularly gaze direction, play a crucial role in enhancing effective communication in social interactions. As social robots increasingly participate in these interactions, they must adapt their gaze based on human…

机器人学 · 计算机科学 2026-02-13 Faezeh Vahedi , Morteza Memari , Ramtin Tabatabaei , Alireza Taheri

Accurate prediction of road accidents remains challenging due to intertwined spatial, temporal, and contextual factors in urban traffic. We propose MSGAT-GRU, a multi-scale graph attention and recurrent model that jointly captures localized…

机器学习 · 计算机科学 2025-09-23 Thrinadh Pinjala , Aswin Ram Kumar Gannina , Debasis Dwibedy

Next-token prediction is the fundamental principle for training large language models (LLMs), and reinforcement learning (RL) further enhances their reasoning performance. As an effective way to model language, image, video, and other…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Zuyao Chen , Jinlin Wu , Zhen Lei , Marc Pollefeys , Chang Wen Chen

Reinforcement learning (RL) has been used in a range of simulated real-world tasks, e.g., sensor coordination, traffic light control, and on-demand mobility services. However, real world deployments are rare, as RL struggles with dynamic…

机器学习 · 计算机科学 2021-12-02 Alberto Castagna , Ivana Dusparic

Gait recognition is an important biometric for human identification at a distance, particularly under low-resolution or unconstrained environments. Current works typically focus on either 2D representations (e.g., silhouettes and skeletons)…

计算机视觉与模式识别 · 计算机科学 2026-04-23 Zhao-Yang Wang , Zhimin Shao , Anirudh Nanduri , Basudha Pal , Laura McDaniel , Jieneng Chen , Rama Chellappa

Gait emotion recognition plays a crucial role in the intelligent system. Most of the existing methods recognize emotions by focusing on local actions over time. However, they ignore that the effective distances of different emotions in the…

计算机视觉与模式识别 · 计算机科学 2022-09-20 Yunfei Yin , Li Jing , Faliang Huang , Guangchao Yang , Zhuowei Wang

Nowadays, Transformers and Graph Convolutional Networks (GCNs) are the prevailing techniques for 3D human pose estimation. However, Transformer-based methods either ignore the spatial neighborhood relationships between the joints when used…

计算机视觉与模式识别 · 计算机科学 2025-05-05 Kamel Aouaidjia , Aofan Li , Wenhao Zhang , Chongsheng Zhang

{Recognizing human interactions is essential for social robots as it enables them to navigate safely and naturally in shared environments. Conventional robotic systems however often focus on obstacle avoidance, neglecting social cues…

机器人学 · 计算机科学 2025-10-20 Thanh Long Nguyen , Duc Phu Nguyen , Thanh Thao Ton Nu , Quan Le , Thuan Hoang Tran , Manh Duong Phung

Click-through rate prediction plays an important role in the field of recommender system and many other applications. Existing methods mainly extract user interests from user historical behaviors. However, behavioral sequences only contain…

信息检索 · 计算机科学 2021-09-28 Yunfei Chu , Xiaofu Chang , Kunyang Jia , Jingzhen Zhou , Hongxia Yang

Vehicular platooning promises transformative improvements in transportation efficiency and safety through the coordination of multi-vehicle formations enabled by Vehicle-to-Everything (V2X) communication. However, the distributed nature of…

密码学与安全 · 计算机科学 2025-12-24 Konstantinos Kalogiannis , Ahmed Mohamed Hussain , Hexu Li , Panos Papadimitratos

Skeleton-based gait recognition models usually suffer from the robustness problem, as the Rank-1 accuracy varies from 90\% in normal walking cases to 70\% in walking with coats cases. In this work, we propose a state-of-the-art robust…

计算机视觉与模式识别 · 计算机科学 2022-04-11 Cun Zhang , Xing-Peng Chen , Guo-Qiang Han , Xiang-Jie Liu

Transformers have increasingly outperformed gated RNNs in obtaining new state-of-the-art results on supervised tasks involving text sequences. Inspired by this trend, we study the question of how Transformer-based models can improve the…

机器学习 · 计算机科学 2020-08-19 Ricky Loynd , Roland Fernandez , Asli Celikyilmaz , Adith Swaminathan , Matthew Hausknecht