English
Related papers

Related papers: Beyond Appearance: Transformer-based Person Identi…

200 papers

Listeners use short interjections, so-called backchannels, to signify attention or express agreement. The automatic analysis of this behavior is of key importance for human conversation analysis and interactive conversational agents.…

Computer Vision and Pattern Recognition · Computer Science 2023-06-05 Ahmed Amer , Chirag Bhuvaneshwara , Gowtham K. Addluri , Mohammed M. Shaik , Vedant Bonde , Philipp Müller

In recent years, the transformer architecture has become the de facto standard for machine learning algorithms applied to natural language processing and computer vision. Despite notable evidence of successful deployment of this…

Robotics · Computer Science 2024-08-13 Carmelo Sferrazza , Dun-Ming Huang , Fangchen Liu , Jongmin Lee , Pieter Abbeel

The past year has witnessed the rapid development of applying the Transformer module to vision problems. While some researchers have demonstrated that Transformer-based models enjoy a favorable ability of fitting data, there are still…

Computer Vision and Pattern Recognition · Computer Science 2021-12-21 Zhengsu Chen , Lingxi Xie , Jianwei Niu , Xuefeng Liu , Longhui Wei , Qi Tian

Self-supervised pre-training of a speech foundation model, followed by supervised fine-tuning, has shown impressive quality improvements on automatic speech recognition (ASR) tasks. Fine-tuning separate foundation models for many downstream…

Machine Learning · Computer Science 2022-11-08 Zhouyuan Huo , Khe Chai Sim , Bo Li , Dongseong Hwang , Tara N. Sainath , Trevor Strohman

Comprehending human motion is a fundamental challenge for developing Human-Robot Collaborative applications. Computer vision researchers have addressed this field by only focusing on reducing error in predictions, but not taking into…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Esteve Valls Mascaro , Shuo Ma , Hyemin Ahn , Dongheui Lee

Sequential user modeling, a critical task in personalized recommender systems, focuses on predicting the next item a user would prefer, requiring a deep understanding of user behavior sequences. Despite the remarkable success of…

Artificial Intelligence · Computer Science 2023-10-10 Hao Wang , Jianxun Lian , Mingqi Wu , Haoxuan Li , Jiajun Fan , Wanyue Xu , Chaozhuo Li , Xing Xie

Recent advances in sensing technologies require the design and development of pattern recognition models capable of processing spatiotemporal data efficiently. In this study, we propose a spatially and temporally aware tensor-based neural…

Efficiently modeling spatio-temporal (ST) physical processes and observations presents a challenging problem for the deep learning community. Many recent studies have concentrated on meticulously reconciling various advantages, leading to…

Artificial Intelligence · Computer Science 2024-06-04 Hao Wu , Yuxuan Liang , Wei Xiong , Zhengyang Zhou , Wei Huang , Shilong Wang , Kun Wang

Modern AI models are increasingly being used as theoretical tools to study human cognition. One dominant approach is to evaluate whether human-derived measures are predicted by a model's output: that is, the end-product of a forward pass.…

Artificial Intelligence · Computer Science 2025-05-20 Jennifer Hu , Michael A. Lepori , Michael Franke

Most of the proposed person re-identification algorithms conduct supervised training and testing on single labeled datasets with small size, so directly deploying these trained models to a large-scale real-world camera network may lead to…

Computer Vision and Pattern Recognition · Computer Science 2018-03-21 Jianming Lv , Weihang Chen , Qing Li , Can Yang

Previous methods for dynamic facial expression in the wild are mainly based on Convolutional Neural Networks (CNNs), whose local operations ignore the long-range dependencies in videos. To solve this problem, we propose the spatio-temporal…

Computer Vision and Pattern Recognition · Computer Science 2022-05-11 Fuyan Ma , Bin Sun , Shutao Li

Anticipating the motion of all humans in dynamic environments such as homes and offices is critical to enable safe and effective robot navigation. Such spaces remain challenging as humans do not follow strict rules of motion and there are…

Robotics · Computer Science 2023-10-02 Tim Salzmann , Lewis Chiang , Markus Ryll , Dorsa Sadigh , Carolina Parada , Alex Bewley

Transformer based language models exhibit intelligent behaviors such as understanding natural language, recognizing patterns, acquiring knowledge, reasoning, planning, reflecting and using tools. This paper explores how their underlying…

Machine Learning · Computer Science 2023-11-15 Sumeet S. Singh

Although Vision Transformer (ViT) has achieved significant success in computer vision, it does not perform well in dense prediction tasks due to the lack of inner-patch information interaction and the limited diversity of feature scale.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-28 Chunlong Xia , Xinliang Wang , Feng Lv , Xin Hao , Yifeng Shi

Temporal expressions in text play a significant role in language understanding and correctly identifying them is fundamental to various retrieval and natural language processing systems. Previous works have slowly shifted from rule-based to…

Computation and Language · Computer Science 2022-01-25 Satya Almasian , Dennis Aumiller , Michael Gertz

Human action understanding is a fundamental and challenging task in computer vision. Although there exists tremendous research on this area, most works focus on action recognition, while action retrieval has received less attention. In this…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Hongsong Wang , Jianhua Zhao , Jie Gui

In person re-identification (re-ID) task, it is still challenging to learn discriminative representation by deep learning, due to limited data. Generally speaking, the model will get better performance when increasing the amount of data.…

Computer Vision and Pattern Recognition · Computer Science 2023-03-01 Wen Li , Cheng Zou , Meng Wang , Furong Xu , Jianan Zhao , Ruobing Zheng , Yuan Cheng , Wei Chu

Transformer-based approaches have been successfully proposed for 3D human pose estimation (HPE) from 2D pose sequence and achieved state-of-the-art (SOTA) performance. However, current SOTAs have difficulties in modeling spatial-temporal…

Computer Vision and Pattern Recognition · Computer Science 2023-01-19 Xiaoye Qian , Youbao Tang , Ning Zhang , Mei Han , Jing Xiao , Ming-Chun Huang , Ruei-Sung Lin

Encoder transformer models compress information from all tokens in a sequence into a single [CLS] token to represent global context. This approach risks diluting fine-grained or hierarchical features, leading to information loss in…

Computation and Language · Computer Science 2025-09-23 Asif Shahriar , Rifat Shahriyar , M Saifur Rahman

This paper focuses on the task of 4D shape reconstruction from a sequence of point clouds. Despite the recent success achieved by extending deep implicit representations into 4D space, it is still a great challenge in two respects, i.e. how…

Computer Vision and Pattern Recognition · Computer Science 2021-03-31 Jiapeng Tang , Dan Xu , Kui Jia , Lei Zhang