English
Related papers

Related papers: 3M-TRANSFORMER: A Multi-Stage Multi-Stream Multimo…

200 papers

Comprehending human motion is a fundamental challenge for developing Human-Robot Collaborative applications. Computer vision researchers have addressed this field by only focusing on reducing error in predictions, but not taking into…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Esteve Valls Mascaro , Shuo Ma , Hyemin Ahn , Dongheui Lee

Previous multi-task dense prediction studies developed complex pipelines such as multi-modal distillations in multiple stages or searching for task relational contexts for each task. The core insight beyond these methods is to maximize the…

Computer Vision and Pattern Recognition · Computer Science 2023-04-10 Yangyang Xu , Xiangtai Li , Haobo Yuan , Yibo Yang , Lefei Zhang

The Transformer and its variants have been proven to be efficient sequence learners in many different domains. Despite their staggering success, a critical issue has been the enormous number of parameters that must be trained (ranging from…

Machine Learning · Computer Science 2021-10-28 Subhabrata Dutta , Tanya Gautam , Soumen Chakrabarti , Tanmoy Chakraborty

Speech emotion recognition plays a crucial role in human-machine interaction systems. Recently various optimized Transformers have been successfully applied to speech emotion recognition. However, the existing Transformer architectures…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-02 Zhipeng Li , Xiaofen Xing , Yuanbo Fang , Weibin Zhang , Hengsheng Fan , Xiangmin Xu

Multimodal sentiment analysis in videos is a key task in many real-world applications, which usually requires integrating multimodal streams including visual, verbal and acoustic behaviors. To improve the robustness of multimodal fusion,…

Computer Vision and Pattern Recognition · Computer Science 2022-06-20 Lianyang Ma , Yu Yao , Tao Liang , Tongliang Liu

In the realm of computer vision and robotics, embodied agents are expected to explore their environment and carry out human instructions. This necessitates the ability to fully understand 3D scenes given their first-person observations and…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Tai Wang , Xiaohan Mao , Chenming Zhu , Runsen Xu , Ruiyuan Lyu , Peisen Li , Xiao Chen , Wenwei Zhang , Kai Chen , Tianfan Xue , Xihui Liu , Cewu Lu , Dahua Lin , Jiangmiao Pang

Human pose forecasting is a challenging problem involving complex human body motion and posture dynamics. In cases that there are multiple people in the environment, one's motion may also be influenced by the motion and dynamic movements of…

Computer Vision and Pattern Recognition · Computer Science 2022-08-31 Edward Vendrow , Satyajit Kumar , Ehsan Adeli , Hamid Rezatofighi

In the domain of multivariate forecasting, transformer models stand out as powerful apparatus, displaying exceptional capabilities in handling messy datasets from real-world contexts. However, the inherent complexity of these datasets,…

Machine Learning · Computer Science 2024-03-08 Jingjing Xu , Caesar Wu , Yuan-Fang Li , Pascal Bouvry

Pre-trained conversation models (PCMs) have achieved promising progress in recent years. However, existing PCMs for Task-oriented dialog (TOD) are insufficient for capturing the sequential nature of the TOD-related tasks, as well as for…

Computation and Language · Computer Science 2023-10-03 Lucen Zhong , Hengtong Lu , Caixia Yuan , Xiaojie Wang , Jiashen Sun , Ke Zeng , Guanglu Wan

Deep-learning-based clinical decision support using structured electronic health records (EHR) has been an active research area for predicting risks of mortality and diseases. Meanwhile, large amounts of narrative clinical notes provide…

Computation and Language · Computer Science 2023-05-10 Weimin Lyu , Xinyu Dong , Rachel Wong , Songzhu Zheng , Kayley Abell-Hart , Fusheng Wang , Chao Chen

Egocentric video generation with fine-grained control through body motion is a key requirement towards embodied AI agents that can simulate, predict, and plan actions. In this work, we propose EgoControl, a pose-controllable video diffusion…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Enrico Pallotta , Sina Mokhtarzadeh Azar , Lars Doorenbos , Serdar Ozsoy , Umar Iqbal , Juergen Gall

Turn-taking is a fundamental component of spoken dialogue, however conventional studies mostly involve dyadic settings. This work focuses on applying voice activity projection (VAP) to predict upcoming turn-taking in triadic multi-party…

Computation and Language · Computer Science 2025-10-06 Mikey Elmers , Koji Inoue , Divesh Lala , Tatsuya Kawahara

Transformer is a promising neural network learner, and has achieved great success in various machine learning tasks. Thanks to the recent prevalence of multimodal applications and big data, Transformer-based multimodal learning has become a…

Computer Vision and Pattern Recognition · Computer Science 2023-05-11 Peng Xu , Xiatian Zhu , David A. Clifton

Transformers have gained increasing popularity in a wide range of applications, including Natural Language Processing (NLP), Computer Vision and Speech Recognition, because of their powerful representational capacity. However, harnessing…

Egocentric action anticipation is the task of predicting the future actions a camera wearer will likely perform based on past video observations. While in a real-world system it is fundamental to output such predictions before the action…

Computer Vision and Pattern Recognition · Computer Science 2022-05-11 Antonino Furnari , Giovanni Maria Farinella

We introduce EgoSim, a closed-loop egocentric world simulator that generates spatially consistent interaction videos and persistently updates the underlying 3D scene state for continuous simulation. Existing egocentric simulators either…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Jinkun Hao , Mingda Jia , Ruiyan Wang , Xihui Liu , Ran Yi , Lizhuang Ma , Jiangmiao Pang , Xudong Xu

Scanpath prediction in 360{\deg} images can help realize rapid rendering and better user interaction in Virtual/Augmented Reality applications. However, existing scanpath prediction models for 360{\deg} images execute scanpath prediction on…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Rong Quan , Yantao Lai , Mengyu Qiu , Dong Liang

Video diffusion models have recently achieved remarkable progress in realism and controllability. However, achieving seamless video translation across different perspectives, such as first-person (egocentric) and third-person (exocentric),…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Quanjian Song , Yiren Song , Kelly Peng , Yuan Gao , Mike Zheng Shou

Deep neural networks are often applied to medical images to automate the problem of medical diagnosis. However, a more clinically relevant question that practitioners usually face is how to predict the future trajectory of a disease.…

Image and Video Processing · Electrical Eng. & Systems 2023-09-20 Huy Hoang Nguyen , Matthew B. Blaschko , Simo Saarakkala , Aleksei Tiulpin

Transformer encoder-decoder models have achieved great performance in dialogue generation tasks, however, their inability to process long dialogue history often leads to truncation of the context To address this problem, we propose a novel…

Computation and Language · Computer Science 2023-05-24 Qingyang Wu , Zhou Yu