中文
相关论文

相关论文: A Multimodal Seq2Seq Transformer for Predicting Br…

200 篇论文

Predicting the behaviors of other agents on the road is critical for autonomous driving to ensure safety and efficiency. However, the challenging part is how to represent the social interactions between agents and output different possible…

机器人学 · 计算机科学 2021-09-15 Zhiyu Huang , Xiaoyu Mo , Chen Lv

We present the early-stage design and implementation of a multimodal, real-time communication analysis system intended as a foundational interaction layer for adaptive VR training. The system integrates five parallel processing streams: (1)…

Understanding complex animal behaviors hinges on deciphering the neural activity patterns within brain circuits, making the ability to forecast neural activity crucial for developing predictive models of brain dynamics. This capability…

Multimodal dimensional emotion recognition has drawn a great attention from the affective computing community and numerous schemes have been extensively investigated, making a significant progress in this area. However, several questions…

计算机视觉与模式识别 · 计算机科学 2020-04-29 Dung Nguyen , Duc Thanh Nguyen , Rui Zeng , Thanh Thi Nguyen , Son N. Tran , Thin Nguyen , Sridha Sridharan , Clinton Fookes

Recent advances in omni-modal large language models have enabled remarkable progress in joint vision-audio understanding. However, prevailing architectures rely on modality-specific encoders with a \emph{video-coarse, audio-dense} design --…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Detao Bai , Shimin Yao , Weixuan Chen , Chengen Lai , Yuanming Li , Zhiheng Ma , Xihan Wei

Brain decoding, understood as the process of mapping brain activities to the stimuli that generated them, has been an active research area in the last years. In the case of language stimuli, recent studies have shown that it is possible to…

计算与语言 · 计算机科学 2020-11-12 Nicolas Affolter , Beni Egressy , Damian Pascual , Roger Wattenhofer

Recent research on representation learning has proved the merits of multi-modal clues for robust semantic segmentation. Nevertheless, a flexible pretrain-and-finetune pipeline for multiple visual modalities remains unexplored. In this…

计算机视觉与模式识别 · 计算机科学 2025-09-19 Bo-Wen Yin , Jiao-Long Cao , Xuying Zhang , Yuming Chen , Ming-Ming Cheng , Qibin Hou

Randomly masking and predicting word tokens has been a successful approach in pre-training language models for a variety of downstream tasks. In this work, we observe that the same idea also applies naturally to sequential decision making,…

Human emotions are difficult to convey through words and are often abstracted in the process; however, electroencephalogram (EEG) signals can offer a more direct lens into emotional brain activity. Recent studies show that deep learning…

神经元与认知 · 定量生物学 2025-11-19 Nilay Kumar , Priyansh Bhandari , G. Maragatham

Brain decoding is a key neuroscience field that reconstructs the visual stimuli from brain activity with fMRI, which helps illuminate how the brain represents the world. fMRI-to-image reconstruction has achieved impressive progress by…

神经元与认知 · 定量生物学 2025-10-27 Guoying Sun , Weiyu Guo , Tong Shao , Yang Yang , Haijin Zeng , Jie Liu , Jingyong Su

Despite impressive recent advances in text-to-image diffusion models, obtaining high-quality images often requires prompt engineering by humans who have developed expertise in using them. In this work, we present NeuroPrompts, an adaptive…

人工智能 · 计算机科学 2024-04-09 Shachar Rosenman , Vasudev Lal , Phillip Howard

Image captioning aims to automatically generate a natural language description of a given image, and most state-of-the-art models have adopted an encoder-decoder framework. The framework consists of a convolution neural network (CNN)-based…

计算机视觉与模式识别 · 计算机科学 2019-05-21 Jun Yu , Jing Li , Zhou Yu , Qingming Huang

As mobile robots increasingly operate in environments shared with humans, proactively anticipating human motion rather than responding reactively is critical for preempting collisions during close-proximity navigation, while maintaining…

人机交互 · 计算机科学 2025-11-24 Xiaoshan Zhou , Carol C. Menassa , Vineet R. Kamat

Predicting brain activity in response to naturalistic, multimodal stimuli is a key challenge in computational neuroscience. While encoding models are becoming more powerful, their ability to generalize to truly novel contexts remains a…

Decoding visual content from fMRI signals recorded while a person views images, and specifically answering questions about the seen images, is a long-standing challenge. While significant progress has been made in recent years in visual…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Roman Beliy , Matias Cosarinsky , Oliver Heinimann , Navve Wasserman , Michal Irani

Brain encoding and decoding aims to understand the relationship between external stimuli and brain activities, and is a fundamental problem in neuroscience. In this article, we study latent embedding alignment for brain encoding and…

统计方法学 · 统计学 2026-03-24 Shuoxun Xu , Zhanhao Yan , Lexin Li

We address prevailing challenges of the brain-powered research, departing from the observation that the literature hardly recover accurate spatial information and require subject-specific models. To address these challenges, we propose…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Weihao Xia , Raoul de Charette , Cengiz Öztireli , Jing-Hao Xue

fMRI semantic category understanding using linguistic encoding models attempt to learn a forward mapping that relates stimuli to the corresponding brain activation. Classical encoding models use linear multi-variate methods to predict the…

机器学习 · 计算机科学 2018-12-04 Subba Reddy Oota , Adithya Avvaru , Naresh Manwani , Raju S. Bapi

In this work we tackle the task of video-based audio-visual emotion recognition, within the premises of the 2nd Workshop and Competition on Affective Behavior Analysis in-the-wild (ABAW2). Poor illumination conditions, head/body orientation…

计算机视觉与模式识别 · 计算机科学 2022-11-04 Panagiotis Antoniadis , Ioannis Pikoulis , Panagiotis P. Filntisis , Petros Maragos

Reconstructing human dynamic vision from brain activity is a challenging task with great scientific significance. Although prior video reconstruction methods have made substantial progress, they still suffer from several limitations,…

计算机视觉与模式识别 · 计算机科学 2025-02-20 Yizhuo Lu , Changde Du , Chong Wang , Xuanliu Zhu , Liuyun Jiang , Xujin Li , Huiguang He