English
Related papers

Related papers: VIBE: Video-Input Brain Encoder for fMRI Response …

200 papers

Dynamic Magnetic Resonance Imaging (MRI) of the vocal tract has become an increasingly adopted imaging modality for speech motor studies. Beyond image signals, systematic data loss, noise pollution, and audio file corruption can occur due…

Sound · Computer Science 2025-12-02 Yaxuan Li , Han Jiang , Yifei Ma , Shihua Qin , Jonghye Woo , Fangxu Xing

Online mapping and end-to-end (E2E) planning in autonomous driving remain largely sensor-centric, leaving rich map priors, including HD/SD vector maps, rasterized SD maps, and satellite imagery, underused because of heterogeneity, pose…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Zongzheng Zhang , Sizhe Zou , Guantian Zheng , Zhenxin Zhu , Yu Gao , Guoxuan Chi , Shuo Wang , Yuwen Heng , Zhigang Sun , Yiru Wang , Hao Sun , Chao Ma , Zhen Li , Anqing Jiang , Hao Zhao

Many motion-centric video analysis tasks, such as atomic actions, detecting atypical motor behavior in individuals with autism, or analyzing articulatory motion in real-time MRI of human speech, require efficient and interpretable temporal…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Hong Nguyen , Dung Tran , Hieu Hoang , Phong Nguyen , Shrikanth Narayanan

Sound speed heterogeneities can create aberrations in B-mode ultrasound images by inducing tissue-dependent delays and diffractive effects that conventional beamforming does not incorporate. By using the Fourier split-step method to…

Medical Physics · Physics 2026-05-01 Rehman Ali , Trevor M. Mitcham , Marvin M. Doyley , Nebojsa Duric , Jeremy J. Dahl

This paper introduces the Neural Transcoding Vision Transformer (\modelname), a generative model designed to estimate high-resolution functional Magnetic Resonance Imaging (fMRI) samples from simultaneous Electroencephalography (EEG) data.…

Image and Video Processing · Electrical Eng. & Systems 2024-09-19 Romeo Lanzino , Federico Fontana , Luigi Cinque , Francesco Scarcello , Atsuto Maki

Micro-gesture recognition and behavior-based emotion prediction are both highly challenging tasks that require modeling subtle, fine-grained human behaviors, primarily leveraging video and skeletal pose data. In this work, we present two…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Arman Martirosyan , Shahane Tigranyan , Maria Razzhivina , Artak Aslanyan , Nazgul Salikhova , Ilya Makarov , Andrey Savchenko , Aram Avetisyan

The Synesthetic Variational Autoencoder (SynVAE) introduced in this research is able to learn a consistent mapping between visual and auditive sensory modalities in the absence of paired datasets. A quantitative evaluation on MNIST as well…

Computer Vision and Pattern Recognition · Computer Science 2019-09-15 Maximilian Müller-Eberstein , Nanne van Noord

In this paper, we present our solutions for the 5th Workshop and Competition on Affective Behavior Analysis in-the-wild (ABAW), which includes four sub-challenges of Valence-Arousal (VA) Estimation, Expression (Expr) Classification, Action…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Ziyang Zhang , Liuwei An , Zishun Cui , Ao xu , Tengteng Dong , Yueqi Jiang , Jingyi Shi , Xin Liu , Xiao Sun , Meng Wang

Content providers increasingly replace traditional constant bitrate with variable bitrate (VBR) encoding in real-time video communication systems for better video quality. However, VBR encoding often leads to large and frequent bitrate…

Multimedia · Computer Science 2023-07-10 Zicheng Zhang , Hao Chen , Xun Cao , Zhan Ma

We examine "vibe coding": an emerging programming paradigm where developers primarily write code by interacting with code-generating large language models rather than writing code directly. We present the first empirical study of vibe…

Human-Computer Interaction · Computer Science 2025-10-06 Advait Sarkar , Ian Drosos

Continuous brain-computer interfaces (BCIs) that decode motion trajectories from imagined movement offer intuitive motor control, yet how feedback modality and longitudinal training shape neural representations and decoding performance…

Human-Computer Interaction · Computer Science 2026-05-29 Niall McShane , Attila Korik , Karl McCreadie , Naomi Du Bois , Darryl Charles , Damien Coyle

Despite participants engaging in unimodal stimuli, such as watching images or silent videos, recent work has demonstrated that multi-modal Transformer models can predict visual brain activity impressively well, even with incongruent…

Neurons and Cognition · Quantitative Biology 2025-05-27 Subba Reddy Oota , Khushbu Pahwa , Mounika Marreddy , Maneesh Singh , Manish Gupta , Bapi S. Raju

The connection between brain activity and corresponding visual stimuli is crucial in comprehending the human brain. While deep generative models have exhibited advancement in recovering brain recordings by generating images conditioned on…

Computer Vision and Pattern Recognition · Computer Science 2023-06-06 Xuelin Qian , Yikai Wang , Yanwei Fu , Xinwei Sun , Xiangyang Xue , Jianfeng Feng

The remarkable progress of Multimodal Large Language Models (MLLMs) has attracted increasing attention to extend them to physical entities like legged robot. This typically requires MLLMs to not only grasp multimodal understanding…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Gen Luo , Ganlin Yang , Ziyang Gong , Guanzhou Chen , Haonan Duan , Erfei Cui , Ronglei Tong , Zhi Hou , Tianyi Zhang , Zhe Chen , Shenglong Ye , Lewei Lu , Jingbo Wang , Wenhai Wang , Jifeng Dai , Yu Qiao , Rongrong Ji , Xizhou Zhu

We introduce VideoPrism, a general-purpose video encoder that tackles diverse video understanding tasks with a single frozen model. We pretrain VideoPrism on a heterogeneous corpus containing 36M high-quality video-caption pairs and 582M…

Recently, a generative variational autoencoder (VAE) has been proposed for speech enhancement to model speech statistics. However, this approach only uses clean speech in the training phase, making the estimation particularly sensitive to…

Audio and Speech Processing · Electrical Eng. & Systems 2021-05-18 Huajian Fang , Guillaume Carbajal , Stefan Wermter , Timo Gerkmann

In the context of long-term video understanding with large multimodal models, many frameworks have been proposed. Although transformer-based visual compressors and memory-augmented approaches are often used to process long videos, they…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Sosuke Yamao , Natsuki Miyahara , Yuankai Qi , Shun Takeuchi

Speech enhancement plays an essential role in various applications, and the integration of visual information has been demonstrated to bring substantial advantages. However, the majority of current research concentrates on the examination…

Sound · Computer Science 2025-04-03 Xinyuan Qian , Jiaran Gao , Yaodan Zhang , Qiquan Zhang , Hexin Liu , Leibny Paola Garcia , Haizhou Li

How to learn discriminative video representation from unlabeled videos is challenging but crucial for video analysis. The latest attempts seek to learn a representation model by predicting the appearance contents in the masked regions.…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Xinyu Sun , Peihao Chen , Liangwei Chen , Changhao Li , Thomas H. Li , Mingkui Tan , Chuang Gan

Physiological signals such as electrocardiograms (ECG) and electroencephalograms (EEG) provide complementary insights into human health and cognition, yet multi-modal integration is challenging due to limited multi-modal labeled data, and…