English
Related papers

Related papers: Insights from the Algonauts 2025 Winners

200 papers

For a considerable time, deep convolutional neural networks (DCNNs) have reached human benchmark performance in object recognition. On that account, computational neuroscience and the field of machine learning have started to attribute…

Computer Vision and Pattern Recognition · Computer Science 2020-12-15 Leonard E. van Dyck , Walter R. Gruber

Neuroscience and Artificial Intelligence (AI) have made significant progress in the past few years but have only been loosely inter-connected. Based on a workshop held in August 2025, we identify current and future areas of synergism…

Artificial Intelligence · Computer Science 2026-01-29 Jean-Marc Fellous , Gert Cauwenberghs , Cornelia Fermüller , Yulia Sandamisrkaya , Terrence Sejnowski

While computer vision models have made incredible strides in static image recognition, they still do not match human performance in tasks that require the understanding of complex, dynamic motion. This is notably true for real-world…

Neurons and Cognition · Quantitative Biology 2025-04-09 Jacob Yeung , Andrew F. Luo , Gabriel Sarch , Margaret M. Henderson , Deva Ramanan , Michael J. Tarr

This paper provides a review of the NTIRE 2025 challenge on real-world face restoration, highlighting the proposed solutions and the resulting outcomes. The challenge focuses on generating natural, realistic outputs while maintaining…

In this paper, we introduce our submissions for the tasks of trimmed activity recognition (Kinetics) and trimmed event recognition (Moments in Time) for Activitynet Challenge 2018. In the two tasks, non-local neural networks and temporal…

Computer Vision and Pattern Recognition · Computer Science 2018-06-13 Xiaoteng Zhang , Yixin Bao , Feiyun Zhang , Kai Hu , Yicheng Wang , Liang Zhu , Qinzhu He , Yining Lin , Jie Shao , Yao Peng

Most research decoding brain signals into images, often using them as priors for generative models, has focused only on visual content. This overlooks the brain's natural ability to integrate auditory and visual information, for instance,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Jianxiong Gao , Yichang Liu , Baofeng Yang , Jianfeng Feng , Yanwei Fu

This paper proposes a novel deep learning framework for multi-modal motion prediction. The framework consists of three parts: recurrent neural networks to process the target agent's motion process, convolutional neural networks to process…

Robotics · Computer Science 2022-07-05 Zhiyu Huang , Xiaoyu Mo , Chen Lv

This paper presents a novel deep neural network (DNN) for multimodal fusion of audio, video and text modalities for emotion recognition. The proposed DNN architecture has independent and shared layers which aim to learn the representation…

Computer Vision and Pattern Recognition · Computer Science 2019-07-09 Juan D. S. Ortega , Mohammed Senoussaoui , Eric Granger , Marco Pedersoli , Patrick Cardinal , Alessandro L. Koerich

Document Image Machine Translation (DIMT) seeks to translate text embedded in document images from one language to another by jointly modeling both textual content and page layout, bridging optical character recognition (OCR) and natural…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Yaping Zhang , Yupu Liang , Zhiyang Zhang , Zhiyuan Chen , Lu Xiang , Yang Zhao , Yu Zhou , Chengqing Zong

The advance of speech decoding from non-invasive brain data holds the potential for profound societal impact. Among its most promising applications is the restoration of communication to paralysed individuals affected by speech deficits…

We present VIBE, a two-stage Transformer that fuses multi-modal video, audio, and text features to predict fMRI activity. Representations from open-source models (Qwen2.5, BEATs, Whisper, SlowFast, V-JEPA) are merged by a modality-fusion…

Machine Learning · Computer Science 2025-07-28 Daniel Carlström Schad , Shrey Dixit , Janis Keck , Viktor Studenyak , Aleksandr Shpilevoi , Andrej Bicanski

Humans and animals excel in combining information from multiple sensory modalities, controlling their complex bodies, adapting to growth, failures, or using tools. These capabilities are also highly desirable in robots. They are displayed…

Robotics · Computer Science 2022-11-08 Matej Hoffmann

Dealing with incomplete information is a well studied problem in the context of machine learning and computational intelligence. However, in the context of computer vision, the problem has only been studied in specific scenarios (e.g.,…

Computer Vision and Pattern Recognition · Computer Science 2021-06-25 Sergio Escalera , Marti Soler , Stephane Ayache , Umut Guclu , Jun Wan , Meysam Madadi , Xavier Baro , Hugo Jair Escalante , Isabelle Guyon

Deep Learning has driven recent and exciting progress in computer vision, instilling the belief that these algorithms could solve any visual task. Yet, datasets commonly used to train and test computer vision algorithms have pervasive…

Computer Vision and Pattern Recognition · Computer Science 2020-04-01 Vincent Jacquot , Zhuofan Ying , Gabriel Kreiman

This work examines the findings of the NTIRE 2025 Shadow Removal Challenge. A total of 306 participants have registered, with 17 teams successfully submitting their solutions during the final evaluation phase. Following the last two…

The Dialogic Robot Competition 2023 (DRC2023) is a competition for humanoid robots (android robots that closely resemble humans) to compete in interactive capabilities. This is the third year of the competition. The top four teams from the…

Robotics · Computer Science 2024-01-17 Ryuichiro Higashinaka , Takashi Minato , Hiromitsu Nishizaki , Takayuki Nagai

The neural underpinning of the biological visual system is challenging to study experimentally, in particular as the neuronal activity becomes increasingly nonlinear with respect to visual input. Artificial neural networks (ANNs) can serve…

Understanding and comprehending video content is crucial for many real-world applications such as search and recommendation systems. While recent progress of deep learning has boosted performance on various tasks using visual cues, deep…

Artificial Intelligence · Computer Science 2021-08-24 Hung-Ting Su , Po-Wei Shen , Bing-Chen Tsai , Wen-Feng Cheng , Ke-Jyun Wang , Winston H. Hsu

Autonomous driving is a multi-task problem requiring a deep understanding of the visual environment. End-to-end autonomous systems have attracted increasing interest as a method of learning to drive without exhaustively programming…

Computer Vision and Pattern Recognition · Computer Science 2019-09-12 Alexander Makrigiorgos , Ali Shafti , Alex Harston , Julien Gerard , A. Aldo Faisal

Motivation: Behavioral observations are an important resource in the study and evaluation of psychological phenomena, but it is costly, time-consuming, and susceptible to bias. Thus, we aim to automate coding of human behavior for use in…

Computer Vision and Pattern Recognition · Computer Science 2022-05-13 Nicole N. Lønfeldt , Flavia D. Frumosu , A. -R. Cecilie Mora-Jensen , Nicklas Leander Lund , Sneha Das , A. Katrine Pagsberg , Line K. H. Clemmensen