English
Related papers

Related papers: The Algonauts Project 2023 Challenge: UARK-UAlbany…

200 papers

Vision Language Navigation (VLN) requires agents to follow natural language instructions by grounding them in sequential visual observations over long horizons. Explicit reasoning could enhance temporal consistency and perception action…

In this report, we present our solution to the multi-task robustness track of the 1st Visual Continual Learning (VCL) Challenge at ICCV 2023 Workshop. We propose a vanilla framework named UniNet that seamlessly combines various visual…

Computer Vision and Pattern Recognition · Computer Science 2024-02-28 Zehui Chen , Qiuchen Wang , Zhenyu Li , Jiaming Liu , Shanghang Zhang , Feng Zhao

In this work we tackle the task of video-based audio-visual emotion recognition, within the premises of the 2nd Workshop and Competition on Affective Behavior Analysis in-the-wild (ABAW2). Poor illumination conditions, head/body orientation…

Computer Vision and Pattern Recognition · Computer Science 2022-11-04 Panagiotis Antoniadis , Ioannis Pikoulis , Panagiotis P. Filntisis , Petros Maragos

Building universal user representations that capture the essential aspects of user behavior is a crucial task for modern machine learning systems. In real-world applications, a user's historical interactions often serve as the foundation…

Information Retrieval · Computer Science 2025-08-12 Anton Klenitskiy , Artem Fatkulin , Daria Denisova , Anton Pembek , Alexey Vasilev

Many functional and structural neuroimaging studies call for accurate morphometric segmentation of different brain structures starting from image intensity values of MRI scans. Current automatic (multi-) atlas-based segmentation strategies…

Image and Video Processing · Electrical Eng. & Systems 2019-09-27 Dennis Bontempi , Sergio Benini , Alberto Signoroni , Michele Svanera , Lars Muckli

Background: A universal unanswered question in neuroscience and machine learning is whether computers can decode the patterns of the human brain. Multi-Voxels Pattern Analysis (MVPA) is a critical tool for addressing this question. However,…

Machine Learning · Statistics 2017-10-06 Muhammad Yousefnezhad , Daoqiang Zhang

We present a deep learning-based multi-task approach for head pose estimation in images. We contribute with a network architecture and training strategy that harness the strong dependencies among face pose, alignment and visibility, to…

Computer Vision and Pattern Recognition · Computer Science 2022-02-07 Roberto Valle , José Miguel Buenaposada , Luis Baumela

Image denoising can remove natural noise that widely exists in images captured by multimedia devices due to low-quality imaging sensors, unstable image transmission processes, or low light conditions. Recent works also find that image…

Computer Vision and Pattern Recognition · Computer Science 2021-08-25 Yupeng Cheng , Qing Guo , Felix Juefei-Xu , Wei Feng , Shang-Wei Lin , Weisi Lin , Yang Liu

Decoding visual stimuli from neural population activity is crucial for understanding the brain and for applications in brain-machine interfaces. However, such biological data is often scarce, particularly in primates or humans, where…

Machine Learning · Computer Science 2025-10-24 Jan Sobotka , Luca Baroni , Ján Antolík

A myriad of algorithms for the automatic analysis of brain MR images is available to support clinicians in their decision-making. For brain tumor patients, the image acquisition time series typically starts with an already pathological…

Image and Video Processing · Electrical Eng. & Systems 2024-09-24 Florian Kofler , Felix Meissen , Felix Steinbauer , Robert Graf , Stefan K Ehrlich , Annika Reinke , Eva Oswald , Diana Waldmannstetter , Florian Hoelzl , Izabela Horvath , Oezguen Turgut , Suprosanna Shit , Christina Bukas , Kaiyuan Yang , Johannes C. Paetzold , Ezequiel de da Rosa , Isra Mekki , Shankeeth Vinayahalingam , Hasan Kassem , Juexin Zhang , Ke Chen , Ying Weng , Alicia Durrer , Philippe C. Cattin , Julia Wolleb , M. S. Sadique , M. M. Rahman , W. Farzana , A. Temtam , K. M. Iftekharuddin , Maruf Adewole , Syed Muhammad Anwar , Ujjwal Baid , Anastasia Janas , Anahita Fathi Kazerooni , Dominic LaBella , Hongwei Bran Li , Ahmed W Moawad , Gian-Marco Conte , Keyvan Farahani , James Eddy , Micah Sheller , Sarthak Pati , Alexandros Karagyris , Alejandro Aristizabal , Timothy Bergquist , Verena Chung , Russell Takeshi Shinohara , Farouk Dako , Walter Wiggins , Zachary Reitman , Chunhao Wang , Xinyang Liu , Zhifan Jiang , Elaine Johanson , Zeke Meier , Ariana Familiar , Christos Davatzikos , John Freymann , Justin Kirby , Michel Bilello , Hassan M Fathallah-Shaykh , Roland Wiest , Jan Kirschke , Rivka R Colen , Aikaterini Kotrotsou , Pamela Lamontagne , Daniel Marcus , Mikhail Milchenko , Arash Nazeri , Marc-André Weber , Abhishek Mahajan , Suyash Mohan , John Mongan , Christopher Hess , Soonmee Cha , Javier Villanueva-Meyer , Errol Colak , Priscila Crivellaro , Andras Jakab , Abiodun Fatade , Olubukola Omidiji , Rachel Akinola Lagos , O O Olatunji , Goldey Khanna , John Kirkpatrick , Michelle Alonso-Basanta , Arif Rashid , Miriam Bornhorst , Ali Nabavizadeh , Natasha Lepore , Joshua Palmer , Antonio Porras , Jake Albrecht , Udunna Anazodo , Mariam Aboian , Evan Calabrese , Jeffrey David Rudie , Marius George Linguraru , Juan Eugenio Iglesias , Koen Van Leemput , Spyridon Bakas , Benedikt Wiestler , Ivan Ezhov , Marie Piraud , Bjoern H Menze

Recent advances in multimodal vision-language-action (VLA) models have revolutionized traditional robot learning, enabling systems to interpret vision, language, and action in unified frameworks for complex task planning. However, mastering…

Robotics · Computer Science 2025-06-12 Hongjun Wu , Heng Zhang , Pengsong Zhang , Jin Wang , Cong Wang

This paper presents a unified Vision-Language Pre-training (VLP) model. The model is unified in that (1) it can be fine-tuned for either vision-language generation (e.g., image captioning) or understanding (e.g., visual question answering)…

Computer Vision and Pattern Recognition · Computer Science 2019-12-05 Luowei Zhou , Hamid Palangi , Lei Zhang , Houdong Hu , Jason J. Corso , Jianfeng Gao

Electroencephalography (EEG) signals are known to manifest differential patterns when individuals visually concentrate on different objects. In this work, we present an end-to-end digital fabrication system, Brain2Object, to print the 3D…

Human-Computer Interaction · Computer Science 2020-06-18 Xiang Zhang , Lina Yao , Chaoran Huang , Salil S. Kanhere , Dalin Zhang , Yu Zhang

Over many decades, researchers working in object recognition have longed for an end-to-end automated system that will simply accept 2D or 3D image or videos as inputs and output the labels of objects in the input data. Computer vision…

Computer Vision and Pattern Recognition · Computer Science 2016-01-29 Rama Chellappa , Jun-Cheng Chen , Rajeev Ranjan , Swami Sankaranarayanan , Amit Kumar , Vishal M. Patel , Carlos D. Castillo

Global pandemic due to the spread of COVID-19 has post challenges in a new dimension on facial recognition, where people start to wear masks. Under such condition, the authors consider utilizing machine learning in image inpainting to…

Computer Vision and Pattern Recognition · Computer Science 2023-09-18 Zhengyang Han , Zehao Jiang , Yuan Ju

In this paper, the solution of HYU MLLAB KT Team to the Multimodal Algorithmic Reasoning Task: SMART-101 CVPR 2024 Challenge is presented. Beyond conventional visual question-answering problems, the SMART-101 challenge aims to achieve…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Jinwoo Ahn , Junhyeok Park , Min-Jun Kim , Kang-Hyeon Kim , So-Yeong Sohn , Yun-Ji Lee , Du-Seong Chang , Yu-Jung Heo , Eun-Sol Kim

In this paper, we introduce a novel visual representation learning which relies on a handful of adaptively learned tokens, and which is applicable to both image and video understanding tasks. Instead of relying on hand-designed splitting…

Computer Vision and Pattern Recognition · Computer Science 2022-04-05 Michael S. Ryoo , AJ Piergiovanni , Anurag Arnab , Mostafa Dehghani , Anelia Angelova

In this report, we describe the technical details of our submission to the EPIC-SOUNDS Audio-Based Interaction Recognition Challenge 2023, by Team "AcieLee" (username: Yuqi\_Li). The task is to classify the audio caused by interactions…

Sound · Computer Science 2023-06-16 Yuqi Li , Yizhi Luo , Xiaoshuai Hao , Chuanguang Yang , Zhulin An , Dantong Song , Wei Yi

At present, and increasingly so in the future, much of the captured visual content will not be seen by humans. Instead, it will be used for automated machine vision analytics and may require occasional human viewing. Examples of such…

Image and Video Processing · Electrical Eng. & Systems 2022-04-13 Hyomin Choi , Ivan V. Bajic

Understanding how biological visual systems process information is challenging due to the complex nonlinear relationship between neuronal responses and high-dimensional visual input. Artificial neural networks have already improved our…

‹ Prev 1 4 5 6 7 8 10 Next ›