English
Related papers

Related papers: Hierarchical Granularity Alignment and State Space…

200 papers

Recent vision-language pre-training models have exhibited remarkable generalization ability in zero-shot recognition tasks. Previous open-vocabulary 3D scene understanding methods mostly focus on training 3D models using either image or…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Ruihuang Li , Zhengqiang Zhang , Chenhang He , Zhiyuan Ma , Vishal M. Patel , Lei Zhang

Accurate Autism Spectrum Disorder (ASD) diagnosis is vital for early intervention. This study presents a hybrid deep learning framework combining Vision Transformers (ViT) and Vision Mamba to detect ASD using eye-tracking data. The model…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Wafaa Kasri , Yassine Himeur , Abigail Copiaco , Wathiq Mansoor , Ammar Albanna , Valsamma Eapen

Hyperspectral anomaly detection (HAD) aims to identify rare and irregular targets in high-dimensional hyperspectral images (HSIs), which are often noisy and unlabelled data. Existing deep learning methods either fail to capture long-range…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Aayushma Pant , Lakpa Tamang , Tsz-Kwan Lee , Sunil Aryal

Recently, a novel visual state space (VSS) model, referred to as Mamba, has demonstrated significant progress in modeling long sequences with linear complexity, comparable to Transformer models, thereby enhancing its adaptability for…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Tao Wang , Tiecheng Bai , Chao Xu , Bin Liu , Erlei Zhang , Jiyun Huang , Hongming Zhang

Understanding neural responses to visual stimuli remains challenging due to the inherent complexity of brain representations and the modality gap between neural data and visual inputs. Existing methods, mainly based on reducing neural…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Weihang You , Hanqi Jiang , Yi Pan , Junhao Chen , Tianming Liu , Fei Dou

Retrieval-augmented generation (RAG) systems have predominantly focused on text-based retrieval, limiting their effectiveness in handling visually-rich documents that encompass text, images, tables, and charts. To bridge this gap, we…

Information Retrieval · Computer Science 2025-05-07 Mingjun Xu , Zehui Wang , Hengxing Cai , Renxin Zhong

Most state-of-the-art approaches for Facial Action Unit (AU) detection rely upon evaluating facial expressions from static frames, encoding a snapshot of heightened facial activity. In real-world interactions, however, facial expressions…

Computer Vision and Pattern Recognition · Computer Science 2021-03-05 Nikhil Churamani , Sinan Kalkan , Hatice Gunes

Audio question answering (AQA), acting as a widely used proxy task to explore scene understanding, has got more attention. The AQA is challenging for it requires comprehensive temporal reasoning from different scales' events of an audio…

Sound · Computer Science 2023-05-30 Guangyao Li , Yixin Xu , Di Hu

As we exceed upon the procedures for modelling the different aspects of behaviour, expression recognition has become a key field of research in Human Computer Interactions. Expression recognition in the wild is a very interesting problem…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-03 Sowmya Rasipuram , Junaid Hamid Bhat , Anutosh Maitra

Facial Action Units (AUs) represent a set of facial muscular activities and various combinations of AUs can represent a wide range of emotions. AU recognition is often used in many applications, including marketing, healthcare, education,…

Computer Vision and Pattern Recognition · Computer Science 2021-07-09 Junya Saito , Xiaoyu Mi , Akiyoshi Uchida , Sachihiro Youoku , Takahisa Yamamoto , Kentaro Murase , Osafumi Nakayama

In this paper, we present our solution and experiment result for the Multi-Task Learning Challenge of the 7th Affective Behavior Analysis in-the-wild(ABAW7) Competition. This challenge consists of three tasks: action unit detection, facial…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Xiaodong Li , Wenchao Du , Hongyu Yang

Deep learning has revolutionized medical imaging by providing innovative solutions to complex healthcare challenges. Traditional models often struggle to dynamically adjust feature importance, resulting in suboptimal representation,…

Image and Video Processing · Electrical Eng. & Systems 2024-04-29 Kazi Shahriar Sanjid , Md. Tanzim Hossain , Md. Shakib Shahariar Junayed , M. Monir Uddin

We propose a Neural Hidden Markov Model (HMM) with Adaptive Granularity Attention (AGA) for high-frequency order flow modeling. The model addresses the challenge of capturing multi-scale temporal dynamics in financial markets, where…

Statistical Finance · Quantitative Finance 2026-03-24 Tianzuo Hu

Human affective behavior analysis plays a vital role in human-computer interaction (HCI) systems. In this paper, we introduce our submission to the CVPR 2023 Competition on Affective Behavior Analysis in-the-wild (ABAW). We propose a…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Jun Yu , Renda Li , Zhongpeng Cai , Gongpeng Zhao , Guochen Xie , Jichao Zhu , Wangyuan Zhu

With the rapid proliferation of information across digital platforms, stance detection has emerged as a pivotal challenge in social media analysis. While most of the existing approaches focus solely on textual data, real-world social media…

Computer Vision and Pattern Recognition · Computer Science 2025-09-11 Lata Pangtey , Omkar Kabde , Shahid Shafi Dar , Nagendra Kumar

Vision foundation models like DINOv2 demonstrate remarkable potential in medical imaging despite their origin in natural image domains. However, their design inherently works best for uni-modal image analysis, limiting their effectiveness…

Image and Video Processing · Electrical Eng. & Systems 2025-09-09 Daniel Scholz , Ayhan Can Erdur , Viktoria Ehm , Anke Meyer-Baese , Jan C. Peeken , Daniel Rueckert , Benedikt Wiestler

Point clouds and RGB images are two general perceptional sources in autonomous driving. The former can provide accurate localization of objects, and the latter is denser and richer in semantic information. Recently, AutoAlign presents a…

Computer Vision and Pattern Recognition · Computer Science 2022-07-22 Zehui Chen , Zhenyu Li , Shiquan Zhang , Liangji Fang , Qinhong Jiang , Feng Zhao

We present ASAP, a new framework for detecting and grounding multi-modal media manipulation (DGM4).Upon thorough examination, we observe that accurate fine-grained cross-modal semantic alignment between the image and text is vital for…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Zhenxing Zhang , Yaxiong Wang , Lechao Cheng , Zhun Zhong , Dan Guo , Meng Wang

Recent advances in multi-modal detection have significantly improved detection accuracy in challenging environments (e.g., low light, overexposure). By integrating RGB with modalities such as thermal and depth, multi-modal fusion increases…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Xiaofan Yang , Yubin Liu , Wei Pan , Guoqing Chu , Junming Zhang , Jie Zhao , Zhuoqi Man , Xuanming Cao

Handling lengthy context is crucial for enhancing the recognition and understanding capabilities of multimodal large language models (MLLMs) in applications such as processing high-resolution images or high frame rate videos. The rise in…

Computer Vision and Pattern Recognition · Computer Science 2024-11-14 Jianing Zhou , Han Li , Shuai Zhang , Ning Xie , Ruijie Wang , Xiaohan Nie , Sheng Liu , Lingyun Wang
‹ Prev 1 3 4 5 6 7 10 Next ›