English
Related papers

Related papers: CoLoRSMamba: Conditional LoRA-Steered Mamba for Su…

200 papers

Long-range 3D object detection remains challenging because LiDAR observations become highly sparse and fragmented in the far field, making reliable context modeling difficult for existing detectors. To address this issue, recent state space…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Cheng Lu , Mingqian Ji , Shanshan Zhang , Zhihao Li , Jian Yang

Mamba, a State Space Model (SSM) that accelerates training by recasting recurrence as a parallel scan, has recently emerged as a linearly-scaling alternative to self-attention. Because of its unidirectional nature, each state in Mamba only…

Computer Vision and Pattern Recognition · Computer Science 2025-11-13 Jingwei Zhang , Xi Han , Hong Qin , Mahdi S. Hosseini , Dimitris Samaras

StyleMamba has recently demonstrated efficient text-driven image style transfer by leveraging state-space models (SSMs) and masked directional losses. In this paper, we extend the StyleMamba framework to handle video sequences. We propose…

Graphics · Computer Science 2025-07-31 Chao Li , Minsu Park , Cristina Rossi , Zhuang Li

Continuous Emotion Recognition (CER) plays a crucial role in intelligent human-computer interaction, mental health monitoring, and autonomous driving. Emotion modeling based on the Valence-Arousal (VA) space enables a more nuanced…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Yuheng Liang , Zheyu Wang , Feng Liu , Mingzhou Liu , Yu Yao

Multi-modal fusion holds great promise for integrating information from different modalities. However, due to a lack of consideration for modal consistency, existing multi-modal fusion methods in the field of remote sensing still face…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Mingxiang Cao , Weiying Xie , Xin Zhang , Jiaqing Zhang , Kai Jiang , Jie Lei , Yunsong Li

This work introduces a new task, text-conditioned selective video-to-audio (V2A) generation, which produces only the user-intended sound from a multi-object video. This capability is especially crucial in multimedia production, where audio…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Junwon Lee , Juhan Nam , Jiyoung Lee

This report provides an architecture-led analysis of two modern vision-language models (VLMs), Qwen2.5-VL-7B-Instruct and Llama-4-Scout-17B-16E-Instruct, and explains how their architectural properties map to a practical video-to-artifact…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Thomson Tong , Diba Darooneh

Recently, deception detection on human videos is an eye-catching techniques and can serve lots applications. AI model in this domain demonstrates the high accuracy, but AI tends to be a non-interpretable black box. We introduce an…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Shun-Wen Hsiao , Cheng-Yuan Sun

Video demoireing aims to remove undesirable interference patterns that arise during the capture of screen content, restoring artifact-free frames while maintaining temporal consistency. Existing video demoireing methods typically utilize…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Shuning Xu , Xina Liu , Binbin Song , Xiangyu Chen , Qiubo Chen , Jiantao Zhou

Low-Rank Adaptation (LoRA) has emerged as a widely adopted technique in text-to-image models, enabling precise rendering of multiple distinct elements, such as characters and styles, in multi-concept image generation. However, current…

Computer Vision and Pattern Recognition · Computer Science 2025-02-10 Xiandong Zou , Mingzhu Shen , Christos-Savvas Bouganis , Yiren Zhao

In this paper, we propose a Vision-Audio-Language Omni-peRception pretraining model (VALOR) for multi-modal understanding and generation. Different from widely-studied vision-language pretraining models, VALOR jointly models relationships…

Machine Learning · Computer Science 2025-01-07 Jing Liu , Sihan Chen , Xingjian He , Longteng Guo , Xinxin Zhu , Weining Wang , Jinhui Tang

Existing attacks against multimodal language models (MLLMs) primarily communicate instructions through text accompanied by adversarial images. In contrast, we exploit the capabilities of MLLMs to interpret non-textual instructions,…

Cryptography and Security · Computer Science 2025-06-03 Jiahui Geng , Thy Thy Tran , Preslav Nakov , Iryna Gurevych

Recently, a novel visual state space (VSS) model, referred to as Mamba, has demonstrated significant progress in modeling long sequences with linear complexity, comparable to Transformer models, thereby enhancing its adaptability for…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Tao Wang , Tiecheng Bai , Chao Xu , Bin Liu , Erlei Zhang , Jiyun Huang , Hongming Zhang

The depth/thermal information is beneficial for detecting salient object with conventional RGB images. However, in dual-modal salient object detection (SOD) model, the robustness against noisy inputs and modality missing is crucial but…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Shuang Hao , Chunlin Zhong , He Tang

Recent Mamba-based image restoration methods have achieved promising results but remain limited by fixed scanning patterns and inefficient feature utilization. Conventional Mamba architectures rely on predetermined paths that cannot adapt…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Han Hu , Zhuoran Zheng , Liang Li , Chen Lyu

Low-dose computed tomography (LDCT) reduces radiation exposure but often degrades image quality, potentially compromising diagnostic accuracy. Existing deep learning-based denoising methods focus primarily on pixel-level mappings,…

Image and Video Processing · Electrical Eng. & Systems 2025-07-09 Zhihao Chen , Tao Chen , Chenhui Wang , Qi Gao , Huidong Xie , Chuang Niu , Ge Wang , Hongming Shan

Temporal video grounding (TVG) is a critical task in video content understanding, requiring precise alignment between video content and natural language instructions. Despite significant advancements, existing methods face challenges in…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Wenrui Li , Xiaopeng Hong , Ruiqin Xiong , Xiaopeng Fan

Vision-Language Models (VLMs) have achieved remarkable success in various multi-modal tasks, but they are often bottlenecked by the limited context window and high computational cost of processing high-resolution image inputs and videos.…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Xubing Ye , Yukang Gan , Xiaoke Huang , Yixiao Ge , Yansong Tang

Vision-Language Models (VLMs) have shown strong performance in tasks like visual question answering and multimodal text generation, but their effectiveness in scientific domains such as materials science remains limited. While some machine…

Machine Learning · Computer Science 2025-11-11 An Vuong , Minh-Hao Van , Prateek Verma , Chen Zhao , Xintao Wu

Previous methods for predicting room acoustic parameters and speech quality metrics have focused on the single-channel case, where room acoustics and Mean Opinion Score (MOS) are predicted for a single recording device. However,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-14 Jozef Coldenhoff , Andrew Harper , Paul Kendrick , Tijana Stojkovic , Milos Cernak