English
Related papers

Related papers: $\mathtt{M^3VIR}$: A Large-Scale Multi-Modality Mu…

200 papers

Text-to-video (T2V) diffusion models have recently achieved impressive visual quality, yet most systems still generate silent clips and treat audio as a secondary concern. Existing audio-video generation pipelines typically decompose the…

Long videos, ranging from minutes to hours, present significant challenges for current Multi-modal Large Language Models (MLLMs) due to their complex events, diverse scenes, and long-range dependencies. Direct encoding of such videos is…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Zizhong Li , Haopeng Zhang , Jiawei Zhang

Detecting a diverse range of objects under various driving scenarios is essential for the effectiveness of autonomous driving systems. However, the real-world data collected often lacks the necessary diversity presenting a long-tail…

Computer Vision and Pattern Recognition · Computer Science 2024-03-25 Aqeel Anwar , Tae Eun Choe , Zian Wang , Sanja Fidler , Minwoo Park

The rapid advancement in AI-generated video synthesis has led to a growth demand for standardized and effective evaluation metrics. Existing metrics lack a unified framework for systematically categorizing methodologies, limiting a holistic…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Xinhao Xiang , Xiao Liu , Zizhong Li , Zhuosheng Liu , Jiawei Zhang

Despite the commercial abundance of UAVs, aerial data acquisition remains challenging, and the existing Asia and North America-centric open-source UAV datasets are small-scale or low-resolution and lack diversity in scene contextuality.…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Aritra Dutta , Srijan Das , Jacob Nielsen , Rajatsubhra Chakraborty , Mubarak Shah

Visual designers naturally draw inspiration from multiple visual references, combining diverse elements and aesthetic principles to create artwork. However, current image generative frameworks predominantly rely on single-source inputs --…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Ruoxi Chen , Dongping Chen , Siyuan Wu , Sinan Wang , Shiyun Lang , Petr Sushko , Gaoyang Jiang , Yao Wan , Ranjay Krishna

We introduce MVSplat360, a feed-forward approach for 360{\deg} novel view synthesis (NVS) of diverse real-world scenes, using only sparse observations. This setting is inherently ill-posed due to minimal overlap among input views and…

Computer Vision and Pattern Recognition · Computer Science 2024-11-08 Yuedong Chen , Chuanxia Zheng , Haofei Xu , Bohan Zhuang , Andrea Vedaldi , Tat-Jen Cham , Jianfei Cai

We present a novel task for cross-dataset visual grounding in 3D scenes (Cross3DVG), which overcomes limitations of existing 3D visual grounding models, specifically their restricted 3D resources and consequent tendencies of overfitting a…

Computer Vision and Pattern Recognition · Computer Science 2024-02-08 Taiki Miyanishi , Daichi Azuma , Shuhei Kurita , Motoki Kawanabe

Understanding and analyzing video actions are essential for producing insightful and contextualized descriptions, especially for video-based applications like intelligent monitoring and autonomous systems. The proposed work introduces a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Lakshita Agarwal , Bindu Verma

Autoregressive (AR) models have recently shown strong performance in image generation, where a critical component is the visual tokenizer (VT) that maps continuous pixel inputs to discrete token sequences. The quality of the VT largely…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Huawei Lin , Tong Geng , Zhaozhuo Xu , Weijie Zhao

In recent years, artificial intelligence (AI)-driven video generation has gained significant attention. Consequently, there is a growing need for accurate video quality assessment (VQA) metrics to evaluate the perceptual quality of…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Zhichao Zhang , Wei Sun , Xinyue Li , Jun Jia , Xiongkuo Min , Zicheng Zhang , Chunyi Li , Zijian Chen , Puyi Wang , Fengyu Sun , Shangling Jui , Guangtao Zhai

Recent advances in Vision-Language Models (VLMs) have achieved impressive progress in multimodal mathematical reasoning. Yet, how much visual information truly contributes to reasoning remains unclear. Existing benchmarks report strong…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Yuandong Wang , Yao Cui , Yuxin Zhao , Zhen Yang , Yangfu Zhu , Zhenzhou Shao

This paper presents NTIRE 2026, the 3rd Restore Any Image Model (RAIM) challenge on multi-exposure image fusion in dynamic scenes. We introduce a benchmark that targets a practical yet difficult HDR imaging setting, where exposure…

The convergence of 3D geometric perception and video synthesis has created an unprecedented demand for large-scale video data that is rich in both semantic and spatio-temporal information. While existing datasets have advanced either 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Yunnan Wang , Kecheng Zheng , Jianyuan Wang , Minghao Chen , David Novotny , Christian Rupprecht , Yinghao Xu , Xing Zhu , Wenjun Zeng , Xin Jin , Yujun Shen

Composed Image Retrieval (CoIR) has recently gained popularity as a task that considers both text and image queries together, to search for relevant images in a database. Most CoIR approaches require manually annotated datasets, comprising…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Lucas Ventura , Antoine Yang , Cordelia Schmid , Gül Varol

Generating multi-view images from human instructions is crucial for 3D content creation. The primary challenges involve maintaining consistency across multiple views and effectively synthesizing shapes and textures under diverse conditions.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 JiaKui Hu , Yuxiao Yang , Jialun Liu , Jinbo Wu , Chen Zhao , Yanye Lu

Fully immersive experiences that tightly integrate 6-DoF visual and auditory interaction are essential for virtual and augmented reality. While such experiences can be achieved through computer-generated content, constructing them directly…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Zhengxian Yang , Shengqi Wang , Shi Pan , Hongshuai Li , Haoxiang Wang , Lin Li , Guanjun Li , Zhengqi Wen , Borong Lin , Jianhua Tao , Tao Yu

Compared with an extensive list of automotive radar datasets that support autonomous driving, indoor radar datasets are scarce at a smaller scale in the format of low-resolution radar point clouds and usually under an open-space single-room…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 M. Mahbubur Rahman , Ryoma Yataka , Sorachi Kato , Pu Perry Wang , Peizhao Li , Adriano Cardace , Petros Boufounos

There has been increasing interest in smart factories powered by robotics systems to tackle repetitive, laborious tasks. One impactful yet challenging task in robotics-powered smart factory applications is robotic grasping: using robotic…

Computer Vision and Pattern Recognition · Computer Science 2022-08-31 Yuhao Chen , E. Zhixuan Zeng , Maximilian Gilles , Alexander Wong

Recent advances in deep generative models have lead to remarkable progress in synthesizing high quality images. Following their successful application in image processing and representation learning, an important next step is to consider…

Computer Vision and Pattern Recognition · Computer Science 2019-03-28 Thomas Unterthiner , Sjoerd van Steenkiste , Karol Kurach , Raphael Marinier , Marcin Michalski , Sylvain Gelly