中文
相关论文

相关论文: GeoMag: Geometric-Aware Video Motion Magnification…

200 篇论文

Language-guided grasping has emerged as a promising paradigm for enabling robots to identify and manipulate target objects through natural language instructions, yet it remains highly challenging in cluttered or occluded scenes. Existing…

机器人学 · 计算机科学 2026-02-05 Rui Tang , Guankun Wang , Long Bai , Huxin Gao , Jiewen Lai , Chi Kit Ng , Jiazheng Wang , Fan Zhang , Hongliang Ren

This paper does not introduce a novel method but instead establishes a straightforward, incremental, yet essential baseline for video temporal grounding (VTG), a core capability in video understanding. While multimodal large language models…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Jun Zhang , Teng Wang , Yuying Ge , Yixiao Ge , Xinhao Li , Ying Shan , Limin Wang

Leveraging the universal representations of pre-trained LLMs and MLLMs offers a promising path toward brain foundation models. However, visually-evoked EEG datasets remain scarce, leading existing methods to align neural signals mainly with…

人工智能 · 计算机科学 2026-05-26 Jun-Yu Pan , Yansen Wang , Enze Zhang , Bao-Liang Lu , Wei-Long Zheng , Dongsheng Li

The task of video geolocalization aims to determine the precise GPS coordinates of a video's origin and map its trajectory; with applications in forensics, social media, and exploration. Existing classification-based approaches operate at a…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Parth Parag Kulkarni , Rohit Gupta , Prakash Chandra Chhipa , Mubarak Shah

Video depth estimation extends monocular prediction into the temporal domain to ensure coherence. However, existing methods often suffer from spatial blurring in fine-detail regions and temporal inconsistencies. We argue that current…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Yuecheng Liu , Junda Cheng , Longliang Liu , Wenjing Liao , Hanrui Cheng , Yuzhou Wang , Xin Yang

The online reconstruction of dynamic scenes from multi-view streaming videos faces significant challenges in training, rendering and storage efficiency. Harnessing superior learning speed and real-time rendering capabilities, 3D Gaussian…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Qiankun Gao , Jiarui Meng , Chengxiang Wen , Jie Chen , Jian Zhang

3D-aware image synthesis aims to generate images of objects from multiple views by learning a 3D representation. However, one key challenge remains: existing approaches lack geometry constraints, hence usually fail to generate multi-view…

计算机视觉与模式识别 · 计算机科学 2022-04-14 Xuanmeng Zhang , Zhedong Zheng , Daiheng Gao , Bang Zhang , Pan Pan , Yi Yang

Estimating the state of an environment from high-dimensional, multimodal, and noisy observations is a fundamental challenge in reinforcement learning (RL). Traditional approaches rely on probabilistic models to account for the uncertainty,…

机器学习 · 计算机科学 2026-02-13 Alfredo Reichlin , Adriano Pacciarelli , Danica Kragic , Miguel Vasco

In this paper, we propose to go beyond the well-established approach to vision-based localization that relies on visual descriptor matching between a query image and a 3D point cloud. While matching keypoints via visual descriptors makes…

计算机视觉与模式识别 · 计算机科学 2022-08-02 Qunjie Zhou , Sérgio Agostinho , Aljosa Osep , Laura Leal-Taixé

High-quality novel view synthesis (NVS) from real-world videos is crucial for applications such as cultural heritage preservation, digital twins, and immersive media. However, real-world videos typically contain long sequences with…

计算机视觉与模式识别 · 计算机科学 2026-02-11 Hojun Song , Heejung Choi , Aro Kim , Chae-yeong Song , Gahyeon Kim , Soo Ye Kim , Jaehyup Lee , Sang-hyo Park

Vision-language models (VLM) excel at general understanding yet remain weak at dynamic spatial reasoning (DSR), i.e., reasoning about the evolvement of object geometry and relationship in 3D space over time, largely due to the scarcity of…

计算机视觉与模式识别 · 计算机科学 2025-12-24 Shengchao Zhou , Yuxin Chen , Yuying Ge , Wei Huang , Jiehong Lin , Ying Shan , Xiaojuan Qi

Despite recent progress in calibration-free monocular SLAM via 3D vision foundation models, scale drift remains severe on long sequences. Motion-agnostic partitioning breaks contextual coherence and causes zero-motion drift, while…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Zhuang Xiong , Chen Zhang , Qingshan Xu , Wenbing Tao

Visual Geo-localization (VG) is the task of estimating the position where a given photo was taken by comparing it with a large database of images of known locations. To investigate how existing techniques would perform on a real-world…

计算机视觉与模式识别 · 计算机科学 2022-04-08 Gabriele Berton , Carlo Masone , Barbara Caputo

The rapid growth of 3D digital content necessitates expandable recognition systems for open-world scenarios. However, existing 3D class-incremental learning methods struggle under extreme data scarcity due to geometric misalignment and…

计算机视觉与模式识别 · 计算机科学 2025-09-23 Tuo Xiang , Xuemiao Xu , Bangzhen Liu , Jinyi Li , Yong Li , Shengfeng He

The development of video large multimodal models (LMMs) has been hindered by the difficulty of curating large amounts of high-quality raw data from the web. To address this, we propose an alternative approach by creating a high-quality…

计算机视觉与模式识别 · 计算机科学 2025-08-04 Yuanhan Zhang , Jinming Wu , Wei Li , Bo Li , Zejun Ma , Ziwei Liu , Chunyuan Li

Large language models (LLMs) have demonstrated strong reasoning capabilities in text-based mathematical problem solving; however, when adapted to visual reasoning tasks, particularly geometric problem solving, their performance…

人工智能 · 计算机科学 2025-10-28 Nannan Shi , Chuanyu Qin , Shipeng Song , Man Luo

Learning visuomotor policies from scarce expert demonstrations remains a core challenge in robotic manipulation. A primary hurdle lies in distilling high-dimensional RGB representations into control-relevant geometry without overfitting.…

机器人学 · 计算机科学 2026-05-18 Davide Buoso , Andrea Protopapa , Stefano Di Carlo , Francesca Pistilli , Giuseppe Averta

Recent neuro-symbolic geometry theorem provers have made significant progress on Euclidean problems by coupling neural guidance with symbolic verification. However, most existing systems operate almost exclusively in a symbolic space,…

人工智能 · 计算机科学 2026-02-24 Minfeng Zhu , Zi Wang , Sizhe Ji , Zhengtong Du , Shengqiang Tai , Junming Ke , Xiao Deng , Zanlang Yin , Xiuqi Huang , Heyu Wang , Wei Chen

Molecular modeling, a central topic in quantum mechanics, aims to accurately calculate the properties and simulate the behaviors of molecular systems. The molecular model is governed by physical laws, which impose geometric constraints such…

机器学习 · 计算机科学 2024-06-25 Tianlang Chen , Shengjie Luo , Di He , Shuxin Zheng , Tie-Yan Liu , Liwei Wang

Reconstructing dynamic 4D scenes from monocular videos is a fundamental yet challenging task. While recent 3D foundation models provide strong geometric priors, their performance significantly degrades in dynamic environments. This…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Ying Zang , Xuanyi Liu , Yidong Han , Deyi Ji , Chaotao Ding , Yuanqi Hu , Qi Zhu , Xuanfu Li , Jin Ma , Lingyun Sun , Tianrun Chen , Lanyun Zhu