中文
相关论文

相关论文: TorchSpatial: A Location Encoding Framework and Be…

200 篇论文

Vision-language models (VLMs) have demonstrated remarkable capabilities in understanding and reasoning about visual content, but significant challenges persist in tasks requiring cross-viewpoint understanding and spatial reasoning. We…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Dingming Li , Hongxing Li , Zixuan Wang , Yuchen Yan , Hang Zhang , Siqi Chen , Guiyang Hou , Shengpei Jiang , Wenqi Zhang , Yongliang Shen , Weiming Lu , Yueting Zhuang

The rapid growth of streaming media and e-commerce has driven advancements in recommendation systems, particularly Sequential Recommendation Systems (SRS). These systems employ users' interaction histories to predict future preferences.…

信息检索 · 计算机科学 2025-01-22 Alejo Lopez-Avila , Jinhua Du , Abbas Shimary , Ze Li

Multimodal large language models (MLLMs) have advanced from image-level reasoning to pixel-level grounding, but extending these capabilities to videos remains challenging as models must achieve spatial precision and temporally consistent…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Mohamad Alansari , Naufal Suryanto , Divya Velayudhan , Sajid Javed , Naoufel Werghi , Muzammal Naseer

Unsupervised text representation learning (TRL) is a fundamental task in natural language processing, which is beneficial for improving search and recommendations with the web's unlabeled texts. A recent empirical study finds that the…

计算与语言 · 计算机科学 2025-10-14 Ruize An , Richong Zhang , Zhijie Nie , Zhanyu Wu , Yanzhao Zhang , Dingkun Long

Spatial understanding is a crucial capability that enables robots to perceive their surroundings, reason about their environment, and interact with it meaningfully. In modern robotics, these capabilities are increasingly provided by…

计算机视觉与模式识别 · 计算机科学 2026-02-19 Chan Hee Song , Valts Blukis , Jonathan Tremblay , Stephen Tyree , Yu Su , Stan Birchfield

Generating learning-friendly representations for points in a 2D space is a fundamental and long-standing problem in machine learning. Recently, multi-scale encoding schemes (such as Space2Vec) were proposed to directly encode any point in…

计算机视觉与模式识别 · 计算机科学 2022-01-26 Gengchen Mai , Yao Xuan , Wenyun Zuo , Krzysztof Janowicz , Ni Lao

The band selection in the hyperspectral image (HSI) data processing is an important task considering its effect on the computational complexity and accuracy. In this work, we propose a novel framework for the band selection problem:…

计算机视觉与模式识别 · 计算机科学 2024-09-10 Mete Ahishali , Serkan Kiranyaz , Iftikhar Ahmad , Moncef Gabbouj

The function approximators employed by traditional image-based Deep Reinforcement Learning (DRL) algorithms usually lack a temporal learning component and instead focus on learning the spatial component. We propose a technique, Temporal…

机器学习 · 计算机科学 2021-10-28 Deepak George Thomas , Tichakorn Wongpiromsarn , Ali Jannesari

Robust prediction of citywide traffic flows at different time periods plays a crucial role in intelligent transportation systems. While previous work has made great efforts to model spatio-temporal correlations, existing methods still…

机器学习 · 计算机科学 2024-03-07 Jiahao Ji , Jingyuan Wang , Chao Huang , Junjie Wu , Boren Xu , Zhenhe Wu , Junbo Zhang , Yu Zheng

Generating learning-friendly representations for points in space is a fundamental and long-standing problem in ML. Recently, multi-scale encoding schemes (such as Space2Vec and NeRF) were proposed to directly encode any point in 2D/3D…

计算机视觉与模式识别 · 计算机科学 2023-07-04 Gengchen Mai , Yao Xuan , Wenyun Zuo , Yutong He , Jiaming Song , Stefano Ermon , Krzysztof Janowicz , Ni Lao

Spatial representation learning is essential for GeoAI applications such as urban analytics, enabling the encoding of shapes, locations, and spatial relationships (topological and distance-based) of geo-entities like points, polylines, and…

计算机视觉与模式识别 · 计算机科学 2025-08-28 Chen Chu , Cyrus Shahabi

With the development of earth observation technology, massive amounts of remote sensing (RS) images are acquired. To find useful information from these images, cross-modal RS image-voice retrieval provides a new insight. This paper aims to…

多媒体 · 计算机科学 2022-01-05 Hailong Ning , Bin Zhao , Yuan Yuan

Learning and recognition is a fundamental process performed in many robot operations such as mapping and localization. The majority of approaches share some common characteristics, such as attempting to extract salient features, landmarks…

机器人学 · 计算机科学 2017-07-21 Adam Jacobson , Walter Scheirer , Michael Milford

Super-resolution (SR) techniques aim to enhance data resolution, enabling the retrieval of finer details, and improving the overall quality and fidelity of the data representation. There is growing interest in applying SR methods to complex…

计算机视觉与模式识别 · 计算机科学 2025-05-12 Pu Ren , N. Benjamin Erichson , Junyi Guo , Shashank Subramanian , Omer San , Zarija Lukic , Michael W. Mahoney

We present a new framework for self-supervised representation learning by formulating it as a ranking problem in an image retrieval context on a large number of random views (augmentations) obtained from images. Our work is based on two…

计算机视觉与模式识别 · 计算机科学 2020-11-23 Ali Varamesh , Ali Diba , Tinne Tuytelaars , Luc Van Gool

As deep spatio-temporal neural networks are increasingly utilised in urban computing contexts, the deployment of such methods can have a direct impact on users of critical urban infrastructure, such as public transport, emergency services,…

机器学习 · 计算机科学 2025-08-12 Sichen Zhao , Wei Shao , Jeffrey Chan , Ziqi Xu , Flora Salim

Supervised learning for semantic segmentation requires a large number of labeled samples, which is difficult to obtain in the field of remote sensing. Self-supervised learning (SSL), can be used to solve such problems by pre-training a…

计算机视觉与模式识别 · 计算机科学 2022-02-01 Haifeng Li , Yi Li , Guo Zhang , Ruoyun Liu , Haozhe Huang , Qing Zhu , Chao Tao

To date, various 3D scene understanding tasks still lack practical and generalizable pre-trained models, primarily due to the intricate nature of 3D scene understanding tasks and their immense variations introduced by camera views,…

计算机视觉与模式识别 · 计算机科学 2021-09-02 Siyuan Huang , Yichen Xie , Song-Chun Zhu , Yixin Zhu

In this paper, we claim that spatial understanding is the keypoint in robot manipulation, and propose SpatialVLA to explore effective spatial representations for the robot foundation model. Specifically, we introduce Ego3D Position Encoding…

机器人学 · 计算机科学 2025-05-20 Delin Qu , Haoming Song , Qizhi Chen , Yuanqi Yao , Xinyi Ye , Yan Ding , Zhigang Wang , JiaYuan Gu , Bin Zhao , Dong Wang , Xuelong Li

We present a self-supervised learning (SSL) method suitable for semi-global tasks such as object detection and semantic segmentation. We enforce local consistency between self-learned features, representing corresponding image locations of…

计算机视觉与模式识别 · 计算机科学 2022-12-09 Ashraful Islam , Ben Lundell , Harpreet Sawhney , Sudipta Sinha , Peter Morales , Richard J. Radke