English
Related papers

Related papers: Poly2Vec: Polymorphic Fourier-Based Encoding of Ge…

200 papers

Accurate property data for chemical elements is crucial for materials design and manufacturing, but many of them are difficult to measure directly due to equipment constraints. While traditional methods use the properties of other elements…

Computation and Language · Computer Science 2025-10-20 Yuanhao Li , Keyuan Lai , Tianqi Wang , Qihao Liu , Jiawei Ma , Yuan-Chao Hu

Multimodal Large Language Models have advanced AI in applications like text-to-video generation and visual question answering. These models rely on visual encoders to convert non-text data into vectors, but current encoders either lack…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Junjie Li , Jianghong Ma , Xiaofeng Zhang , Yuhang Li , Jianyang Shi

The development of computer vision algorithms for Unmanned Aerial Vehicle (UAV) applications in urban environments heavily relies on the availability of large-scale datasets with accurate annotations. However, collecting and annotating…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Francesco Barbato , Matteo Caligiuri , Pietro Zanuttigh

Cross-view object geo-localization has recently gained attention due to potential applications. Existing methods aim to capture spatial dependencies of query objects between different views through attention mechanisms to obtain spatial…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Xingtao Ling Yingying Zhu

We demonstrate the application of machine learning for rapid and accurate extraction of plasmonic particles cluster geometries from hyperspectral image data via a dual variational autoencoder (dual-VAE). In this approach, the information is…

Disordered Systems and Neural Networks · Physics 2022-08-09 Muammer Y. Yaman , Sergei V. Kalinin , Kathryn N. Guye , David Ginger , Maxim Ziatdinov

Maintaining an up-to-date map to reflect recent changes in the scene is very important, particularly in situations involving repeated traversals by a robot operating in an environment over an extended period. Undetected changes may cause a…

Multi-modal cross-view place recognition remains a fundamental challenge in computer vision and robotics due to the severe viewpoint, modality, and spatial-structure discrepancies between ground observations and aerial references. To…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Zhengyi Xu , Yuhang Ming , Zhihao Zhan , Hanyu Zhu , Javier Civera , Wanzeng Kong

Detecting diverse objects within complex indoor 3D point clouds presents significant challenges for robotic perception, particularly with varied object shapes, clutter, and the co-existence of static and dynamic elements where traditional…

Robotics · Computer Science 2025-07-24 Haichuan Li , Changda Tian , Panos Trahanias , Tomi Westerlund

Precise spatial reasoning is fundamental to robotic manipulation, yet the visual backbones of current vision-language-action (VLA) models are predominantly pretrained on 2D image data without explicit 3D geometric supervision, resulting in…

Unsupervised representation learning techniques, such as learning word embeddings, have had a significant impact on the field of natural language processing. Similar representation learning techniques have not yet become commonplace in the…

Computer Vision and Pattern Recognition · Computer Science 2021-02-09 Joël Bachmann , Kenneth Blomqvist , Julian Förster , Roland Siegwart

Different from Visual Question Answering task that requires to answer only one question about an image, Visual Dialogue involves multiple questions which cover a broad range of visual content that could be related to any objects,…

Computer Vision and Pattern Recognition · Computer Science 2019-11-19 Xiaoze Jiang , Jing Yu , Zengchang Qin , Yingying Zhuang , Xingxing Zhang , Yue Hu , Qi Wu

Many adaptations of transformers have emerged to address the single-modal vision tasks, where self-attention modules are stacked to handle input sources like images. Intuitively, feeding multiple modalities of data to vision transformers…

Computer Vision and Pattern Recognition · Computer Science 2022-07-18 Yikai Wang , Xinghao Chen , Lele Cao , Wenbing Huang , Fuchun Sun , Yunhe Wang

This paper introduces GeloVec, a new CNN-based attention smoothing framework for semantic segmentation that addresses critical limitations in conventional approaches. While existing attention-backed segmentation methods suffer from boundary…

Computer Vision and Pattern Recognition · Computer Science 2025-05-05 Boris Kriuk , Matey Yordanov

Recent progress in geospatial foundation models highlights the importance of learning general-purpose representations for real-world locations, particularly points-of-interest (POIs) where human activity concentrates. Existing approaches,…

Machine Learning · Computer Science 2026-03-06 Maria Despoina Siampou , Shushman Choudhury , Shang-Ling Hsu , Neha Arora , Cyrus Shahabi

The lack of quality labeled data is one of the main bottlenecks for training Deep Learning models. As the task increases in complexity, there is a higher penalty for overfitting and unstable learning. The typical paradigm employed today is…

Computer Vision and Pattern Recognition · Computer Science 2023-09-08 Priyam Mazumdar , Aiman Soliman , Volodymyr Kindratenko , Luigi Marini , Kenton McHenry

During the last decades, we have witnessed a surge of interests of learning a low-dimensional space with discriminative information from one single view. Even though most of them can achieve satisfactory performance in some certain…

Machine Learning · Computer Science 2019-05-21 Lin Feng , Xiangzhu Meng , Huibing Wang

While Transformers have demonstrated remarkable potential in modeling Partial Differential Equations (PDEs), modeling large-scale unstructured meshes with complex geometries remains a significant challenge. Existing efficient architectures…

Machine Learning · Computer Science 2026-05-01 Zhuo Zhang , Xi Yang , Ying Miao , Xiaobin Hu , Yifu Gao , Yuan Zhao , Yong Yang , Canqun Yang , Boocheong Khoo

The availability of city-scale Lidar maps enables the potential of city-scale place recognition using mobile cameras. However, the city-scale Lidar maps generally need to be compressed for storage efficiency, which increases the difficulty…

Computer Vision and Pattern Recognition · Computer Science 2024-02-27 Xudong Cai , Yongcai Wang , Zhe Huang , Yu Shao , Deying Li

Accurate 6D object pose estimation is an important task for a variety of robotic applications such as grasping or localization. It is a challenging task due to object symmetries, clutter and occlusion, but it becomes more challenging when…

Computer Vision and Pattern Recognition · Computer Science 2022-11-28 Thomas Jantos , Mohamed Amin Hamdad , Wolfgang Granig , Stephan Weiss , Jan Steinbrener

While 6D object pose estimation has recently made a huge leap forward, most methods can still only handle a single or a handful of different objects, which limits their applications. To circumvent this problem, category-level object pose…

Computer Vision and Pattern Recognition · Computer Science 2022-03-18 Yan Di , Ruida Zhang , Zhiqiang Lou , Fabian Manhardt , Xiangyang Ji , Nassir Navab , Federico Tombari