中文
相关论文

相关论文: InfoGeo: Information-Theoretic Object-Centric Lear…

200 篇论文

In this work, we present an information-theoretic framework that formulates cross-lingual language model pre-training as maximizing mutual information between multilingual-multi-granularity texts. The unified view helps us to better…

计算与语言 · 计算机科学 2021-04-08 Zewen Chi , Li Dong , Furu Wei , Nan Yang , Saksham Singhal , Wenhui Wang , Xia Song , Xian-Ling Mao , Heyan Huang , Ming Zhou

Visual Question Answering (VQA) presents a unique challenge by requiring models to understand and reason about visual content to answer questions accurately. Existing VQA models often struggle with biases introduced by the training data,…

计算机视觉与模式识别 · 计算机科学 2025-09-26 Zhifei Li , Feng Qiu , Yiran Wang , Yujing Xia , Kui Xiao , Miao Zhang , Yan Zhang

In this work, we aim at an important but less explored problem of a simple yet effective backbone specific for cross-view geo-localization task. Existing methods for cross-view geo-localization tasks are frequently characterized by 1)…

计算机视觉与模式识别 · 计算机科学 2023-02-06 Yingying Zhu , Hongji Yang , Yuxin Lu , Qiang Huang

Unmanned Aerial Vehicles (UAVs) play an increasingly critical role in Intelligence, Surveillance, and Reconnaissance (ISR) missions such as border patrolling and criminal detection, thanks to their ability to access remote areas and…

图像与视频处理 · 电气工程与系统科学 2024-10-16 Niloufar Mehrabi , Sayed Pedram Haeri Boroujeni , Jenna Hofseth , Abolfazl Razi , Long Cheng , Manveen Kaur , James Martin , Rahul Amin

The rapid growth of 3D digital content necessitates expandable recognition systems for open-world scenarios. However, existing 3D class-incremental learning methods struggle under extreme data scarcity due to geometric misalignment and…

计算机视觉与模式识别 · 计算机科学 2025-09-23 Tuo Xiang , Xuemiao Xu , Bangzhen Liu , Jinyi Li , Yong Li , Shengfeng He

Unmanned Aerial Vehicles (UAVs) represent a new frontier in a wide range of monitoring and research applications. To fully leverage their potential, a key challenge is planning missions for efficient data acquisition in complex…

机器人学 · 计算机科学 2020-01-10 Marija Popovic , Teresa Vidal-Calleja , Gregory Hitz , Jen Jen Chung , Inkyu Sa , Roland Siegwart , Juan Nieto

The rapid proliferation of unmanned aerial vehicles (UAVs) has highlighted the importance of robust and efficient object detection in diverse aerial scenarios. Detecting small objects under complex conditions, however, remains a significant…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Kunwei Lv , Zhiren Xiao , Hang Ren , Ping Lan

Supervised learning, while prevalent for information cascade modeling, often requires abundant labeled data in training, and the trained model is not easy to generalize across tasks and datasets. It often learns task-specific…

社会与信息网络 · 计算机科学 2022-02-22 Xovee Xu , Fan Zhou , Kunpeng Zhang , Siyuan Liu

Remote sensing visual grounding (RSVG) aims to localize objects in remote sensing imagery according to natural language expressions. Previous methods typically rely on sentence-level vision-language alignment, which struggles to exploit…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Ke Li , Ting Wang , Di Wang , Yongshan Zhu , Yiming Zhang , Tao Lei , Quan Wang

Real-time 3D object detection is crucial for autonomous cars. Achieving promising performance with high efficiency, voxel-based approaches have received considerable attention. However, previous methods model the input space with features…

计算机视觉与模式识别 · 计算机科学 2020-07-20 Jun Wang , Shiyi Lan , Mingfei Gao , Larry S. Davis

Incomplete multi-view clustering (IMVC) aims to cluster multi-view data that are only partially available. This poses two main challenges: effectively leveraging multi-view information and mitigating the impact of missing views. Prevailing…

机器学习 · 计算机科学 2024-07-15 Ge Teng , Ting Mao , Chen Shen , Xiang Tian , Xuesong Liu , Yaowu Chen , Jieping Ye

Cross-modal Thermal Geo-localization (TG) provides a robust, all-weather solution for Unmanned Aerial Vehicles (UAVs) in Global Navigation Satellite System (GNSS)-denied environments. However, profound thermal-visible modality gaps…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Xiaoran Zhang , Yu Liu , Jinyu Liang , Kangqiushi Li , Zhiwei Huang , Huaxin Xiao

This paper explores the use of applying a deep learning approach for 3D object detection to compute the relative position of an Unmanned Aerial Vehicle (UAV) from an Unmanned Ground Vehicle (UGV) equipped with a LiDAR sensor in a GPS-denied…

机器人学 · 计算机科学 2025-04-10 Uthman Olawoye , Jason N. Gross

Traditional search engines on World Wide Web (WWW) focus essentially on relevance ranking at the page level. But this lead to missing innumerable structured information about real-world objects embedded in static Web pages and online Web…

信息检索 · 计算机科学 2011-07-19 Dr. Pushpa R. Suri , Harmunish Taneja

Monitoring sustainable development goals requires accurate and timely socioeconomic statistics, while ubiquitous and frequently-updated urban imagery in web like satellite/street view images has emerged as an important source for…

计算机视觉与模式识别 · 计算机科学 2023-02-28 Yu Liu , Xin Zhang , Jingtao Ding , Yanxin Xi , Yong Li

The bird's-eye-view (BEV) representation allows robust learning of multiple tasks for autonomous driving including road layout estimation and 3D object detection. However, contemporary methods for unified road layout estimation and 3D…

计算机视觉与模式识别 · 计算机科学 2022-09-20 Curie Kim , Ue-Hwan Kim

LiDAR-based place recognition (LPR) is one of the most crucial components of autonomous vehicles to identify previously visited places in GPS-denied environments. Most existing LPR methods use mundane representations of the input point…

计算机视觉与模式识别 · 计算机科学 2023-10-09 Junyi Ma , Guangming Xiong , Jingyi Xu , Xieyuanli Chen

Although Multimodal Large Language Models (MLLMs) have advanced rapidly, they still face notable challenges in fine-grained multi-image understanding, often exhibiting spatial hallucination, attention leakage, and failures in object…

计算机视觉与模式识别 · 计算机科学 2026-04-27 Lihao Zheng , Zhenwei Shao , Yu Zhou , Yan Yang , Xintian Shen , Jiawei Chen , Hao Ma , Tao Wei

Vision-language alignment is crucial for various downstream tasks such as cross-modal generation and retrieval. Previous multimodal approaches like CLIP utilize InfoNCE to maximize mutual information, primarily aligning pairwise samples…

机器学习 · 计算机科学 2026-02-25 Wenzhe Yin , Zehao Xiao , Pan Zhou , Shujian Yu , Jiayi Shen , Jan-Jakob Sonke , Efstratios Gavves

Recently, with the prevalence of large-scale image dataset, the co-occurrence information among classes becomes rich, calling for a new way to exploit it to facilitate inference. In this paper, we propose Obj-GloVe, a generic scene-based…

计算机视觉与模式识别 · 计算机科学 2019-07-03 Canwen Xu , Zhenzhong Chen , Chenliang Li