中文
相关论文

相关论文: RS-WorldModel: a Unified Model for Remote Sensing …

200 篇论文

Understanding and forecasting the scene evolutions deeply affect the exploration and decision of embodied agents. While traditional methods simulate scene evolutions through trajectory prediction of potential instances, current works use…

计算机视觉与模式识别 · 计算机科学 2025-05-12 Zhang Zhang , Qiang Zhang , Wei Cui , Shuai Shi , Yijie Guo , Gang Han , Wen Zhao , Jingkai Sun , Jiahang Cao , Jiaxu Wang , Hao Cheng , Xiaozhu Ju , Zhengping Che , Renjing Xu , Jian Tang

Reinforcement learning has been demonstrated as a flexible and effective approach for learning a range of continuous control tasks, such as those used by robots to manipulate objects in their environment. But in robotics particularly,…

机器人学 · 计算机科学 2022-10-25 Tuluhan Akbulut , Max Merlin , Shane Parr , Benedict Quartey , Skye Thompson

Embodied AI requires agents that perceive, act, and anticipate how actions reshape future world states. World models serve as internal simulators that capture environment dynamics, enabling forward and counterfactual rollouts to support…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Xinqing Li , Xin He , Le Zhang , Min Wu , Xiaoli Li , Yun Liu

In the last decade, the rapid development of deep learning (DL) has made it possible to perform automatic, accurate, and robust Change Detection (CD) on large volumes of Remote Sensing Images (RSIs). However, despite advances in CD methods,…

计算机视觉与模式识别 · 计算机科学 2025-02-06 Lei Ding , Danfeng Hong , Maofan Zhao , Hongruixuan Chen , Chenyu Li , Jie Deng , Naoto Yokoya , Lorenzo Bruzzone , Jocelyn Chanussot

Recent advances in remote sensing have led to an increase in the number of available foundation models; each trained on different modalities, datasets, and objectives, yet capturing only part of the vast geospatial knowledge landscape.…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Joelle Hanna , Damian Falk , Stella X. Yu , Damian Borth

Deep learning has largely reshaped remote sensing (RS) research for aerial image understanding and made a great success. Nevertheless, most of the existing deep models are initialized with the ImageNet pretrained weights. Since natural…

计算机视觉与模式识别 · 计算机科学 2023-07-19 Di Wang , Jing Zhang , Bo Du , Gui-Song Xia , Dacheng Tao

Existing robot video world models are typically trained with low-level objectives such as reconstruction and perceptual similarity, which are poorly aligned with the capabilities that matter most for robot decision making, including…

Dynamic Scene Graph Generation (DSGG) models how object relations evolve over time in videos. However, existing methods are trained only on annotated object pairs and lack guidance for non-related pairs, making it difficult to identify…

计算机视觉与模式识别 · 计算机科学 2025-11-13 Hae-Won Jo , Yeong-Jun Cho

Remote sensing (RS) images contain numerous objects of different scales, which poses significant challenges for the RS image change captioning (RSICC) task to identify visual changes of interest in complex scenes and describe them via…

计算机视觉与模式识别 · 计算机科学 2023-12-19 Chenyang Liu , Jiajun Yang , Zipeng Qi , Zhengxia Zou , Zhenwei Shi

Recently, Referring Remote Sensing Image Segmentation (RRSIS) has aroused wide attention. To handle drastic scale variation of remote targets, existing methods only use the full image as input and nest the saliency-preferring techniques of…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Jiaxing Yang , Lihe Zhang , Huchuan Lu

Obtaining accurate photometric redshift estimations is an important aspect of cosmology, remaining a prerequisite of many analyses. In creating novel methods to produce redshift estimations, there has been a shift towards using machine…

天体物理仪器与方法 · 物理学 2021-07-07 Ben Henghes , Connor Pettitt , Jeyan Thiyagalingam , Tony Hey , Ofer Lahav

Pre-trained models have demonstrated exceptional generalization capabilities in time-series forecasting; however, adapting them to evolving data distributions remains a significant challenge. A key hurdle lies in accessing the original…

机器学习 · 计算机科学 2025-11-18 Tianyi Yin , Jingwei Wang , Chenze Wang , Han Wang , Jiexuan Cai , Min Liu , Yunlong Ma , Kun Gao , Yuting Song , Weiming Shen

World models predict state transitions in response to actions and are increasingly developed across diverse modalities. However, standard training objectives such as maximum likelihood estimation (MLE) often misalign with task-specific…

机器学习 · 计算机科学 2025-10-28 Jialong Wu , Shaofeng Yin , Ningya Feng , Mingsheng Long

The generation and enhancement of satellite imagery are critical in remote sensing, requiring high-quality, detailed images for accurate analysis. This research introduces a two-stage diffusion model methodology for synthesizing…

计算机视觉与模式识别 · 计算机科学 2024-10-08 Ahmad Sebaq , Mohamed ElHelw

Biological intelligence systems of animals perceive the world by integrating information in different modalities and processing simultaneously for various tasks. In contrast, current machine learning research follows a task-specific…

计算机视觉与模式识别 · 计算机科学 2021-12-03 Xizhou Zhu , Jinguo Zhu , Hao Li , Xiaoshi Wu , Xiaogang Wang , Hongsheng Li , Xiaohua Wang , Jifeng Dai

We introduce HY-World 2.0, a multi-modal world model framework that advances our prior project HY-World 1.0. HY-World 2.0 accommodates diverse input modalities, including text prompts, single-view images, multi-view images, and videos, and…

As remote sensing (RS) data obtained from different sensors become available largely and openly, multimodal data processing and analysis techniques have been garnering increasing interest in the RS and geoscience community. However, due to…

计算机视觉与模式识别 · 计算机科学 2021-05-24 Danfeng Hong , Jingliang Hu , Jing Yao , Jocelyn Chanussot , Xiao Xiang Zhu

In this paper, we present a robust and low complexity deep learning model for Remote Sensing Image Classification (RSIC), the task of identifying the scene of a remote sensing image. In particular, we firstly evaluate different low…

计算机视觉与模式识别 · 计算机科学 2022-12-13 Cam Le , Lam Pham , Nghia NVN , Truong Nguyen , Le Hong Trang

Foundation Models (FMs) are increasingly integrated into remote sensing (RS) pipelines. These models include unimodal vision encoders and multimodal architectures. FMs are adapted to diverse perception tasks, such as image classification,…

计算机视觉与模式识别 · 计算机科学 2026-03-12 Binger Chen , Tacettin Emre Bök , Behnood Rasti , Volker Markl , Begüm Demir

Visual imitation learning enables robotic agents to acquire skills by observing expert demonstration videos. In the one-shot setting, the agent generates a policy after observing a single expert demonstration without additional fine-tuning.…

机器人学 · 计算机科学 2026-01-01 Raktim Gautam Goswami , Prashanth Krishnamurthy , Yann LeCun , Farshad Khorrami