中文
相关论文

相关论文: TerraMind: Large-Scale Generative Multimodality fo…

200 篇论文

Multimodal aerial data are used to monitor natural systems, and machine learning can significantly accelerate the classification of landscape features within such imagery to benefit ecology and conservation. It remains under-explored,…

计算机视觉与模式识别 · 计算机科学 2024-10-08 Lucia Gordon , Nico Lang , Catherine Ressijac , Andrew Davies

We present TartanGround, a large-scale, multi-modal dataset to advance the perception and autonomy of ground robots operating in diverse environments. This dataset, collected in various photorealistic simulation environments includes…

机器人学 · 计算机科学 2025-07-31 Manthan Patel , Fan Yang , Yuheng Qiu , Cesar Cadena , Sebastian Scherer , Marco Hutter , Wenshan Wang

We propose to build omni-modal intelligence, which is capable of understanding any modality and learning universal representations. In specific, we propose a scalable pretraining paradigm, named Multimodal Context (MiCo), which can scale up…

计算机视觉与模式识别 · 计算机科学 2024-06-14 Yiyuan Zhang , Handong Li , Jing Liu , Xiangyu Yue

Managing natural resources and mitigating risks from floods, droughts, wildfires, and landslides require models that can accurately predict climate-driven land-surface responses. Traditional models often struggle with spatial generalization…

机器学习 · 计算机科学 2026-02-03 Nicholas Kraabel , Jiangtao Liu , Yuchen Bian , Daniel Kifer , Chaopeng Shen

Multimodal magnetic resonance imaging (MRI) constitutes the first line of investigation for clinicians in the care of brain tumors, providing crucial insights for surgery planning, treatment monitoring, and biomarker identification.…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Lucas Robinet , Ahmad Berjaoui , Elizabeth Cohen-Jonathan Moyal

Vision-Language Models (VLMs) typically assume a uniform spatial fidelity across the entire field of view of visual inputs, dedicating equal precision to even the uninformative regions. By contrast, human vision is neither uniform nor…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Soumyaratna Debnath , Bui Duc Manh , Zinan Liu , Lin Wang

Holistic 3D modeling of molecularly defined brain structures is crucial for understanding complex brain functions. Using emerging tissue profiling technologies, researchers charted comprehensive atlases of mammalian brain with sub-cellular…

计算机视觉与模式识别 · 计算机科学 2025-12-15 Jiqing Wu , Ingrid Berg , Yawei Li , Ender Konukoglu , Viktor H. Koelzer

Multimodal sensing systems are increasingly prevalent in various real-world applications. Most existing multimodal learning approaches heavily rely on training with a large amount of synchronized, complete multimodal data. However, such a…

机器学习 · 计算机科学 2025-03-06 Xiaomin Ouyang , Jason Wu , Tomoyoshi Kimura , Yihan Lin , Gunjan Verma , Tarek Abdelzaher , Mani Srivastava

Despite the recent achievements made in the multi-modal emotion recognition task, two problems still exist and have not been well investigated: 1) the relationship between different emotion categories are not utilized, which leads to…

计算与语言 · 计算机科学 2020-10-08 Wenliang Dai , Zihan Liu , Tiezheng Yu , Pascale Fung

Advances in machine learning over the past decade have resulted in a proliferation of algorithmic applications for encoding, characterizing, and acting on complex data that may contain many high dimensional features. Recently, the emergence…

Multimodal emotion recognition (MMER) systems typically outperform unimodal systems by leveraging the inter- and intra-modal relationships between, e.g., visual, textual, physiological, and auditory modalities. This paper proposes an MMER…

计算机视觉与模式识别 · 计算机科学 2024-04-23 Paul Waligora , Haseeb Aslam , Osama Zeeshan , Soufiane Belharbi , Alessandro Lameiras Koerich , Marco Pedersoli , Simon Bacon , Eric Granger

Multimodal learning is a rapidly growing research field that has revolutionized multitasking and generative modeling in AI. While much of the research has focused on dealing with unstructured data (e.g., language, images, audio, or video),…

人工智能 · 计算机科学 2024-03-11 Marco D Alessandro , Enrique Calabrés , Mikel Elkano

Generative foundation models have advanced large-scale text-driven natural image generation, becoming a prominent research trend across various vertical domains. However, in the remote sensing field, there is still a lack of research on…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Chenyang Liu , Keyan Chen , Rui Zhao , Zhengxia Zou , Zhenwei Shi

Significant efforts have been directed towards adapting self-supervised multimodal learning for Earth observation applications. However, most current methods produce coarse patch-sized embeddings, limiting their effectiveness and…

Predicting multiple real-world tasks in a single model often requires a particularly diverse feature space. Multimodal (MM) models aim to extract the synergistic predictive potential of multiple data types to create a shared feature space…

Recently, emotion recognition based on physiological signals has emerged as a field with intensive research. The utilization of multi-modal, multi-channel physiological signals has significantly improved the performance of emotion…

多媒体 · 计算机科学 2023-08-22 Xinda Li

We present UniBind, a flexible and efficient approach that learns a unified representation space for seven diverse modalities -- images, text, audio, point cloud, thermal, video, and event data. Existing works, eg., ImageBind, treat the…

计算机视觉与模式识别 · 计算机科学 2024-03-20 Yuanhuiyi Lyu , Xu Zheng , Jiazhou Zhou , Lin Wang

This paper presents Omni-View, which extends the unified multimodal understanding and generation to 3D scenes based on multiview images, exploring the principle that "generation facilitates understanding". Consisting of understanding model,…

计算机视觉与模式识别 · 计算机科学 2026-02-02 JiaKui Hu , Shanshan Zhao , Qing-Guo Chen , Xuerui Qiu , Jialun Liu , Zhao Xu , Weihua Luo , Kaifu Zhang , Yanye Lu

Recent developments in spatial omics technologies have enabled the generation of high dimensional molecular data, such as transcriptomes, proteomes, and epigenomes, within their spatial tissue context, either through coprofiling on the same…

Global data assimilation enables weather forecasting at all scales and provides valuable data for studying the Earth system. However, the computational demands of physics-based algorithms used in operational systems limits the volume and…

机器学习 · 计算机科学 2024-07-17 Thomas J. Vandal , Kate Duffy , Daniel McDuff , Yoni Nachmany , Chris Hartshorn