中文
相关论文

相关论文: LUCAS-MEGA: A Large-Scale Multimodal Dataset for R…

200 篇论文

Scientific figure interpretation is a crucial capability for AI-driven scientific assistants built on advanced Large Vision Language Models. However, current datasets and benchmarks primarily focus on simple charts or other relatively…

Multi-modal learning is a fast growing area in artificial intelligence. It tries to help machines understand complex things by combining information from different sources, like images, text, and audio. By using the strengths of each…

We develop a deep learning based convolutional-regression model that estimates the volumetric soil moisture content in the top ~5 cm of soil. Input predictors include Sentinel-1 (active radar), Sentinel-2 (optical imagery), and SMAP…

大气与海洋物理 · 物理学 2023-10-17 Vishal Batchu , Grey Nearing , Varun Gulshan

Multimodal datasets contain an enormous amount of relational information, which grows exponentially with the introduction of new modalities. Learning representations in such a scenario is inherently complex due to the presence of multiple…

机器学习 · 计算机科学 2019-09-24 Devanshu Arya , Stevan Rudinac , Marcel Worring

Forecasting urban phenomena such as housing prices and public health indicators requires the effective integration of various geospatial data. Current methods primarily utilize task-specific models, while recent foundation models for…

机器学习 · 计算机科学 2025-10-16 Dominik J. Mühlematter , Lin Che , Ye Hong , Martin Raubal , Nina Wiedemann

Lecture slide presentations, a sequence of pages that contain text and figures accompanied by speech, are constructed and presented carefully in order to optimally transfer knowledge to students. Previous studies in multimedia and…

人工智能 · 计算机科学 2022-08-18 Dong Won Lee , Chaitanya Ahuja , Paul Pu Liang , Sanika Natu , Louis-Philippe Morency

Road surface classification (RSC) is a key enabler for environment-aware predictive maintenance systems. However, existing RSC techniques often fail to generalize beyond narrow operational conditions due to limited sensing modalities and…

Terrain modeling has traditionally relied on procedural techniques, which often require extensive domain expertise and handcrafted rules. In this paper, we present MESA - a novel data-centric alternative by training a diffusion model on…

图形学 · 计算机科学 2025-04-15 Paul Borne--Pons , Mikolaj Czerkawski , Rosalie Martin , Romain Rouffet

Human-machine interaction has been around for several decades now, with new applications emerging every day. One of the major goals that remain to be achieved is designing an interaction similar to how a human interacts with another human.…

人机交互 · 计算机科学 2022-12-27 Tauheed Khan Mohd , Nicole Nguyen , Ahmad Y Javaid

Large language models (LLMs) are increasingly grounded in sensor data to perceive and reason about human physiology and the physical world. However, accurately interpreting heterogeneous multimodal sensor data remains a fundamental…

人工智能 · 计算机科学 2026-01-13 Hyungjun Yoon , Mohammad Malekzadeh , Sung-Ju Lee , Fahim Kawsar , Lorena Qendro

Meeting the increasing global demand for food security and sustainable farming requires intelligent crop recommendation systems that operate in real time. Traditional soil analysis techniques are often slow, labor-intensive, and not…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Vishal Pandey , Ranjita Das , Debasmita Biswas

Recently, Vision-Language Models (VLMs) have achieved remarkable progress in multimodal tasks, and multimodal instruction data serves as the foundation for enhancing VLM capabilities. Despite the availability of several open-source…

Multi-modal multi-view action recognition is a rapidly growing field in computer vision, offering significant potential for applications in surveillance. However, current datasets often fail to address real-world challenges such as…

计算机视觉与模式识别 · 计算机科学 2025-05-08 Trung Thanh Nguyen , Yasutomo Kawanishi , Vijay John , Takahiro Komamizu , Ichiro Ide

Geometry problem-solving remains a significant challenge for Large Multimodal Models (LMMs), requiring not only global shape recognition but also attention to intricate local relationships related to geometric theory. To address this, we…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Linger Deng , Yuliang Liu , Wenwen Yu , Zujia Zhang , Jianzhong Ju , Zhenbo Luo , Xiang Bai

Plant classification is vital for ecological conservation and agricultural productivity, enhancing our understanding of plant growth dynamics and aiding species preservation. The advent of deep learning (DL) techniques has revolutionized…

计算机视觉与模式识别 · 计算机科学 2025-08-06 Alfreds Lapkovskis , Natalia Nefedova , Ali Beikmohammadi

Multimodal remote sensing image (MRSI) matching is pivotal for cross-modal fusion, localization, and object detection, but it faces severe challenges due to geometric, radiometric, and viewpoint discrepancies across imaging modalities.…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Peihao Wu , Yongxiang Yao , Wenfei Zhang , Dong Wei , Yi Wan , Yansheng Li , Yongjun Zhang

Learning holistic computational representations in physical, chemical or biological systems requires the ability to process information from different distributions and modalities within the same model. Thus, the demand for multimodal…

机器学习 · 计算机科学 2025-04-17 Konstantin Hemker , Nikola Simidjievski , Mateja Jamnik

The current availability of soil moisture data over large areas comes from satellite remote sensing technologies (i.e., radar-based systems), but these data have coarse resolution and often exhibit large spatial information gaps. Where data…

机器学习 · 计算机科学 2019-05-22 Danny Rorabaugh , Mario Guevara , Ricardo Llamas , Joy Kitson , Rodrigo Vargas , Michela Taufer

Multimodal learning, a rapidly evolving field in artificial intelligence, seeks to construct more versatile and robust systems by integrating and analyzing diverse types of data, including text, images, audio, and video. Inspired by the…

Meta-learning, or learning to learn, is a machine learning approach that utilizes prior learning experiences to expedite the learning process on unseen tasks. As a data-driven approach, meta-learning requires meta-features that represent…

机器学习 · 计算机科学 2021-01-12 Hadi S. Jomaa , Lars Schmidt-Thieme , Josif Grabocka