English
Related papers

Related papers: S1-MMAlign: A Large-Scale, Multi-Disciplinary Data…

200 papers

Recent advances in multimodal large language models (MLLMs) have accelerated progress in domain-oriented AI, yet their development in geoscience and remote sensing (RS) remains constrained by distinctive challenges: wide-ranging…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Aoran Xiao , Shihao Cheng , Yonghao Xu , Yexian Ren , Hongruixuan Chen , Naoto Yokoya

We present MM1.5, a new family of multimodal large language models (MLLMs) designed to enhance capabilities in text-rich image understanding, visual referring and grounding, and multi-image reasoning. Building upon the MM1 architecture,…

Fake news detection remains a challenging problem due to the complex interplay between textual misinformation, manipulated images, and external knowledge reasoning. While existing approaches have achieved notable results in verifying…

Computer Vision and Pattern Recognition · Computer Science 2025-10-20 Tuan-Vinh La , Minh-Hieu Nguyen , Minh-Son Dao

Faces and humans are crucial elements in social interaction and are widely included in everyday photos and videos. Therefore, a deep understanding of faces and humans will enable multi-modal assistants to achieve improved response quality…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Lixiong Qin , Shilong Ou , Miaoxuan Zhang , Jiangning Wei , Yuhang Zhang , Xiaoshuai Song , Yuchen Liu , Mei Wang , Weiran Xu

Leveraging Multi-modal Large Language Models (MLLMs) to accelerate frontier scientific research is promising, yet how to rigorously evaluate such systems remains unclear. Existing benchmarks mainly focus on single-document understanding,…

Artificial Intelligence · Computer Science 2026-04-14 Lei Xiong , Huaying Yuan , Zheng Liu , Zhao Cao , Zhicheng Dou

Nowadays, metadata information is often given by the authors themselves upon submission. However, a significant part of already existing research papers have missing or incomplete metadata information. German scientific papers come in a…

Information Retrieval · Computer Science 2021-11-11 Azeddine Bouabdallah , Jorge Gavilan , Jennifer Gerbl , Prayuth Patumcharoenpol

Massive web-crawled image-text datasets lay the foundation for recent progress in multimodal learning. These datasets are designed with the goal of training a model to do well on standard computer vision benchmarks, many of which, however,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Thao Nguyen , Matthew Wallingford , Sebastin Santy , Wei-Chiu Ma , Sewoong Oh , Ludwig Schmidt , Pang Wei Koh , Ranjay Krishna

Large Language Models (LLMs) are being explored for applications in scientific research, including their capabilities to synthesize literature, answer research questions, generate research ideas, and even conduct computational experiments.…

Assessing scientific claims requires identifying, extracting, and reasoning with multimodal data expressed in information-rich figures in scientific literature. Despite the large body of work in scientific QA, figure captioning, and other…

Computation and Language · Computer Science 2025-07-31 Yash Kumar Lal , Manikanta Bandham , Mohammad Saqib Hasan , Apoorva Kashi , Mahnaz Koupaee , Niranjan Balasubramanian

Reasoning over sequences of images remains a challenge for multimodal large language models (MLLMs). While recent models incorporate multi-image data during pre-training, they still struggle to recognize sequential structures, often…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Danae Sánchez Villegas , Ingo Ziegler , Desmond Elliott

The rapid evolution of Multimodal Large Language Models (MLLMs) has brought substantial advancements in artificial intelligence, significantly enhancing the capability to understand and generate multimodal content. While prior studies have…

Artificial Intelligence · Computer Science 2024-09-30 Lin Li , Guikun Chen , Hanrong Shi , Jun Xiao , Long Chen

Recent advances in multimodal large language models (MLLMs) have demonstrated strong capabilities in understanding general visual content. However, these general-domain MLLMs perform poorly in face perception tasks, often producing…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Jingzhi Li , Changjiang Luo , Ruoyu Chen , Hua Zhang , Wenqi Ren , Jianhou Gan , Xiaochun Cao

The lack of interpretability in the field of medical image analysis has significant ethical and legal implications. Existing interpretable methods in this domain encounter several challenges, including dependency on specific models,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-30 Lijie Hu , Songning Lai , Wenshuo Chen , Hongru Xiao , Hongbin Lin , Lu Yu , Jingfeng Zhang , Di Wang

Scientific documents contain complex multimodal structures, which makes evidence localization and scientific reasoning in Document Visual Question Answering particularly challenging. However, most existing benchmarks evaluate models only at…

Databases · Computer Science 2026-03-31 Wenhan Yu , Zhaoxi Zhang , Wang Chen , Guanqiang Qi , Weikang Li , Lei Sha , Deguo Xia , Jizhou Huang

Multimodal Large Language Models (MLLMs) excel in general domains but struggle with complex, real-world science. We posit that polymer science, an interdisciplinary field spanning chemistry, physics, biology, and engineering, is an ideal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Wanhao Liu , Weida Wang , Jiaqing Xie , Suorong Yang , Jue Wang , Benteng Chen , Guangtao Mei , Zonglin Yang , Shufei Zhang , Yuchun Mo , Lang Cheng , Jin Zeng , Houqiang Li , Wanli Ouyang , Yuqiang Li

Astronomical image interpretation presents a significant challenge for applying multimodal large language models (MLLMs) to specialized scientific tasks. Existing benchmarks focus on general multimodal capabilities but fail to capture the…

Instrumentation and Methods for Astrophysics · Physics 2025-10-22 Jinghang Shi , Xiaoyu Tang , Yang Huang , Yuyang Li , Xiao Kong , Yanxia Zhang , Caizhan Yue

Figures of speech such as metaphors, similes, and idioms are integral parts of human communication. They are ubiquitous in many forms of discourse, allowing people to convey complex, abstract ideas and evoke emotion. As figurative forms are…

Computation and Language · Computer Science 2023-11-28 Ron Yosef , Yonatan Bitton , Dafna Shahaf

The explosion of scientific literature has made the efficient and accurate extraction of structured data a critical component for advancing scientific knowledge and supporting evidence-based decision-making. However, existing tools often…

Human-Computer Interaction · Computer Science 2025-11-06 Xingbo Wang , Samantha L. Huey , Rui Sheng , Saurabh Mehta , Fei Wang

Recent advancements in Unified Multimodal Models (UMMs) have enabled remarkable image understanding and generation capabilities. However, while models like Gemini-2.5-Flash-Image show emerging abilities to reason over multiple related…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Mingrui Wu , Hang Liu , Jiayi Ji , Xiaoshuai Sun , Rongrong Ji

Recent progress in large language model (LLM) reasoning has focused on domains like mathematics and coding, where abundant high-quality data and objective evaluation metrics are readily available. In contrast, progress in LLM reasoning…

Artificial Intelligence · Computer Science 2026-01-12 Tengxiao Liu , Deepak Nathani , Zekun Li , Kevin Yang , William Yang Wang
‹ Prev 1 8 9 10 Next ›