中文
相关论文

相关论文: MolRecBench-Wild: A Real-World Benchmark for Optic…

200 篇论文

We investigate 17 benchmarks (e.g. SugarCREPE, VALSE) commonly used for measuring compositional understanding capabilities of vision-language models (VLMs). We scrutinize design choices in their construction, including data source (e.g.…

计算机视觉与模式识别 · 计算机科学 2025-06-11 Vishaal Udandarao , Mehdi Cherti , Shyamgopal Karthik , Jenia Jitsev , Samuel Albanie , Matthias Bethge

This paper introduces SA-OOSC, a multimodal large language models (MLLM)-distilled semantic communication framework that achieves efficient semantic coding with scenario-aware importance allocations. This approach addresses a critical…

信号处理 · 电气工程与系统科学 2025-09-10 Feifan Zhang , Yuyang Du , Yifan Xiang , Xiaoyan Liu , Soung Chang Liew

Molecular property prediction is a fundamental task in computational chemistry with critical applications in drug discovery and materials science. While recent works have explored Large Language Models (LLMs) for this task, they primarily…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Deepan Adak , Yogesh Singh Rawat , Shruti Vyas

Scene Text Recognition (STR) remains challenging due to real-world complexities, where decoupled visual-linguistic optimization in existing frameworks amplifies error propagation through cross-modal misalignment. Visual encoders exhibit…

计算机视觉与模式识别 · 计算机科学 2025-11-12 Lixu Sun , Nurmemet Yolwas , Wushour Silamu

Weakly-Supervised Semantic Segmentation (WSSS) methods with image-level labels generally train a classification network to generate the Class Activation Maps (CAMs) as the initial coarse segmentation labels. However, current WSSS methods…

计算机视觉与模式识别 · 计算机科学 2022-02-11 Lixiang Ru , Bo Du , Yibing Zhan , Chen Wu

Multi-subject personalized generation presents unique challenges in maintaining identity fidelity and semantic coherence when synthesizing images conditioned on multiple reference subjects. Existing methods often suffer from identity…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Dong She , Siming Fu , Mushui Liu , Qiaoqiao Jin , Hualiang Wang , Mu Liu , Jidong Jiang

Object-centric representation learning offers the potential to overcome limitations of image-level representations by explicitly parsing image scenes into their constituent components. While image-level representations typically lack…

计算机视觉与模式识别 · 计算机科学 2023-08-30 Nathan Drenkow , Mathias Unberath

Multi-label image recognition with incomplete labels is a challenging yet vital task in computer vision, which faces two fundamental challenges: learning semantic-aware features and recovering missing labels. In this paper, we propose a…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Zhi-Fen He , Ren-Dong Xie , Bo Li , Bin Liu , Jin-Yan Hu

Optical character recognition (OCR) has advanced rapidly with the rise of vision-language models, yet evaluation has remained concentrated on a small cluster of high- and mid-resource scripts. We introduce GlotOCR Bench, a comprehensive…

计算与语言 · 计算机科学 2026-04-15 Amir Hossein Kargaran , Nafiseh Nikeghbal , Jana Diesner , François Yvon , Hinrich Schütze

Recent advances in neural image compression (NIC) have produced models that are starting to outperform classic codecs. While this has led to growing excitement about using NIC in real-world applications, the successful adoption of any…

图像与视频处理 · 电气工程与系统科学 2023-10-31 Kelsey Lieberman , James Diffenderfer , Charles Godfrey , Bhavya Kailkhura

Representation based classification (RC) methods such as sparse RC (SRC) have shown great potential in face recognition in recent years. Most previous RC methods are based on the conventional regression models, such as lasso regression,…

计算机视觉与模式识别 · 计算机科学 2017-11-17 Yulong Wang , Yuan Yan Tang , Luoqing Li , Hong Chen

Retrosynthesis prediction is fundamental to drug discovery and chemical synthesis, requiring the identification of reactants that can produce a target molecule. Current template-free methods struggle to capture the structural invariance…

机器学习 · 计算机科学 2025-10-21 Jiaxi Zhuang , Yu Zhang , Aimin Zhou , Ying Qian

Oracle bone script (OBS), as China's earliest mature writing system, present significant challenges in automatic recognition due to their complex pictographic structures and divergence from modern Chinese characters. We introduce…

计算机视觉与模式识别 · 计算机科学 2024-11-28 Hanqi Jiang , Yi Pan , Junhao Chen , Zhengliang Liu , Yifan Zhou , Peng Shu , Yiwei Li , Huaqin Zhao , Stephen Mihm , Lewis C Howe , Tianming Liu

Open-vocabulary 3D scene understanding (OV-3D) aims to localize and classify novel objects beyond the closed set of object classes. However, existing approaches and benchmarks primarily focus on the open vocabulary problem within the…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Youjun Zhao , Jiaying Lin , Shuquan Ye , Qianshi Pang , Rynson W. H. Lau

Recognizing materials in real-world images is a challenging task. Real-world materials have rich surface texture, geometry, lighting conditions, and clutter, which combine to make the problem particularly difficult. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2015-04-15 Sean Bell , Paul Upchurch , Noah Snavely , Kavita Bala

We present FireRed-OCR, a systematic framework to specialize general VLMs into high-performance OCR models. Large Vision-Language Models (VLMs) have demonstrated impressive general capabilities but frequently suffer from ``structural…

Recently, deep learning based methods have revolutionized remote sensing image segmentation. However, these methods usually rely on a pre-defined semantic class set, thus needing additional image annotation and model training when adapting…

计算机视觉与模式识别 · 计算机科学 2024-12-30 Chengyang Ye , Yunzhi Zhuge , Pingping Zhang

Recent advances in video generation have opened new avenues for macroscopic simulation of complex dynamic systems, but their application to microscopic phenomena remains largely unexplored. Microscale simulation holds great promise for…

人工智能 · 计算机科学 2026-03-03 Rongsheng Wang , Minghao Wu , Hongru Zhou , Zhihan Yu , Zhenyang Cai , Junying Chen , Benyou Wang

Large language models (LLMs) have demonstrated several emergent behaviors with scale, including reasoning and fluency in long-form text generation. However, they continue to struggle with tasks requiring precise spatial and positional…

Large Multimodal Models (LMMs) have recently shown strong performance on Optical Character Recognition (OCR) tasks, demonstrating their promising capability in document literacy. However, their effectiveness in real-world applications…