中文
相关论文

相关论文: FUNSD: A Dataset for Form Understanding in Noisy S…

200 篇论文

Document understanding in real-world applications often requires processing heterogeneous, multi-page document packets containing multiple documents stitched together. Despite recent advances in visual document understanding, the…

Objects that undergo non-rigid deformation are common in the real world. A typical and challenging example is the human faces. While various techniques have been developed for deformable shape registration and classification, benchmarks…

计算机视觉与模式识别 · 计算机科学 2018-07-11 Gareth Andrews , Sam Endean , Roberto Dyke , Yukun Lai , Gwenno Ffrancon , Gary KL Tam

Understanding and extracting of information from large documents, such as business opportunities, academic articles, medical documents and technical reports, poses challenges not present in short documents. Such large documents may be…

计算与语言 · 计算机科学 2019-10-10 Muhammad Mahbubur Rahman , Tim Finin

This paper considers recognizing products from daily photos, which is an important problem in real-world applications but also challenging due to background clutters, category diversities, noisy labels, etc. We address this problem by two…

计算机视觉与模式识别 · 计算机科学 2019-07-29 Qing Li , Xiaojiang Peng , Liangliang Cao , Wenbin Du , Hao Xing , Yu Qiao

We present the Stanford Question Answering Dataset (SQuAD), a new reading comprehension dataset consisting of 100,000+ questions posed by crowdworkers on a set of Wikipedia articles, where the answer to each question is a segment of text…

计算与语言 · 计算机科学 2016-10-12 Pranav Rajpurkar , Jian Zhang , Konstantin Lopyrev , Percy Liang

Purpose: Automated ultrasound image analysis is challenging due to anatomical complexity and limited annotated data. To tackle this, we take a data-centric approach, assembling the largest public ultrasound segmentation dataset and training…

图像与视频处理 · 电气工程与系统科学 2025-11-13 Adrien Meyer , Aditya Murali , Farahdiba Zarin , Didier Mutter , Nicolas Padoy

Ultrasound imaging reveals eye morphology and aids in diagnosing and treating eye diseases. However, interpreting diagnostic reports requires specialized physicians. We present a labeled ophthalmic dataset for the precise analysis and the…

计算机视觉与模式识别 · 计算机科学 2024-07-29 Jing Wang , Junyan Fan , Meng Zhou , Yanzhu Zhang , Mingyu Shi

Sonar images are relevant for advancing underwater exploration, autonomous navigation, and ecosystem monitoring. However, the progress depends on data availability. The scarcity of publicly available, well-annotated sonar image datasets…

We present DocFormer -- a multi-modal transformer based architecture for the task of Visual Document Understanding (VDU). VDU is a challenging problem which aims to understand documents in their varied formats (forms, receipts etc.) and…

计算机视觉与模式识别 · 计算机科学 2021-09-21 Srikar Appalaraju , Bhavan Jasani , Bhargava Urala Kota , Yusheng Xie , R. Manmatha

Named entity recognition (NER) models often struggle with noisy inputs, such as those with spelling mistakes or errors generated by Optical Character Recognition processes, and learning a robust NER model is challenging. Existing robust NER…

计算与语言 · 计算机科学 2024-07-29 Chaoyi Ai , Yong Jiang , Shen Huang , Pengjun Xie , Kewei Tu

We present MOFI, Manifold OF Images, a new vision foundation model designed to learn image representations from noisy entity annotated images. MOFI differs from previous work in two key aspects: (i) pre-training data, and (ii) training…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Wentao Wu , Aleksei Timofeev , Chen Chen , Bowen Zhang , Kun Duan , Shuangning Liu , Yantao Zheng , Jonathon Shlens , Xianzhi Du , Zhe Gan , Yinfei Yang

Given the central role of charts in scientific, business, and communication contexts, enhancing the chart understanding capabilities of vision-language models (VLMs) has become increasingly critical. A key limitation of existing VLMs lies…

计算机视觉与模式识别 · 计算机科学 2025-10-17 Jiangning Zhu , Yuxing Zhou , Zheng Wang , Juntao Yao , Yima Gu , Yuhui Yuan , Shixia Liu

Robustness to label noise within data is a significant challenge in federated learning (FL). From the data-centric perspective, the data quality of distributed datasets can not be guaranteed since annotations of different clients contain…

计算机视觉与模式识别 · 计算机科学 2025-05-13 Xuefeng Jiang , Jia Li , Nannan Wu , Zhiyuan Wu , Xujing Li , Sheng Sun , Gang Xu , Yuwei Wang , Qi Li , Min Liu

Automatic Emotion Detection (ED) aims to build systems to identify users' emotions automatically. This field has the potential to enhance HCI, creating an individualised experience for the user. However, ED systems tend to perform poorly on…

人机交互 · 计算机科学 2023-07-27 Annanda Sousa , Karen Young , Mathieu D'aquin , Manel Zarrouk , Jennifer Holloway

Large-scale public datasets with high-quality annotations are rarely available for intelligent medical imaging research, due to data privacy concerns and the cost of annotations. In this paper, we release SynFundus-1M, a high-quality…

计算机视觉与模式识别 · 计算机科学 2024-03-15 Fangxin Shang , Jie Fu , Yehui Yang , Haifeng Huang , Junwei Liu , Lei Ma

Large sense-annotated datasets are increasingly necessary for training deep supervised systems in Word Sense Disambiguation. However, gathering high-quality sense-annotated data for as many instances as possible is a laborious and expensive…

计算与语言 · 计算机科学 2020-03-16 Tommaso Pasini , Jose Camacho-Collados

This paper presents a new synthetic dataset of ID and travel documents, called SIDTD. The SIDTD dataset is created to help training and evaluating forged ID documents detection systems. Such a dataset has become a necessity as ID documents…

计算机视觉与模式识别 · 计算机科学 2024-01-04 Carlos Boned , Maxime Talarmain , Nabil Ghanmi , Guillaume Chiron , Sanket Biswas , Ahmad Montaser Awal , Oriol Ramos Terrades

This work presents a unified knowledge protocol, called UKnow, which facilitates knowledge-based studies from the perspective of data. Particularly focusing on visual and linguistic modalities, we categorize data knowledge into five unit…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Biao Gong , Shuai Tan , Yutong Feng , Xiaoying Xie , Yuyuan Li , Chaochao Chen , Kecheng Zheng , Yujun Shen , Deli Zhao

Current vision systems are trained on huge datasets, and these datasets come with costs: curation is expensive, they inherit human biases, and there are concerns over privacy and usage rights. To counter these costs, interest has surged in…

计算机视觉与模式识别 · 计算机科学 2022-05-02 Manel Baradad , Jonas Wulff , Tongzhou Wang , Phillip Isola , Antonio Torralba

Document Layout Analysis, which is the task of identifying different semantic regions inside of a document page, is a subject of great interest for both computer scientists and humanities scholars as it represents a fundamental step towards…

计算机视觉与模式识别 · 计算机科学 2024-01-17 Silvia Zottin , Axel De Nardin , Emanuela Colombi , Claudio Piciarelli , Filippo Pavan , Gian Luca Foresti