中文
相关论文

相关论文: ACL-Fig: A Dataset for Scientific Figure Classific…

200 篇论文

With the rapid growth of the Natural Language Processing (NLP) field, a vast variety of Large Language Models (LLMs) continue to emerge for diverse NLP tasks. As more papers are published, researchers and developers face the challenge of…

计算与语言 · 计算机科学 2024-11-26 Shengwei Tian , Lifeng Han , Goran Nenadic

Scientific writing involves retrieving, summarizing, and citing relevant papers, which can be time-consuming processes in large and rapidly evolving fields. By making these processes inter-operable, natural language processing (NLP)…

计算与语言 · 计算机科学 2023-11-07 Nianlong Gu , Richard H. R. Hahnloser

This paper presents a hierarchical classification system that automatically categorizes a scholarly publication using its abstract into a three-tier hierarchical label set (discipline, field, subfield) in a multi-class setting. This system…

数字图书馆 · 计算机科学 2024-07-26 Susie Xi Rao , Peter H. Egger , Ce Zhang

The growing volume of digital images necessitates advanced systems for efficient categorization and retrieval, presenting a significant challenge in database management and information retrieval. This paper introduces PICS (Pipeline for…

计算机视觉与模式识别 · 计算机科学 2024-02-16 Grant Rosario , David Noever

Memes have emerged as a powerful form of communication, integrating visual and textual elements to convey humor, satire, and cultural messages. Existing research has focused primarily on aspects such as emotion classification, meme…

机器学习 · 计算机科学 2025-01-24 Shiling Deng , Serge Belongie , Peter Ebert Christensen

Scientific figure captioning is a complex task that requires generating contextually appropriate descriptions of visual content. However, existing methods often fall short by utilizing incomplete information, treating the task solely as…

In the digital age, seeking health advice on the Internet has become a common practice. At the same time, determining the trustworthiness of online medical content is increasingly challenging. Fact-checking has emerged as an approach to…

计算与语言 · 计算机科学 2024-03-26 Juraj Vladika , Phillip Schneider , Florian Matthes

We formalize and enable the task of open tree decomposition, which segments an image into hierarchical trees of visual components with unconstrained granularity and flexibility. Specifically, we provide the foundation benchmark for this new…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Junhyub Lee , Seunghun Chae , Hyosu Kim

The last decade has seen widespread adoption of Machine Learning (ML) components in software systems. This has occurred in nearly every domain, from natural language processing to computer vision. These ML components range from relatively…

Artificial Intelligence for scientific applications increasingly requires training large models on data that cannot be centralized due to privacy constraints, data sovereignty, or the sheer volume of data generated. Federated learning (FL)…

机器学习 · 计算机科学 2026-03-23 Yijiang Li , Zilinghan Li , Kyle Chard , Ian Foster , Todd Munson , Ravi Madduri , Kibaek Kim

Existing Multi-Turn Composed Image Retrieval (MTCIR) datasets lack dialogue-history consistency and are restricted to the fashion domain. To address these limitations, we construct CIRCLED by extending FashionIQ, CIRR, and CIRCO. In…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Tomohisa Takeda , Yu-Chieh Lin , Yuji Nozawa , Youyang Ng , Osamu Torii , Yusuke Matsui

We contribute the largest publicly available dataset of naturally occurring factual claims for the purpose of automatic claim verification. It is collected from 26 fact checking websites in English, paired with textual sources and rich…

There is growing interest in systems that generate captions for scientific figures. However, assessing these systems output poses a significant challenge. Human evaluation requires academic expertise and is costly, while automatic…

计算与语言 · 计算机科学 2023-10-25 Ting-Yao Hsu , Chieh-Yang Huang , Ryan Rossi , Sungchul Kim , C. Lee Giles , Ting-Hao K. Huang

Scholarly data are largely fragmented across siloed databases with divergent metadata and missing linkages among them. We present the Science Data Lake, a locally-deployable infrastructure built on DuckDB and simple Parquet files that…

数字图书馆 · 计算机科学 2026-03-04 Jonas Wilinski

Compared to natural images, understanding scientific figures is particularly hard for machines. However, there is a valuable source of information in scientific literature that until now has remained untapped: the correspondence between a…

人工智能 · 计算机科学 2019-09-20 Jose Manuel Gomez-Perez , Raul Ortega

Nailfold capillaroscopy is widely used in assessing health conditions, highlighting the pressing need for an automated nailfold capillary analysis system. In this study, we present a pioneering effort in constructing a comprehensive…

图像与视频处理 · 电气工程与系统科学 2024-03-15 Linxi Zhao , Jiankai Tang , Dongyu Chen , Xiaohong Liu , Yong Zhou , Yuanchun Shi , Guangyu Wang , Yuntao Wang

The rapid evolution of scientific fields introduces challenges in organizing and retrieving scientific literature. While expert-curated taxonomies have traditionally addressed this need, the process is time-consuming and expensive.…

计算与语言 · 计算机科学 2025-06-13 Priyanka Kargupta , Nan Zhang , Yunyi Zhang , Rui Zhang , Prasenjit Mitra , Jiawei Han

Publication databases rely on accurate metadata extraction from diverse web sources, yet variations in web layouts and data formats present challenges for metadata providers. This paper introduces CRAWLDoc, a new method for contextual…

计算与语言 · 计算机科学 2025-06-05 Fabian Karl , Ansgar Scherp

We introduce SciQAG, a novel framework for automatically generating high-quality science question-answer pairs from a large corpus of scientific literature based on large language models (LLMs). SciQAG consists of a QA generator and a QA…

计算与语言 · 计算机科学 2024-07-11 Yuwei Wan , Yixuan Liu , Aswathy Ajith , Clara Grazian , Bram Hoex , Wenjie Zhang , Chunyu Kit , Tong Xie , Ian Foster

Vision and vision-language applications of neural networks, such as image classification and captioning, rely on large-scale annotated datasets that require non-trivial data-collecting processes. This time-consuming endeavor hinders the…

‹ 上一页 1 8 9 10 下一页 ›