中文
相关论文

相关论文: WCXB: A Multi-Type Web Content Extraction Benchmar…

200 篇论文

Web applications (web apps) have become a key arena for large language models (LLMs) to demonstrate their code generation capabilities and commercial potential. However, building a benchmark for LLM-generated web apps remains challenging…

软件工程 · 计算机科学 2026-03-17 Chenxu Liu , Yingjie Fu , Wei Yang , Ying Zhang , Tao Xie

Web accessibility aims to ensure that web content and services are usable by people with diverse abilities. In recent years, Large Language Models (LLMs) have been increasingly explored to support accessibility-related tasks on the web,…

数字图书馆 · 计算机科学 2026-05-15 Wajdi Aljedaani , Rubel Hassan Mollik

Detecting problematic content, such as hate speech, is a multifaceted and ever-changing task, influenced by social dynamics, user populations, diversity of sources, and evolving language. There has been significant efforts, both in academia…

计算与语言 · 计算机科学 2023-10-09 Ali Omrani , Alireza S. Ziabari , Preni Golazizian , Jeffrey Sorensen , Morteza Dehghani

Multilingual falsehoods threaten information integrity worldwide, yet detection benchmarks remain confined to English or a few high-resource languages, leaving low-resource linguistic communities without robust defense tools. We introduce…

In this paper, we introduce a new NLP task -- generating short factual articles with references for queries by mining supporting evidence from the Web. In this task, called WebBrain, the ultimate goal is to generate a fluent, informative,…

计算与语言 · 计算机科学 2023-04-11 Hongjing Qian , Yutao Zhu , Zhicheng Dou , Haoqi Gu , Xinyu Zhang , Zheng Liu , Ruofei Lai , Zhao Cao , Jian-Yun Nie , Ji-Rong Wen

Online misinformation poses an escalating threat, amplified by the Internet's open nature and increasingly capable LLMs that generate persuasive yet deceptive content. Existing misinformation detection methods typically focus on either…

Structured information extraction from scientific literature is crucial for capturing core concepts and emerging trends in specialized fields. While existing datasets aid model development, most focus on specific publication sections due to…

计算与语言 · 计算机科学 2026-04-06 Decheng Duan , Yingyi Zhang , Jitong Peng , Chengzhi Zhang

As the information contained within the web is increasing day by day, organizing this information could be a necessary requirement.The data mining process is to extract information from a data set and transform it into an understandable…

信息检索 · 计算机科学 2014-05-22 Prabhjot Kaur

Reliably extracting tables from PDFs is essential for large-scale scientific data mining and knowledge base construction, yet existing evaluation approaches rely on rule-based metrics that fail to capture semantic equivalence of table…

计算机视觉与模式识别 · 计算机科学 2026-03-20 Pius Horn , Janis Keuper

Clickbait headlines degrade the quality of online information and undermine user trust. We present a hybrid approach to clickbait detection that combines transformer-based text embeddings with linguistically motivated informativeness…

计算与语言 · 计算机科学 2026-02-23 Wojciech Michaluk , Tymoteusz Urban , Mateusz Kubita , Soveatin Kuntur , Anna Wroblewska

Benchmark datasets have a significant impact on accelerating research in programming language tasks. In this paper, we introduce CodeXGLUE, a benchmark dataset to foster machine learning research for program understanding and generation.…

In this paper we introduce a new publicly available dataset for verification against textual sources, FEVER: Fact Extraction and VERification. It consists of 185,445 claims generated by altering sentences extracted from Wikipedia and…

计算与语言 · 计算机科学 2018-12-19 James Thorne , Andreas Vlachos , Christos Christodoulopoulos , Arpit Mittal

Most of the existing information extraction frameworks (Wadden et al., 2019; Veysehet al., 2020) focus on sentence-level tasks and are hardly able to capture the consolidated information from a given document. In our endeavour to generate…

计算与语言 · 计算机科学 2021-06-22 Debanjana Kar , Sudeshna Sarkar , Pawan Goyal

We present a dataset of 833k paragraphs extracted from CC-BY licensed scientific publications, classified into four categories: acknowledgments, data mentions, software/code mentions, and clinical trial mentions. The paragraphs are…

计算与语言 · 计算机科学 2025-10-28 Eric Jeangirard

Text summarization is crucial for mitigating information overload across domains like journalism, medicine, and business. This research evaluates summarization performance across 17 large language models (OpenAI, Google, Anthropic,…

计算与语言 · 计算机科学 2025-04-08 Anantharaman Janakiraman , Behnaz Ghoraani

Data from online job postings are difficult to access and are not built in a standard or transparent manner. Data included in the standard taxonomy and occupational information database (O*NET) are updated infrequently and based on small…

计算机与社会 · 计算机科学 2025-10-03 Stephen Meisenbacher , Svetlozar Nestorov , Peter Norlander

This paper focuses on a traditional relation extraction task in the context of limited annotated data and a narrow knowledge domain. We explore this task with a clinical corpus consisting of 200 breast cancer follow-up treatment letters in…

机器学习 · 计算机科学 2019-04-25 Jiyu Chen , Karin Verspoor , Zenan Zhai

We introduce a state-of-the-art approach for URL categorization that leverages the power of Large Language Models (LLMs) to address the primary objectives of web content filtering: safeguarding organizations from legal and ethical risks,…

机器学习 · 计算机科学 2023-05-11 Tamás Vörös , Sean Paul Bergeron , Konstantin Berlin

Content-dense news report important factual information about an event in direct, succinct manner. Information seeking applications such as information extraction, question answering and summarization normally assume all text they deal with…

计算与语言 · 计算机科学 2017-04-04 Yinfei Yang , Ani Nenkova

Extracting structured information from HTML documents is a long-studied problem with a broad range of applications, including knowledge base construction, faceted search, and personalized recommendation. Prior works rely on a few…

信息检索 · 计算机科学 2022-08-30 Ritesh Sarkhel , Binxuan Huang , Colin Lockard , Prashant Shiralkar