English
Related papers

Related papers: WCXB: A Multi-Type Web Content Extraction Benchmar…

200 papers

Web applications (web apps) have become a key arena for large language models (LLMs) to demonstrate their code generation capabilities and commercial potential. However, building a benchmark for LLM-generated web apps remains challenging…

Software Engineering · Computer Science 2026-03-17 Chenxu Liu , Yingjie Fu , Wei Yang , Ying Zhang , Tao Xie

Web accessibility aims to ensure that web content and services are usable by people with diverse abilities. In recent years, Large Language Models (LLMs) have been increasingly explored to support accessibility-related tasks on the web,…

Digital Libraries · Computer Science 2026-05-15 Wajdi Aljedaani , Rubel Hassan Mollik

Detecting problematic content, such as hate speech, is a multifaceted and ever-changing task, influenced by social dynamics, user populations, diversity of sources, and evolving language. There has been significant efforts, both in academia…

Computation and Language · Computer Science 2023-10-09 Ali Omrani , Alireza S. Ziabari , Preni Golazizian , Jeffrey Sorensen , Morteza Dehghani

Multilingual falsehoods threaten information integrity worldwide, yet detection benchmarks remain confined to English or a few high-resource languages, leaving low-resource linguistic communities without robust defense tools. We introduce…

Computation and Language · Computer Science 2026-03-03 Jason Lucas , Matt Murtagh-White , Adaku Uchendu , Ali Al-Lawati , Michiharu Yamashita , Dominik Macko , Ivan Srba , Robert Moro , Dongwon Lee

In this paper, we introduce a new NLP task -- generating short factual articles with references for queries by mining supporting evidence from the Web. In this task, called WebBrain, the ultimate goal is to generate a fluent, informative,…

Computation and Language · Computer Science 2023-04-11 Hongjing Qian , Yutao Zhu , Zhicheng Dou , Haoqi Gu , Xinyu Zhang , Zheng Liu , Ruofei Lai , Zhao Cao , Jian-Yun Nie , Ji-Rong Wen

Online misinformation poses an escalating threat, amplified by the Internet's open nature and increasingly capable LLMs that generate persuasive yet deceptive content. Existing misinformation detection methods typically focus on either…

Structured information extraction from scientific literature is crucial for capturing core concepts and emerging trends in specialized fields. While existing datasets aid model development, most focus on specific publication sections due to…

Computation and Language · Computer Science 2026-04-06 Decheng Duan , Yingyi Zhang , Jitong Peng , Chengzhi Zhang

As the information contained within the web is increasing day by day, organizing this information could be a necessary requirement.The data mining process is to extract information from a data set and transform it into an understandable…

Information Retrieval · Computer Science 2014-05-22 Prabhjot Kaur

Reliably extracting tables from PDFs is essential for large-scale scientific data mining and knowledge base construction, yet existing evaluation approaches rely on rule-based metrics that fail to capture semantic equivalence of table…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Pius Horn , Janis Keuper

Clickbait headlines degrade the quality of online information and undermine user trust. We present a hybrid approach to clickbait detection that combines transformer-based text embeddings with linguistically motivated informativeness…

Computation and Language · Computer Science 2026-02-23 Wojciech Michaluk , Tymoteusz Urban , Mateusz Kubita , Soveatin Kuntur , Anna Wroblewska

Benchmark datasets have a significant impact on accelerating research in programming language tasks. In this paper, we introduce CodeXGLUE, a benchmark dataset to foster machine learning research for program understanding and generation.…

In this paper we introduce a new publicly available dataset for verification against textual sources, FEVER: Fact Extraction and VERification. It consists of 185,445 claims generated by altering sentences extracted from Wikipedia and…

Computation and Language · Computer Science 2018-12-19 James Thorne , Andreas Vlachos , Christos Christodoulopoulos , Arpit Mittal

Most of the existing information extraction frameworks (Wadden et al., 2019; Veysehet al., 2020) focus on sentence-level tasks and are hardly able to capture the consolidated information from a given document. In our endeavour to generate…

Computation and Language · Computer Science 2021-06-22 Debanjana Kar , Sudeshna Sarkar , Pawan Goyal

We present a dataset of 833k paragraphs extracted from CC-BY licensed scientific publications, classified into four categories: acknowledgments, data mentions, software/code mentions, and clinical trial mentions. The paragraphs are…

Computation and Language · Computer Science 2025-10-28 Eric Jeangirard

Text summarization is crucial for mitigating information overload across domains like journalism, medicine, and business. This research evaluates summarization performance across 17 large language models (OpenAI, Google, Anthropic,…

Computation and Language · Computer Science 2025-04-08 Anantharaman Janakiraman , Behnaz Ghoraani

Data from online job postings are difficult to access and are not built in a standard or transparent manner. Data included in the standard taxonomy and occupational information database (O*NET) are updated infrequently and based on small…

Computers and Society · Computer Science 2025-10-03 Stephen Meisenbacher , Svetlozar Nestorov , Peter Norlander

This paper focuses on a traditional relation extraction task in the context of limited annotated data and a narrow knowledge domain. We explore this task with a clinical corpus consisting of 200 breast cancer follow-up treatment letters in…

Machine Learning · Computer Science 2019-04-25 Jiyu Chen , Karin Verspoor , Zenan Zhai

We introduce a state-of-the-art approach for URL categorization that leverages the power of Large Language Models (LLMs) to address the primary objectives of web content filtering: safeguarding organizations from legal and ethical risks,…

Machine Learning · Computer Science 2023-05-11 Tamás Vörös , Sean Paul Bergeron , Konstantin Berlin

Content-dense news report important factual information about an event in direct, succinct manner. Information seeking applications such as information extraction, question answering and summarization normally assume all text they deal with…

Computation and Language · Computer Science 2017-04-04 Yinfei Yang , Ani Nenkova

Extracting structured information from HTML documents is a long-studied problem with a broad range of applications, including knowledge base construction, faceted search, and personalized recommendation. Prior works rely on a few…

Information Retrieval · Computer Science 2022-08-30 Ritesh Sarkhel , Binxuan Huang , Colin Lockard , Prashant Shiralkar