中文
相关论文

相关论文: Is a Document Educational or Just Wikipedia-Style?…

200 篇论文

Large-scale models are pretrained on massive web-crawled datasets containing documents of mixed quality, making data filtering essential. A popular method is Classifier-based Quality Filtering (CQF), which trains a binary classifier to…

机器学习 · 计算机科学 2025-10-03 Thiziri Nait Saada , Louis Bethune , Michal Klein , David Grangier , Marco Cuturi , Pierre Ablin

We propose an edit-centric approach to assess Wikipedia article quality as a complementary alternative to current full document-based techniques. Our model consists of a main classifier equipped with an auxiliary generative module which,…

计算与语言 · 计算机科学 2019-09-20 Edison Marrese-Taylor , Pablo Loyola , Yutaka Matsuo

Wikipedia articles aim to be definitive sources of encyclopedic content. Yet, only 0.6% of Wikipedia articles have high quality according to its quality scale due to insufficient number of Wikipedia editors and enormous number of articles.…

社会与信息网络 · 计算机科学 2021-08-06 Sumit Asthana , Sabrina Tobar Thommel , Aaron Lee Halfaker , Nikola Banovic

Language models increasingly rely on massive web dumps for diverse text data. However, these sources are rife with undesirable content. As such, resources like Wikipedia, books, and newswire often serve as anchors for automatically…

Millions of people irrespective of socioeconomic and demographic backgrounds, depend on Wikipedia articles everyday for keeping themselves informed regarding popular as well as obscure topics. Articles have been categorized by editors into…

社会与信息网络 · 计算机科学 2020-10-15 Bhanu Prakash Reddy , Sasi Bhusan , Soumya Sarkar , Animesh Mukherjee

Wikipedia is the largest web repository of free knowledge. Volunteer editors devote time and effort to creating and expanding articles in more than 300 language editions. As content quality varies from article to article, editors also spend…

计算机与社会 · 计算机科学 2024-04-16 Paramita Das , Isaac Johnson , Diego Saez-Trumper , Pablo Aragón

Wikipedia has been turned into an immensely popular crowd-sourced encyclopedia for information dissemination on numerous versatile topics in the form of subscription free content. It allows anyone to contribute so that the articles remain…

社会与信息网络 · 计算机科学 2021-11-03 Paramita Das , Bhanu Prakash Reddy Guda , Sasi Bhusan Seelaboyina , Soumya Sarkar , Animesh Mukherjee

Wikipedia's perceived high quality and broad language coverage have established it as a fundamental resource in NLP. However, in recent years, such assumptions of high quality have become the subject of scrutiny in low-resource and…

With the development of deep learning and natural language processing techniques, pre-trained language models have been widely used to solve information retrieval (IR) problems. Benefiting from the pre-training and fine-tuning paradigm,…

信息检索 · 计算机科学 2024-01-02 Weihang Su , Qingyao Ai , Xiangsheng Li , Jia Chen , Yiqun Liu , Xiaolong Wu , Shengluan Hou

The use of domain knowledge is generally found to improve query efficiency in content filtering applications. In particular, tangible benefits have been achieved when using knowledge-based approaches within more specialized fields, such as…

信息检索 · 计算机科学 2015-03-17 Pekka Malo , Pyry Siitari , Oskar Ahlgren , Jyrki Wallenius , Pekka Korhonen

The quality of a document is affected by various factors, including grammaticality, readability, stylistics, and expertise depth, making the task of document quality assessment a complex one. In this paper, we explore this task in the…

计算与语言 · 计算机科学 2019-01-15 Aili Shen , Bahar Salehi , Timothy Baldwin , Jianzhong Qi

Encyclopedic queries express the intent of obtaining information typically available in encyclopedias, such as biographical, geographical or historical facts. In this paper, we train a classifier for detecting the encyclopedic intent of web…

信息检索 · 计算机科学 2015-12-01 Pedro Saleiro , Luís Sarmento

Nowadays, thanks to Web 2.0 technologies, people have the possibility to generate and spread contents on different social media in a very easy way. In this context, the evaluation of the quality of the information that is available online…

计算与语言 · 计算机科学 2018-12-10 Elias Bassani , Marco Viviani

Large Language Model (LLM) pre-training exhausts an ever growing compute budget, yet recent research has demonstrated that careful document selection enables comparable model quality with only a fraction of the FLOPs. Inspired by efforts…

计算与语言 · 计算机科学 2024-06-10 Xiang Kong , Tom Gunter , Ruoming Pang

Well curated, large-scale corpora of social media posts containing broad public opinion offer an alternative data source to complement traditional surveys. While surveys are effective at collecting representative samples and are capable of…

计算与语言 · 计算机科学 2025-02-14 Michael V. Arnold , Peter Sheridan Dodds , Christopher M. Danforth

Identifying critical research within the growing body of academic work is an intrinsic aspect of conducting quality research. Systematic review processes used in evidence-based medicine formalise this as a procedure that must be followed in…

数字图书馆 · 计算机科学 2024-10-14 John Hawkins , David Tivey

Recent work demonstrates that filtering harmful content from pretraining data improves model safety without degrading capabilities. We propose a natural extension: do it again. A model trained on filtered data can filter the corpus further;…

人工智能 · 计算机科学 2026-02-04 Robin Young

We introduce a state-of-the-art approach for URL categorization that leverages the power of Large Language Models (LLMs) to address the primary objectives of web content filtering: safeguarding organizations from legal and ethical risks,…

机器学习 · 计算机科学 2023-05-11 Tamás Vörös , Sean Paul Bergeron , Konstantin Berlin

Auditing language-model outputs often requires more than judging correctness: an auditor may need to identify which source document most likely supports the knowledge expressed in a response. We study this as pinpoint provenance: given a…

人工智能 · 计算机科学 2026-05-08 Xiaomin Li , Andrzej Banburski-Fahey , Jaron Lanier

Fine-tuning-based unlearning methods prevail for preventing targeted harmful, sensitive, or copyrighted information within large language models while preserving overall capabilities. However, the true effectiveness of these methods is…

计算与语言 · 计算机科学 2024-10-16 Yihuai Hong , Yuelin Zou , Lijie Hu , Ziqian Zeng , Di Wang , Haiqin Yang
‹ 上一页 1 2 3 10 下一页 ›