English
Related papers

Related papers: Language-Agnostic Modeling of Source Reliability o…

200 papers

INTRODUCTION: Wikipedia is a major source of information, particularly for medical and health content, citing over 4 million scholarly publications. However, the representation of research-based knowledge across different languages on…

Digital Libraries · Computer Science 2025-01-17 Michael Taylor , Roisi Proven , Carlos Areia

The proliferation of fake news has had far-reaching implications on politics, the economy, and society at large. While Fake news detection methods have been employed to mitigate this issue, they primarily depend on two essential elements:…

Computation and Language · Computer Science 2024-03-18 Guanghua Li , Wensheng Lu , Wei Zhang , Defu Lian , Kezhong Lu , Rui Mao , Kai Shu , Hao Liao

Dictionaries are often developed using tools that save to Extensible Markup Language (XML)-based standards. These standards often allow high-level repeating elements to represent lexical entries, and utilize descendants of these repeating…

Computation and Language · Computer Science 2016-02-18 Paul Rodrigues , David Zajic , David Doermann , Michael Bloodgood , Peng Ye

One of the most impressive human endeavors of the past two decades is the collection and categorization of human knowledge in the free and accessible format that is Wikipedia. In this work we ask what makes a term worthy of entering this…

Computation and Language · Computer Science 2020-09-18 Yonatan Bilu , Shai Gretz , Edo Cohen , Noam Slonim

LLMs are known to store vast amounts of knowledge in their parametric memory. However, learning and recalling facts from this memory is known to be unreliable, depending largely on the prevalence of particular facts in the training data and…

Computation and Language · Computer Science 2025-08-14 Jessy Lin , Vincent-Pierre Berges , Xilun Chen , Wen-Tau Yih , Gargi Ghosh , Barlas Oğuz

Verifiable generation requires large language models (LLMs) to cite source documents supporting their outputs, thereby improve output transparency and trustworthiness. Yet, previous work mainly targets the generation of sentence-level…

Computation and Language · Computer Science 2024-06-11 Shuyang Cao , Lu Wang

Despite growing efforts to halt distasteful content on social media, multilingualism has added a new dimension to this problem. The scarcity of resources makes the challenge even greater when it comes to low-resource languages. This work…

Social and Information Networks · Computer Science 2024-10-30 Mohammad Zia Ur Rehman , Somya Mehta , Kuldeep Singh , Kunal Kaushik , Nagendra Kumar

Automated disinformation generation is often listed as an important risk associated with large language models (LLMs). The theoretical ability to flood the information space with disinformation content might have dramatic consequences for…

Computation and Language · Computer Science 2025-06-13 Ivan Vykopal , Matúš Pikuliak , Ivan Srba , Robert Moro , Dominik Macko , Maria Bielikova

The reproducibility of scientific articles is central to the advancement of science. Despite this importance, evaluating reproducibility remains challenging due to the scarcity of ground truth data. Predictive models can address this…

Digital Libraries · Computer Science 2024-10-25 Akhil Pandey Akella , Sagnik Ray Choudhury , David Koop , Hamed Alhoori

Discourse analysis allows us to attain inferences of a text document that extend beyond the sentence-level. The current performance of discourse models is very low on texts outside of the training distribution's coverage, diminishing the…

Computation and Language · Computer Science 2022-03-23 Katherine Atwell , Anthony Sicilia , Seong Jae Hwang , Malihe Alikhani

Although Wikipedia is the largest multilingual encyclopedia, it remains inherently incomplete. There is a significant disparity in the quality of content between high-resource languages (HRLs, e.g., English) and low-resource languages…

Computation and Language · Computer Science 2024-12-10 Paramita Das , Amartya Roy , Ritabrata Chakraborty , Animesh Mukherjee

The frequency of a web search keyword generally reflects the degree of public interest in a particular subject matter. Search logs are therefore useful resources for trend analysis. However, access to search logs is typically restricted to…

Social and Information Networks · Computer Science 2015-09-09 Mitsuo Yoshida , Yuki Arase , Takaaki Tsunoda , Mikio Yamamoto

We study text reuse related to Wikipedia at scale by compiling the first corpus of text reuse cases within Wikipedia as well as without (i.e., reuse of Wikipedia text in a sample of the Common Crawl). To discover reuse beyond verbatim copy…

Information Retrieval · Computer Science 2018-12-24 Milad Alshomary , Michael Völske , Tristan Licht , Henning Wachsmuth , Benno Stein , Matthias Hagen , Martin Potthast

Wikipedia serves as a good example of how editors collaborate to form and maintain an article. The relationship between editors, derived from their sequence of editing activity, results in a directed network structure called the revision…

Social and Information Networks · Computer Science 2019-04-18 James R. Ashford , Liam D. Turner , Roger M. Whitaker , Alun Preece , Diane Felmlee , Don Towsley

A thermodynamic framework is presented to characterize the evolution of efficiency, order, and quality in social content production systems, and this framework is applied to the analysis of Wikipedia. Contributing editors are characterized…

Social and Information Networks · Computer Science 2012-04-18 Huan-Kai Peng , Ying Zhang , Peter Pirolli , Tad Hogg

Counterfactual explanations provide human-understandable reasoning for AI-made decisions by describing minimal changes to input features that would alter a model's prediction. To be truly useful in practice, such explanations must be…

Machine Learning · Computer Science 2025-08-15 Asiful Arefeen , Shovito Barua Soumma , Hassan Ghasemzadeh

As Web sites are now ordinary products, it is necessary to explicit the notion of quality of a Web site. The quality of a site may be linked to the easiness of accessibility and also to other criteria such as the fact that the site is up to…

Information Retrieval · Computer Science 2007-05-23 Thierry Despeyroux

Media seems to have become more partisan, often providing a biased coverage of news catering to the interest of specific groups. It is therefore essential to identify credible information content that provides an objective narrative of an…

Artificial Intelligence · Computer Science 2017-05-16 Subhabrata Mukherjee , Gerhard Weikum

Knowledge and language understanding of models evaluated through question answering (QA) has been usually studied on static snapshots of knowledge, like Wikipedia. However, our world is dynamic, evolves over time, and our models' knowledge…

Search engines increasingly leverage large language models (LLMs) to generate direct answers, and AI chatbots now access the Internet for fresh data. As information curators for billions of users, LLMs must assess the accuracy and…

Computation and Language · Computer Science 2025-05-23 Kai-Cheng Yang , Filippo Menczer
‹ Prev 1 8 9 10 Next ›